John Hughes

Hi, I'm John 👋

🛡️ AI Control @ Anthropic 🤖 Vibe Coding 🥽 FPV Drones 🤘 Live Music 🥁 Drums 🎾 Tennis 🧖 Spa

I work at Anthropic on the AI control team, where my focus is making increasingly autonomous AI agents safe. Most recently I built the safeguards behind Claude Code's auto mode. I got my start in AI safety through the MATS Program in Summer 2023, supervised by Ethan Perez, working on scalable oversight and adversarial robustness. Our paper on debate came out of that and won best paper at ICML 2024.

I also helped run the technical onboarding for the Anthropic Fellows programme. My notes on empirical research workflows and presenting results came out of that, though LLM progress has already dated a good chunk of the tooling advice.

Before AI safety I was a machine learning engineer and manager at Speechmatics. In another life I was a cox, which mostly meant shouting at rowers a bunch, and you can watch that here. The rest of this site is my work and hobbies (and some AI generated art!). Thanks for visiting!

Featured Work

Anthropic glyph artwork from the auto mode post

How We Built Claude Code Auto Mode: A Safer Way to Skip Permissions

March 25th 2026 Anthropic Engineering

My main project on Anthropic's AI control team. Auto mode lets Claude Code work without a permission prompt on every action, behind a two layer defence. A prompt injection probe screens tool outputs, and a transcript classifier judges each action against safety criteria before it runs. It catches overeager agents and honest mistakes at a 0.4% false positive rate on real traffic.

Dead-leaf mimic butterfly, artwork from the paper

Why Do Some Language Models Fake Alignment While Others Don't?

June 22nd 2025 NeurIPS 2025 · Spotlight

Abhay Sheshadri*, John Hughes*, Julian Michael, Alex Mallen, Arun Jose, Janus, Fabien Roger

Claude 3 Opus selectively complies with a helpful-only training objective to avoid having its behaviour modified. We extend this analysis to 25 frontier chat models and find that only five show a compliance gap between training and deployment, with only Claude 3 Opus's gap consistently motivated by preserving its goals. We also investigate why most models don't fake alignment. It is not simply a lack of capability, since many base models fake alignment some of the time, and post-training can either eliminate or amplify it.

Best-of-N Jailbreaking artwork

Best-of-N Jailbreaking

December 4th 2024 NeurIPS 2025

John Hughes*, Sara Price*, Aengus Lynch*, Rylan Schaeffer, Fazl Barez, Sanmi Koyejo, Henry Sleight, Erik Jones, Ethan Perez†, Mrinank Sharma†

We introduce Best-of-N (BoN) Jailbreaking, a straightforward algorithm that effectively jailbreaks AI systems across modalities by sampling prompt variations with simple augmentations. Achieving high attack success rates on models like GPT-4o and Claude 3.5 Sonnet, BoN also bypasses advanced defenses and extends to other modalities (vision and audio). Its effectiveness increases with more samples and the scaling behaviour follows a power-law, highlighting significant vulnerabilities of AI to subtle input changes.

LLM debate artwork

Debating with More Persuasive LLMs Leads to More Truthful Answers

February 9th 2024 ICML 2024 Oral 🏆 Best Paper Award

Akbir Khan*, John Hughes*, Dan Valentine*, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel†, Ethan Perez†

We investigated the potential of weaker language models (non-experts) to assess the correctness of stronger models (experts) via LLM debate, demonstrating a significant improvement in accuracy for both non-expert models and humans in the QuALITY comprehension task. Pioneered the optimisation of expert debaters for persuasiveness in an unsupervised manner, leading to enhanced non-expert capabilities in identifying accurate answers during debates.

Other Work

Measuring the Ability of Opus 4.5 to Fool Narrow Classifiers

May 2nd 2026 LessWrong

Fabien Roger, John Hughes

How well a frontier model can fool the monitors watching it. Current monitors are not trivial to jailbreak, though specialised edge cases do get through.

Alignment Faking Revisited: Improved Classifiers and Open Source Extensions

April 8th 2025 Anthropic Alignment Science LessWrong

John Hughes, Abhay Sheshadri

Much stronger alignment faking classifiers, with an AUROC of 0.9 against the original 0.6, plus open sourced classifiers, datasets and fine-tuned models.

Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach

December 3rd 2024 NeurIPS 2024 AdvML Frontiers · Oral SoLaR

Tony T. Wang*, John Hughes*, Henry Sleight, Rylan Schaeffer, Rajashree Agrawal, Fazl Barez, Mrinank Sharma, Jesse Mu, Nir Shavit, Ethan Perez†

Prohibiting a single behaviour turns out to be hard. Our transcript classifier defense beats safety training and input/output classifiers, but it still occasionally fails.

Supporting Contributions

Agentic Misalignment in Summer 2026

July 13th 2026 Anthropic Alignment Science

Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded Outputs

January 20th 2026 ICLR 2026

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents

June 17th 2025 Anthropic Research

How Do Large Language Monkeys Get Their Power (Laws)?

February 24th 2025 ICML 2025

Looking Inward: Language Models Can Learn About Themselves by Introspection

October 17th 2024 ICLR 2025

Failures to Find Transferable Image Jailbreaks Between Vision-Language Models

July 21st 2024 ICLR 2025 NeurIPS 2024 Workshops · Best Paper

Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data

April 1st 2024 COLM 2024

Hierarchical Quantised Autoencoders

February 19th 2020 NeurIPS 2020

Earlier Work in Speech Recognition

Flow, an API for Building Voice Interactions

July 2024 Speechmatics

Flow pairs real-time speech recognition with an LLM and text to speech, so you can build voice interactions into pretty much anything. I ran the projects hooking LLMs up to our audio systems that fed into it.

Ursa: The World's Most Accurate Speech-to-Text

March 8th 2023 Speechmatics

Ursa beat Microsoft and OpenAI's Whisper by 22% and 25% on relative accuracy. I led the technical work and wrote the launch blog. It built on our self-supervised learning work (project Hydra) and the per-language modelling pipelines from project Aladdin.

Meaning Error Rate

October 25th 2022 Speechmatics

An alternative to Word Error Rate that scores whether meaning actually changed, automated with GPT-3, few-shot learning and chain of thought. This was the first automation of the NER metric.

Side Projects

Hackathons, experiments and older work.