← Front page

The Cognitive Revolution · Thursday, July 9, 2026

Counterfactual Reflection: A New Training Method for AI Alignment

A novel training method called 'counterfactual reflection' has been introduced, where AI models are paused mid-task and trained to respond according to desired values. This method involves interrupting the model and guiding it toward a 'constitutionally right' response, which is then used for supervised training. Anthropic's research suggests this technique helps embed desired concepts into the J-space, improving overall model behavior.

companyAnthropic

The tape

3 quotes
A training method the paper called counterfactual reflection.
Basically, what they do is pause the model mid-task and then do supervised training on once interrupted, asking it like, what should we be doing here? Like, what's the constitutionally right thing to be doing in this moment?
Train on that, and it seems to allow the model, it seems to cause the model to bring into this J-space, this kind of global workspace memory type space, the concepts that Anthropic wants it to have on reflection.
Heard on The Cognitive Revolution — “AI:AM Highlights: Exploring the J-Space, AI Superforecasters, SambaNova's Chips, & LTX Video Gen, published Thursday, July 9, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.10
Counterfactual Reflection: A New Training Method for AI Alignment — Heardvine