← Front page

How I AI · Thursday, July 9, 2026

Speaker's Benchmark Favors GPT 5.6 Soul for Output Quality

The speaker's "Claire weighted index" for evaluating AI models, which balances LLM judging with personal taste (70% speaker, 30% LLM), found GPT 5.6 Soul to be the favorite. They praised Soul's output, particularly for prototyping and design, noting it "output the best work" and had the "highest taste score."

The tape

2 quotes
And so if you look at that 70-30 split, your girl loves 5.6 Soul. She just does. It had highest taste score by a significant amount. So I just thought it output the best work.
I did blind taste test these. And so I do really feel like it did a good job and I will give you a couple examples of that.
Heard on How I AI — “GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark, published Thursday, July 9, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.05
Speaker's Benchmark Favors GPT 5.6 Soul for Output Quality — Heardvine