← Front page

How I AI · Tuesday, June 30, 2026

Disagreement Between Human and AI Judges Highlighted

The benchmark revealed discrepancies between human subjective ratings and AI-driven evaluations, particularly concerning code quality and adherence to constraints. The host suggests that AI judges may lack the 'taste' and nuanced understanding of human evaluators.

The tape

2 quotes
What got flagged on the automated results? Well, things that the AI and I could not get flagged. So AI, I think was looking for broken working code. It ignored constraints, it was incomplete. Whereas I was just eyeball really the first screenshot.
I don't think these models are spiky enough when it comes to how they evaluate output. And I think like models are, let's say sloppy. And I don't think they have that vision of taste, uniqueness, what it looks like to the human eye.
Heard on How I AI — “Sonnet 5 review: I ran 64 generations to find out if it's worth it, published Tuesday, June 30, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.02
Disagreement Between Human and AI Judges Highlighted — Heardvine