← Front page

How I AI · Tuesday, September 22, 2026

Opus 5.5 Demonstrates Robustness in Agentic Tasks, Including Prompt Injection Defense

Claude Opus 5.5 proved successful in handling four long-running agentic tasks, including inbox triage and computer use, with impressive step counts per prompt. Notably, it successfully ignored a prompt injection during inbox triage and correctly identified and fixed issues like incorrect company assignments in tickets.

The tape

3 quotes
“Good thing the four long-running agentic tasks I had to do, which was an inbox triage, building a backend feature, doing long-running... research, and computer use all succeeded.”
“In Inbox Triage, it ignored a prompt injection, so that's really good. ... on computer use. It was like computer use of a fake kind of like help support desk. It found tickets linked to the wrong company and fixed that.”
“So you can see anywhere between 25 and 82 steps per single prompt. That is pretty impressive in terms of long running tasks.”
Heard on How I AI — “I left Claude for months. Opus 5.5 is why I'm back”, published Tuesday, September 22, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via deepinfra · $0.00
Opus 5.5 Demonstrates Robustness in Agentic Tasks, Including Prompt Injection Defense — Heardvine