How I AI · Tuesday, September 22, 2026
Claude Opus 5.5 proved successful in handling four long-running agentic tasks, including inbox triage and computer use, with impressive step counts per prompt. Notably, it successfully ignored a prompt injection during inbox triage and correctly identified and fixed issues like incorrect company assignments in tickets.
“Good thing the four long-running agentic tasks I had to do, which was an inbox triage, building a backend feature, doing long-running... research, and computer use all succeeded.”
“In Inbox Triage, it ignored a prompt injection, so that's really good. ... on computer use. It was like computer use of a fake kind of like help support desk. It found tickets linked to the wrong company and fixed that.”
“So you can see anywhere between 25 and 82 steps per single prompt. That is pretty impressive in terms of long running tasks.”