Claude Code + Opus 5.5 High
Anthropic / Claude CodeAn early lead on High-effort implementation: the six-project Claude Code comparison favors Opus, with model-level benchmark corroboration and positive daily-use reports. Broken-flow reports and limited long-running evidence keep this provisional.
Model API$4 in $20 out per 1M tokens- Best for
- High-effort implementation and everyday repository changes, with review and tests
- Price & access
- Claude paid plan or API billing; select Opus 5.5 High. Cost per finished task varies by workload.
- Watch out for
- You need proven unattended reliability or a result guaranteed across frameworks and effort settings
- Restrictions
- Moderate
- Evidence
- AI Coding Daily: Opus High leads Astra High on the Sep 25 snapshot; AA High-vs-High supports coding capability. Every reports useful patches and serious failures at mixed efforts. These are external tests; exact harness builds and broader replication remain gaps.
- System version
- September 2026 Claude Code; exact build undisclosed in external test
- Tested access
- Claude paid plan or API billing; High effort; default safeguards
Sources and observations (4)
- AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026
September 25 leaderboard snapshot: Claude Code with Opus 5.5 High leads the six-project aggregate, ahead of Codex Astra High, with lower measured average API cost and task time. Perfect published behavioral subtotal; strong quality grading. Actual evaluation dated September 23. Limited projects and repetitions, automated quality judging, harness build undisclosed; publisher-specific sponsorship unknown. This directly tests High in Claude Code; not a debugging-specialty result.
- Artificial AnalysisSource record · Observed Sep 26, 2026
AA September 26 comparison: Opus 5.5 adaptive High with default fallback leads Astra High on Terminal-Bench 4.0 and SciCode, but Astra xhigh remains ahead of Opus High on Terminal-Bench. AA's model evaluation is contextual for Claude Code, not a Claude Code harness score. Intelligence-index cost per task is slightly higher than Astra High despite lower token prices; do not promise universally cheaper completion.
- Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026
Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Claude Code + Opus 5.5 High: creator preferred its outputs on all three builds for detail, prompt fidelity and gameplay. Supports creative coding quality; does not establish superiority on repository repair or every coding task. Creator still prefers Astra for everyday refactors/architecture. Scope: creative coding support within the broader coding rubric; repository performance remains unmeasured here.
- Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026
Every September 22 Opus 5.5 Vibe Check: original Claude Code/app tests support daily implementation and concise patches, with some builders switching from Fable/Codex. Contrary evidence includes broken core app flows, uncalled required service and destructive verification behavior. Examples use mixed efforts (including xhigh), so they are contextual for High, not matched High victories. Every received Anthropic pre-launch access; editorial independence asserted by Every, no independent-replication claim.
Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.
New evidence since this ranking (1)
These observations may support or challenge the existing recommendation. They have not changed its order, score, or verification date. Multiple articles from one publisher count as one source.
- AI Coding Daily — original coding benchmarkSource record · Reviewed Oct 3, 2026
October 3 table: Opus 5.5 High / Claude Code; $0.79, 3m10s, 67.41/70. Added bug-finding task changes denominator from 60; not a longitudinal improvement. Exact build and added-task rubric incompletely disclosed. One publisher; no standalone debugging qualification. Preserve prior observations.