Current AI rankings

Best AI Models & Tools

Coding / Writing / Images / Video

AI changes fast. Find the best models and tools for what you want to create—and see why they lead.

Know when the recommendations change.

Meaningful AI updates, plus On the Horizon: new releases we’re watching. About once a week. Only when there’s something worth sharing.

Scored rankings combine task quality and practical factors; small score gaps are close calls. Image and video use editorial first, second and third places, with reasons, tradeoffs and confidence rather than invented numerical scores. How we compare →

Choose by what you want to make. You don’t need to know how to code.

Models, coding tools, and builders: what’s the difference?
The model
Astra is the AI doing the reasoning. The tools it can use depend on where you access it.
The coding tool
Codex gives a model tools to work on project files, run tests, and fix problems. “Codex + Astra” names the combination being compared.
The builder
A platform such as Lovable brings building, databases, and publishing into one service. We judge the whole experience, including how easily you can maintain the result.

And ChatGPT Work? It’s a broader workspace for completing tasks; Codex provides the coding capabilities. Work and Astra aren’t competing models. The model, available tools, and setup together affect the result.

A tool can appear in more than one category. A website primarily presents information; a web app lets people do things with data. Some projects do both.

Choose a coding partner.

For building custom features or working on an existing project, with more control over the code.

The one to beat

Best Coding Overall

Our strongest all-around choice for planning, building, testing, and maintaining code. We weigh quality, reliability, workflow, and cost together.

CorrectnessRepository understandingTool useMaintainabilityWorkflow reliability

This is a provisional recommendation based on current external benchmarks and hands-on comparisons. The model, reasoning effort and coding tools matter. Broader testing may change the order; we reassess it within two days.

1

Claude Code + Opus 5.5 High

Anthropic / Claude Code

An early lead on High-effort implementation: the six-project Claude Code comparison favors Opus, with model-level benchmark corroboration and positive daily-use reports. Broken-flow reports and limited long-running evidence keep this provisional.

Model API$4 in $20 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
Best for
High-effort implementation and everyday repository changes, with review and tests
Price & access
Claude paid plan or API billing; select Opus 5.5 High. Cost per finished task varies by workload.
Watch out for
You need proven unattended reliability or a result guaranteed across frameworks and effort settings
Restrictions
Moderate
Evidence
AI Coding Daily: Opus High leads Astra High on the Sep 25 snapshot; AA High-vs-High supports coding capability. Every reports useful patches and serious failures at mixed efforts. These are external tests; exact harness builds and broader replication remain gaps.
System version
September 2026 Claude Code; exact build undisclosed in external test
Tested access
Claude paid plan or API billing; High effort; default safeguards
Sources and observations (4)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot: Claude Code with Opus 5.5 High leads the six-project aggregate, ahead of Codex Astra High, with lower measured average API cost and task time. Perfect published behavioral subtotal; strong quality grading. Actual evaluation dated September 23. Limited projects and repetitions, automated quality judging, harness build undisclosed; publisher-specific sponsorship unknown. This directly tests High in Claude Code; not a debugging-specialty result.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 26 comparison: Opus 5.5 adaptive High with default fallback leads Astra High on Terminal-Bench 4.0 and SciCode, but Astra xhigh remains ahead of Opus High on Terminal-Bench. AA's model evaluation is contextual for Claude Code, not a Claude Code harness score. Intelligence-index cost per task is slightly higher than Astra High despite lower token prices; do not promise universally cheaper completion.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Claude Code + Opus 5.5 High: creator preferred its outputs on all three builds for detail, prompt fidelity and gameplay. Supports creative coding quality; does not establish superiority on repository repair or every coding task. Creator still prefers Astra for everyday refactors/architecture. Scope: creative coding support within the broader coding rubric; repository performance remains unmeasured here.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 Vibe Check: original Claude Code/app tests support daily implementation and concise patches, with some builders switching from Fable/Codex. Contrary evidence includes broken core app flows, uncalled required service and destructive verification behavior. Examples use mixed efforts (including xhigh), so they are contextual for High, not matched High victories. Every received Anthropic pre-launch access; editorial independence asserted by Every, no independent-replication claim.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

New evidence since this ranking (1)

These observations may support or challenge the existing recommendation. They have not changed its order, score, or verification date. Multiple articles from one publisher count as one source.

  • AI Coding Daily — original coding benchmarkSource record · Reviewed Oct 3, 2026

    October 3 table: Opus 5.5 High / Claude Code; $0.79, 3m10s, 67.41/70. Added bug-finding task changes denominator from 60; not a longitudinal improvement. Exact build and added-task rubric incompletely disclosed. One publisher; no standalone debugging qualification. Preserve prior observations.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 23, 2026Official product / access information ↗
2

Codex + GPT-6 Astra

OpenAI / Codex

A close alternative: Astra High completes the benchmark's behavioral tasks but trails Opus High overall. Broader Astra evidence remains strong; xhigh terminal results must be kept separate from High. Existing access and value advantages remain useful.

Model API$10 in $50 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
When to choose Codex + GPT-6 Astra
Best for
Codex repository work, terminal tasks and users choosing effort to match the job
Price & access
Codex, ChatGPT Work or API; access and cost depend on plan and effort.
Watch out for
You want the leading High-effort result on the current six-project implementation suite
Restrictions
Low
Evidence
AI Coding Daily High-vs-High favors Opus on aggregate quality, time and cost. AA's High coding results also favor Opus; AA xhigh terminal results favor Astra. Every adds mixed-effort throughput and production-refactoring context.
System version
September 2026 production harness
Tested access
Codex, ChatGPT Work, or API
Sources and observations (5)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot explicitly tests Codex CLI + GPT-6 Astra High and Medium separately. High trails Claude Code + Opus 5.5 High on aggregate quality; both clear the behavior tasks. High is slower and costlier on this suite. High evidence is a configuration-specific observation within the historically broad Astra entry, not evidence about xhigh/max settings.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 26: Astra High trails Opus 5.5 High on terminal and scientific coding, while Astra xhigh retains a terminal advantage over Opus High. AA evaluates a different harness from Codex; settings must stay explicit. Astra retains stronger AutomationBench results. The broad Astra configuration supports choosing xhigh for hard terminal work, not claiming its High setting wins.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Codex + GPT-6 Astra, explicitly High in this comparison: Opus won creator preference on all three creative builds; Astra was faster in reported runs and cheapest on two. Creator still favors Astra for everyday refactors/architecture, a reported workflow preference rather than a measured debugging test. Do not transfer to xhigh or other efforts. Scope: creative coding support within the broader coding rubric; repository performance remains unmeasured here.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 review retains Astra for time-bounded work and reports higher raw throughput in one coding experiment, while Opus met more latency budgets. Scope is one experiment with mixed settings, not an Astra High benchmark. September 3 production-rewrite benchmark remains separate and used Astra extra-high.

  • Nate Herk — GPT-6 Astra vs Fable 5.1 practical comparisonSource page · Observed Sep 6, 2026

    Both configurations completed substantial practical work. The comparison favored Codex + Astra more often for day-to-day execution while preserving Claude Code + Fable 5.1 as stronger on selected tasks.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗
3

Claude Code + Fable 5.1

Anthropic / Claude Code

Still a strong choice for difficult implementation and careful codebase judgment. The newer direct comparisons reduce its claim to a general generation lead, while Every still prefers Fable for some hard problems; effort settings differ across tests.

Model API$10 in $50 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
When to choose Claude Code + Fable 5.1
Best for
Careful implementation and difficult existing-codebase problems
Price & access
Claude subscription or API billing; heavier usage can be costly.
Watch out for
Your priority is the strongest current High-effort implementation result or lowest task cost
Restrictions
Moderate
Evidence
Every retains Fable for some difficult tasks; AI Coding Daily tests Fable Medium, not High/Max. AA's earlier frontier SciCode result also uses a different effort. These comparisons support a cautious contender, not a uniform apples-to-apples defeat.
System version
September 2026 production release
Tested access
Claude Max or API billing
Sources and observations (5)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot: Claude Code + Fable 5.1 Medium trails Opus 5.5 High and Astra High on the six-project aggregate. The Fable row is Medium: it does not measure Fable High/Max. Retain that scope and use separate evidence for hard-problem frontier capability.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 22 original Opus 5.5 analysis: Opus max surpasses Fable 5.1's prior SciCode result and overall index, but this max comparison is contextual and cannot be relabeled Opus High versus Fable High. Fable remains a frontier coding model with premium pricing. Retain the older direct Fable configuration observations alongside this version-specific challenge.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Claude Code + Fable 5.1, explicitly High in this comparison: Opus won creator preference across all three creative builds. Fable showed speed but missed some requirements. These are task-level observations, not a universal third-place claim or evidence for other effort settings. Scope: creative coding support within the broader coding rubric; repository performance remains unmeasured here.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 review: some daily coding moves to Opus for value, but two testers still prefer Fable 5.1 on the hardest coding problems. Ruby comparison used Fable High versus Opus xhigh; do not treat effort as matched. This supports retaining Fable as a demanding-task alternative. Anthropic supplied pre-launch access to Opus; keep that relationship disclosed.

  • Nate Herk — GPT-6 Astra vs Fable 5.1 practical comparisonSource page · Observed Sep 6, 2026

    Both configurations completed substantial practical work. The comparison favored Codex + Astra more often for day-to-day execution while preserving Claude Code + Fable 5.1 as stronger on selected tasks.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗
Advanced comparisonsAutonomous coding · Writing code · Fixing & improving code

Why doesn’t the overall winner win every category?

Overall balances several kinds of work, reliability, workflow, and cost. A specialist can produce a more refined implementation or a better fix while another tool offers the strongest all-around combination. Builder rankings also judge the finished product and publishing experience, so their winners can differ.

How we weigh the evidence →

For delegating a multi-step development task: the agent plans, edits files, runs tests, and works through errors with limited supervision.

Task completionPlanningTool useLong-horizon reliabilityRecovery from errors

This is a provisional recommendation based on current external benchmarks and hands-on comparisons. The model, reasoning effort and coding tools matter. Broader testing may change the order; we reassess it within two days.

1

Codex + GPT-6 Astra

OpenAI / Codex

Retains the broader agentic lead through terminal and production-work evidence, including Astra xhigh. Opus High beats Astra High in the current direct coding comparisons, so this is a configuration-sensitive recommendation, not a blanket High-versus-High victory.

Model API$10 in $50 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
Best for
Codex repository work, terminal tasks and users choosing effort to match the job
Price & access
Codex, ChatGPT Work or API; access and cost depend on plan and effort.
Watch out for
You want the leading High-effort result on the current six-project implementation suite
Restrictions
Low
Evidence
AI Coding Daily High-vs-High favors Opus on aggregate quality, time and cost. AA's High coding results also favor Opus; AA xhigh terminal results favor Astra. Every adds mixed-effort throughput and production-refactoring context.
System version
September 2026 production harness
Tested access
Codex, ChatGPT Work, or API
Sources and observations (5)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot explicitly tests Codex CLI + GPT-6 Astra High and Medium separately. High trails Claude Code + Opus 5.5 High on aggregate quality; both clear the behavior tasks. High is slower and costlier on this suite. High evidence is a configuration-specific observation within the historically broad Astra entry, not evidence about xhigh/max settings.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 26: Astra High trails Opus 5.5 High on terminal and scientific coding, while Astra xhigh retains a terminal advantage over Opus High. AA evaluates a different harness from Codex; settings must stay explicit. Astra retains stronger AutomationBench results. The broad Astra configuration supports choosing xhigh for hard terminal work, not claiming its High setting wins.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Codex + GPT-6 Astra, explicitly High in this comparison: Opus won creator preference on all three creative builds; Astra was faster in reported runs and cheapest on two. Creator still favors Astra for everyday refactors/architecture, a reported workflow preference rather than a measured debugging test. Do not transfer to xhigh or other efforts. Scope: extended autonomous tool-using builds; not controlled repository repair. Deployment reminder/intervention limits one-shot claims.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 review retains Astra for time-bounded work and reports higher raw throughput in one coding experiment, while Opus met more latency budgets. Scope is one experiment with mixed settings, not an Astra High benchmark. September 3 production-rewrite benchmark remains separate and used Astra extra-high.

  • OpenAI Codex changelogSource page · Observed Sep 18, 2026

    Codex CLI 0.155.0 added live reasoning-status updates, agent-task archiving and worktree ownership, Touch ID verification for local MCP requests, recoverable daemon updates, and authentication recovery; 0.155.1 corrected reasoning-summary defaults for providers that reject them. These are material harness workflow and reliability changes, not direct model-quality evidence.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗
2

Claude Code + Opus 5.5 High

Anthropic / Claude Code

A strong new agentic challenger with better High-effort coding results in current comparisons. Reports of broken flows and a deleted verification artifact, plus thinner sustained-task evidence, keep it behind Astra's broader record for now.

Model API$4 in $20 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
When to choose Claude Code + Opus 5.5 High
Best for
High-effort implementation and everyday repository changes, with review and tests
Price & access
Claude paid plan or API billing; select Opus 5.5 High. Cost per finished task varies by workload.
Watch out for
You need proven unattended reliability or a result guaranteed across frameworks and effort settings
Restrictions
Moderate
Evidence
AI Coding Daily: Opus High leads Astra High on the Sep 25 snapshot; AA High-vs-High supports coding capability. Every reports useful patches and serious failures at mixed efforts. These are external tests; exact harness builds and broader replication remain gaps.
System version
September 2026 Claude Code; exact build undisclosed in external test
Tested access
Claude paid plan or API billing; High effort; default safeguards
Sources and observations (4)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot: Claude Code with Opus 5.5 High leads the six-project aggregate, ahead of Codex Astra High, with lower measured average API cost and task time. Perfect published behavioral subtotal; strong quality grading. Actual evaluation dated September 23. Limited projects and repetitions, automated quality judging, harness build undisclosed; publisher-specific sponsorship unknown. This directly tests High in Claude Code; not a debugging-specialty result.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 26 comparison: Opus 5.5 adaptive High with default fallback leads Astra High on Terminal-Bench 4.0 and SciCode, but Astra xhigh remains ahead of Opus High on Terminal-Bench. AA's model evaluation is contextual for Claude Code, not a Claude Code harness score. Intelligence-index cost per task is slightly higher than Astra High despite lower token prices; do not promise universally cheaper completion.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Claude Code + Opus 5.5 High: creator preferred its outputs on all three builds for detail, prompt fidelity and gameplay. Supports creative coding quality; does not establish superiority on repository repair or every coding task. Creator still prefers Astra for everyday refactors/architecture. Scope: extended autonomous tool-using builds; not controlled repository repair. Deployment reminder/intervention limits one-shot claims.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 Vibe Check: original Claude Code/app tests support daily implementation and concise patches, with some builders switching from Fable/Codex. Contrary evidence includes broken core app flows, uncalled required service and destructive verification behavior. Examples use mixed efforts (including xhigh), so they are contextual for High, not matched High victories. Every received Anthropic pre-launch access; editorial independence asserted by Every, no independent-replication claim.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 23, 2026Official product / access information ↗
3

Claude Code + Fable 5.1

Anthropic / Claude Code

Still a strong choice for difficult implementation and careful codebase judgment. The newer direct comparisons reduce its claim to a general generation lead, while Every still prefers Fable for some hard problems; effort settings differ across tests.

Model API$10 in $50 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
When to choose Claude Code + Fable 5.1
Best for
Careful implementation and difficult existing-codebase problems
Price & access
Claude subscription or API billing; heavier usage can be costly.
Watch out for
Your priority is the strongest current High-effort implementation result or lowest task cost
Restrictions
Moderate
Evidence
Every retains Fable for some difficult tasks; AI Coding Daily tests Fable Medium, not High/Max. AA's earlier frontier SciCode result also uses a different effort. These comparisons support a cautious contender, not a uniform apples-to-apples defeat.
System version
September 2026 production release
Tested access
Claude Max or API billing
Sources and observations (5)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot: Claude Code + Fable 5.1 Medium trails Opus 5.5 High and Astra High on the six-project aggregate. The Fable row is Medium: it does not measure Fable High/Max. Retain that scope and use separate evidence for hard-problem frontier capability.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 22 original Opus 5.5 analysis: Opus max surpasses Fable 5.1's prior SciCode result and overall index, but this max comparison is contextual and cannot be relabeled Opus High versus Fable High. Fable remains a frontier coding model with premium pricing. Retain the older direct Fable configuration observations alongside this version-specific challenge.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Claude Code + Fable 5.1, explicitly High in this comparison: Opus won creator preference across all three creative builds. Fable showed speed but missed some requirements. These are task-level observations, not a universal third-place claim or evidence for other effort settings. Scope: extended autonomous tool-using builds; not controlled repository repair. Deployment reminder/intervention limits one-shot claims.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 review: some daily coding moves to Opus for value, but two testers still prefer Fable 5.1 on the hardest coding problems. Ruby comparison used Fable High versus Opus xhigh; do not treat effort as matched. This supports retaining Fable as a demanding-task alternative. Anthropic supplied pre-launch access to Opus; keep that relationship disclosed.

  • Claude Code changelogSource page · Observed Sep 18, 2026

    Claude Code 2.1.274–2.1.276 added critical memory warnings, configurable MCP startup waits, synced claude.ai skills/plugins, marketplace plugin installation, improved queued-message control, and then repaired a proxy/gateway regression introduced in 2.1.275. These are material agent workflow and reliability changes, not direct task-quality benchmark evidence.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗

For turning clear instructions into correct, maintainable code. The focus is the implementation itself.

Functional correctnessSpecification adherenceCode qualityLanguage rangeEfficiency

This is a provisional recommendation based on current external benchmarks and hands-on comparisons. The model, reasoning effort and coding tools matter. Broader testing may change the order; we reassess it within two days.

1

Claude Code + Opus 5.5 High

Anthropic / Claude Code

An early lead on High-effort implementation: the six-project Claude Code comparison favors Opus, with model-level benchmark corroboration and positive daily-use reports. Broken-flow reports and limited long-running evidence keep this provisional.

Model API$4 in $20 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
Best for
High-effort implementation and everyday repository changes, with review and tests
Price & access
Claude paid plan or API billing; select Opus 5.5 High. Cost per finished task varies by workload.
Watch out for
You need proven unattended reliability or a result guaranteed across frameworks and effort settings
Restrictions
Moderate
Evidence
AI Coding Daily: Opus High leads Astra High on the Sep 25 snapshot; AA High-vs-High supports coding capability. Every reports useful patches and serious failures at mixed efforts. These are external tests; exact harness builds and broader replication remain gaps.
System version
September 2026 Claude Code; exact build undisclosed in external test
Tested access
Claude paid plan or API billing; High effort; default safeguards
Sources and observations (4)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot: Claude Code with Opus 5.5 High leads the six-project aggregate, ahead of Codex Astra High, with lower measured average API cost and task time. Perfect published behavioral subtotal; strong quality grading. Actual evaluation dated September 23. Limited projects and repetitions, automated quality judging, harness build undisclosed; publisher-specific sponsorship unknown. This directly tests High in Claude Code; not a debugging-specialty result.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 26 comparison: Opus 5.5 adaptive High with default fallback leads Astra High on Terminal-Bench 4.0 and SciCode, but Astra xhigh remains ahead of Opus High on Terminal-Bench. AA's model evaluation is contextual for Claude Code, not a Claude Code harness score. Intelligence-index cost per task is slightly higher than Astra High despite lower token prices; do not promise universally cheaper completion.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Claude Code + Opus 5.5 High: creator preferred its outputs on all three builds for detail, prompt fidelity and gameplay. Supports creative coding quality; does not establish superiority on repository repair or every coding task. Creator still prefers Astra for everyday refactors/architecture. Scope: generated code/output fidelity across three tasks; auxiliary image tools must not be credited to native model generation.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 Vibe Check: original Claude Code/app tests support daily implementation and concise patches, with some builders switching from Fable/Codex. Contrary evidence includes broken core app flows, uncalled required service and destructive verification behavior. Examples use mixed efforts (including xhigh), so they are contextual for High, not matched High victories. Every received Anthropic pre-launch access; editorial independence asserted by Every, no independent-replication claim.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 23, 2026Official product / access information ↗
2

Codex + GPT-6 Astra

OpenAI / Codex

A close alternative: Astra High completes the benchmark's behavioral tasks but trails Opus High overall. Broader Astra evidence remains strong; xhigh terminal results must be kept separate from High. Existing access and value advantages remain useful.

Model API$10 in $50 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
When to choose Codex + GPT-6 Astra
Best for
Codex repository work, terminal tasks and users choosing effort to match the job
Price & access
Codex, ChatGPT Work or API; access and cost depend on plan and effort.
Watch out for
You want the leading High-effort result on the current six-project implementation suite
Restrictions
Low
Evidence
AI Coding Daily High-vs-High favors Opus on aggregate quality, time and cost. AA's High coding results also favor Opus; AA xhigh terminal results favor Astra. Every adds mixed-effort throughput and production-refactoring context.
System version
September 2026 production harness
Tested access
Codex, ChatGPT Work, or API
Sources and observations (5)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot explicitly tests Codex CLI + GPT-6 Astra High and Medium separately. High trails Claude Code + Opus 5.5 High on aggregate quality; both clear the behavior tasks. High is slower and costlier on this suite. High evidence is a configuration-specific observation within the historically broad Astra entry, not evidence about xhigh/max settings.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 26: Astra High trails Opus 5.5 High on terminal and scientific coding, while Astra xhigh retains a terminal advantage over Opus High. AA evaluates a different harness from Codex; settings must stay explicit. Astra retains stronger AutomationBench results. The broad Astra configuration supports choosing xhigh for hard terminal work, not claiming its High setting wins.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Codex + GPT-6 Astra, explicitly High in this comparison: Opus won creator preference on all three creative builds; Astra was faster in reported runs and cheapest on two. Creator still favors Astra for everyday refactors/architecture, a reported workflow preference rather than a measured debugging test. Do not transfer to xhigh or other efforts. Scope: generated code/output fidelity across three tasks; auxiliary image tools must not be credited to native model generation.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 review retains Astra for time-bounded work and reports higher raw throughput in one coding experiment, while Opus met more latency budgets. Scope is one experiment with mixed settings, not an Astra High benchmark. September 3 production-rewrite benchmark remains separate and used Astra extra-high.

  • Nate Herk — GPT-6 Astra vs Fable 5.1 practical comparisonSource page · Observed Sep 6, 2026

    The software-build task compared the exact Codex + Astra and Claude Code + Fable 5.1 configurations rather than attributing results to naked models.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗
3

Claude Code + Fable 5.1

Anthropic / Claude Code

Still a strong choice for difficult implementation and careful codebase judgment. The newer direct comparisons reduce its claim to a general generation lead, while Every still prefers Fable for some hard problems; effort settings differ across tests.

Model API$10 in $50 out per 1M tokens
Multiple* Early evidence — position may changeTested configuration
When to choose Claude Code + Fable 5.1
Best for
Careful implementation and difficult existing-codebase problems
Price & access
Claude subscription or API billing; heavier usage can be costly.
Watch out for
Your priority is the strongest current High-effort implementation result or lowest task cost
Restrictions
Moderate
Evidence
Every retains Fable for some difficult tasks; AI Coding Daily tests Fable Medium, not High/Max. AA's earlier frontier SciCode result also uses a different effort. These comparisons support a cautious contender, not a uniform apples-to-apples defeat.
System version
September 2026 production release
Tested access
Claude Max or API billing
Sources and observations (5)
  • AI Coding Daily — original coding benchmarkSource record · Observed Sep 26, 2026

    September 25 leaderboard snapshot: Claude Code + Fable 5.1 Medium trails Opus 5.5 High and Astra High on the six-project aggregate. The Fable row is Medium: it does not measure Fable High/Max. Retain that scope and use separate evidence for hard-problem frontier capability.

  • Artificial AnalysisSource record · Observed Sep 26, 2026

    AA September 22 original Opus 5.5 analysis: Opus max surpasses Fable 5.1's prior SciCode result and overall index, but this max comparison is contextual and cannot be relabeled Opus High versus Fable High. Fable remains a frontier coding model with premium pricing. Retain the older direct Fable configuration observations alongside this version-specific challenge.

  • Pat Simmons / AI for Mortals — practical coding comparisonsSource record · Observed Sep 23, 2026

    Original comparison VxzdNX6mNSQ and companion AI for Mortals results article; tests September 23, 2026. Same prompts, High effort on all models in Claude Code/Codex, three creative builds (code-only illustration, launch site/film, 3D game). Exact harness builds undisclosed; shared gpt-image-2 where allowed. Single creator's practical comparison; not a BMN reproduction, objective repository benchmark, or separate publisher from the article. Claude Code + Fable 5.1, explicitly High in this comparison: Opus won creator preference across all three creative builds. Fable showed speed but missed some requirements. These are task-level observations, not a universal third-place claim or evidence for other effort settings. Scope: generated code/output fidelity across three tasks; auxiliary image tools must not be credited to native model generation.

  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource record · Observed Sep 22, 2026

    Every September 22 Opus 5.5 review: some daily coding moves to Opus for value, but two testers still prefer Fable 5.1 on the hardest coding problems. Ruby comparison used Fable High versus Opus xhigh; do not treat effort as matched. This supports retaining Fable as a demanding-task alternative. Anthropic supplied pre-launch access to Opus; keep that relationship disclosed.

  • Nate Herk — GPT-6 Astra vs Fable 5.1 practical comparisonSource page · Observed Sep 6, 2026

    The software-build task compared the exact Codex + Astra and Claude Code + Fable 5.1 configurations rather than attributing results to naked models.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 26, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗

For finding the cause of a bug or improving an existing project without breaking what already works.

DiagnosisFix correctnessRegression avoidanceRefactoring qualityExplanation
1

Claude Code + Fable 5.1

Anthropic / Claude Code

Retains the lead through codebase judgment, disciplined edits, and current product-quality preference.

Model API$10 in $50 out per 1M tokens
MultipleWeighted score: 92.2Tested configuration
Best for
Careful implementation, large codebases, and strong product judgment
Price & access
Claude Max or API billing; heavy usage can become expensive.
Watch out for
Speed and cost per completed task matter more than refinement
Restrictions
Moderate
Evidence
Current independent benchmarks and practitioner testing support frontier quality and strong product judgment, with higher cost and slower task completion.
System version
September 2026 production release
Tested access
Claude Max or API billing
Sources and observations (1)

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗
2

Codex + GPT-6 Astra

OpenAI / Codex

Strong fresh production-rewrite and efficiency evidence secure second, with restraint issues preventing first.

Model API$10 in $50 out per 1M tokens
MultipleWeighted score: 90.6Tested configuration
When to choose Codex + GPT-6 Astra
Best for
Fast agentic repository work, terminal tasks, and cost-aware frontier coding
Price & access
Codex, ChatGPT Work, or API; frontier access remains plan-dependent.
Watch out for
You need the broadest immediate availability or strongly prefer conservative product judgment
Restrictions
Low
Evidence
Current Artificial Analysis, terminal, Every, and matched practitioner evidence support frontier performance, speed, and cost efficiency.
System version
September 2026 production harness
Tested access
Codex, ChatGPT Work, or API
Sources and observations (2)
  • Every — current coding-agent Vibe Checks and Senior Engineer BenchmarkSource page · Observed Sep 3, 2026

    Astra's 71/100 production-code rewrite materially exceeded both Fable 5.1 runs, but Every documented overbuilding and stopping-judgment weaknesses.

  • Artificial AnalysisReferenced source page · Observed Sep 12, 2026

    Referenced in the published evaluation or explanation. A specific source observation is not attached here; this link is context, not a substitute for a documented comparison.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗
3

Claude Code + Opus 5

Anthropic / Claude Code

A dependable third-place debugger with mature root-cause analysis.

Model API$5 in $25 out per 1M tokens
MultipleWeighted score: 88.8Tested configuration

Supporting source links for this published explanation are still being reconciled.

When to choose Claude Code + Opus 5
Best for
Mature repository architecture, diagnosis, and regression-conscious changes
Price & access
Claude subscription or API billing.
Watch out for
You want the newest performance leader or best cost efficiency
Restrictions
Moderate
Evidence
Independent repository and coding-agent evidence supports mature, dependable engineering performance.
System version
September 2026 production release
Tested access
Claude Max or API billing
Ranking last verified: Sep 12, 2026 · Configuration tested: Sep 7, 2026Official product / access information ↗
Also excellent

Cursor + Fable 5.1

Cursor / Anthropic

A close alternative to third place (0.2 weighted points apart). This gap does not establish a meaningful quality difference.

Model API$10 in $50 out per 1M tokens
Weighted score: 88.6

From an idea to a working product.

Build, revise, fix, and run a working product—without needing to be a developer.

What makes a good builder? →
Practical consensus

Best Website Builder

Create a portfolio, business homepage, or publication. We look for polished design, easy editing, and a reliable way to publish.

Creator consensusFinished websitesOwnership & publishingEditing & recoveryEase & running costs
1

Codex + GPT-6 Astra

OpenAI / Codex

The provisional design leader for custom websites, producing unusually strong first-pass visual hierarchy, layering, motion, and brand fit in current hands-on testing.

Model API$10 in $50 out per 1M tokens
MultipleWeighted score: 92.3Provisional: evidence still developingTested configuration
Best for
Distinctive custom marketing sites with code ownership and high visual ambition
Price & access
Codex, ChatGPT Work, and API; API starts at $10/M input and $50/M output
Watch out for
You need a visual editor that a nontechnical owner can maintain alone
Restrictions
Low
Evidence
Nate Herk's reference-site test and current independent design testing support the lead; full Creator Consensus is not yet complete.
System version
September 2026 production harness
Tested access
Codex, ChatGPT Work, or API
Sources and observations (1)

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 10, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗
2

Framer

Framer

A leading provisional design-first website builder with fast AI generation, strong responsive control, polished motion, and an approachable visual canvas.

MultipleWeighted score: 91.0Provisional: evidence still developingTested configurationPlatform-selected AI

Supporting source links for this published explanation are still being reconciled.

When to choose Framer
Best for
Marketing sites and portfolios where visual polish matters most
Price & access
Free entry tier with paid hosting and collaboration plans
Watch out for
You need a complex application backend or unrestricted source-code ownership
Restrictions
Low
Evidence
Current specialist website-builder assessments and first-party feature evidence support a top-three result; full creator consensus is pending.
System version
September 2026 production release
Tested access
Current public paid plan
AI model
Selected automatically by Framer. The exact model—or combination of models—may change as the service evolves.
Ranking last verified: Sep 10, 2026 · Configuration tested: Sep 12, 2026Official product / access information ↗
3

Claude Code + Fable 5.1

Anthropic / Claude Code

A provisional top-three custom-site configuration with excellent final polish after iteration, though it required more rounds than Astra in the matched comparison.

Model API$10 in $50 out per 1M tokens
MultipleWeighted score: 90.5Provisional: evidence still developingTested configuration
When to choose Claude Code + Fable 5.1
Best for
Carefully iterated custom sites and teams comfortable directing a coding agent
Price & access
Claude Max or API billing
Watch out for
You want the strongest one-shot starting point or a visual no-code editor
Restrictions
Moderate
Evidence
Nate Herk's matched reference-site test preserves the exact Claude Code + Fable 5.1 configuration; additional independent creator ballots are pending.
System version
September 2026 production release
Tested access
Claude Max or API billing
Sources and observations (1)

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 10, 2026 · Configuration tested: Sep 6, 2026Official product / access information ↗
Practical consensus

Best Web App Builder

Build something people use in a browser, such as a booking system, CRM, or writing tool—with accounts, saved data, and working features.

Creator consensusWorking productsOwnership & publishingEase & revisionRecovery & running costs
1

Lovable

Lovable

Three qualifying practitioner verdicts converge on Lovable, with finished-product proof and strong ownership.

MultipleWeighted score: 97.0Tested configurationPlatform-selected AI
Best for
Polished full-stack MVPs with auth, data, deployment, and GitHub ownership
Price & access
See current public free and paid plans
Watch out for
You need broad framework choice or deep manual control
Restrictions
Low
Evidence
Three qualifying verdicts rank Lovable first; two are independent, and one demonstrates a functioning deployed project tracker.
System version
September 2026 production release
Tested access
Current public paid plan
AI model
Selected automatically by Lovable. The exact model—or combination of models—may change as the service evolves.
Sources and observations (3)
  • LOW/CODE Agency — 2026 no-code builder comparisonSource page · Observed Sep 3, 2026

    LOW/CODE's explicitly numbered vibe-coding section places Lovable first and calls it the best AI app builder for nontechnical founders, based on recent internal testing across real workflows.

  • Vibe Coding Academy — 2026 AI app builder guideSource page · Observed May 1, 2026

    Vibe Coding Academy names Lovable the strongest all-around AI web app builder and demonstrates a deployed project tracker with authentication, database-backed projects and tasks, and iterative dashboard changes.

  • Till Freitag — Base44 vs Lovable agency comparisonSource page · Observed Mar 17, 2026

    After using both builders on several agency projects and comparing a CRM dashboard workflow, Till Freitag prefers Lovable for production and client work because of code and backend portability; its Lovable services create a material commercial-interest discount.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026 · Configuration tested: Sep 12, 2026Official product / access information ↗
2

Bolt

StackBlitz

The explicit consensus runner-up, pairing fast iteration with exportable code.

MultipleWeighted score: 49.3Tested configurationPlatform-selected AI
When to choose Bolt
Best for
Rapid full-stack prototypes with code control
Price & access
See current public free and paid plans
Watch out for
You want the most forgiving all-in-one backend setup
Restrictions
Low
Evidence
LOW/CODE ranks Bolt second; current first-party evidence supports code ownership and deployment.
System version
September 2026 production release
Tested access
Current public paid plan
AI model
Selected automatically by Bolt. The exact model—or combination of models—may change as the service evolves.
Sources and observations (1)

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026 · Configuration tested: Sep 12, 2026Official product / access information ↗
3

Base44

Wix / Base44

The easiest integrated option for internal and data-heavy apps, held back by backend lock-in.

MultipleWeighted score: 37.8Tested configurationPlatform-selected AI
When to choose Base44
Best for
Internal tools and business apps for nontechnical users
Price & access
See current public free and paid plans
Watch out for
Portable backend code is essential
Restrictions
Moderate
Evidence
LOW/CODE ranks Base44 third; practitioner evidence supports completion and ease while documenting weaker portability.
System version
September 2026 production release
Tested access
Current public paid plan
AI model
Selected automatically by Base44. The exact model—or combination of models—may change as the service evolves.
Sources and observations (1)

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026 · Configuration tested: Sep 12, 2026Official product / access information ↗
Practical consensus

Best Mobile App Builder

Create an installable iPhone or Android app. Device testing, code ownership, and a practical path to app stores matter here.

Creator consensusWorking mobile appsCode ownershipStore pathwayEase, recovery & value
1

FlutterFlow

FlutterFlow

The strongest provisional mobile-app result because it combines AI assistance with a mature visual Flutter workflow, testing tools, code export, and native-store publishing.

MultipleWeighted score: 89.7Provisional: evidence still developingTested configurationPlatform-selected AI
Best for
Cross-platform iOS and Android apps that must reach native stores
Price & access
Paid FlutterFlow plan required for key export and deployment workflows
Watch out for
You want a purely conversational builder with no visual-tool learning curve
Restrictions
Low
Evidence
Verified native-publishing capability and current specialist signals support the lead; more qualified same-project creator comparisons are required.
System version
September 2026 production release
Tested access
Current public paid plan
AI model
Selected automatically by FlutterFlow. The exact model—or combination of models—may change as the service evolves.
Sources and observations (3)
  • LOW/CODE Agency — 2026 no-code builder comparisonSource page · Observed Sep 3, 2026

    The agency identifies FlutterFlow as its strongest fit for cross-platform native apps needing source ownership, based on recent hands-on evaluation and production experience. The article does not publish a category-specific ranked top-three ballot.

  • The App Builder Report — Rork vs FlutterFlowSource page · Observed Aug 14, 2026

    The comparison explicitly judges FlutterFlow more capable than Rork for repeat building, custom logic, code ownership, and mature store deployment. FlutterFlow was assessed from documentation and independent coverage rather than the author's direct use; the publisher's Rork affiliate relationship discounts the verdict.

  • Unite.AI — Superapp hands-on reviewSource page · Observed Jul 31, 2026

    In the same hands-on review, FlutterFlow is characterized as the more mature choice for cross-platform control and customization. The comparison is useful but affiliate-discounted and is not a complete mobile top-three ballot.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 10, 2026 · Configuration tested: Sep 12, 2026Official product / access information ↗
2

Draftbit

Draftbit

A strong provisional mobile runner-up with React Native and Expo foundations, visual editing, code ownership, previewing, and practical publishing control.

MultipleWeighted score: 85.3Provisional: evidence still developingTested configurationPlatform-selected AI
When to choose Draftbit
Best for
Teams wanting a visual mobile builder with exportable React Native code
Price & access
Paid plans unlock production collaboration and export workflows
Watch out for
You want the largest no-code community or the simplest one-click path
Restrictions
Low
Evidence
First-party workflow verification and current mobile-builder comparisons support eligibility, but not yet a settled consensus.
System version
September 2026 production release
Tested access
Current public paid plan
AI model
Selected automatically by Draftbit. The exact model—or combination of models—may change as the service evolves.
Sources and observations (1)
  • LOW/CODE Agency — 2026 no-code builder comparisonSource page · Observed Sep 3, 2026

    The agency identifies Draftbit as a strong React Native option for semi-technical teams prioritizing source ownership. The article does not publish a category-specific ranked top-three ballot.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 10, 2026 · Configuration tested: Sep 12, 2026Official product / access information ↗
3

Replit Agent

Replit

The broadest provisional all-in-one builder across web and mobile workflows, with an integrated environment, deployment, database, and routed agent stack.

MultipleWeighted score: 85.2Provisional: evidence still developingTested configurationPlatform-selected AI

Supporting source links for this published explanation are still being reconciled.

When to choose Replit Agent
Best for
Builders who want one workspace from prompt through deployment
Price & access
Replit Core or Teams with usage-based agent costs
Watch out for
You need predictable credit use or the most polished first-pass design
Restrictions
Low
Evidence
Current product capabilities and comparative builder evidence support the breadth advantage; the Creator Consensus threshold remains incomplete.
System version
September 2026 production release
Tested access
Replit Core or Teams
AI model
Selected automatically by Replit Agent. The exact model—or combination of models—may change as the service evolves.
Ranking last verified: Sep 10, 2026 · Configuration tested: Sep 7, 2026Official product / access information ↗

API pricing

Leading text & coding models. US dollars per million tokens.

Prices checked Sep 26, 2026

Category winners + On the Horizon

ModelInput 0 — $10Output 0 — $50
Input$2
Output$10
Input$4
Output$20
Input$10
Output$50
Input$10
Output$50

Each column has its own scale. Standard uncached input and output rates; short-context pricing where applicable.

Input is what you send; output is what the model generates, including billable thinking. More thinking can mean more tokens at the same rate. These are API costs, not monthly subscriptions or a price per completed task.

Sources, thinking levels & pricing details

Compare standard paid API access. Caching, batch discounts, fast service, regional routing, tool fees and taxes can change the bill. Coding tools may add their own charges. Low, medium and high thinking are provider-specific settings, not comparable amounts of reasoning.

Gemini 3.8 Flash

Paid standard introductory rate through December 31, 2026; scheduled to become $1.50 input / $7.50 output on January 1, 2027. Output includes thinking tokens.

Provider pricing ↗Checked 2026-09-26

GPT-6 Sol

Standard rate for prompts up to 272K input tokens. Above 272K, the full request costs 2× the input rate and 1.5× the output rate. Low, medium and high reasoning share these token rates; extra reasoning adds billable output tokens.

Provider pricing ↗Checked 2026-09-26

Claude Opus 5.5

Global standard rate across the supported context window. Low, medium and high effort share these token rates; thinking adds billable output. Fast mode costs extra.

Provider pricing ↗Checked 2026-09-26

GPT-5.6 Sol

Promotional standard short-context rate, available at least through November 21, 2026. Above 272K input tokens, the full request costs 2× input and 1.5× output. Thinking adds billable output tokens.

Provider pricing ↗Checked 2026-09-26

Claude Opus 5

Global standard rate across the supported context window. Thinking is included in billed output; fast mode costs extra.

Provider pricing ↗Checked 2026-09-26

Claude Fable 5.1

Global standard rate across the supported context window. Thinking is included in billed output; effort changes token use, not this rate.

Provider pricing ↗Checked 2026-09-26

GPT-6 Astra

Standard rate for prompts up to 272K input tokens. Above 272K, the full request costs 2× the input rate and 1.5× the output rate. Low, medium and high reasoning share these token rates; extra reasoning adds billable output tokens.

Provider pricing ↗Checked 2026-09-26

Thinking billing: OpenAI · Anthropic · Google

On the Horizon

New announcements we’re following. A confirmed launch is not a recommendation: independent comparisons and BMN’s category reviews still matter.

Black Forest LabsAwaiting BMN review

FLUX 3 Image

Announced Oct 1, 2026

Black Forest Labs released the image-generation and editing endpoint of FLUX 3 on October 1. Public API and playground access, up to ten references, bounding-box composition and up to 4K output are documented. These are supplier claims, not a BMN quality verdict.

What we’re checking

October 4 official release notes, Kingy launch article, Creative AI News pixel audit and Fuser comparison checked. FLUX3 registered47820. New sources321/322 approved for methods/discovery only: vendor-selected samples, mixed sizes/dates, few trials and unresolved relationships prevent independent corroboration. No supported new order for editing, artistry or photography. Seek matched independent tests; recheck by October6.

Original announcement ↗

Source checked Oct 4, 2026 · Next review target Oct 6, 2026

GoogleAwaiting BMN review

Gemini 4 Argon

Announced Sep 30, 2026

Google announced Gemini 4 Argon for complex long-horizon coding, enterprise knowledge work and cyber defense. Access is currently limited to trusted cyber defenders through the Fairwind Program; Google says broader developer, enterprise and consumer access will follow. Vendor performance claims remain attributed.

What we’re checking

October4 official announcement, Gemini API changelog and Arena listing checked. Access remains restricted to trusted Fairwind users; exact generally usable API configuration is not verified. Arena listing does not establish production access. No writing screen is eligible yet. Reassess by October6.

Original announcement ↗

Source checked Oct 4, 2026 · Next review target Oct 6, 2026

AnthropicEvaluation in progress

Claude Sonnet 5.5

Announced Sep 28, 2026

Anthropic announced Claude Sonnet 5.5 on September 28, 2026. The public API model is claude-sonnet-5-5; vendor capability claims remain attributed and do not establish a BMN ranking.

What we’re checking

October4 official page, Vals, CodeRabbit, Flavio and AI Coding Daily checked. API High vs Claude Code Medium defaults and external High/Max harnesses kept separate. Writing results reused: Fiction screened out; Screenwriting suite82 processed, no approval pending. Replit Power configuration47823 now registered, untested. No supported coding podium displacement; need second matched comparative publisher. Reassess by October6.

Original announcement ↗

Source checked Oct 4, 2026 · Next review target Oct 6, 2026

Know when the recommendations change.

Meaningful AI updates, plus On the Horizon: new releases we’re watching. About once a week. Only when there’s something worth sharing.

September 26, 2026

Opus 5.5 High enters the coding podium.

Today’s external-evidence review puts Claude Code + Opus 5.5 High first in Best Coding Overall and Code Generation, and second in Agentic Coding. These early recommendations carry an asterisk and a two-day review cycle. Exact effort settings, contrary results, and test limitations remain visible. Coding reviews proceed independently of pending writing approvals.

September 25, 2026

Clear recommendations. Reasons you can inspect.

Image and video reviews now support explicit first, second and third places, with category-specific criteria and published tiebreak decisions. Photorealism covers everyday phone snapshots as well as professional photography. Confidence and evidence gaps remain visible alongside our recommendations.

September 22, 2026

Image and video evidence, with its limits.

New source observations identify tested variants, task-specific findings, and conflicts with existing recommendations. They appear separately from the evidence behind the older rankings. Generic video summaries now receive the same provisional treatment as image summaries; scores and ranking-verification dates remain unchanged.

September 21, 2026

Clearer comparisons. Evidence you can inspect.

Larger text, direct category links, and expandable comparisons make the rankings easier to use. Evidence gaps are now visible, including inherited image placements awaiting version-specific confirmation. Ranking order and verification dates have not been advanced by this presentation update.

September 10, 2026

ChatGPT Images 2.5 replaces GPT Image 2.

OpenAI’s new image release adds faster generation, stronger subject preservation, more reliable multi-turn edits, and new direct-editing tools. The release initially inherited GPT Image 2’s placements. Those inherited placements are provisional until version-specific comparative evidence is documented. GPT-6 Astra’s access and API pricing have also been brought up to date.

September 8, 2026

Coding now ranks tested configurations.

Every Coding result identifies the system, underlying model or platform-selected AI, access plan, system version, and test date. Technical awards cleared the first supervised sweep; Builder awards are provisional until the full Practical Consensus threshold is met.

Evidence in.
A useful answer out.

We assess outside findings for the task and tested configuration. No single benchmark, arena, reviewer, or company determines a winner. These weights apply when comparable evidence supports scored rankings. Image and video use ranked editorial recommendations with explicit tiebreakers, confidence and evidence limitations; they have no numerical scores. Builders use a separate profile: 20% working-product proof, 60% practitioner consensus, 10% ownership and publishing, and 10% ease and value.

Read the full methodology →
Quality

Category performance

Judge the task: artistic expression, photographic believability, motion, or editing precision.

Evidence

Comparisons that fit

Blind arenas and practitioner comparisons inform the decision, with exact versions and task limits preserved.

Tiebreaks

A clear order

Consistency, instruction fidelity, creative control, breadth, then a documented editorial judgment.

Buying advice

Price & access

Separate from quality-specialty ordering. Best Overall considers practical workflow and value after quality.

Confidence

Limits in view

A recommendation can be useful without being indisputable. We explain close calls and missing comparisons.

Gate

Complete candidate pool

Every contender in the defined, dated evaluation field must be evaluated. New discoveries enter through an eligibility review.