Current AI rankings

Best AI for Writing

Writing / Coding / Images / Video

Compare AI for business writing, fiction prose, and screenwriting, with test results, practical tradeoffs, and supporting evidence.

Know when the recommendations change.

Meaningful AI updates, plus On the Horizon: new releases we’re watching. About once a week. Only when there’s something worth sharing.

Scored rankings combine task quality and practical factors; small score gaps are close calls. Qualified image and video recommendations show useful options, tested configurations, and unresolved questions without numerical scores. How we compare →

The one to beat

Best Writing Overall

The strongest practical writing model across quality, control, versatility, access, and reliability.

Writing qualityInstruction controlRevision strengthLong-form coherenceAccess & reliability
1

Claude Fable 5.1

Anthropic

The strongest current balance of nuanced writing, instruction control, revision judgment, and sustained long-form quality.

MultipleWeighted score: 91.5
Best for
High-stakes long-form writing, voice-sensitive drafts, and demanding revision
Price & access
Claude subscription and API; premium API cost
Watch out for
Cost and throughput matter more than the last increment of prose quality
Restrictions
Moderate
Evidence
Leads the current Arena writing preference view and ties Astra atop Artificial Analysis intelligence, while outperforming Astra on AA-Briefcase and GDPval-AA v2.
Sources and observations (2)
  • LMArena leaderboardSource page · Observed Sep 12, 2026

    Claude Fable 5.1 (Max) currently leads Arena's writing/preference view at 13.85% ±1.92%.

  • Artificial AnalysisSource page · Observed Sep 12, 2026

    Claude Fable 5.1 (max with fallback) ties GPT-6 Astra for first in Intelligence Index v4.3 at 53 and leads Astra on AA-Briefcase and GDPval-AA v2.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026Official product / access information ↗
2

GPT-6 Astra

OpenAI

Nearly matches the writing leader while offering stronger cost efficiency, factual discipline, and concise execution.

MultipleWeighted score: 90.0
When to choose GPT-6 Astra
Best for
Versatile professional writing, structured revision, and quality-conscious users who also value efficiency
Price & access
ChatGPT Work and API; access remains plan-dependent
Watch out for
You want the most consistently distinctive narrative voice
Restrictions
Moderate
Evidence
Places second in Arena's current writing preference view, ties Fable 5.1 on the Artificial Analysis Intelligence Index, and costs substantially less per evaluated task.
Sources and observations (2)
  • LMArena leaderboardSource page · Observed Sep 12, 2026

    GPT-6 Astra (Max) currently places second at 12.39% ±2.60%; its interval overlaps the leader's.

  • Artificial AnalysisSource page · Observed Sep 12, 2026

    GPT-6 Astra (max) ties Claude Fable 5.1 at 53 while costing about 40% as much per Intelligence Index task.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026Official product / access information ↗
3

Claude Opus 5

Anthropic

Exceptional prose and editorial judgment earn third, but premium cost and weaker value keep it behind Astra.

MultipleWeighted score: 85.8
When to choose Claude Opus 5
Best for
Careful prose, high-stakes editing, and polished long-form work
Price & access
Claude subscription and API; premium pricing
Watch out for
You need economical high-volume drafting
Restrictions
Moderate
Evidence
Current Arena preference places Opus 5 directly behind Fable 5.1 and Astra; independent frontier evidence supports quality, though the value profile is weaker.
Sources and observations (1)
  • LMArena leaderboardSource page · Observed Sep 12, 2026

    Claude Opus 5 (High) currently places third at 11.06% ±1.70%.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026Official product / access information ↗
Also excellent

Gemini 3.8 Flash

Google

A close alternative to third place (0.1 weighted points apart). This gap does not establish a meaningful quality difference.

Weighted score: 85.7

Clear, persuasive professional writing that needs less cleanup before it can be used.

Clarity & structurePersuasionTone controlFactual disciplineRewriting
1

Claude Fable 5.1

Anthropic

The strongest current model for nuanced, persuasive, high-stakes business documents, with a clear lead on demanding knowledge-work evaluations.

MultipleWeighted score: 91.0
Best for
Strategy memos, proposals, executive briefs, and sensitive persuasive writing
Price & access
Claude subscription and API; premium API cost
Watch out for
High-volume routine copy makes cost the dominant concern
Restrictions
Moderate
Evidence
Leads Astra on AA-Briefcase and GDPval-AA v2, while current preference and practical comparison evidence support superior tone and document judgment.
Sources and observations (1)
  • Artificial AnalysisSource page · Observed Sep 12, 2026

    Claude Fable 5.1 leads GPT-6 Astra on AA-Briefcase (1662 vs 1562) and GDPval-AA v2 (1764 vs 1580), directly supporting demanding knowledge-work quality.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026Official product / access information ↗
2

GPT-5.6 Sol

OpenAI

The strongest practical value choice: polished presentation, dependable revision, mature access, and lower cost produce a razor-thin second-place finish.

MultipleWeighted score: 90.0
When to choose GPT-5.6 Sol
Best for
Everyday business documents, presentations, synthesis, and high-volume professional drafting
Price & access
ChatGPT and API; lower token pricing than Astra
Watch out for
You need the highest ceiling for unusually sensitive or persuasive documents
Restrictions
Moderate
Evidence
Artificial Analysis identifies GPT-5.6 Sol as the current presentation-quality leader in AA-Briefcase; broad access and strong economics reinforce its practical business-writing score.
Sources and observations (1)
  • Artificial AnalysisSource page · Observed Sep 12, 2026

    GPT-5.6 Sol remains the presentation-quality leader in AA-Briefcase and retains strong professional-work performance at lower token prices than Astra.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026Official product / access information ↗
3

GPT-6 Astra

OpenAI

Excellent structure, factual discipline, professional-document handling, and stronger efficiency than Fable 5.1, narrowly edged by Sol after all practical factors are weighted.

MultipleWeighted score: 89.8
When to choose GPT-6 Astra
Best for
Structured professional writing, document revision, and business work requiring strong reasoning with controlled cost
Price & access
ChatGPT Work and API; access remains plan-dependent
Watch out for
You prioritize mature broad access or the finest persuasive voice
Restrictions
Moderate
Evidence
Strong AA-Briefcase performance, leadership in AutomationBench-AA, a 31% all-criteria pass rate on GDP.pdf, and substantially lower evaluated-task cost than Fable 5.1.
Sources and observations (1)
  • Artificial AnalysisSource page · Observed Sep 12, 2026

    GPT-6 Astra trails Fable 5.1 on AA-Briefcase and GDPval-AA v2 but leads AutomationBench-AA, passes every GDP.pdf criterion on 31% of attempts, and offers materially lower cost per task.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 12, 2026Official product / access information ↗

Compelling prose with sustained voice, character, narrative control, and originality.

Prose qualityCharacter voiceNarrative controlOriginalityRewriting
1

Claude Opus 5

Anthropic

Claude Opus 5 opened strongest in the qualifier (1st) and held high revision fidelity (2nd). Opus balanced vivid surface detail with controlled narrative pacing; characters behaved believably and the revision exercise showed clear discipline responding to the brief while retaining voice. Its prose is confident and less eccentric than Fable, favoring clarity and reader-forward choices.

MultipleWeighted score: 84.6
Best for
Clear, voice-conscious commercial and literary scenes that need reliable editorial fidelity.
Price & access
Anthropic via OpenRouter; premium API subscription recommended for production use.
Watch out for
You want maximal stylistic risk or hyper-distinctive authorial color — Opus leans toward reader-facing clarity.
Restrictions
Moderate
Evidence
Placed 1st in qualifier and 2nd in final (9 points). Strong original draft presence and high-fidelity revisions; consistent scene control and believable character choices.
Sources and observations (1)
  • Best Model Now — Fiction Prose Blind TestBMN test · Observed Sep 15, 2026

    Fiction Prose Blind Test — Original + Revision: qualifier #1, final #2, 9/10 combined points. qualifier #1 · 1096 words; final #2 · 853 words.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 16, 2026Official product / access information ↗
2

GPT-6 Astra

OpenAI

GPT-6 Astra tied at the top of this suite (2nd in qualifier, 1st in final). Astra showed the cleanest revision discipline and outstanding narrative control in the final: tightened pacing, sharpened character choices, and decisive fidelity to the editorial brief while retaining an engaging voice. The qualifier was very strong but slightly less distinct than Opus in initial surface color.

MultipleWeighted score: 83.0
When to choose GPT-6 Astra
Best for
High-fidelity revisions, disciplined developmental rewrites, and producing publication-ready passages from a clear brief.
Price & access
OpenAI API (gpt-6-astra); currently newer but accessible via standard API channels.
Watch out for
You require highly idiosyncratic literary risk-taking on first draft; Astra favors clarity and control.
Restrictions
Moderate
Evidence
Placed 2nd in qualifier and 1st in final (9 points). Best evidence in this run for revision discipline and balanced voice + narrative control; strong specialist preference signal.
Sources and observations (1)
  • Best Model Now — Fiction Prose Blind TestBMN test · Observed Sep 15, 2026

    Fiction Prose Blind Test — Original + Revision: qualifier #2, final #1, 9/10 combined points. qualifier #2 · 1158 words; final #1 · 758 words.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 16, 2026Official product / access information ↗
3

Claude Fable 5.1

Anthropic

Claude Fable 5.1 produced the most recognizably literary draft in the qualifier — rich detail and a strong, sustained narrator — but its controlled revision ranked third. Strengths: textured scene-setting, consistent character motivation, and a clear narrative throughline. Weaknesses: revision trimmed idiosyncrasies that gave voice its edge and occasionally flattened stakes; revision discipline was competent but cautious, reducing risk rather than sharpened focus.

MultipleWeighted score: 63.5
When to choose Claude Fable 5.1
Best for
Drafting voice-forward scenes, literary fiction where texture and sustained narrator matter.
Price & access
Claude subscription and premium API tier (Anthropic OpenRouter).
Watch out for
You need aggressive, directive revisions that reshape pacing or plot logic instead of polishing voice.
Restrictions
Moderate
Evidence
Placed 3rd in both qualifier and final (6 points). Qualifier showed superior scene texture and voice; final showed solid but conservative fidelity to the brief, losing some voice-specific risks.
Sources and observations (1)
  • Best Model Now — Fiction Prose Blind TestBMN test · Observed Sep 15, 2026

    Fiction Prose Blind Test — Original + Revision: qualifier #3, final #3, 6/10 combined points. qualifier #3 · 1120 words; final #3 · 820 words.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 16, 2026Official product / access information ↗

Professional dramatic writing with strong scenes, distinct voices, subtext, and clean formatting.

Dialogue & subtextScene constructionPacingCharacter voiceFormatting & revision
1

GPT-6 Astra

OpenAI

GPT-6 Astra (OpenAI) placed 1st in both the qualifier and the final (12 points, normalized 100). Both drafts were complete, well-structured, and controlled tone and dialogue across nearly 2k total words. The outputs used the model's Low reasoning setting and required no visible structural corrections in blind review — a clear specialist-level showing for scene construction and dialogue under pressure.

MultipleWeighted score: 90.0
Best for
Full scene drafts, character-driven dialogue, early-structure-first passes.
Price & access
OpenAI subscription / premium API tier.
Watch out for
This scene-writing test does not establish performance on a complete screenplay or a dedicated rewriting task.
Restrictions
Moderate
Evidence
Qualifier: placement 1, 938 words, Low reasoning, 79s generation. Final: placement 1, 1,019 words. Dominant across both scenes; high consistency and scene control.
Sources and observations (1)
  • Best Model Now — Screenwriting Blind TestBMN test · Observed Sep 11, 2026

    Screenwriting Blind Test — Qualifier + Final: qualifier #1, final #1, 12/12 combined points. qualifier #1 · 938 words; final #1 · 1019 words.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 11, 2026Official product / access information ↗
2

Gemini 3.8 Flash

Google

Gemini 3.8 Flash (Google) showed an uneven profile: a weak qualifier (5th) but a strong final (2nd), totaling 7 points (normalized 58.3). Word counts were the longest of the batch, and the final scene demonstrates that Gemini can produce compelling long-form scene work — but inconsistency between scenes is a concern in blind specialist evaluation.

MultipleWeighted score: 68.4
When to choose Gemini 3.8 Flash
Best for
Long-context drafts and extended scene work where breadth and richness matter.
Price & access
Google Cloud / Gemini access.
Watch out for
You need consistently high polish on short turnaround character beats.
Restrictions
Moderate
Evidence
Qualifier: placement 5, 1,362 words. Final: placement 2, 1,261 words. Strong final-scene control, inconsistent qualifier execution suggests variable stability across prompts.
Sources and observations (1)
  • Best Model Now — Screenwriting Blind TestBMN test · Observed Sep 11, 2026

    Screenwriting Blind Test — Qualifier + Final: qualifier #5, final #2, 7/12 combined points. qualifier #5 · 1362 words; final #2 · 1261 words.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 11, 2026Official product / access information ↗
3

Kimi K3

Moonshot AI

Kimi K3 (Moonshot) was steady and reliable: 3rd in qualifier and 3rd in final (8 points, normalized 66.7). Thinking-enabled·High setting produced substantive drafts with good pacing and clear dramatic beats, though prose and dialogue texture were less polished than Astra's.

MultipleWeighted score: 67.4
When to choose Kimi K3
Best for
Balanced scene generation when you want thoughtful internal logic and consistent pacing.
Price & access
Moonshot / provider-specific access.
Watch out for
You prioritize the smoothest line-level dialogue polish above overall structure.
Restrictions
Moderate
Evidence
Qualifier: placement 3, 688 words, Thinking enabled · High, 54s. Final: placement 3, 980 words. Solid across scenes but behind top contender on nuance and line-level choices.
Sources and observations (1)
  • Best Model Now — Screenwriting Blind TestBMN test · Observed Sep 11, 2026

    Screenwriting Blind Test — Qualifier + Final: qualifier #3, final #3, 8/12 combined points. qualifier #3 · 688 words; final #3 · 980 words.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Ranking last verified: Sep 11, 2026Official product / access information ↗

On the Horizon

New announcements we’re following. A confirmed launch is not a recommendation: independent comparisons and BMN’s category reviews still matter.

OpenAIAwaiting BMN review

GPT-6 Sol

Announced Sep 22, 2026

OpenAI has released GPT-6 Sol in ChatGPT Work, Codex, and the API. The company reports stronger coding and professional-work capabilities with lower API pricing than GPT-5.6 Sol.

What we’re checking

Independent, version-specific comparisons for writing quality, revision control, coding reliability, and practical value. Provider results do not establish a BMN category winner.

Original announcement ↗

Source checked Sep 23, 2026 · Next review target Sep 30, 2026

AnthropicAwaiting BMN review

Claude Opus 5.5

Announced Sep 22, 2026

Anthropic has released Claude Opus 5.5. The company reports stronger coding and professional work, clearer communication, and lower typical running costs than Opus 5.

What we’re checking

Independent comparisons with Fable 5.1 and current OpenAI models, plus task-specific writing and coding evidence. We still need matched settings and real workflow costs.

Original announcement ↗

Source checked Sep 23, 2026 · Next review target Sep 30, 2026

Know when the recommendations change.

Meaningful AI updates, plus On the Horizon: new releases we’re watching. About once a week. Only when there’s something worth sharing.

September 23, 2026

Qualified media recommendations are live.

All nine image and video categories now show unscored leading groups or choices by use case, with original comparisons, tested settings, a review of the eligible field, and the evidence still needed. GPT Image 2 and Riverflow Pro-high have separate records. The research interface and published methodology now share the same category-specific weights.

September 22, 2026

Image and video evidence, with its limits.

New source observations identify tested variants, task-specific findings, and conflicts with existing recommendations. They appear separately from the evidence behind the older rankings. Generic video summaries now receive the same provisional treatment as image summaries; scores and ranking-verification dates remain unchanged.

September 21, 2026

Clearer comparisons. Evidence you can inspect.

Larger text, direct category links, and expandable comparisons make the rankings easier to use. Evidence gaps are now visible, including inherited image placements awaiting version-specific confirmation. Ranking order and verification dates have not been advanced by this presentation update.

September 10, 2026

ChatGPT Images 2.5 replaces GPT Image 2.

OpenAI’s new image release adds faster generation, stronger subject preservation, more reliable multi-turn edits, and new direct-editing tools. The release initially inherited GPT Image 2’s placements. Those inherited placements are provisional until version-specific comparative evidence is documented. GPT-6 Astra’s access and API pricing have also been brought up to date.

September 8, 2026

Coding now ranks tested configurations.

Every Coding result identifies the system, underlying model or platform-selected AI, access plan, system version, and test date. Technical awards cleared the first supervised sweep; Builder awards are provisional until the full Practical Consensus threshold is met.

Evidence in.
A useful answer out.

We assess outside findings for the task and tested configuration. No single benchmark, arena, reviewer, or company determines a winner. These weights apply when comparable evidence supports scored rankings. Qualified image and video recommendations have no numerical scores and disclose unresolved comparisons. Builders use a separate profile: 20% working-product proof, 60% practitioner consensus, 10% ownership and publishing, and 10% ease and value.

Read the full methodology →
50–70%

Category performance

The capabilities that matter most for the specific award. Artistry and video use a quality-first profile.

15–20%

Real-world preference

Blind arenas, practitioner testing, and credible specialist consensus—with task fit made explicit.

5–15%

Access & workflow

Availability and usability can break a close call, but cannot rescue materially weaker output.

5–10%

Price & value

Cost matters after a model clears the quality bar; inexpensive mediocrity does not win.

5%

Restrictions

Material refusals, creative limitations, and other constraints that reduce practical usefulness.

Gate

Complete candidate pool

Every contender in the defined, dated evaluation field must be evaluated. New discoveries enter through an eligibility review.