Current AI rankings

Best AI for Images

Writing / Coding / Images / Video

Compare image generation, editing, photorealism, and artistry. Inspect the evidence and the limitations behind each placement.

Know when the recommendations change.

Meaningful AI updates, plus On the Horizon: new releases we’re watching. About once a week. Only when there’s something worth sharing.

Scored rankings combine task quality and practical factors; small score gaps are close calls. Qualified image and video recommendations show useful options, tested configurations, and unresolved questions without numerical scores. How we compare →

Qualified recommendations

Best Image Overall

The strongest all-around image model across quality, control, editing, consistency, and access.

Output qualityCreative controlEditingConsistencyAccess & value

Leading options · no settled order. The current independent generation comparisons support this leading group. The published choices have no implied internal order or invented five-factor scores.

GPT Image 2.5 Sunburst

OpenAI

Sunburst heads both reviewed generation tables. Its Arena entry is preliminary, and generation preference does not establish the best editing workflow or value.

APIProvisional: evidence still developing
Best for
High-quality general generation
Price & access
Paid API; select the named variant and quality setting.
Watch out for
You need a proven winner for editing, cost, or a complete production workflow.
Restrictions
Pending
Evidence
Arena named Sunburst and Artificial Analysis max quality both head their generation pools.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Sunburst: named leaderboard variant. Current generation leaderboard leading group. https://arena.ai/leaderboard/text-to-image Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2.5 Sunburst: named leaderboard variant. Settings: Arena public blind pairwise preference, September 21 snapshot. Limits: Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Sunburst: max quality. Current generation leaderboard leading group. https://artificialanalysis.ai/image/leaderboard/text-to-image Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2.5 Sunburst: max quality. Settings: Artificial Analysis public image preference leaderboard, September 23 snapshot. Limits: Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

GPT Image 2.5 Flare

OpenAI

Flare is near the head of both generation comparisons and merits a parallel trial. These sources do not establish that its speed or cost makes it a better overall workflow.

APIProvisional: evidence still developing
Best for
General generation alongside a Sunburst trial
Price & access
Paid API; select the named variant and quality setting.
Watch out for
You need a proven winner for editing, cost, or a complete production workflow.
Restrictions
Pending
Evidence
Arena named Flare and Artificial Analysis max quality are near the head of their generation pools.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Flare: named leaderboard variant. Current generation leaderboard leading group. https://arena.ai/leaderboard/text-to-image Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2.5 Flare: named leaderboard variant. Settings: Arena public blind pairwise preference, September 21 snapshot. Limits: Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Flare: max quality. Current generation leaderboard leading group. https://artificialanalysis.ai/image/leaderboard/text-to-image Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2.5 Flare: max quality. Settings: Artificial Analysis public image preference leaderboard, September 23 snapshot. Limits: Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

GPT Image 2

OpenAI

The older GPT Image 2 remains strongly competitive in both generation tables. Its separately documented API identity avoids confusing these results with ChatGPT Images 2.5.

APIProvisional: evidence still developing
Best for
Generation using the established GPT Image 2 API
Price & access
Paid API; select the named variant and quality setting.
Watch out for
You need a proven winner for editing, cost, or a complete production workflow.
Restrictions
Pending
Evidence
Arena medium and Artificial Analysis high quality both place GPT Image 2 in the leading generation group.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2: medium quality. Current generation leaderboard leading group. https://arena.ai/leaderboard/text-to-image Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2: medium quality. Settings: Arena public blind pairwise preference, September 21 snapshot. Limits: Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    GPT Image 2: high quality. Current generation leaderboard leading group. https://artificialanalysis.ai/image/leaderboard/text-to-image Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2: high quality. Settings: Artificial Analysis public image preference leaderboard, September 23 snapshot. Limits: Settings differ between publishers; new variants have fewer votes. Generation preference does not establish editing, cost, or workflow superiority.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: Generation preference is narrower than Best Image Overall. Editing, latency, value and production reliability are not matched across the field. Sunburst and Flare are preliminary in Arena; generic ChatGPT Images is not the tested API configuration. Riverflow Pro-high is now included; Design Arena's unspecified Pro configuration remains unmapped.

How we reviewed the field (19 contenders)

Recheck the two leaderboards and exact settings by September 30; seek matched editing, workflow and cost comparisons before naming an overall leader.

  • insufficient evidence

    ChatGPT Images 2.5: ChatGPT product routing is not the named Sunburst/Flare API configuration; do not inherit API preference results.

  • considered

    MAI-Image-2.6: Current generation and editing comparisons support a serious alternative; specialist and workflow results remain incomplete.

  • considered

    Nano Banana 2: The Arena web-search configuration is not an unspecified default; observed preference does not settle full workflow value.

  • considered

    Reve 2.1: Historical preference and cinematic examples remain useful, but the recorded September 27 creation shutdown makes a new ongoing workflow unsuitable.

  • considered

    Muse Image: Current editing comparisons provide support; consumer-product access and missing specialist comparisons limit broader claims.

  • insufficient evidence

    Midjourney V8.2: Curious Refuge directly compares V8.2 with GPT Image 2 and Reve for cinematic style. Few examples, undisclosed commercial relationship and incomplete rival coverage prevent an independent field-wide verdict; absence from Arena is not a loss.

  • considered

    Grok Imagine Image 2.0: Arena's low setting and Artificial Analysis's configuration differ; editing results disagree, so neither is transferred indiscriminately.

  • considered

    Seedream 5.0 Pro: Current generation and editing results were reviewed; they do not establish superiority in the narrower specialist tasks.

  • considered

    Nano Banana Pro: Arena's 2K variant and hands-on compositions are useful; do not equate portrait appeal with all photorealism or inherit unspecified settings.

  • considered

    FLUX.2 [max]: Current generation/editing preference results do not sustain the earlier blanket top-three claim; specialized reference control needs separate comparisons.

  • considered

    Ideogram 4.0: Current quality-setting generation results were reviewed. P-Image examples and separate object-removal products are not evidence for Ideogram 4.0.

  • insufficient evidence

    Krea 2 Large: Generation preference is available but style profiles and personalization are not matched against current rivals.

  • recommended

    GPT Image 2.5 Sunburst: Current named-variant comparisons reviewed; preliminary vote volume and task-specific uncertainty remain visible.

  • recommended

    GPT Image 2.5 Flare: Current named-variant comparisons reviewed; close results do not establish an independent speed/value advantage.

  • considered

    MAI-Image-2.6-Flash: Artificial Analysis editing support is promising; preview Flash is distinct from the base MAI-Image-2.6 configuration.

  • considered

    Nano Banana 2 Lite: Lite efficiency configuration is separately represented; no automatic quality or grounded-workflow inheritance from Nano Banana 2.

  • considered

    Qwen-Image-3.0-Pro: Current generation comparisons include this Pro version; matched specialist and workflow support remains incomplete.

  • recommended

    GPT Image 2: GPT Image 2 is a separately verified API model; Arena medium and Artificial Analysis high settings are not interchangeable.

  • insufficient evidence

    Riverflow 2.5 Pro (high): Official API verifies the Pro-high configuration. Design Arena's bare Pro label omits reasoning level; exact comparative mapping is unresolved, so its portrait/abstract results are not inherited.

Qualified recommendations

Image Editing

Reliable targeted changes without needlessly disturbing the rest of the image.

Edit precisionImage preservationReference consistencyText handlingWorkflow speed

Leading options · no settled order. Current independent editing preference supports Sunburst, Flare and MAI-Image-2.6 as a useful unscored shortlist. This is not a claim that every edit type shares one winner.

GPT Image 2.5 Sunburst

OpenAI

Sunburst leads the reviewed editing preference tables at their tested settings. This supports trying it for edits, while preservation and exact text still need task-specific checks.

APIProvisional: evidence still developing
Best for
General image edits using the named Sunburst setting
Price & access
Paid API; use the named variant and tested quality setting.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena image-edit and Artificial Analysis editing both place the named Sunburst variant at the head of their pools.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Sunburst, named public leaderboard entry. Image edits evaluated by public pairwise visual preference. https://arena.ai/leaderboard/image-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2.5 Sunburst, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Sunburst at max quality. Image edits evaluated by public pairwise visual preference. https://artificialanalysis.ai/image/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2.5 Sunburst at max quality. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

GPT Image 2.5 Flare

OpenAI

Flare is close to the front of both editing comparisons. It is a supported alternative to trial on your source images, without an inferred speed, cost or preservation advantage.

APIProvisional: evidence still developing
Best for
An alternative for iterative image editing
Price & access
Paid API; use the named variant and tested quality setting.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Both current editing leaderboards support Flare; their settings and broader preservation tests are not matched.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Flare, named public leaderboard entry. Image edits evaluated by public pairwise visual preference. https://arena.ai/leaderboard/image-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2.5 Flare, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Flare at max quality. Image edits evaluated by public pairwise visual preference. https://artificialanalysis.ai/image/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2.5 Flare at max quality. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

MAI-Image-2.6

Microsoft AI

MAI-Image-2.6 offers a well-supported alternative: Artificial Analysis places it higher than Arena does. That disagreement is useful evidence to test your actual edit type.

MultipleProvisional: evidence still developing
Best for
Comparing a non-OpenAI editor on your own edit tasks
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
MAI-Image-2.6 is third in Artificial Analysis editing and fifth in Arena's image-edit snapshot; this does not imply every task shares that order.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    MAI-Image-2.6, named public leaderboard entry. Image edits evaluated by public pairwise visual preference. https://arena.ai/leaderboard/image-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test. Original evaluation/snapshot: 2026-09-21. Tested: MAI-Image-2.6, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    MAI-Image-2.6, named public leaderboard entry. Image edits evaluated by public pairwise visual preference. https://artificialanalysis.ai/image/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test. Original evaluation/snapshot: 2026-09-23. Tested: MAI-Image-2.6, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: Arena and Artificial Analysis order MAI, Grok and GPT Image 2 differently and use different configurations. Surgical preservation, text accuracy, consistency and cost lack a common full-field test.

How we reviewed the field (19 contenders)

Compare preservation, text edits and multi-reference consistency with matched settings, and recheck current preference intervals by September 30.

  • insufficient evidence

    ChatGPT Images 2.5: ChatGPT product routing is not the named Sunburst/Flare API configuration; do not inherit API preference results.

  • recommended

    MAI-Image-2.6: Current generation and editing comparisons support a serious alternative; specialist and workflow results remain incomplete.

  • considered

    Nano Banana 2: The Arena web-search configuration is not an unspecified default; observed preference does not settle full workflow value.

  • considered

    Reve 2.1: Historical preference and cinematic examples remain useful, but the recorded September 27 creation shutdown makes a new ongoing workflow unsuitable.

  • considered

    Muse Image: Current editing comparisons provide support; consumer-product access and missing specialist comparisons limit broader claims.

  • insufficient evidence

    Midjourney V8.2: Curious Refuge directly compares V8.2 with GPT Image 2 and Reve for cinematic style. Few examples, undisclosed commercial relationship and incomplete rival coverage prevent an independent field-wide verdict; absence from Arena is not a loss.

  • considered

    Grok Imagine Image 2.0: Arena's low setting and Artificial Analysis's configuration differ; editing results disagree, so neither is transferred indiscriminately.

  • considered

    Seedream 5.0 Pro: Current generation and editing results were reviewed; they do not establish superiority in the narrower specialist tasks.

  • considered

    Nano Banana Pro: Arena's 2K variant and hands-on compositions are useful; do not equate portrait appeal with all photorealism or inherit unspecified settings.

  • considered

    FLUX.2 [max]: Current generation/editing preference results do not sustain the earlier blanket top-three claim; specialized reference control needs separate comparisons.

  • considered

    Ideogram 4.0: Current quality-setting generation results were reviewed. P-Image examples and separate object-removal products are not evidence for Ideogram 4.0.

  • insufficient evidence

    Krea 2 Large: Generation preference is available but style profiles and personalization are not matched against current rivals.

  • recommended

    GPT Image 2.5 Sunburst: Current named-variant comparisons reviewed; preliminary vote volume and task-specific uncertainty remain visible.

  • recommended

    GPT Image 2.5 Flare: Current named-variant comparisons reviewed; close results do not establish an independent speed/value advantage.

  • considered

    MAI-Image-2.6-Flash: Artificial Analysis editing support is promising; preview Flash is distinct from the base MAI-Image-2.6 configuration.

  • considered

    Nano Banana 2 Lite: Lite efficiency configuration is separately represented; no automatic quality or grounded-workflow inheritance from Nano Banana 2.

  • considered

    Qwen-Image-3.0-Pro: Current generation comparisons include this Pro version; matched specialist and workflow support remains incomplete.

  • considered

    GPT Image 2: GPT Image 2 is a separately verified API model; Arena medium and Artificial Analysis high settings are not interchangeable.

  • insufficient evidence

    Riverflow 2.5 Pro (high): Official API verifies the Pro-high configuration. Design Arena's bare Pro label omits reasoning level; exact comparative mapping is unresolved, so its portrait/abstract results are not inherited.

Qualified recommendations

Photorealism

Images that hold up as believable photographs rather than merely looking impressive at a glance.

Faces & anatomyLight & materialsCamera credibilityFine detailArtifact rate

Leading options · no settled order. Arena's Photorealistic filter and Design Arena's Portrait view support these as current candidates. The shortlist reflects photo-oriented preference rather than a universal realism winner.

GPT Image 2.5 Sunburst

OpenAI

Sunburst is near the head of Arena's Photorealistic filter, with uncertainty overlapping Flare. Design Arena's Portrait view offers narrower corroboration.

APIProvisional: evidence still developing
Best for
Photo-oriented scenes; inspect anatomy and materials
Price & access
Paid API; use the named variant and tested quality setting.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena Photorealistic: Sunburst and Flare share rank spread 1–2. Design Arena Portrait also includes Sunburst near the front.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Sunburst, named public leaderboard entry. Photorealistic image preference (Arena) and portrait preference (Design Arena). https://arena.ai/leaderboard/text-to-image/photorealistic Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2.5 Sunburst, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification.

  • Design Arena — media comparisonsblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Sunburst, named public leaderboard entry. Photorealistic image preference (Arena) and portrait preference (Design Arena). https://www.designarena.ai/leaderboard/image Design Arena Portrait tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limitation: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2.5 Sunburst, named public leaderboard entry. Settings: Design Arena Portrait tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limits: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

GPT Image 2.5 Flare

OpenAI

Flare is statistically close to Sunburst in Arena's Photorealistic filter and near the front of Design Arena's Portrait view. Neither result establishes superior realism across all subjects.

APIProvisional: evidence still developing
Best for
Photo-oriented generation alongside Sunburst
Price & access
Paid API; use the named variant and tested quality setting.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena Photorealistic places Flare close to Sunburst; Design Arena Portrait is supplementary, not a full realism test.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Flare, named public leaderboard entry. Photorealistic image preference (Arena) and portrait preference (Design Arena). https://arena.ai/leaderboard/text-to-image/photorealistic Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2.5 Flare, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification.

  • Design Arena — media comparisonsblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Flare, named public leaderboard entry. Photorealistic image preference (Arena) and portrait preference (Design Arena). https://www.designarena.ai/leaderboard/image Design Arena Portrait tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limitation: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2.5 Flare, named public leaderboard entry. Settings: Design Arena Portrait tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limits: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

GPT Image 2

OpenAI

GPT Image 2 remains strong in Arena's Photorealistic filter and heads Design Arena's narrower Portrait view. The differing tasks favor a use-case trial over a universal winner.

APIProvisional: evidence still developing
Best for
Portrait-focused generation with the GPT Image 2 API
Price & access
Paid API; use the named variant and tested quality setting.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena medium is third in the Photorealistic snapshot; Design Arena's Portrait view favors GPT Image 2 under less transparent settings.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2 at medium quality; not ChatGPT Images 2.5. Photorealistic image preference (Arena) and portrait preference (Design Arena). https://arena.ai/leaderboard/text-to-image/photorealistic Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2 at medium quality; not ChatGPT Images 2.5. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification.

  • Design Arena — media comparisonsblind preference · Observed Sep 23, 2026

    GPT Image 2, named public leaderboard entry. Photorealistic image preference (Arena) and portrait preference (Design Arena). https://www.designarena.ai/leaderboard/image Design Arena Portrait tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limitation: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2, named public leaderboard entry. Settings: Design Arena Portrait tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limits: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow 2.5 Pro needs exact public-identity verification.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: Arena's Sunburst and Flare intervals overlap. Design Arena is portrait-specific corroboration with a low inclusion threshold and incomplete settings/sample transparency. Portrait appeal does not settle anatomy, materials or camera plausibility across genres. Riverflow Pro-high is verified, but Design Arena's unspecified Pro setting cannot be inherited.

How we reviewed the field (19 contenders)

Resolve Design Arena's Pro reasoning setting against verified Riverflow Pro-high and obtain non-portrait realism comparisons; recheck photo-filter intervals by September 30.

  • insufficient evidence

    ChatGPT Images 2.5: ChatGPT product routing is not the named Sunburst/Flare API configuration; do not inherit API preference results.

  • considered

    MAI-Image-2.6: Current generation and editing comparisons support a serious alternative; specialist and workflow results remain incomplete.

  • considered

    Nano Banana 2: The Arena web-search configuration is not an unspecified default; observed preference does not settle full workflow value.

  • considered

    Reve 2.1: Historical preference and cinematic examples remain useful, but the recorded September 27 creation shutdown makes a new ongoing workflow unsuitable.

  • considered

    Muse Image: Current editing comparisons provide support; consumer-product access and missing specialist comparisons limit broader claims.

  • insufficient evidence

    Midjourney V8.2: Curious Refuge directly compares V8.2 with GPT Image 2 and Reve for cinematic style. Few examples, undisclosed commercial relationship and incomplete rival coverage prevent an independent field-wide verdict; absence from Arena is not a loss.

  • considered

    Grok Imagine Image 2.0: Arena's low setting and Artificial Analysis's configuration differ; editing results disagree, so neither is transferred indiscriminately.

  • considered

    Seedream 5.0 Pro: Current generation and editing results were reviewed; they do not establish superiority in the narrower specialist tasks.

  • considered

    Nano Banana Pro: Arena's 2K variant and hands-on compositions are useful; do not equate portrait appeal with all photorealism or inherit unspecified settings.

  • considered

    FLUX.2 [max]: Current generation/editing preference results do not sustain the earlier blanket top-three claim; specialized reference control needs separate comparisons.

  • considered

    Ideogram 4.0: Current quality-setting generation results were reviewed. P-Image examples and separate object-removal products are not evidence for Ideogram 4.0.

  • insufficient evidence

    Krea 2 Large: Generation preference is available but style profiles and personalization are not matched against current rivals.

  • recommended

    GPT Image 2.5 Sunburst: Current named-variant comparisons reviewed; preliminary vote volume and task-specific uncertainty remain visible.

  • recommended

    GPT Image 2.5 Flare: Current named-variant comparisons reviewed; close results do not establish an independent speed/value advantage.

  • considered

    MAI-Image-2.6-Flash: Artificial Analysis editing support is promising; preview Flash is distinct from the base MAI-Image-2.6 configuration.

  • considered

    Nano Banana 2 Lite: Lite efficiency configuration is separately represented; no automatic quality or grounded-workflow inheritance from Nano Banana 2.

  • considered

    Qwen-Image-3.0-Pro: Current generation comparisons include this Pro version; matched specialist and workflow support remains incomplete.

  • recommended

    GPT Image 2: GPT Image 2 is a separately verified API model; Arena medium and Artificial Analysis high settings are not interchangeable.

  • insufficient evidence

    Riverflow 2.5 Pro (high): Official API verifies the Pro-high configuration. Design Arena's bare Pro label omits reasoning level; exact comparative mapping is unresolved, so its portrait/abstract results are not inherited.

Qualified recommendations

Artistry

Distinctive visual taste, composition, style control, and expressive range.

CompositionVisual tasteStylistic rangeOriginalityCreative control

Leading options · no settled order. Arena's Art filter supports a close leading group among these tested variants; Design Arena's Abstract view adds limited corroboration. These are unranked starting points for prompt-driven art.

GPT Image 2.5 Sunburst

OpenAI

Sunburst belongs to Arena's close Art-filter group and leads Design Arena's Abstract view. This supports prompt-driven artwork, not a universal judgment of taste or personalization.

APIProvisional: evidence still developing
Best for
Prompt-driven art and abstract concepts
Price & access
Paid API; use the named variant and tested quality setting.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena Art gives Sunburst rank spread 1–3; Design Arena Abstract provides limited corroboration.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Sunburst, named public leaderboard entry. Art-filter preference (Arena) and Abstract-view preference (Design Arena). https://arena.ai/leaderboard/text-to-image/art Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2.5 Sunburst, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender.

  • Design Arena — media comparisonsblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Sunburst, named public leaderboard entry. Art-filter preference (Arena) and Abstract-view preference (Design Arena). https://www.designarena.ai/leaderboard/image Design Arena Abstract tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limitation: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2.5 Sunburst, named public leaderboard entry. Settings: Design Arena Abstract tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limits: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

GPT Image 2.5 Flare

OpenAI

Flare sits in the same overlapping Art-filter group as Sunburst and GPT Image 2. Design Arena's Abstract view also rates it strongly, without resolving stylistic range.

APIProvisional: evidence still developing
Best for
Exploring art prompts alongside Sunburst
Price & access
Paid API; use the named variant and tested quality setting.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena Art intervals overlap for the three selected variants; Design Arena Abstract places Flare near the front.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Flare, named public leaderboard entry. Art-filter preference (Arena) and Abstract-view preference (Design Arena). https://arena.ai/leaderboard/text-to-image/art Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2.5 Flare, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender.

  • Design Arena — media comparisonsblind preference · Observed Sep 23, 2026

    GPT Image 2.5 Flare, named public leaderboard entry. Art-filter preference (Arena) and Abstract-view preference (Design Arena). https://www.designarena.ai/leaderboard/image Design Arena Abstract tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limitation: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2.5 Flare, named public leaderboard entry. Settings: Design Arena Abstract tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limits: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

GPT Image 2

OpenAI

GPT Image 2 remains in Arena's overlapping Art-filter group and performs strongly in Design Arena's Abstract view. Midjourney's cinematic-style case remains an unresolved alternative.

APIProvisional: evidence still developing
Best for
Controlled art prompts using the established image API
Price & access
Paid API; use the named variant and tested quality setting.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena medium has rank spread 1–3 in Art; Design Arena Abstract includes GPT Image 2 near the front, with Riverflow Pro ahead under unspecified settings.
Sources and observations (2)
  • Arena image leaderboards — current routesblind preference · Observed Sep 23, 2026

    GPT Image 2 at medium quality; not ChatGPT Images 2.5. Art-filter preference (Arena) and Abstract-view preference (Design Arena). https://arena.ai/leaderboard/text-to-image/art Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender. Original evaluation/snapshot: 2026-09-21. Tested: GPT Image 2 at medium quality; not ChatGPT Images 2.5. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender.

  • Design Arena — media comparisonsblind preference · Observed Sep 23, 2026

    GPT Image 2, named public leaderboard entry. Art-filter preference (Arena) and Abstract-view preference (Design Arena). https://www.designarena.ai/leaderboard/image Design Arena Abstract tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limitation: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender. Original evaluation/snapshot: 2026-09-23. Tested: GPT Image 2, named public leaderboard entry. Settings: Design Arena Abstract tab; provider defaults, exact settings and usable sample counts not fully disclosed. Limits: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow 2.5 Pro is an unresolved external contender.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: All three Arena Art rank intervals overlap. Abstract preference covers only part of artistry. Midjourney V8.2 remains a credible cinematic-style alternative, but its few specialist examples do not establish a broad independent comparison. Personalization and stylistic range remain under-tested; Riverflow Pro-high is registered, but its mapping to Design Arena's unspecified Pro remains unresolved.

How we reviewed the field (19 contenders)

Seek independent V8.2 versus current API-variant comparisons across styles and controlled personalization; resolve Riverflow's tested reasoning setting and recheck Art results by September 30.

  • insufficient evidence

    ChatGPT Images 2.5: ChatGPT product routing is not the named Sunburst/Flare API configuration; do not inherit API preference results.

  • considered

    MAI-Image-2.6: Current generation and editing comparisons support a serious alternative; specialist and workflow results remain incomplete.

  • considered

    Nano Banana 2: The Arena web-search configuration is not an unspecified default; observed preference does not settle full workflow value.

  • considered

    Reve 2.1: Historical preference and cinematic examples remain useful, but the recorded September 27 creation shutdown makes a new ongoing workflow unsuitable.

  • considered

    Muse Image: Current editing comparisons provide support; consumer-product access and missing specialist comparisons limit broader claims.

  • insufficient evidence

    Midjourney V8.2: Curious Refuge directly compares V8.2 with GPT Image 2 and Reve for cinematic style. Few examples, undisclosed commercial relationship and incomplete rival coverage prevent an independent field-wide verdict; absence from Arena is not a loss.

  • considered

    Grok Imagine Image 2.0: Arena's low setting and Artificial Analysis's configuration differ; editing results disagree, so neither is transferred indiscriminately.

  • considered

    Seedream 5.0 Pro: Current generation and editing results were reviewed; they do not establish superiority in the narrower specialist tasks.

  • considered

    Nano Banana Pro: Arena's 2K variant and hands-on compositions are useful; do not equate portrait appeal with all photorealism or inherit unspecified settings.

  • considered

    FLUX.2 [max]: Current generation/editing preference results do not sustain the earlier blanket top-three claim; specialized reference control needs separate comparisons.

  • considered

    Ideogram 4.0: Current quality-setting generation results were reviewed. P-Image examples and separate object-removal products are not evidence for Ideogram 4.0.

  • insufficient evidence

    Krea 2 Large: Generation preference is available but style profiles and personalization are not matched against current rivals.

  • recommended

    GPT Image 2.5 Sunburst: Current named-variant comparisons reviewed; preliminary vote volume and task-specific uncertainty remain visible.

  • recommended

    GPT Image 2.5 Flare: Current named-variant comparisons reviewed; close results do not establish an independent speed/value advantage.

  • considered

    MAI-Image-2.6-Flash: Artificial Analysis editing support is promising; preview Flash is distinct from the base MAI-Image-2.6 configuration.

  • considered

    Nano Banana 2 Lite: Lite efficiency configuration is separately represented; no automatic quality or grounded-workflow inheritance from Nano Banana 2.

  • considered

    Qwen-Image-3.0-Pro: Current generation comparisons include this Pro version; matched specialist and workflow support remains incomplete.

  • recommended

    GPT Image 2: GPT Image 2 is a separately verified API model; Arena medium and Artificial Analysis high settings are not interchangeable.

  • insufficient evidence

    Riverflow 2.5 Pro (high): Official API verifies the Pro-high configuration. Design Arena's bare Pro label omits reasoning level; exact comparative mapping is unresolved, so its portrait/abstract results are not inherited.

On the Horizon

New announcements we’re following. A confirmed launch is not a recommendation: independent comparisons and BMN’s category reviews still matter.

OpenAIAwaiting BMN review

GPT-6 Sol

Announced Sep 22, 2026

OpenAI has released GPT-6 Sol in ChatGPT Work, Codex, and the API. The company reports stronger coding and professional-work capabilities with lower API pricing than GPT-5.6 Sol.

What we’re checking

Independent, version-specific comparisons for writing quality, revision control, coding reliability, and practical value. Provider results do not establish a BMN category winner.

Original announcement ↗

Source checked Sep 23, 2026 · Next review target Sep 30, 2026

AnthropicAwaiting BMN review

Claude Opus 5.5

Announced Sep 22, 2026

Anthropic has released Claude Opus 5.5. The company reports stronger coding and professional work, clearer communication, and lower typical running costs than Opus 5.

What we’re checking

Independent comparisons with Fable 5.1 and current OpenAI models, plus task-specific writing and coding evidence. We still need matched settings and real workflow costs.

Original announcement ↗

Source checked Sep 23, 2026 · Next review target Sep 30, 2026

Know when the recommendations change.

Meaningful AI updates, plus On the Horizon: new releases we’re watching. About once a week. Only when there’s something worth sharing.

September 23, 2026

Qualified media recommendations are live.

All nine image and video categories now show unscored leading groups or choices by use case, with original comparisons, tested settings, a review of the eligible field, and the evidence still needed. GPT Image 2 and Riverflow Pro-high have separate records. The research interface and published methodology now share the same category-specific weights.

September 22, 2026

Image and video evidence, with its limits.

New source observations identify tested variants, task-specific findings, and conflicts with existing recommendations. They appear separately from the evidence behind the older rankings. Generic video summaries now receive the same provisional treatment as image summaries; scores and ranking-verification dates remain unchanged.

September 21, 2026

Clearer comparisons. Evidence you can inspect.

Larger text, direct category links, and expandable comparisons make the rankings easier to use. Evidence gaps are now visible, including inherited image placements awaiting version-specific confirmation. Ranking order and verification dates have not been advanced by this presentation update.

September 10, 2026

ChatGPT Images 2.5 replaces GPT Image 2.

OpenAI’s new image release adds faster generation, stronger subject preservation, more reliable multi-turn edits, and new direct-editing tools. The release initially inherited GPT Image 2’s placements. Those inherited placements are provisional until version-specific comparative evidence is documented. GPT-6 Astra’s access and API pricing have also been brought up to date.

September 8, 2026

Coding now ranks tested configurations.

Every Coding result identifies the system, underlying model or platform-selected AI, access plan, system version, and test date. Technical awards cleared the first supervised sweep; Builder awards are provisional until the full Practical Consensus threshold is met.

Evidence in.
A useful answer out.

We assess outside findings for the task and tested configuration. No single benchmark, arena, reviewer, or company determines a winner. These weights apply when comparable evidence supports scored rankings. Qualified image and video recommendations have no numerical scores and disclose unresolved comparisons. Builders use a separate profile: 20% working-product proof, 60% practitioner consensus, 10% ownership and publishing, and 10% ease and value.

Read the full methodology →
50–70%

Category performance

The capabilities that matter most for the specific award. Artistry and video use a quality-first profile.

15–20%

Real-world preference

Blind arenas, practitioner testing, and credible specialist consensus—with task fit made explicit.

5–15%

Access & workflow

Availability and usability can break a close call, but cannot rescue materially weaker output.

5–10%

Price & value

Cost matters after a model clears the quality bar; inexpensive mediocrity does not win.

5%

Restrictions

Material refusals, creative limitations, and other constraints that reduce practical usefulness.

Gate

Complete candidate pool

Every contender in the defined, dated evaluation field must be evaluated. New discoveries enter through an eligibility review.