Current AI rankings

Best AI for Video

Writing / Coding / Images / Video

Compare AI video tools for generation, image animation, editing, and production capabilities, with access details and evidence.

Know when the recommendations change.

Meaningful AI updates, plus On the Horizon: new releases we’re watching. About once a week. Only when there’s something worth sharing.

Scored rankings combine task quality and practical factors; small score gaps are close calls. Qualified image and video recommendations show useful options, tested configurations, and unresolved questions without numerical scores. How we compare →

Qualified recommendations

Best Video Overall

The best practical blend of motion quality, control, consistency, access, and workflow.

Motion qualityCreative controlVisual consistencyOutput qualityAccess & workflow

Recommendations by use case. These models have useful generation and editing evidence, but no single matched full-production trial establishes an overall winner. This is a production shortlist grounded in narrower comparative tasks.

Wan 3.0

Alibaba

Wan 3.0 has cross-source generation and editing support. Specialist examples suggest reference-heavy scene potential but also cinematic and lip-sync weaknesses; those examples remain supplementary.

MultipleProvisional: evidence still developing
Best for
Generation plus footage transformations
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena and Artificial Analysis support exact Wan 3.0 generation/editing configurations; this is not a complete production-quality or feature-breadth score.
Sources and observations (3)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Wan 3.0, named public leaderboard entry. Existing-footage editing preference used as one bounded production-workflow signal. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched. Original evaluation/snapshot: 2026-09-21. Tested: Wan 3.0, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    Wan 3.0, named public leaderboard entry. Existing-footage editing preference used as one bounded production-workflow signal. https://artificialanalysis.ai/video/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched. Original evaluation/snapshot: 2026-09-23. Tested: Wan 3.0, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    Wan 3.0 named text-to-video entry; keep audio pools separate. Text-to-video preference as one production-workflow signal. https://artificialanalysis.ai/video/leaderboard/text-to-video Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth. Original evaluation/snapshot: 2026-09-23. Tested: Wan 3.0 named text-to-video entry; keep audio pools separate. Settings: Artificial Analysis September 23 audio/no-audio pools considered separately. Limits: Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

MiniMax H3

MiniMax

Hosted base H3 combines strong source-image animation with independently corroborated editing support. Treat its local version and fal's Max configuration as separate workflow choices.

MultipleProvisional: evidence still developing
Best for
Reference-image animation plus footage revisions
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena image-to-video and editing plus Artificial Analysis editing support base H3; none of these results transfers to Max or local H3.
Sources and observations (3)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    MiniMax H3, named public leaderboard entry. Existing-footage editing preference used as one bounded production-workflow signal. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched. Original evaluation/snapshot: 2026-09-21. Tested: MiniMax H3, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    MiniMax H3, named public leaderboard entry. Existing-footage editing preference used as one bounded production-workflow signal. https://artificialanalysis.ai/video/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched. Original evaluation/snapshot: 2026-09-23. Tested: MiniMax H3, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched.

  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    MiniMax H3 base hosted entry; not fal Max or local H3. Image-to-video preference as one production-workflow signal. https://arena.ai/leaderboard/image-to-video Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth. Original evaluation/snapshot: 2026-09-21. Tested: MiniMax H3 base hosted entry; not fal Max or local H3. Settings: Arena September 21 snapshot and named configuration. Limits: Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Dreamina Seedance 2.5

ByteDance Seed

Seedance 2.5 merits a filmmaking trial: Arena supports its image animation and editing, while specialist cinematic examples add narrower context. The 1080p Dreamina workflow is not the 720p leaderboard setting.

MultipleProvisional: evidence still developing
Best for
Filmmaking trials combining animation and scene revision
Price & access
Dreamina and selected partners; access and settings vary by region.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena supports exact Seedance 2.5 image-to-video and editing. Independent cross-publisher support for its complete filmmaking workflow remains missing.
Sources and observations (2)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Dreamina Seedance 2.5, named public leaderboard entry. Existing-footage editing preference used as one bounded production-workflow signal. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched. Original evaluation/snapshot: 2026-09-21. Tested: Dreamina Seedance 2.5, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched.

  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Dreamina Seedance 2.5 at 720p in Arena. Image-to-video preference as one production-workflow signal. https://arena.ai/leaderboard/image-to-video Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth. Original evaluation/snapshot: 2026-09-21. Tested: Dreamina Seedance 2.5 at 720p in Arena. Settings: Arena September 21 snapshot and named configuration. Limits: Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: Editing preference is one component of production usefulness, not a complete filmmaking score. Seedance 2.5's promising specialist examples are supplementary: rerolls, settings and commercial relationships are incomplete. Costs, acting, sound, continuity and failure recovery remain unmatched.

How we reviewed the field (15 contenders)

Seek finished-workflow comparisons with disclosed rerolls, resolution, audio, cost per usable shot and commercial relationships; review by September 30.

  • considered

    Gemini Omni Flash: Only explicit Omni 1.1 results support this record. Unversioned older Omni Flash editing and Artificial Analysis results were excluded.

  • recommended

    Wan 3.0: Current generation and editing preference results are strong; cinematic performance and lip-sync findings vary by task.

  • recommended

    MiniMax H3: Base H3 has current generation and editing support. Hosted and local open-weight configurations differ; Max results are not inherited.

  • considered

    MiniMax H3 Max (fal): Artificial Analysis supports the exact fal H3 Max generation configuration; base-H3 editing evidence cannot validate Max.

  • considered

    Dreamina Seedance 2.0: Current 720p generation results and narrower lip-sync examples were considered separately from Seedance 2.5.

  • recommended

    Dreamina Seedance 2.5: Arena generation and editing plus specialist filmmaking examples support task-specific interest; 720p leaderboard and 1080p Dreamina results remain distinct.

  • insufficient evidence

    Kling 3.0: Base, Pro/Standard and Omni identities are not interchangeable. Current exact-configuration production/editing comparisons remain incomplete.

  • considered

    Veo 3.1: Generation comparisons exist, but Fast/Quality and resolution distinctions prevent a blanket editing or production claim.

  • considered

    FLUX 3 Video: Preview generation has comparative support. Separate video editing, upscaled output and announced controls do not prove reliable production breadth.

  • insufficient evidence

    Runway Gen-4.5: Generation evidence is available; Aleph editing results belong to another model and cannot establish Gen-4.5 editing strength.

  • insufficient evidence

    Luma Ray 3.2: Exact 3.2 support remains limited; Ray 3 or 3.14 results cannot be inherited by this record.

  • insufficient evidence

    HappyHorse 1.1: Exact 1.1 generation support does not inherit HappyHorse 1.0 editing comparisons.

  • insufficient evidence

    Vidu Q3 Pro: Pro endpoint settings and current matched production comparisons remain insufficiently documented.

  • insufficient evidence

    LTX-2.5: Current 2.5 capabilities are documented, but older 2.3 tests do not establish quality; beta editing and export claims require direct comparative evidence.

  • considered

    Grok Imagine Video 1.5: Base Grok 1.5 generation is distinct from Arena's agent configuration; agent placement was not inherited.

Qualified recommendations

Text-to-Video

Creates coherent, compelling footage directly from a written description.

Motion coherenceScene interpretationCamera controlVisual qualityArtifact rate

Leading options · no settled order. Exact Omni 1.1 in Arena and Wan 3.0/H3 Max in Artificial Analysis support a qualified shortlist. Different fields, audio pools and overlapping intervals prevent a shared total order.

Gemini Omni Flash

Google

Explicit Omni 1.1 Flash sits near the head of Arena's text-to-video pool, with overlapping uncertainty among leading entries. Older unversioned Omni results were excluded.

MultipleProvisional: evidence still developing
Best for
Text-led generation with explicit Omni 1.1
Price & access
Use the explicit Omni 1.1 configuration; product routing may differ.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena September 21 supports the exact 1.1 Flash entry; Artificial Analysis's older unversioned Omni is not corroboration.
Sources and observations (1)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Gemini Omni 1.1 Flash, explicit 1.1 entry only. Text-to-video public preference; retain audio/no-audio and resolution distinctions. https://arena.ai/leaderboard/text-to-video Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial. Original evaluation/snapshot: 2026-09-21. Tested: Gemini Omni 1.1 Flash, explicit 1.1 entry only. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Wan 3.0

Alibaba

Wan 3.0 performs strongly in Artificial Analysis's separate audio and no-audio pools and remains competitive in Arena. Pool differences prevent a single combined rank.

MultipleProvisional: evidence still developing
Best for
Text-led generation with the required audio mode
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Artificial Analysis favors Wan in the no-audio pool and places it near the head with audio; Arena includes Wan 3.0 separately.
Sources and observations (2)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Wan 3.0, named public leaderboard entry. Text-to-video public preference; retain audio/no-audio and resolution distinctions. https://arena.ai/leaderboard/text-to-video Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial. Original evaluation/snapshot: 2026-09-21. Tested: Wan 3.0, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    Wan 3.0, named public leaderboard entry. Text-to-video public preference; retain audio/no-audio and resolution distinctions. https://artificialanalysis.ai/video/leaderboard/text-to-video Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial. Original evaluation/snapshot: 2026-09-23. Tested: Wan 3.0, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

MiniMax H3 Max (fal)

fal / MiniMax

fal's H3 Max is a strong candidate in Artificial Analysis's with-audio text-to-video pool. It is a distinct configuration, so base-H3 observations cannot be copied to it.

APIProvisional: evidence still developing
Best for
Hosted text-to-video through fal's H3 Max endpoint
Price & access
Paid fal API; choose H3 Max, not base H3.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Artificial Analysis's with-audio pool places H3 Max near the head; no exact H3 Max counterpart was established in Arena.
Sources and observations (1)
  • Artificial Analysisblind preference · Observed Sep 23, 2026

    MiniMax H3 Max (fal), distinct post-trained hosted configuration. Text-to-video public preference; retain audio/no-audio and resolution distinctions. https://artificialanalysis.ai/video/leaderboard/text-to-video Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial. Original evaluation/snapshot: 2026-09-23. Tested: MiniMax H3 Max (fal), distinct post-trained hosted configuration. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: Arena's top intervals overlap and include different configurations from Artificial Analysis. Artificial Analysis's unversioned Omni is not evidence for Omni 1.1. H3 Max is fal's distinct configuration. This shortlist does not imply the omitted close contenders lost a matched trial.

How we reviewed the field (15 contenders)

Obtain same-version, same-resolution, same-audio comparisons for Omni 1.1, Wan 3.0, H3 Max, FLUX 3 and Seedance 2.5; recheck by September 30.

  • recommended

    Gemini Omni Flash: Only explicit Omni 1.1 results support this record. Unversioned older Omni Flash editing and Artificial Analysis results were excluded.

  • recommended

    Wan 3.0: Current generation and editing preference results are strong; cinematic performance and lip-sync findings vary by task.

  • considered

    MiniMax H3: Base H3 has current generation and editing support. Hosted and local open-weight configurations differ; Max results are not inherited.

  • recommended

    MiniMax H3 Max (fal): Artificial Analysis supports the exact fal H3 Max generation configuration; base-H3 editing evidence cannot validate Max.

  • considered

    Dreamina Seedance 2.0: Current 720p generation results and narrower lip-sync examples were considered separately from Seedance 2.5.

  • considered

    Dreamina Seedance 2.5: Arena generation and editing plus specialist filmmaking examples support task-specific interest; 720p leaderboard and 1080p Dreamina results remain distinct.

  • insufficient evidence

    Kling 3.0: Base, Pro/Standard and Omni identities are not interchangeable. Current exact-configuration production/editing comparisons remain incomplete.

  • considered

    Veo 3.1: Generation comparisons exist, but Fast/Quality and resolution distinctions prevent a blanket editing or production claim.

  • considered

    FLUX 3 Video: Preview generation has comparative support. Separate video editing, upscaled output and announced controls do not prove reliable production breadth.

  • insufficient evidence

    Runway Gen-4.5: Generation evidence is available; Aleph editing results belong to another model and cannot establish Gen-4.5 editing strength.

  • insufficient evidence

    Luma Ray 3.2: Exact 3.2 support remains limited; Ray 3 or 3.14 results cannot be inherited by this record.

  • insufficient evidence

    HappyHorse 1.1: Exact 1.1 generation support does not inherit HappyHorse 1.0 editing comparisons.

  • insufficient evidence

    Vidu Q3 Pro: Pro endpoint settings and current matched production comparisons remain insufficiently documented.

  • insufficient evidence

    LTX-2.5: Current 2.5 capabilities are documented, but older 2.3 tests do not establish quality; beta editing and export claims require direct comparative evidence.

  • considered

    Grok Imagine Video 1.5: Base Grok 1.5 generation is distinct from Arena's agent configuration; agent placement was not inherited.

Qualified recommendations

Image-to-Video

Animates a reference image while preserving subject identity and visual intent.

Source preservationNatural motionCamera controlTemporal stabilityArtifact rate

Leading options · no settled order. Arena supports base H3 and exact Omni 1.1; Artificial Analysis's audio pool supports fal H3 Max and base H3. Their different configurations support a leading group, not an interchangeable family-level score.

MiniMax H3

MiniMax

Base H3 heads Arena's image-to-video snapshot and is strong in Artificial Analysis's audio pool. Try this hosted configuration separately from fal Max and local H3.

MultipleProvisional: evidence still developing
Best for
Animating a source image with hosted base H3
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena favors base H3; Artificial Analysis also supports it, while audio/no-audio settings change the order.
Sources and observations (2)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    MiniMax H3, named public leaderboard entry. Image-to-video public preference with reference-image input. https://arena.ai/leaderboard/image-to-video Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded. Original evaluation/snapshot: 2026-09-21. Tested: MiniMax H3, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    MiniMax H3, named public leaderboard entry. Image-to-video public preference with reference-image input. https://artificialanalysis.ai/video/leaderboard/image-to-video Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded. Original evaluation/snapshot: 2026-09-23. Tested: MiniMax H3, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

MiniMax H3 Max (fal)

fal / MiniMax

H3 Max heads Artificial Analysis's with-audio image-to-video pool. This supports the fal configuration specifically, without inheriting base-H3 editing or local-workflow claims.

APIProvisional: evidence still developing
Best for
Reference-image animation with audio through fal
Price & access
Paid fal API; choose H3 Max, not base H3.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Artificial Analysis supports the exact fal H3 Max entry in its with-audio image-to-video pool.
Sources and observations (1)
  • Artificial Analysisblind preference · Observed Sep 23, 2026

    MiniMax H3 Max (fal), distinct post-trained hosted configuration. Image-to-video public preference with reference-image input. https://artificialanalysis.ai/video/leaderboard/image-to-video Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded. Original evaluation/snapshot: 2026-09-23. Tested: MiniMax H3 Max (fal), distinct post-trained hosted configuration. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Gemini Omni Flash

Google

Exact Omni 1.1 is close to the front of Arena's image-to-video snapshot. Broad uncertainty and differing audio pools leave Wan and Seedance as serious alternatives.

MultipleProvisional: evidence still developing
Best for
Reference-image animation with explicit Omni 1.1
Price & access
Use the explicit Omni 1.1 configuration; product routing may differ.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena's explicit 1.1 entry supports this choice; unversioned older Omni results elsewhere were excluded.
Sources and observations (1)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Gemini Omni 1.1 Flash, explicit 1.1 entry only. Image-to-video public preference with reference-image input. https://arena.ai/leaderboard/image-to-video Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded. Original evaluation/snapshot: 2026-09-21. Tested: Gemini Omni 1.1 Flash, explicit 1.1 entry only. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: Audio and no-audio pools give different orders. Wan 3.0 is a serious alternative, especially in Artificial Analysis's no-audio pool. Base H3, fal Max and local H3 are distinct; unversioned Omni results were excluded.

How we reviewed the field (15 contenders)

Compare source preservation and motion using matched audio/resolution settings across H3, H3 Max, Omni 1.1, Wan 3.0 and Seedance 2.5 by September 30.

  • recommended

    Gemini Omni Flash: Only explicit Omni 1.1 results support this record. Unversioned older Omni Flash editing and Artificial Analysis results were excluded.

  • considered

    Wan 3.0: Current generation and editing preference results are strong; cinematic performance and lip-sync findings vary by task.

  • recommended

    MiniMax H3: Base H3 has current generation and editing support. Hosted and local open-weight configurations differ; Max results are not inherited.

  • recommended

    MiniMax H3 Max (fal): Artificial Analysis supports the exact fal H3 Max generation configuration; base-H3 editing evidence cannot validate Max.

  • considered

    Dreamina Seedance 2.0: Current 720p generation results and narrower lip-sync examples were considered separately from Seedance 2.5.

  • considered

    Dreamina Seedance 2.5: Arena generation and editing plus specialist filmmaking examples support task-specific interest; 720p leaderboard and 1080p Dreamina results remain distinct.

  • insufficient evidence

    Kling 3.0: Base, Pro/Standard and Omni identities are not interchangeable. Current exact-configuration production/editing comparisons remain incomplete.

  • considered

    Veo 3.1: Generation comparisons exist, but Fast/Quality and resolution distinctions prevent a blanket editing or production claim.

  • considered

    FLUX 3 Video: Preview generation has comparative support. Separate video editing, upscaled output and announced controls do not prove reliable production breadth.

  • insufficient evidence

    Runway Gen-4.5: Generation evidence is available; Aleph editing results belong to another model and cannot establish Gen-4.5 editing strength.

  • insufficient evidence

    Luma Ray 3.2: Exact 3.2 support remains limited; Ray 3 or 3.14 results cannot be inherited by this record.

  • insufficient evidence

    HappyHorse 1.1: Exact 1.1 generation support does not inherit HappyHorse 1.0 editing comparisons.

  • insufficient evidence

    Vidu Q3 Pro: Pro endpoint settings and current matched production comparisons remain insufficiently documented.

  • insufficient evidence

    LTX-2.5: Current 2.5 capabilities are documented, but older 2.3 tests do not establish quality; beta editing and export claims require direct comparative evidence.

  • considered

    Grok Imagine Video 1.5: Base Grok 1.5 generation is distinct from Arena's agent configuration; agent placement was not inherited.

Qualified recommendations

Video-to-Video Editing

Transforms or edits footage while maintaining continuity and directability.

Edit precisionContinuityReference retentionTemporal stabilityWorkflow control

Leading options · no settled order. Arena's Wan 3.0, Seedance 2.5 and base H3 editing results are close; Artificial Analysis separately corroborates Wan and H3. Publish the group without converting leaderboard order into a definitive winner.

Wan 3.0

Alibaba

Wan 3.0 has corroborated editing preference support in Arena and Artificial Analysis. Arena's overlapping intervals do not establish a decisive win over Seedance 2.5 or H3.

MultipleProvisional: evidence still developing
Best for
General footage transformations
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Wan heads both reviewed editing tables, but Arena's uncertainty overlaps the other selected entries.
Sources and observations (2)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Wan 3.0, named public leaderboard entry. Editing or transforming existing video; public output preference. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks. Original evaluation/snapshot: 2026-09-21. Tested: Wan 3.0, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    Wan 3.0, named public leaderboard entry. Editing or transforming existing video; public output preference. https://artificialanalysis.ai/video/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks. Original evaluation/snapshot: 2026-09-23. Tested: Wan 3.0, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Dreamina Seedance 2.5

ByteDance Seed

Seedance 2.5 is close to Wan and H3 in Arena's editing pool. Its camera-angle demonstrations are useful supplementary examples, not independent proof of universal preservation.

MultipleProvisional: evidence still developing
Best for
Camera and scene-revision trials with preservation checks
Price & access
Dreamina and selected partners; access and settings vary by region.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena editing supports exact Seedance 2.5 with overlapping intervals; a second exact-version independent editing comparison remains missing.
Sources and observations (1)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Dreamina Seedance 2.5, named public leaderboard entry. Editing or transforming existing video; public output preference. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks. Original evaluation/snapshot: 2026-09-21. Tested: Dreamina Seedance 2.5, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

MiniMax H3

MiniMax

Base H3 has editing support from both preference publishers. This supports testing hosted H3 on footage changes, while keeping identity preservation and the Max configuration separate.

MultipleProvisional: evidence still developing
Best for
Footage revisions using hosted base H3
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need a definitive winner across settings or tasks not covered by these comparisons.
Restrictions
Pending
Evidence
Arena and Artificial Analysis both support base H3 editing; no Max or local-model editing result is inherited.
Sources and observations (2)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    MiniMax H3, named public leaderboard entry. Editing or transforming existing video; public output preference. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks. Original evaluation/snapshot: 2026-09-21. Tested: MiniMax H3, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    MiniMax H3, named public leaderboard entry. Editing or transforming existing video; public output preference. https://artificialanalysis.ai/video/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks. Original evaluation/snapshot: 2026-09-23. Tested: MiniMax H3, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: Arena's September 21 snapshot has roughly 429–962 votes for these entries and broadly overlapping intervals. Artificial Analysis does not independently corroborate Seedance 2.5 here. Existing-footage transformations, camera changes and preservation are not interchangeable tasks.

How we reviewed the field (15 contenders)

Seek an independent exact-version Seedance 2.5 comparison and controlled identity/preservation tests; verify separate FLUX fast, Aleph and Omni editors by September 30.

  • considered

    Gemini Omni Flash: Only explicit Omni 1.1 results support this record. Unversioned older Omni Flash editing and Artificial Analysis results were excluded.

  • recommended

    Wan 3.0: Current generation and editing preference results are strong; cinematic performance and lip-sync findings vary by task.

  • recommended

    MiniMax H3: Base H3 has current generation and editing support. Hosted and local open-weight configurations differ; Max results are not inherited.

  • considered

    MiniMax H3 Max (fal): Artificial Analysis supports the exact fal H3 Max generation configuration; base-H3 editing evidence cannot validate Max.

  • considered

    Dreamina Seedance 2.0: Current 720p generation results and narrower lip-sync examples were considered separately from Seedance 2.5.

  • recommended

    Dreamina Seedance 2.5: Arena generation and editing plus specialist filmmaking examples support task-specific interest; 720p leaderboard and 1080p Dreamina results remain distinct.

  • insufficient evidence

    Kling 3.0: Base, Pro/Standard and Omni identities are not interchangeable. Current exact-configuration production/editing comparisons remain incomplete.

  • considered

    Veo 3.1: Generation comparisons exist, but Fast/Quality and resolution distinctions prevent a blanket editing or production claim.

  • considered

    FLUX 3 Video: Preview generation has comparative support. Separate video editing, upscaled output and announced controls do not prove reliable production breadth.

  • insufficient evidence

    Runway Gen-4.5: Generation evidence is available; Aleph editing results belong to another model and cannot establish Gen-4.5 editing strength.

  • insufficient evidence

    Luma Ray 3.2: Exact 3.2 support remains limited; Ray 3 or 3.14 results cannot be inherited by this record.

  • insufficient evidence

    HappyHorse 1.1: Exact 1.1 generation support does not inherit HappyHorse 1.0 editing comparisons.

  • insufficient evidence

    Vidu Q3 Pro: Pro endpoint settings and current matched production comparisons remain insufficiently documented.

  • insufficient evidence

    LTX-2.5: Current 2.5 capabilities are documented, but older 2.3 tests do not establish quality; beta editing and export claims require direct comparative evidence.

  • insufficient evidence

    FLUX Video Edit [fast]: The separate fast editing endpoint has official documentation but no verified matching independent comparison; generic FLUX 3 Video Edit results were excluded.

Qualified recommendations

Capabilities

The broadest genuinely useful production toolkit—not a checklist of novelty features.

Multiple referencesUseful durationSound generationDirecting controlsResolution & export

Recommendations by use case. Replace the unsupported broadest-toolkit ordering with a qualified shortlist for demonstrated generation and footage-revision workflows. Comparative editing evidence supports useful function; feature-count leadership is unproven.

Wan 3.0

Alibaba

Wan 3.0 has comparative support for both generation and existing-footage revision, making it a useful workflow candidate. Its reliability across references, sound and long scenes is not established.

MultipleProvisional: evidence still developing
Best for
Generation plus footage transformations
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need verified leadership across every production feature.
Restrictions
Pending
Evidence
Arena and Artificial Analysis support exact Wan 3.0 generation/editing configurations; this is not a complete production-quality or feature-breadth score.
Sources and observations (3)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Wan 3.0, named public leaderboard entry. Comparative footage-revision output as evidence of one useful production capability. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth. Original evaluation/snapshot: 2026-09-21. Tested: Wan 3.0, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    Wan 3.0, named public leaderboard entry. Comparative footage-revision output as evidence of one useful production capability. https://artificialanalysis.ai/video/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth. Original evaluation/snapshot: 2026-09-23. Tested: Wan 3.0, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    Wan 3.0 named text-to-video entry; keep audio pools separate. Text-to-video preference as one production-workflow signal. https://artificialanalysis.ai/video/leaderboard/text-to-video Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth. Original evaluation/snapshot: 2026-09-23. Tested: Wan 3.0 named text-to-video entry; keep audio pools separate. Settings: Artificial Analysis September 23 audio/no-audio pools considered separately. Limits: Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

MiniMax H3

MiniMax

Hosted base H3 has useful image-animation and footage-editing support. This demonstrates more than a feature announcement, while leaving local deployment, sound controls and native export as separate questions.

MultipleProvisional: evidence still developing
Best for
Reference-image animation plus footage revisions
Price & access
Use the exact linked hosted model; current plans and pricing vary.
Watch out for
You need verified leadership across every production feature.
Restrictions
Pending
Evidence
Arena image-to-video and editing plus Artificial Analysis editing support base H3; none of these results transfers to Max or local H3.
Sources and observations (3)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    MiniMax H3, named public leaderboard entry. Comparative footage-revision output as evidence of one useful production capability. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth. Original evaluation/snapshot: 2026-09-21. Tested: MiniMax H3, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth.

  • Artificial Analysisblind preference · Observed Sep 23, 2026

    MiniMax H3, named public leaderboard entry. Comparative footage-revision output as evidence of one useful production capability. https://artificialanalysis.ai/video/leaderboard/editing Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limitation: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth. Original evaluation/snapshot: 2026-09-23. Tested: MiniMax H3, named public leaderboard entry. Settings: Artificial Analysis current category pool; audio/no-audio pools remain separate. Provider-recommended settings are not assumed matched to Arena. Limits: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth.

  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    MiniMax H3 base hosted entry; not fal Max or local H3. Image-to-video preference as one production-workflow signal. https://arena.ai/leaderboard/image-to-video Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth. Original evaluation/snapshot: 2026-09-21. Tested: MiniMax H3 base hosted entry; not fal Max or local H3. Settings: Arena September 21 snapshot and named configuration. Limits: Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Dreamina Seedance 2.5

ByteDance Seed

Seedance 2.5 has comparative support for image animation and video editing. Its multi-shot filmmaking demonstrations motivate a trial, but do not establish the broadest reliable toolkit.

MultipleProvisional: evidence still developing
Best for
Filmmaking trials combining animation and scene revision
Price & access
Dreamina and selected partners; access and settings vary by region.
Watch out for
You need verified leadership across every production feature.
Restrictions
Pending
Evidence
Arena supports exact Seedance 2.5 image-to-video and editing. Independent cross-publisher support for its complete filmmaking workflow remains missing.
Sources and observations (2)
  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Dreamina Seedance 2.5, named public leaderboard entry. Comparative footage-revision output as evidence of one useful production capability. https://arena.ai/leaderboard/video-edit Arena September 21 category snapshot; retain named settings and reported uncertainty. Limitation: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth. Original evaluation/snapshot: 2026-09-21. Tested: Dreamina Seedance 2.5, named public leaderboard entry. Settings: Arena September 21 category snapshot; retain named settings and reported uncertainty. Limits: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth.

  • Arena video leaderboards — current routesblind preference · Observed Sep 23, 2026

    Dreamina Seedance 2.5 at 720p in Arena. Image-to-video preference as one production-workflow signal. https://arena.ai/leaderboard/image-to-video Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth. Original evaluation/snapshot: 2026-09-21. Tested: Dreamina Seedance 2.5 at 720p in Arena. Settings: Arena September 21 snapshot and named configuration. Limits: Generation preference does not evaluate a complete filmmaking workflow or verify maximum useful feature breadth.

Observation dates record our evidence review, not necessarily the source’s publication date. A source page is identified when the record has no individual article link.

Recommendation reviewed: Sep 23, 2026Official product / access information ↗

Still unresolved: This is not a ranking of maximum duration, reference count, native resolution or sound controls. Available comparisons do not jointly test all those capabilities. Preview promises, separate endpoints, upscaling and local versus hosted models cannot establish reliable breadth.

How we reviewed the field (15 contenders)

Build an exact-endpoint evidence matrix for references, duration, sound, controls and native export; prioritize direct comparative reliability and recheck by September 30.

  • considered

    Gemini Omni Flash: Only explicit Omni 1.1 results support this record. Unversioned older Omni Flash editing and Artificial Analysis results were excluded.

  • recommended

    Wan 3.0: Current generation and editing preference results are strong; cinematic performance and lip-sync findings vary by task.

  • recommended

    MiniMax H3: Base H3 has current generation and editing support. Hosted and local open-weight configurations differ; Max results are not inherited.

  • considered

    MiniMax H3 Max (fal): Artificial Analysis supports the exact fal H3 Max generation configuration; base-H3 editing evidence cannot validate Max.

  • considered

    Dreamina Seedance 2.0: Current 720p generation results and narrower lip-sync examples were considered separately from Seedance 2.5.

  • recommended

    Dreamina Seedance 2.5: Arena generation and editing plus specialist filmmaking examples support task-specific interest; 720p leaderboard and 1080p Dreamina results remain distinct.

  • insufficient evidence

    Kling 3.0: Base, Pro/Standard and Omni identities are not interchangeable. Current exact-configuration production/editing comparisons remain incomplete.

  • considered

    Veo 3.1: Generation comparisons exist, but Fast/Quality and resolution distinctions prevent a blanket editing or production claim.

  • considered

    FLUX 3 Video: Preview generation has comparative support. Separate video editing, upscaled output and announced controls do not prove reliable production breadth.

  • insufficient evidence

    Runway Gen-4.5: Generation evidence is available; Aleph editing results belong to another model and cannot establish Gen-4.5 editing strength.

  • insufficient evidence

    Luma Ray 3.2: Exact 3.2 support remains limited; Ray 3 or 3.14 results cannot be inherited by this record.

  • insufficient evidence

    HappyHorse 1.1: Exact 1.1 generation support does not inherit HappyHorse 1.0 editing comparisons.

  • insufficient evidence

    Vidu Q3 Pro: Pro endpoint settings and current matched production comparisons remain insufficiently documented.

  • insufficient evidence

    LTX-2.5: Current 2.5 capabilities are documented, but older 2.3 tests do not establish quality; beta editing and export claims require direct comparative evidence.

  • considered

    Grok Imagine Video 1.5: Base Grok 1.5 generation is distinct from Arena's agent configuration; agent placement was not inherited.

On the Horizon

New announcements we’re following. A confirmed launch is not a recommendation: independent comparisons and BMN’s category reviews still matter.

OpenAIAwaiting BMN review

GPT-6 Sol

Announced Sep 22, 2026

OpenAI has released GPT-6 Sol in ChatGPT Work, Codex, and the API. The company reports stronger coding and professional-work capabilities with lower API pricing than GPT-5.6 Sol.

What we’re checking

Independent, version-specific comparisons for writing quality, revision control, coding reliability, and practical value. Provider results do not establish a BMN category winner.

Original announcement ↗

Source checked Sep 23, 2026 · Next review target Sep 30, 2026

AnthropicAwaiting BMN review

Claude Opus 5.5

Announced Sep 22, 2026

Anthropic has released Claude Opus 5.5. The company reports stronger coding and professional work, clearer communication, and lower typical running costs than Opus 5.

What we’re checking

Independent comparisons with Fable 5.1 and current OpenAI models, plus task-specific writing and coding evidence. We still need matched settings and real workflow costs.

Original announcement ↗

Source checked Sep 23, 2026 · Next review target Sep 30, 2026

Know when the recommendations change.

Meaningful AI updates, plus On the Horizon: new releases we’re watching. About once a week. Only when there’s something worth sharing.

September 23, 2026

Qualified media recommendations are live.

All nine image and video categories now show unscored leading groups or choices by use case, with original comparisons, tested settings, a review of the eligible field, and the evidence still needed. GPT Image 2 and Riverflow Pro-high have separate records. The research interface and published methodology now share the same category-specific weights.

September 22, 2026

Image and video evidence, with its limits.

New source observations identify tested variants, task-specific findings, and conflicts with existing recommendations. They appear separately from the evidence behind the older rankings. Generic video summaries now receive the same provisional treatment as image summaries; scores and ranking-verification dates remain unchanged.

September 21, 2026

Clearer comparisons. Evidence you can inspect.

Larger text, direct category links, and expandable comparisons make the rankings easier to use. Evidence gaps are now visible, including inherited image placements awaiting version-specific confirmation. Ranking order and verification dates have not been advanced by this presentation update.

September 10, 2026

ChatGPT Images 2.5 replaces GPT Image 2.

OpenAI’s new image release adds faster generation, stronger subject preservation, more reliable multi-turn edits, and new direct-editing tools. The release initially inherited GPT Image 2’s placements. Those inherited placements are provisional until version-specific comparative evidence is documented. GPT-6 Astra’s access and API pricing have also been brought up to date.

September 8, 2026

Coding now ranks tested configurations.

Every Coding result identifies the system, underlying model or platform-selected AI, access plan, system version, and test date. Technical awards cleared the first supervised sweep; Builder awards are provisional until the full Practical Consensus threshold is met.

Evidence in.
A useful answer out.

We assess outside findings for the task and tested configuration. No single benchmark, arena, reviewer, or company determines a winner. These weights apply when comparable evidence supports scored rankings. Qualified image and video recommendations have no numerical scores and disclose unresolved comparisons. Builders use a separate profile: 20% working-product proof, 60% practitioner consensus, 10% ownership and publishing, and 10% ease and value.

Read the full methodology →
50–70%

Category performance

The capabilities that matter most for the specific award. Artistry and video use a quality-first profile.

15–20%

Real-world preference

Blind arenas, practitioner testing, and credible specialist consensus—with task fit made explicit.

5–15%

Access & workflow

Availability and usability can break a close call, but cannot rescue materially weaker output.

5–10%

Price & value

Cost matters after a model clears the quality bar; inexpensive mediocrity does not win.

5%

Restrictions

Material refusals, creative limitations, and other constraints that reduce practical usefulness.

Gate

Complete candidate pool

Every contender in the defined, dated evaluation field must be evaluated. New discoveries enter through an eligibility review.