Best Model Now does not reuse one universal leaderboard. Each tab—and each category within it—changes what we measure, which sources qualify, and how much each factor counts.
One framework, four scoring profiles
Every eligible contender receives a score on the same five practical factors. The weights change to match the category. “Category performance” means the concrete criteria shown on that ranking card—not a generic intelligence score.
Quality firstArtistry; Video Overall, Text-to-Video, Image-to-Video, Video Editing65%20%5%5%5%
Capability firstVideo Capabilities70%15%5%5%5%
Creator ConsensusWeb App, Mobile App, and Website Builders55%25%10%5%5%
What changes by tab
Evidence follows the actual task.
Writing
Writing quality, instruction control, revision strength, coherence, and human judgment. Screenwriting remains provisional until blind dramatic-writing review is complete.
Coding
Technical categories use benchmarks and repository-level system tests. Builder categories use Creator Consensus because end-to-end product quality, iteration, and publishing matter more than isolated code generation.
Image
Blind preference, edit preservation, realism, control, and specialist visual judgment. Artistry deliberately weights output quality more heavily.
Video
Motion coherence, temporal stability, direction, reference preservation, production features, and real access. Quality-first weighting applies to most video awards.
What counts as a contender
We rank the thing a person can actually use.
A pure model can be ranked by itself. A coding system or builder platform is different: its harness, tools, routing, plan, and release can materially change results. We therefore record tested configurations such as Codex + GPT-6 Astra.
When a platform chooses models behind the scenes, the public ranking says Platform-selected AI. That means the service may route a task to one model or several, and the exact stack may change. We rank the observable product configuration, record its version, plan, and test date, and do not pretend a single disclosed model powers it.
Evidence policy
Objective evidence
Benchmarks, arenas, structured leaderboards, and reproducible datasets. They must map meaningfully to the category.
Specialist evidence
Qualified creators and practitioners who show their work, test relevant tasks, and disclose enough context to judge the finding.
BMN testing
Controlled first-party test suites. Prompts, candidate access, outputs, blind rankings, reviewer identity, settings, costs, and dates are recorded. An incomplete suite cannot influence a published ranking.
Screenwriting blind test
One qualifier scene. One finalist scene.
All seven contenders receive the same original-scene qualifier. The top six receive a different original-scene final, and the two rounds contribute equal points. Because one human operator pastes each named output, the test is partially blind: capture is identified, while reading and ranking use shuffled labels without model names. Dedicated rewriting and formatting tests are excluded from this initial sweep; formatting is judged within the scenes.
Generate once in a clean conversation using the exact prompt.
Disable browsing, tools, memory, projects, custom instructions, and rerolls.
Record the product, plan, setting, generation time, and cost when available.
Avoid reviewing outputs during capture; judge them only in the shuffled ranking view.
ContenderBMN Standard
GPT-6 AstraLow reasoning
Claude Fable 5.1High effort · Adaptive thinking
Claude Opus 5High effort · Adaptive thinking
Claude Sonnet 5High effort · Adaptive thinking
Gemini 3.8 FlashDynamic thinking · Medium
DeepSeek V4 ProThinking enabled · High
Kimi K3Thinking enabled · High
Publication gates
A score alone is not enough.
The eligible candidate field must be complete.
Contenders must be named at the level actually tested: model, system, platform, or configuration.
Evidence must meet the category's relevance, independence, freshness, and diversity threshold.
Builder rankings remain provisional until Creator Consensus is complete.
Screenwriting remains provisional until its blind human review is complete.
Dates, versions, access plans, routing, and material limitations must be recorded.
Rankings cannot be purchased. Sponsorship never changes eligibility, scoring, or placement. Material corrections are recorded and rankings are reverified as products change.