For writing that sounds like you.
Long-form thinking, clean structure, and a better first draft when the blank page is the hardest part.
The right model depends on the job. Start with writing, coding, agents, research, or creative work — then see the evidence behind the shortlist.
THE CONSENSUS INDEX
Scores reflect relative performance across public model evaluations. Coverage and evidence confidence are shown alongside the score so a ranking never has to tell the whole story by itself.
| Rank | Model | |||||
|---|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 2026-06-09 | 88.0% | ¥67.21 | ¥336.03 | 89.4 Strong evidence |
| 02 | Claude Opus 5Anthropic | 2026-07-24 | 84.5% | ¥33.60 | ¥168.02 | 86.2 Strong evidence |
| 03 | GPT-5.6 SolOpenAI | 2026-07-09 | 84.5% | ¥33.60 | ¥201.62 | 83.3 Strong evidence |
| 04 | Kimi K3Moonshot AI | 2026-07-16 | 84.5% | ¥20.00 | ¥100.00 | 79.9 Strong evidence |
| 05 | GLM-5.3Z.ai | 2026-08-18 | 60.0% | ¥9.41 | ¥29.57 | 76.9 Moderate evidence |
| 06 | Gemini 3.7 FlashGoogle | 2026-08-13 | 81.0% | ¥5.04 | ¥25.20 | 75.7 Strong evidence |
| 07 | Grok 4.6xAI | 2026-08-12 | 84.5% | ¥13.44 | ¥40.32 | 74.8 Strong evidence |
| 08 | GPT-5.5OpenAI | 2026-04-23 | 80.0% | ¥33.60 | ¥201.62 | 73.9 Moderate evidence |
| 09 | Muse Spark 1.2Meta | 2026-08-05 | 73.0% | ¥8.40 | ¥28.56 | 72.5 Moderate evidence |
| 10 | Qwen3.8 MaxAlibaba | 2026-07-19 | 81.0% | ¥12.00 | ¥36.00 | 71.7 Strong evidence |
| 11 | Claude Opus 4.8Anthropic | 2026-05-28 | 88.0% | ¥33.60 | ¥168.02 | 70.6 Strong evidence |
| 12 | GPT-5.6 TerraOpenAI | 2026-07-09 | 81.0% | ¥13.44 | ¥80.65 | 69.3 Moderate evidence |
| 13 | Claude Opus 4.7Anthropic | 2026-04-16 | 72.0% | ¥33.60 | ¥168.02 | 64.1 Moderate evidence |
| 14 | Muse Spark 1.1Meta | 2026-07-09 | 76.5% | ¥8.40 | ¥28.56 | 61.2 Moderate evidence |
| 15 | DeepSeek V4 Pro 0813DeepSeek | 2026-08-13 | 73.0% | ¥9.00 | ¥27.00 | 58.7 Moderate evidence |
| 16 | GPT-5.4OpenAI | 2026-03-05 | 69.5% | ¥16.80 | ¥100.81 | 57.1 Moderate evidence |
| 17 | Grok 4.5xAI | 2026-07-08 | 88.0% | ¥13.44 | ¥40.32 | 56.4 Strong evidence |
| 18 | Claude Sonnet 5Anthropic | 2026-06-30 | 84.5% | ¥13.44 | ¥67.21 | 55.7 Strong evidence |
| 19 | GPT-5.6 LunaOpenAI | 2026-07-09 | 81.0% | ¥1.34 | ¥8.06 | 51.9 Moderate evidence |
| 20 | Gemini 3.6 FlashGoogle | 2026-07-21 | 81.0% | ¥5.04 | ¥25.20 | 49.9 Moderate evidence |
| 21 | Qwen3.8 27BAlibaba | 2026-08-14 | 51.0% | ¥3.00 | ¥12.00 | 47.9 Limited evidence |
| 22 | Gemini 3.5 FlashGoogle | 2026-05-19 | 81.0% | ¥10.08 | ¥60.49 | 47.6 Strong evidence |
| 23 | GLM-5.2Z.ai | 2026-06-16 | 84.5% | ¥8.00 | ¥28.00 | 46.8 Strong evidence |
| 24 | Claude Opus 4.6Anthropic | 2026-02-05 | 73.0% | ¥33.60 | ¥168.02 | 46.1 Moderate evidence |
| 25 | DeepSeek V4 Flash 0731DeepSeek | 2026-07-31 | 76.5% | ¥3.00 | ¥9.00 | 42.9 Moderate evidence |
| 26 | Gemini 3.1 Pro PreviewGoogle | 2026-02-19 | 88.0% | ¥13.44 | ¥80.65 | 41.1 Strong evidence |
| 27 | Claude Sonnet 4.6Anthropic | 2026-02-17 | 76.5% | ¥20.16 | ¥100.81 | 37.4 Moderate evidence |
| 28 | Qwen3.7 MaxAlibaba | 2026-05-19 | 51.0% | ¥12.00 | ¥36.00 | 34.3 Limited evidence |
| 29 | GPT-5.3 CodexOpenAI | 2026-02-05 | 42.0% | ¥11.76 | ¥94.09 | 32.2 Limited evidence |
| 30 | DeepSeek V4 Pro PreviewDeepSeek | 2026-04-24 | 73.0% | ¥2.92 | ¥5.85 | 31.8 Moderate evidence |
Benchmark results link directly to the original public evaluation boards. Provider pricing is checked against official websites and converted to CNY where necessary.
START WITH THE WORK
There is no universal winner. Pick the job first, then compare the models that make sense for that job.
Long-form thinking, clean structure, and a better first draft when the blank page is the hardest part.
A shortlist for debugging, refactors, code review, and the long middle between “it should work” and production.
Turn rough thinking into clear briefs, polished presentations, and documents that are easier to scan and act on.
Models that make a useful core for tool calls, multi-step plans, research loops, and background execution.
A practical starting set for synthesis, comparison, source-heavy analysis, and making sense of too much information.
Video-specific evaluations are the next lane for this index. For now, start here for storyboards, shot lists, visual direction, and creative prompts.
Editorial starting points based on the current AI Center snapshot — not a replacement for the full evidence table.
READ THE SIGNAL
Each model is placed on a shared capability scale using real head-to-head results across the evaluation set. The final score is mapped from 0 to 100.
Coverage shows how much evidence a model has actually earned. Missing evaluations are not treated as zeroes, and missing source weight is never redistributed.
High, medium, and low confidence describe how complete and stable the evidence is. They do not describe the model’s quality.
A TRANSPARENT METHOD
AI Center’s model index is designed for decisions, not hot takes. A model can lead with less evidence, and a narrow score gap does not automatically mean a meaningful capability gap.
Read the methodology