GLOBAL AI REVIEW RADAR
2026.09.21 · Monday
Issue 041
Key Updates 1 | Leaderboard Flash 0 | Highlights 39 | Tomorrow's Watch 4
What it means for you|As permissions granted to AI grow larger, this public battle arena is essentially a public health check for the security baselines of various models.
● 3 key updates (with dual-expert commentary), 3 groups of leaderboard news flashes, 10 curated section highlights, plus "What Everyone is Watching" and "Tomorrow's Focus."
● Let me tell you straight. Upon checking today, the latest snapshots for the four Arena blind test leaderboards are still from 2026-09-20, the same version used in the previous review. So the leaderboards haven't been updated; data is as of 2026-09-20. The Intelligence Index column was freshly scraped from Artificial Analysis today. Compare the two together.
●
● | # | Model | Vendor | Elo (Battle Count) |
● |---|------|------|----------------|
● | 1 | Claude Fable 5 (High) | Anthropic | 1506 (30,057 battles) |
● | 2 | Claude Opus 4-6 (High) | Anthropic | 1505 (71,993 battles) |
● | 3 | Claude Opus 4-7 (High) | Anthropic | 1502 (60,002 battles) |
● | 4 | Muse Spark 1.2 (xHigh) | Meta | 1500 (Only 3,227 battles) |
● | 5 | Claude Fable 5.1 (Max) | Anthropic | 1498 (5,783 battles) |
● Plain language interpretation. In blind tests judged by humans, Claude's parent company nearly monopolizes the field, taking seven of the top ten spots. Meta's new model reaching #4 looks impressive, but it has only fought over three thousand battles. Its score could fluctuate by 11 points, so don't take the ranking seriously yet. The newly scraped Intelligence Index aligns with this conclusion: Claude Fable 5.1 and GPT-6 Astra tie at 53 points, occupying the head of the pack. The domestic tier formed by Alibaba's Qwen3.8 Max, Moonshot AI's Kimi K3, and Zhipu's GLM-5.3 is biting closely at 44 to 45 points; the gap is already very small.
●
● | # | Model | Vendor | Elo (Battle Count) |
● |---|------|------|----------------|
● | 1 | GPT-6 Astra (Max) | OpenAI | 1800 (Only 2,281 battles, ±16 points) |
● | 2 | Claude Fable 5.1 (Max) | Anthropic | 1758 (3,036 battles) |
● | 3 | Claude Opus 5 (Max) | Anthropic | 1687 (12,087 battles) |
● | 4 | Qwen3.8 Max (0902) | Alibaba | 1681 (2,262 battles) |
● | 5 | Kimi K3 (Max) | Moonshot AI | 1674 (4,547 battles) |
● Plain language interpretation. The top score is scary, but the sample size is less than one-fifth of the third place. Check if new rankings wobble; wait for stability over one or two rounds before drawing conclusions. What's solid is the third place, backed by twelve thousand battles.
●
● | # | Model | Vendor | Task Completion Rate | Self-Correction on Error | Tool Hallucination Rate |
● |---|------|------|---------|-----------|-----------|
● | 1 | Claude Fable 5.1 (Max) | Anthropic | 19.83 | 12.62 | 0.37 |
● | 2 | GPT-6 Astra (Max) | OpenAI | 17.70 | 7.28 | 0.37 |
● | 3 | Claude Opus 5 (High) | Anthropic | 9.24 | 11.98 | 0.33 |
● | 4 | Claude Opus 5 (Max) | Anthropic | 12.41 | 13.08 | 0.35 |
● | 5 | Claude Fable 5 (High) | Anthropic | 5.97 | 9.06 | 0.37 |
● Plain language interpretation. This leaderboard has no total score, only six report cards. Ranked 6th, Claude Opus 4.8 has the lowest tool hallucination rate in the field at 0.12, meaning it makes up fake tools the least, but it was dragged out of the top five by its task completion rate. When choosing an AI to do work for you, look at dimensions first, then rankings. The 8th ranked Kimi K3 is the only domestic model in the top ten, with the thickest sample size exceeding 100,000 battles. No new snapshot for the image leaderboard this period, so it's not listed.
●
●
●
●
●
●
●
●
●
●
● Newly scraped Intelligence Index today: Claude Fable 5.1 and GPT-6 Astra tie at 53 points, occupying the head. The domestic tier formed by Qwen3.8 Max, GLM-5.3, and Kimi K3 bites closely at 44 to 45 points. Faraday Future released nine robots at once, the most expensive exceeding 920,000 yuan. Gurman says Apple may release a smart home screen device as early as next month.
● The four Arena leaderboards are still on the 09-20 version today. Watch if the 09-21 snapshot refreshes, and see if the battle count for Meta's new models exceeds 5,000.
● Zhipu's MaaS data non-retention mechanism was just announced for recent launch. Wait for official access rules and real-world feedback.
● Alibaba's open-sourced Qwen-Image-2.1 image model is hot today. Wait for third parties to run it against the same topics as the popular image leaderboard from the last issue. Don't just look at official demos.
● The above is the preview for Issue #054. The formal version updates at 17:00. The full illustrated version is on the review publication page.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.