GLOBAL AI REVIEW RADAR
2026.09.23 · Wednesday
Issue 043
Key Updates 1 | Leaderboard Flash 0 | Highlights 40 | Tomorrow's Watch 2
What it means for you|Next time you use AI to pick something, have it only do the filtering step, verify the specs yourself on the official site, and click pay and order yourself.
● 3 key updates (with dual expert commentary), 3 leaderboard briefs, 10 section picks, plus What Everyone's Watching and Tomorrow's Watchlist.
● Let me be honest up front. I checked today, and Arena's latest snapshot is still the 2026-09-21 version, the same one used last issue, so the leaderboards are not yet updated, data as of 2026-09-21. Today I also didn't get a new leaderboard to swap in, and I'd rather label it honestly than pass off old data as new. The visual blind-test leaderboard from last issue is pulled this time, replaced by the text blind-test leaderboard not used last issue.
●
● | # | Model | Vendor | Score (calculated from human votes) |
● |---|------|------|----------------|
● | 1 | claude-fable-5-high | Anthropic | 1506 (30,057 matches) |
● | 2 | claude-opus-4-6-high | Anthropic | 1505 (71,993 matches) |
● | 3 | claude-opus-4-7-high | Anthropic | 1502 (60,002 matches) |
● | 4 | muse-spark-1.2 (xHigh) | Meta | 1500 (3,227 matches) |
● | 5 | claude-fable-5.1-max | Anthropic | 1498 (5,783 matches) |
● Plain-language read. The top five are separated by only 8 points, so first and fifth are basically a tie. Fourth is Meta's Muse Spark, with only 3,227 matches, and the leaderboard lists its fluctuation as 11 points, bigger than its gap from third, so don't take the ranking seriously yet. Four of the top five are Anthropic's Claude. Also, this leaderboard was last updated September 13, eight days earlier than the snapshot date.
●
● | # | Model | Vendor | Score (match count in parentheses) |
● |---|------|------|----------------|
● | 1 | gpt-6-astra-max | OpenAI | 1800 (2,281 matches, ±16 points) |
● | 2 | claude-fable-5.1-max | Anthropic | 1758 (3,036 matches) |
● | 3 | claude-opus-5-max | Anthropic | 1687 (12,087 matches) |
● | 4 | qwen3.8-max-0902 | Alibaba | 1681 (2,262 matches) |
● | 5 | kimi-k3-max | Moonshot AI | 1674 (4,547 matches) |
● Plain-language read. First place has 1800 points, opening a 42-point gap over second, which looks solid. Its match count is only 2,281, with ±16 points fluctuation; third has 12,087 matches, with only ±7 points fluctuation. Scores with thin match counts don't hold up to scrutiny. One more thing: this leaderboard was last updated September 11, unchanged for thirteen days, and four of the top five have fewer than five thousand matches.
●
● | # | Model | Vendor | Sample (match count) | Rate of tool misstatement |
● |---|------|------|----------------|----------------|
● | 1 | Claude Fable 5.1 (Max) | Anthropic | 13,320 | 0.37 |
● | 2 | GPT 6 Astra (Max) | OpenAI | 10,372 | 0.37 |
● | 3 | Claude Opus 5 (High) | Anthropic | 24,794 | 0.33 |
● | 4 | Claude Opus 5 (Max) | Anthropic | 19,934 | 0.35 |
● | 5 | Claude Fable 5 (High) | Anthropic | 38,293 | 0.37 |
● Plain-language read. This leaderboard tests AI that can call tools on its own and work through several steps in a row, called agents in the industry. It doesn't rank an overall score, but calculates six items separately: getting things done, saying fewer wrong things, following instructions, and recovering on its own after errors. "Tool misstatement" means it claims a tool call that never happened actually happened. The top five all fall between 0.33 and 0.37 on this item, a tiny difference, and what separates the rankings is the getting-things-done item: first place 19.83, second 17.7. The thickest match count isn't the top spot; fourth played over nineteen thousand matches.
●
●
●
●
●
●
●
●
●
●
● Musk's Grok Bot passed 400,000 users in a month. Microsoft Xbox had another round of layoffs, with several studios merged into Activision. Someone on Hacker News found that Couchsurfing's entire site was rewritten by AI, with images also officially generated. AstroForge handed command of its next asteroid mission to AI.
● The just-released Claude Opus 5.5 and GPT-6 Sol, Luna haven't entered the Arena snapshot yet (the snapshot is stuck at 2026-09-21); wait for the next refresh to see if the rankings hold. ② The Yunqi Conference runs until September 24; third-party benchmarks for AgentCore and Qwen4 aren't out yet, so note the official claims for now. ③ Meta's Muse passed 500,000 users in its first week; wait for third parties to produce actual task completion rates before discussing whether it counts as a shopping gateway.
● The above is the preview version of Issue 056; the full illustrated version is on the review publication page.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.