GLOBAL AI REVIEW RADAR
2026.09.13 · Sunday
Issue 033
Key Updates 1 | Leaderboard Flash 0 | Highlights 20 | Tomorrow's Watch 3
● The top tier of the Intelligence Index is as follows.
● | # | Model | Vendor | Total Score |
● |---|------|------|------|
● | 1 | Claude Fable 5.1 (Max w/ fallback) | Anthropic | 53 |
● | 2 | GPT-6 Astra (Max) | OpenAI | 53 |
● | 3 | Claude Opus 5 (Max) | Anthropic | 51 |
● The first tier is crowded around 53 points. The strongest configurations from both companies differ by only fractions, making it hard to tell who wins by eye. This leaderboard condenses dozens of exams into a single total score; it's good enough for a rough ranking, but choosing tools still requires checking specific metrics based on your scenario.
● In the open-weight leaderboard (models you can download and run locally), GLM-5.3 (Max) leads with 45, followed by Kimi K3 (Max) at 44, and GLM-5.3-Flash at 42. The top three spots are swept by two Chinese companies, trailing closed-source flagships by only about 8 points. A year ago, that gap was over 15 points.
● The speed champion is Celeris-1 at 1313.8 words per second. The cost-efficiency champion is Granite 4.2 3B at one cent per task. Fast, cheap, and strong are three different leaderboards.
● Data taken from artificialanalysis.ai/leaderboards/models. The latest snapshot of the Arena Human Blind Test leaderboard remains September 12, same as last issue, so it wasn't used today.
●
●
●
●
●
●
●
●
●
●
● Vals.ai IOI leaderboard launch, agent "eval version ≠ deployment version" going viral, 42-day robot dog long-term test, AA real-time leaderboard tie at 53 points, Wired abuse archive finale.
●
●
●
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.