Key Updates 3 | Leaderboard Flash 0 | Highlights 0 | Tomorrow's Watch 0
Update 1
| Rank |
Model |
Vendor |
Score |
| 1 |
Claude Opus 5 (max) |
Anthropic |
63 |
| 2 |
Claude Fable 5 |
Anthropic |
62 |
| 3 |
GPT-5.6 Sol (max) |
OpenAI |
61 |
| 4 |
Grok 4.6 (high) |
SpaceXAI |
61 |
| 5 |
Kimi K3 (max) |
Moonshot AI |
60 |
| 6 |
GLM-5.3 (max) |
Zhipu |
60 |
| 7 |
Qwen3.8-Max |
Alibaba |
58 |
Plain Talk | This score is an "Average AI IQ" calculated by weighting 9 standard exams
| Rank |
Model |
Vendor |
Score |
| 1 |
Claude Opus 5 (max) |
Anthropic |
63 |
| 2 |
Claude Fable 5 |
Anthropic |
62 |
| 3 |
GPT-5.6 Sol (max) |
OpenAI |
61 |
| 4 |
Grok 4.6 (high) |
SpaceXAI |
61 |
| 5 |
Kimi K3 (max) |
Moonshot AI |
60 |
| 6 |
GLM-5.3 (max) |
Zhipu |
60 |
| 7 |
Qwen3.8-Max |
Alibaba |
58 |
Plain Talk | This score is an "Average AI IQ" calculated by weighting 9 standard exams. Today's capture shows Anthropic sweeping the top two. OpenAI and xAI tie for the third tier. Chinese camp's Kimi K3 and GLM-5.3 both squeeze into the 60-point tier; Qwen3.8-Max touches 58. The score gap between top US and Chinese models has converged to single digits.
Update 2
| Rank |
Model |
Vendor |
Elo |
| 1 |
Claude Opus 5 (max) |
Anthropic |
1692 |
| 2 |
Kimi K3 (max) |
Moonshot AI |
1674 |
| 3 |
Qwen3.8-Max |
Alibaba |
1667 |
| 4 |
Claude Opus 5 (high) |
Anthropic |
1663 |
| 5 |
Grok 4.6 (high) |
SpaceXAI |
1631 |
Plain Talk | This is a leaderboard where developers anonymously vote on "whose code is more usable." Biggest news in the latest snapshot: Two of the top three are Chinese models
| Rank |
Model |
Vendor |
Elo |
| 1 |
Claude Opus 5 (max) |
Anthropic |
1692 |
| 2 |
Kimi K3 (max) |
Moonshot AI |
1674 |
| 3 |
Qwen3.8-Max |
Alibaba |
1667 |
| 4 |
Claude Opus 5 (high) |
Anthropic |
1663 |
| 5 |
Grok 4.6 (high) |
SpaceXAI |
1631 |
Plain Talk | This is a leaderboard where developers anonymously vote on "whose code is more usable." Biggest news in the latest snapshot: Two of the top three are Chinese models. Kimi K3 is second, Qwen3.8-Max third, biting closely at first-place Claude. DeepSeek's V4 Pro also squeezed into the top ten. Programming capabilities of Chinese open-source models are now on the same starting line as top closed-source models.
Update 3
| Rank |
Model |
Vendor |
Elo |
| 1 |
GPT-Image-2 (medium) |
OpenAI |
1463 |
| 2 |
Grok Imagine 2.0 (low) |
SpaceXAI |
1439 |
| 3 |
MAI-Image 2.6 (Preview) |
Microsoft |
1420 |
| 4 |
Muse Image |
Meta |
1406 |
| 5 |
Seedream 5.0 Pro |
ByteDance |
1394 |
Plain Talk | Image editing blind test is "Let AI edit images per your instructions, compare whose edits please you most." GPT-Image-2 sits firmly first with nearly 200k battles
| Rank |
Model |
Vendor |
Elo |
| 1 |
GPT-Image-2 (medium) |
OpenAI |
1463 |
| 2 |
Grok Imagine 2.0 (low) |
SpaceXAI |
1439 |
| 3 |
MAI-Image 2.6 (Preview) |
Microsoft |
1420 |
| 4 |
Muse Image |
Meta |
1406 |
| 5 |
Seedream 5.0 Pro |
ByteDance |
1394 |
Plain Talk | Image editing blind test is "Let AI edit images per your instructions, compare whose edits please you most." GPT-Image-2 sits firmly first with nearly 200k battles. Microsoft's preview model parachuted straight to third, but with only 5k+ battles, the score isn't stable yet—just take a look. Highest domestic rank is ByteDance's Seedream 5.0 Pro at sixth. For image editing, it's still American companies' turf.
Section Highlights
Pander Score New Benchmark | When users express strong opinions, does AI stick to facts or pander to the user? This new benchmark aims to quantify this. "AI pandering to users" finally has a ruler. Assistants that love marketing hype should be pulled out and tested.
Bloomberg Cross-Comparison of US-China AI | Placing agents from ChatGPT, Gemini, DeepSeek, Kimi into the same coordinate system to compare who is more capable and who is cheaper. One chart clarifies the current landscape. Cross-comparison charts from authoritative media are suitable for saving as reference coordinates.
Window Cleaning Robot Field Test | Wired tested high-end window cleaning robots. Cleaning effect is okay, but slow climbing, missed corners, and issues returning to charge dock are possible. Conclusion: Don't buy, use a rag. Another record of smart hardware failure. Check field tests before buying high-tech household tools.
10 Common Security Flaws in AI-Generated Apps | Hardcoded keys, excessive permissions, unvalidated inputs. A developer names the 10 most common security errors in AI-generated code one by one. AI writing code is indeed fast, but writing secure code still requires human oversight.
AI Output Becomes New Attack Vector | Tech share at Symfony conference proposes that AI-generated content is replacing human input as the new vector for injection and privilege escalation. We used to sanitize user input; now we must validate AI output.
Zeno Local AI Workbench | Developers open-sourced Zeno, focusing on MacBo...
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.
ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)