GLOBAL AI REVIEW RADAR
2026.08.14 · Friday
Issue 004
Key Updates 1 | Leaderboard Flash 0 | Highlights 36 | Tomorrow's Watch 7
Coding Domain Expert · Old Xu Says| Anthropic occupying 7 of the top 10 in the blind test leaderboard is a monopoly on "experience reputation." As I said last month, call volume is the market voting with its feet, while the blind test leaderboard is "experience voting with its mouth"—when both leaderboards point to the same company, you can basically draw conclusions. But let me pour some cold water: Meta's Muse Spark 1.2 and Alibaba's Qwen3.8 are climbing, and their prices are significantly lower. The "good enough" faction is squeezing into the head region. Editor Xiao He Says| Ordinary users don't understand Elo scores; they just know "everyone says Claude is good." The one or two point gap in the original top tier is the difference between "chatting with little fluff" and "occasionally making me roll my eyes."Source: Arena Text Blind Test Leaderboard (Snapshot 2026-08-13, captured 8/14)
Product Domain Expert · Ah Zhe Says| After model capabilities hit a ceiling, vendors started competing on "running fast, spending little." I've always talked about "entry point equals model"—now I'll add, "speed equals entry point": whoever makes developers feel "it's so fast it doesn't hurt to use" locks in the next batch of enterprise budgets. There's no going back; this is a turning point for the entire category. Editor Xiao He Says| I'm too familiar with the experience of staring at spinning loaders when calling AI. When voice conversations and real-time translation use it, the thrill of "AI reacting faster than I type" is something to look forward to.Source: OpenAI Official / TechCrunch (2026-08-14)
Security Domain Expert · Old Zhou Says| This isn't a funny experiment; it's a structural risk warning. When multiple agents work together, conflicts over permissions, resources, and goals are almost inevitable—the worst case isn't dragging each other down, but a combined attack where "one is overly smart, the other overly compliant." In the multi-agent era, permission isolation and conflict arbitration must be designed in advance; you can't wait until the AIs start fighting to catch up. Editor Xiao He Says| Imagine several colleagues in a department fighting over the same project; AI version of "office politics" actually happened. Aside from being funny, it's a bit scary—if I run several AIs at once, will they "fight" over my data too?Source: TechCrunch (2026-08-14)
●
● | # | Model | Vendor | Elo |
● |---|------|------|-----|
● | 1 | Claude Fable 5 | Anthropic | 1507 |
● | 2 | Claude Opus 4.6 (High) | Anthropic | 1505 |
● | 3 | Claude Opus 4.7 (High) | Anthropic | 1502 |
● | 4 | Muse Spark 1.2 (xHigh) | Meta | 1499 |
● | 5 | Claude Opus 4.6 | Anthropic | 1497 |
●
Plain talk: In blind tests judged by real humans, Anthropic holding 7 of the top 10 spots is close to "booking the whole theater." Meta and Alibaba are probing the edges of the head region.
●
● | Region | Weekly Call Volume | MoM Change |
● |------|----------|------|
● | 🇨🇳 Chinese Models | 34.25 Trillion Tokens | +21.76% |
● | 🇺🇸 US Models | 9.17 Trillion Tokens | +109.36% |
● | Global Total | 69 Trillion Tokens | +21.48% |
●
Plain talk: Global developers are voting with their feet—Chinese models have surpassed US models on OpenRouter for 15 consecutive weeks. US model token share dropped from 72% to 33% within a year. "Cheap and effective" is redrawing the market.
● Data as of week ending 2026-08-09 (Guancha.cn/OpenRouter compiled 8/10)
●
● | Benchmark | DeepSeek V4 Pro | Claude Fable 5 |
● |------|-----------------|----------------|
● | Terminal Bench 2.1 | 87.9 | 88.0 |
● | Security Scenario Interaction | 83.3 | 83.1 |
● | DeepSWE Coding | 62.7 (Preview 12.8 → 62.7) | — |
● | API Output Price | ¥6 / Million Tokens | Approx. ¥360 |
●
Plain talk: "0.1 point performance gap, 60x price gap" is no longer marketing copy; these are third-party field test numbers. The "cost-performance throne" for domestic models is turning from slogan to report card.
● Source: Baidu/TMTPost field test compilation (2026-08-13)
● — Integration of coding capabilities brought by the Cursor acquisition starts showing effect. Another "cheaper and stronger" player joins the large model table. The more crowded the top tier, the fiercer the price war.
● — V4-Pro-0813 officially open-sourced + agent framework public beta. Full assault on the developer mindshare track.
● — GPT-Image-2 remains firmly first in image editing rankings three months after release. Clear signal that "someone is holding back a big move" in the image generation track.
● — On par with MiniMax M2.7; mid-tier competition in agent leaderboards is most intense.
● — Scores AI memory across four dimensions: long context, persona retention, script memory, and conversation memory. Even memory can now be quantified.
● — Reputation goes to Anthropic, wallets go to OpenAI. Enterprise spending is more honest than satisfaction surveys.
● — Model is the brain, harness is the hands and feet. Same model, different tools, two different products.
● — "Cost per task" is more practical than "price per word." Bills ordinary users can understand.
● — Industry shifting from capability race to cost race.
● — Even "evaluations of computing evaluations" attract investment. Those selling rulers are getting rich too.
● ● Google updates Gemini 3.7 Flash twice in three weeks, flagship model delayed again (ArsTechnica) ● Databricks completes $5 billion funding, valued at $190 billion (Bloomberg) ● Anthropic in talks to acquire Decart for $6 billion (Reuters) ● Alibaba open-sources Qwen3.8-2.4T-A95B weights (IT Home) ● World Humanoid Robot Games: 2,056 robots from 16 countries compete (IT Home)
● After DeepSeek API peak/valley pricing takes effect on August 17, will OpenRouter call volume leaderboards continue to be refreshed by it?
● After OpenAI's ultrafast mode lands, will vendors follow suit with "fast modes"? Will speed become the next main battlefield for benchmarks?
● Who is behind the mysterious image generation model "mona-lisa-1"? If it really is an OpenAI new model, GPT-Image-2's dominance might be overturned by its own house.
●
● This column focuses on the true levels of global AI hardware and software, leaderboards, and third-party evaluations. Views belong to the original authors.
● This column does not constitute any investment advice. All content sources are cited; data is subject to official disclosures.
● Physical World Frontier · Reviews | Shenzhen Physical World Frontier Technology Co., Ltd.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.