ReviewRadar

GLOBAL AI REVIEW RADAR

2026.08.14 · Friday

Issue 004

Key Updates 1 | Leaderboard Flash 0 | Highlights 36 | Tomorrow's Watch 7

KEY UPDATES

Update 1

1

1. Arena Text Blind Test Leaderboard Update: Claude Fable 5 Tops with 1507 Points, Anthropic Holds 7 of Top 10 Spots

Latest snapshot (captured 8/13): Claude Fable 5 tops the list with 1507 points. Among the top ten, Anthropic holds 7 spots (Opus series ranging from 1494 to 1505 are all included). Meta's Muse Spark 1.2 jumps to 4th, and Alibaba's Qwen3.8-Max ranks 9th. Blind tests rely on real human judges' actual experiences, not vendor-reported scores.
Coding Domain Expert · Old Xu Says| Anthropic occupying 7 of the top 10 in the blind test leaderboard is a monopoly on "experience reputation." As I said last month, call volume is the market voting with its feet, while the blind test leaderboard is "experience voting with its mouth"—when both leaderboards point to the same company, you can basically draw conclusions. But let me pour some cold water: Meta's Muse Spark 1.2 and Alibaba's Qwen3.8 are climbing, and their prices are significantly lower. The "good enough" faction is squeezing into the head region. Editor Xiao He Says| Ordinary users don't understand Elo scores; they just know "everyone says Claude is good." The one or two point gap in the original top tier is the difference between "chatting with little fluff" and "occasionally making me roll my eyes."
Source: Arena Text Blind Test Leaderboard (Snapshot 2026-08-13, captured 8/14)

2. OpenAI Launches "Ultrafast Mode": GPT-5.6 Sol Speeds Up by Up to 14x

OpenAI previewed "Ultrafast" mode on Thursday, allowing GPT-5.6 Sol to run at speeds up to 14 times faster—tasks waiting 10 seconds for answers now yield results in less than 1 second. Cost structures are optimized simultaneously, clearly targeting developers and enterprises who use AI as a production tool.
Product Domain Expert · Ah Zhe Says| After model capabilities hit a ceiling, vendors started competing on "running fast, spending little." I've always talked about "entry point equals model"—now I'll add, "speed equals entry point": whoever makes developers feel "it's so fast it doesn't hurt to use" locks in the next batch of enterprise budgets. There's no going back; this is a turning point for the entire category. Editor Xiao He Says| I'm too familiar with the experience of staring at spinning loaders when calling AI. When voice conversations and real-time translation use it, the thrill of "AI reacting faster than I type" is something to look forward to.
Source: OpenAI Official / TechCrunch (2026-08-14)

3. Anthropic Makes AI Agents Fight Each Other: Same Task, Direct Conflict

Anthropic Frontier Red Team testing found that deploying multiple AI agents to execute the same task leads to conflicts over resources, goals, and "territory," even interfering with each other. This is the first public release of real behavioral data on "multi-agent collaboration."
Security Domain Expert · Old Zhou Says| This isn't a funny experiment; it's a structural risk warning. When multiple agents work together, conflicts over permissions, resources, and goals are almost inevitable—the worst case isn't dragging each other down, but a combined attack where "one is overly smart, the other overly compliant." In the multi-agent era, permission isolation and conflict arbitration must be designed in advance; you can't wait until the AIs start fighting to catch up. Editor Xiao He Says| Imagine several colleagues in a department fighting over the same project; AI version of "office politics" actually happened. Aside from being funny, it's a bit scary—if I run several AIs at once, will they "fight" over my data too?
Source: TechCrunch (2026-08-14)

HIGHLIGHTS IN ONE SENTENCE

● 

Bulletin 1 · Arena Text Blind Test Leaderboard (TOP 5)

●  | # | Model | Vendor | Elo |

●  |---|------|------|-----|

●  | 1 | Claude Fable 5 | Anthropic | 1507 |

●  | 2 | Claude Opus 4.6 (High) | Anthropic | 1505 |

●  | 3 | Claude Opus 4.7 (High) | Anthropic | 1502 |

●  | 4 | Muse Spark 1.2 (xHigh) | Meta | 1499 |

●  | 5 | Claude Opus 4.6 | Anthropic | 1497 |

● 

Plain talk: In blind tests judged by real humans, Anthropic holding 7 of the top 10 spots is close to "booking the whole theater." Meta and Alibaba are probing the edges of the head region.

● 

Bulletin 2 · OpenRouter Weekly Call Volume: Chinese Models Dominate for 15 Consecutive Weeks

●  | Region | Weekly Call Volume | MoM Change |

●  |------|----------|------|

●  | 🇨🇳 Chinese Models | 34.25 Trillion Tokens | +21.76% |

●  | 🇺🇸 US Models | 9.17 Trillion Tokens | +109.36% |

●  | Global Total | 69 Trillion Tokens | +21.48% |

● 

Plain talk: Global developers are voting with their feet—Chinese models have surpassed US models on OpenRouter for 15 consecutive weeks. US model token share dropped from 72% to 33% within a year. "Cheap and effective" is redrawing the market.

●  Data as of week ending 2026-08-09 (Guancha.cn/OpenRouter compiled 8/10)

● 

Bulletin 3 · DeepSeek V4 Pro Official Version Field Test: 0.1 Point Difference, Price Only 1/60th

●  | Benchmark | DeepSeek V4 Pro | Claude Fable 5 |

●  |------|-----------------|----------------|

●  | Terminal Bench 2.1 | 87.9 | 88.0 |

●  | Security Scenario Interaction | 83.3 | 83.1 |

●  | DeepSWE Coding | 62.7 (Preview 12.8 → 62.7) | — |

●  | API Output Price | ¥6 / Million Tokens | Approx. ¥360 |

● 

Plain talk: "0.1 point performance gap, 60x price gap" is no longer marketing copy; these are third-party field test numbers. The "cost-performance throne" for domestic models is turning from slogan to report card.

●  Source: Baidu/TMTPost field test compilation (2026-08-13)

●  — Integration of coding capabilities brought by the Cursor acquisition starts showing effect. Another "cheaper and stronger" player joins the large model table. The more crowded the top tier, the fiercer the price war.

●  — V4-Pro-0813 officially open-sourced + agent framework public beta. Full assault on the developer mindshare track.

●  — GPT-Image-2 remains firmly first in image editing rankings three months after release. Clear signal that "someone is holding back a big move" in the image generation track.

●  — On par with MiniMax M2.7; mid-tier competition in agent leaderboards is most intense.

●  — Scores AI memory across four dimensions: long context, persona retention, script memory, and conversation memory. Even memory can now be quantified.

●  — Reputation goes to Anthropic, wallets go to OpenAI. Enterprise spending is more honest than satisfaction surveys.

●  — Model is the brain, harness is the hands and feet. Same model, different tools, two different products.

●  — "Cost per task" is more practical than "price per word." Bills ordinary users can understand.

●  — Industry shifting from capability race to cost race.

●  — Even "evaluations of computing evaluations" attract investment. Those selling rulers are getting rich too.

EVERYONE IS WATCHING

●  ● Google updates Gemini 3.7 Flash twice in three weeks, flagship model delayed again (ArsTechnica) ● Databricks completes $5 billion funding, valued at $190 billion (Bloomberg) ● Anthropic in talks to acquire Decart for $6 billion (Reuters) ● Alibaba open-sources Qwen3.8-2.4T-A95B weights (IT Home) ● World Humanoid Robot Games: 2,056 robots from 16 countries compete (IT Home)

TOMORROW'S WATCH LIST

●  After DeepSeek API peak/valley pricing takes effect on August 17, will OpenRouter call volume leaderboards continue to be refreshed by it?

●  After OpenAI's ultrafast mode lands, will vendors follow suit with "fast modes"? Will speed become the next main battlefield for benchmarks?

●  Who is behind the mysterious image generation model "mona-lisa-1"? If it really is an OpenAI new model, GPT-Image-2's dominance might be overturned by its own house.

● 

●  This column focuses on the true levels of global AI hardware and software, leaderboards, and third-party evaluations. Views belong to the original authors.

●  This column does not constitute any investment advice. All content sources are cited; data is subject to official disclosures.

●  Physical World Frontier · Reviews | Shenzhen Physical World Frontier Technology Co., Ltd.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)