ReviewRadar

GLOBAL AI REVIEW RADAR

2026.08.11 · Tuesday

Issue 001

Key Updates 10 | Leaderboard Flash 0 | Highlights 0 | Tomorrow's Watch 0

KEY UPDATES

Update 1OpenRouter's weekly rankings show that last week (8/3-8/9), total global large AI model call volume was 69 trillion tokens

OpenRouter's weekly rankings show that last week (8/3-8/9), total global large AI model call volume was 69 trillion tokens. Chinese models hit 34.25 trillion, surpassing the US for the 15th consecutive week. DeepSeek V4 Flash took the global #1 spot with 8.83 trillion; the top four were all Chinese models. The AI most used by developers worldwide right now is made by Chinese teams.

Update 2Anthropic announced that new Claude models embed watermarks invisible to the naked eye within generated text, marking the content from start to finish

Anthropic announced that new Claude models embed watermarks invisible to the naked eye within generated text, marking the content from start to finish. In the future, you'll be able to use tools to determine if a piece of text is AI-generated, helping identify AI content online and prevent deepfake-style fraud.

What it means for you|Developers worldwide currently use Chinese-made models most, with DeepSeek V4 Flash leading global rankings.

Update 3Five universities jointly released the first evaluation of robot three-view world models, providing a unified scoring st

Five universities jointly released the first evaluation of robot three-view world models, providing a unified scoring standard for "how robots see the world," with rankings continuously updated. Finally, there's a third-party reference for evaluating which robot is smarter, rather than just manufacturer self-praise.

What it means for you|You can soon use tools to identify AI-generated text and prevent deepfake-style fraud online.

Update 4OpenRouter Weekly Call Volume DeepSeek V4 Flash 8.83 Trillion Top 4 all Chinese models; China surpasses US for 15 consec

OpenRouter Weekly Call Volume DeepSeek V4 Flash 8.83 Trillion Top 4 all Chinese models; China surpasses US for 15 consecutive weeks

What it means for you|There is now a third-party reference for evaluating which robot is smarter, not just manufacturer claims.

Update 5LMArena Text-to-Image GPT Image 2 (1340 Elo) Microsoft MAI-Image-2.6 lands suddenly at #2

LMArena Text-to-Image GPT Image 2 (1340 Elo) Microsoft MAI-Image-2.6 lands suddenly at #2

Update 6AA Intelligence Index Claude Opus 5 (63 points) Domestic open-source Kimi K3 scores 60, becoming "Open Source #1"

AA Intelligence Index Claude Opus 5 (63 points) Domestic open-source Kimi K3 scores 60, becoming "Open Source #1"

Update 7Plain Language Interpretation: Anthropic still leads in overall capability, but free and user-friendly domestic mode

Plain Language Interpretation: Anthropic still leads in overall capability, but free and user-friendly domestic models are gradually narrowing the gap. For ordinary people: If you want to experience top-tier AI, there's another no-cost option.

Update 8Microsoft MAI-Image-2.6 Lands Suddenly at #2 on Text-to-Image Leaderboard — Updated every three weeks, closely chasi

Microsoft MAI-Image-2.6 Lands Suddenly at #2 on Text-to-Image Leaderboard — Updated every three weeks, closely chasing OpenAI's GPT Image. AI art generation is shifting from a toy to a must-contest entry point for giants. Large AI Model Weekly Ranking: Alibaba qwen3.8-max Breaks into Top 6 — Multiple domestic models collectively gain ground; "chasing" is turning into "running alongside." Century-Old Riemann Hypothesis Broken by Record-High Performance from a Certain New Claude Model — Mathematical reasoning steps up again, but the model isn't public yet; treat it as a capability preview for now. 3B Small Model Om AI Edge VLX Makes a Comeback — Achieves precise perception of the physical world with small parameters; AI running on phones is getting smarter. NVIDIA Open-Sources Magpie TTS — Multilingual low-latency voice model; voice avatars are accelerating into applications. More Than Half of AI-Generated Patches Are Flawed — Research shows AI writes code fast but not necessarily correctly; human oversight remains indispensable. Research: AI Reads Text Better Than It Listens to Voice — "Can read" and "can listen" are two different things; voice AI still has a long way to go. New Anthropic Research: AI Hasn't Actually Gotten Stronger — Even those building AI are questioning whether evaluations truly measure progress. Traces of Google Gemini 3.7 Flash Exposed — Less than a month after 3.6 release, another update is coming; AI iteration speed is terrifyingly fast. Zhipu ZCode Fully Upgraded — Domestic coding AI also competes on "multi-agent collaboration," with competitive pricing.

What it means for you|If you want to experience top-tier AI, there is another no-cost option available.

Update 9DeepSeek V4 Flash tops OpenRouter with 8.83 trillion tokens in a single week; China surpasses US for 15 consecutive week

DeepSeek V4 Flash tops OpenRouter with 8.83 trillion tokens in a single week; China surpasses US for 15 consecutive weeks (East Money) LMArena Text-to-Image Leaderboard: Microsoft MAI-Image-2.6 lands suddenly at #2 (IT Home) First Robot "World Model" Evaluation Released by Five Major Universities (QbitAI) Century-Old Riemann Hypothesis Broken by Record-High Performance from a Certain New Claude Model (QbitAI) Alibaba qwen3.8-max Breaks into Top 6 Overall Rankings (IT Home) New Claude Models Embed Invisible Watermarks Throughout, Making AI Text Identifiable (QbitAI) NVIDIA Open-Sources Magpie TTS Multilingual Low-Latency Voice Model (HuggingFace) 3B Small Model Om AI Edge VLX Achieves Precise Physical World Perception with Small Parameters (QbitAI)

Update 10<img src="https://bbs.physixfrontier.com/images/emoji/twi

<img src="https://bbs.physixfrontier.com/images/emoji/twi

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)