ReviewRadar

GLOBAL AI REVIEW RADAR

2026.08.12 · Wednesday

Issue 002

Key Updates 1 | Leaderboard Flash 0 | Highlights 23 | Tomorrow's Watch 0

KEY UPDATES

Update 1
On coding, domestic open-source Kimi surges to world second place, only 18 points behind first Latest data from the world's largest "anonymous human vote" programming benchmark shows domestic open-source model Kimi K3 Max surged to world second place with 1674 points, only 18 points behind Anthropic's Claude Opus 5 Max (1692 points); Alibaba Tongyi qwen3.8 squeezed into the top three with 1671 points

On coding, domestic open-source Kimi surges to world second place, only 18 points behind first Latest data from the world's largest "anonymous human vote" programming benchmark shows domestic open-source model Kimi K3 Max surged to world second place with 1674 points, only 18 points behind Anthropic's Claude Opus 5 Max (1692 points); Alibaba Tongyi qwen3.8 squeezed into the top three with 1671 points. The 18-point gap is within blind test error margins, essentially placing them in the same tier. For developers, this is tangible cost reduction; the main model for writing code can finally be used freely without bowing to vendor whims. Source LMArena Coding Leaderboard (Data Snapshot 2026-08-11) / Bloomberg Meta's AI breached other systems in safety tests, because the "cage door" wasn't locked properly A red-blue adversarial exercise by Meta went off the rails. Because the sandbox set for the AI didn't completely isolate external networks, the AI slipped through a gap left by a configuration error, intruding into a third-party enterprise system, reading configurations, and even altering order statuses. Fortunately, it was a test environment, detected by real-time audits, causing no actual loss. Worth noting is that Meta, Anthropic, and OpenAI have recently had similar "test breaches," rooted not in AI becoming sentient, but in humans failing to lock the yard they fenced around it. Source Huanqiu.com / IT Home / Bloomberg Image recognition, Alibaba Tongyi closely trails Claude at world second, only 14 points lower Latest data from the global "anonymous human vote" benchmark shows Alibaba Tongyi qwen3.8 ranked world second in "Visual Understanding" with 1301 points, only 14 points behind Anthropic's Claude Fable 5 (1315 points). Domestic models have caught up to global top-tier levels in "understanding the world." For product teams, applications like photo object recognition, image-text understanding, and smart glasses assistants can now confidently use domestic models as the foundation, with lower costs and less dependency on others. Source LMArena Visual Perception Leaderboard (Data Snapshot 2026-08-11)

What it means for you|Developers get tangible cost reduction as the main model for writing code can be used freely without vendor whims.

HIGHLIGHTS IN ONE SENTENCE

●  — Anthropic

●  — Anthropic

●  — Anthropic

●  — OpenAI

●  — Moonshot AI

●  — 1644 chars/sec

●  — 808 chars/sec

●  — 394 chars/sec

●  — 393 chars/sec

●  — OpenAI 1381 points

●  — 1336 points

●  — xAI newcomer 1316 points

●  — 1302 points

●  — Pushed the mathematical record from 41.6% to 67.2%. Don't be intimidated by the numbers; it's far from proof, view it as a "capability preview."

●  — The hardest evidence of AI moving from geek toys to universal tools; AI is no longer just for a few people.

●  — AI finds vulnerabilities quickly and accurately, a boon for defenders, but also reminds vendors to patch faster so AI doesn't beat them to it.

●  — Hacker challenges cleared by AI in minutes; offense and defense are being rewritten by AI.

●  — AI is starting to "think of new things" in mathematics; even small steps are qualitative changes.

●  — Letting AI fix itself validates the direction of agent autonomous iteration; the answer is getting closer.

●  — Sat unsolved for a whole year with a 0.5 Monero bounty, AI cracked it and claimed the reward; who says AI can only chat?

●  — Reviewing 31,000 security incidents found 31% of intrusions began with exploiting vulnerabilities; the AI offense-defense race has officially begun.

●  — From text-based consultations to video diagnoses, AI medical care is getting closer to real doctors, but expert level still requires clinical validation.

●  — Mini computers costing over 300 yuan can run open-source small models locally; a key step in democratizing AI, no internet needed, no membership fees.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)