GLOBAL AI REVIEW RADAR
2026.08.13 · Thursday
Issue 003
Key Updates 0 | Leaderboard Flash 3 | Highlights 25 | Tomorrow's Watch 6
● ① AI Agent Blind Test TOP5: Claude Fable 5 (High) 12.04 | Claude Opus 5 (High) 11.: ① AI Agent Blind Test TOP5: Claude Fable 5 (High) 12.04 | Claude Opus 5 (High) 11.96 | Claude Opus 5 (Max) 11.92 | GPT-5.6 Sol (xHigh) 10.72 | Kimi K3 (Max) 10.43. In blind tests judged by humans, Anthropic has become the "ever-victorious general," with Kimi K3 representing China; the leaderboard also tests "can it fix itself if it messes up," which is more practical than chat ability.
● ② AA "Intelligence Index" TOP5: Claude Opus 5 (max) 63 | Claude Fable 5 62 | GPT-5: ② AA "Intelligence Index" TOP5: Claude Opus 5 (max) 63 | Claude Fable 5 62 | GPT-5.6 Sol (max) 61 | Grok 4.6 (high) 61 | Kimi K3 (max) 60. Competing on overall strength vs. cost-performance, the top open-source model is still Kimi K3—the gap between open-source and closed-source has shrunk from a "generational gap" to "just a few points."
● ③ LiveBench Total Score Leaderboard (Anti-Cheating Exam Hall): Claude Fable 5 Max Effo: ③ LiveBench Total Score Leaderboard (Anti-Cheating Exam Hall): Claude Fable 5 Max Effort 83.0 | GPT-5.6 Sol Max Effort 81.0 | GPT-5.5 Thinking 80.2 | Claude 5 Opus Thinking 80.1 | Kimi K3 (Open Source) 79.2 | DeepSeek V4 Pro 0813 lands at 77.4, single-task cost $0.044, lowest in the field. "Cheap and large portion" is now written into the official report card.
●
● LiveBench is known as the "anti-cheating exam hall"—questions change every six months, preventing data leakage, making scores more reliable than typical leaderboards. On the latest leaderboard, DeepSeek V4 Pro official version landed with a score of 77.4, just 6 points behind the top tier (Claude Fable 5 at 83.0); even more impressive is the cost, averaging only 4.4 cents per successfully completed task (about 0.3 RMB), the lowest in the field. If you want to save money while getting near-top-tier AI, this number is worth remembering.
●
Programming Expert · Lao Xu says | A score of 77.4 would have sold for ten times the price half a year ago. DeepSeek's strategy hasn't changed: achieve 90% of the capability at one-tenth the price, turning "good enough" into "truly great." I said last month that call volume is the hard metric of the market voting with its feet—this landing on LiveBench writes both "cheap" and "strong" into the official report card. However, let me pour some cold water: leaderboard scores are one thing, stability when integrated into real projects is another; wait for third-party tests first.
●
Editor Xiao He says | I've used DeepSeek to write weekly reports and schedule trips; honestly, it's not much different from models costing several times more. What impressed me most was its fast response and lack of verbosity; it's still online when I'm editing drafts at 3 AM. At a price of 0.3 RMB per question, I don't even need to agonize over "should I use less today."
● Source: LiveBench (2026-06-25 version leaderboard) / IT Home
●
● If you're looking for the "most capable working AI" rather than the "best chatting AI," check the Arena agent blind test leaderboard. Human judges, brand-blind, head-to-head battles: Claude Fable 5 and Opus 5 series take the top three spots, OpenAI's GPT-5.6 Sol ranks fourth, Moonshot AI's Kimi K3 ranks fifth with 10.43 points, the highest ranking for a Chinese model. Among the six dimensions on the leaderboard, there's also "command recovery," testing whether an agent can fix itself after messing up.
●
Security Expert · Lao Zhou says | The agent leaderboard is harder to game than chat leaderboards—it tests "reliability in doing work." Kimi breaking into the top five is a solid signal that domestic models are truly catching up in the "working" dimension. The inclusion of "command recovery" as a test subject indicates the industry is starting to treat "whether AI can self-correct after errors" as a hard metric. This is also a sense of boundaries: the stronger the capability, the more we must prevent "overstepping." Permissions should still be set to the minimum necessary; don't grant excessive access just because scores are high.
●
Editor Xiao He says | I've been trying out agents these past two days to help organize meeting minutes and follow up on tasks. The experience is: it can handle 80% of the work, leaving 20% for me to cover—but at least it admits when it fails. The one or two point differences on the leaderboard might translate to just a single sentence difference in daily use.
● Source: Arena Agent Blind Test Leaderboard (Snapshot 2026-08-11, captured 8/12)
●
● Where is the limit for open-source small models? An independent evaluation team conducted 123 types of "red team" tests, totaling 391 records, on Ant Group's open-sourced Ling 3.0 Tiny (free version): jailbreak protection, sensitive content, logical consistency, instruction following... checked thoroughly like a physical exam. Some areas held up, others failed on the spot; the report clearly lists strengths and weaknesses—this is the transparency dividend of open-source models "daring to let anyone poke around."
●
Security Expert · Lao Zhou says | This kind of third-party full-body red team check-up is worth far more than vendor-bragged benchmark scores. 391 records across 123 categories, with conclusions for each category, is like showing users their underwear. I've always emphasized boundaries and guardrails—small models aren't as powerful, so they're actually more prone to failing at the edges. An open-source model daring to be tested this way is itself an attitude; the weaknesses exposed in the report are exactly the "manual" users should know before using it.
●
Editor Xiao He says | For the first time, I feel I can understand an AI's "physical exam report": which items held up, which failed on the spot, all clear. This approach of putting flaws on display actually makes me more willing to use it—at least I know where its bottom line is.
● Source: lateos.ai Red Team Report / Hacker News
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
● After the hype around DeepSeek V4 Pro's official release, will it continue to refresh the various sub-leaderboards on LiveBench? How long can the "cost-performance throne" of 77.4 points last?
● The agent blind test leaderboard adds a "command recovery" dimension; can domestic models leverage this question to advance further and break into the top three?
● Earnings season enters deep waters: After Cerebras plunges 14%, will the market reassess the valuation system for AI chip companies?
●
● Frontier Review · Global AI Evaluation Radar|Data sourced from public leaderboards and reports, subject to official disclosures
● Frontier Review | Shenzhen Frontier Technology Co., Ltd.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.