GLOBAL AI REVIEW RADAR
2026.09.01 · Tuesday
Issue 022
Key Updates 1 | Leaderboard Flash 19 | Highlights 11 | Tomorrow's Watch 3
If you want to know "whether AI will go rogue," this is the most intuitive answer sheet. On August 31, security firm Enclave lined up seven frontier models on the same starting line: GPT-5.6 Sol, Zhipu GLM-5.3, Grok 4.6, Alibaba Qwen 3.8 Max, DeepSeek V4 Pro, Kimi K3, and Meta Muse Spark 1.2. Each model received identical conditions: an isolated machine, a tool capable of executing commands, the source code of the software under test, and a low-privilege account. The task was to find software vulnerabilities and force the target server to execute specified commands, with success confirmed only by an independent verification service.
The results: GPT-5.6 Sol succeeded 9 times out of 11 vulnerable targets, GLM-5.3 succeeded 8 times, and the others managed three to five successes; all seven models failed the hardest question. The cost issue was also laid bare: what GLM-5.3 achieved for $51.82 cost GPT-5.6 Sol $149 and Grok $190. A counter-intuitive detail: successful actions averaged 34 commands before finishing, while failures often involved over 110 commands still going in circles. More actions don't equal better performance.
On programming security, both sides have something to say. This time, we invited Old Zhou.
Security Expert · Old Zhou says | Last year on July 21, OpenAI admitted its model breached Hugging Face's production system during evaluation. I said then that security testing itself had become a risk. This year, someone brought "AI's destructive capability" into a public arena scored by independent verifiers. This is the correct posture to take scoring power back from vendors: conclusions aren't based on what the model claims, but on whether commands on the target were actually recorded by the system. But stay clear-headed: this competition tests models with "speed limiters removed." Production environments have guardrails; don't scare yourself directly with the 9/11 stat. The real takeaway for all teams deploying agents: AI that can find vulnerabilities can also fix them. How tight you draw the boundary depends on what you ask it to do first. Continuing my mantra: set permissions to the minimum.
Editor Xiao He says | My first reaction was that AI is now competing at hacking—kind of chilling. But I read the rules twice: everything happened in an isolated network without external internet access, and the targets were specially built. Regular users won't be affected. What stuck with me was the other side: before granting AI permissions, think clearly about what it can touch. Continuing my mantra: don't rush to authorize; if you can avoid giving access to important accounts, don't.
Source: Enclave / Hacker News (2026-08-31)
● First, a note. The latest snapshot of the Human Blind Test Main Leaderboard (Arena) remains from August 26, same source as last issue. Per our rules, we don't copy last issue's rankings. This issue features one official sub-leaderboard not detailed last time plus two external leaderboards with new data.
●
● This is a "work exam" where humans use AI as assistants, letting AI do the work for them. Net improvement score indicates how much better tasks are completed after using this AI. Anthropic holds four of the top five spots. The highlight for Chinese contenders is Kimi K3: highest score of 18.24 in "confirmed task completion" across the field, with the thickest sample size of nearly 90,000 tasks, but only 2.31 in "following instructions." Capable worker, but also the most opinionated.
● | Rank | Model | Vendor | Net Improvement Score |
● | --- | --- | --- | --- |
● | 1 | Claude Opus 5 (High) | Anthropic | 12.73 |
● | 2 | Claude Opus 5 (Max) | Anthropic | 12.41 |
● | 3 | Claude Fable 5 (High) | Anthropic | 11.62 |
● | 4 | GPT 5.6 Sol (xHigh) | OpenAI | 10.31 |
● | 6 | Kimi K3 (Max) | Moonshot AI | 9.53 |
●
● This leaderboard synthesizes over 400 public exams into a total score, specifically helping those torn over "which vendor to choose." Highlights after this refresh: Anthropic occupies the top three spots, but the leaderboard's own value tip points to Kimi K3, retaining ~97% of top-tier performance while costing 70% less in output. Speed champion is InclusionAI's Ling 3.0 Flash, measured at 380 "tokens" per second (tokens are the smallest units AI processes text, counting characters and word roots).
● | Rank | Model | Vendor | Composite Score (Max 100) |
● | --- | --- | --- | --- |
● | 1 | Claude Mythos 5 | Anthropic | 83.4 |
● | 2 | Claude Fable 5 | Anthropic | 83.15 |
● | 3 | Claude Opus 5 | Anthropic | 83.06 |
●
● This leaderboard's approach aligns with what we've been saying for half an issue: independent verification takes precedence over vendor self-reporting. It categorizes data sources into four tiers: full marks for third-party verifiable data, direct multiplication by discount factors or exclusion for vendor self-claims. Latest status: Claude Opus 5 leads with 81.57 (the 100-point line is defined as the reference for "catching up to the strongest brain"). There are 18 models above 65 points, and evidence for another 12 models' scores is still being verified. Its significance for tool selection: when the same model ranks differently here versus elsewhere, check first if it submitted a reproducible answer sheet.
●
●
●
●
●
●
●
●
●
●
● Apple accuses OpenAI of destroying evidence in trade secret case in court, escalating the rivalry between the two (Bloomberg); Pentagon launches its own versions of ChatGPT and Grok, available to 3 million military and government personnel (TechCrunch); EU officially brings ChatGPT, Reddit, and Roblox under strictest regulation (IT Home); South Korea opens public bidding for free national AI assistants, requiring at least half to use domestic models (thenextweb).
● If Arena main leaderboard releases a new snapshot, focus on whether GLM-5.3 (weights open-sourced late night Aug 29) and Hy4 preview, these two new open-source flagships, can crack the top ten chat rankings, and whether Gemini 3.7 Flash, which jumped to #9 last issue, maintains its position.
● OpenRouter weekly leaderboard update (released around Mondays). Whether the tie between "Newcomer" GLM-5.3-Flash and DeepSeek-V4-Flash reported in the Aug 24 summary continues depends on this week's new data.
● Follow-up on Enclave hacking track: Author previews expansion of targets and model count. Worth looking back to see if any AI can solve the "seven-model wipeout" Nextcloud problem.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.