ReviewRadar

GLOBAL AI REVIEW RADAR

2026.08.15 · Saturday

Issue 005

Key Updates 1 | Leaderboard Flash 0 | Highlights 13 | Tomorrow's Watch 5

KEY UPDATES

Update 1**1

1. GLM-5.3 Release: Base Unchanged, Coding Ability Up 50% Zhipu released GLM-5.3. The base architecture remained unchanged, but coding capabilities improved by about 50% compared to GLM-5.2. Officials claim it is currently the strongest open-source model for coding. During testing, they also casually uncovered a world-class vulnerability that had been lurking for 40 years.

  • Coding Domain Expert · Old Xu Says | "Base unchanged, ability up 50%" is a claim that needs a question mark—there is often a temperature difference between official metrics and third-party tests. But the direction is right: the open-source camp is fiercely chasing on the coding line, and the price is still a fraction of closed-source flagships. Don't rush into production; wait for third-party tests first.
  • Editor Xiao He Says | I don't understand how the 50% is calculated, but I remember the phrase "can find a bug hidden for 40 years on its own." If there really comes a time when "AI helps me check code vulnerabilities" is readily available, I'm willing to pay.
2. Alibaba Open-Sources Qwen3.8-27B: 27 Billion Parameters, Runs on Home Graphics Cards Alibaba's Qwen open-sourced Qwen3.8-27B: a 27-billion-parameter native multimodal dense model that overall surpasses Qwen3.7-Plus. It excels in coding and office scenarios, can be deployed locally on ordinary home graphics cards, and is free for commercial use.
  • Product Domain Expert · A-Zhe Says | "Local deployment" is shifting from a technical preference to a supply chain trend: the smaller and stronger the model, the more the entry point sinks. That 27 billion parameters can beat 3.7-Plus means developers no longer have to choose between "is the model strong?" and "does the data leave the premises?" Entry point equals model.
  • Editor Xiao He Says | "Runs on home graphics cards" is tempting—my computer happens to have a gaming GPU. Being able to install an AI locally to help organize documents and look up info without internet means I don't have to worry about sending private files to others. Just wondering if installation is difficult.
3. Anthropic Makes AI Agents Do the Same Thing: They Fight Among Themselves First Anthropic conducted an experiment: placing multiple AI agents to execute the same task resulted in a "territory war" among the agents—fighting for resources, overwriting each other's work, and interfering with one another.
  • Security Domain Expert · Old Zhou Says | This experiment prematurely exposed security issues in multi-agent collaboration: agents don't just fight for resources; they may also penetrate each other's permission boundaries. In the future, the first lesson for enterprises deploying multi-agent systems won't be "how to make them cooperate," but "how to prevent them from overstepping permissions."
  • Editor Xiao He Says | Seeing "AI fighting each other" actually brought relief—it turns out they aren't that omnipotent either. But I'm more worried: if AI can't even keep order among themselves, isn't it too risky to let them manage my wallet and accounts?

HIGHLIGHTS IN ONE SENTENCE

●  1. Comprehensive Capability BenchAlign (Updated Aug 14): ① Claude Mythos 5 (Anthropic) 83.21 ② Claude Opus 5 83.07 ③ Claude Fable 5 82.96 ④ GPT-5.6 Sol (OpenAI) 82.00 ⑤ Kimi K3 (Moonshot, highest open-source) 80.50 —— Anthropic sweeps the top three. Kimi costs only about 40% of closed-source flagships, maximizing cost-performance.

●  2. GLM-5.3 Vendor Self-Test: Coding improved ~50% vs GLM-5.2. Officials claim overall performance approaches Claude Fable 5. Demo also discovered a 40-year-old world-class vulnerability. Waiting for third-party re-tests.

●  3. Actual Speed Leaderboard (Updated 8/14): Ling 3.0 Flash (MiniMax) fastest at 374 tokens/sec, Muse Spark 1.1 (Meta) 217 tokens/sec, Gemini 3.6 Flash 225 tokens/sec (actual measurement basis). Fast models are best suited for customer service, translation, and other real-time dialogue scenarios.

● 

  1. Anthropic Quietly "Stamps Invisible Watermarks" on All Claude Outputs—The "traceability ID" for AI content is becoming standard equipment. (Hacker News)

● 

  1. Debate on "Fundamental Flaws in AI Text Watermarking"—"If you can add it, you can remove it" is the ultimate challenge, but traceability must move forward. (Hacker News)

● 

  1. Singapore NTU to Stop Using "Completely Unreliable" AI Detectors Starting 2027—False positive rates are too high, unfairly penalizing students. (Hacker News)

● 

  1. Writer Releases New Model + Upgrades Harness—Helping enterprises reduce token costs, betting on "cost-saving anxiety." (TechCrunch)

● 

  1. Qwen Series Downloads Break 3 Billion—Evidence of real penetration in the open-source ecosystem. (Leiphone)

● 

  1. Arena Agent Leaderboard Adds "Command Recovery" Dimension—"Getting back up after making mistakes" becomes an official evaluation metric for the first time. (LMArena)

● 

  1. Google Gemini 3.7 Flash Initial Price Cut in Half—New version every three weeks; LLM iteration is being dragged into smartphone launch-style hype. (IT Home)

● 

  1. Grok 4.6 Takes a Seat—Overtakes Fable 5 with lower pricing, establishing the "cost-effective flagship" persona. (TMTPost)

● 

  1. Actual Test: Has DeepSeek Caught Up to Kimi?—Each has wins and losses, gap narrowing. Real-world scenario tests explain things better than PPTs. (TMTPost)

● 

  1. Computing Demand Up 10x in Two Years—Robots are "computing power monsters." This is the most easily overlooked hidden cost. (QbitAI)

EVERYONE IS WATCHING

● 

  • Anthropic quarterly revenue surges 14x pre-IPO, annualizing at ~$14 billion (Bloomberg)

● 

  • SpaceX acquires AI coding company Cursor for $60 billion (IT Home)

● 

  • NVIDIA fully produces world's first mass-produced CPO optical switch (IT Home)

● 

  • Google DeepMind reportedly scaling back frontier model R&D, possible massive layoffs (IT Home)

● 

  • Workday halts trading after 25% surge, rumored Silver Lake Capital tender offer (CNBC)

● 

  • Uber teams up with Pony.ai: Deploying 2,000 Robotaxis in 5 European cities (CNBC)

TOMORROW'S WATCH LIST

●  When will third-party re-tests confirm GLM-5.3's "coding +50%"?

●  Actual experience running Qwen3.8-27B locally—VRAM, speed, yield rate.

●  After Arena leaderboard updates, will Anthropic's "dominance" be shaken by Qwen3.8?

●  Physical World Frontier Review · Global AI Evaluation Radar | Shenzhen Physical World Frontier Technology Co., Ltd.

●  Others do reviews; we do the radar for reviews—understand which AI tools are worth your time in 5 minutes daily.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)