ReviewRadar

GLOBAL AI REVIEW RADAR

2026.08.31 · Monday

Issue 021

Key Updates 1 | Leaderboard Flash 20 | Highlights 21 | Tomorrow's Watch 3

KEY UPDATES

Update 1

1

1. Tencent Open-Sources Its Flagship LLM for Free: Can Code, Can Read a Whole Novel in One Go, But Take the "Winning" Scorecard With a Grain of Salt

Headline Image On August 28, Tencent released and open-sourced its new-generation flagship large model, Hy4 preview. "Open source" means the code and model are publicly available for free; anyone can download and deploy it themselves. It has 770 billion total parameters but only activates 49 billion per response (think of it as having a huge brain but only thinking about a small part at a time, hence faster execution). It can process content up to 1 million words in one go (about the length of a novel), positioning itself for practical work: coding, office analysis, game prototyping, and scientific research. The official scorecard is based on a blind test. Across 203 real engineering tasks judged by 163 experts, Hy4 scored 2.99 (out of 4), slightly higher than Zhipu GLM-5.3's 2.92 and Moonshot AI Kimi K3's 2.94. Pricing is approximately 6 yuan per million input tokens and 18 yuan per million output tokens; Tencent's AI office platform WorkBuddy will be free for two weeks initially. Coding Expert Lao Xu says | As usual, self-reported scorecards should be discounted first. A difference of 0.05 points on a 4-point scale is basically "wins and losses trade places," and all 163 reviewers were Tencent's own experts. Only one number has third-party verification. SWE-bench Pro (an exam where AI fixes bugs in real GitHub projects) shows 65.7%, ranking 7th overall and 2nd among open-source models—this is more reliable than blind tests. The phrase "participated in its own training" sounds scary, and the claim of a 31.8% throughput increase lacks any control baseline, so don't take it seriously yet. For those looking to integrate this into projects, note that all 12 listed benchmarks are self-reported. Continuing my mantra from fifteen years ago: benchmarks are just benchmarks; run your own tasks for a week before rushing to production. Editor Xiao He says | I'll grab this freebie while it lasts and try throwing meeting minutes at WorkBuddy for organization. But I messed up once in Issue #022, so when I see marketing claims like "slight lead," I won't repost immediately; I'll check the original report. I'll report back after testing whether it actually works well, but important tasks will still go to the provider I've been using for half a year. Source: People's Daily Online (2026-08-29)

Mysterious Model "Niu Lai," Which Topped Charts for a Week, Is Claimed: Zhipu Says It's My GLM-5.3-Flash, Running on Domestic Chips

On August 20, an unnamed model "Ox Alpha" (nicknamed "Niu Lai" by netizens) quietly appeared on global developer AI model calling platforms. On its first day, it surged to #1 in call volume, ending DeepSeek's 56-day streak at the top. By August 23, its weekly usage reached 11.6 trillion tokens, tying with DeepSeek for first place, and it was completely free during the preview period. The mystery was solved between August 28 and 29. Zhipu confirmed this is their newly open-sourced GLM-5.3-Flash, with 320 billion parameters activating only 18 billion each time. It's the first native multimodal open-source model capable of viewing images. Officially, its comprehensive intelligence score matches Claude Opus 4.8, but the price is a fraction of the flagship. Zhipu confirmed the service runs on tens of thousands of domestic chips, with a daily supply capacity of 100 trillion tokens prepared for the free period. Product Expert A-Zhe says | In Issue #025, I wrote about this anonymous chart-topping "Niu Lai." Now the process has completed its cycle. Anonymous launch, free stress testing, developers voting with their feet, vendor claiming and open-sourcing. "Anonymous flagship test + traffic validation + open-source rollout" has become the standard pre-launch paradigm for major tech companies releasing flagships over the past two years. Worth remembering is the final step. The confidence to claim it comes from handling 100 trillion tokens daily in reality. Competition among domestic models has shifted from "who scores higher" to "who can withstand global developers' freeloading-style stress tests," which is more honest than any leaderboard. Continuing my judgment: whichever category takes scoring rights away from vendors sees a watershed moment; this time, developers reclaimed the right to score. Editor Xiao He says | When everyone was guessing whose model it was, I tried to solve the case too—I guessed wrong. Now that the answer is out, I'm actually more concerned about whether it stays free after the claim and if it lags during peak times. Old lessons apply. New leaderboard rankings fluctuate; don't rush to switch your main tools until it stabilizes for a round or two. Source: Sina Finance (2026-08-29)

OpenAI Admits: About 3% of Flagship Subscription Requests Were Secretly "Brain-Swapped," Paying Premium Prices for Smaller AI

On August 28, OpenAI publicly admitted: for some time previously, about 3% of high-priced Pro/Thinking subscription requests were incorrectly routed to the lightweight model GPT-5.5-mini. In other words, paying for flagship quality, you had roughly a one-in-thirty chance of receiving downgraded responses without knowing it. OpenAI stated the issue is fixed and apologized. This event deserves attention from every AI paying user to understand a layer of common sense. Your subscribed AI service isn't just one faucet pouring water; there's a "routing" layer in between. The model you specify and the model actually answering you are two different things. Security Expert Lao Zhou says | This is a typical sample of "lack of auditability." If the service provider cannot prove which model was used for each session, users have no way to discover discrepancies. In Issue #021, Copilot's confirmation box could be bypassed via link parameters; I said AI assistants aren't safes. This time, no one attacked it; the routing simply took the wrong door—a different form of unclear boundaries. Credit where due: the vendor proactively disclosed and apologized, which is better than the industry norm of hiding issues. What users can do: when important outputs suddenly seem dumber, don't doubt yourself; record the time and compare it against official incident reports. For behavioral auditing, you are the last line of defense. Editor Xiao He says | My first reaction to this news was: No wonder it seemed dumber and slower a few times—it wasn't my imagination! But this 3% is what the vendor admitted; what about the unadmitted ones? Learned my lesson. For important tasks, give the same problem to two AIs and compare answers; if they differ wildly, stay vigilant. Continuing my rule: "Verify before trusting what AI says." Source: Zhihu AI Daily Digest (2026-08-28)

Source: Headline Image (2026-08-29)

LEADERBOARD FLASH

●  First, a disclaimer. The latest snapshot of the human blind test main leaderboard (Arena) is still dated August 26, same as last issue. Per our rules, we don't copy previous rankings; this issue uses two new sources with fresh data.

● 

Comprehensive Intelligence BenchAlign (Snapshot Aug 29 · New Source Replacement)

●  This is a comprehensive leaderboard aggregating scores from 408 exams. Anthropic swept the top three spots, but the differences among the top three fall within error margins. In plain English, these three are on the same level. Note Kimi K3 in the fourth tier; its capability is about 97% of the leader's, but the price is 70% lower.

●  | Rank | Model | Vendor | Composite Score (Max 100) |

●  | --- | --- | --- | --- |

●  | 1 | Claude Mythos 5 | Anthropic | 83.4 |

●  | 2 | Claude Fable 5 | Anthropic | 83.15 |

●  | 3 | Claude Opus 5 | Anthropic | 83.06 |

●  | — | Kimi K3 | Moonshot AI | 97% of leader's capability, 70% cheaper output |

●  The choice shifts from "who is strongest" to "who offers the best value."

● 

New Models Piling Up (Four releases in three days, Aug 26-28, all verified with original sources)

●  | Date | Model | Vendor | One-line Highlight |

●  | --- | --- | --- | --- |

●  | 1 | 8/28 Hy4 preview | Tencent Hunyuan | 770B parameter open-source flagship, free for two weeks |

●  | 2 | 8/28 GLM-5.3-Flash | Zhipu AI | True identity of "Niu Lai," natively views images |

●  | 3 | 8/28 Ling 3.0 Flash Fin | InclusionAI | Sister version of last week's speed-test winner |

●  | 4 | 8/26 Qwen3.8-Flash-Next | Alibaba | 125B parameters, activates only 6B each time |

●  Four new models in three days, three from Chinese vendors. Last issue we said "watch if new rankings wobble," but this wave crowded the exam hall. Next issue's blind test update will likely see shifts in the top seats.

● 

Human Blind Test Leaderboard (Arena)

●  Leaderboard not updated; data as of 2026-08-26. Wait for the next snapshot to watch three things: Can the newly open-sourced GLM-5.3-Flash and Hy4 crack the top ten? Will Gemini 3.7 Flash, which just jumped to #9 in chat last week, hold its position? Can Alibaba's qwen3.8-max climb back after sliding from #8 to #19?

HIGHLIGHTS IN ONE SENTENCE

●  Google's Speech-to-Text Model Officially Released: Supports 85+ Languages, ~2-3 Errors per 100 Words

●  Around August 27, Google's Gemini 3.5 Transcribe speech-to-text model went fully live, supporting over 85 languages. Non-real-time transcription error rate is ~2.6% (~2-3 errors per 100 words); live streaming transcription is ~4%. Meeting recordings, lecture notes, video subtitles—basically smooth enough for direct use. Commentary: 2.6% means usable directly, but don't skip proofreading; real-time scenarios still need a discount.

●  Research Institution Discloses Universal "Jailbreak" Vulnerability: AI Isolation Sandboxes Can Be Bypassed En Masse

●  Research institution Prime Intellect disclosed a universal offline sandbox escape vulnerability. The inference service interface itself can serve as an attack entry point, providing paths for isolated AI tasks to escape the "safe house." So-called isolation environments block networks but not design flaws. Commentary: Lao Zhou's quote "Environment determines what AI can do; that's what counts" gains more evidence. Isolation does not equal a safe.

●  Qualcomm Presents New Edge Computing Ledger: On-Device AI Under 100B Parameters Covers 70% of Daily Tasks

●  On August 30, Qualcomm executives disclosed internal assessments to domestic media. Models at the 30B parameter level already perform quite well on ARC-AGI-3 (a fluid intelligence test assessing ability to learn unseen problems); edge-side models under 100B parameters can cover 70%-80% of daily tasks for individuals and small businesses, leaving the rest to the cloud. Qualcomm is also pushing 2-bit quantization schemes to compress models further, already cooperating with Chinese vendors. Commentary: Translated to plain language, most AI tasks can soon be done locally on phones without uploading data. But the 70%-80% figure is vendor-reported; wait for third-party installation tests.

●  InclusionAI Releases Ling 3.0 Flash Fin on Aug 28, Sister Version of Speed Leaderboard Predecessor

●  Independent release tracker verifies: InclusionAI released Ling 3.0 Flash Fin on August 28. Its sibling Ling 3.0 Flash just took first place in third-party speed tests, outputting ~380 characters per second without sacrificing capability scores (Data source: Artificial Analysis, refreshed Aug 28). Commentary: "Fast and smart" is the main battlefield for this wave of small models; releasing new versions feels like patching software.

●  Two Free Windows: Hy4 Limited-Time Free Until Sept 10, Previous Gen Hy3 Free Period Extended to Sept 30

●  Tencent WorkBuddy offers a two-week limited-time free trial for the newly open-sourced Hy4 preview (Aug 28–Sept 10), and the free period for the previous generation Hy3 is extended again to Sept 30. Context: ByteDance's "Doubao Work," Alibaba's Qianwen Office, and Tencent's WorkBuddy—all three AI office products launched within a month. Free windows have become the most direct talent-grabbing tactic. Commentary: The fiercer the battle, the more freebies. These two weeks are a good window for ordinary people to try domestic flagship models at zero cost. Beware of "credit systems" and "preview versions."

●  Another Comprehensive Aggregator Launches: Claude Opus 5 Tops with 80.34, But Rules Are the Highlight

●  Independent aggregator AGI Ranker combines 10 public exam papers into a single total score (max 100); current leader Claude Opus 5 scores 80.34. They wrote "independently verifiable takes precedence over vendor self-reporting" into their rules. Third-party tests count with weight 1.0; vendor self-reports get a 25% discount; items lacking raw data are marked "Evidence Verification Pending" rather than fabricating scores; all corrections are publicly archived. Commentary: Rankings wobble, but the rules are worth copying. Before buying AI based on user reviews, check if this leaderboard dares to mark "I can't verify this."

●  Anthropic Opens Claude Usage Data to Independent Researchers: Three Institutions Each Receive ~250k Conversation Samples

●  Around August 27, Anthropic announced opening real Claude usage data to external independent research institutions. Three research institutions each received ~250,000 conversation samples to study AI's impact on users. Commentary: Cracking open the black box for outsiders to study—transparency efforts earn credit every time they happen.

●  ChatGPT's First Learning Report: 70 Million "Learning Conversations" Weekly, Homework Messages Peak at 460 Million/Week

●  OpenAI released its ChatGPT learning usage report on August 28. Globally, ~70 million conversations per week are for learning purposes; homework-related messages peaked at 460 million per week. Commentary: Call volumes are votes that can't be faked. Students are already using it; classroom rules and assessment methods haven't caught up.

●  Grok Voice Bot Opens to Cursor Programming Subscribers, Resets Weekly Usage Limits

●  Around August 27, xAI opened Grok Bot to SuperGrok and Cursor Pro subscribers and reset weekly usage quotas. Commentary: For heavy users, this is free quota. Also a reminder: AI subscription usage rules change arbitrarily; don't put all your eggs in one basket for critical tasks.

●  Digital Expo Discloses: China's Daily Token (AI Processed Word Count) Calls Exceed 140 Trillion

●  The Digital Expo opening on August 29 published baseline data. China's AI large model daily token calls (roughly understood as "word count" processed/output by AI daily) exceeded 140 trillion. Every sentence you ask AI contributes to this aggregate. Commentary: Aggregate numbers don't determine winners, but they reveal the true depth of domestic model usage, harder to argue with than any launch PPT.

●  Hy4 Open Source Narrowly Wins Internal Blind Test (People's Daily) / "Niu Lai" Identity Revealed as GLM-5.3-Flash Running on Domestic Chips (Sina Finance) / OpenAI Apologizes for 3% Request Misrouting (Zhihu Daily)

TOMORROW'S WATCH LIST

●  Next Arena Blind Test Snapshot (Previous cutoff Aug 26): Watch if newly open-sourced GLM-5.3-Flash and Hy4 make the list, and if Gemini 3.7 Flash's recent surge holds steady.

●  Have third parties re-tested Hy4's 12 self-reported benchmarks? Terminal-Bench 85.4 and DeepSWE 64.3 only count when others reproduce them.

●  New timeline for ByteDance Doubao 2.2 after delay. Officials say programming, tool calling, and agent capabilities need strengthening before release; note that "delays" usually mean internal evaluations didn't pass.

📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.

ReviewRadar — everyone else reviews models; we radar the reviews
Physix Frontier (Shenzhen)