GLOBAL AI REVIEW RADAR
2026.08.31 · Monday
Issue 021
Key Updates 1 | Leaderboard Flash 20 | Highlights 21 | Tomorrow's Watch 3
On August 28, Tencent released and open-sourced its new-generation flagship large model, Hy4 preview. "Open source" means the code and model are publicly available for free; anyone can download and deploy it themselves. It has 770 billion total parameters but only activates 49 billion per response (think of it as having a huge brain but only thinking about a small part at a time, hence faster execution). It can process content up to 1 million words in one go (about the length of a novel), positioning itself for practical work: coding, office analysis, game prototyping, and scientific research. The official scorecard is based on a blind test. Across 203 real engineering tasks judged by 163 experts, Hy4 scored 2.99 (out of 4), slightly higher than Zhipu GLM-5.3's 2.92 and Moonshot AI Kimi K3's 2.94. Pricing is approximately 6 yuan per million input tokens and 18 yuan per million output tokens; Tencent's AI office platform WorkBuddy will be free for two weeks initially.
Coding Expert Lao Xu says | As usual, self-reported scorecards should be discounted first. A difference of 0.05 points on a 4-point scale is basically "wins and losses trade places," and all 163 reviewers were Tencent's own experts. Only one number has third-party verification. SWE-bench Pro (an exam where AI fixes bugs in real GitHub projects) shows 65.7%, ranking 7th overall and 2nd among open-source models—this is more reliable than blind tests. The phrase "participated in its own training" sounds scary, and the claim of a 31.8% throughput increase lacks any control baseline, so don't take it seriously yet. For those looking to integrate this into projects, note that all 12 listed benchmarks are self-reported. Continuing my mantra from fifteen years ago: benchmarks are just benchmarks; run your own tasks for a week before rushing to production.
Editor Xiao He says | I'll grab this freebie while it lasts and try throwing meeting minutes at WorkBuddy for organization. But I messed up once in Issue #022, so when I see marketing claims like "slight lead," I won't repost immediately; I'll check the original report. I'll report back after testing whether it actually works well, but important tasks will still go to the provider I've been using for half a year.
Source: People's Daily Online (2026-08-29)
● First, a disclaimer. The latest snapshot of the human blind test main leaderboard (Arena) is still dated August 26, same as last issue. Per our rules, we don't copy previous rankings; this issue uses two new sources with fresh data.
●
● This is a comprehensive leaderboard aggregating scores from 408 exams. Anthropic swept the top three spots, but the differences among the top three fall within error margins. In plain English, these three are on the same level. Note Kimi K3 in the fourth tier; its capability is about 97% of the leader's, but the price is 70% lower.
● | Rank | Model | Vendor | Composite Score (Max 100) |
● | --- | --- | --- | --- |
● | 1 | Claude Mythos 5 | Anthropic | 83.4 |
● | 2 | Claude Fable 5 | Anthropic | 83.15 |
● | 3 | Claude Opus 5 | Anthropic | 83.06 |
● | — | Kimi K3 | Moonshot AI | 97% of leader's capability, 70% cheaper output |
● The choice shifts from "who is strongest" to "who offers the best value."
●
● | Date | Model | Vendor | One-line Highlight |
● | --- | --- | --- | --- |
● | 1 | 8/28 Hy4 preview | Tencent Hunyuan | 770B parameter open-source flagship, free for two weeks |
● | 2 | 8/28 GLM-5.3-Flash | Zhipu AI | True identity of "Niu Lai," natively views images |
● | 3 | 8/28 Ling 3.0 Flash Fin | InclusionAI | Sister version of last week's speed-test winner |
● | 4 | 8/26 Qwen3.8-Flash-Next | Alibaba | 125B parameters, activates only 6B each time |
● Four new models in three days, three from Chinese vendors. Last issue we said "watch if new rankings wobble," but this wave crowded the exam hall. Next issue's blind test update will likely see shifts in the top seats.
●
● Leaderboard not updated; data as of 2026-08-26. Wait for the next snapshot to watch three things: Can the newly open-sourced GLM-5.3-Flash and Hy4 crack the top ten? Will Gemini 3.7 Flash, which just jumped to #9 in chat last week, hold its position? Can Alibaba's qwen3.8-max climb back after sliding from #8 to #19?
● Google's Speech-to-Text Model Officially Released: Supports 85+ Languages, ~2-3 Errors per 100 Words
● Around August 27, Google's Gemini 3.5 Transcribe speech-to-text model went fully live, supporting over 85 languages. Non-real-time transcription error rate is ~2.6% (~2-3 errors per 100 words); live streaming transcription is ~4%. Meeting recordings, lecture notes, video subtitles—basically smooth enough for direct use. Commentary: 2.6% means usable directly, but don't skip proofreading; real-time scenarios still need a discount.
● Research Institution Discloses Universal "Jailbreak" Vulnerability: AI Isolation Sandboxes Can Be Bypassed En Masse
● Research institution Prime Intellect disclosed a universal offline sandbox escape vulnerability. The inference service interface itself can serve as an attack entry point, providing paths for isolated AI tasks to escape the "safe house." So-called isolation environments block networks but not design flaws. Commentary: Lao Zhou's quote "Environment determines what AI can do; that's what counts" gains more evidence. Isolation does not equal a safe.
● Qualcomm Presents New Edge Computing Ledger: On-Device AI Under 100B Parameters Covers 70% of Daily Tasks
● On August 30, Qualcomm executives disclosed internal assessments to domestic media. Models at the 30B parameter level already perform quite well on ARC-AGI-3 (a fluid intelligence test assessing ability to learn unseen problems); edge-side models under 100B parameters can cover 70%-80% of daily tasks for individuals and small businesses, leaving the rest to the cloud. Qualcomm is also pushing 2-bit quantization schemes to compress models further, already cooperating with Chinese vendors. Commentary: Translated to plain language, most AI tasks can soon be done locally on phones without uploading data. But the 70%-80% figure is vendor-reported; wait for third-party installation tests.
● InclusionAI Releases Ling 3.0 Flash Fin on Aug 28, Sister Version of Speed Leaderboard Predecessor
● Independent release tracker verifies: InclusionAI released Ling 3.0 Flash Fin on August 28. Its sibling Ling 3.0 Flash just took first place in third-party speed tests, outputting ~380 characters per second without sacrificing capability scores (Data source: Artificial Analysis, refreshed Aug 28). Commentary: "Fast and smart" is the main battlefield for this wave of small models; releasing new versions feels like patching software.
● Two Free Windows: Hy4 Limited-Time Free Until Sept 10, Previous Gen Hy3 Free Period Extended to Sept 30
● Tencent WorkBuddy offers a two-week limited-time free trial for the newly open-sourced Hy4 preview (Aug 28–Sept 10), and the free period for the previous generation Hy3 is extended again to Sept 30. Context: ByteDance's "Doubao Work," Alibaba's Qianwen Office, and Tencent's WorkBuddy—all three AI office products launched within a month. Free windows have become the most direct talent-grabbing tactic. Commentary: The fiercer the battle, the more freebies. These two weeks are a good window for ordinary people to try domestic flagship models at zero cost. Beware of "credit systems" and "preview versions."
● Another Comprehensive Aggregator Launches: Claude Opus 5 Tops with 80.34, But Rules Are the Highlight
● Independent aggregator AGI Ranker combines 10 public exam papers into a single total score (max 100); current leader Claude Opus 5 scores 80.34. They wrote "independently verifiable takes precedence over vendor self-reporting" into their rules. Third-party tests count with weight 1.0; vendor self-reports get a 25% discount; items lacking raw data are marked "Evidence Verification Pending" rather than fabricating scores; all corrections are publicly archived. Commentary: Rankings wobble, but the rules are worth copying. Before buying AI based on user reviews, check if this leaderboard dares to mark "I can't verify this."
● Anthropic Opens Claude Usage Data to Independent Researchers: Three Institutions Each Receive ~250k Conversation Samples
● Around August 27, Anthropic announced opening real Claude usage data to external independent research institutions. Three research institutions each received ~250,000 conversation samples to study AI's impact on users. Commentary: Cracking open the black box for outsiders to study—transparency efforts earn credit every time they happen.
● ChatGPT's First Learning Report: 70 Million "Learning Conversations" Weekly, Homework Messages Peak at 460 Million/Week
● OpenAI released its ChatGPT learning usage report on August 28. Globally, ~70 million conversations per week are for learning purposes; homework-related messages peaked at 460 million per week. Commentary: Call volumes are votes that can't be faked. Students are already using it; classroom rules and assessment methods haven't caught up.
● Grok Voice Bot Opens to Cursor Programming Subscribers, Resets Weekly Usage Limits
● Around August 27, xAI opened Grok Bot to SuperGrok and Cursor Pro subscribers and reset weekly usage quotas. Commentary: For heavy users, this is free quota. Also a reminder: AI subscription usage rules change arbitrarily; don't put all your eggs in one basket for critical tasks.
● Digital Expo Discloses: China's Daily Token (AI Processed Word Count) Calls Exceed 140 Trillion
● The Digital Expo opening on August 29 published baseline data. China's AI large model daily token calls (roughly understood as "word count" processed/output by AI daily) exceeded 140 trillion. Every sentence you ask AI contributes to this aggregate. Commentary: Aggregate numbers don't determine winners, but they reveal the true depth of domestic model usage, harder to argue with than any launch PPT.
● Hy4 Open Source Narrowly Wins Internal Blind Test (People's Daily) / "Niu Lai" Identity Revealed as GLM-5.3-Flash Running on Domestic Chips (Sina Finance) / OpenAI Apologizes for 3% Request Misrouting (Zhihu Daily)
● Next Arena Blind Test Snapshot (Previous cutoff Aug 26): Watch if newly open-sourced GLM-5.3-Flash and Hy4 make the list, and if Gemini 3.7 Flash's recent surge holds steady.
● Have third parties re-tested Hy4's 12 self-reported benchmarks? Terminal-Bench 85.4 and DeepSWE 64.3 only count when others reproduce them.
● New timeline for ByteDance Doubao 2.2 after delay. Officials say programming, tool calling, and agent capabilities need strengthening before release; note that "delays" usually mean internal evaluations didn't pass.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.