GLOBAL AI REVIEW RADAR
2026.08.28 · Friday
Issue 018
Key Updates 22 | Leaderboard Flash 0 | Highlights 0 | Tomorrow's Watch 0
Voice AI Costing 2 Cents Per Minute Tops "Hard Question Quick Answer" Blind Test ThunderPhone's new architecture focuses on voice AI. Official figures state cost is ~2 cents per minute. It took first place in the Big Bench Audio voice QA exam. This exam tests if AI answers tricky questions fast and accurately enough. Voice assistants have long been "clear recognition, slow answers, confused by interruptions." This new architecture aims to solve "understanding well and answering fast." Topping the list is an official self-test metric. Whether it holds up in noisy, accented, or interrupted scenarios requires third-party real-world testing. Product Expert Ah Zhe says|Scoring voice entries by "getting things done" adds another sample. Previously voice leaderboards competed on recognition accuracy; now they compete on truly answering questions clearly. Voice assistants are rolling from "can hear" to "can do." 2 cents per minute aims to bring voice AI into daily consumer scenarios; the direction is right. Entry point battles are never won by whoever has lower scores, but by whoever makes users willing to speak daily. Per usual rule, don't trust launch events; try speaking yourself before judging. Editor Xiao He says|As someone who heavily trips over voice AI, I care most about accents, interruptions, and cost. 2 cents per minute sounds cheap, but cheapness must be built on "not being stupid." If it gets two out of three questions wrong, saved money is wasted. I'll consider recommending it to friends after someone does a daily test with continuous speech and interruptions. Source|Hacker News / ThunderPhone (2026-08-27)
Office AI Assistants Can Also Be Tricked by "A Few Lines in a Spreadsheet" AI helping with spreadsheets, emails, and reimbursements is trending. But a public demo proved malicious instructions can hide in spreadsheet cells. When AI reads the sheet, it gets "brainwashed" and executes unauthorized operations. The principle isn't complex: AI can't distinguish data from commands in a table. Someone disguised attack instructions as normal text in the table, and AI accepted it all. Warning for ordinary people: Don't throw unknown files to AI for processing, especially office assistants with account permissions. Security Expert Old Zhou says|This isn't an occasional bug; it's another manifestation of an old problem. AI can't distinguish data from instructions; boundaries are unclear, so attackers stuff things into the gaps. Tables, documents, and web links can all be poisoning entry points because, to AI, they are all read-in text. My stance remains unchanged: set permissions to minimum; don't rush to grant account permissions to office assistants. This was a demo, proving the attack surface exists. Wait for real cases and fixes before relaxing. Editor Xiao He says|After watching the demo, I got chills. I really do throw messy tables to AI for organizing. I've said before: don't click unknown links, humans or AI. Now add: don't accept unknown files, humans or AI. Don't enable valuable account permissions unless necessary. I've noted this since Issue 021 and will continue to note it. Source|shiftmag.dev (Hacker News, 2026-08-27)
What it means for you|Do not throw unknown files to AI for processing, especially office assistants with account permissions.
Install a Monitored "Isolation Sandbox" for AI Coding Assistants More people let AI write code, but consequences of AI messing with files or running random commands fall on you. Sandy is an open-source tool that locks AI coding assistants in a sandbox with monitoring and policy control. Every step AI takes is recorded; actions beyond allowed scope are blocked. For developers, it's like putting a dashcam and permission lock on the AI assistant. You dare let it work freely while seeing what it did. Programming Expert Old Xu says|I agree with the direction. The stronger AI's coding ability, the more critical knowing which files it touched. Sandbox plus monitoring is a pragmatic way to swap "trust" for "verifiable." I've seen many such tools. Key points: ease of integration with existing projects, whether monitoring slows development, and how many leaks occur. Don't rush to production; run a small project for a week to check stability. Pretty docs aren't as good as rolling through real projects. Editor Xiao He says|I don't code, but I get the logic. Locking AI in a cage where it can work but bad deeds are stopped is better than letting it roam wild. Before handing keys to AI, think clearly about which doors it can open and whether to retrieve them after use. I've noted this since Issue 024. Even if this tool isn't for me, it's a reassuring direction for those afraid of AI causing trouble. Source|Hacker News / GitHub (2026-08-28)
What it means for you|Sandy locks AI coding assistants in a sandbox with monitoring, blocking actions beyond allowed scope.
Arena Human Blind Test Snapshot not updated this week (data as of Aug 26, same as last issue). This issue uses three new exams to rank.
① Terminal Operation Exam (Terminal-Bench 2.1, Official Verified List as of 08-19)
Tests if AI can open terminal, modify code, and run tasks itself. On the official verified list, Claude Code + Fable 5 tops at 83.8%. Anthropic-submitted Claude Code + Opus 4.8 ranks 5th at 78.9%. Newly open-sourced model Ornith-1.5 self-reported 86.1, but that's team-tested (average of five runs), not directly comparable to official lists. Discount AI self-reported scores before looking.
What it means for you|Discount AI self-reported scores before looking, as they are not directly comparable to official lists.
② Coding Ability Exam (SWE-bench Verified, Data Updated 08-26)
Human experts pick bugs from real open-source projects for AI to fix. Correctness is obvious. Claude Opus 5 Thinking High Config scored 96, ranking 1st. Claude Fable 5 two tiers scored 95, tying for 2nd and 3rd. Top three swept by Claude family. Domestic camp closes in tight: DeepSeek-V4-Pro scored 80.6, ranking 12th, free for commercial use. Qwen Qwen3.7 Max scored 80.4, ranking 13th.
What it means for you|DeepSeek-V4-Pro scored 80.6, ranking 12th, and is free for commercial use.
③ Blind Test Code Score List (LMArena Coding, Snapshot 08-26)
Humans judge, AI fights blindly. Kimi's highest tier kimi-k3-max scored 1542, ranking 6th, continuing domestic models' first entry into top ten comprehensive blind tests. Claude Opus 5 High Config scored 1533, ranking 7th. Zhipu GLM-5.3-Max scored 1531, ranking 10th. Blind test scores are just reference coordinates; real utility depends on your specific tasks.
What it means for you|Blind test scores are just reference coordinates; real utility depends on your specific tasks.
Pydantic AI adds "type insurance" for Python AI apps, reducing pitfalls of incorrect AI return formats during development.
What it means for you|Pydantic AI reduces pitfalls of incorrect AI return formats during development.
Apronagents gives each AI coding assistant a disposable independent repo, discarded after use, preventing cross-contamination.
What it means for you|Apronagents prevents cross-contamination by giving each AI coding assistant a disposable independent repo.
Gantree lets AI run long tasks for hours without humans staring at chat boxes.
What it means for you|Gantree lets AI run long tasks for hours without humans staring at chat boxes.
KinoPipe encapsulates video editing into services AI can call directly, no guessing command lines for editing videos.
What it means for you|KinoPipe allows editing videos without guessing command lines by encapsulating services AI can call.
SCQOS is the world's first public challenge; AI must prove it should act before acting.
What it means for you|SCQOS requires AI to prove it should act before acting.
Jailbox is an offline black-box VM; AI and untrusted code run in locked environments, explosions don't affect outside.
What it means for you|Jailbox ensures explosions from untrusted code do not affect the outside environment.
LetItLoop allows AI long tasks to resume from breakpoints after crashes, no need to restart from scratch.
What it means for you|LetItLoop allows resuming from breakpoints after crashes, so no need to restart from scratch.
Relay Q uses one microphone to make AI understand you; Wired real-world test rated it close to specialized voice software.
What it means for you|Relay Q uses one microphone to make AI understand you, rated close to specialized software.
deepagents is LangChain's open-source AI framework; AI can plan tasks, call tools, and complete multi-step work itself.
What it means for you|deepagents allows AI to plan tasks, call tools, and complete multi-step work itself.
GateOnAI calculated 4 million real compatibility connections among 2866 AI tools. Data speaks to whether they can be stitched together.
What it means for you|GateOnAI data speaks to whether 2866 AI tools can be stitched together.
Anthropic tests letting Claude operate robots and lab instruments | Nvidia's 70% growth expectation ignites AI rally | Hugging Face pushes $399 open-source duck robot | OpenAI CEO admits public hates data centers
① Ornith-1.5 three-tier weights rolling out; see if third-party re-tests reproduce its self-reported 86.1 ② After new voice assistant architectures, see who integrates into mass products first ③ For table injection attacks, see if major vendors release targeted protections Full version at daily.physixfrontier.com/review/
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.