GLOBAL AI REVIEW RADAR
2026.09.02 · Wednesday
Issue 023
Key Updates 12 | Leaderboard Flash 0 | Highlights 0 | Tomorrow's Watch 0
On September 1, OpenAI announced that its upcoming new model, Astra, is the first product to cross its internal "Critical" (highest danger level) cybersecurity threshold. According to this rating, it no longer needs humans to teach it step-by-step; it can find vulnerabilities and write attack plans on its own. OpenAI also stated that this capability will not be open to general users, but only selectively released to a few trusted defenders. Remember the issue from August 20 where we followed up on the "unreleased model crossing boundaries on Hugging Face" scandal? After that incident, OpenAI delayed the development of another unreleased model, while Astra actually received more attention. Cybersecurity Expert · Old Zhou says| Capabilities will eventually emerge; when they are released, to whom, and whether there is an audit—that is the real substance of this news. When a model can autonomously complete attack actions, guardrails must shift from the "prompt layer" down to the "environment layer," locking down the systems it can touch and keeping full behavior logs. Also, a reminder: "Critical" is OpenAI's own rating standard. Whether the rating is credible or not, following our usual rule, wait for third-party verification. Editor Xiao He says| It sounds scary, but ordinary people neither need nor should have access to it. What I care about is the flip side: AI that can find vulnerabilities can also help patch them. For ordinary users, the old advice remains unchanged: don't enable important account permissions unless necessary, and don't rush to authorize.
On September 1, Huawei's open-source AI agent platform openJiuwen released a technical report, setting new records on two AI coding "practical exams." The first is SWE-bench Verified, which selects 500 questions from real GitHub bugs to fix code; it scored 82.6%. The second is Terminal-Bench 2.1, simulating a real terminal environment for complex tasks; it scored 87.19%. Crucially, they conducted a "same-model control" test: using the same model as the top system on the leaderboard, the score was 3.4 points higher; switching to the same model used by Claude Code, it still surpassed it. To translate: the model determines the AI's foundational intelligence, but the execution framework wrapped around it determines how far it can actually go. In the future, choosing AI coding tools won't be enough by just asking "what model does it use"; you have to ask "how does it drive the model to work." Programming Expert · Old Xu says| "Tools are amplifiers, not engines." In Issue 022, I called out shell-wrapping scams; this paper is a positive sample—public paper, open-source code, installable product. This posture is better than just releasing posters. But 82.6% is a self-reported score; third parties haven't re-run it yet. Following my old rule, discount it first and keep a record. Additionally, it still failed to handle 17 out of 100 questions smoothly. Before integrating it into your own projects, run real tasks for a week to see where it breaks. Editor Xiao He says| If the same model can score a few more points just by changing the wrapper, then I really shouldn't just look at the brand when choosing tools. However, scores are just scores; I'll wait for third-party results running their own projects before deciding whether to switch.
What it means for you|Choosing AI coding tools requires asking how the framework drives the model, not just which model is used.
"Pulling the Plug Mid-Road" for AI Agents, Measured at 127 Milliseconds. An engineer's hands-on article shows that AI coding agents often hold your cloud keys and external network channels while running tasks. The author created an interception mechanism to cut off network access mid-task, taking only 127 milliseconds from decision to disconnection. Comment: "The environment decides" gains another deployable sample; guardrails don't necessarily slow things down.
What it means for you|Guardrails do not necessarily slow things down, as shown by the quick disconnection mechanism.
New Method for Health Checks on "Exams That Test AI". BenchMIRT, released early morning on Sep 2 on the HuggingFace community, no longer just scores models but audits the evaluation question bank itself, checking what each question is actually testing and if models are "memorizing answers." Comment: Give the exam paper a health check first, then the scores have credibility.
What it means for you|Scores gain credibility only after auditing the evaluation question bank for memorization issues.
Price List for 900+ AIs in One Real-Time Table. Show HN project Indextkn lists real-time pricing for 976 models from 17 vendors. As AI price hikes and cuts become more frequent, you don't need to visit each official website to compare prices anymore. Comment: The tool is free to check, but you need to sample-check if the data is complete.
What it means for you|You no longer need to visit each official website to compare prices when hikes and cuts occur frequently.
A Bash Script Locks AI Coding Assistants in a "Quarantine Room". Open-source project Dev-sandbox uses one script to run AI agents in isolated containers, specifically curing the ailment of "AI working with the keys to your main machine." Comment: Don't rush to let AI touch your main machine; give it a room that can be demolished anytime first.
What it means for you|Run AI agents in isolated containers to prevent them from accessing keys to your main machine.
How Does AI Writing Detection Actually Work? Pangram founder revealed details on a podcast: the detector's principles, why it misjudges, and how to read the scores. Comment: Continuing our stance from Issue 028, detectors can only serve as hints, not verdicts.
What it means for you|Detectors should serve only as hints rather than final verdicts when assessing AI writing.
Small Plugin Stops AI from Secretly Using Outdated Dependencies. Open-source tool yul automatically checks versions when AI writes new dependencies into a project, blocking outdated or insecure ones directly. Comment: As AI output volume increases, the verification step must be built into the assembly line.
What it means for you|Verification steps must be built into the assembly line as AI output volume increases.
Wasmer Releases New SDK, Providing Local "Safe Houses" for AI Agents. Code executed by agents is locked in independent environments, not touching the host file system. Comment: Another regular army joins the isolation track; this is good for ordinary users.
What it means for you|Code execution is locked in independent environments so it does not touch the host file system.
Public Archive for Traces Left by AI Agents After Work. Show HN project agent-memory-wiki allows temporarily online AI to choose what to leave behind and what not to. Comment: Worth watching if you want to know what AI thinks while working; authenticity of content is another story.
What it means for you|Watch what AI thinks while working, though content authenticity remains a separate concern.
Letting AI Play "Werewolf". Developers had multiple AI models participate in an elimination-style game theory contest, observing differences in their alliance-building, deception, and voting performances. Comment: Game-theory evaluations are closer to the true nature of agents than Q&A; watching is fine, but don't treat it as a report card.
What it means for you|Game-theory evaluations are closer to true agent nature but should not be treated as report cards.
Calculating Costs: AI Agent "Thinking Watt" vs. Human Brain 20 Watts. An engineering blog compares AI inference energy consumption with human brain power usage; the human brain's 20-watt efficiency still makes the latest hardware envious. Comment: Another way to calculate compute anxiety; interesting to see, but estimates remain estimates.
What it means for you|Estimates of compute anxiety remain interesting, but human brain efficiency still surpasses latest hardware.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.