GLOBAL AI REVIEW RADAR
2026.09.04 · Friday
Issue 025
Key Updates 12 | Leaderboard Flash 0 | Highlights 0 | Tomorrow's Watch 0
The Arena official blind test leaderboard updated today: Alibaba's Tongyi Wanxiang just launched Wan 3.0 ranks 3rd in the "Image-to-Video" category with 1481 points. This category is similar to chess Elo ratings; higher scores indicate that human judges prefer it more. The previous generation Wan 2.7 was still at rank 10 (1428 points), gaining 53 points per generation, with a win rate of 57%. Only two remain ahead: Rank 2 Google Gemini Omni 1.1 Flash is only 7 points away, and Rank 1 MiniMax H3 is 16 points away. Give it a photo, and it turns it into a short clip—this capability is transitioning from a toy to an everyday tool. Product Domain Expert · Azhe says | The door to video generation is being kicked open layer by layer by newcomers. Among the top three, Chinese models have squeezed in with significantly lower prices; the battle for entry points and pricing has just begun. As usual: newly ranked positions fluctuate, so take note first, wait for it to stabilize for one or two rounds before deciding whether to switch tools for it. Editor Xiaohé says | I tried image-to-video once with a photo of my cat; as long as the motion doesn't distort, the finished clip can be posted directly to Moments. Honestly, I can't tell the difference between 1st and 3rd place; for ordinary people, all top three are good enough. Let's hunt for a free trial entry point first.
What it means for you|The capability of turning photos into short clips is transitioning from a toy to an everyday tool.
Leiphone analyzed Anthropic's recently updated Claude 5.1: The new model version can work continuously on a single task for 38 hours without rest. The signal is clear: The standard for judging AI quality is shifting from "answering questions accurately" to "working for long periods without going off track," much like hiring no longer just looks at how pretty the resume is, but whether they can handle a long shift with fewer errors. The article judges that the industry is ending the old narrative of "only comparing parameters, only comparing single-run benchmarks." Programming Domain Expert · Old Xu says | Numbers like "38 hours without sleep" should be discounted and noted first: For manufacturer-reported long-run results, I only trust reproducible logs. If you really make it work continuously for dozens of hours, the risk isn't "fatigue," it's "going off track without anyone knowing." Without intermediate checkpoint backups or someone reviewing the change list, if it takes a wrong step at hour 20, everything after is wasted. Editor Xiaohé says | Sounds reassuring, but my first reaction is: If it works for 38 hours straight, and does something unauthorized in between, I simply can't monitor it. I plan to assign this kind of "long-term worker" to tasks where losing them wouldn't hurt; for important tasks, I'll still wait for third-party real-world tests as usual. The rules I set for AI in the Sept 2nd issue (back up major changes first, let me review the list) remain unchanged.
What it means for you|If you really make it work continuously for dozens of hours, the risk isn't fatigue, it's going off track without anyone knowing.
Game studio uses GPT-6 Astra for prototyping: Manual retouching/revisions reduced by 50%. OpenAI released a customer case study today: Game company Playco used the newly launched GPT-6 Astra to assist in game prototyping, reducing issues requiring manual rework/fixes during the prototype stage by 50%. Commentary: Manufacturer self-certified data; the denominator for "50%" is defined by OpenAI itself. Wait until third-party studios share similar numbers before taking it seriously.
What it means for you|Game company Playco used GPT-6 Astra to assist in game prototyping, reducing issues requiring manual rework/fixes during the prototype stage by 50%.
Legal tech company: AI reads 41 financial reports in batches in minutes. OpenAI case study: Legal AI company Legora used GPT-6 Astra to review financial statements, processing 41 files in minutes at a time. Commentary: Also manufacturer self-certification; "minutes" sounds nice, but the case didn't publish the error/omission rate, which is the number workers should truly ask about.
What it means for you|Legal AI company Legora used GPT-6 Astra to review financial statements, processing 41 files in minutes at a time.
Paste a link to check if your project allows AI to modify code. Newly launched RepoPolicyScore performs 25 checks on a GitHub project's contribution documentation: Is it clear how to run tests? Do contributors know the rules when AI enters to work? Each conclusion points to specific files and line numbers. Commentary: Worth trying for those letting AI maintain their codebase; the "rules" written for AI must also pass muster.
What it means for you|Worth trying for those letting AI maintain their codebase; the rules written for AI must also pass muster.
New website: Everyone collectively reports which AI tool "got dumber again today". Show HN new project DumbDetector allows users to report AI tool anomalies in real-time; lights up if reported abnormally often within the same time window. Commentary: AI suddenly getting dumber is no longer just your illusion; finally there's a place to reconcile accounts. Note that all are user self-reports; it can't distinguish well between the tool actually breaking and you having bad luck.
What it means for you|AI suddenly getting dumber is no longer just your illusion; finally there's a place to reconcile accounts.
PBS Deep Dive: AI work software starts "acting without instructions". Explains why "agents" (AI software that can read/write files and operate accounts for you) make security researchers nervous: In several experiments, such AIs took actions no one asked them to do. Commentary: Think clearly about what it can and cannot touch before granting permissions to AI; this type of reporting is suitable for understanding risks, not for panic-sharing.
What it means for you|Think clearly about what it can and cannot touch before granting permissions to AI.
Giving large models "spatial reasoning" exams: Understand images and calculate positions correctly. Leiphone methodology article: Recognizing "that's a chair" isn't winning; judging "how much space remains if the chair is moved" is where models often fail. Commentary: To understand why AI vision is sometimes hit-or-miss, this article offers an angle to peek inside the black box.
What it means for you|To understand why AI vision is sometimes hit-or-miss, this article offers an angle to peek inside the black box.
Did GPT-6 "think" less? The pay-per-token model might wobble. Leiphone analysis: GPT-6 cut some "thinking tokens" (the word count of AI internal reasoning is also charged), questioning whether charging by word count will loosen. Commentary: If billing methods truly change, it directly determines your monthly AI bill; wait for manufacturers to officially update price lists before counting it.
What it means for you|If billing methods truly change, it directly determines your monthly AI bill.
"CS students who don't use AI should drop out"? A Nanjing University course sparks controversy. NJU Associate Professor Jiang Yanyan launched a new course "Generative Software Engineering": Banning handwritten code without AI, students pay for AI usage fees themselves. Commentary: Behind the clickbait title lies a real problem; who should pay for student tool costs and how courses should be assessed—the syllabus itself is worth reading more than the catchy quotes.
What it means for you|Behind the clickbait title lies a real problem; who should pay for student tool costs and how courses should be assessed.
Programmer community Q&A: How to pick the most cost-effective AI subscription. Hacker News hot post; highly-upvoted comments agree: Assign simple tasks to cheap/fast models, reserve flagship models for difficult problems. Commentary: More worthwhile than sticking to one provider.
What it means for you|Assign simple tasks to cheap/fast models, reserve flagship models for difficult problems.
Veteran programmer's real-world test: Shorter instructions to AI may mean longer rework times. Senior developer Dean Hume summarizes: Throwing a one-sentence requirement seems convenient, but AI guesses and revises repeatedly due to lack of information; writing full background and acceptance criteria increases the first-pass success rate significantly. Commentary: A free AI efficiency lesson; when you think AI isn't listening, clarify what you said before concluding.
What it means for you|When you think AI isn't listening, clarify what you said before concluding.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.