GLOBAL AI REVIEW RADAR
2026.08.18 · Tuesday
Issue 008
Key Updates 1 | Leaderboard Flash 3 | Highlights 17 | Tomorrow's Watch 7
① AI is now better at "persuading people" than human experts An arXiv study claims that AI systems have outperformed human experts in persuasion tasks
① AI is now better at "persuading people" than human experts An arXiv study claims that AI systems have outperformed human experts in persuasion tasks. AI has no emotions, doesn't get tired, and can tailor rhetoric for each individual—persuasion is the underlying capability common to all business, politics, and scams. (Source: arXiv / Hacker News) ② Official deep dive into Claude's invisible watermarking Anthropic explains the principle of invisible text watermarking: fine-tuning word selection probabilities embeds a "signature" into the text stream, which persists even after copying or rewriting. Quality impact and centralized verification rights are two points of controversy. (Source: TheVerge / Guardian) ③ Agent costs to rise 5x in five years Enterprise-grade AI agent costs are expected to inflate by approximately 5 times by 2028. Engineering capabilities to save tokens will become key to scaling—this is also why the model routing layer is valued at $7 billion. (Source: The Register)
● OpenRouter Latest Weekly Chart (8/10-8/16): Global usage reached 75.3 trillion token: OpenRouter Latest Weekly Chart (8/10-8/16): Global usage reached 75.3 trillion tokens (+9.13%), with China accounting for 36.84 trillion, maintaining the global #1 spot for 16 consecutive weeks. DeepSeek V4 Flash official version topped the chart for two consecutive weeks with 11.2 trillion tokens, Tencent Hy3 came second with 9.95 trillion, and GPT-5.6 Luna entered the top three for the first time with 5.31 trillion. (Source: Daily Economic News 08-17)
● Cost-Performance Chart: DeepSeek V4 Flash costs approximately 1/40th of GPT-5.6 Sol: Cost-Performance Chart: DeepSeek V4 Flash costs approximately 1/40th of GPT-5.6 Sol and 1/100th of Claude Fable 5 per task (Artificial Analysis data).
● New Benchmark: CladBench launched—536 questions across 12 categories, specifically: New Benchmark: CladBench launched—536 questions across 12 categories, specifically testing large models' mastery of UK and EU building regulations.
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
● See how DeepSeek's peak/off-peak pricing affects next week's OpenRouter weekly chart
● World Robot Conference opens on 8/19; Unitree's "Superman" robot makes its live debut
● Claude's invisible watermark rolls out in gray scale; can third parties detect it?
● After GPT-5.6 Luna enters the top three, the tug-of-war between open-source and closed-source call volumes will become clear
●
● Views belong to the original authors; data is subject to official disclosures. This column focuses on the true level of AI software and hardware and does not constitute any investment advice.
● Frontier Tech Review · Global AI Evaluation Radar | Shenzhen Frontier Tech Co., Ltd.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.