GLOBAL AI REVIEW RADAR
2026.08.29 · Saturday
Issue 019
Key Updates 1 | Leaderboard Flash 4 | Highlights 16 | Tomorrow's Watch 4
The official Arena (formerly LMArena) account announced that Gemini 3.7 Flash (High) has landed on three leaderboards. It ranks 20th overall in the Agentic Leaderboard; the previous generation, Gemini 3.6 Flash (High), was ranked 35th, so this is a jump of 15 positions. Its tier for long-horizon tasks is now on par with Claude Sonnet 4.6 and GPT-5.6 Terra (xHigh). In the WebDev sub-leaderboard for coding, it jumped from 19th to 8th. On the Text Chat leaderboard, it sits at 9th with a score of 1490. Regarding pricing, until the end of 2026, input costs $0.75 per million tokens and output costs $3.75. Starting January 1, 2027, prices will revert to $1.50 and $7.50 respectively.
Programming Expert · Old Xu says | Moving from 35th to 20th on the Agentic Leaderboard and 19th to 8th on WebDev represents real ranking changes. But these new ranks have only been on the board for a day or two, and the sample size isn't sufficient yet. Let's wait for it to stabilize for a round or two first. The vendor's own report cards from the batch released on August 13 (FrontierCode 43.6%, DeepSWE 65.3%) should be discounted as usual. If you want to verify, run your own real projects through it.
Editor Xiao He says | So Google's latest model improved its rankings across all three areas: working, coding, and chatting, and there's a price discount until year-end. However, I just learned last issue that new leaderboard rankings fluctuate, so don't rush to switch your entire stack to Google. Public leaderboards are essentially buyer reviews before purchase. The leaderboard data is current as of August 26, while the new rankings were announced on August 28, leaving a two-day gap.
● Text Chat Leaderboard:: Claude Fable 5 (1508 points) leads. Anthropic holds four of the top five spots, with Meta's Muse Spark 1.2 sandwiched in between. In human blind tests, the Claude family almost monopolizes the board.
● Agentic Leaderboard:: Claude Opus 5 (High) leads with a net improvement of 12.73 points. Anthropic holds four of the top five spots. OpenAI's GPT-5.6 Sol ranks 4th. Domestic Kimi K3 ranks 6th.
● Coding Leaderboard:: Claude Opus 5 (Max) is first with 1691 points. Open-source contenders Kimi K3 and Qwen3.8 Max broke into the top three.
● New in this issue:: Gemini 3.7 Flash landed at 8th on the WebDev sub-leaderboard on August 28. See Key Updates for details.
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
●
● This publication is a preview version. The official version updates at 5 PM on the same day. Leaderboard data is current as of 2026-08-26.
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.