GLOBAL AI REVIEW RADAR
2026.08.21 · Friday
Issue 011
Key Updates 1 | Leaderboard Flash 20 | Highlights 9 | Tomorrow's Watch 0
What it means for you|AI responds faster and costs less on the same hardware, making local offline applications less laggy for ordinary users.
●
● | # | Model | Vendor | Pass Rate |
● |---|------|------|--------|
● | 1 | Pine Voice Preview | Pine AI | 80.2% |
● | 2 | Pine Voice Preview | Pine AI | 75.4% |
● | 3 | grok-voice-think-fast-1.0 | xAI | 74.8% |
● | 4 | grok-voice-think-fast-2.0 | xAI | 62.5% |
● | 5 | qwen3.5-omni-plus-realtime | Alibaba | 53.7% |
● In plain language, a new leaderboard testing "whether AI voice customer service can actually get things done." Pine AI's voice model took first place with an 80.2% pass rate; the second place is also them (same model, different config). One lab occupying the top two spots—this leaderboard is still early, so don't treat the rankings as final conclusions yet. Grok Voice is stuck at 74.8%, and OpenAI's voice model only has 51.2%; the gap is significant.
●
● | # | Model | Vendor | Pass Rate |
● |---|------|------|--------|
● | 1 | Qwen 3.8 Max | Alibaba | 55.2% |
● | 2 | Claude Opus 5 | Anthropic | 48.7% |
● | 3 | Grok 4.5 | xAI | 47.9% |
● | 4 | GPT-5.6-sol | OpenAI | 46.9% |
● | 5 | GPT-5.5 | OpenAI | 44.6% |
● In plain language, the text version of the exam from the same τ-bench family, asking AI to act as a bank customer service agent in a knowledge base of 700 documents, getting the user's task done in one go. Alibaba's Qwen 3.8 Max ranks first with 55.2%, marking the first time a domestic model tops this "actually getting things done" dimension. Getting all four steps right isn't easy anyway; even the leader can't chew down more than half.
●
● Arena's latest snapshot (08-20) is identical to the previous issue; there is no new battle data, so this week can only be marked as not updated. Recap highlights: In human blind voting, Claude Fable 5 defends its throne with 1507 Elo; Anthropic holds four seats in the top ten. Meta's muse-spark-1.2 skyrocketed to 4th place with only 3264 matches; the sample size is too small, and the ranking needs further observation. Wait for the next snapshot to see who wins.
●
●
●
●
●
●
●
●
●
📖 Full leaderboard tables and expert commentary live in the Chinese edition of the review magazine.