● First, the ground rules: Arena's several boards are the same snapshot as last issue, with no new changes this period, and the actual update dates differ across groups (the text and coding boards are frozen at 2026-09-25, the agentic board at 2026-09-27, and the image board at 2026-09-13). The new numbers we can report today come from Artificial Analysis's composite intelligence score.
● Composite Intelligence Score (scraped today, data as of 2026-09-29)
● | # | Model | Vendor | Composite Intelligence Index |
● | --- | --- | --- | --- |
● | 1 | Claude Opus 5.5 (max) | Anthropic | 58 |
● | 2 | Claude Opus 5.5 (xhigh) | Anthropic | 56 |
● | 3 | Claude Sonnet 5.5 (max) | Anthropic | 56 |
● | 4 | Claude Opus 5.5 (high) | Anthropic | 54 |
● | 5 | Claude Fable 5.1 (max) | Anthropic | 53 |
● This board puts each vendor's models through the same batch of questions and scores them; the higher the score, the more well-rounded the performance. The top five are all Anthropic's Claude. OpenAI's GPT-6 Astra is 53, tied with fifth place; the highest domestic one is Xiaomi's MiMo-V2.6-Pro at 46.
● Human Blind-Vote Text Board (data as of 2026-09-25)
● | # | Model | Vendor | Blind-test score (matches) |
● | --- | --- | --- | --- |
● | 1 | claude-opus-5.5-high | Anthropic | 1509 (2,307 matches) |
● | 2 | claude-opus-4-6-high | Anthropic | 1505 (76,518 matches) |
● | 3 | claude-fable-5-high | Anthropic | 1504 (36,462 matches) |
● | 4 | claude-opus-4-7-high | Anthropic | 1502 (64,007 matches) |
● | 5 | claude-fable-5.1-max | Anthropic | 1501 (9,942 matches) |
● This board is voted on by humans one-on-one — two people get two answers to the same question, and neither knows which company made which. The top five are all Anthropic. Pay attention to the match counts in parentheses: No. 1 has only accumulated 2,307 votes, with a margin of plus or minus 12 points, so its rank is still wobbling; No. 2 has 76,000 votes, so it's more reliable.
● Agentic Capability Board (data as of 2026-09-27)
● | # | Model | Vendor | Net improvement score |
● | --- | --- | --- | --- |
● | 1 | Claude Fable 5.1 (Max) | Anthropic | 13.8 |
● | 2 | Claude Opus 5.5 (High) | Anthropic | 12.2 |
● | 3 | GPT 6 Astra (Max) | OpenAI | 10.3 |
● | 4 | Claude Opus 5 (Max) | Anthropic | 9.6 |
● | 5 | Claude Opus 5 (High) | Anthropic | 9.5 |
● This group tests whether, after letting AI actually get hands-on (clicking the mouse, editing files), it gets the job done. The score is called net improvement score, meaning how much better the result is than before. First place is Anthropic's Fable 5.1, second is its Opus 5.5, and third is OpenAI's GPT-6 Astra.
●
- AI chat is narrowing our knowledge more and more. Researchers at the University of Copenhagen in Denmark raised a warning: more and more people rely on AI chat to understand the world, and over time the knowledge we encounter will concentrate on fewer and fewer topics, with vast blank areas no one looks at anymore. One-line take: don't treat AI as a library — it's better as a helper for looking up the occasional question.
●
- OpenAI is about to give its AI assistant some remedial lessons. This Verge piece says OpenAI was the first to popularize chat AI, but in the "can actually do work on its own" category, it's fallen behind. The article notes its annual developer conference is approaching, and outsiders are watching what it brings to fill this gap. One-line take: being strong at chatting doesn't equal being strong at doing work — those are two different things; when picking a tool, first figure out which one you need.
●
- What kind of AI product counts as "serious"? This blog post aired a complaint: today's AI products look lively, but most don't take their own words seriously — they profess some principle in words, yet you can't see it in the product. The author casually lists a few "what it would look like if done seriously." One-line take: to judge whether a tool is reliable, first check whether what it does and what it advertises are the same thing.
●
- Saying "change this" to the screen — which "this" should the AI understand? When there's a lot on screen and you casually say "change this," the AI may well guess wrong about what you're pointing at. This article discusses: in interfaces made for AI, how to make the act of "selecting" clearer and less error-prone. One-line take: the more clearly you point, the less it guesses wrong.
●
- Someone gave AI an inner life that "thinks even when no one's talking to it." An open-source project lets AI keep its own "state" when it's not chatting with anyone: it remembers, ponders what it is, gets curious. The whole thing runs on your own computer. One-line take: sounds novel, but "having an inner life" and "being reliable" are two different things — treat it as an experiment for now.
●
- AI assistants spinning in place and unable to get out — now there's a tool to catch it early. When you have AI do long tasks, it sometimes gets stuck in a loop doing the same thing over and over. A small open-source tool specifically watches for this "spinning in place" and can raise an alarm when it starts spinning, saving you from burning time and money for nothing. One-line take: the scariest thing about long tasks isn't slowness — it's that it's going in circles and you don't know.
●
- Whether your website is visible to AI — there's a small tool to check. More and more people find things through AI rather than a search box. This small open-source tool goes in this new direction: it scans your site and picks out the problems that keep AI and search engines from grabbing your content. One-line take: only if AI can grab it can it be written into its answers.
●
- One afternoon, using AI to build an app that runs on four Apple devices. A developer shared an experience: from the first line to it running, it took about five hours, producing an app that works on four Apple platforms including iPhone and iPad. He said the whole process was treating AI as a partner while he watched the key spots himself. One-line take: one person's one-afternoon output shouldn't be taken as a team's delivery speed.
●
- Give AI assistants a "work badge" to clarify who they are and what they can do. An identity management product added a spot specifically for AI assistants: like a company issuing badges and permissions to employees, it registers a separate identity for each AI that can work on its own and limits what it can touch. One-line take: only once you can call it by name and state its permissions can you talk about controlling it.
●
- Turn a stack of legal documents into searchable, verifiable electronic text. This tool handles scanned legal documents: it first recognizes which article each paragraph is, converts it into searchable text, and keeps a link back to the original. The author stresses one thing: when checking legal provisions, first make sure the source matches up. One-line take: it finds the original text for you, but the step of judging right from wrong is still up to you.
●
- The mystery of Alibaba AI's growth: cloud business becomes the barometer of its transformation (Huxiu)
●
- Modal Labs, in the AI model hosting business, is close to completing a $750 million funding round at a $15.75 billion valuation (TechCrunch)
●
- Investor Vinod Khosla predicts: most robotics companies' valuations will fall by 2030 (The Information)
●
- A media outlet did the math: AI's hidden $3 trillion in costs could weigh on the global economy (Telegraph)
●
- Meta launches a new enterprise AI platform, tapping MongoDB's former CEO to lead it (ITHome)