Written pieces with no video, each a sourced answer to one question about which model to use, at what effort and at what cost. A question we keep re-reading has one page, updated in place, with each reading kept on its history page. The method page says who writes them and how they are checked.
On Artificial Analysis's leaderboard as we read it on 4 October 2026, 102 current model settings have a score and a cost per task above zero. Only 15 of them are beaten by no other setting on score and cost at once: GPT-6 Luna at every reasoning setting from low to max, Xiaomi's MiMo-V2.6-Flash and MiMo-V2.6-Pro, GPT-6.1 Sol at every setting, and Claude Opus 5.5 at high, xhigh and max.
We gave four image models the same eight small edits, one after another, on a portrait, a product shot and a shopfront, and measured how much of the picture we never asked about still looks the same. After seven edits, FLUX 3 Image still had 93% of that area unchanged on average. GPT Image 2.5 Sunburst, Nano Banana Pro and Grok Imagine Image 2.0 had between 5% and 6%.
On Artificial Analysis's Intelligence Index, read on 4 October 2026, Claude Opus 5.5 scores higher than Claude Fable 5.1 at every reasoning setting from medium to max, by 2.3 points to 4.3 points, and costs less per task at every setting. Fable leads only at low, where it costs 4.3× as much.