model fatıgue
A measurement of ours

Which image model leaves the rest of your picture alone? Eight small edits, measured

Published 4 Oct 2026Measured by us, 3 October 2026

The short answer

We gave four image models the same eight small edits, one after another, on a portrait, a product shot and a shopfront, and measured how much of the picture we never asked about still looks the same. After seven edits, FLUX 3 Image still had 93% of that area unchanged on average. GPT Image 2.5 Sunburst, Nano Banana Pro and Grok Imagine Image 2.0 had between 5% and 6%.

The untouched area that still looks the same, after one edit and after seven

Share of the pixels we never asked about that stay within one just-noticeable difference of the original, averaged over three images; outlined after one edit, filled after seven

after one editafter seven
0%20%40%60%80%100%share unchangedchange, pointsFLUX 3 ImageFLUX 3 Image: after one edit 99% (measured 3 October 2026), after seven 93% (measured 3 October 2026) · Model Fatigue, our measurement−5.8GPT Image 2.5 SunburstGPT Image 2.5 Sunburst: after one edit 54% (measured 3 October 2026), after seven 6% (measured 3 October 2026) · Model Fatigue, our measurement−48.4Nano Banana ProNano Banana Pro: after one edit 65% (measured 3 October 2026), after seven 5% (measured 3 October 2026) · Model Fatigue, our measurement−60.3Grok Imagine Image 2.0Grok Imagine Image 2.0: after one edit 55% (measured 3 October 2026), after seven 6% (measured 3 October 2026) · Model Fatigue, our measurement−49.3
0%20%40%60%80%100%share unchangedchange, pointsFLUX 3 ImageFLUX 3 Image: after one edit 99% (measured 3 October 2026), after seven 93% (measured 3 October 2026) · Model Fatigue, our measurement−5.8GPT Image 2.5 SunburstGPT Image 2.5 Sunburst: after one edit 54% (measured 3 October 2026), after seven 6% (measured 3 October 2026) · Model Fatigue, our measurement−48.4Nano Banana ProNano Banana Pro: after one edit 65% (measured 3 October 2026), after seven 5% (measured 3 October 2026) · Model Fatigue, our measurement−60.3Grok Imagine Image 2.0Grok Imagine Image 2.0: after one edit 55% (measured 3 October 2026), after seven 6% (measured 3 October 2026) · Model Fatigue, our measurement−49.3
Our run of 3 October 2026: one chain of eight edits per model and image, each model at its defaults. Every turn is scored against the original.
The numbers in this chart
after one editafter sevenchange, points
FLUX 3 Image99%93%−5.8 points
GPT Image 2.5 Sunburst54%6%−48.4 points
Nano Banana Pro65%5%−60.3 points
Grok Imagine Image 2.055%6%−49.3 points

GPT Image 2.5 Sunburst is #1 on Artificial Analysis's image-editing board, which ranks models by people's votes between two edits of the same picture. FLUX 3 Image isn't on that board yet. It was also the cheapest on average in our run, at $0.025 an edit.

It is one run per model and image, at each model's default settings. FLUX 3 and GPT Image 2.5 Sunburst both ignored the one edit that asked for new text.

What we did

If you use an image model to change one thing in a picture, you want the rest of the picture to stay as it was, and you often want to make a second change and a third. We measured how well four models do that.

We made three pictures with a fifth model, so no model edited its own work: a fictional woman at a desk, a bottle of skincare serum with props around it, and a Berlin bakery front with a sign and a chalkboard. Each model then got the same eight instructions for each picture, one after another, each naming one region and ending "Change nothing else in the image." Every edit started from the picture the previous one returned.

After each edit we compared the picture with the original, looking only at the area no instruction had touched so far, and counted the share of its pixels whose colour is within one just-noticeable difference of the original. The protocol was written down and committed before the first paid edit, and amended once after a pilot, to call FLUX 3 through fal because OpenRouter refused it.

Three of the four repaint the picture

FLUX 3 Image kept 99% of the untouched area unchanged after one edit and 93% after seven, on average across the three pictures. The other three changed a large part of it with the first edit: GPT Image 2.5 Sunburst kept 54%, Nano Banana Pro 65% and Grok Imagine Image 2.0 55%. The repainting builds up with each edit, and after seven all three were at 5% or 6%.

The pattern is the same on all three pictures, as the table shows. FLUX 3's one large drop is the product shot's last edit, which turns the whole wall behind the bottle blue and leaves little of the picture outside the edited area. FLUX 3 changed the colour of what was left as well, and by our rule that counts as change.

Every model and image, after one, four, seven and eight edits

Share of the untouched area within one just-noticeable difference of the original

modelimageafter one editafter fourafter sevenafter eight
FLUX 3 Imageportrait99%95%91%91%
FLUX 3 Imageproduct shot100%98%95%12%
FLUX 3 Imageshopfront97%91%92%91%
GPT Image 2.5 Sunburstportrait63%23%4%3%
GPT Image 2.5 Sunburstproduct shot62%20%5%3%
GPT Image 2.5 Sunburstshopfront37%16%7%5%
Nano Banana Proportrait69%24%6%5%
Nano Banana Proproduct shot88%30%3%3%
Nano Banana Proshopfront38%13%6%5%
Grok Imagine Image 2.0portrait55%12%4%3%
Grok Imagine Image 2.0product shot78%51%9%8%
Grok Imagine Image 2.0shopfront32%9%4%3%
Our measurement. The product shot's eighth edit turns the whole wall blue, which leaves little of the picture outside the edited area; that is why FLUX 3's product figure falls at the last edit.

The woman in the portrait stays recognisable to face recognition with every model: after eight edits her face similarity to the original is 0.99 with FLUX 3, 0.74 with Nano Banana Pro, 0.66 with Grok Imagine and 0.59 with GPT Image 2.5 Sunburst, all above 0.363, the same-person threshold our protocol takes from OpenCV for this recogniser. Above that line is not the same as unchanged: with the three models that repaint, her similarity falls with every few edits, while with FLUX 3 it stays near 1.00.

The board and the measurement disagree

Artificial Analysis's image-editing board ranks 80 models by Elo from people's votes between two edits of the same picture, made from the same instruction. GPT Image 2.5 Sunburst is #1 there, Grok Imagine Image 2.0 #10 and Nano Banana Pro #13. FLUX 3 Image isn't on the board yet.

A vote on one edit rewards the result people prefer, and it never sees a second edit; Artificial Analysis doesn't say whether voters see the original picture to judge what else changed. Both measurements can be right at once: GPT Image 2.5 Sunburst can make the edit people prefer and still change 46% of the rest of the picture in that one edit, on our average. The board's row for it is at max quality, and ours ran at the model's defaults, so the two aren't the same setting.

Did they do what was asked?

A model that ignores an instruction keeps the picture perfectly, so we checked every edit by eye. FLUX 3 and GPT Image 2.5 Sunburst did 23 of the 24, and the one each missed is the same: changing a price on the bakery's chalkboard, the only edit that asked for new text. Nano Banana Pro and Grok Imagine did all 24, though asked to empty the shop window, each built new shelving in it. And asked to whiten the books on the top shelf, FLUX 3 and Grok Imagine whitened the books on every shelf, which our first look at the pictures missed. Part of that extra change, on the second shelf, is outside every region we asked about, and the pixel scores count it. The lower shelves fall inside the region around the lamp, edited at the fourth step, which the score stops looking at from then on, so FLUX 3's portrait score counts only some of the over-edit.

The four models side by side

Board rank is Artificial Analysis's; cost, edits and face similarity are our run; the last two columns are from Artificial Analysis's board and the makers' own pages

modelediting boardcost per editedits doneface similarity at the endopen weightswho owns the output
FLUX 3 Imagenot on the board$0.025230.99licensed for self-hosting on request (BFL)you (BFL developer terms)
GPT Image 2.5 Sunburst#1$0.019 to $0.074230.59none listedyou (OpenAI Services Agreement)
Nano Banana Pro#13$0.136240.74none listedGoogle won't claim it (Gemini API terms)
Grok Imagine Image 2.0#10$0.070240.66none listedyou (xAI enterprise terms)
Cost per edit is what each call cost at the model's defaults: as billed through OpenRouter for three models, and for FLUX 3 on fal, which reports no cost, fal's list price for the output size. GPT Image 2.5 Sunburst's price varied from call to call with the number of image tokens it billed for its output, though every output was the same size, so it shows a range. Face similarity is the SFace cosine to the original face, where a score of one means identical. Open weights: Artificial Analysis's board for the three models on it, Black Forest Labs' FLUX 3 page for FLUX 3. Who owns the output: each maker's own API or developer terms, read 4 October 2026 (for BFL, both its general and its EU developer terms say so). We called the models through fal and OpenRouter and haven't read their terms.

FLUX 3 was also the cheapest on average, at $0.025 an edit, and Nano Banana Pro the dearest, at $0.136. The 24 edits for each of the four models cost $6.65 in all, and the whole run $6.83.

What this doesn't tell you

It is one run per model and picture, and we sent no random seed, so a second run would give somewhat different numbers. We'd be surprised if it closed a gap this size between FLUX 3 and the rest, but we haven't run it twice.

Every model ran at its defaults with a plain-language instruction. FLUX 3 can take boxes that tell it what to keep, and our protocol notes a fidelity setting for GPT Image 2.5 that OpenRouter doesn't pass on; we used neither. Either could change that model's numbers.

There was one text edit, and it split the models two against two, so this says little about editing text in pictures.

Keeping the rest of the picture is one thing you might want from an editor. A model that repaints more can still give you the picture you wanted.

Every number

These are all 119 figures behind this article, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one we first published.

Model Fatigue, our measurement

Our edit-drift run (scores, compliance check and spend ledger; derived/drift.py) · measured 3 October 2026
FLUX 3 Image, the portrait: share of the untouched area within one just-noticeable difference of the original after 1 edit99%
FLUX 3 Image, the product shot: share of the untouched area within one just-noticeable difference of the original after 1 edit100%
FLUX 3 Image, the shopfront: share of the untouched area within one just-noticeable difference of the original after 1 edit97%
FLUX 3 Image: the same share after 1 edit, averaged over the three images (our arithmetic)99%
FLUX 3 Image, the portrait: share of the untouched area within one just-noticeable difference of the original after 4 edits95%
FLUX 3 Image, the product shot: share of the untouched area within one just-noticeable difference of the original after 4 edits98%
FLUX 3 Image, the shopfront: share of the untouched area within one just-noticeable difference of the original after 4 edits91%
FLUX 3 Image: the same share after 4 edits, averaged over the three images (our arithmetic)94%
FLUX 3 Image, the portrait: share of the untouched area within one just-noticeable difference of the original after 7 edits91%
FLUX 3 Image, the product shot: share of the untouched area within one just-noticeable difference of the original after 7 edits95%
FLUX 3 Image, the shopfront: share of the untouched area within one just-noticeable difference of the original after 7 edits92%
FLUX 3 Image: the same share after 7 edits, averaged over the three images (our arithmetic)93%
FLUX 3 Image, the portrait: share of the untouched area within one just-noticeable difference of the original after 8 edits91%
FLUX 3 Image, the product shot: share of the untouched area within one just-noticeable difference of the original after 8 edits12%
FLUX 3 Image, the shopfront: share of the untouched area within one just-noticeable difference of the original after 8 edits91%
FLUX 3 Image: the same share after 8 edits, averaged over the three images (our arithmetic)65%
FLUX 3 Image: face similarity to the original after 1 edit (SFace cosine; 1 is identical)1.00
FLUX 3 Image: face similarity to the original after 4 edits (SFace cosine; 1 is identical)0.99
FLUX 3 Image: face similarity to the original after 8 edits (SFace cosine; 1 is identical)0.99
FLUX 3 Image: edits done, checked by eye23
FLUX 3 Image: edits asked for24
FLUX 3 Image: edits done that changed more than asked, checked by eye (with the corrections in derived/compliance-corrections.csv)1
FLUX 3 Image: cost per edit at its defaults, mean over the 24 calls (USD)$0.025
FLUX 3 Image: cheapest single edit (USD)$0.025
FLUX 3 Image: dearest single edit (USD)$0.025
GPT Image 2.5 Sunburst, the portrait: share of the untouched area within one just-noticeable difference of the original after 1 edit63%
GPT Image 2.5 Sunburst, the product shot: share of the untouched area within one just-noticeable difference of the original after 1 edit62%
GPT Image 2.5 Sunburst, the shopfront: share of the untouched area within one just-noticeable difference of the original after 1 edit37%
GPT Image 2.5 Sunburst: the same share after 1 edit, averaged over the three images (our arithmetic)54%
GPT Image 2.5 Sunburst, the portrait: share of the untouched area within one just-noticeable difference of the original after 4 edits23%
GPT Image 2.5 Sunburst, the product shot: share of the untouched area within one just-noticeable difference of the original after 4 edits20%
GPT Image 2.5 Sunburst, the shopfront: share of the untouched area within one just-noticeable difference of the original after 4 edits16%
GPT Image 2.5 Sunburst: the same share after 4 edits, averaged over the three images (our arithmetic)20%
GPT Image 2.5 Sunburst, the portrait: share of the untouched area within one just-noticeable difference of the original after 7 edits4%
GPT Image 2.5 Sunburst, the product shot: share of the untouched area within one just-noticeable difference of the original after 7 edits5%
GPT Image 2.5 Sunburst, the shopfront: share of the untouched area within one just-noticeable difference of the original after 7 edits7%
GPT Image 2.5 Sunburst: the same share after 7 edits, averaged over the three images (our arithmetic)6%
GPT Image 2.5 Sunburst, the portrait: share of the untouched area within one just-noticeable difference of the original after 8 edits3%
GPT Image 2.5 Sunburst, the product shot: share of the untouched area within one just-noticeable difference of the original after 8 edits3%
GPT Image 2.5 Sunburst, the shopfront: share of the untouched area within one just-noticeable difference of the original after 8 edits5%
GPT Image 2.5 Sunburst: the same share after 8 edits, averaged over the three images (our arithmetic)4%
GPT Image 2.5 Sunburst: face similarity to the original after 1 edit (SFace cosine; 1 is identical)0.92
GPT Image 2.5 Sunburst: face similarity to the original after 4 edits (SFace cosine; 1 is identical)0.77
GPT Image 2.5 Sunburst: face similarity to the original after 8 edits (SFace cosine; 1 is identical)0.59
GPT Image 2.5 Sunburst: edits done, checked by eye23
GPT Image 2.5 Sunburst: edits asked for24
GPT Image 2.5 Sunburst: edits done that changed more than asked, checked by eye (with the corrections in derived/compliance-corrections.csv)0
GPT Image 2.5 Sunburst: cost per edit at its defaults, mean over the 24 calls (USD)$0.046
GPT Image 2.5 Sunburst: cheapest single edit (USD)$0.019
GPT Image 2.5 Sunburst: dearest single edit (USD)$0.074
Nano Banana Pro, the portrait: share of the untouched area within one just-noticeable difference of the original after 1 edit69%
Nano Banana Pro, the product shot: share of the untouched area within one just-noticeable difference of the original after 1 edit88%
Nano Banana Pro, the shopfront: share of the untouched area within one just-noticeable difference of the original after 1 edit38%
Nano Banana Pro: the same share after 1 edit, averaged over the three images (our arithmetic)65%
Nano Banana Pro, the portrait: share of the untouched area within one just-noticeable difference of the original after 4 edits24%
Nano Banana Pro, the product shot: share of the untouched area within one just-noticeable difference of the original after 4 edits30%
Nano Banana Pro, the shopfront: share of the untouched area within one just-noticeable difference of the original after 4 edits13%
Nano Banana Pro: the same share after 4 edits, averaged over the three images (our arithmetic)22%
Nano Banana Pro, the portrait: share of the untouched area within one just-noticeable difference of the original after 7 edits6%
Nano Banana Pro, the product shot: share of the untouched area within one just-noticeable difference of the original after 7 edits3%
Nano Banana Pro, the shopfront: share of the untouched area within one just-noticeable difference of the original after 7 edits6%
Nano Banana Pro: the same share after 7 edits, averaged over the three images (our arithmetic)5%
Nano Banana Pro, the portrait: share of the untouched area within one just-noticeable difference of the original after 8 edits5%
Nano Banana Pro, the product shot: share of the untouched area within one just-noticeable difference of the original after 8 edits3%
Nano Banana Pro, the shopfront: share of the untouched area within one just-noticeable difference of the original after 8 edits5%
Nano Banana Pro: the same share after 8 edits, averaged over the three images (our arithmetic)4%
Nano Banana Pro: face similarity to the original after 1 edit (SFace cosine; 1 is identical)0.97
Nano Banana Pro: face similarity to the original after 4 edits (SFace cosine; 1 is identical)0.89
Nano Banana Pro: face similarity to the original after 8 edits (SFace cosine; 1 is identical)0.74
Nano Banana Pro: edits done, checked by eye24
Nano Banana Pro: edits asked for24
Nano Banana Pro: edits done that changed more than asked, checked by eye (with the corrections in derived/compliance-corrections.csv)1
Nano Banana Pro: cost per edit at its defaults, mean over the 24 calls (USD)$0.136
Nano Banana Pro: cheapest single edit (USD)$0.136
Nano Banana Pro: dearest single edit (USD)$0.136
Grok Imagine Image 2.0, the portrait: share of the untouched area within one just-noticeable difference of the original after 1 edit55%
Grok Imagine Image 2.0, the product shot: share of the untouched area within one just-noticeable difference of the original after 1 edit78%
Grok Imagine Image 2.0, the shopfront: share of the untouched area within one just-noticeable difference of the original after 1 edit32%
Grok Imagine Image 2.0: the same share after 1 edit, averaged over the three images (our arithmetic)55%
Grok Imagine Image 2.0, the portrait: share of the untouched area within one just-noticeable difference of the original after 4 edits12%
Grok Imagine Image 2.0, the product shot: share of the untouched area within one just-noticeable difference of the original after 4 edits51%
Grok Imagine Image 2.0, the shopfront: share of the untouched area within one just-noticeable difference of the original after 4 edits9%
Grok Imagine Image 2.0: the same share after 4 edits, averaged over the three images (our arithmetic)24%
Grok Imagine Image 2.0, the portrait: share of the untouched area within one just-noticeable difference of the original after 7 edits4%
Grok Imagine Image 2.0, the product shot: share of the untouched area within one just-noticeable difference of the original after 7 edits9%
Grok Imagine Image 2.0, the shopfront: share of the untouched area within one just-noticeable difference of the original after 7 edits4%
Grok Imagine Image 2.0: the same share after 7 edits, averaged over the three images (our arithmetic)6%
Grok Imagine Image 2.0, the portrait: share of the untouched area within one just-noticeable difference of the original after 8 edits3%
Grok Imagine Image 2.0, the product shot: share of the untouched area within one just-noticeable difference of the original after 8 edits8%
Grok Imagine Image 2.0, the shopfront: share of the untouched area within one just-noticeable difference of the original after 8 edits3%
Grok Imagine Image 2.0: the same share after 8 edits, averaged over the three images (our arithmetic)5%
Grok Imagine Image 2.0: face similarity to the original after 1 edit (SFace cosine; 1 is identical)0.95
Grok Imagine Image 2.0: face similarity to the original after 4 edits (SFace cosine; 1 is identical)0.85
Grok Imagine Image 2.0: face similarity to the original after 8 edits (SFace cosine; 1 is identical)0.66
Grok Imagine Image 2.0: edits done, checked by eye24
Grok Imagine Image 2.0: edits asked for24
Grok Imagine Image 2.0: edits done that changed more than asked, checked by eye (with the corrections in derived/compliance-corrections.csv)2
Grok Imagine Image 2.0: cost per edit at its defaults, mean over the 24 calls (USD)$0.070
Grok Imagine Image 2.0: cheapest single edit (USD)$0.070
Grok Imagine Image 2.0: dearest single edit (USD)$0.070
What the 96 edits cost in all (USD)$6.65
What the whole run cost, the source images included (USD)$6.83
FLUX 3 Image: change in that average share from 1 edit to 7, in percentage points (our arithmetic)−5.8 points
FLUX 3 Image: share of the untouched area that changed after one edit, averaged over the three images (our arithmetic)1%
GPT Image 2.5 Sunburst: change in that average share from 1 edit to 7, in percentage points (our arithmetic)−48.4 points
GPT Image 2.5 Sunburst: share of the untouched area that changed after one edit, averaged over the three images (our arithmetic)46%
Nano Banana Pro: change in that average share from 1 edit to 7, in percentage points (our arithmetic)−60.3 points
Nano Banana Pro: share of the untouched area that changed after one edit, averaged over the three images (our arithmetic)35%
Grok Imagine Image 2.0: change in that average share from 1 edit to 7, in percentage points (our arithmetic)−49.3 points
Grok Imagine Image 2.0: share of the untouched area that changed after one edit, averaged over the three images (our arithmetic)45%

Artificial Analysis

Artificial Analysis, image editing leaderboard · read 4 Oct 2026, 08:51 CEST
GPT Image 2.5 Sunburst: rank on Artificial Analysis's image-editing board#1 ⟳
GPT Image 2.5 Sunburst: Elo on Artificial Analysis's image-editing board1182 ⟳
Nano Banana Pro: rank on Artificial Analysis's image-editing board#13 ⟳
Nano Banana Pro: Elo on Artificial Analysis's image-editing board1099 ⟳
Grok Imagine Image 2.0: rank on Artificial Analysis's image-editing board#10 ⟳
Grok Imagine Image 2.0: Elo on Artificial Analysis's image-editing board1107 ⟳
Models on Artificial Analysis's image-editing board (our count)80 ⟳

Model Fatigue, our measurement

The edit-drift protocol, frozen before the paid run · frozen 3 October 2026
The same-person threshold for SFace cosine that the protocol takes from OpenCV0.363
The colour difference (CIE76 ΔE) the protocol counts as one just-noticeable difference2.3

Sources

These are the pages this article draws on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.