The short answer
We gave four image models the same eight small edits, one after another, on a portrait, a product shot and a shopfront, and measured how much of the picture we never asked about still looks the same. After seven edits, FLUX 3 Image still had 93% of that area unchanged on average. GPT Image 2.5 Sunburst, Nano Banana Pro and Grok Imagine Image 2.0 had between 5% and 6%.
The untouched area that still looks the same, after one edit and after seven
Share of the pixels we never asked about that stay within one just-noticeable difference of the original, averaged over three images; outlined after one edit, filled after seven
The numbers in this chart
| after one edit | after seven | change, points | |
|---|---|---|---|
| FLUX 3 Image | 99% | 93% | −5.8 points |
| GPT Image 2.5 Sunburst | 54% | 6% | −48.4 points |
| Nano Banana Pro | 65% | 5% | −60.3 points |
| Grok Imagine Image 2.0 | 55% | 6% | −49.3 points |
GPT Image 2.5 Sunburst is #1 on Artificial Analysis's image-editing board, which ranks models by people's votes between two edits of the same picture. FLUX 3 Image isn't on that board yet. It was also the cheapest on average in our run, at $0.025 an edit.
It is one run per model and image, at each model's default settings. FLUX 3 and GPT Image 2.5 Sunburst both ignored the one edit that asked for new text.
What we did
If you use an image model to change one thing in a picture, you want the rest of the picture to stay as it was, and you often want to make a second change and a third. We measured how well four models do that.
We made three pictures with a fifth model, so no model edited its own work: a fictional woman at a desk, a bottle of skincare serum with props around it, and a Berlin bakery front with a sign and a chalkboard. Each model then got the same eight instructions for each picture, one after another, each naming one region and ending "Change nothing else in the image." Every edit started from the picture the previous one returned.
After each edit we compared the picture with the original, looking only at the area no instruction had touched so far, and counted the share of its pixels whose colour is within one just-noticeable difference of the original. The protocol was written down and committed before the first paid edit, and amended once after a pilot, to call FLUX 3 through fal because OpenRouter refused it.
Three of the four repaint the picture
FLUX 3 Image kept 99% of the untouched area unchanged after one edit and 93% after seven, on average across the three pictures. The other three changed a large part of it with the first edit: GPT Image 2.5 Sunburst kept 54%, Nano Banana Pro 65% and Grok Imagine Image 2.0 55%. The repainting builds up with each edit, and after seven all three were at 5% or 6%.
The pattern is the same on all three pictures, as the table shows. FLUX 3's one large drop is the product shot's last edit, which turns the whole wall behind the bottle blue and leaves little of the picture outside the edited area. FLUX 3 changed the colour of what was left as well, and by our rule that counts as change.
Every model and image, after one, four, seven and eight edits
Share of the untouched area within one just-noticeable difference of the original
| model | image | after one edit | after four | after seven | after eight |
|---|---|---|---|---|---|
| FLUX 3 Image | portrait | 99% | 95% | 91% | 91% |
| FLUX 3 Image | product shot | 100% | 98% | 95% | 12% |
| FLUX 3 Image | shopfront | 97% | 91% | 92% | 91% |
| GPT Image 2.5 Sunburst | portrait | 63% | 23% | 4% | 3% |
| GPT Image 2.5 Sunburst | product shot | 62% | 20% | 5% | 3% |
| GPT Image 2.5 Sunburst | shopfront | 37% | 16% | 7% | 5% |
| Nano Banana Pro | portrait | 69% | 24% | 6% | 5% |
| Nano Banana Pro | product shot | 88% | 30% | 3% | 3% |
| Nano Banana Pro | shopfront | 38% | 13% | 6% | 5% |
| Grok Imagine Image 2.0 | portrait | 55% | 12% | 4% | 3% |
| Grok Imagine Image 2.0 | product shot | 78% | 51% | 9% | 8% |
| Grok Imagine Image 2.0 | shopfront | 32% | 9% | 4% | 3% |
The woman in the portrait stays recognisable to face recognition with every model: after eight edits her face similarity to the original is 0.99 with FLUX 3, 0.74 with Nano Banana Pro, 0.66 with Grok Imagine and 0.59 with GPT Image 2.5 Sunburst, all above 0.363, the same-person threshold our protocol takes from OpenCV for this recogniser. Above that line is not the same as unchanged: with the three models that repaint, her similarity falls with every few edits, while with FLUX 3 it stays near 1.00.
The board and the measurement disagree
Artificial Analysis's image-editing board ranks 80 models by Elo from people's votes between two edits of the same picture, made from the same instruction. GPT Image 2.5 Sunburst is #1 there, Grok Imagine Image 2.0 #10 and Nano Banana Pro #13. FLUX 3 Image isn't on the board yet.
A vote on one edit rewards the result people prefer, and it never sees a second edit; Artificial Analysis doesn't say whether voters see the original picture to judge what else changed. Both measurements can be right at once: GPT Image 2.5 Sunburst can make the edit people prefer and still change 46% of the rest of the picture in that one edit, on our average. The board's row for it is at max quality, and ours ran at the model's defaults, so the two aren't the same setting.
Did they do what was asked?
A model that ignores an instruction keeps the picture perfectly, so we checked every edit by eye. FLUX 3 and GPT Image 2.5 Sunburst did 23 of the 24, and the one each missed is the same: changing a price on the bakery's chalkboard, the only edit that asked for new text. Nano Banana Pro and Grok Imagine did all 24, though asked to empty the shop window, each built new shelving in it. And asked to whiten the books on the top shelf, FLUX 3 and Grok Imagine whitened the books on every shelf, which our first look at the pictures missed. Part of that extra change, on the second shelf, is outside every region we asked about, and the pixel scores count it. The lower shelves fall inside the region around the lamp, edited at the fourth step, which the score stops looking at from then on, so FLUX 3's portrait score counts only some of the over-edit.
The four models side by side
Board rank is Artificial Analysis's; cost, edits and face similarity are our run; the last two columns are from Artificial Analysis's board and the makers' own pages
| model | editing board | cost per edit | edits done | face similarity at the end | open weights | who owns the output |
|---|---|---|---|---|---|---|
| FLUX 3 Image | not on the board | $0.025 | 23 | 0.99 | licensed for self-hosting on request (BFL) | you (BFL developer terms) |
| GPT Image 2.5 Sunburst | #1 | $0.019 to $0.074 | 23 | 0.59 | none listed | you (OpenAI Services Agreement) |
| Nano Banana Pro | #13 | $0.136 | 24 | 0.74 | none listed | Google won't claim it (Gemini API terms) |
| Grok Imagine Image 2.0 | #10 | $0.070 | 24 | 0.66 | none listed | you (xAI enterprise terms) |
FLUX 3 was also the cheapest on average, at $0.025 an edit, and Nano Banana Pro the dearest, at $0.136. The 24 edits for each of the four models cost $6.65 in all, and the whole run $6.83.
What this doesn't tell you
It is one run per model and picture, and we sent no random seed, so a second run would give somewhat different numbers. We'd be surprised if it closed a gap this size between FLUX 3 and the rest, but we haven't run it twice.
Every model ran at its defaults with a plain-language instruction. FLUX 3 can take boxes that tell it what to keep, and our protocol notes a fidelity setting for GPT Image 2.5 that OpenRouter doesn't pass on; we used neither. Either could change that model's numbers.
There was one text edit, and it split the models two against two, so this says little about editing text in pictures.
Keeping the rest of the picture is one thing you might want from an editor. A model that repaints more can still give you the picture you wanted.