When this post went up in April 2026, GPT Image 2 had just topped the image leaderboard by 241 points. That was the story: a record-breaking gap, one model far ahead of everything else.
Three months later the gap is 37 points, five models that did not exist in the original table are now in the top 10, and Midjourney has fallen to 89th place while Grok Imagine sits at 12th. This is the refreshed version. Every number below was re-verified on 27 July 2026.
The leaderboard, and why the source changed
The original version of this post was built on an Arena.ai snapshot from 22 April 2026. Arena has since restructured. Its image leaderboard is no longer reachable at a stable URL, the tab that should hold it renders client-side without producing a table, and the site rate-limited me when I kept probing. I could not re-verify a single Arena figure, so I stopped citing them rather than presenting three-month-old numbers as current.
Every figure in this post now comes from the Artificial Analysis Text to Image leaderboard, which the post already used as a secondary source. It ranks 140 models and publishes, for each one, an Elo rating with a 95% confidence interval, the number of votes behind it, the release date, and the current API price per 1,000 images. That last column is why it is a better primary source for this post than a pure preference board.
Elo works the same way it does in chess. Two models generate images from the same prompt, a human picks the one they prefer without knowing which is which, and ratings move based on who beat whom. A higher Elo means more people preferred that model's output in head-to-head matchups.
It has three blind spots worth holding in mind while you read the table. It does not weight cost, so a model at $211 per 1,000 images ranks against one at $12 as if they were the same purchase. It does not weight speed. And it under-samples anything without a public API, because the people voting in volume are mostly testing models they can call programmatically. That third point does most of the work in the Midjourney section below.
The full comparison table
The current top 10, plus the two models this post's readers actually search for: Grok Imagine at 12th and Midjourney v7 at 89th. Rank is shown so the jump is visible rather than hidden. Elo, samples, release date and API pricing are taken directly from Artificial Analysis on 27 July 2026. "Best for" is my own assessment, not a leaderboard measurement.
| Rank | Model | Creator | Elo | Samples | Released | API / 1k imgs | Best for |
|---|---|---|---|---|---|---|---|
| 1 | GPT Image 2 (high) | OpenAI | 1,338 | 21,121 | Apr 2026 | $211.0 | Text-heavy marketing assets |
| 2 | Reve 2.1 | Reve | 1,301 | 9,382 | Jul 2026 | $24.0 | Top quality at pipeline cost |
| 3 | MAI-Image-2.5 | Microsoft AI | 1,270 | 9,366 | Jun 2026 | $48.1 | Microsoft stack teams |
| 4 | Nano Banana 2 Lite | 1,262 | 9,298 | Jun 2026 | Coming soon | Not priced yet, one to watch | |
| 5 | Nano Banana 2 | 1,261 | 19,377 | Feb 2026 | $67.0 | General work on Gemini | |
| 6 | GPT Image 1.5 (high) | OpenAI | 1,260 | 12,219 | Dec 2025 | $133.0 | Cheaper OpenAI tier |
| 7 | HiDream-O1-Image-1.5 | HiDream | 1,246 | 12,408 | Jun 2026 | $80.0 | Hosted, open weight siblings |
| 8 | Seedream 5.0 Pro | Bytedance | 1,239 | 8,151 | Jul 2026 | $90.0 | Newest, still thin on votes |
| 9 | Nano Banana Pro | 1,223 | 11,527 | Nov 2025 | $134.0 | Higher-fidelity Gemini | |
| 10 | Cosmos3-Super-Text2Image | NVIDIA | 1,218 | 9,828 | May 2026 | Coming soon | Open weights. Self-hosting |
| 12 | grok-imagine-image-quality | xAI | 1,203 | 15,493 | Apr 2026 | $50.0 | Quality at a quarter of the price |
| 89 | Midjourney v7 Alpha | Midjourney | 1,068 | 4,331 | Apr 2025 | No API | Directed creative work, no API |
Source: Artificial Analysis Text to Image leaderboard, captured 27 July 2026. Prices are per 1,000 images, so $211.0 is roughly $0.21 per image and $24.0 is roughly $0.024. Artificial Analysis lists Grok Imagine's creator as SpaceXAI; the model is xAI's. Midjourney has no public API and requires a subscription of $10 to $120 per month. "Coming soon" means the model is ranked but not yet priced for API access.
Also ranked: the models you searched for
Several models from the original table have been pushed out of the top 10 by newer releases, and a few readers arrive looking specifically for them. Here is where they actually sit today.
| Rank | Model | Elo | Released | API / 1k imgs |
|---|---|---|---|---|
| 13 | Recraft V4.1 Utility | 1,203 | May 2026 | $35.0 |
| 15 | FLUX.2 [max] | 1,195 | Dec 2025 | $70.0 |
| 16 | Seedream 4.0 | 1,190 | Sept 2025 | $30.0 |
| 22 | FLUX.2 [pro] | 1,185 | Nov 2025 | $30.0 |
| 24 | grok-imagine-image | 1,179 | Jan 2026 | $20.0 |
| 27 | Qwen Image 2.0 Pro | 1,172 | Apr 2026 | $75.0 |
| 28 | Ideogram 4.0 (Quality) open weights | 1,172 | Jun 2026 | $100.0 |
| 29 | Imagen 4 Ultra | 1,171 | Jun 2025 | $60.0 |
| 38 | FLUX.2 [dev] Turbo open weights | 1,153 | Dec 2025 | $8.0 |
| 85 | Ideogram 3.0 | 1,077 | Mar 2025 | $60.0 |
| 110 | Stable Diffusion 3.5 Large Turbo open weights | 1,022 | Oct 2024 | $40.0 |
Two of these deserve a flag. FLUX.2 [dev] Turbo at 38th for $8 per 1,000 images is the best quality-per-dollar entry on the whole board, and it ships open weights. And Ideogram 4.0 (Quality) at 28th sits 57 Elo points above the Ideogram 3.0 this post originally listed, so if you came here for Ideogram's text rendering, you want the 4.0 line.
How the 241-point gap became 37
In April, GPT Image 2's lead looked structural. Positions 2 through 10 spanned roughly 100 points and the leader sat 241 clear of all of them. The obvious read was that OpenAI had found something nobody else had.
The board today says otherwise. GPT Image 2 (high) still leads at 1,338, but Reve 2.1 arrived in July at 1,301 and the rest of the top 10 is packed into the 100 points below it. What closed the gap was not GPT Image 2 slipping. It was five new models landing in three months: Reve 2.1, Microsoft's MAI-Image-2.5, Nano Banana 2 Lite, HiDream-O1-Image-1.5 and Seedream 5.0 Pro. None of them existed in the original table.
The commercial consequence is sharper than the ranking one. Second place now costs $24 per 1,000 images against $211 for first. You are paying roughly nine times as much for 37 Elo points, which is inside the range where a blind voter would struggle to call the winner consistently. In April, paying the premium for the leader was defensible on quality grounds. In July it is a much harder argument to make for anything running at volume.
It is also worth reading the confidence intervals. Ranks 3 through 6 sit within 10 points of each other with intervals of plus or minus 7 to 10, which is why Artificial Analysis publishes a rank range rather than a single position for most models. Treat the top 10 as a group, not an order.
For the full breakdown of GPT Image 2's capabilities, limitations, and pricing, I covered everything in the companion post: GPT Image 2 Is Here: What Changed, What It Costs, and What Builders Should Know.
Grok Imagine is #12. Midjourney is #89.
This is the most striking line in the data, and the one most likely to be misread. Grok Imagine's quality tier ranks 12th at 1,203. Midjourney v7 Alpha ranks 89th of 140 at 1,068. A model that costs $50 per 1,000 images through an API is beating the tool that most brand and editorial designers still open first.
That does not mean Midjourney got worse. Three things explain most of the distance.
Elo measures general preference, not artistic quality. Voters compare two outputs on generic prompts, usually without a brief, a house style, or a reference. That test rewards clean, literal, well-composed images. It does not reward the thing Midjourney is actually good at, which is holding a directed aesthetic across a set.
Midjourney has no public API. It collected 4,331 votes. GPT Image 2 collected 21,121 and Grok Imagine's quality tier collected 15,493. Models you can call in a loop get tested in a loop. Models behind a subscription and a chat interface do not, which both thins the sample and skews who is voting.
v7 Alpha shipped in April 2025. It is the oldest model in the main table by nine months. Everything above it in the top 10 was released between November 2025 and July 2026. Comparing a fifteen-month-old model to a two-month-old one and concluding the older one is bad skips the obvious explanation.
The honest reading is that the two models are optimised for different jobs and the leaderboard only scores one of them. If your work is a directed creative brief, Midjourney's rank tells you very little. If your work is an automated pipeline, Midjourney's rank is almost beside the point, because the missing API already disqualified it.
Category winners
GPT Image 2 (high) (OpenAI). Elo 1,338 across 21,121 samples, the largest vote count on the board. Still first, but by 37 points rather than 241.
Reve 2.1. Second place at 1,301 for $24 per 1,000 images. Nothing else on the board gets that close to the leader at that price.
GPT Image 2, on OpenAI's claimed 99% accuracy at launch. Specialist runner-up: Ideogram 4.0 (Quality), 28th at 1,172, not the 3.0 version at 85th.
Midjourney v7, still, despite ranking 89th. Blind preference voting on generic prompts does not measure directed aesthetic control. Read the section above before treating the rank as a verdict.
Recraft V4.1 Utility, 13th at 1,203 for $35 per 1,000 images. Recraft is still the only family with true SVG export, and 4.1 sits 68 points above the V4 this post originally listed.
FLUX.2 [dev] Turbo. 38th at 1,153 for $8 per 1,000 images, with open weights. Runner-up: grok-imagine-image at 24th for $20.
NVIDIA Cosmos3-Super-Text2Image. 10th at 1,218, the first open weight model to reach the top 10 on this board. API pricing is not live yet, so self-hosting is the route today.
Reve 2.1 (Reve). Released in July 2026 and straight into second place. Worth a test if you last evaluated image models before this summer.
Picking the right model: a decision framework
Leaderboards rank models. Decision frameworks pick the right one. Here is how I think about model selection for production work, updated for the July board.
GPT Image 2, or Ideogram 4.0 (Quality) if you want the specialist. Everything else is still unreliable for text-heavy assets.
Midjourney v7, and ignore the 89th place ranking for this use case. If you need API access, FLUX.2 [max] at 15th is the closest alternative.
Reve 2.1 ($24 per 1,000) or FLUX.2 [pro] ($30). Both hold high positions at a price that survives volume. GPT Image 2 at $211 rarely does.
Recraft V4.1. Still the only family that exports true SVG, and nothing else in this list comes close for vector work.
FLUX.2 [dev] Turbo at $8 per 1,000 images is the best cheap option that still ranks well. grok-imagine-image at $20 is the next step up. Below that, quality drops faster than price does.
This is the category that changed most. NVIDIA Cosmos3 is now 10th, and FLUX.2 [dev], HiDream's Dev line, Ideogram 4.0 and Qwen all ship open weights. You no longer give up much to run your own hardware.
Google: Nano Banana 2 (5th). OpenAI: GPT Image 2 (1st) or GPT Image 1.5 (6th) at 63% of the price. Microsoft: MAI-Image-2.5 (3rd). All three are inside the top 10, so staying put costs you very little.
What I would pick for client work
I build AI systems for businesses. Most of my work is text-based automation (content pipelines, lead scoring, reporting), but image generation increasingly shows up in client workflows. Here is how I would split the choices today, with the July pricing in front of me.
Marketing assets with text (social posts, email banners, infographics): GPT Image 2. Text accuracy is still the reason, and it is the first image model I trust for assets that need readable copy. At $0.21 per image it is a considered purchase rather than a default, so I use it where the text has to be right and something cheaper everywhere else. I covered the full capability breakdown in the GPT Image 2 launch post.
Creative brand work (mood boards, brand photography, art direction): Midjourney v7. The 89th place ranking has not changed my answer here, for the reasons in the section above. The lack of API access is a real limitation for automation, but for directed creative work it is still the tool I reach for.
High-volume automated generation (product thumbnails, catalogue images, batch processing): Reve 2.1 or FLUX.2 [pro]. This is the pick that changed most in three months. Both sit high on the board at $24 to $30 per 1,000 images, which is where the maths works for a pipeline generating thousands of assets.
Logo and vector needs: Recraft V4.1. True SVG export is a genuine differentiator. Nothing else in this list can do what Recraft does for vector work.
Anything with a data residency or cost floor: FLUX.2 [dev] or NVIDIA Cosmos3, self-hosted. Open weights reaching 10th place is the quietest important change on this board, and it makes self-hosting a real option rather than a compromise.
If you are evaluating which image model fits your business workflow, or if you need help integrating image generation into an existing automation stack, the AI Consulting and Roadmapping service is where I help teams make these decisions.
Sources and credits
All leaderboard data on this page was verified on 27 July 2026. Elo ratings move as votes accumulate and new models arrive, so treat the exact numbers as a snapshot rather than a fixed state. The ranking order in the top 10 is inside the margin of error and should be read as a group.
- Artificial Analysis, Text to Image Leaderboard (primary source, 140 models, captured 27 July 2026). Elo, 95% confidence intervals, sample counts, release dates and API pricing.
- Arena.ai (original primary source for the April 2026 version of this post). Arena has restructured and its image leaderboard is no longer reachable at a stable URL, so those figures have been removed rather than carried forward unverified.
- ElevenLabs GPT Image 2 comparison (YouTube).
- OpenAI (GPT Image 2 launch claims, including the 99% text rendering figure). Covered in full in the GPT Image 2 launch post.
This post does not cover DALL-E. OpenAI's current image models are the GPT Image line, which is what appears on the leaderboard and in the table above.