i
DATAIST
News · 2026-09-20

OpenArt ranks AI models by creative task, not one universal score

@neuronium_ai @neuronium_ai

OpenArt has launched OpenArt Arena, a public benchmark that ranks image and video models by creative task rather than by one universal score. The four-year-old startup, founded by former Google employees Coco Mao and John Qiao, is targeting a practical problem: teams choosing between GPT Image 2 from OpenAI, Google’s Nano Banana family, ByteDance’s Seedream and Seedance, Alibaba’s Wan, xAI’s Grok Imagine and other fast-moving systems. The first results put Seedance 2.5 at the top of most video boards, but the more important test is whether OpenArt can make a commercially embedded benchmark independent enough to guide procurement decisions.

Cover: OpenArt ranks AI models by creative task, not one universal score

What Arena measures

OpenArt says Arena will rank models for specific creative disciplines, including:

Animation
Advertising
Graphic design
E-commerce
Film
Motion design
Video editing
Lip sync
General tasks

The methodology uses blind, side-by-side comparisons. Reviewers see two outputs without model names and choose the better result. OpenArt then combines those preferences with the Bradley–Terry statistical model from 1952, which estimates the hidden probability that one option will beat another in direct comparisons.

The judging system has two groups:

Creative Experts Council — a smaller group that includes Emmy-winning animation director William Lau, creative technologist Willonius “King Willonius” Hatcher, marketing executive David Shing, and leaders and educators from Edelman and UCLA.
“Taste leaders” — a planned group of roughly 800–1000 OpenArt users, members of external creative communities and working professionals, including people from advertising.

The number of 800 to 1000 is a planned group size, not a confirmed count of people who completed the launch evaluations. OpenArt has not disclosed the final number of judges, the total number of pairwise comparisons or the number of prompts in each initial benchmark.

The company plans to publish the methodology and some prompts. It will keep some active prompts private so model developers cannot tune their systems directly against the test.

That trade-off is familiar. A completely open and static benchmark eventually measures how well developers optimize for the benchmark, rather than how well a model generalizes to new creative work. But closed prompts create a different problem: outsiders must trust that the hidden tasks represent professional workflows and that outputs are generated, selected and compared consistently.

OpenArt itself operates in the middle of this model market. Its platform provides access to many third-party image, video, audio and 3D systems, alongside products such as Director, a conversational workflow for making multi-scene videos through descriptions of the desired atmosphere.

The first rankings

The clearest signal from the launch is video.

ByteDance’s Seedance 2.5, a closed model available through an API with no published weights, took first place in the overall video ranking with 1 081 points. Alibaba’s Wan 3.0 followed with 1 004, ahead of ByteDance’s closed Seedance 2.0 at 1 000.

Seedance 2.5 also led four of the five video boards provided at launch. It won film with 1 049 points and motion design with 1 059. In lip sync, it scored 1 063, ahead of Seedance 2.0 at 1 000 and Seedance 2.0 Mini at 994.

1 081overall video
1 049film
1 059motion design

Video editing was the exception:

Wan 3.0 — first place, 1 034 points
Seedance 2.5 — second place, 1 033 points
Google Gemini Omni Flash — third place, 1 021 points

That one-point gap is exactly why the scores need to be read with their 95% confidence intervals. The ranking shows a narrow observed difference, not proof of a meaningful quality lead.

Wan 3.0 also complicates the usual open-versus-closed comparison. Alibaba calls it open and has published a repository under the Apache 2.0 license, but the current repository contains documentation and a license rather than model weights. It therefore cannot yet be treated as a conventional open release that customers can run on their own servers.

The image boards are less concentrated. GPT Image 2 from OpenAI won graphic design and image editing, scoring 1 000 in both. Seedream 5.0 Pro led film images with 1 014 points and e-commerce images with 1 004. In the overall image ranking, Seedream 5.0 Pro scored 1 010, ten points ahead of GPT Image 2.

OpenAI updated the model to GPT-Images-2.5 a week before the results, so at least part of the image ranking is already due for revision.

Why task-specific rankings matter

A single image or video leaderboard is convenient, but it is a poor procurement tool. A marketing department may need accurate packaging and logos, an internal studio may need a visual preproduction tool, and a post-production team may need to alter existing footage without damaging the original material.

OpenArt says the evaluation criteria change with the board. Film may emphasize camera movement, lighting, cinematic quality and realistic skin. Advertising may put more weight on readable text, accurate logos, product similarity and object placement.

That is a more useful model of enterprise adoption than asking which system is “best” at images or video. Stella Guan, OpenArt’s head of growth and operations, said the company wants Arena to become an industry standard for choosing a model for a particular image or video task.

Creators describe the same constraint from the other direction. The creator @BLVCKL!GHT told VentureBeat that looping animations and dynamic scenes require different tools; he might choose Minimax for one project and Seedance 2.5 for another. Clients may demand the newest models, but budget ultimately narrows the practical choice.

The launch results support that view:

Seedance 2.5 looks like a strong general-purpose video model, but it did not win every video task.
Wan 3.0 narrowly won video editing.
GPT Image 2 led graphic design and image editing.
Seedream 5.0 Pro led film and e-commerce images.

My reading is that Arena’s most valuable output is not a champion. It is a map of where the ranking changes when the work changes.

This is not a new idea. Contra Labs’ Human Creativity benchmark evaluates working creative professionals on landing pages, advertising images, brand images and product videos, with separate stages for ideation, layout and refinement. Arena.ai ranks image models across commercial design, cinematic imagery, portraits, text rendering, 3D and art, using large volumes of real user prompts. Artificial Analysis covers areas including marketing, e-commerce, film, animation, games, architecture, real estate, UI/UX, social media and user-generated content, while also publishing price, speed and openness data.

OpenArt’s distinction is organizational rather than foundational. It starts with a professional task, builds a dedicated prompt set and criteria, then asks selected judges to compare outputs. Arena.ai largely groups prompts according to community usage. Artificial Analysis goes broader on image categories and splits video mainly by generation mode or technical process. Contra focuses on a continuing research program around where model performance changes during a workflow.

OpenArt combines task-based boards for both images and video with a selected expert council, a broader judging group and a direct route from ranking to model access inside its commercial platform.

The missing denominator

The benchmark’s biggest weakness is not the statistical method. It is the amount of information missing around the statistics.

OpenArt has not yet published:

The number of prompts
The number of judges who actually completed evaluations
The total number of votes
The total number of pairwise comparisons

That makes it difficult to compare the launch with Contra’s published evaluator and comparison counts or Arena.ai’s large vote totals. A planned panel of 800–1000 people sounds substantial, but it does not tell readers how many people judged a specific board, how many comparisons each completed or how many prompts they saw.

OpenArt’s choice to keep some prompts private is defensible. I think the company needs that protection if it wants to measure generalization rather than benchmark-specific optimization. But it cannot use secrecy to avoid reporting the size and composition of the evidence. Those details are necessary for readers to interpret a one-point difference, a confidence interval or a category win.

The commercial context makes the omission more important. OpenArt says it brings together more than 100 models, including Seedance, Veo, Kling, Wan, GPT Image, Nano Banana, Seedream and Grok Imagine. It also sells access to many of those models and is building products such as Director on top of them.

Guan said Arena will appear inside OpenArt’s model-selection interface. That means the company will influence both the decision about which model to use and the subsequent generation work.

This does not establish that the launch results are biased. The published leaders do not appear to include an OpenArt-owned base model for image or video generation. But the incentive is clear: a user who chooses the wrong model may blame the platform, while a user who follows Arena’s recommendation stays inside OpenArt’s product.

Artificial Analysis explicitly describes itself as independent and says model providers do not pay for inclusion or favorable results. OpenArt is structurally different because it is a commercial creative platform that provides access to the systems it ranks.

The right conclusion is not to dismiss Arena. It is to treat it as a vendor-produced decision aid, not as a neutral authority. What I’d want to see next is the same discipline OpenArt asks of model developers: stable reporting on sample sizes, completed evaluations, comparison counts, confidence intervals and changes between benchmark versions.

Where Arena fits in a workflow

For a corporate creative team, the useful question is not which model won. It is whether a board resembles the team’s actual work:

Product images with exact packaging
Advertising materials with typography
Cinematic scenes with camera movement
Edits that preserve a client’s original footage

OpenArt’s task boards point in that direction. The company also plans to add business data such as price, content moderation and intellectual-property protection, and says it wants to show price-to-quality ratios rather than rank models without considering budget.

That matters because deployment constraints sit beside output quality:

Whether the model can run on a company’s own servers
Whether data can remain in a chosen region
Whether the system can be tuned
How dependent the team becomes on a provider or API
What the total cost of ownership looks like

For now, the best use of Arena is as a shortlist generator. A production team can identify promising models, then test them on its own brands, references, legal constraints and delivery standards. No public board can replace that internal evaluation.

OpenArt is trying to make model selection part of the creative workflow itself. That could make Arena more useful than a detached leaderboard. It also means the benchmark’s credibility will be judged alongside its commercial position: the closer the ranking gets to a purchasing decision, the less acceptable it becomes to leave the evidence behind the ranking underspecified.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X