What Arena measures
OpenArt says Arena will rank models for specific creative disciplines, including:
The methodology uses blind, side-by-side comparisons. Reviewers see two outputs without model names and choose the better result. OpenArt then combines those preferences with the Bradley–Terry statistical model from 1952, which estimates the hidden probability that one option will beat another in direct comparisons.
The judging system has two groups:
The number of 800 to 1000 is a planned group size, not a confirmed count of people who completed the launch evaluations. OpenArt has not disclosed the final number of judges, the total number of pairwise comparisons or the number of prompts in each initial benchmark.
The company plans to publish the methodology and some prompts. It will keep some active prompts private so model developers cannot tune their systems directly against the test.
That trade-off is familiar. A completely open and static benchmark eventually measures how well developers optimize for the benchmark, rather than how well a model generalizes to new creative work. But closed prompts create a different problem: outsiders must trust that the hidden tasks represent professional workflows and that outputs are generated, selected and compared consistently.
OpenArt itself operates in the middle of this model market. Its platform provides access to many third-party image, video, audio and 3D systems, alongside products such as Director, a conversational workflow for making multi-scene videos through descriptions of the desired atmosphere.
The first rankings
The clearest signal from the launch is video.
ByteDance’s Seedance 2.5, a closed model available through an API with no published weights, took first place in the overall video ranking with 1 081 points. Alibaba’s Wan 3.0 followed with 1 004, ahead of ByteDance’s closed Seedance 2.0 at 1 000.
Seedance 2.5 also led four of the five video boards provided at launch. It won film with 1 049 points and motion design with 1 059. In lip sync, it scored 1 063, ahead of Seedance 2.0 at 1 000 and Seedance 2.0 Mini at 994.
Video editing was the exception:
That one-point gap is exactly why the scores need to be read with their 95% confidence intervals. The ranking shows a narrow observed difference, not proof of a meaningful quality lead.
Wan 3.0 also complicates the usual open-versus-closed comparison. Alibaba calls it open and has published a repository under the Apache 2.0 license, but the current repository contains documentation and a license rather than model weights. It therefore cannot yet be treated as a conventional open release that customers can run on their own servers.
The image boards are less concentrated. GPT Image 2 from OpenAI won graphic design and image editing, scoring 1 000 in both. Seedream 5.0 Pro led film images with 1 014 points and e-commerce images with 1 004. In the overall image ranking, Seedream 5.0 Pro scored 1 010, ten points ahead of GPT Image 2.
OpenAI updated the model to GPT-Images-2.5 a week before the results, so at least part of the image ranking is already due for revision.
Why task-specific rankings matter
A single image or video leaderboard is convenient, but it is a poor procurement tool. A marketing department may need accurate packaging and logos, an internal studio may need a visual preproduction tool, and a post-production team may need to alter existing footage without damaging the original material.
OpenArt says the evaluation criteria change with the board. Film may emphasize camera movement, lighting, cinematic quality and realistic skin. Advertising may put more weight on readable text, accurate logos, product similarity and object placement.
That is a more useful model of enterprise adoption than asking which system is “best” at images or video. Stella Guan, OpenArt’s head of growth and operations, said the company wants Arena to become an industry standard for choosing a model for a particular image or video task.
Creators describe the same constraint from the other direction. The creator @BLVCKL!GHT told VentureBeat that looping animations and dynamic scenes require different tools; he might choose Minimax for one project and Seedance 2.5 for another. Clients may demand the newest models, but budget ultimately narrows the practical choice.
The launch results support that view:
My reading is that Arena’s most valuable output is not a champion. It is a map of where the ranking changes when the work changes.
This is not a new idea. Contra Labs’ Human Creativity benchmark evaluates working creative professionals on landing pages, advertising images, brand images and product videos, with separate stages for ideation, layout and refinement. Arena.ai ranks image models across commercial design, cinematic imagery, portraits, text rendering, 3D and art, using large volumes of real user prompts. Artificial Analysis covers areas including marketing, e-commerce, film, animation, games, architecture, real estate, UI/UX, social media and user-generated content, while also publishing price, speed and openness data.
OpenArt’s distinction is organizational rather than foundational. It starts with a professional task, builds a dedicated prompt set and criteria, then asks selected judges to compare outputs. Arena.ai largely groups prompts according to community usage. Artificial Analysis goes broader on image categories and splits video mainly by generation mode or technical process. Contra focuses on a continuing research program around where model performance changes during a workflow.
OpenArt combines task-based boards for both images and video with a selected expert council, a broader judging group and a direct route from ranking to model access inside its commercial platform.
The missing denominator
The benchmark’s biggest weakness is not the statistical method. It is the amount of information missing around the statistics.
OpenArt has not yet published:
That makes it difficult to compare the launch with Contra’s published evaluator and comparison counts or Arena.ai’s large vote totals. A planned panel of 800–1000 people sounds substantial, but it does not tell readers how many people judged a specific board, how many comparisons each completed or how many prompts they saw.
OpenArt’s choice to keep some prompts private is defensible. I think the company needs that protection if it wants to measure generalization rather than benchmark-specific optimization. But it cannot use secrecy to avoid reporting the size and composition of the evidence. Those details are necessary for readers to interpret a one-point difference, a confidence interval or a category win.
The commercial context makes the omission more important. OpenArt says it brings together more than 100 models, including Seedance, Veo, Kling, Wan, GPT Image, Nano Banana, Seedream and Grok Imagine. It also sells access to many of those models and is building products such as Director on top of them.
Guan said Arena will appear inside OpenArt’s model-selection interface. That means the company will influence both the decision about which model to use and the subsequent generation work.
This does not establish that the launch results are biased. The published leaders do not appear to include an OpenArt-owned base model for image or video generation. But the incentive is clear: a user who chooses the wrong model may blame the platform, while a user who follows Arena’s recommendation stays inside OpenArt’s product.
Artificial Analysis explicitly describes itself as independent and says model providers do not pay for inclusion or favorable results. OpenArt is structurally different because it is a commercial creative platform that provides access to the systems it ranks.
The right conclusion is not to dismiss Arena. It is to treat it as a vendor-produced decision aid, not as a neutral authority. What I’d want to see next is the same discipline OpenArt asks of model developers: stable reporting on sample sizes, completed evaluations, comparison counts, confidence intervals and changes between benchmark versions.
Where Arena fits in a workflow
For a corporate creative team, the useful question is not which model won. It is whether a board resembles the team’s actual work:
OpenArt’s task boards point in that direction. The company also plans to add business data such as price, content moderation and intellectual-property protection, and says it wants to show price-to-quality ratios rather than rank models without considering budget.
That matters because deployment constraints sit beside output quality:
For now, the best use of Arena is as a shortlist generator. A production team can identify promising models, then test them on its own brands, references, legal constraints and delivery standards. No public board can replace that internal evaluation.
OpenArt is trying to make model selection part of the creative workflow itself. That could make Arena more useful than a detached leaderboard. It also means the benchmark’s credibility will be judged alongside its commercial position: the closer the ranking gets to a purchasing decision, the less acceptable it becomes to leave the evidence behind the ranking underspecified.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X