No model beats 50% on shopping in the new ACE consumer benchmark

While AI handles logic problems and writes code with confidence, in real life people increasingly ask it about something far more mundane: what to buy at the store, what to substitute for an ingredient, how to fix a leaking faucet, which build to pick in a game. And here an inconvenient fact surfaces: everyday requests are simple only in the telling. They hang on context, on current data from the web, on safety, compatibility, prices and links. And on the ability not to make things up.
The authors of The AI Consumer Index (ACE) propose measuring exactly that: how ready AI models are to carry out high-value consumer tasks, where an error costs time, money or risk.

A benchmark about ordinary people, not olympiad questions
ACE is built as a set of scenarios across four domains: Shopping, Food, Gaming and DIY. The full version keeps 400 tasks hidden (100 per domain) — deliberately, so models cannot learn the answers in advance. A small open dataset of 80 cases is published separately for researchers.
The key idea: the tasks are not abstract. Every prompt carries a persona — who the user is and the conditions they are acting under. Without it, an answer easily comes out averaged, correct on paper and useless in practice. The authors admit that for the sake of fair scoring they appended a short specification of expectations to the end of each prompt: people rarely phrase things that way in real life, but for comparing models it removes ambiguity.

How they built it
The dataset was assembled by hand by domain specialists: chefs and nutritionists, builders and engineers, professional players and developers, stylists and shopping-media editors. 47 experts took part in all, and every case went through several rounds of review.
Crucially, for each query the experts write a rubric — a set of checkable criteria. For example: are there concrete steps, are the constraints respected, are safety warnings given, does the product match the requirements, are links provided.

The most interesting part is how they catch plausible fabrications
ACE does not stop at checking that an answer looks reasonable. For Shopping and Gaming, many criteria require grounding — a connection to reality: the model has to rely on web sources it actually found rather than invent them.
Scoring is hierarchical. First comes the core of the task — criteria asking whether the model really solves the user's problem rather than piling up helpful advice that misses the point. Fail those and the task is zeroed out. That guards against an answer collecting points on secondary details while never doing the main thing.
After that, if a criterion requires grounding, a penalty rule kicks in: backed by sources is a plus, unbacked is a minus. Lying with confidence becomes worse than honestly making no claim.

What the results showed
The authors ran 10 leading AI models with web search enabled. Each prompt was executed 8 times to smooth out randomness, and a separate LLM judge graded the answers. That came to 32,000 answers and more than 220,000 criterion checks.
The best averages across all domains land around 56%: the leader is GPT 5 (Thinking = High) — 56.1%, with o3 Pro — 55.2% and GPT 5.1 — 55.1% just behind. The spread across domains is the telling part. In Food the models are on firmer ground (70.1% for the leader), while in Shopping even the best result is under 50% (45.4%). That tracks closely with what users sense intuitively: AI gives decent recipes and general advice, but shopping comes down to exact prices, availability, product versions and links that actually work.
Separately, the authors show the gap between how well models satisfy the stated requirements and how well they can substantiate facts. Some models behave as if they are determined to be useful at any cost — including the cost of invention.

Why it matters and where the benchmark pushes the industry
ACE brings out an uncomfortable truth: being able to talk and being able to help are not the same skill. In consumer scenarios the value sits in the details: the exact price, a working link, a part that actually fits, the right safety warning, tastes and restrictions taken into account. That is precisely where models stumble, because generating a plausible answer is easier than carefully verifying a fact.
The authors are candid about the limits: grounding checks can misfire given how varied websites are, the internet shifts and forces regular re-runs, and publishing the dataset and the methodology in principle makes it possible to tune to the test. But on the whole ACE reads as a step toward a more grown-up evaluation of AI as a dependable personal assistant for everyone.
AI paper breakdowns
Every day we read the new AI papers and retell what matters in plain language — no hype, no filler. If you want to see where AI agents are heading before everyone else, subscribe.
New breakdowns every day.
On Telegram