Six widely used language models answered questions about global development indicators correctly 21.2% of the time. The figure comes from a UNICEF benchmark covering more than 133,000 answers, and UNICEF chief statistician João Pedro Azevedo gave it to reporters on a virtual briefing alongside the launch of a shared UN statistics platform built on Google's Data Commons. Twenty-six UN entities have joined, data from nearly 20 is available at launch, and the organization wants 80% of the statistical datasets across the UN system on the platform by 2027.
The models tested were OpenAI's GPT-4o and GPT-4o-mini, Anthropic's Claude Sonnet 4.5 and Haiku 4.5, and Google's Gemini 2.5 Flash and Gemini 2.0 Flash. In roughly three cases out of five they produced no usable number at all, often sidestepping a precise answer. When the same questions went back to the same model versions about two days later, the models that gave a number both times agreed with their own earlier answer only about half the time.
That second finding is the harder one. An accuracy score can be improved with better data; a model that contradicts itself across two days on a fixed question is unstable in a way that no amount of source quality fixes.
The study is a working paper being prepared for journal submission and has not been peer reviewed. UNICEF says it will publish the methodology, code and data with it.
The timing is not academic. UNICEF's website draws more than 6 million visits a month and is among the organization's most-used resources. Between January 1 and September 14, referrals from links inside ChatGPT answers rose 67% year over year, Azevedo said. Those referrals accounted for 6.4% of all sessions in that period, and UNICEF puts the share of its total visits arriving from AI assistants at roughly 10%.
Set the two numbers next to each other and the problem states itself: about a tenth of the traffic to one of the UN's main public data properties now arrives through systems that, asked cold, get development statistics right one time in five.
The platform underneath the fix is Google's Data Commons, launched in 2018 to pull public datasets from scattered sources into a common structure. Last year it gained MCP support, the standard that lets AI agents query Data Commons directly for statistics and their provenance. The UN version runs in an instance under UN management. Google.org contributed $2 million for the capability building and technical support behind the base infrastructure, and Prem Ramaswami, who leads Google's Data Commons team, told TechCrunch the intent is for the UN to maintain, operate and scale it internally from here.
Shantanu Mukherjee, acting director of the UN Statistics Division, said the system unites data from this many agencies for the first time and makes it available to AI at large scale, with broad capabilities and a flexible structure.
At launch Google used a train-the-trainers approach: staff who learned the system pass it on to other teams. Ramaswami said UN system staff got up to speed quickly.
Every indicator carries its origin with it, so a user can trace a number an AI system produced back to the original UN source. Azevedo called that especially important as people lean harder on AI tools to find and interpret information.
Agents can do more than fetch single indicators. Google demonstrated a system connected to UN data over MCP that combines several indicators into dashboards, charts and written analysis without anyone manually locating and reconciling the underlying datasets. In one demo, Google asked the system to find the effect of the US President's Emergency Plan for AIDS Relief on Africa; it pulled relevant UN indicators including HIV infections, AIDS deaths and life expectancy, and generated an infographic.
Two parts of this look thinner than the launch suggests. The first is the money. $2 million covers capability building and technical support, not a permanent home for the statistical output of 26 UN agencies, and the plan already hands maintenance and scaling to the UN itself. Whatever this is, it is not the UN buying infrastructure; it is Google seeding a standard and leaving.
The second is what the benchmark measured. Asking a model a development question with no source attached tests recall, and recall of specific figures is a known weak spot, not a discovery. A retrieval layer addresses the retrieval half. It does not address the half where three answers in five contained no number at all. Refusal is the honest failure mode here, and the quiet risk of wiring models to a clean, authoritative source is that refusal turns into a confident figure — right most of the time, wrong the rest, with a provenance link attached either way.
Notably absent from the announcement is any commitment from an AI company to use the thing. MCP support means an agent can query UN data; nothing makes ChatGPT, Gemini or Claude reach for it when someone asks how many children are out of school. Google funded the platform and built the standard, which gives one vendor an obvious reason to connect it. The other two were in the benchmark, not in the announcement.
Ramaswami's closing caveat — that models can misread nuance, so a person should check the result before citing or publishing it — is sound, and it is also the one part of this system with no engineering behind it. The 67% growth in ChatGPT referrals counts only the readers who clicked through to look.