i
DATAIST
Review · 2026-08-08

85.7% of what AI cites about a brand comes from someone else's site

85.7% of what AI cites about a brand comes from someone else's site

Who actually tells AI about your brand

When you ask an LLM about a company, the model rarely answers from memory. It pulls sources off the web first, then assembles an answer. And that is where it gets interesting: a brand's reputation with AI is shaped less by the company's own site than by everyone else's.

A new study is about what exactly the models cite. Instead of asking "what did the AI say about the brand?", the researcher asks a different question: "where did the AI get that in the first place?"

For marketing, for search and for everything now filed under generative engine optimization, this matters. If the model mostly reads your site, that calls for one strategy. If it reads Wikipedia, YouTube, job boards and business media, that calls for another.

The paper's conclusion is simple: AI builds its picture of a brand mainly out of the external web. And that web is lopsided: a few domains carry enormous weight, while a long tail of sources is all but invisible.

What was studied

The study is assembled from three Rankfor AI datasets. Together they cover 128 brands, 12 home markets, 13 languages and 167,551 citations with URLs — links where a specific source domain can be identified.

The author was after four questions:

🟠 how many links point to brands' own sites and how many to outside resources

🟣 how concentrated the pool of sources is

🟠 which domains lead by language and by market

🟣 how differently the various models behave

The method was fairly direct. Every link was reduced to a domain, Wikipedia's language subdomains were merged into a single `wikipedia.org`, and each source was then tagged by type: company site, Wikipedia, news, social media, government site and so on.

There is one important technical detail. With Gemini, some links ran through a Google redirect service. Leave that as it is and you get a false picture, in which the leading source appears to be a Google utility domain. The author reconstructed the real domains separately, from the fields holding link titles. Without that cleanup, the per-model and per-language results would have been wrong.

86% of a brand's AI reputation comes from outside

The most striking number in the paper: 85.7% of citations point to third-party sites. Brands' own domains account for only 14.3%.

So when AI answers "who can I trust", "which service should I pick" or "what is known about this company", it is mostly leaning on what other people have written about the brand.

Even for the brands cited through their own sites more often than most, self-citation stays a minority share. And for a number of companies, their own sites never showed up among the sources at all.

The practical reading is blunt. If you want to influence what AI says about your company, one good corporate site is not enough. You need mentions, reviews, profiles, articles and listings on the sites the models actually read.

AI's sources are heavily concentrated

The second finding: web sources are distributed very unevenly.

🟣 80% of all citations come from roughly 18% of domains

🟠 half of all citations fit into a few hundred hosts

🟣 the remaining thousands of domains form a long tail with little visibility

The authors also checked the distribution against Zipf's law. It is the classic shape for many information systems: a few leaders take the lion's share of attention, and the drop-off after them is steep.

Why does that matter? Because competing for visibility in AI search means competing to be inside the small group of domains the models read most often. If your brand is prominent on those sites, your chance of shaping AI answers grows out of all proportion.

The breakdown by source type is skewed in an equally legible way. The largest category is the open third-party web: local media, industry sites, blogs, marketplaces, directories, everything that did not fit a tidy label. It accounts for 77% of URL citations. Wikipedia supplies about 4%, company sites about 16%, and every other type is tiny on its own.

Wikipedia wins almost everywhere

Look at the most-cited domains by language and the picture is remarkably stable. Wikipedia is the number one domain in 11 of 12 languages.

This is one of the paper's most actionable results. Wikipedia has no fashionable aura about it, but for LLMs it remains the default reference work. If a brand's article is thin, out of date or missing, that directly shapes the material AI has to build an answer from.

There is exactly one exception: Lithuanian. There, Wikipedia is beaten by the local business outlet `vz.lt`.

It is a good illustration of how AI's common layer really is global while its edges shift by market. The model almost always leans on Wikipedia, but inside a particular language environment local business media can push it aside, if they are prominent and relevant enough.

In the overall domain ranking, the top looks like this: Wikipedia, YouTube, Statista, then individual brand sites, Reddit, European government resources and other large platforms.

Out of this comes a clear picture: AI assembles a brand not only from official corporate pages but from a mix of encyclopedia, video, statistics sites, forums and local publications.

Poland is the outlier: AI reads job boards

The most interesting local anomaly in the paper is the Polish market. There, across 46 national brands, the most-cited domain turned out to be YouTube, not Wikipedia.

Something else is more interesting still. Four Polish recruitment and careers sites together produced 637 citations, against 297 for Polish Wikipedia. Job platforms beat the encyclopedia by more than two to one.

What does that mean in practice? If you operate in a market where the brand is actively discussed in the context of hiring, working conditions, salaries, vacancies and employee reviews, AI may be assembling its image of the company from exactly there. Not from the "About us" page, but from the environment around you as an employer.

For a business, that is an uncomfortable but useful signal. A brand's AI reputation is not only product, PR and news. It is also how the labor market describes you.

Models cite differently

A separate section of the paper covers the differences between models. There is no single "AI" behavior. Different systems read and cite differently.

🟠 Perplexity cites the most

🟣 Perplexity covers the widest set of domains

🟠 GPT draws on a narrower set of sources

🟣 Gemini, without redirect cleanup, gives a distorted picture of its own sources

By the paper's numbers, Perplexity produced 90,276 of the 131,514 citations in the main part of the corpus and pointed to nearly 16,000 domains. That is well ahead of the other models in the sample. It also had the highest share of links to brands' own sites.

Gemini showed a different kind of problem: if you do not unpack the redirect links, you might conclude that the model barely links to company sites at all. Restore the real domains and the picture changes.

Which gives a rule for any audit: if you want to understand how AI sees your brand, you cannot look at one model only, and you cannot trust raw links without normalization. Otherwise you are measuring data-pipeline artifacts rather than how the system actually behaves.

Why this matters

Over the past two years, plenty of people have started measuring brand visibility in LLMs through the answers: were you mentioned or not, recommended or not, in what tone. This paper shows it is more useful to look one level down, at the source layer.

Because an answer can be rephrased. A source is harder to move. If the model reads Wikipedia, YouTube, the local press and job boards over and over, that is where the raw material for future answers is forming.

For companies, several direct conclusions follow:

🟣 your own site is necessary, but it is not AI's main source

🟠 working on Wikipedia remains a baseline task in almost every market

🟣 local media and specialist platforms can matter more than global resources

🟠 your reputation as an employer shapes what AI says about the brand

🟣 visibility has to be checked market by market and model by model, not in one averaged slice

There is a broader point too. AI search is becoming the new intermediary between the user and the web. Which makes "who ended up in the sources" steadily as important a question as "who ended up at the top of conventional search".

Limitations worth keeping in mind

The paper describes several limitations, carefully.

First, the three datasets were not collected in quite the same way. This is not one perfectly uniform table but a merged corpus with varying depth of labeling.

Second, not every source was available as a clean URL. Part of the Central and Eastern European data was labeled by keyword rather than by domain, so the main analysis effectively rests on two large subsets.

Third, the "brand's own site" heuristic hinges on the brand name appearing in the domain. That is sensible but imperfect: it will miss some cases and count others debatably.

Finally, the study measures how models behave when they cite, not a brand's actual reputation in people's eyes. Those are not the same thing. But in a world where AI answers are increasingly the first contact anyone has with a brand, that measurement already matters a great deal on its own.

The bottom line

If you want to influence what LLMs say about your company, you have to look not only at the answers but at the sources underneath them.

The paper shows three durable things. One: AI mostly reads third-party sites, not brand sites. Two: influence is concentrated in a small number of domains. Three: the global picture is built around Wikipedia almost everywhere, but local markets can shift the balance sharply.

A practical logic follows. The work is not only on your own content but on the external web: encyclopedias, business media, video platforms, industry directories, careers sites, local publications. That is where LLMs most often "read" a brand.

Visibility in AI search looks less and less like classic promotion of one site on one domain. It is now the management of your presence across the network of sources the model assembles its version of reality from.

AI papers in plain words

Every day we read the new AI papers and retell what matters in plain language — no hype, no filler. If you want to see where AI agents are heading before everyone else, subscribe.

New breakdowns every day

On Telegram