Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study
According to BusySeed, 88.4% of the company recommendations AI assistants make for a buyer question come from only one of three engines, and just 3.2% are made by all three.
According to BusySeed, 88.4% of the company recommendations AI assistants make for a buyer question come from only one of three engines, and just 3.2% are made by all three.
What to take away
- Across 56,819 company-prompt pairs, 88.4% were named by only one of the three engines, and just 3.2% were named by all three.
- Agreement is not binary and not rare in the middle: 8.4% of pairs were named by exactly two engines, showing a gradient rather than a simple visible/invisible split.
- Visibility rates differ by engine: ChatGPT surfaced 42.7% of companies it was asked about, Claude 39.9%, and Gemini 33.6%.
- No company in the 30,564-company cohort was invisible on every engine (0% never appeared anywhere), but 96% appeared on at least one engine without appearing on all of them.
- Average mention position is close across engines (Gemini 4.85, ChatGPT 5.25, Claude 5.38), so where engines disagree, it is mostly about whether a company is named, not where.
- A single 'AI visibility score' built from one engine will systematically miss the 88.4% of company mentions that only show up on a different engine.
Key findings
- 88.4% of company recommendations were made by only one of the 3 engines
- 3.2% were made by all 3 engines for the same question
- 5.18 average mention position where they do appear, across 65,241 appearances
- 42.7% of the cohort is visible on ChatGPT
- 39.9% of the cohort is visible on Claude
- 33.6% of the cohort is visible on Gemini
Why this question matters
A buyer researching, say, a blockchain testing vendor or a commercial cleaning service increasingly starts with one AI assistant, not four. They ask ChatGPT, or they ask Claude, or they ask Gemini. They rarely ask all three and compare notes. That means whichever companies a given engine happens to surface become, for that buyer, the entire visible market. If a company's presence in AI-generated answers varies sharply by which engine is asked, then a company's fate in this new channel depends partly on an accident of which assistant the buyer opened.
This matters most to two groups. The first is companies trying to manage their presence in AI answers, who need to know whether optimizing for one engine's behavior helps or does nothing for the others. The second is anyone buying or building an 'AI visibility' monitoring product, who needs to know whether a single-engine score is a reasonable proxy for the whole market or a measurement of one-third of it.
The question is also a check on a common assumption: that large language models, trained on overlapping web corpora and asked the same question, will converge on a similar shortlist of companies the way three independent analysts reading the same industry reports might converge on the same handful of leaders. If that assumption held, a single engine's answer would be a reasonable stand-in for 'what AI recommends.' If it does not hold, the entire idea of a portable, engine-agnostic AI visibility score needs to be qualified with which engine, or engines, it is actually measuring.
This study does not ask whether the companies named are good companies, or whether being named causes more business. It asks a narrower, checkable question: given the same prompt, do three engines converge on the same names. The answer, measured across 30,564 companies and 170,457 answers, is that they mostly do not, and the rest of this study describes the shape of that disagreement in more usable terms than 'mostly.'
How the measurement works
The design starts from a fixed prompt set, 2,500 distinct prompts covering commercial-investigation, informational, navigational, and transactional intents, spanning categories such as SaaS, Legal Services, Real Estate Investment, and Manufacturing Services. Each prompt was sent to all three engines, ChatGPT, Claude, and Gemini, using a single model version per engine (1 version tracked for ChatGPT, 1 for Claude, 1 for Gemini), so differences in results reflect differences between engines rather than between model generations within an engine.
Runs were collected across a measurement window bounded by two timestamps in 2026, and 3 mitigations were applied during collection to reduce known sources of noise, such as retrying failed calls and normalizing company names before matching. The raw log contains 170,457 rows; after filtering to rows with a clean, parseable answer, 170,457 remained, all logged with an 'ok' status (170,457 of them), meaning none were dropped for parsing failure in the final count.
For every usable answer, the study recorded which companies were named and, when named, the ordinal position of the mention within the answer (first company mentioned, second, and so on). This produced two separate measurements that are easy to conflate but answer different questions. The first is whether a company appears at all on a given engine for a given prompt, which is the basis for the agreement analysis. The second is where in the answer it appears when it does, which is the basis for position analysis such as 5.18 as the overall average position across 65,241 scored mentions.
The unit of comparison for agreement is the company-prompt pair: for each prompt, which companies did each engine name, and how many of the three engines named each one. This is deliberately stricter than asking 'did the engines cover similar topics,' because it requires exact company identity to match, not just category or theme. A pair counts as agreed only if the same company name, after normalization, was produced by more than one engine for the same prompt.
The headline result, and its limits
Across 56,819 company-prompt pairs where at least one engine named a company, the breakdown by number of agreeing engines is:
| Engines naming the pair | Share of pairs |
|---|---|
| Exactly one | 88.4% (50,209 pairs) |
| Exactly two | 8.4% (4,798 pairs) |
| All three | 3.2% (1,812 pairs) |
The plain reading is that sole mentions dominate. 88.4% of pairs were named by exactly one engine, and full three-way agreement happened on only 3.2% of pairs.
What this does not mean is that the engines are randomly guessing or that most companies are invisible. Separately, 0% of the 30,564-company cohort never appeared on any engine at all, meaning the disagreement is concentrated among companies that do get named somewhere, just not everywhere. And 96% of companies were visible on at least one engine without being visible on all three, which is the group most exposed to the choice-of-engine problem described in the first section.
It also does not mean the engines are working from different facts about the world. A more likely explanation, consistent with the pattern, is that each engine draws from a differently weighted slice of its training and retrieval sources and applies its own shortlist logic, so two engines can both have accurate information about a company and still not both choose to name it in a given answer, especially in categories with many plausible candidates. Distinguishing 'the engine doesn't know about this company' from 'the engine knows but didn't pick it' would require probing each engine directly about specific companies, which this study did not do; it observed only what engines volunteered in response to prompts a buyer would plausibly write.
A second limit: the study measures naming, not ranking quality or accuracy. A company could be named correctly by one engine and correctly omitted by another because the other engine judged it a weaker fit for that specific prompt. Low agreement is consistent with both 'engines see different parts of the market' and 'engines make different but equally defensible editorial choices.' This study's data cannot fully separate those two stories, though the category-level consistency described later leans toward the former.
Is this a gradient or a cliff
The shape of a result changes what follows from it. If visibility were binary, either a company is known to AI engines or it is not, the practical response would be simple: get past the threshold. If it is a gradient, the response has to be about degree and about which engine, not about crossing a single line.
The data support a gradient. The three-way split in the previous section, roughly 88.4% sole mentions, 8.4% two-engine mentions, and 3.2% three-engine mentions, is not a spike at zero and a spike at three with nothing in between. There is a real, populated middle tier of pairs named by exactly two engines. And per-engine visibility rates form their own gradient rather than clustering at one value: 42.7% on ChatGPT, 39.9% on Claude, 33.6% on Gemini.
The near-zero rate of complete invisibility reinforces this. 0% of the cohort's 30,564 companies never appeared on any engine, which means the practical question for almost every company in this dataset is not 'am I visible at all' but 'am I visible on the engine this particular buyer happens to be using.' That is a meaningfully different problem to solve. A binary visibility problem has a single fix: produce more of the content that gets a company noticed at all. A gradient, engine-specific problem means a company can be doing everything right for one engine's retrieval and ranking behavior and still be absent from another's answer to the identical question.
This also means a company cannot infer its standing on Gemini from its standing on ChatGPT. The position data adds a secondary nuance: once named, companies land in a similar spot regardless of engine, averaging 5.25 on ChatGPT, 5.38 on Claude, and 4.85 on Gemini, against an overall average of 5.18. So the gradient is almost entirely about the yes/no decision to name a company, not about how prominently it gets named once an engine decides to include it.
What the engines appear to be working from
This study did not log source citations for this dataset (no per-source citation data was captured), so it cannot say directly which pages, directories, or reviews each engine drew on when deciding whom to name. What it can say is inferred from the pattern of answers themselves.
Coverage was effectively complete at the category level: across every category examined, from small ones like Healthcare IT Services (9 companies) and Fintech Testing Services (24 companies) to large ones like SaaS (994 companies) and Food Service Business (819 companies), the visible share of companies sits at 100% of companies with at least one visible-somewhere mention in every category checked. That uniform, near-total category coverage suggests the engines are not simply blind to entire industries or company sizes; the disagreement documented above is happening inside categories, between specific competing companies, not between categories.
Prompt intent also shapes how often any company gets named, though less than might be expected. Appearance rates by prompt type are close to each other: 38.6% for commercial-investigation prompts (of 21,885 answers), 38.7% for informational prompts, 36.9% for navigational prompts, and 38.3% for transactional prompts, by far the largest bucket with 124,338 answers and 47,628 appearances. The similarity across intents suggests that whether a company gets named is driven less by what kind of question was asked and more by which specific companies each engine's underlying sources happen to surface for that topic.
The honest limit here is that 'what the engine is reading' is an inference from output patterns, not a direct observation of retrieval logs or training data. A company that wants to know why it was named by one engine and not another cannot get that answer from this dataset; it can only see that the split exists and how large it is.
Where the engines disagree, and why a single score misleads
The practical consequence of the agreement numbers is aimed squarely at anyone buying or building a single 'AI visibility score.' If 88.4% of company-prompt pairs are named by only one engine, then a monitoring tool that checks only ChatGPT, for instance, is not a slightly-incomplete version of the full picture. It is closer to a one-third sample dressed up as the whole.
The per-engine visibility rates make the size of that gap concrete: 42.7% of companies appear on ChatGPT, 39.9% on Claude, and 33.6% on Gemini. A company sitting at, say, strong ChatGPT visibility and weak Gemini visibility would score very differently depending on which single engine a monitoring product happened to track, even though nothing about the company changed between the two measurements.
This is not simply engines disagreeing at random. The five companies that showed the strongest presence across the dataset, The Home Edit, Junk King, BMM Testlabs, Ecobee, and Visit El Paso, a set of 5 companies spanning organizing services, junk removal, gaming testing, smart home devices, and tourism, span very different categories, which suggests that broad visibility is achievable in different market types, not confined to one kind of business or one kind of engine's apparent preference.
What this means for a buyer of an AI visibility product: ask which engine, or engines, the score is built from, and treat a single-engine score as a measurement of that engine's answer behavior, not of 'AI' as a whole. A score built from all three engines, weighted or reported separately, will track the 3.2% of pairs where engines agree and the much larger share where they do not, rather than collapsing both into one number that hides the disagreement. The alternative explanation worth naming honestly: some of this disagreement may reflect legitimate differences in what counts as a good answer for a given prompt, not just gaps in each engine's knowledge. Nothing in this dataset separates 'engine doesn't know about this company' from 'engine knows but chose a different answer,' and a company investigating its own case would need engine-specific probing to tell the two apart.
What a company on the wrong side of this should do
A company that finds itself named by one engine and not the others has a narrower, more tractable problem than 'improve AI visibility' suggests. The data point to a few concrete responses.
First, check per-engine visibility separately rather than accepting a single blended score. Given that 42.7%, 39.9%, and 33.6% are visibility rates that move independently, a company should know its standing on each engine, not just an average across them. An average of a high and a low score looks identical to a flat medium score, and the fix for those two situations is different.
Second, treat the 0% near-zero rate of total invisibility as evidence that the floor is not the problem for almost any company in a covered category. The realistic goal for most companies is not 'get discovered by AI at all,' which is already happening; it is 'close the gap between the engine where I appear and the engines where I don't.' That reframes the work from broad content production toward figuring out what is different about the sources or signals feeding the engine where the company is absent.
Third, do not assume position work matters more than naming work. Because average mention position is close across engines (5.18 overall), a company already being named is not obviously being buried by one engine relative to another. The bigger lever is the binary naming decision itself, not incremental position improvement once named.
Fourth, watch category context. Since visible coverage is at or near total in every category examined, a company that is absent from an engine's answers is not absent because the category itself is under-covered; it is competing against specific named alternatives within a well-covered category. That argues for competitive analysis, comparing against the companies that do get named for the same prompts, over generic visibility campaigns.
Finally, recognize the limits of self-diagnosis. Because this study cannot see each engine's underlying sources, a company cannot fully reverse-engineer why one engine omits it. What it can do is monitor consistently across all three engines over time and treat any engine-specific gap as a hypothesis to test with targeted content and outreach, then remeasure.
How to read the published dataset yourself
The underlying dataset is built from 170,457 logged rows, narrowed to 170,457 usable answers, all with an 'ok' status, collected across a window bounded by two UTC timestamps in 2026. Each row ties one prompt, one engine, and one answer together, with company mentions and their positions extracted from the answer text.
Three tables matter most for independent checking. The first is the pair agreement table, which groups company-prompt pairs by how many of the 3 engines named them; use the pairs count columns, not just the percentages, if you want to know how thin or thick a slice you are looking at, since 1,812 pairs is a much smaller base than 50,209 pairs even though both get summarized as percentages.
The second is the per-category breakdown, useful for anyone who only cares about their own vertical. Categories range enormously in size, from Healthcare IT Services at 9 companies to Digital Marketing at 827 companies, and a category-level visible rate calculated on a nine-company base carries far less weight than the same rate on an eight-hundred-company base, even when both render as the same percentage.
The third is the per-engine visibility and position tables, which should always be read side by side rather than blended, given the disagreement documented in this study. If you build your own summary statistic from this data, report it per engine first, and only then consider whether an aggregate is meaningful for your purpose.
One caution for reuse: this dataset used a single model version per engine (1 for ChatGPT, 1 for Claude, 1 for Gemini), captured across a fixed window. Engine behavior changes as providers update models, so a rerun months later on updated model versions should be expected to shift these numbers, possibly substantially, and should not be treated as invalidating this snapshot or as directly comparable without checking model versions on both sides.
Findings in depth
Each of these has its own page, written to stand on its own.
When you ask ChatGPT, Claude, and Gemini the same question, how often do they name the same company
How often do ChatGPT, Claude, and Gemini name the same company for the same prompt
Full three-way agreement is rare; most company mentions come from just one engine.
Read this finding →Are there companies that never get named by any of the three engines
Is any company invisible to AI engines entirely
Almost none: across the full cohort, the rate of appearing nowhere on any engine is effectively zero.
Read this finding →Do ChatGPT, Claude, and Gemini show companies at similar rates, or does one engine surface far more than the others
Which AI engine names the most companies, ChatGPT, Claude, or Gemini
ChatGPT named companies most often in this dataset, Gemini least often, with a real gap between them.
Read this finding →Do commercial, informational, navigational, and transactional prompts produce different company appearance rates
Does the type of question asked change how often companies get named
Appearance rates are close across all four prompt intents, so what drives naming is mostly which companies each engine knows or picks, not what kind of question was asked.
Read this finding →How this sits against other published work
Most published work on AI answer engines asks whether a citation is accurate. We asked a narrower and, we think, more basic question: whether engines even agree with each other about which companies to name at all, independent of whether any single answer is right or wrong. That framing puts this study in a different lane from the two literatures it otherwise resembles.
The academic anchor for work on how companies get surfaced in generative answers is the GEO paper (Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande), published on arXiv and later at ACM SIGKDD. GEO treats visibility as something that can be optimized: the paper reports that its interventions can lift a source's visibility in generative engine responses by up to 40%. That result presumes a stable target to optimize toward, one engine's ranking of sources for a given query. Our data complicate that premise for anyone trying to optimize across engines rather than within one. We logged 56,819 company-prompt pairs where at least one of ChatGPT, Claude, or Gemini named a company, and 88.4% of those pairs were named by exactly one engine. Only 3.2% were named by all three. If engines drew on a shared, stable notion of "the best answer," we would expect far higher three-way overlap than that. GEO's benchmark, GEO-bench, does not test cross-engine agreement, so the two findings are not in tension, but they are in a productive contrast: GEO shows visibility is optimizable per engine, and our data show that optimizing for one engine does not transfer to the others.
The Tow Center's study, covered by the Columbia Journalism Review and summarized by Nieman Lab, is the closest existing empirical parallel to ours in method: it ran repeated prompts against multiple engines (eight, including ChatGPT Search, Perplexity, Gemini, and Copilot) and scored the results, finding that across 1,600 test queries the engines failed to retrieve correct information more than 60% of the time, with Perplexity performing best among the eight. That study measured whether a cited source was the right one for a news claim. We did not score correctness at all; we measured whether three engines converged on the same company names for a commercial prompt. The two are not directly comparable, and a reader should not average them, but they point the same direction: engines are inconsistent, whether the object is a news citation or a company name. The Tow Center's low agreement on correctness and our low agreement on company selection (88.4% sole-mention rate) are consistent with a shared underlying cause, each engine's retrieval and ranking layer is a separate system built by a separate company, but we cannot rule out an alternative: our prompts may simply have more legitimate right answers than a factual news query does, so disagreement here is less a failure than a difference in retrieval breadth. Distinguishing those two explanations would require running the same prompts through a task with a single verifiable correct answer, which is what the Tow Center did and we did not.
Industry sources such as the 5WPR Citation Source Index and boringmarketing.com's AI Visibility Statistics report per-engine citation rates from unspecified methodologies; we cite them as context, not validation, since neither discloses a comparable sampling frame. G2 and Gartner's buyer surveys establish why this matters commercially: buyers act on what these engines say, so disagreement between engines is not a curiosity but a source selection problem for every vendor watching their own visibility.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- GEO: Generative Engine Optimization arXiv, 2023 Introduces GEO-bench and reports that generative engine optimization techniques can raise a source's visibility in generative engine answers by up to 40%, establishing that citation behavior in these systems is measurable and manipulable.
- AI Platform Citation Source Index 2026 5WPR via PR Newswire, 2026 Industry citation-source ranking across five engines, cited for its figure on Reddit's dominance as a cited source and its scale of citations analyzed.
- GEO: Generative Engine Optimization Princeton University, Collaborate repository, 2023 Institutional archival record of the GEO paper, used to confirm authorship and publication context.
- AI Visibility Statistics (2026) boringmarketing.com, 2026 Independent industry dataset reporting brand citation rates by engine, cited for comparison against our per-engine visibility rates.
- New G2 Research: Half of B2B Software Buyers Now Start Their Research With AI Chatbots G2 via PR Newswire, 2026 Survey evidence that buyers act on AI chatbot recommendations, establishing why engine disagreement in company naming has commercial consequence.
- GEO: Generative Engine Optimization ACM SIGKDD (Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining), 2024 Peer-reviewed proceedings version of the GEO paper, cited as the field's foundational benchmark for measuring visibility in generative engine responses.
- Tow Center's Latest Report on AI Search Engines Columbia Journalism School, 2025 Institutional news page confirming the scope and release of the Tow Center report.
- Generative engine optimization Wikipedia, 2026 Background and definitional context for the term GEO as used in this related-work discussion.
- AI search engines fail to produce accurate citations in over 60% of tests, according to new Tow Center study Nieman Lab, 2025 Summary of the Tow Center study reporting per-engine failure rates, including Perplexity's lowest failure rate among the eight engines tested.
- Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights Gartner, 2026 Survey data on how buyers source and validate vendor information, including average number of information sources used, cited as context for why cross-engine agreement matters to buyers.
- GEO: Generative Engine Optimization OpenReview, 2023 Review version of the GEO paper, cited for its framing that generative engines need a purpose-built visibility metric distinct from search-engine ranking metrics.
- AI Search Has a Citation Problem Columbia Journalism Review / Tow Center for Digital Journalism, 2025 Primary independent study testing eight AI search engines on citation accuracy for news content, the closest existing benchmark to our cross-engine comparison.
Terms used in this study
- Company-prompt pair
- One company and one prompt considered together. If a company is named by at least one engine in response to a given prompt, that combination counts as one pair. Agreement is measured by how many of the three engines named that same pair.
- Appearance rate
- The share of answers, within a given group such as a prompt type, in which at least one tracked company was named by the engine.
- Visibility rate
- The share of companies in the cohort that were named by a specific engine at least once across the study's prompts.
- Mention position
- The ordinal place a company holds within an engine's answer when it lists multiple companies, for example first, second, or third mentioned. A lower average position means companies tend to be named earlier in the answer.
- Sole mention
- A company-prompt pair named by exactly one of the three engines and not named by the other two for that same prompt.
- Cohort
- The full set of distinct companies that appeared at least once, on any engine, anywhere in the study's usable answers.
- Usable answer
- A logged engine response that passed quality checks, meaning it could be parsed cleanly enough to extract company names and positions, and was marked with an 'ok' status.
- Prompt intent type
- One of four categories, commercial-investigation, informational, navigational, or transactional, describing the kind of question a prompt represents, based on how a real buyer might phrase it.
- Category
- A business vertical or service type used to group companies and prompts, such as Blockchain Testing Services or Landscaping Services.
- Visible somewhere, not everywhere
- A company that was named by at least one engine but not by all three, the group most affected by which single engine a buyer happens to consult.
Questions about this study
What did this study actually measure?
So do the engines agree with each other or not?
Does that mean most companies are invisible to two out of three engines?
Why would three engines answering the same prompt give different lists?
Is this just because one engine writes shorter answers than the others?
When a company does get named by more than one engine, does it show up in the same spot?
Data and method
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number on this page can be recomputed from it.
Limitations we volunteer
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
How to cite this study
Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study. BusySeed, 2026-05-23. https://busyseed.com/research/cross-engine-disagreement-buyer-questions
About BusySeed
BusySeed is a data-driven growth marketing agency that measures and improves how brands appear in AI-generated answers.
More BusySeed research
The 50 fastest-growing SaaS companies barely exist in AI answers
According to BusySeed, 5 of 50 companies studied (10%) never appear when buyers ask AI assistants about their own category.
The State of AI Search, July 2026
According to BusySeed, in July 2026, 88.1% of the company recommendations AI assistants made for a buyer question came from only one of three engines, and just 3.3% were made by all three.
Does the Company AI Recommends First Stay the Same? A 15-Week Volatility Study
According to BusySeed, the company an AI assistant names first for a category changes from one week to the next in 56.4% of week-to-week transitions, and 82.3% of the question and engine pairs tracked saw their top answer change at least once.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
