← / BlogResearch
researchai-citationscitation-indexoriginal-research

We asked AI 450 times which project management tool to buy. Notion never came up.

An original study of 450 AI answers across ChatGPT, Claude and Perplexity. A review blog got cited more than any product, Notion and Linear were never mentioned once, and the three engines agreed on almost nothing.

5 min read

TL;DR

  • We put 50 buyer questions to three AI engines, three times each — 450 answers, 2026-09-25. Every domain each answer cited was recorded.
  • A review publication was cited more than any product. thedigitalprojectmanager.com appeared in 36% of answers; the highest-scoring actual tool, Wrike, managed 21%.
  • Notion and Linear were never cited once. Trello and Basecamp appeared in one answer each, out of 450.
  • The engines barely agree. Of 606 domains cited, 480 appeared on only one engine and just 43 appeared on all three.
  • Being a big brand did not predict being recommended. Whatever these engines are rewarding, it is not market share.

Anyone selling project management software has a rough idea of who the market leaders are. We wanted to know something different: when a buyer asks an AI assistant which tool to use, whose website actually gets cited as the source? So we asked, systematically, and counted.

How we ran it

Fifty questions, written to sound like a real buyer rather than a keyword — "which project management tool should I use if my team works across five time zones", "what do people complain about with these tools", "how much does this cost for a 10-person agency". No brand names in the questions: the point was to see who the engines volunteer unprompted.

Each question went to ChatGPT, Claude and Perplexity, three times each, with web search enabled. Three runs because these systems are not deterministic — ask twice, get two different answers. 450 calls completed successfully out of 453 attempted.

Google Gemini is excluded. Its API returns citations as redirect wrappers rather than publisher domains, so there is nothing to count. That is a limitation of our instrument, not a statement about Gemini.

A review blog beat every product

domainshare of answerswhat it is
1thedigitalprojectmanager.com36.0%review publication
2ones.com28.9%vendor
3wrike.com21.1%vendor
4celoxis.com14.9%vendor
5teamwork.com14.2%vendor
6monday.com13.3%vendor

The top result is not a product. It is a publication that writes about products, and it was cited in more than a third of every answer we collected — top-three on all three engines independently.

Widen it and the pattern holds: 58% of all answers cited at least one review publication — project-management.com, cloudwards.net, techrepublic.com and similar. When an AI assistant answers "which tool should I buy", it is largely relaying what reviewers have written, not what vendors have published about themselves.

Brand size did not predict citation

This is the part that should worry a marketing team:

  • Notion: zero citations in 450 answers. Not one.
  • Linear: zero.
  • Trello: one. Basecamp: one.
  • Asana: 7.8%, ranked 18th.
  • monday.com: 13.3%, ranked 6th — the best-performing name most buyers would recognise.

Meanwhile ones.com and celoxis.com — tools with a fraction of the brand recognition — finished 2nd and 4th.

We cannot tell you why from this data, and we are not going to guess. Correlation with any particular on-site or off-site factor is not something 450 answers in one category can establish. What the data does show is that the ranking AI engines produce is not the ranking the market would produce, and that the gap is large.

The engines are not interchangeable

Of 606 distinct domains cited across the study, only 43 were cited by all three engines. 480 — nearly four in five — appeared on exactly one.

Some individual splits are stark:

  • Reddit: 25% of ChatGPT answers, 4% of Perplexity, 0% of Claude.
  • Asana: 22% on Perplexity, 1% on Claude, 0% on ChatGPT.
  • goodday.work: 23% on Claude, 0% on Perplexity.

So "how visible are we in AI search" is not a single question. A brand can be well represented on one engine and absent from another, and any advice that treats AI search as one channel — including plenty of advice about Reddit — is generalising from one engine's behaviour.

What we are not claiming

Three honest limits, because a study without them is marketing.

One category, one day. This is project management software on 2026-09-25. It is a data point, not a law. We would expect different dynamics in categories with different publishing ecosystems.

Answers move between runs. Across 447 repeated pairs, two runs of the same question on the same engine shared only 57% of their cited domains. The rankings above are averages over three runs, which is why we ran three — but roughly four in ten citations change if you ask again.

No causal claim. We measured who gets cited. We did not measure why, and nothing here shows that any particular action causes citation. Anyone telling you otherwise from data like this is selling something.

FAQ

Does being cited by an AI actually bring traffic?

Often not directly — AI answers are largely zero-click. The citation matters as evidence that the engine treats your site as a credible source for that question, which is a different and usually earlier thing than a visit.

Why three runs per question?

Because one run is noise. The same question asked twice returns meaningfully different sources — we measured 57% overlap between repeats. Any study reporting single-shot results is reporting variance as if it were a finding.

Can I see the raw data?

Yes — all of it is public, CC0. The 450-row dataset, the exact 50 questions, and the two discarded OpenAI collection passes are all there, along with the methodology notes. Check it, and disagree with us if the numbers say something different to you.

Sources


Want to know whether AI cites you in your category? Run a free scan at CiteFlow — we will show you which sites got cited instead.