Skip to content

Getting Cited by Perplexity: What Actually Gets You There

Ask Perplexity a commercial question — the best plumber in a mid-sized city, the most reliable project management tool for a five-person team — and the top citation is often not a business’s own website. It’s a Reddit thread, sometimes years old, written by a stranger with no stake in the answer. That’s not an occasional glitch. Profound’s analysis of more than 4 billion AI citations found Reddit is Perplexity’s single most-cited domain — ranked ahead of every company site, every review platform, everything.

Why does Perplexity favor Reddit over a company’s own site?

Perplexity doesn’t rank a list of links the way a search engine does. It retrieves a set of pages, then writes an answer and attaches citations to specific claims inside it. That process rewards a source that says something a model can lift cleanly: a direct opinion, a specific comparison, a recent first-hand account. A company’s homepage usually says the opposite of that — broad claims about quality, written to be true for any customer rather than specific to the one question being asked. A Reddit reply to “which of these two did you actually end up using” answers exactly the question in front of it, in one paragraph, from someone with no reason to oversell it. That fits Perplexity’s retrieval far better than most marketing copy does, which is a large part of why Reddit turns up first so often.

What does PerplexityBot actually do?

Perplexity’s own developer documentation is specific about the job: PerplexityBot crawls public pages to build the index Perplexity’s search and answer features draw on — and it states plainly that this crawler is not used to gather training data for Perplexity’s underlying models. That’s one function, done ahead of time, on a schedule nobody outside Perplexity controls.

Is PerplexityBot the only Perplexity crawler?

No, and the second one works differently in a way that matters. Perplexity names a separate agent, Perplexity-User, that fetches a specific page live, at the exact moment a person’s question sends Perplexity there — not a scheduled pass, a one-off request tied to one conversation. Perplexity’s documentation draws the compliance line between them explicitly: PerplexityBot follows robots.txt, while Perplexity-User “generally ignores robots.txt rules” because a person, not a scheduled job, asked for that page. The same split exists for Anthropic’s crawlers around Claude — a training or indexing bot that respects the file, and a live, on-demand fetcher that doesn’t answer to it the same way.

Does blocking PerplexityBot stop a page from being cited?

Only partly, and it’s worth being precise about which part. Disallowing PerplexityBot in robots.txt keeps a page out of the index that most everyday answers draw from — which is where the large majority of citations come from, since most questions aren’t naming one exact page. It does not stop Perplexity-User from fetching and citing that same page on the spot, if a specific question happens to point Perplexity directly at it, because that fetcher isn’t reading robots.txt the way the scheduled crawler does. Blocking the bot lowers the odds of ordinary, ambient citation. It doesn’t close the door completely, and a site owner who disallows PerplexityBot expecting total exclusion is working from an incomplete model of how the two agents behave.

How do you check whether PerplexityBot can actually reach a site?

Reading robots.txt only shows what a site intends to permit. It says nothing about what the server does when a request actually arrives. The audit engine behind this site — the full breakdown of what it checks and how lives in the audit methodology — checks both, separately: it reads robots.txt for what’s declared, and for a named set of crawlers — PerplexityBot among them — it also sends a live request identifying itself with that crawler’s real user agent, to see what the server does when that specific name shows up. The two answers don’t always match. A CDN’s bot-management layer can quietly block a crawler that robots.txt allows, and nobody reading the text file would ever know. A robots.txt checker shows what the file says for free; catching a server that’s quietly out of step with it takes an actual probe like the one above.

Does a citation actually take, beyond having the right robots.txt line?

Getting the crawler access right removes a barrier. It doesn’t manufacture the kind of content Perplexity prefers to cite. The evidence above points to a specific, unglamorous answer: pages that state a narrow claim plainly, backed by something concrete, read closer to what wins a citation than a general homepage does — and genuine discussion elsewhere, in places a model actually retrieves from, does more work than any on-page change can. None of that is something a robots.txt edit produces on its own, and none of it is something this or any other audit tool can manufacture for a site — it can only measure whether the crawler access and the page structure are in the way.

That last part is a real limit, not a caveat added for balance. A free scan here doesn’t run AI model probes at all; a deep scan does, but even then it questions general-purpose chat models about what they already know about a brand — not Perplexity, because Perplexity doesn’t hold fixed, queryable knowledge the way those models do. Its answer to “have you heard of this business” depends on what its index has retrieved in that moment, which can be different an hour later. Measuring that in the same way would mean re-running the question indefinitely, against a system built to change.

Frequently asked questions

Does PerplexityBot train Perplexity's AI models?

No. Perplexity's own developer documentation states PerplexityBot is not used to crawl content for training foundation models. It builds the index Perplexity's search and answer features retrieve from, which is a separate job from training a model.

What's the difference between PerplexityBot and Perplexity-User?

PerplexityBot is a scheduled crawler that builds Perplexity's search index ahead of time. Perplexity-User fetches a specific page on demand, at the moment a person's question sends Perplexity there. Perplexity's documentation says PerplexityBot follows robots.txt, while Perplexity-User generally does not, because a person's live request triggered the fetch.

Does blocking PerplexityBot in robots.txt stop Perplexity from citing a page?

Not entirely. It removes the page from Perplexity's pre-built index, so it won't surface in the ordinary run of answers. It doesn't stop Perplexity-User from fetching that exact page live if a specific question sends it there, because that fetcher generally ignores robots.txt.

Why does Perplexity cite Reddit so often?

Profound's analysis of over 4 billion AI citations found Reddit is Perplexity's single most-cited domain, ranked ahead of every other source. Perplexity's retrieval favors recent, specific, first-hand discussion, and a Reddit thread of real replies fits that shape more often than a company's own marketing page does.

Can a free SEO scan tell you whether Perplexity's model already knows about a business?

Not directly. A free scan doesn't run AI model probes at all — that's a deep-scan feature, and even then the models it questions are general-purpose chat models, not Perplexity itself, because Perplexity answers from live retrieval rather than fixed training-time knowledge that can be queried the same way.