Perplexity is the one AI tool that shows its work. Every answer carries numbered citations that link straight back to the pages it used, which means you can check, in about a minute, whether your site is in the pool of sources it draws from. No other AI system makes the diagnosis this easy, and that changes what “optimizing” for it looks like: less guesswork, more reading the receipts.

The catch is that getting into that pool depends on things most guides skip. Perplexity runs its own crawler, publishes its own IP ranges, and has been at the centre of a public fight about crawler behaviour. Meanwhile, the infrastructure sitting in front of your website may be quietly turning it away. This piece covers all of it, in the order that matters.

How does Perplexity decide which pages to cite?

Perplexity searches the live web for each question rather than answering from training data alone, pulls a handful of pages, extracts passages from them, and builds an answer with numbered footnotes pointing at those exact sources. To be cited, your page has to be reachable by its crawler, relevant to the query, and written so a specific passage can be lifted without distortion.

That architecture has two practical consequences. First, freshness carries real weight, because the system reads the web at query time rather than recalling what it learned months ago. A page with last year’s prices competes badly against one updated last week. Second, access is binary: if Perplexity’s crawler cannot fetch your pages, nothing else on this list matters, no matter how good the writing is.

The scale is worth a number. Perplexity CEO Aravind Srinivas said onstage at Bloomberg’s Tech Summit that the platform handled 780 million queries in May 2025 and was growing more than 20% month over month, as reported by TechCrunch. Treat that as a snapshot from a fast-moving company rather than a current figure.

Is your CDN blocking Perplexity without telling you?

Check this before you touch your content. If your site sits behind Cloudflare, AI crawler access may be governed by a dashboard setting you never chose. Cloudflare announced on 1 July 2025 that every new domain signing up would be asked whether to allow AI crawlers, so the answer given during setup — possibly by your developer, years-ago-you, or a default — now decides whether Perplexity can read you.

This catches out a lot of small US and UK business sites, and it is invisible from the outside. Your robots.txt can say exactly the right thing while a network-level rule blocks the request before it reaches your server. In Cloudflare’s own words from that press release, the aim was that “every new domain starts with the default of control, and eliminates the need for webpage owners to manually configure their settings to opt out.”

The rules are still moving. In a blog post published 1 July 2026, Cloudflare said that from 15 September 2026, for all new domains onboarding to Cloudflare, bots classified as Training and Agent will be blocked by default on pages that display ads, while Search remains allowed. That distinction matters more than it looks: Perplexity’s indexing crawler falls on the Search side, but the fetch that happens when a live user asks Perplexity about your business behaves like an Agent. Block the second category and you can sit in the index while being unreadable at the moment a prospect asks.

What to actually do: open your CDN’s bot or AI crawler controls and confirm what is allowed, then check your server logs for the user agents below. A setting that says “allowed” and logs that show nothing are two different states.

Which Perplexity crawlers should you allow?

Perplexity documents two user agents that behave differently. PerplexityBot is the search crawler, it respects robots.txt, and allowing it is what gets your pages into the pool Perplexity searches. Perplexity-User is the fetcher that visits a page because a person asked a question. Perplexity’s documentation states it “generally ignores robots.txt rules.”

User agent What it does Respects robots.txt Published IP list
PerplexityBot Surfaces and links websites in Perplexity search results; per Perplexity’s docs, “not used to crawl content for AI foundation models” Yes perplexity.com/perplexitybot.json
Perplexity-User Visits a page in real time when a user’s question requires it No — “generally ignores robots.txt rules” perplexity.com/perplexity-user.json

Both descriptions and both IP endpoints come from Perplexity’s crawler documentation. Those published IP ranges are the useful part for verification: if something in your logs claims to be PerplexityBot but its IP is not on that list, it is not Perplexity. The minimum viable robots.txt entry is three lines:

User-agent: PerplexityBot
Allow: /

If you have ever pasted in a blanket “block all AI bots” rule — a popular move in 2024 and 2025 — this is where it bites. You can hold that position on principle, but you cannot hold it and expect citations. The broader crawler list for other tools is in how to optimize your website for AI search.

Five steps to get cited by Perplexity

None of these require a rebuild. They are ordered by leverage: the first one can flip you from invisible to eligible in an afternoon, the last one compounds slowly.

  1. Clear the access path end to end. CDN setting, then robots.txt, then server logs. Confirm PerplexityBot requests are arriving and returning 200s, not 403s. One check, not three projects — and the only step where a single wrong line makes everything else pointless.
  2. Keep your buying-decision pages current. Prices, service areas, availability, contact details. Live retrieval rewards pages that are right today, and a stale price is worse than no price: it makes you a source that gets burned by being cited.
  3. Write passages that survive extraction. A sentence carrying one concrete fact — a price, a timeline, a specific figure — that still makes sense pulled out of its paragraph is exactly what fits under a numbered footnote. Vague positioning language has nothing to extract.
  4. Make your business facts identical everywhere. Name, address, phone, hours and core services matching across your site, your Google Business Profile and the directories. Conflicting facts give a system that cites publicly a reason to pick a source it can stand behind. More on this in AI visibility for small and local businesses.
  5. Read the citations on your own category’s questions. Ask Perplexity what your customers ask, look at the numbered sources, then open the pages that got quoted and note what type of content earned it: a definition, a price table, a comparison. That tells you what to publish next with far more precision than any generic checklist.

Notice what is absent: there is no way to buy a citation slot. The same is true of ChatGPT recommendations, which is one reason the work overlaps so heavily — see why ChatGPT doesn’t recommend your business.

What about the crawler disputes and the publisher payouts?

Both are real, both are worth knowing, and neither changes what you should do. In August 2025 Cloudflare accused Perplexity of stealth crawling; Perplexity disputed the framing. Separately, Perplexity launched a revenue-share programme for publishers. If you have seen the headlines and wondered whether Perplexity is a partner or an adversary, the honest answer is that the industry has not settled it.

The dispute, in the specifics. In a post dated 4 August 2025, Cloudflare said it “de-listed them as a verified bot and added heuristics to our managed rules that block this stealth crawling,” alleging Perplexity used “a generic browser intended to impersonate Google Chrome on macOS when their declared crawler was blocked,” rotating IPs and ASNs “across tens of thousands of domains and millions of requests per day.” Perplexity’s response argued that user-driven fetching is categorically different from automated crawling, said Cloudflare had misattributed traffic from BrowserBase, a third-party cloud browser service, and wrote that Cloudflare’s leadership “is either dangerously misinformed on the basics of AI, or simply more flair than cloud.” The actionable residue: verify bot identity against the published IP lists rather than trusting a user-agent string, and read your logs rather than assuming.

The payouts, in the specifics. Perplexity announced Comet Plus in August 2025, a $5-a-month tier feeding a publisher revenue pool. Jessica Chan, Perplexity’s head of publisher partnerships, told Digiday that subscription revenue is pooled, Perplexity keeps 20% and 80% goes to participants, from an initial pool of $42.5 million, split across human visits, search citations and agent actions. Publishers apply by emailing publishers@perplexity.ai. Be realistic: this is built for publishers with real content volume, and a five-page service site is not the target. Context for the ecosystem, not a revenue line for you.

One more number is worth sitting with. Cloudflare’s analysis published 29 August 2025 tracked crawls per referred visitor and found Perplexity climbing from 54 to 195 between January and July 2025, against Google moving from 3.8 to 5.4 over the same period. Perplexity was far more referral-efficient than OpenAI (1,091) or Anthropic (38,065) in that data, but the direction is clear: AI citations are not a like-for-like replacement for search clicks. Worth pursuing, yes. Expecting them to replace organic traffic one-for-one is not supported by any data anyone has published.

How do you check whether Perplexity cites you today?

Ask Perplexity the questions your customers ask — once with your business name, once without — in a fresh or signed-out session, and read the numbered sources under each answer. Log whether your domain appears, and whose does. Because the citations are explicit, this is observation rather than inference, and it takes an evening.

Do it more than once. AI answers vary between runs, so a single check tells you almost nothing and a pattern across several tells you a lot. The method for turning that into a number you can compare over time is in how to measure AI visibility, and the full prompt set for testing across four tools at once is in how to check if AI recommends your business.

When is this worth handing to someone else?

Checking Perplexity yourself tells you about one tool. Your customers use four. If you want the whole picture scored rather than sampled, that is what our AI visibility audit does: 39 criteria across six areas, the same customer questions run through ChatGPT, Gemini, Perplexity and Claude, and a prioritized fix list with an expert reading of why each gap exists. It starts at PLN 499 net (roughly $125), and if we don’t find at least five things worth fixing, you get a full refund.

If you’re not sure it applies to you yet, a free 20-minute consultation will tell you honestly — including when the answer is “not yet, fix your website first.” Or start broader with the free 3-minute marketing audit, which scores your whole marketing before you narrow in on AI.

FAQ

Does Perplexity use Google’s or Bing’s index? Perplexity runs its own crawler, PerplexityBot, with its own published IP ranges, and its documentation describes it as the bot that surfaces and links websites in Perplexity search results. That independence cuts both ways: a site that is well-indexed by Google can still be missing from Perplexity if PerplexityBot is blocked at the CDN or in robots.txt, and fixing one does not fix the other.

Do Perplexity citations actually send traffic? More than AI systems that hide their sources, because the citations are visible and clickable, but far less than search rankings. Cloudflare’s August 2025 data put Perplexity at 195 crawls per referred visitor in July 2025 versus 5.4 for Google. Citations are worth having, and they are a visibility and credibility play rather than a traffic channel that replaces organic search.

If I already optimized for ChatGPT, do I need to do this separately? Mostly no. Crawler access, factual consistency, concrete on-page detail and third-party trust signals work on both systems, so the foundation is shared. The Perplexity-specific additions are narrow: allow PerplexityBot specifically, and keep your key pages genuinely current, since live retrieval weights freshness more heavily than systems leaning on training data.

Should I block AI crawlers instead, on principle? That is a legitimate choice, and Cloudflare has built the tooling to make it easy — but it is a choice, not a default you should drift into. Blocking means opting out of the answers your prospects are reading. Decide it deliberately, and if you’re blocking to protect content from model training rather than from search, check whether your rules distinguish the two, because many blanket rules don’t.