OverreachPI-0021
Answer engine accused of using disguised crawlers to get around sites' robots.txt and firewall blocks
Cloudflare reported that after sites blocked Perplexity's declared crawlers, requests continued from undeclared agents that imitated an ordinary Chrome browser and rotated network addresses, and that content from test domains set to refuse all bots still appeared in Perplexity answers. Perplexity disputed the analysis, saying much of the traffic came from a third-party browser service and that its fetches are made on behalf of users.
- Harm
- Harm: No harm reportedAdds nothing to the index.
- Control
- Control level 3, Limit bypassReported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
Under review: the facts may change as more is reported. The rating may change; every change is logged below.
Sources
How we know
4 sources · single source. Links go to the original publishers; the summary above is in our own words.
- primaryRestricted-domain tests and crawler observationsCloudflare · Aug. 4, 2025blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-we…
- primaryAgents or Bots? Making Sense of AI on the Open WebPerplexity · Aug. 4, 2025perplexity.ai/hub/blog/agents-or-bots-making-sense-of-ai-on-the-open-web
- newsCloudflare Accuses AI Startup of 'Stealth Crawling Behavior' Across Millions of SitesTechRepublic · Aug. 5, 2025techrepublic.com/article/news-cloudflare-accuses-perplexity-stealth-crawling-vi…
- newsPerplexity hits back after Cloudflare slams its online scraping toolsTechRadar · Aug. 6, 2025techradar.com/pro/perplexity-hits-back-after-cloudflare-slams-its-online-scrapi…
Why this rating
No harm reported; control failure level 3
Two separate assessments. Only documented harm can count toward the index.
Observed harm
No harm reportedContent was fetched from sites that had blocked the crawler; no burden or loss was documented (methodology 4b).
Not a finding that no harm occurred.
Disputed: Perplexity disputes that the crawlers were its own, attributing the traffic to a third-party browser service it occasionally uses.
The harm scale
- 1 Negligible Inconvenience, easily remedied.
- 2 Minor Limited, recoverable harm.
- 3 Moderate Material harm needing significant effort to remedy.
- 4 Severe Severe harm to health, rights, property or essential services.
- 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.
Control assessment
Limit bypassAs claimed: automated agents bypassed robots.txt and WAF rules set by site owners, which the methodology treats as acting outside the rules set for the agent.
How: Got around a working safeguard.
Reported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
The control scale
- 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
- 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
- 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
- 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
- 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.
Rating rationale
Control 3 reflects the restriction bypass Cloudflare reports in tests and customer logs. Perplexity disputes attribution; this is not a finding of autonomous intent. Retained as Single source with operator attribution and under review.
Effect on the index
It is one of the control failures at the floor of the Aug. 4 reading, which did not change
The reading for the week to Aug. 4, 2025, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.
Not counted: no harm reported. 3 other records behind the reading for that week.
The arithmetic
| Step | With it | Without |
|---|---|---|
| Counts toward the index?documented, external, eligible evidence | No | — |
| Worst documented harm, ksets the band | — (none counting) | — (none counting) |
| Harms at that level, nposition in the band | 0 | 0 |
| Highest control level breachedsets the reading only when no harm counts | 3 | 3 |
| Readingrounded down | 3 Control failures only | 3 Control failures only |
Inside the window of 5 weekly readings
| Week to | Reading | Band |
|---|---|---|
| Aug. 4, 2025 | 3 | Control failures only |
| Aug. 11, 2025 | 3 | Control failures only |
| Aug. 18, 2025 | 3 | Control failures only |
| Aug. 25, 2025 | 3 | Control failures only |
| Sept. 1, 2025 | 3 | Control failures only |
Revisions
What we changed
5 logged. Every change to a rating is logged here, with the reason.
- v5Oct. 6, 2026
Ratings confirmed by the editor.
- v4Oct. 3, 2026
Audit corrections: impact unknown → not reported, since getting around limits on public resources is not harm unless a burden or loss is documented (methodology 4b) and none is. Status resolved → unknown as of 6 Aug 2025: Cloudflare blocked the pattern on its own network and Perplexity disputes attribution. Added Perplexity's 4 Aug 2025 rebuttal; removed the Harvard TagTeam mirror (a duplicate of Cloudflare's post, now unavailable); TechRepublic and TechRadar dated. Control unchanged.
- v3Sept. 30, 2026
Rated: impact unknown; control type restriction bypass.
- v2Sept. 30, 2026
Beta evidence review: Retained with explicit source provenance and bounded score basis.
- v1Sept. 30, 2026
Backfilled from public reporting.
Cite and share
Use this record
Citation
Paperclip Index. “Answer engine accused of using disguised crawlers to get around sites' robots.txt and firewall blocks.” Record PI-0021. Reported Aug. 4, 2025; updated Oct. 6, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0021