Field note

Are AI-scraping lawsuits changing what crawlers can take? The 2026 cases

Two 2026 cases sharpened the fight over AI crawler access: News Corp sued Brave (filed 22 July 2026) alleging it disguised crawlers to evade publisher blocks, and a federal judge on 31 July 2026 largely refused to dismiss Reddit's 'industrial-scale' scraping suit against Perplexity. Here is what each case alleges, how licensing deals run alongside the litigation, and what the access fight means for whether your content gets cited.

Buffy Editorial2026-08-13 · 5 min read

The 2026 AI-scraping lawsuits are a fight over how crawlers obtain content, not whether AI can cite you — and the early rulings favour publishers. In July 2026 two cases raised the stakes: News Corp sued Brave (filed 22 July) for allegedly disguising crawlers to evade blocks, and a federal judge on 31 July largely refused to dismiss Reddit's "industrial-scale" scraping suit against Perplexity. Both turn on bypassing access controls, which is the same behaviour our crawler guidance has long said a self-declared user-agent can't be trusted to reveal.

Last reviewed: 13 August 2026. Every legal fact below is attributed and dated; complaints and interim rulings are not verdicts, so read them as the state of play in mid-2026, not settled law. This piece explains what each case alleges, how licensing deals run in parallel, and what the access fight means for your AI visibility.

What did News Corp accuse Brave of?

News Corp filed suit against Brave on 22 July 2026, alleging that Brave disguised its web crawlers to evade publisher blocks and then scraped and sold copyrighted News Corp content to AI companies before March 2025. The complaint frames the disguise as the wrongdoing, not just the copying.

"They have shamelessly stolen and then perfidiously profited from that pilfering by illicitly fencing our journalists' work." — News Corp CEO Robert Thomson, on the Brave complaint (as reported, July 2026)

The disguise allegation is the durable part. A block only works if the blocked party identifies itself honestly, and the corpus has documented repeatedly that a crawler's name is self-reported and trivially spoofed. What is new is a large publisher asking a court to treat evading the block by hiding identity as unlawful. If that theory holds, covert access becomes a legal liability, not a grey-area growth tactic. The claims are unproven; Brave has not conceded them.

What happened in Reddit v Perplexity?

On 31 July 2026, a federal judge largely rejected Perplexity and SerpApi's motion to dismiss Reddit's lawsuit, letting the core claims proceed. The court found Reddit had "plausibly pleaded that Perplexity conspired with at least one of the three data scrapers to bypass access controls."

The specifics that make it citable:

  • What Reddit alleges: "industrial-scale" bypassing of its technical protections to obtain Reddit content for an AI answer engine.
  • Who is named: Perplexity, plus data-scraping suppliers SerpApi, Oxylabs and AWMProxy as co-defendants.
  • The stage: surviving a motion to dismiss means the case continues; it is not a finding of liability.

The one-line read: a court has now let a "bypassing access controls" theory proceed against both an AI engine and its scraping suppliers. That extends legal exposure down the supply chain, to the vendors who actually do the fetching, which is exactly the "scraping vendor wearing a costume" pattern the corpus has flagged.

How do licensing deals fit alongside the lawsuits?

They are the other half of the same market: publishers are simultaneously suing some operators and licensing to others. The reported deals give the fight a price tag. All figures are as reported mid-2026; most deals publish no terms.

Deal Reported terms Note
News Corp — OpenAI $250M+ over 5 years Largest single reported figure
News Corp — Meta Up to $50M/year (≥3 years) Same publisher, multiple buyers
Amazon — The New York Times $20–25M/year Ongoing content licence
Google — Reddit Undisclosed Reddit content for Gemini training
Nine — Microsoft Undisclosed (2026) Terms not published

Source: Press Gazette and AI Business reporting, mid-2026; figures where a publisher or filing disclosed them. The pattern is a multi-buyer licensing market forming next to active litigation — the same publisher (News Corp) both licenses to OpenAI and Meta and sues Brave. Read it as leverage: content owners are being paid where access is negotiated, and going to court where they say it was taken.

What does the access fight mean for your AI visibility?

Mostly it reinforces a strategy the corpus already recommends: be reachable to identifiable, well-behaved crawlers, and enforce access by verified identity, not by trusting a name. The lawsuits do not threaten citation of content a crawler was permitted to read; they target covert access. Three practical implications:

  1. Blocking is a real trade-off, now with legal weather around it. If you block AI crawlers, the operators that respect the block lose the ability to cite you, while litigation pressure pushes the rest toward licensed access over evasion. Decide per crawler class rather than blanket-blocking.
  2. Enforce by verified identity. Because a self-declared name is spoofable, confirm crawlers against published IP ranges with a verified-bot check, per seeing which AI bots crawl your site. That is the same signal these cases turn on.
  3. Licensing is becoming a lever, not just a defence. Large publishers are monetising access directly; the emerging pay-per-answer and charge-the-crawler models are the infrastructure version of the same shift for everyone else.

The through-line: the 2026 cases are narrowing the space for covert scraping, which makes legitimate, identifiable crawling the norm engines are being pushed toward. For a brand that wants to be cited, that is the access posture to optimise for.

Buffy Intel tracks whether your content is actually being reached and cited across AI engines — the citation side of the same access story these lawsuits are about. See where you show up, and where a block or a missed crawl is quietly costing you answers. Questions: [email protected].

Frequently asked

What is News Corp's lawsuit against Brave about?

News Corp filed suit against Brave on 22 July 2026, alleging that Brave disguised its web crawlers to evade publisher blocks and then scraped and sold copyrighted News Corp content to AI companies before March 2025. In News Corp CEO Robert Thomson's words, the company said its journalists' work was 'shamelessly stolen'. The legal theory that matters for the wider web is the disguise allegation: it treats evading a block by hiding a crawler's identity as an unlawful act, not merely a technical workaround. The claims are unproven and Brave has not conceded them, so treat this as a filed complaint, dated mid-2026, not a ruling.

Did Reddit win its case against Perplexity?

Not yet, but it cleared the first hurdle. On 31 July 2026 a federal judge largely rejected Perplexity and SerpApi's motion to dismiss Reddit's suit, finding Reddit had 'plausibly pleaded that Perplexity conspired with at least one of the three data scrapers to bypass access controls'. Reddit alleges 'industrial-scale' bypassing of its technical protections, with data-scraping suppliers SerpApi, Oxylabs and AWMProxy named as co-defendants. Surviving a motion to dismiss means the case proceeds; it is not a finding of liability. Facts are as reported mid-2026 and will move as the case develops.

Do these lawsuits mean AI engines will stop citing my content?

No. The suits target how some operators obtained content (evasion and bypassing access controls), not the act of citing a source that a crawler was allowed to read. The likely direction is toward licensed or clearly-permitted access rather than covert scraping, which rewards being reachable to identifiable, well-behaved crawlers. If anything, the practical lesson for a brand that wants citations is the opposite of blocking everything: keep legitimate AI crawlers able to reach you, and enforce access by verified identity rather than trusting a self-declared name.