In This Article

Back to blog

Access Denied: The Web’s New Default - Big Data LDN 2026 Recap

News

Vytautas Glambinskas

Last updated - ‐ 7 min read

Public web data was easy to reach for many years. If a page was online, you could usually read it, whether you were a person or a scraper. That’s changing in 2026. Websites are closing their doors to automated traffic, and the rules for who gets access are being rewritten.

That shift was the focus of “Access Denied: The Web’s New Default,” my talk at Big Data LDN 2026. I explained why the web is moving from open by default to closed by default, and what that means for every team that relies on public web data.

Here’s a quick recap of the key points.

The Deal That Kept the Web Open

For more than 25 years, the web ran on a clear trade. Publishers created content. Google crawled and indexed it. Search results sent clicks back to the publisher, who made money from ads or sales.

The click was the currency. It kept publishers motivated to publish and share.

IMG 1.webp

However, that trade no longer works the way it used to. AI-generated answers can now respond to questions directly, so fewer people need to visit the original source.

Here are some numbers that prove it:

  • US Google searches that end without a click grew from 60.5% in 2024 to 68% in 2026
  • AI Overviews have been linked to a 58% drop in clicks for top-ranking pages
  • Reach PLC, the publisher behind the Mirror, Express, and Daily Star, has lost 43% of its share price since 2024

Publishers’ content still feeds answers and summaries, but in many cases they no longer see any value from it.

Why Websites Are Pushing Back

Look at how many pages each crawler reads for every visitor it sends back, based on Cloudflare Radar data from mid-2026:

IMG 2.webp

In other words, Google reads about five pages for every visitor it sends back. Anthropic’s crawler reads 1,782. Cloudflare also reports that more than half of AI crawl traffic re-fetches pages that haven’t changed. For a publisher, that means bandwidth and server load with nothing in return.

For that reason, websites are starting to limit bot access in three ways.

1. Blocking by Default

Robots.txt never really enforced anything. Google followed the contents of this file because it wanted the traffic that came with the deal. Once crawlers started collecting content for answers instead of referrals, they had little reason to follow it.

So blocking moved to the network level, where requests can actually be refused. Two things changed as a result:

  • Crawlers are judged by what they do, not who they are - search, agents, training, and data collection are now separate decisions.
  • The default flipped - with robots.txt, sites were open unless they blocked you. Now they block bots at the network level and only let in the ones they approve.

IMG 3.webp

This matters even more because a single infrastructure provider sits in front of roughly 20% of the internet. When its defaults change, much of the web changes with it.

2. Treating All Bots as Suspicious

On June 3, 2026, automated traffic passed human traffic across Cloudflare’s network. Bots made up 57.5% of HTML requests, according to Cloudflare Radar , and humans made up 42.5%. Forecasts had predicted this for 2027, which means it arrived 18 months earlier.

IMG 4.webp

Agentic traffic, meaning software acting for a person in real time, grew by thousands of percent, though from a very small base.

The problem is that good bots and bad bots often look the same. The easiest response for a website is to treat every bot as a threat. That means legitimate, compliant data collection can then look exactly like an attack.

3. Moving Content Out of Reach

This isn’t about blocking crawlers on purpose. The content is simply moving to places crawlers can’t reach:

  • Behind logins - prices, listings, and reviews now sit inside accounts and apps
  • Inside AI answers - a generated reply has no URL and no archive, so there’s nothing to collect
  • Rendered per visitor - two people can see different prices and rankings on the same page

Each of these makes sense as a product decision. Together, they leave less public data to work with, creating the same effect as blocking.

Ready to get started?
Register now

The New System: Identify, Declare, Pay

Access isn’t going away. What’s changing is how much you can access anonymously and for free. The new system comes down to three steps:

1. Identify: who is asking?

For 30 years, a crawler identified itself with a user-agent string that anyone could fake. Now requests can be signed with a published public key (Web Bot Auth), which is much harder to fake.

2. Declare: what will you do with the data?

Robots.txt was mostly one-sided and opt-out by design. New standards such as Content Signals and RSL 1.0 give both sides a way to publish machine-readable terms. Search, answers, and training are treated as separate questions.

3. Pay: per request, not per deal.

The click used to be the payment. Now payment can happen inside the request itself. HTTP status code 402, “Payment Required,” was reserved in 1997 and went unused for 28 years. In 2026, it finally started being used, including through the x402 standard.

IMG 5.webp

Data Collection Has Its Own Door

In July 2026, Cloudflare published its bot categories . Alongside Search, Agent, and Training, there’s a category called Data Collection, covering “price scraping, competitive intelligence, third-party analytics.”

In other words, the largest bot management company on the internet lists price scraping as a behavior you declare, not as abuse. It’s also treated differently. When Cloudflare’s new defaults took effect on September 15, training and agent crawling were blocked on pages that carry ads. Data Collection wasn’t.

So there’s a way in: sites that use Cloudflare now have a recognized path for data collection bots. To use it, declare your purpose, sign your requests, and get listed. However, being identified also makes you easier to block. Still, it’s a more stable option long term.

As detection improves, approaches that rely on staying undetected require constant adjustment, while a recognized bot can maintain its status as the rules evolve.

Who Wins and Who Quietly Loses

A closing web doesn’t affect everyone equally.

Large companies will be able to adapt. They can sign licenses, absorb per-request costs, and send people to help write the standards.

Smaller players will feel the effects most. A price comparison startup, a regional retailer tracking competitors, a university research team, a small NGO. Anyone whose view of their market depended on free access to public pages.

For them, it won’t happen overnight. They’ll just see less of their market over time.

What You Can Do

If you use web data:

  • Know what you depend on - list every dashboard, model, and decision that relies on data from other websites, and note which bot category each one falls into.
  • Identify yourself instead of hiding - register with verified bot programs, sign your requests, and declare your purpose. Websites are more likely to allow access when they know what a bot is doing.
  • Treat data access like a supplier - web data access now comes with contracts, prices, owners, and renewal dates.
  • Use more than one source - don’t let one company’s settings become your single point of failure.

If you publish anything:

  • Audit your CDN defaults - a crawler that does several jobs gets the strictest rule you’ve set. If you block training, you may also block search crawlers on pages with ads. Some sites only found out from a drop in traffic.
  • Publish clear terms instead of blocking blindly - state what you allow instead of blocking everything.
  • Measure agent traffic - most analytics tools can’t separate agents from humans, so check yours before deciding.
  • Make someone responsible - if nobody owns these decisions, your CDN provider makes them for you.

The Takeaway

I closed the talk with one line: “The open web was never a right. It was a business model. And it is being repriced.”

For data teams, the smart move is to prepare now, be transparent about how you collect data, and build a setup that doesn’t depend on a single door staying open.

IPRoyal’s Residential Proxies and Web Unblocker can be part of that setup, helping your team collect accurate, localized data as the rules continue to change.

Create Account
Share on
Article by IPRoyal
Meet our writers
Data News in Your Inbox

No spam whatsoever, just pure data gathering news, trending topics and useful links. Unsubscribe anytime.

No spam. Unsubscribe anytime.

Related articles