Dominic Feron

The Gate Behind the Gate

Cloudflare's crawler rules show how a monopoly can be pressured without winning its customers, and why that still does not make the web decentralized.

A crawler arrives at the edge of a small website. It asks for a recipe, a product review, or the answer to a question somebody spent an afternoon working out.

The request looks ordinary. The argument begins only after the server asks what the crawler plans to do with the answer.

Search indexing? Come in.

AI training? Not necessarily.

On September 15, Cloudflare will apply that distinction more sharply. New domains with advertising pages will block crawlers classified as Training or Agent by default while allowing Search.

Customers who already chose to block AI training will also block mixed-purpose crawlers under the stricter rule unless they opt out. Cloudflare names Googlebot, Applebot, and BingBot as examples.

This is narrower than the viral version of the story. Cloudflare is not switching off Google for every small customer. Site owners retain a choice. Search remains an allowed category, and the default applies in defined cases.

Still, the narrower version reveals something more useful than another David and Goliath tale.

You do not always weaken a monopoly by persuading its customers to leave.

Sometimes you make it separate things its power had allowed it to combine.

Google’s search advantage is a loop. A broad crawl builds a better index. A better index attracts more searches. Those searches produce more signals and give publishers more reasons to remain visible.

A new search engine must build the index without the users and win the users without the index. This is the sort of business problem that looks much cleaner on a whiteboard than in a bank account.

Cloudflare attacks a different point in the loop. It does not need to build a better search engine. It sits between crawlers and a large population of sites. That lets it turn millions of weak individual preferences into one enforceable rule.

A recipe blogger cannot make Google redesign a crawler. An infrastructure provider can at least make the current design expensive.

The disputed bundle is purpose. Cloudflare classifies some bots as serving both Search and Training, then applies the strictest selected rule. Google says publishers can use the Google-Extended token to restrict future Gemini training without affecting inclusion or ranking in Search.

The two companies are not even describing the boundary in quite the same way.

That disagreement matters. The technical question of which bot fetched a page hides the economic question of which rights came with the fetch. Indexing, training, answering, and acting on a user’s behalf may use similar plumbing. They do not create the same bargain for the publisher.

The old bargain is already fraying. In Pew’s study of browsing activity in March 2025, people clicked a traditional result on 8 percent of Google visits that showed an AI summary. They did so on 15 percent of visits without a summary. A source inside the summary drew a click on only 1 percent of visits.

Search can still provide valuable traffic. But the price publishers thought they were paying for discovery has changed.

Cloudflare’s move therefore looks more like collective bargaining than insurgent competition. Small sites do not have to coordinate, negotiate, or understand crawler architecture in detail. A shared intermediary supplies enforcement and changes the default.

The incumbent then faces one concentrated counterparty instead of a million scattered complaints.

This is one of the few reliable ways to challenge a network-effect monopoly: find another layer where dependence runs in the opposite direction. Phone networks can be made to accept number portability. Software platforms can be required to expose interfaces. A search crawler can be asked to declare separate purposes.

The remedy does not recreate the incumbent’s network. It reduces the penalty for resisting it.

I find that more convincing than waiting for a heroic competitor to build the same flywheel from zero. I may also be wrong about how much pressure Cloudflare can sustain. Google could change its crawler behavior. Publishers could opt out of the block. The dispute could settle into a new checkbox that few people notice.

There is a larger problem too. Cloudflare can play this role because it has become a gatekeeper of its own. The power to aggregate small publishers’ choices is also the power to impose a private classification across a large part of the web.

We have not removed the gate. We have found the gate behind it.

That is why a corporate standoff is not yet a cure for monopoly. It becomes structural only when the separation can survive the company that first imposed it. It needs to become a portable standard, a market norm, a contract right, or law.

Otherwise Google may blink, but the internet remains governed by whichever concentrated layer can still close the door.