Cover photo

AI Crawlers Are Turning Websites Into Metered Suppliers

Publishers and online businesses are gaining tools to measure, block and charge machine visitors, but pricing can also reduce discovery.

The web’s familiar business exchange is becoming less reliable. Search engines traditionally copied information from websites, then sent people back through links. Those visits could become advertising impressions, subscriptions or sales.

AI crawlers often complete only the first half of that exchange. They fetch pages to train models, build search indexes or answer a user’s question inside another product. The website still supplies the material, but the person may never arrive.

This is now an operating issue, not just a dispute between publishers and AI companies. Product catalogs, travel data, documentation, research archives and specialist databases can all become inputs to an automated service. New infrastructure is giving their owners a choice that was difficult to enforce before. Let the machine in, block it, or ask it to pay.

The traffic is not one thing

Cloudflare introduced an Attribution Business Insights dashboard in July to show customers how bots use their sites. It separates crawlers used for model training from those used for search or an agent acting on a person’s request. It also compares how frequently each operator crawls a site with how many visitors it refers back.

That separation is essential. A bot fetching a page because a customer asked an assistant to compare two products may create a commercial opportunity. A training crawler collecting thousands of pages produces no immediate visitor. Treating both as generic AI traffic hides the difference in value.

The measurements are still evolving. Cloudflare initially reported that 52 percent of crawler requests in June were for training. An independent recheck on August 1 found that Cloudflare had reclassified part of that traffic as mixed purpose. The corrected series put training at 44.54 percent of AI crawler requests in July, still the largest single declared purpose. Search and user-triggered requests together accounted for about 14.23 percent.

The correction does not reverse the business problem, but it is a warning against building policy around one headline number. A site needs its own traffic data before deciding what access is worth.

From access rule to payment rule

Cloudflare is also developing a Monetization Gateway that can charge for web pages, datasets, APIs and tools used by AI agents. It relies on x402, an open payment protocol named after the web’s largely unused “402 Payment Required” status code.

The sequence is simple in principle. An automated client requests a protected resource. The server returns a price and payment instructions. The client pays, repeats the request with proof, and receives the resource after verification. Cloudflare plans to perform the metering and payment check at the edge of its network, before the request reaches the customer’s server.

The first version will use stablecoins, digital tokens designed to track conventional currencies. That makes fractions-of-a-cent payments possible without requiring every agent to open an account with every website. Sellers would be able to retain the tokens or redeem them for conventional money.

This is not a finished mass-market payment channel. The gateway currently has an early-access waitlist, and the important buyers still have to support the protocol. Cloudflare’s position between websites and visitors also gives it a commercial interest in making metered access a new infrastructure category.

Other approaches are forming alongside it. The Really Simple Licensing standard lets publishers declare terms such as free use, attribution, subscriptions, payment per crawl or payment when content contributes to an AI-generated answer. Unlike an enforced payment gate, a licensing declaration still depends on crawler compliance or an infrastructure provider willing to block noncompliant requests.

What a business can do now

The immediate task is not to put a price on every page. It is to identify which digital assets lose value when they are copied and which gain value through wider discovery.

A public product page may work best when search engines, shopping assistants and customer-triggered agents can read it freely. The business wants the product considered, even if the buyer arrives through a new interface. A continuously updated pricing feed, proprietary comparison table or specialist research archive is different. It may deserve authenticated access, usage limits or a fee because it directly improves another company’s product.

That suggests a layered policy. Keep marketing information accessible. Measure which bots fetch it and whether they produce referrals or attributable sales. Place high-cost or high-value data behind clearer terms. Rate-limit repetitive requests that add server expense without creating a useful relationship. Reserve payment gates for material that has a plausible machine customer.

Pricing also requires experimentation. Charging too little can fail to cover data creation and payment administration. Charging too much can make an agent select a rival source or omit the business entirely. A per-request fee may suit a live API, while a license or revenue share may fit an archive whose value appears across many answers.

A market without a settled price

Independent analysis from Brookings warns that AI licensing could reproduce the same concentration that shaped search and social media. Large publishers can negotiate private agreements. Smaller sites may depend on collective licensing groups or intermediaries that control measurement, enforcement and payment.

There are technical uncertainties too. Crawlers can disguise themselves, ignore declared rules or obtain data from third parties. Stablecoin settlement introduces accounting and compliance work. A paid request proves that a transaction occurred, but it does not by itself show how the purchased information will be stored, combined or reused.

The practical change is that website access is becoming a business decision rather than a default. For years, most organizations treated automated crawling as a technical matter handled through a small text file and firewall rules. Now the same traffic can affect distribution, infrastructure cost, data licensing and sales.

The most useful response is neither a universal block nor an open door. It is a clearer inventory of what the site offers, who consumes it and what the business receives in return. Metered access will matter only when that exchange is valuable enough for both sides to keep making it.

Sources