Reading Time: 6 minutes
When we use computers, it often feels like we never come to the end of something. Occasionally an app will be closed down—RIP Google Reader and Microsoft Bob—but we get used to versioning. A product is created, captures an audience, and then increments while it remains solvent, whether or not it is profitable. There is inertia that keeps people using it (and, increasingly dark patterns) and the inertia is broken only when the product becomes replaceable. That is the place I am at, finally, with the Google search engine.
Those of you who have read the blog will have seen me post occasionally about the travails of battling with the bots and scrapers. If anything, it has been the occasional reader who has reached out to me to say “I can’t see the site because of your firewall” that has kept me tinkering with access at all. It would be far easier to just require every request to face a firewall prompt but that really isn’t the internet that I want to inhabit. Cloudflare estimates that half of the internet’s crawler traffic are training bots.

One of the particular challenges of this problem is to let in the search engine crawlers while blocking the crawlers one doesn’t want. A few years ago, that was pretty easy: a request from Googlebot or Bingbot was a request that would be indexed and fed back through Google and the search engines that use the Bing index. Now, not so much. I’ve written already about the crawlers who are emulating the search engine crawlers in order to bypass any filtering.
But it’s not just them. Alphabet and Microsoft now leverage those crawlers for artificial intelligence training. Those are the mixed use crawlers that Cloudflare highlights. You cannot allow crawl access to your site for search without also providing crawling access for training. While there are ways to opt out of explicit training bots from Google at least, they do not guarantee that search crawled content is also kept out of of the training pool.
There are days when you realize that you’re not the only one battling a problem but you may not have a good idea of how others are handling it. It was good, then, to see that some very large internet publishers are rethinking their position in relation to the Google search engine. The traffic that the product partnership with Google brought has diminished. The value of the product as it iterates forward is no longer there and there is no partnership; it’s become one-sided.
This has clarified for me that, perhaps, being available on Google isn’t that relevant any longer. I have, by default, followed the advice Google gives for search engine optimization: produce high-quality, unique content. Okay, maybe the quality is not that great but it’s definitely authentically me. And, while I still get SEO hacks emailing me with proposals to get myself into the top 10 results—is that even a thing any more with AI results?—I have not modified my approach.
Increasingly, though, the visits from discernible people coming through from Google are few and far between. That may be because the AI scrapers have figured out how to scrape Google results (literally, just search for “scrape Google SERP“). In a way, it is funny. The Northern District of California just tossed out a Google case against a scraper in a way that reminds me of the West Publishing court losses against Matthew Bender and Hyperlaw.
The Court agrees with SerpApi in part. To the extent that Google Search results do not contain any copyrighted content, SearchGuard cannot be said to effectively control access to a work protected under the Copyright Act. Here, Google alleges that SearchGuard controls access to Google Search results, which are compilations of publicly-available information that Google obtains from the internet and organizes for presentation to users on google.com based on relevance….
However, Google does not allege that google.com or the Google Search results displayed therein are protected under the Copyright Act. Importantly, Google alleges that Google Search results are “often” accompanied by a “Knowledge Panel” that may contain some copyrighted content that Google licenses from third parties, such as copyrighted images. Google does not allege that the “Knowledge Panel” is always included in Google Search results, or that the Knowledge Panel, if included in the Search results, always contains copyrighted content. See id. ¶¶ 14-16. Accordingly, Google’s allegations indicate a mix of content, some with copyrighted material and others without.
Google v. SerpAPI, N.D. Cal., 7/20/2026, 4:25-cv-10826-YGR, p. 11
The content thieves like OpenAI and Anthropic have tried to defend their copyright infringement as fair use. Now they are being harvested themselves in so-called distillation attacks, where the attacker uses the AI itself to extract information for a new, smaller LLM. I have zero sympathy although it is interesting to see how the naming goes in this ourobouros. Distillation is a legitimate process in AI development—Anthropic points that out in its post about suing some Chinese distillers—so this will cause confusion. It’s like how the marketers have taken “human in the loop” from AI training to talk about how an end user has to be responsible for their AI use.
One thing with another, I’ve decided to just block the training crawlers full stop. That includes Google and Bing, so you will no longer find this site indexed by them. I regret having to do this a bit but not really that much. It’s a bit like leaving (or avoiding) platforms like Meta’s Facebook or Instagram or X or X’s Grok. What is the value when your contributions, whatever they are, are monetized for such evil purposes. The value exchange with search engines has been imbalanced for a long time; it was going to snap at some point. I guess that time is now for me. I can now see, on the other side, a point at which Google is no longer a search engine that crawls the known internet but, instead, is one that is blocked or paywalled. Its diminution may create space for new options to blossom.
But wait? How will people find this blog?
Sow for a Harvest
Probably the same way they have in the past. There are people who follow the website in their news reader: Feedly, Inoreader, RSS Owl, FreshRSS. Some folks pass around links to posts or add them to their own posts, so visitors come from a referral by someone who has read the post. I get a spike of traffic when AALL publishes a link in their daily newsletter. My content is linked in to LexBlog and, now, HeinOnline so it’s visible even if my primary site isn’t.
This organic traffic is probably the best. As some of you have noticed, I have started to add links to my navigation, at the bottom of the menus, with links to sites that I rely on. We used to have web rings of curated links and I think that there is a possibility that sort of interconnectivity will return, in light of the diminishment of search.
And this doesn’t have to be a perpetual outcome. If the future brings some new options, I’ll be keeping an eye on those. I like the idea that someone can serendipitously find my site when they need help building an irish dance stage or finishing a level of a computer game. One of the players that is doing interesting things is Cloudflare, which is currently my firewall.
Cloudflare does not come to the internet with clean hands. They are notorious as a company that serves up protection for bad actors, policing it with mixed results. I’m not sure I could continue to blog though without their free tier firewall. They are proactively creating tools that even someone like me, who is not monetizing anything and for whom any additional investment to just be present is likely to make me walk away, can utilize.
Right now, the AI-related firewall tools are iterating quickly. I find them a bit confusing and they are not always better than using the existing tools. For example, the list that I get from Known Agents (f/k/a Dark Visitors) of AI crawlers is far more extensive than the list on Cloudflare. Then, if you toggle to block the crawler here, all Cloudflare does is write a new rule with that block in it. But if you’re on the free tier, you only get 5 rules; an optimized rule will include a lot more than just this very basic block.

They also provide some curated collections of bots, so perhaps this is where they capture more of the ankle biters. This list is easy to apply in one of your custom rules. However, I find that the list of curated collections keeps changing and growing so you need to check back periodically to make sure you’re still blocking the ones you intend and add any new ones.

On September 15, 2025, they are going to be blocking AI training crawlers by default. They have started to add some fine-tuning for firewall users. For the most part, since I’m blocking in other ways, I’m not sure this will have much value for me. As long as it doesn’t impact my rules, I figure more is better.
On September 15, 2026, we’ll be setting new defaults for each of these three classifications. For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.
Your site, your rules: new AI traffic options for all customers, Cloudflare.com, July 1, 2026
I also appreciate that they are trying to come up with a way to monetize crawling so that the crawled can be compensated for the resource usage of the crawler. The idea is that, instead of just serving a 403 forbidden error, they get a slightly different server error that indicates payment is possible. Cloudflare serves as the payment manager. I’m not really into a lot of the other AI-supportive policies that they’re building, but I can see how other people might feel like they are necessary.
Cloudflare is also playing both sides, so they are optimizing their systems for AI agents to be able to interact with your dashboard. Unless you block AI agents.

There is always a potential risk with sticking with a large company. Like Cloudflare for a firewall. Or like WordPress for a content management system. For now, these risks are manageable and both of those companies embracing the AI side of the internet make sense to me. Those aren’t my choices but I have no employees nor financial stake in it.
Here’s an automated curl request from a user with a French IP address. They made about 300 URL requests in 10 seconds (not 300 resource requests). Frankly, until I don’t see this sort of thing, I’ll be skeptical that the technology companies have made enough changes.

At the same time, there is no driver for me to stay within what has become a deeply extractive internet. It doesn’t harm me at all to fail to appear on Google’s search engine results pages or in Claude or CoPilot answer set. I’m a bit sorry that people who might have found some value on this site may not but they are making their choices to use tools that are focused on extraction. I will continue to punch small holes for search engines like DuckDuckGo’s non-AI, Ecosia, or Qwant, the default search for EU computers. In other words, for search engines that are still about search.