The only largest AI crawler on my web site over the previous day was not an AI crawler. It arrived roughly 1,500 occasions under Common Crawl’s name; it despatched again nothing, and what it needed was my SSH keys.
I went wanting due to a quantity.
Cloudflare’s CFO Instructed Analysts Machine Site visitors Might Attain 1,000 Occasions Human Site visitors
Cloudflare’s Chief Monetary Officer, Thomas Seifert, advised analysts on the corporate’s second-quarter earnings name that “if the present developments proceed, we predict in 5 years, non-human site visitors might be as a lot as 1,000 occasions as a lot as human site visitors.” Then the road that may seize the headlines: “people might be a rounding error on the web, not as a result of human site visitors goes down, however that’s simply how briskly we’re seeing non-human site visitors develop.”
Two issues price saying earlier than anybody reaches for the pitchforks. First, Seifert added his personal caveat, unprompted: “with the massive caveat that I’ve known as it unsuitable at each level alongside the best way.” Cloudflare beforehand anticipated machine traffic to pass human traffic in 2027, and it occurred in Could 2026. His errors have run towards beneathestimating, which is the strongest argument for taking the projection critically.
Second, the underlying measurement is actual. Cloudflare’s personal submit revealed the identical week says fewer than half of all HTML web page requests now come from a human. I’ve no argument with that. The machine guests are actual and they’re the entire topic of this web site.
The argument is about what the quantity counts.
What One Day of Crawler Site visitors on My Personal Web site Seems to be Like
I pulled Cloudflare’s AI crawler view for nohacks.co for the 24 hours ending the night of August 7. About 3,000 requests, of which roughly a 3rd had been unsuccessful, a determine up greater than 1,000% on the earlier interval.
By crawler: CCBot 1,510. ChatGPT-Person 375. ClaudeBot 296. Googlebot 245. PetalBot 107. 13 others sharing 353 between them.

CCBot is Widespread Crawl’s crawler, the long-running non-profit web archive whose corpus educated a great share of the fashions everybody now argues about. On paper, it being my largest customer is unremarkable.
Then I exported the paths.
It Requested for My SSH Keys, Not My Articles
Listed below are the most-requested paths in that AI crawler site visitors, with request counts, precisely as they got here out of the export:
/.ssh/known_hosts(42 requests)/phpinfo.php(31 requests)/.boto(30 requests)/.env.manufacturing(29 requests)/.vscode/launch.json(28 requests)/.env.check(27 requests)/firebase-service-account.json(26 requests)/.gitconfig(24 requests)/server/.env(24 requests)
It continues like that for 100 paths: /id_rsa, /id_ecdsa, /private-key, /ssl/localhost.key, /key.json, /serviceAccountKey.json, /.aws/config, /actuator/configprops, /api/v1/env, /Dockerfile, /values.yaml, and /@fs/proc/self/environ, which is an try at a identified path-traversal bug in a growth server.
Throughout these hundred paths: 1,028 requests, 6.7 MB transferred, and 0 referrals. The variety of requests to something I’ve really written rounds to nothing. The closest it got here to my content material was /weblog/wp-login.php, a WordPress login probe geared toward a web site that has by no means run WordPress, and two requests for /weblog/null.
That final element issues greater than it seems to be. No matter that is, it isn’t studying my pages earlier than it asks for issues. It’s working via an inventory, the identical listing it really works via all over the place, and my web site is a row in a loop.
This can be a credential scanner. Widespread Crawl follows hyperlinks and fetches pages, and it has no purpose to ask a podcast web site for its Firebase service account key.
I couldn’t confirm the supply addresses to show impersonation, as a result of per-request IP knowledge is just not one thing I can attain on my plan. Widespread Crawl publishes the check: real CCBot site visitors comes from documented deal with blocks and reverse-resolves to hostnames ending in crawl.commoncrawl.org. Somebody with these logs can settle it in a minute. What I can say is what arrived, what it requested for, and the way it was labelled: Cloudflare’s AI dashboard attributes this to Widespread Crawl because the operator, and counts each request towards my AI crawler totals.
Which results in the half that unsettles me most. I went on the lookout for these requests in my safety occasions and located nothing in any respect, as a result of the safety log solely data requests that journey a rule. I’m not blocking this site visitors, so it passes via, will get served, and leaves no mark. It seems in precisely one place on my complete dashboard: the AI crawler view, sitting within the listing beside ChatGPT-Person and Googlebot, beneath the title of a nonprofit analysis archive. A credential scanner is absolutely legible to me as agent site visitors and utterly invisible as a safety occasion.
2 of These Paths Are New, and They Are the Ones I Hold Pondering About
Buried in that listing are /.mcp.json, requested 30 occasions, and /.proceed/config.json, requested 24.
These two are agent tooling configuration: an MCP server definition and a coding assistant’s settings file. Each routinely maintain API keys and entry tokens, as a result of that’s what you set in them to let an agent attain your providers.
Somebody has added agent credentials to the usual secret-scanning wordlist. The identical automated sweep that has been asking each web site on the web for /.env since roughly endlessly now additionally asks for the file that lists which instruments your brokers can name and what they authenticate with. No one introduced that, and it occurred quick. When you run something agentic, the wordlist arrived earlier than most individuals completed writing their first MCP server.
Cloudflare Printed the Correction Itself, the Similar Week
The strongest counterweight to the earnings-call framing is in Cloudflare’s personal engineering writing from the identical week.
Their agentic-internet post says a number of site visitors from well-behaved bots is re-fetching pages that haven’t modified, and that this runs to billions of requests. Of their phrases, “an unlimited quantity of machine effort, connected to no consequence in any respect.”
Machine effort and machine demand are totally different portions. My very own logs are a sharper model of the identical level than I anticipated to seek out: the biggest single contributor to my machine site visitors was not merely ineffective, it was hostile, and it nonetheless counted.
Meta crawling your web site and by no means sending something again is the definition of ineffective site visitors in case you are the one who owns the web site. I wrote about that break up on August 1. A scanner sporting a analysis crawler’s title whereas it hunts in your cloud credentials is a class beneath that, and each land in the identical bar on the identical chart.
So when the graph climbs, the query for a web site proprietor is what the site visitors really is.
Assist Create the Drawback, Market the Drawback, Promote the Resolution
It’s clear what Cloudflare is positioning itself as right here, and it must be known as out. Assist create the issue, market the issue, promote the options. Within the first week of August alone: a bot-traffic projection on the earnings name, a weblog submit quantifying how a lot of the net is now not human, an agent-readiness scanner to inform you that you’re not prepared, an AI-visibility product to attain you, a bridge to reveal your web site’s instruments to brokers, and a default that starts blocking some of those agents in September except you resolve in any other case.
Each a kind of merchandise is an inexpensive response to one thing actual. That’s what makes the sample price noticing relatively than dismissing. The corporate measuring the issue, framing the issue, and promoting the repair is one firm, they usually now personal each the meter and the valve.
I wish to watch out right here, as a result of I’ve backed a number of what Cloudflare has performed. Pay-per-crawl was the suitable concept. Content material Independence Day was the suitable concept. Giving web site homeowners an actual alternative over which machines get in beats a courtroom deciding it for them, which is what I argued when the Ninth Circuit took up that question on August 4.
All of that may be true directly. Cloudflare can do some good issues, some directionally good issues, and a few issues that look sketchy, on the similar time. Most firms can. The error is deciding they’re the great guys or the dangerous guys after which studying every part they do via it.
Go and Take a look at Your Personal Logs
Take the site visitors numbers critically and take the framing with the salt it deserves. Machines are the vast majority of requests. That’s measured, and it’s true.
Then open your own crawler analytics and skim the paths, not the totals. Mine advised me three issues I didn’t know this morning: that my largest AI crawler was a scanner, that it was burning megabytes of my bandwidth on nothing, and that the wordlist it really works from now consists of the config information it thinks my agent tooling lives in.
None of that element is in anyone’s projection. The amount is. Fifteen hundred of those arrived at one small web site in a single day, each one in every of them counting towards the thousand-to-one Seifert described to analysts, and never one in every of them needed something I wrote.
Extra Sources:
This submit was initially revealed on No Hacks.
Featured Picture: Lightspring/Shutterstock
#Cloudflare #Machine #Site visitors #Hit #1000x #Human #Site visitors #Years
