Every few weeks somebody forwards me a hosting alert about bandwidth, or a plugin email about "bot traffic detected", and asks a version of the same question: who is crawling my website when nobody is looking? I used to answer from experience. This time I counted.
I pulled 97 days of raw access logs off my own server — 603,905 lines, 19 MB, 15 June to 19 September 2026 — and sorted every request by who made it. Our own cron jobs, uptime monitor and tooling are excluded; counting yourself as a visitor is how you get a big number and a wrong story. What was left: 84,065 requests from 25 crawlers that identify themselves by name, 867 a day, against 377 published pages.
Then I sorted them into categories. AI crawlers made 33,283 of those requests. Googlebot made 3,311. Ten to one, in favour of the bots that send nobody to your website.
- 25 named crawlers, 84,065 requests, 97 days, one small business site.
- AI crawlers: 33,283. Search engines: 27,015. SEO and sales tools: 22,138.
- Googlebot came eighth. Meta's crawler came first, with 12,084 visits.
- All 73 pages I published in that window were fetched by an AI crawler within a day. Googlebot got to 66 of them.
- 443 requests claiming to be Googlebot — 11.8% — were not Googlebot. You can check this yourself in two commands.
- Four crawlers never asked for robots.txt once, including the busiest one.
- The whole lot cost 1.5 GB of bandwidth in three months. Do not buy anything.
Who is actually crawling my website?
Mostly companies you never hired and cannot invoice. Over 97 days, 25 named crawlers made 84,065 requests: 33,283 from AI crawlers, 27,015 from search engines, 22,138 from SEO and sales tools, 1,629 from social link previews. Googlebot was eighth on that list.
Here is the whole census. Each row counts requests carrying that crawler's user agent, which is all a log can tell you on its own — the Googlebot row gets checked properly further down, and 443 of it turns out not to be Google. "Pages" is how many of my 377 published pages each crawler asked for, a better measure of interest than raw hits.
| Crawler | Type | Requests | Per day | Addresses | Pages |
|---|---|---|---|---|---|
| Meta AI crawler | AI | 12,084 | 124.6 | 443 | 347 |
| AhrefsBot | SEO tool | 7,515 | 77.5 | 2,583 | 375 |
| Bingbot | Search | 7,175 | 74 | 374 | 353 |
| PetalBot (Huawei) | Search | 6,699 | 69.1 | 9 | 363 |
| SemrushBot | SEO tool | 5,342 | 55.1 | 63 | 356 |
| Other SEO and sales tools | SEO tool | 4,812 | 49.6 | 385 | 367 |
| ClaudeBot (Anthropic) | AI | 4,238 | 43.7 | 195 | 374 |
| Googlebot | Search | 3,763 | 38.8 | 392 | 356 |
| Other AI crawlers | AI | 3,441 | 35.5 | 134 | 344 |
| Bytespider (TikTok) | AI | 3,336 | 34.4 | 426 | 18 |
| Amazonbot | AI | 3,311 | 34.1 | 443 | 368 |
| SeznamBot | Search | 2,713 | 28 | 44 | 167 |
| MJ12bot (Majestic) | SEO tool | 2,557 | 26.4 | 162 | 351 |
| Google, other crawlers | Search | 2,258 | 23.3 | 114 | 342 |
Two rows are worth holding on to. PetalBot, Huawei's search crawler, made 6,699 requests from nine addresses. AhrefsBot made 7,515 from 2,583 — about three requests per address. That second number comes back the moment anyone suggests blocking bots by IP.

Do AI crawlers really visit more than Google?
On my site, by a wide margin. AI crawlers made 33,283 requests against 27,015 from every search engine combined, and 10.1 times Googlebot's own 3,311. Meta's crawler alone came 12,084 times — more than three times Googlebot's entire crawl.
Before anyone builds a strategy on that: a crawl is not a reader. Googlebot's 3,311 visits belong to a search engine that sent real people here in the same window. Most of those 33,283 AI fetches are attached to nothing I can measure.
I checked whether the gap was an artefact of my own measurement before believing it. Raw request counts can be undercounted (see the last section), so I used a number that is harder to distort: how many different URLs each crawler asked for. AhrefsBot fetched 1,290, ClaudeBot 1,131, Meta 1,115 — Googlebot 769. The ordering survives.
How often does Googlebot actually come?
Every single day — 34.1 verified visits a day on average, reaching 356 of my 377 pages. That is a healthy crawl for a small site. The interesting part is where the volume went: 393 of those requests landed on a redirect and 89 on a page that no longer exists.
That is 12.8% of Googlebot's time here spent on URLs that are not pages — self-inflicted, and the one number in this article you can improve this afternoon. Old links inside your own site, stale URLs in your sitemap and redirects chained to other redirects are what produce it. The cleanup is in why old pages still show up in Google and what internal links actually do.
How fast does anything find a brand-new page?
Faster than the folklore says, and the AI crawlers are fastest. I matched the publish date of all 73 pages I put out in this window against the first time each crawler asked for that URL. Every one of the 73 was fetched by an AI crawler within a day. Not most. All of them.
| Who found it | Within a day | Median delay | Never came |
|---|---|---|---|
| Any AI crawler | 73 of 73 | 0 days | 0 |
| Googlebot | 66 of 73 | 0 days | 3 |
| Bingbot | 33 of 73 | 2 days | 4 |
Being fetched is not being indexed, and certainly not being ranked — I measured that gap in how long Google takes to index a new page. But it kills one common worry stone dead: if your new page is not showing up, the problem is almost never that nothing came to look at it.
Treat the three "never came" entries with suspicion rather than alarm. Two of them have impressions in Search Console, so Google did crawl them and my log never saw it — a hole in my instrument that gets its own section below.
Which bots bother to read robots.txt?
Some read it obsessively, some have never looked. ClaudeBot requested robots.txt 1,054 times across 4,238 visits — one check for every four requests. MJ12bot, 1,004 times. SemrushBot, 955. Googlebot, 192. And four crawlers responsible for over 18,000 requests between them never asked once in 97 days.
- Never requested robots.txt: Meta's AI crawler (12,084 requests), Amazonbot (3,311), GPTBot (1,680), Baiduspider (1,394).
- Checked it constantly: ClaudeBot (1,054 checks), MJ12bot (1,004), SemrushBot (955), OAI-SearchBot (670).
One of those zeroes has an innocent explanation: OpenAI documents that when a site allows both of its bots, "we may use the results from just one crawl for both use cases", and OAI-SearchBot did fetch robots.txt 670 times. GPTBot's zero is probably efficiency. Meta's 12,084 requests without a single check, and Amazon's 3,311, I cannot explain away from this log.
Here is the part most owners have backwards: robots.txt is not a lock. RFC 9309, the specification that defines it, says the rules are "not a form of access authorization" in as many words, and tells crawlers not to use a cached copy for more than 24 hours. A crawler that ignores your robots.txt breaks a convention, not a control.

Is one in nine "Googlebots" a fake?
On my log, yes. 3,763 requests in this window carried a Googlebot user agent. 443 of them, from 363 different addresses, were not Google at all — 11.8%. They failed both tests: the address sat outside Google's published crawler ranges, and Google's own reverse DNS check said no.
The impostors were not one busy attacker. They were 363 addresses making one or two requests each from ordinary hosting providers — an AWS instance, a Vultr box, a colocation machine. Somebody's scraper, wearing a name that gets let through. Thirteen requests went further and claimed to be "Google-Extended", a crawler Google's own documentation says has no user agent string at all.
Checking takes two lookups, and it is the most useful thing in this article. Google publishes the method:
- Run
hoston the address from your log —host 66.249.66.1, with your address in place of that one. - Check that the name it returns ends in
googlebot.com,google.comorgoogleusercontent.com. - Run
hoston that name and confirm it gives back the same address you started with.
If step two or three fails, that visitor is not Google, whatever it typed. Microsoft documents the same trick for Bingbot, and OpenAI, Anthropic and Meta all publish their address ranges.
The busiest "crawler" in my log was lying
My first pass had a bucket for feed readers and WordPress.com, and it topped the whole table with 23,984 requests — more than Meta, more than double Googlebot. Every user agent in it said Jetpack or WordPress.com.
Then I looked at what it was asking for. 23,960 of the 23,984 requests were POST /xmlrpc.php and nothing else, from 421 addresses. And the version numbers inside the user agent — Jetpack 12.0, 12.1, 12.5, 13.0 crossed with WordPress 6.1 through 6.4 — produced 16 version pairs, all appearing on the same day, every day. One website's plugin reports one version. This was an XML-RPC password-guessing run wearing a plugin's name, which is why it sits in its own bucket and in none of the numbers above.
The lesson outlives my log: a row in a report is not a fact until you look at what is inside it. Ship the first pass and the headline becomes "WordPress.com is the biggest crawler on the web", which is nonsense. More of that pattern in what probe traffic on a small site actually looks like.
Should I block the AI crawlers?
Not with one blunt rule, because "AI crawler" covers at least three different jobs and blocking the wrong one costs you traffic. Every statement in the table below comes from the company's own documentation, not from an SEO blog.
| robots.txt token | Blocking it stops | Blocking it does NOT stop |
|---|---|---|
| GPTBot | Your content training OpenAI's models | Appearing in ChatGPT's search answers |
| OAI-SearchBot | Appearing in ChatGPT search answers | Model training — that is GPTBot |
| Google-Extended | Gemini training and grounding | Inclusion or ranking in Google Search |
| ClaudeBot | Future content entering training data | Claude-User fetching a page a person asked for |
Two of those deserve spelling out. Google states that Google-Extended "does not impact a site's inclusion in Google Search", and is not a ranking signal either — so blocking it is a training decision and nothing else. OpenAI states that sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers". For a service business, that second one is a referral channel you are switching off.
What definitely does not work is blocking by address. AhrefsBot used 2,583 of them in 97 days on one small site; Applebot used 854. Any IP list you write is stale by the weekend, and the real cost arrives the day you block a customer's office by mistake. Name the crawler in robots.txt instead.
What is all of this costing me?
Less than the plugin that offers to stop it. All 84,065 crawler requests together moved 1.5 GB over 97 days — about 470 MB a month on a site of 377 pages, inside the cheapest hosting plan anyone sells.
Crawler traffic becomes a real cost in one situation: when a bot repeatedly hits a page your server builds from scratch — a search results page, a filtered product list, a calendar. Static pages cost nothing to serve a hundred times. If your host is complaining about CPU rather than bandwidth, that is a caching problem, not a security one, and not a reason to buy anything. Junk on the hosting bill is its own genre; I catalogued it in how hosting companies quietly bill you $400 a year.
Why your access log undercounts everything
My HTML sits behind a ten-minute cache. Rather than assume what that does to the log, I tested it: four requests for the same page under a unique user agent, one with a cache-busting query string and three without. The cache-busting request appears in the log. The three cache hits appear nowhere. They were served, they returned 200, and the server that writes my log never knew.
So every number here is a floor, and one page proves it. My log says Googlebot never fetched find-your-website-competitors — not once in 97 days. Search Console says that page earned 252 impressions and a click in the same window, which it cannot do uncrawled. Google came, the cache answered, my log stayed blank.
The bias runs the other way much harder. My log recorded 36,097 page views from agents that looked like ordinary browsers. Analytics, which only counts a browser that actually ran JavaScript, recorded 2,386 in the same window — a factor of fifteen. A Chrome user-agent string is the cheapest disguise on the internet, and most of what looks human in a raw log is not.
A server log is very good at telling you who visits, because well-behaved crawlers volunteer their names and you can verify the important ones. It is bad at telling you how many, in both directions. For how many, use Search Console and your analytics, read together.
How to read your own log in twenty minutes
- Find it. cPanel calls it Raw Access Logs, Plesk calls it Logs, and on managed WordPress hosting you may have to ask support. You want the raw file, not the pretty dashboard built from it.
- Check whether a cache sits in front. Request one page twice and look for an
X-Cache-Statusheader. If the second one says HIT and your log has one line, you now know your numbers are a floor. - Count by user agent. Sort, group, count. Almost all of the value is in the top twenty rows.
- Bucket them: search engines, AI crawlers, SEO tools, social previews, unknown. Then open each bucket and look at what it actually requested, before you believe any of it.
- Verify the Googlebot lines with the two
hostcommands above. - Count the non-200 responses in the search-engine rows. That is crawl spent on nothing, and it is yours to fix.
- Compare the "human" total with your analytics. A large gap is normal; it tells you how much of your log is automation in a browser's clothing.
- Write down the date and the totals. One reading is trivia. Two readings three months apart is a trend, and the second one takes five minutes.
Do not buy anything because of this article. Nothing in 97 days of my logs justifies a bot-blocking subscription on a small business site. If you take one action, make it step five — verify the Googlebot lines — and if you take two, fix the 12.8% of crawl landing on dead URLs. Both are free.
Frequently asked questions
Who is actually crawling my website?
Mostly companies you never hired. Over 97 days on my own site, 25 named crawlers made 84,065 requests: 33,283 from AI crawlers, 27,015 from search engines, 22,138 from SEO and sales tools, and 1,629 from social link previews. Googlebot was eighth on that list, with 3,311 verified visits.
Do AI crawlers visit more than Google does?
On my site, by a wide margin. AI crawlers made 33,283 requests in 97 days against Googlebot’s 3,311 — about ten to one. Meta’s crawler alone came 12,084 times. This is one small business site, not the whole web, but the gap is too large to be noise.
Is bot traffic bad for my website?
Usually not. All 84,065 crawler requests together moved 1.5 GB in three months, about 470 MB a month, which fits inside the smallest hosting plan sold. Bot traffic becomes a problem when a bot hammers a slow database page, not because it exists.
How do I know a visit from Googlebot is really Googlebot?
Run a reverse DNS lookup on the address, check that the name ends in googlebot.com, google.com or googleusercontent.com, then look that name up again and confirm it returns the same address. Google documents exactly this. On my log, 443 of 3,763 requests claiming to be Googlebot failed it.
Should I block AI crawlers in robots.txt?
Only if you know which one you are blocking. GPTBot controls training. OAI-SearchBot controls whether you appear in ChatGPT search answers. Google-Extended controls Gemini training but, in Google’s own words, does not impact inclusion in Google Search. A blanket block costs you the referral side too.
Do crawlers have to obey robots.txt?
No. The standard is voluntary — RFC 9309 says plainly that the rules are "not a form of access authorization". On my log four crawlers never requested robots.txt once in 97 days, including the one that visited most. Well-behaved bots read it; nothing forces the rest.
Why does my access log show far more visitors than Google Analytics?
Because a browser user-agent string is the cheapest disguise there is. My log recorded 36,097 page views from agents that looked like ordinary browsers. Analytics, which only counts a browser that actually ran JavaScript, recorded 2,386 in the same window. Trust the smaller number.
Sources, read on 20 September 2026: Google — Verify requests from Google crawlers and fetchers; Google — Common crawlers (Google-Extended); OpenAI — Overview of OpenAI crawlers; Anthropic — Does Anthropic crawl data from the web; RFC 9309 — Robots Exclusion Protocol. Traffic figures are from this site's own raw SSL access logs, 15 June to 19 September 2026.
Want to know who is crawling your site?
Send me your domain. I'll tell you which crawlers can reach it, how much of Google's crawl is landing on dead URLs, and whether anything in your robots.txt is costing you traffic — in plain English, no invoice attached.
Get My Free Quote


