Skip to content

Guide

Your hosting bill went up and your traffic did not.

If you are paying overage charges while your analytics show flat visitor numbers, those two facts are not in conflict. They are measuring different things. Your analytics counts people. Your hosting plan counts requests. In 2026 the gap between those two numbers is larger than it has ever been.

~7,000 Crawler hits, 7 days One site, one measurement
~30,000 Extrapolated monthly All billable, none of them customers
0 Of it visible in GA4 Analytics filters bots by design
1 wk Minimum useful window Crawl patterns are uneven day to day

01 · What changed

Three things happened at roughly the same time.

A

AI companies started crawling at scale

Training data collection, retrieval for live answers, and indexing for AI search products all mean crawlers, and there are far more of them than there were two years ago. New ones launch constantly.

B

They crawl more aggressively than search engines

Googlebot has spent twenty years learning to be polite: it throttles, it respects crawl budget, it backs off when your server slows. Many newer crawlers do not have that discipline, and some of them re-crawl content that has not changed.

C

Hosting moved to visit-based billing

Managed WordPress hosting is now commonly priced by monthly visits rather than storage or bandwidth. That pricing model made sense when a visit meant a person. If the shape of your traffic has changed this much, the fix may be a differently sized hosting plan rather than a bigger one.

The result: a line item that used to be invisible is now on your invoice.

02 · The numbers

What it looked like when we measured it.

On one recruitment site we measured, AI crawlers alone accounted for roughly 7,000 hits over a seven-day period. Extrapolated across a month that is around 30,000 visits, all billable against a plan priced by visit count, none of them a potential customer.

That client had been exceeding their monthly plan allowance and paying the overage every single month. The conversation they were having internally was about upgrading the plan. The actual problem was that a third of what they were paying for was not people.

Your numbers will be different. The point is that you cannot know until you measure, and almost nobody has measured.

03 · Blind spot

Why you cannot see this in Google Analytics.

Google Analytics filters known bots out by design. That is normally the correct behavior, because you want your marketing reports to reflect humans.

But it means your analytics is structurally incapable of showing you this problem. You will look at flat traffic and a rising bill and reasonably conclude that your host has changed their pricing.

To see bot traffic you need visibility at the edge: a CDN or reverse proxy sitting in front of your site, reporting on every request before it reaches your server. Cloudflare analytics will break traffic down by bot versus human and identify individual crawlers by name. So will most equivalent services.

If nothing sits in front of your site, step one is putting something there, and the visibility is worth as much as the filtering. It is the first thing we put in place on any WordPress security engagement.

04 · Sorting the crawlers

Treating bots as one category is where most advice goes wrong.

Keep, always

Search engine crawlers

Googlebot, Bingbot and their verified variants. Blocking these is far more expensive than any hosting overage, and the damage is delayed and hard to attribute: nothing breaks, rankings just decline over weeks. We have dealt with an environment where search crawlers were being blocked by the host own bot handling, undetected, for months.

Probably keep

The AI assistants that send traffic back

If people are discovering your business through an AI search tool, blocking its crawler removes you from that surface. Whether that trade is worth making is a commercial judgment, not a technical one. Check your referral data first. If an AI product is sending you real visitors, its crawler is paying rent.

Block

Everything that returns nothing

SEO tool crawlers for platforms you do not use. Regional crawlers serving markets you do not sell into. Scrapers and content aggregators with no identifiable product behind them. Anything generating a high volume of 404s, which is usually a sign of either a badly behaved crawler or a problem on your own site worth investigating separately.

05 · What to do

Five steps, and the fourth is the one people skip.

01

Measure first

Get edge-level analytics in place and look at a full week. You want the human-to-bot split, the top crawlers by request volume, and what they are requesting. A week is the minimum useful window because crawl patterns are uneven day to day.

02

Decide, do not default

Go through the list and make a deliberate call on each significant crawler. Keep, keep for now, or block. If this is a client site, the client makes that call, because it is a decision about which surfaces their business appears on.

03

Block by name, at the edge

Block specific crawlers rather than raising a global bot protection setting. The global setting is one click and it is the wrong click. Block at the edge rather than in WordPress: a plugin-based block still requires PHP to load and process the request, so you pay the resource cost anyway. This is the whole of our bot and AI crawler control service, done with your approval on each crawler.

04

Verify the traffic is actually passing through your rules

If your DNS is not configured so that all traffic routes through your CDN, some requests reach your origin server directly, untouched by any rule you have written. We have seen exactly this: blocks deployed correctly, traffic barely moving, because a portion of it was never passing through Cloudflare at all.

RuleAfter deploying blocks, confirm the effect in your analytics. If the numbers do not move, suspect your DNS before you suspect your rules.

05

Re-check quarterly

New crawlers appear constantly and existing ones change behavior. This is a maintenance item, not a one-time fix.

What about robots.txt?

robots.txt is a request for voluntary compliance. Reputable crawlers honor it. The ones costing you the most frequently do not, and compliance among AI crawlers specifically is inconsistent and changes without notice.

It is worth maintaining, because for the well-behaved crawlers it is the polite and effective mechanism. It is not a control. If you need a crawler to stop, block it at the edge.

This guide is reference material behind our WordPress security work, and the quarterly re-check it describes is part of our maintenance service.

Common questions

Should I just upgrade my hosting plan?+

Only after you know what you are buying capacity for. If your traffic is genuinely growing, upgrade. If a third of it is crawlers that will never convert, upgrading means paying more every month, forever, to serve robots faster. Measure first, then decide.

Is blocking AI crawlers bad for my business?+

It depends entirely on which ones. Blocking a crawler whose product sends you referral traffic costs you visibility. Blocking a scraper that takes your content and returns nothing costs you nothing. Treating them as one category is the mistake.

Will Google penalize me for blocking bots?+

No, provided you do not block Googlebot. Google has no position on whether you allow other companies crawlers. The risk is entirely about accidentally catching search crawlers in a blanket rule, which is why blocking should be specific and verified.

How much traffic is bots on a typical site?+

It varies enormously by site type and content. Content-heavy sites attract more crawling than transactional ones. Rather than working from an industry average, measure your own, because the average tells you nothing actionable about your invoice.

My host says they handle bot filtering. Should I still do this?+

Probably. Host-level filtering is generic, invisible to you, and not adjustable. You cannot see what it is blocking or make exceptions, and in at least one case we have handled, it was blocking search engine crawlers. Edge-level control gives you both visibility and the ability to make exceptions.

Can bot traffic actually take a site down?+

Yes. Sustained crawler load consumes server resources the same way real traffic does. When a site hits its resource ceiling it starts shedding work, and what it sheds first is usually third-party API calls rather than public pages. So the site looks fine while your CRM sync, your booking system or your payment integration quietly fails.

Keep reading

Find out what share of your bill is robots.

If you are paying an overage every month and your audience has not grown, it is worth an hour to find out why.

Book a scoping call
Ali Demirci
Ali Demirci
Founder
A week of edge-level data
The minimum that tells you anything.
The crawlers, named and ranked
By what they are costing you.
A decision, not a default
You make the call on each one.