01 · What changed
Three things happened at roughly the same time.
AI companies started crawling at scale
Training data collection, retrieval for live answers, and indexing for AI search products all mean crawlers, and there are far more of them than there were two years ago. New ones launch constantly.
They crawl more aggressively than search engines
Googlebot has spent twenty years learning to be polite: it throttles, it respects crawl budget, it backs off when your server slows. Many newer crawlers do not have that discipline, and some of them re-crawl content that has not changed.
Hosting moved to visit-based billing
Managed WordPress hosting is now commonly priced by monthly visits rather than storage or bandwidth. That pricing model made sense when a visit meant a person. If the shape of your traffic has changed this much, the fix may be a differently sized hosting plan rather than a bigger one.
The result: a line item that used to be invisible is now on your invoice.
03 · Blind spot
Why you cannot see this in Google Analytics.
Google Analytics filters known bots out by design. That is normally the correct behavior, because you want your marketing reports to reflect humans.
But it means your analytics is structurally incapable of showing you this problem. You will look at flat traffic and a rising bill and reasonably conclude that your host has changed their pricing.
To see bot traffic you need visibility at the edge: a CDN or reverse proxy sitting in front of your site, reporting on every request before it reaches your server. Cloudflare analytics will break traffic down by bot versus human and identify individual crawlers by name. So will most equivalent services.
If nothing sits in front of your site, step one is putting something there, and the visibility is worth as much as the filtering. It is the first thing we put in place on any WordPress security engagement.
04 · Sorting the crawlers
Treating bots as one category is where most advice goes wrong.
Keep, always
Search engine crawlers
Googlebot, Bingbot and their verified variants. Blocking these is far more expensive than any hosting overage, and the damage is delayed and hard to attribute: nothing breaks, rankings just decline over weeks. We have dealt with an environment where search crawlers were being blocked by the host own bot handling, undetected, for months.
Probably keep
The AI assistants that send traffic back
If people are discovering your business through an AI search tool, blocking its crawler removes you from that surface. Whether that trade is worth making is a commercial judgment, not a technical one. Check your referral data first. If an AI product is sending you real visitors, its crawler is paying rent.
Block
Everything that returns nothing
SEO tool crawlers for platforms you do not use. Regional crawlers serving markets you do not sell into. Scrapers and content aggregators with no identifiable product behind them. Anything generating a high volume of 404s, which is usually a sign of either a badly behaved crawler or a problem on your own site worth investigating separately.
05 · What to do
Five steps, and the fourth is the one people skip.
Measure first
Get edge-level analytics in place and look at a full week. You want the human-to-bot split, the top crawlers by request volume, and what they are requesting. A week is the minimum useful window because crawl patterns are uneven day to day.
Decide, do not default
Go through the list and make a deliberate call on each significant crawler. Keep, keep for now, or block. If this is a client site, the client makes that call, because it is a decision about which surfaces their business appears on.
Block by name, at the edge
Block specific crawlers rather than raising a global bot protection setting. The global setting is one click and it is the wrong click. Block at the edge rather than in WordPress: a plugin-based block still requires PHP to load and process the request, so you pay the resource cost anyway. This is the whole of our bot and AI crawler control service, done with your approval on each crawler.
Verify the traffic is actually passing through your rules
If your DNS is not configured so that all traffic routes through your CDN, some requests reach your origin server directly, untouched by any rule you have written. We have seen exactly this: blocks deployed correctly, traffic barely moving, because a portion of it was never passing through Cloudflare at all.
RuleAfter deploying blocks, confirm the effect in your analytics. If the numbers do not move, suspect your DNS before you suspect your rules.
Re-check quarterly
New crawlers appear constantly and existing ones change behavior. This is a maintenance item, not a one-time fix.
What about robots.txt?
robots.txt is a request for voluntary compliance. Reputable crawlers honor it. The ones costing you the most frequently do not, and compliance among AI crawlers specifically is inconsistent and changes without notice.
It is worth maintaining, because for the well-behaved crawlers it is the polite and effective mechanism. It is not a control. If you need a crawler to stop, block it at the edge.
This guide is reference material behind our WordPress security work, and the quarterly re-check it describes is part of our maintenance service.
Common questions
Should I just upgrade my hosting plan?+
Only after you know what you are buying capacity for. If your traffic is genuinely growing, upgrade. If a third of it is crawlers that will never convert, upgrading means paying more every month, forever, to serve robots faster. Measure first, then decide.
Is blocking AI crawlers bad for my business?+
It depends entirely on which ones. Blocking a crawler whose product sends you referral traffic costs you visibility. Blocking a scraper that takes your content and returns nothing costs you nothing. Treating them as one category is the mistake.
Will Google penalize me for blocking bots?+
No, provided you do not block Googlebot. Google has no position on whether you allow other companies crawlers. The risk is entirely about accidentally catching search crawlers in a blanket rule, which is why blocking should be specific and verified.
How much traffic is bots on a typical site?+
It varies enormously by site type and content. Content-heavy sites attract more crawling than transactional ones. Rather than working from an industry average, measure your own, because the average tells you nothing actionable about your invoice.
My host says they handle bot filtering. Should I still do this?+
Probably. Host-level filtering is generic, invisible to you, and not adjustable. You cannot see what it is blocking or make exceptions, and in at least one case we have handled, it was blocking search engine crawlers. Edge-level control gives you both visibility and the ability to make exceptions.
Can bot traffic actually take a site down?+
Yes. Sustained crawler load consumes server resources the same way real traffic does. When a site hits its resource ceiling it starts shedding work, and what it sheds first is usually third-party API calls rather than public pages. So the site looks fine while your CRM sync, your booking system or your payment integration quietly fails.