A new wave of bots are swarming the web, and they’re coming for your bandwidth

Posted by Tiger Marketing on 20 November 2024

If you run a website in 2024, you’ll be familiar with the concept of a crawler: an automated browser of the web that gathers information typically for a search engine to index, i.e. they make your site discoverable on the web. This has been true since the late 1990s, when bots proliferated during the dotcom bubble. But now, new and hungrier bots are fuelling a fresh bubble – one that costs you bandwidth and could be scraping copyrighted content in order to to build increasingly complex libraries of data for ever-advancing AIs as big corporations fight for dominance in an emergent market.

Below is our guide to the current evolution of web crawlers in the age of AI, how that could affect or even harm your website and business, and how you can easily tackle this issue pre-emptively.

The dotcom bubble was similar to this new wave of AI, except back in the ’90s, site owners desperately wanted to be discoverable, so crawling was a welcome concept. The SEO industry that arose alongside the bubble spawned thousands of guides to being as crawlable as possible. That’s not to say the crawled data wasn’t also monetizable for the crawling company – they weren’t funnelling expensive computer resources into indexing the web altruistically. The most successful example is Google, who swiftly turned their omniscience into a business model, with a monopolistic ad empire overlaying their search engine with increasingly stealthy sponsored results.

Source: Statista – Bot traffic, at 49.6% of all web traffic now sits on the cusp of overtaking human users in 2025

Thanks to their crawling, Google learned a lot about all sorts of industries, which likely gave them the edge when it came to creating their own products and services. Google has broadened its business to offer software, hardware, hosting services and even attempted a social network. You simply can’t doubt the value of their vast oversight of the web, and plenty of other businesses do similar things, albeit at a fraction of their coverage. The huge storage and processing power needed to actually do anything useful with that much data at any usable scale is mind-boggling. Google transfers more than 1.2 Exabytes of data everyday – that’s 1 billion gigabytes.

Until recently there was no efficient way to sort that information. You’d need some way of compressing the data down whilst keeping it interpretable, and that interpreter would have to be able to group similar information together that’s presented in radically different ways across a multitude of source websites in different phrasings, dialects, and synonyms. You’d need some kind of… hmm, what would it be – something like a Large Language Model (LLM), perhaps?

Enter BERT

Before ChatGPT graced us with its presence, language models were already in use on Google Search. In 2018, BERT (Bidirectional encoder representations from transformers) was released in a major update. It wasn’t an interface where a user would enter a prompt into and get an answer out of; instead it was a modification to Google Search that tried to process your query as natural language instead of a combination of simple keywords. To highlight what we mean by that, consider the following search phrase:

“math practice books for adults”

None

Prior to BERT, the phrase was read by Google as a simple string of keywords to match. After BERT, each word in the string gave context to the others. Meaning was read ‘bidirectionally’, amending its understanding and weighting of the whole phrase as it goes.

The presence of ‘for’ adjusts the ‘math practice books’ massively, and ‘adults’ specifies the phrase further. As pictured above, the results for this type of search have improved considerably because of BERT. Under the surface, BERT was also quicker at processing queries too –because it had to be, given it was reading phrases every which way. That was the early version of an LLM at work, where the words in its corpus of data were ‘transformed’ and ‘encoded’ as ‘representations’ of related concepts.

“Google’s large language model BERT enriched its understanding of user intent, and made search more powerful”

If you were to delve into the model, it’d be a cloud of number values that represent the words it contains, and ‘for’ might appear several times across the model in different ways by its relation to other encoded words. It wasn’t a simple 1-to-1 words to values. Think of it like the human brain: a word you hear spoken triggering a chain reaction of synapses that conjure its definition and related memories. BERT allowed Google to do that with search queries, enriching its understanding of user intent, and making search more powerful.

Now that we have more LLMs than we know what to do with – from Google, OpenAI, PerplexityAI, plus the private projects that we know Apple and Meta are working on – there’s suddenly a whole lot more we can do with all the data on the internet. So the value of data on the internet just sky-rocketed. Which means the volume of crawlers and bots on the web has sky-rocketed too. Google might have a monopoly on search, but its a fresh frontier for AI.

How does this affect you?

It affects you in two ways:

  • The volume of bots visiting your site
  • The capacity of those bots to devour your content

Crawlers aren’t simple ‘headless’ bots that check your site HTML anymore. Thanks to the power of huge banks of specialised NVIDIA graphics cards in data centres across the world, they are fully rendering your website, JavaScript and all, to experience as close to how a human user might as possible.

That means they’re using your hosting bandwidth freely, opening accordion content, playing videos, listening to audio, searching for common phrases and interacting in every way you’d want a real user to. To employ a metaphor, a big multinational agricultural conglomerate is ploughing a self-driving combine harvester through your garden vegetable patch without as much as a ‘thank you’ for your fresh produce. You may think this is an extreme metaphor, but depending on your content’s perceived value to an LLM, it might be accurate.

A site called Game UI Database saw this exact scenario play out recently with crawlers from OpenAI. They were seeing 70GB of data being transferred from their hosting server every ten minutes. Their hosting was their own, and free to a degree (electricity for a busy server still adds up) but if they’d hosted with Amazon, they estimated the cost would be £850 per day. To add insult to injury, the level of scraping was even taking pages offline when the server was overloaded, much like a DDoS (distributed denial of service) attack would.

Part of the problem is a lack of technical understanding from website owners. We like crawlers, they make sure our sites are discoverable on search engines, or ones from services like SEMRush help us scan our sites for SEO issues so we can fix any discoverability issues ourselves. We want crawlers, and they shouldn’t be all that invasive. But now they’re coming for our content, and if you value your public content in any way, be it expertise, advice, best practices, or your creativity, then your value is being mined by bots.

“If you value your public content in any way – be it expertise, advice, best practices, or creativity – then that value is being mined by bots”

Detecting is even harder too. Google Analytics used to track bot traffic and even allow you to tick a box to filter it out. Now, the latest and much maligned update Google Analytics 4 does the filtering automatically and gives no option to track and segment it as you used to be able to. This means that if you are being hit by a large volume of AI bot traffic then GA4 will keep schtum about it.

Monitoring for bots with Matomo

When we investigated our own bot traffic levels we used a privacy-centric, self-hosted Analytics platform called Matomo which has a bot tracker plugin to log any crawlers we could identify by User Agent. That plugin comes with an inbuilt list of user agents to keep an eye out for. We uploaded an extra list of the latest common AI bots to supplement this.

Fortunately for us, AI bots don’t appear to be gracing us frequently yet. We don’t have a vast database of learning resources like Game UI Database does.

Pingdom and SEMRush are our own tools keeping an eye on the site’s health. Then, Bing, Apple, and Facebook send their own bots alongside traffic, often to help provide instant previews and extra bits of tracking for their users and systems.

If we’d had AI bots stop by they’d be distinct from these – the above businesses having the following:

  • For Facebook it’s the ‘Meta-ExternalAgent’ that trains their in-house AI models off of your website data
  • For Apple it’s the ‘Applebot-Extended’ bot
  • Microsoft uses Chat-GPT and doesn’t appear to be doing their own AI crawling (yet), so for them it’s ‘GPTBot’ or ‘ChatGPT User’
  • And though they’re not in the pie above, Google’s is the ‘Google-Extended’ bot

What you can do about it

Game UI Database opted for an IP ban targeting the OpenAI crawler, which is equivalent to starting a game of whack-a-mole with a multi-billion dollar company, which can respond by changing their IP, or sending the bot out from a dynamic IP. Other corporations can catch on to the value of the database and pile on too.

As of now, the best solution we’ve found is to log everything and ask nicely. We acknowledge that big corporations are very greedy, but people are pushing back: other big businesses are protecting their content, and individual artists and creators are suing the AI companies. As a result the corporations are treading lightly, preferring partnership deals with the likes of Reddit to gather data contractually instead of just through scraping. So, what you can do is implement a tracking solution that monitors bot traffic on your site, and upload a robots.txt file that lists all of the bot crawler’s User Agents under a disallow rule. They don’t have to follow your request, but at least you’ll have the receipts when they do crawl your site. Plus, you can take more extreme measure like IP bans if you know who the culprits are. There’s even a plugin for WordPress that links your robots.txt to an always up-to-date list of crawlers, so you have the edge in that game of whack-a-mole.

Quickly add to your existing robots.txt from this up-to-date repository on GitHub. Or, use the same repository creator’s API or WordPress plugin to keep your robots.txt up to date constantly.

Recent posts...

ROAR Louder

Why pay for one marketer when you can have a full team?

Get in touch if you want to be heard