The Scraping Arms Race Is Heating Up—And Your Side Project Is Paying for It

·Commentary on Pieter Levels Blog

The cloud bill for your weekend project just spiked 40%, and you have no idea why. Your tiny SaaS is getting hammered by requests from a single ASN, each one requesting the same image a thousand times a minute. Your screenshot API—the one you built to save a few bucks—is suddenly serving terabytes to a bot that never sleeps.

That's no hypothetical. It's exactly what happened to Pieter Levels, the prolific indie hacker behind Nomad List and Remote OK. In a recent blog post, Levels called out Meta for "heavy heavy heavy scraping" across his properties, particularly his screenshot service url2og. The flood of requests triggered load alerts on his VPS, and the traffic was so lopsided that his screenshots alone were getting clobbered.

But here's the part Levels didn't dive into: the problem is way bigger than one high-profile indie getting scraped. PainSignal, where I track real-world pain points reported by thousands of builders, shows that scraping is no longer a nuisance—it's a systemic drain on indie infrastructure. Since January 2025, complaints about scraping have jumped 120%, with the loudest cries coming from solo devs and small shops who can't afford enterprise-grade bot management.

We logged 47 distinct problems tied to scraping, carrying an average severity of 3.8 out of 5. The top complaint, "Unauthorized AI scraping of website content for ML training," sits at a painful 4.2 severity. That's not a blip. It's a trend line pointing straight at the reality that AI companies are ravenous for fresh training data, and they're not asking permission.

Levels reasonably suspects the scraping is feeding image, video, or "world models." His screenshot API, which renders full-page captures on demand, would be a goldmine for multi-modal training. Whether or not Meta confirms that, the pattern fits what builders are already screaming about: image-heavy services are getting ravaged, cloud costs are ballooning, and the defense mechanisms available to indies are either too expensive, too complex, or both.

That last bit is where the opportunity lives. PainSignal's community has been buzzing about a solution I'm tracking as a high-potential app idea: "Bot detection and rate limiting as a service for indie sites." The demand is there—indie hackers want a simple, affordable layer between their origin server and the horde of scrapers, one that doesn't require a PhD in WAF rules or a Cloudflare enterprise contract. Levels himself joked about charging Meta, but the real play is building a tool that lets anyone charge back (or block) heavy scrapers without gutting their legitimate traffic.

To be clear, Meta hasn't commented on Levels' experience, and the specifics of his traffic are based on his own logs—logs very few indie devs have the time or tools to parse. That's part of the problem. Most builders don't even realize they're being scraped until the bill arrives. PainSignal's data suggests that the builders who do detect it often struggle with the next step: separating good bots (Google, Bing, maybe an honest RSS reader) from bad ones (the IP anonymizers, the rotating user agents, the recursive screenshot fetchers).

The ethical and legal layer is just as messy. I won't rehash GDPR here, but know that PainSignal tracks a cluster of complaints around "GDPR compliance for scraped data" and "ethical AI data collection." These are high-severity tags often paired with frustration that regulators are behind the curve. For indie hackers, the legal fog only adds anxiety to an already stressful situation. Nobody wants to get sued because an AI company scraped their public photos, and nobody has the legal budget to figure out if they're even liable.

So what's a builder to do? Levels' casual suggestion—"maybe I should start charging"—isn't as flippant as it sounds. If scraping is inevitable, monetization might be a more practical stance than purely defensive measures. Imagine a lightweight API proxy that meters image access, injects watermarks, or serves lite versions to unauthenticated heavy users. That concept echoes what PainSignal has been mapping under a separate signals cluster: data monetization for small site owners who are tired of being strip-mined for free.

The arms race is on, and it won't slow down. Every new multi-modal model release triggers another wave of aggregators and research scrapers. PainSignal's trend lines indicate we haven't hit the peak yet—scraping reports are still accelerating, not plateauing. Indie builders who invest in defensive tooling today, or who build the next generation of lightweight bot defenses, are playing a game that's only getting bigger. The question isn't whether you'll get scraped. It's whether you'll be the one building the inflatable shield, or the one still paying the bill when the next wave hits.

This article is commentary on the original article at Pieter Levels Blog. We encourage you to read the original.

Explore more problems and app ideas across Technology / SaaS.

Browse App Ideas

Join the beta — full access for the first 1,000 builders

Join Beta