The Web Needs a Front Door
How the world slowly is losing access to a resource that was meant to be open
The web’s original breakthrough was radically low-friction access. CERN released its core technology royalty-free, allowing anyone with an internet connection to publish information, link to others, and build without asking a central platform for permission.
Then, decision by decision, the deal changed. Each change had a reasonable justification, based on security, convenience, or performance. Hosting consolidated into large data centers. CDNs and security platforms increasingly sat between websites and their users.
None of these decisions announced themselves as “we are closing the open web.” But add them up and that’s what happened. The web didn’t die. It got concentrated. Today, a relatively small group of search, browser, cloud, CDN and security companies exert disproportionate influence over which automated users can access the web, under what conditions, and at what cost: Google, Cloudflare, Amazon, Microsoft, and several others.
Here’s the part that should bother you. Some of the companies that now control automated access to the web benefited enormously from its openness. They are now imposing restrictions that make it harder for new companies to follow comparable paths.
U.S. merger policy has recently become more permissive in some areas, even as major antitrust cases against dominant technology companies continue. That policy uncertainty gives incumbents more room to entrench themselves through product and infrastructure decisions that rarely receive the scrutiny of a formal acquisition.
The hypocrisy
Google operates one of the largest web-crawling systems in history. It built one of the world’s most valuable businesses by crawling public webpages, storing information about them in a proprietary index, and monetizing access to that index through search advertising. The web even developed a voluntary protocol, robots.txt, through which sites communicate access preferences to automated crawlers. Google did not invent that protocol, but it has become its most consequential participant.
Google was allowed to build a multi-trillion-dollar company around a proprietary index created by crawling the public web. Now they are pulling up the ladder.
And there’s a causal chain here. Because Google does not make its underlying search index generally available, competitors and new products must crawl the web themselves or purchase access from another provider. That duplication likely contributes to the growth in automated traffic.
Cloudflare positions itself as the wall that protects websites from scrapers. It also launched its own crawler, and if you spend one minute on any web search and you’ll find users openly discussing how they run mass scraping through Workers (here are Firecrawl’s docs showing how to scrape with Cloudflare Workers). Cloudflare can therefore profit from both controlling automated access and supplying infrastructure used to perform it…selling both the virus and the panacea.
Amazon monitors constantly for price and product intelligence. It watches competitors and adjusts offers. But when an agent shows up to help a user comparison shop or make a purchase on their behalf, Amazon blocks it, as they did by suing Perplexity. Scraping for me, not for thee.
These companies also quietly benefit from bot traffic. More automated load pushes customers into premium tiers. Ad systems can still collect revenue from fraudulent clicks that evade their detection systems. The incumbents enjoy every benefit of a machine-readable web while working to make sure nobody else does.
The ad problem
Part of why this matters is that the web’s default business model is rotten and the alternatives keep getting strangled.
Ads pay for most of the free web and ads are a bad deal. They’re intrusive and annoying. They violate your privacy by design. And the same targeting machinery that sells sugar and cigarettes now delivers microtargeted political propaganda, allowing savvy marketers to swing elections. We’ve known this for years. We tolerate it because nothing else seems to exist at scale…and the terms of service are so long and so hidden that most people do not read them.
Something else does exist. Consented bandwidth sharing. A user agrees to share a slice of unused bandwidth and in exchange the apps they love stay free. It costs the user nothing they actually care about. No interruptions, no surveillance profile, and no manipulation engine.
Look at LG’s smart TV platform. Security researchers scanned the app catalogs and found proxy SDKs in 42.5% of LG’s apps. Amazon already bans this category. Roku reportedly pulled it and the apps that depended on it disappeared. If LG drew the same line tomorrow, nearly half its catalog would need a new way to pay for itself or vanish. Not because the apps are bad, but because many of them may have no equally viable way to pay for themselves. That’s the strongest evidence I can offer that these SDKs aren’t some parasite on the ecosystem. They’re holding it up.
Security companies object that proxy SDKs are dangerous because they’re silent. I disagree with the framing. Without consent, silence is malware. With consent, silence is the entire point. A monetization method you never have to think about, that never interrupts you, that never profiles you, is the best version of paying for software that most users will ever get.
It’s also similar to the web’s original promise: connect and be connected.
Why there are no standards
Why does an industry this consequential still lack a common, enforceable standard?
For most of its life, it was too small to attract scrutiny and too difficult to explain to outsiders. Neither is true anymore. The two largest providers each claim roughly $400 million in annual revenue. Add the rest of the proxy market, scraping infrastructure and alternative data, and the industry is already worth billions. AI is accelerating its growth.
The second reason is uglier: the natural standards-setter would normally be the market leaders, and ours are disqualified.
Both businesses grew out of free VPNs, one of which reportedly had up to 46 million users as part of a botnet. In late 2014, the later began selling business access to users’ devices as exit nodes. The practice became widely known only after the network was used in a 2015 DDoS attack, prompting the company to revise its FAQ. Checking the internet archives shows there wasn’t a clear link between the VPN and the proxy usage until after the attack.
Consent clarified after exposure is not meaningful consent. It is a legal defense.
The company has since rebranded and accumulated audits and certifications. But those credentials do not answer the core questions around user consent and approved client uses.
A company that built its lead through inadequate disclosure should not be trusted to write the industry’s consent rules without proving consent at the user and device level. Its approach to standards remains unilateral “follow our rules or stay out.”
The industry needs something different: a collaborative, transparent standard that applies equally to everyone.
The vacuum gets filled by the wrong people
Here’s what happens when an industry refuses to set its own standards: legislatures do it instead. And legislatures do not understand this technology.
Look at New York’s Stealth Crawler Prohibition Act. The political logic that produced it is easy to reconstruct. AI is causing problems. Scraping is part of AI. Our voters hate AI. Pass the bill. A nonsensical bill that passed both chambers nearly unanimously awaits delivery to the governor.
Here are a few of the problems with this legislation:
The definition of a “Crawler” is dangerously expansive, encompassing any tool that scans or retrieves web data, like price comparison, search indexing, security scanning, and media monitoring. This is well beyond the scope of AI.
The “Covered news source” designation is so broad it could capture simple company blogs, provided they have a basic corrections policy and a small New York readership.
The bill authorizes pre-action identification subpoenas while specifying little in the bill itself about the evidentiary showing a publisher must make before infrastructure providers are ordered to identify a customer.
While titans like Google and OpenAI have the capital to license their way around these walls, the rest of the industry is effectively locked out.
What that logic misses is that crawling is load-bearing for the entire internet. Search runs on it. Price comparison runs on it. Security research runs on it. Accessibility tools run on it. Your favorite app that tells you when flight prices drop runs on it. AI didn’t create some new crawling problem. It poured fuel on a symptom of a much older disease: an industry that never built accepted standards, regulated by people who can’t tell the load-bearing walls from the termites.
These are the real stakes of the standards conversation. It isn’t industry hygiene. It’s whether the rules for a foundational internet technology get written by people who understand it, or by people reacting to a headline.
The front door
Something we’ve been saying for years at Massive was poignantly echoed by Riley from Spur on a recent webinar: without a compliant path forward, we’ll be left with only criminal activity and criminal networks.
Riley's right, and the enforcement record proves it. These IP networks are hydras. In January, Google itself went after IPIDEA, one of the largest criminal proxy networks in the world, with court orders, domain seizures, and Play Store enforcement.
IPIDEA is still operating. If the most powerful company on the internet can't kill one network, enforcement alone was never going to work. As long as demand exists and no legitimate path exists, the supply will be criminal.
The answer is a front door. Build a compliant path that’s real, practical, and open to everyone, and legitimate companies will take it. Nobody climbs through the window when the front door works. The criminal networks lose their customers, which kills them faster than any takedown.
A front door isn’t abstract. It looks like this:
100% consented supply. Every user in the network chose to be there and can leave.
Blocklists that protect users from unwanted or illegal traffic tunneling through their devices, traffic that could create personal liability or degrade their own connection.
KYC on clients, so you know who’s using the network and for what.
Standards bodies developing guidelines for how consent, sourcing, and traffic controls should work.
Platforms writing rules that proxy SDKs must follow to be listed, rather than choosing between ignoring the category and banning it outright.
Where I stand
Full disclosure: I run Massive, a company in this industry, so I have skin in this game. Here’s our record so you can judge the messenger: born out of a consent network, never resold another provider’s IP pool, always maintained a blocklist, one of the only providers that KYCs clients.
We helped antivirus vendors write their rules for how monetization SDKs should behave, and I want to thank Dennis and the team at AppEsteem for their work creating a front door with us for Windows devices.
I’m not writing this to say we’re the standard. I’m writing it because someone in this industry has to start the conversation in public and the companies with the longest track records of breaking the rules can’t be the ones to write them.
We’ll publish our own standards soon as a living document. There will be a form so we can collect feedback and improve them over time. Hold us to them and hold everyone else to them too.

