Web Crawling for AI Research in 2026: The Infrastructure Nobody Budgets For

Every agent that reads the live web needs a fetch layer. What breaks at scale, why the proxy tier decides your success rate, and how to choose between managed extraction APIs and building it yourself.

Weekly AI tool reviews from a CTO who tests them. No fluff.


Ask an AI assistant to research something on the live web and it either returns a confident summary or it returns nothing useful. The difference rarely comes down to the model. It comes down to whether the fetch layer underneath actually retrieved the pages.

That layer stays invisible until it fails. Then it fails in ways that look like model problems: thin answers, stale citations, an agent that insists a page carries no content when your browser renders it fine.

I run production agents that read the web, some daily and some weekly.

Why Fetching Broke

Most valuable pages render in JavaScript. A plain HTTP GET returns a shell with no content. Any crawler built on a simple request library sees an empty page and reports success, which is worse than reporting failure.

Bot detection improved faster than crawlers did. Cloudflare, Akamai, and PerimeterX now fingerprint TLS handshakes, header ordering, and browser behavior. A default Python user agent gets blocked before the first byte of HTML arrives.

Rate limits tightened. Sites that tolerated a request a second in 2023 now throttle at a tenth of that, and the throttle often arrives as a 200 response containing a challenge page rather than as a 429 you can detect.

Terms of service hardened. Publishers added crawler clauses aimed squarely at AI training and retrieval. The legal surface deserves a real read before you build anything at volume.

The Three Layers

Treat retrieval as three separate concerns. Teams that collapse them into one library discover the seams under load.

Fetch. Getting bytes back from an origin that may not want to give them to you. Proxies, headers, TLS fingerprints, retry policy, and rendering all live here.

Extract. Turning HTML into something a model can use. Boilerplate removal, main-content detection, and conversion to markdown or structured fields.

Orchestrate. Deciding what to fetch, in what order, how often, and what to do when a page fails. Scheduling, deduplication, and freshness policy.

Where the Proxy Tier Decides Everything

Your success rate is mostly a function of your egress IP. Datacenter addresses from the major clouds sit on block lists that most protected sites consult first. The same request from a residential address often passes.

The tiers, in ascending cost and success rate:

Datacenter proxies. Cheapest, fastest, and blocked by anything with real protection. Fine for APIs and unprotected sites.

Residential proxies. Traffic routed through consumer connections. Materially higher success rates against bot detection, and priced by bandwidth rather than by request, which changes how you think about page weight.

Mobile proxies. Carrier-grade addresses shared by many real users, which makes blocking them expensive for the site. Highest success rate and highest cost.

MarsProxies covers the residential and mobile tiers, and the practical advantage over rolling your own pool comes down to IP rotation and pool health rather than raw price. A proxy pool degrades as addresses get flagged, and maintaining freshness is the actual work.

Budget by bandwidth, not by page. A modern article with images and tracking can weigh several megabytes. Requesting rendered pages when you need text multiplies your bill by an order of magnitude for no gain.

Managed Extraction Versus Building It

Managed APIs hand you a URL and return clean markdown. Firecrawl, Jina Reader, and the extraction endpoints inside several search APIs occupy this tier. You pay per page and inherit their proxy pool, their rendering, and their blocking problems.

Build it yourself with Playwright or a headless browser, your own proxy contract, and your own extraction. You gain control over rendering behavior, cost structure, and what happens on failure.

The honest threshold sits around volume and specificity. Below a few thousand pages a month against ordinary sites, managed APIs cost less than the engineering time to maintain an alternative. Above that, or when you need pages that managed providers already fail on, the economics invert.

One pattern beats both in most agent workloads. Try the cheap path first, escalate on failure. Plain fetch, then rendered fetch, then rendered fetch through a residential proxy. Most pages resolve at step one, and you pay the expensive path only where you must.

The Failure Modes Worth Instrumenting

Silent empty extraction. The fetch returns 200, the extractor returns 40 characters of navigation text, and the agent summarizes nothing. Assert on a minimum content length and treat a violation as a failure, because nothing else will tell you.

Challenge pages returning 200. Bot walls frequently answer with a success code and a body containing a CAPTCHA. Check for challenge markers rather than trusting the status.

Stale cache serving old facts. Caching cuts cost and breaks freshness. Set the TTL against how fast the content actually changes rather than against a default.

Retry storms. A blocked domain that retries aggressively burns proxy bandwidth and deepens the block. Retry only what a retry can clear, and back off per domain rather than globally.

Cost with no attribution. Bandwidth-priced proxies produce a single monthly number unless you tag requests by agent, job, and domain. Tag from day one.

Read the robots file and honor it by default. A crawler that ignores robots invites both a block and a conversation you do not want.

Rate limit yourself below what the site tolerates. The polite number costs you almost nothing and removes most of the reason anyone would look at your traffic.

Identify your crawler honestly in the user agent when you operate at any scale, with a contact address. Site operators who can reach you usually ask before they block.

Treat terms of service as a real constraint on commercial use, redistribution, and training. The retrieval question and the reuse question carry different answers.

What I Would Build First

Start with a managed extraction API and instrument it. Log success rate, content length, and cost per domain for a month.

That data tells you whether you have a crawling problem at all. Most teams discover their failures concentrate on a handful of domains, which turns an infrastructure project into a narrow escalation path for those specific sites.

Add a residential proxy tier only where the data says you need it. Buying the expensive tier for every request is the most common way teams overspend on retrieval.

Freshness, Which Decides Cost More Than Volume

Most retrieval budgets are wasted on re-fetching pages that did not change. A crawler that refreshes everything nightly spends the same money whether the corpus moved or not.

Set the interval against observed change rate rather than against a policy. Track a content hash per URL and measure how often it actually differs. Documentation changes monthly, pricing pages change quarterly, and news changes hourly. Treating them identically overspends on two of the three.

Use conditional requests where the origin honors them. ETag and If-Modified-Since turn an unchanged page into a cheap 304 rather than a full transfer, which matters directly when your proxy bills by bandwidth.

Separate discovery from refresh. Finding new URLs and updating known ones are different jobs with different cadences, and merging them means you either discover too slowly or refresh too often.

The Takeaway

Retrieval quality sets the ceiling on every agent that reads the live web, and it fails quietly enough that teams blame the model for months. The fix runs through instrumentation before infrastructure: measure success rate and content length per domain, escalate the cheap path to the expensive one only where it fails, and price your proxy tier against bandwidth rather than against requests.

Share this article

Get more like this.

Weekly AI tool reviews and practical implementation guides, delivered straight to your inbox.

No spam. Unsubscribe anytime.