Raspbytes

web scraping

Retry Strategies for Large-Scale Web Data Collection

Learn how to design scalable retry strategies for reliable web data collection using backoff, jitter, retry budgets, and failure classification.

API-first

Simple integration

Observable

Clear run history

Flexible

On demand or scheduled

web scraping14 min read

At small scale, retry logic looks simple.

A request fails. Wait a moment. Try again.

At large scale, that approach can become one of the fastest ways to make a web data collection system less reliable.

When millions of requests move through a collection pipeline, failures are inevitable. Connections time out. Target servers return errors. Rate limits appear. Proxies fail. Pages respond slowly. DNS resolution occasionally breaks. A request that normally completes in two seconds might suddenly take twenty.

The important question is therefore not whether failures happen, but what the system should do after they happen.

A well-designed retry strategy improves collection completeness without creating unnecessary traffic, consuming excessive proxy bandwidth, or allowing failed requests to overwhelm healthy workloads.

That requires treating retries as part of the system architecture rather than as a simple loop around an HTTP request.

Not Every Failure Deserves a Retry

The first mistake many systems make is assuming every unsuccessful request should be retried.

Some failures are temporary.

Others are effectively permanent.

Consider a few examples:

  • A connection timeout may disappear on the next attempt.

  • An HTTP 429 Too Many Requests response may succeed after waiting.

  • An HTTP 503 Service Unavailable response may indicate temporary server pressure.

  • An HTTP 404 Not Found response will usually remain a 404.

  • A malformed URL is unlikely to become valid after waiting five seconds.

Retrying permanent failures wastes resources and can significantly increase the amount of traffic generated by a large collection operation.

The first component of a retry system should therefore be failure classification.

Instead of thinking in terms of simply "success" and "failure," a collector can classify outcomes such as:

Transient failures are conditions that may disappear naturally, including network timeouts, connection resets and temporary upstream errors.

Rate-limit failures indicate that the target is asking the client to reduce its request rate.

Access failures can indicate that the current request context—IP, session, headers or browser state—is unsuitable.

Permanent failures include conditions where repeating the same request is unlikely to change the outcome.

Application failures occur when a request technically succeeds but the returned content is unusable. A page might return HTTP 200 while containing an error page, challenge page or incomplete response.

The classification determines what happens next.

That decision is considerably more important than simply choosing how many retry attempts should be allowed.

The Problem With Immediate Retries

Imagine a collector sending 10,000 requests per second.

A target service experiences a temporary problem and 30% of requests begin failing.

If every failed request is immediately retried, the system may suddenly generate another 3,000 requests per second.

If those retries also fail and are retried again, traffic increases further.

The retry mechanism has now amplified the original failure.

This phenomenon is sometimes described as a retry storm. It occurs when failures trigger additional traffic at exactly the moment a downstream system is already struggling.

Large collection systems therefore need to control not only whether requests are retried but when those retries occur.

Exponential Backoff

One of the most common techniques is exponential backoff.

Instead of retrying at a constant interval, the delay grows after each failed attempt.

For example:

Attempt 1 → immediate request
Attempt 2 → wait 1 second
Attempt 3 → wait 2 seconds
Attempt 4 → wait 4 seconds
Attempt 5 → wait 8 seconds

A simplified model might be:

delay = base_delay × 2^attempt

The exact numbers are less important than the behaviour.

Repeated failures cause the system to progressively reduce pressure on the target.

This works particularly well for temporary infrastructure problems because it gives the downstream service time to recover rather than repeatedly hitting it with identical requests.

But exponential backoff creates another problem when thousands of requests fail simultaneously.

They may all retry at roughly the same time.

That is where jitter becomes important.

Why Jitter Matters

Suppose 50,000 requests receive errors within the same second.

If every request uses the same exponential schedule, a large portion of them might retry one second later.

Then again two seconds later.

Then four seconds later.

The traffic pattern becomes synchronized.

Instead of solving the retry storm, the system has simply converted it into periodic bursts.

Jitter introduces randomness into retry timing.

For example, rather than every request waiting exactly four seconds, requests might wait somewhere between two and six seconds.

The retry load becomes distributed across time.

There are several possible jitter algorithms, including full jitter, equal jitter, and decorrelated jitter. The precise implementation matters less than the architectural principle:

Large groups of failed requests should not automatically become large groups of synchronized retries.

At sufficient scale, a small amount of randomness can substantially improve system stability.

Retry Budgets Matter More Than Retry Counts

A common configuration might specify:

max_retries = 3

That is useful, but it does not fully control retry behaviour.

Imagine processing ten million URLs.

If 20% initially fail and each receives three additional attempts, the system could generate millions of extra requests.

Those retries consume:

  • network capacity

  • proxy bandwidth

  • connection pool capacity

  • CPU and memory

  • browser capacity for browser-based collection

  • target request allowances

  • time that could have been spent processing healthy work

This is why large-scale systems often benefit from the concept of a retry budget.

Instead of considering each URL independently, the collector also limits how much total capacity can be consumed by retries.

For example, the system might allow retry traffic to represent only a certain percentage of overall request throughput.

If failures suddenly spike, retries are slowed or queued rather than being allowed to consume the entire system.

The objective is simple:

Retries should improve data completeness without threatening the throughput of healthy requests.

Respect Retry-After

Some HTTP responses provide explicit information about when another request should be attempted.

The most important example is the Retry-After header, commonly associated with responses such as 429 and 503.

If a target tells the collector to wait before sending another request, blindly applying a generic retry schedule may be counterproductive.

A more sophisticated policy can therefore prioritize server-provided instructions:

if Retry-After exists:
    use Retry-After
else:
    use calculated backoff

This also highlights a broader principle.

Retry logic should respond to information from the target rather than treating every failure identically.

Retries and Proxy Rotation

Web data collection introduces another dimension that ordinary API clients do not always encounter: proxy infrastructure.

Suppose a request fails while using proxy A.

Should the retry use proxy A again?

Or should it switch to proxy B?

The answer depends on the failure.

A connection failure associated with the proxy itself may justify selecting another endpoint.

A target-side rate limit may indicate that changing the IP or slowing traffic is appropriate.

A server-side 500 error, however, may have nothing to do with the proxy.

Automatically rotating IP addresses after every unsuccessful request can therefore be wasteful.

A better approach is failure-aware proxy rotation.

The retry system considers the likely source of the failure before deciding whether network identity should change.

For example:

Connection failure
→ retry using another healthy proxy

Target 429
→ reduce request rate and potentially rotate according to policy

Target 503
→ backoff before retrying

404
→ normally do not retry

This is particularly important when residential proxies are involved because unnecessary retries translate directly into additional bandwidth consumption.

Proxy services such as residential or datacenter proxy networks can provide the network layer for distributed collection, but application-level retry decisions still need to be made by the collection architecture itself.

Separate Retry Queues From Primary Work

Another useful architectural pattern is separating new work from retry work.

Consider a crawler with a queue containing:

URL A
URL B
URL C
URL D

If URL A repeatedly fails and immediately returns to the front of the same queue, problematic URLs can consume capacity that should be processing new requests.

At large scale, thousands of problematic URLs can produce the same effect.

A cleaner architecture might look like:

Primary Queue
     │
     ▼
Fetch Workers
     │
     ├── Success ──→ Processing
     │
     └── Retryable Failure
                 │
                 ▼
             Retry Queue
                 │
            delayed retry
                 │
                 ▼
             Fetch Workers

The retry queue can support delayed delivery based on the calculated backoff time.

This separation also makes it easier to enforce retry budgets, prioritize fresh work, and observe retry behaviour independently.

Think Beyond Individual URLs

A major improvement occurs when retry systems stop treating every request as an isolated event.

Suppose 5,000 URLs from the same domain suddenly start returning 503.

Individually, each request appears to need a retry.

Collectively, however, the system has learned something more useful:

the domain itself may currently be unhealthy.

Continuing to retry each URL independently wastes requests.

Large collection systems can therefore maintain health information at multiple levels:

Request
   ↓
Proxy
   ↓
Domain
   ↓
Collection Job

If failure rates for a domain exceed a threshold, the scheduler might temporarily reduce concurrency or pause new requests to that domain.

This is similar to the circuit-breaker pattern used in distributed systems.

Instead of repeatedly discovering that a service is unavailable, the collector temporarily stops sending requests and periodically tests whether conditions have recovered.

Adaptive Concurrency Can Prevent Failures Before Retries Are Needed

Retries are reactive.

Another approach is to reduce the number of failures that occur in the first place.

Imagine a domain handling 50 concurrent requests successfully but beginning to return 429 responses when concurrency reaches 200.

A collector that ignores this signal might continue operating at 200 concurrent requests while constantly retrying failures.

That is inefficient.

A more adaptive system can reduce concurrency when rate limits or latency increase.

For example:

Healthy responses
      ↓
gradually increase concurrency

Increasing latency / 429 responses
      ↓
reduce concurrency

The goal is to discover a sustainable operating range for each target.

Retry logic then becomes a safety mechanism rather than the primary method of controlling traffic.

Browser Retries Are More Expensive

The economics change further when data collection involves full browser sessions.

Restarting an HTTP request is relatively cheap.

Restarting a browser workflow may not be.

A browser retry could involve:

Launch session
→ load page
→ execute JavaScript
→ wait for network activity
→ interact with page
→ extract data

Repeating that entire sequence consumes substantially more compute, bandwidth, and time.

Browser-based collectors should therefore consider retry granularity.

If an extraction step fails, does the entire browser session need to restart?

If a navigation times out, can the current session attempt another navigation?

If authentication state has already been established, should it be preserved?

Services that expose remote browser infrastructure through Playwright or Puppeteer can simplify browser provisioning, but collection logic should still avoid unnecessarily recreating expensive sessions.

The more expensive the unit of work, the more important precise retry decisions become.

Idempotency Becomes Important

Retries also create a data problem.

Suppose a request succeeds, but the collector fails before recording that success.

The job scheduler may assume the request failed and run it again.

Now the same record may be extracted twice.

For purely read-oriented HTTP requests, this does not usually affect the target, but it can affect downstream datasets.

Large pipelines therefore need some concept of idempotency or deduplication.

A record might be uniquely identified using combinations such as:

job_id + URL

or:

source + external_record_id

The exact mechanism depends on the workload, but the objective is the same:

Retrying collection should not accidentally create duplicate data.

This becomes especially important when downstream processing triggers additional actions such as enrichment, notifications, or database updates.

Dead-Letter Queues Prevent Infinite Work

Eventually, the system must stop trying.

A URL that has failed repeatedly should not circulate through retry queues forever.

After the retry policy is exhausted, the request can move into a dead-letter queue (DLQ).

A DLQ preserves failed work for later inspection without allowing it to interfere with active collection.

Entries might contain information such as:

URL
job ID
attempt count
last HTTP status
failure category
proxy metadata
timestamps
last error

The dead-letter queue becomes extremely useful operationally.

If thousands of requests suddenly appear there with the same error, the problem may not be those URLs individually. It could indicate a parser change, proxy problem, authentication issue, or target-side change.

Observability Is Part of Retry Design

Retry systems are difficult to improve if their behaviour cannot be measured.

Simply tracking the number of failed requests is not enough.

Useful metrics include:

initial success rate
retry success rate
attempts per successful request
retry traffic percentage
failure rate by domain
failure rate by status code
failure rate by proxy pool
time spent waiting for retries
dead-letter queue volume

One particularly valuable measurement is recovery by attempt number.

Suppose the data shows:

Initial attempt     91.0% success
Retry 1              6.5% recovered
Retry 2              1.2% recovered
Retry 3              0.2% recovered
Retry 4              0.02% recovered

That tells you something important.

The fourth retry may be generating significant traffic while recovering almost no additional data.

Retry limits should therefore be informed by observed recovery rates rather than chosen arbitrarily.

Retry Policies Should Be Target-Aware

There is rarely one retry configuration that works equally well across every website.

Different targets behave differently.

One domain may tolerate significant concurrency but occasionally produce network timeouts.

Another may aggressively return 429 responses.

Another may respond slowly during particular periods.

A mature collector can therefore maintain policies by domain or workload.

Conceptually:

Domain A
max attempts: 3
base backoff: 1s
concurrency: 100

Domain B
max attempts: 5
base backoff: 5s
concurrency: 20

Domain C
max attempts: 2
base backoff: 2s
concurrency: 50

These values do not necessarily need to be manually configured forever. They can increasingly be derived from historical performance.

The system gradually learns which targets recover quickly, which require longer cooldown periods, and which failures rarely recover at all.

A Practical Retry Architecture

Putting these ideas together, a collection pipeline might look something like this:

                   ┌───────────────┐
                   │  Primary Queue│
                   └───────┬───────┘
                           │
                           ▼
                    ┌─────────────┐
                    │ Fetch Layer │
                    └──────┬──────┘
                           │
             ┌─────────────┴─────────────┐
             │                           │
          Success                     Failure
             │                           │
             ▼                           ▼
        Processing              Failure Classifier
                                         │
                         ┌───────────────┼──────────────┐
                         │               │              │
                      Retryable       Permanent      Rate Limited
                         │               │              │
                         ▼               ▼              ▼
                  Backoff + Jitter      DLQ       Domain Throttle
                         │                              │
                         └──────────────┬───────────────┘
                                        ▼
                                   Retry Queue

The important point is that retries are not just an HTTP-client configuration.

They interact with scheduling, proxy management, concurrency control, data integrity, and observability.

The Goal Is Not Maximum Retry Success

It is tempting to measure a retry system purely by how many failed requests it eventually recovers.

But maximizing recovery at any cost is rarely the right objective.

Suppose increasing the retry limit from three attempts to eight improves dataset completeness from 99.5% to 99.7%, while increasing total traffic by 25%.

Whether that trade-off is worthwhile depends on the workload.

For some datasets, the additional completeness may be essential.

For others, it may be economically irrational.

A better way to think about retry performance is:

Useful data recovered
────────────────────────────
Cost of additional attempts

That cost includes infrastructure, proxy bandwidth, processing time and pressure placed on the target.

The optimal retry policy therefore depends on the value of the missing data as much as the technical characteristics of the failure.

Retries Are a Scheduling Problem

At small scale, retries can reasonably be implemented as a few lines around an HTTP client.

At large scale, they become something much broader.

They determine how capacity is distributed between new work and failed work. They influence proxy consumption. They interact with rate limits. They affect browser infrastructure costs. They determine how quickly a system responds when an entire domain begins failing.

The most resilient collection architectures therefore treat retries as a scheduling and resource-management problem, not simply an error-handling mechanism.

Classify failures before retrying them. Introduce exponential backoff and jitter. Limit retry traffic with budgets. React to domain-level signals rather than individual URLs alone. Separate retries from primary workloads. Measure how much value each additional attempt actually produces.

The objective is not to retry everything until it works.

It is to recover the requests worth recovering while keeping the rest of the collection system stable.

Retry Strategies for Large-Scale Web Data Collection | Raspbytes