Scraping

Proxy Selection as a Multi-Armed Bandit Problem

Which proxy configuration should the next request use? Treat it as a bandit problem. In our simulation, adaptive selection came within 2% of perfect hindsight.

James Meadow

James Meadow

October 2, 2026 · 9 min read

Most scraping setups choose a proxy configuration once, at design time, and keep it. Residential with city targeting for this site, ISP proxies for that one, datacenter where it works. The choice is usually made by trying a few options for an afternoon, and it is rarely revisited until something breaks.

That is a decision made under uncertainty, repeated thousands of times a day, with feedback after every request. It is exactly the shape of problem that statisticians call a multi-armed bandit, and there are well-understood algorithms for it. This guide explains the framing, gives a short tested implementation, and shows what it did in a simulation where the best option changed halfway through.

Key takeaways

  • Each proxy configuration you could use for a target is an “arm”; each request is a pull; a good response is a reward. The goal is the lowest cost per good response, not the highest success rate at any price.
  • Thompson sampling handles the trade-off between using the best-known option and testing the others, with a few lines of code and no tuning schedule.
  • Targets change, so old evidence should fade. A discount factor gives the selector a memory of roughly the last thousand outcomes.
  • In our simulation, discounted Thompson sampling came within 2% of a strategy that knew the true rates in advance, while a fixed “safe” choice cost 25% more per good response and sticking with the early winner cost 57% more.
  • Count only genuinely good responses as successes, and limit how much exploration a sensitive target sees.

The framing

Picture a row of slot machines, each paying out with an unknown probability. Every pull teaches you something about one machine, and every pull spent learning about a poor machine is a pull not spent on a good one. That tension between exploiting what you know and exploring what you do not is the multi-armed bandit problem.

For proxy selection, the mapping is direct:

Bandit termIn scraping
ArmA proxy configuration for one target: proxy type, targeting, session setting
PullOne request
RewardA response that contains the data you wanted
Cost of a pullBandwidth, retries and time spent on that request
Non-stationarityThe target changing its defences, or a pool’s reputation shifting

Two features make the scraping version harder than the textbook one. Arms cost different amounts, so the best arm is the one with the lowest cost per success, not the highest success rate. And the rates move: a configuration that works today can be blocked tomorrow, which is the pattern behind subnet contagion and the reason to build a target health score.

The selector

Thompson sampling keeps a probability distribution over each arm’s success rate, a Beta distribution built from its successes and failures. Before each request it draws one plausible rate for every arm and picks the arm that looks cheapest per success under those draws. Arms with little evidence have wide distributions, so they occasionally draw a high rate and get tried. Arms with lots of evidence draw close to their true rate. Exploration fades on its own as the evidence builds.

The version below adds two things: cost per attempt, and a discount that decays old evidence towards the prior so the selector can notice change.

import random


class ProxySelector:
    """Pick a proxy configuration per request with Thompson sampling, minimising cost per successful response.

    Each configuration keeps a Beta(successes, failures) belief about its success rate. A discount below 1
    lets old outcomes fade, so the selector notices when a configuration that used to work starts failing.
    """

    def __init__(self, costs, discount=0.999, prior=(1.0, 1.0), rng=None):
        self.costs = dict(costs)                       # configuration -> cost per attempt
        self.discount = discount
        self.prior = prior
        self.wins = {arm: prior[0] for arm in self.costs}
        self.losses = {arm: prior[1] for arm in self.costs}
        self.rng = rng or random.Random()

    def choose(self):
        """Sample a plausible success rate for each configuration and take the cheapest per expected success."""
        def sampled_cost_per_success(arm):
            rate = self.rng.betavariate(self.wins[arm], self.losses[arm])
            return self.costs[arm] / max(rate, 1e-9)
        return min(self.costs, key=sampled_cost_per_success)

    def update(self, arm, success):
        """Record one outcome. Every belief decays towards the prior first, so recent results count most."""
        for a in self.costs:
            self.wins[a] = self.prior[0] + self.discount * (self.wins[a] - self.prior[0])
            self.losses[a] = self.prior[1] + self.discount * (self.losses[a] - self.prior[1])
        if success:
            self.wins[arm] += 1
        else:
            self.losses[arm] += 1

In use, call choose() before each request, send the request through the chosen configuration, and call update() with whether the response was good. A discount of 0.999 gives an effective memory of about a thousand outcomes; 0.995 about two hundred.

The simulation

To see how the selector behaves when the answer changes, we simulated one target with four configurations. The success rates and costs below are assumptions chosen to make the trade-offs visible, not measurements of any provider or network:

ConfigurationAssumed success rateAssumed cost per attemptCost per success
Datacenter25%0.52.00
ISP80%, falling to 30% after request 5,0000.91.13, then 3.00
Residential, country targeting92%1.31.41
Residential, city targeting95%1.61.68

Costs are relative units covering bandwidth and the overhead of an attempt. For the first 5,000 requests ISP is the cheapest way to get a good response; then the target starts blocking it, and residential with country targeting becomes the best choice. Each strategy made 20,000 requests, and we repeated every strategy 100 times with different random seeds.

StrategyCost per good responseAgainst perfect hindsightGood responsesRequests to adapt after the change
Perfect hindsight (knows the true rates)1.348baseline17,8060
Thompson sampling, discount 0.9991.375+2.0%16,694about 800
Thompson sampling, discount 0.9951.389+3.0%15,704about 1,075
Thompson sampling, no discount1.428+5.9%15,933about 3,500
Epsilon-greedy, 10% exploration1.441+6.9%15,848about 2,800
Thompson sampling, discount 0.981.443+7.0%13,834never settled
Always residential, city targeting1.684+24.9%19,000not applicable
Round robin across all four1.689+25.3%12,731not applicable
Always ISP, the early winner2.118+57.1%8,501not applicable

Figures are averages over 100 runs; for every strategy, the middle 90% of runs spanned less than 2.5% of its average. “Requests to adapt” is the median number of requests after the change until 90% of a 200-request window went to the new best configuration.

What the results say

Fixed choices are expensive in both directions. Always using the most reliable configuration delivered the most good responses but paid 25% more for each one. Always using the configuration that won early was the worst strategy of all once the target changed, at 57% over the baseline.

Learning is not enough without forgetting. Undiscounted Thompson sampling took about 3,500 requests after the change to stop trusting the evidence it had built up for the old winner. A discount of 0.999 cut that to about 800.

Forgetting too fast is its own problem. At a discount of 0.98, the selector remembered only about 50 outcomes, kept re-testing the poor options, and never settled. The useful range depends on traffic: the memory should cover enough requests to tell the configurations apart, and no more.

The objective decides the answer. The bandit minimised cost per good response. If what matters is the most data regardless of cost, or a deadline, the reward has to say so; otherwise the selector will optimise the wrong thing very efficiently.

Using it on real traffic

  • Define success strictly. An HTTP 200 with a block page or empty results is a failure. Feed the selector the same checks you use for the silent failure rate, or it will learn to prefer configurations that fail quietly.
  • Run one selector per target. Success rates differ by site, so a shared selector averages away exactly the differences you want it to learn.
  • Cap exploration on sensitive targets. Every exploratory request through a poor configuration is a request a site may count against you. Remove configurations that are clearly unsuitable rather than letting the selector rediscover it.
  • Keep sessions coherent. Choose the configuration per session, not per request, for anything that depends on state, such as logins or carts, as sticky versus rotating sessions explains.
  • Use real costs. On bandwidth-billed proxies, cost per attempt depends on page weight and retries; the numbers behind cost per clean record are the right inputs.
  • Keep a floor of load balancing. A bandit decides which configuration to use, not how to spread requests within it; that is still the job of rotation, concurrency and retry architecture.

The bottom line

Choosing a proxy configuration once and keeping it is a bet that neither your costs nor the target will change. Treating each request as a small experiment, with Thompson sampling and a sensible memory, turns that bet into a measurement that keeps updating.

In our simulation, that came within 2% of knowing the answer in advance, adapted within about 800 requests when the best option changed, and avoided the 25% to 57% premium of the fixed strategies. The rates and costs were assumptions; the code is short enough to run against your own.

Sources and references

  • Daniel Russo and colleagues, A Tutorial on Thompson Sampling, 2017.
  • Simulation run by Shifter on 2 October 2026 using the code above; success rates and costs are assumptions, not measurements of any network.

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started