Demand & Paid Media · established evidence
Budget Allocation Under Uncertainty: What Multi-Armed Bandits Teach About the Explore/Exploit Problem in Paid Media
When a business has a few thousand dollars a month to split across Google, Meta, and a local surface, it faces a decision most reporting cannot help with: how much to put behind the channel that looks best today, and how much to keep testing the ones it has barely measured. That is not a marketing preference. It is a named problem in decision theory, the explore/exploit tradeoff, and its formal model is the multi-armed bandit. Bandit algorithms, and Thompson sampling in particular, give a principled way to allocate a constrained budget across options whose returns are still uncertain, with proven mathematical guarantees, which is exactly the position a small advertiser is in long before it has the data to run a full marketing mix model. This article shows where the peer-reviewed theory maps cleanly onto paid media and where the application is still emerging, and draws the one operating conclusion it actually supports.
A small budget is a bet placed before the data arrives
A business spending a few thousand dollars a month across two or three paid surfaces is in a specific and underappreciated bind. It does not have enough conversions per channel, per month, to say with any confidence which channel is genuinely earning its keep. A marketing mix model, the instrument built to answer that question, needs years of variation in spend across many channels before it can separate their contributions. The money, meanwhile, has to be split now, this week, this month, on evidence that is thin by construction.
The instinct is to treat this as a data problem to be waited out, or to hand the split to the platform and hope its automation reads the tea leaves correctly. Neither is quite right. Allocating a fixed budget across options whose returns are still uncertain is not a gap in your reporting. It is a well-studied problem with a formal name and a mathematics of its own, and knowing that mathematics changes what a sensible budget process looks like.
The explore/exploit problem, stated precisely
Picture a row of slot machines, each with an unknown and different payout. Every time you pull a lever you spend a coin and receive a reward drawn from that machine's hidden distribution. You want the most reward you can get before the coins run out. This is the multi-armed bandit, and the tension at its heart is exact: with each pull you can exploit the machine that has looked best so far, or explore a machine you have sampled less, in case it is quietly the better one.
The two errors are symmetric and both costly. Exploit too early and you commit the budget to an arm that only looked good on a handful of noisy pulls. Explore too long and you spend the budget learning which arm was best instead of earning from it. The explore/exploit problem is the formal statement of that tradeoff, and it maps onto paid media without much translation: each channel, campaign, keyword, or creative is an arm; each dollar is a pull; the reward is the return that dollar produced. A budget split is a policy for which arms to pull, made under exactly the uncertainty the bandit model describes.
Thompson sampling: allocate in proportion to the probability of being best
Among the strategies for this problem, one is both old and unusually well-behaved: Thompson sampling. Rather than track a single best-guess return for each arm, it keeps a probability distribution over each arm's true return, and updates that distribution as evidence comes in. The wider the distribution, the less you know about that arm.
The decision rule is elegant. To place the next dollar, draw one random sample from each arm's distribution, and back the arm whose sample came out highest. An arm that is probably the best gets chosen most of the time. An arm that is merely uncertain, with a wide distribution, still gets chosen occasionally, because now and then its sample lands high. Exploration is not bolted on as a fixed rule; it falls out of the uncertainty itself, and it shrinks automatically as more data narrows the distributions. That property, that the algorithm explores exactly as much as its own ignorance warrants and no more, is why Thompson sampling has become a reference method for the explore/exploit problem rather than a curiosity.
Why a budget constraint changes the math
The textbook bandit makes two assumptions that a real ad budget violates. It assumes every pull costs the same, and it assumes you care about total reward over a fixed number of pulls. Advertising is not like that. A click on one channel costs more than a click on another, so pulls have different prices, and you do not have a fixed number of pulls; you have a fixed budget, which is a very different constraint.
This is the setting Xia and colleagues formalized in 2015 as the budgeted multi-armed bandit: each arm has both an unknown reward and an unknown cost, and the objective is to earn the most total reward before a fixed budget is exhausted. They developed Thompson-sampling algorithms for precisely this case and proved regret bounds for them, meaning the shortfall against a hypothetical perfectly informed allocator is provably limited. That result matters here because it is the version of the problem that actually resembles a media budget: not "how many times should I pull," but "given that pulls cost different amounts and I have a finite purse, where should the next dollar go."
From one budget to many channels at once
A real allocator faces a further wrinkle that the classic model sets aside. You are rarely choosing one arm at a time. You are dividing a budget across several channels simultaneously, and the choices interact: what you put behind one surface changes what is left for the others, and a slate of channels funded together is not the same as one lever pulled in isolation.
A 2026 preprint extended the budgeted-bandit method to exactly this multichannel case, casting adaptive advertising budget allocation as a combinatorial-bandit problem, where the algorithm chooses a whole combination of funded arms under a shared budget rather than a single arm. This is the frontier of the mapping: the core budgeted-bandit theory is established and peer-reviewed, while the tailored multichannel-advertising framing is recent and still emerging. The direction is sound; the applied claims at small-advertiser scale are newer than the mathematics they rest on.
Why this is not a marketing mix model
It is worth being precise about what a bandit is and is not replacing. The enterprise instrument for "which channels are actually driving my revenue" is the marketing mix model. Google's open-sourced Meridian tool, generally available since January 2025, is a modern example: it uses Bayesian causal-inference methods, built on published prior-calibration research, to estimate each channel's contribution from the historical record.
The catch is in that last phrase. A marketing mix model reads channel contributions out of years of variation in spend across many channels. A business spending a few thousand dollars a month across two or three surfaces does not have that record and will not have it for a long time. Bandit allocation is not a downgraded MMM; it is the tool for the regime below the one an MMM needs. It answers a narrower, more urgent question, "where should the next dollar go, given only what I have learned so far," and it keeps answering it as evidence accumulates. The two live at different scales of data, and a small advertiser is squarely in bandit territory, not MMM territory.
The reward signal is where the accuracy lives
A bandit is only as good as the reward you feed it. Point it at the wrong number and it will chase that number efficiently, straight into the wrong channels. This is the part of the problem that no algorithm solves for you, and it is where most allocation quietly goes wrong.
The default reward on offer is the platform's last-click return, and a decade of field experiments says that number overstates the causal truth.
Platform-reported return is biased upward
Lewis, Rao, and Reiley showed across three controlled experiments that observational estimates of ad effectiveness are systematically inflated by activity bias: people who happen to be active online are more likely to search, click, and buy whether or not they saw an ad, and the correlation gets miscredited to the advertising. Blake, Nosko, and Tadelis, in one of the largest randomized advertising experiments published, found that paid search ads on eBay's own branded keywords produced no measurable short-term incremental benefit at all. The lesson is not a fixed percentage to subtract; it is that the reward a bandit reads off a dashboard can diverge sharply from the reward the dollar actually caused.
Measuring the real reward is now affordable
The fix is to feed the allocator a causal reward, and that is more accessible than it used to be. Vaver and Koehler's geo-experiment method reads true ad lift by randomizing geographic regions into treatment and control, with no individual-level tracking required, and is framed explicitly to inform budgeting and bidding decisions. The "ghost ads" method of Johnson, Lewis, and Nubbemeyer made counterfactual measurement cheaper still, recording the impressions a control group would have seen; in their retargeting test it measured a lift of 17.2% in site visits and 10.5% in purchases while costing a fraction of a traditional public-service-announcement holdout. Industry analysis, less rigorously, reports that incremental return often runs well below the platform-reported figure, a directional claim consistent with the peer-reviewed work but not itself independently audited. The operating point stands: the allocation loop has to measure the arm's real return, or it steers toward an artifact.
The arms keep moving, because the auction does
There is one more reason a budget cannot be set once and left. A bandit's comfort zone is arms whose reward distributions hold still. Paid-media arms do not. The return on a channel is settled inside an auction whose price depends on your competitors, and the generalized second-price auction behind sponsored search was proven, by Edelman, Ostrovsky, and Schwarz and independently by Varian, to have no dominant-strategy equilibrium. There is no fixed honest bid, and the cost of a click moves as rivals enter, raise budgets, or change strategy.
That makes the arms non-stationary: an arm that paid well last month can pay differently this month for reasons that have nothing to do with your account. A one-time adjustment, however clever, will lag a field that keeps shifting under it. The right posture is a loop that keeps sampling precisely because the world it samples keeps changing. In bandit terms, exploration never fully switches off, because the thing being explored is a moving target.
What the theory says about set-and-forget
Assemble the pieces and the operating implication is hard to avoid. The budget split is a decision under genuine uncertainty. The principled method for it is an explore/exploit loop, formalized for finite budgets and different click costs, with proven guarantees. That loop is only as honest as the reward it is fed, and the trustworthy reward is causal, not last-click. And the arms it allocates across are non-stationary, because the auctions underneath them keep moving. Every one of those facts is documented, and every one of them argues against setting a budget once and walking away.
Platform auto-tuning is itself a kind of bandit. But it is a black-box one, tuning itself against the platform-reported outcomes the field-experiment literature shows can mislead, toward the platform's objective, which is not guaranteed to be your incremental profit. The alternative is not to abandon automation but to run a transparent loop around it: measure the real return, reallocate as evidence arrives, and keep exploring because the field will not hold still. Active, measured budget management is not a stylistic choice over set-and-forget. It is what the explore/exploit problem tells you the work actually is.
The evidence
Key findings, with their sources
-
The budgeted multi-armed bandit formalizes allocating a fixed budget across arms with unknown reward and unknown cost; Thompson-sampling algorithms for this setting were developed with proven regret bounds.
established Xia, Y., et al., "Thompson Sampling for Budgeted Multi-armed Bandits", IJCAI 2015, https://www.ijcai.org/Proceedings/15/Papers/556.pdf.
-
The budgeted-bandit method has been extended to splitting a shared budget across several channels at once, framed as adaptive multichannel advertising budget allocation via combinatorial bandits.
emerging "Adaptive Budget Optimization for Multichannel Advertising Using Combinatorial Bandits", arXiv:2502.02920, 2026.
-
Marketing mix modeling estimates channel contributions from years of spend variation using Bayesian causal-inference methods, a different data regime from the small-budget bandit problem; Google's Meridian MMM has been generally available since January 2025.
established Google, "Meridian is now available to everyone", 2025; methodology in Zhang et al., "Marketing Mix Model Calibration With Bayesian Priors", Google, 2024.
-
A bandit is only as good as its reward signal, and platform-reported value can diverge sharply from causal value: paid search ads on eBay's branded keywords produced no measurable short-term incremental benefit in a large randomized field experiment.
established Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 155-174, 2015 (single-firm study).
-
Observational estimates of ad effectiveness are systematically inflated by activity bias, the pre-existing correlation across a user's online behaviors, shown across three controlled experiments.
established Lewis, R. A., Rao, J. M. & Reiley, D. H., "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW 2011.
-
Counterfactual "ghost ads" measured a lift of 17.2% in site visits and 10.5% in purchases in a retargeting test, at a fraction of the cost of a traditional public-service-announcement holdout, making a causal reward signal affordable.
established Johnson, G. A., Lewis, R. A. & Nubbemeyer, E. I., "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 867-884, 2017.
-
The channels a bandit allocates across are non-stationary because the auction underneath is: the generalized second-price auction has no dominant-strategy equilibrium, so a click's cost and return shift as competitors move.
established Edelman, B., Ostrovsky, M. & Schwarz, M., American Economic Review, 97(1), 242-259, 2007; Varian, H. R., "Position Auctions", International Journal of Industrial Organization, 25(6), 1163-1178, 2007.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| Established | The explore/exploit tradeoff and the budgeted multi-armed bandit are formal, well-studied problems; Thompson sampling solves them with proven regret bounds. Marketing mix modeling is the established instrument at enterprise data scale. | Xia et al. 2015 (IJCAI); Google Meridian / Zhang et al. 2024. |
| Established (measurement of the reward) | Platform-reported return overstates causal return; activity bias and the eBay branded-keyword null are documented; geo-experiments and ghost ads make causal measurement affordable. | Lewis, Rao & Reiley 2011 (WWW); Blake, Nosko & Tadelis 2015 (Econometrica); Vaver & Koehler 2011; Johnson, Lewis & Nubbemeyer 2017 (JMR). |
| Emerging | Applying combinatorial budgeted bandits specifically to small-advertiser multichannel budget splits; the tailored SMB-scale framing is newer than the core theory. | arXiv:2502.02920 (2026), extending the peer-reviewed budgeted-bandit line. |
| Contested / industry-reported | Specific figures for how far incremental return falls below platform-reported return; directionally consistent with the peer-reviewed work but not independently audited. | Practitioner and vendor analysis, 2025-2026; treat as directional pending primary data. |
Reference
Glossary
- Multi-armed bandit
- A decision model in which you repeatedly choose among options ("arms") with unknown payouts, trying to earn the most reward over time while still learning which arm is best. The formal home of the explore/exploit problem.
- Explore/exploit tradeoff
- The tension between exploiting the option that has looked best so far and exploring less-tested options that might be better. Exploit too soon and you lock into a mistake; explore too long and you spend the budget learning instead of earning.
- Thompson sampling
- A bandit strategy that keeps a probability distribution over each arm's true return and picks the next arm by drawing one sample from each and backing the highest. Exploration emerges from uncertainty and shrinks automatically as data accumulates.
- Budgeted multi-armed bandit
- A bandit variant where each pull has a cost and you have a fixed budget rather than a fixed number of pulls, matching a real ad budget where clicks cost different amounts and the purse is finite.
- Regret bound
- A proven limit on how far an algorithm's total reward can fall short of a hypothetical perfectly informed allocator. A regret bound is what makes a bandit method a guarantee rather than a heuristic.
- Non-stationary arm
- An option whose reward distribution changes over time. Paid-media channels are non-stationary because their returns are set inside competitive auctions that shift as rivals move.
- Marketing mix model (MMM)
- A statistical model that estimates each channel's contribution to revenue from years of variation in spend across channels. Built for enterprise data scale, not for a small monthly budget with thin per-channel data.
Straight answers
Frequently asked questions
What is the explore/exploit problem in plain terms?
It is the tension every budget faces: put more behind the channel that looks best on the data you have (exploit), or keep spending to test channels you have barely measured in case one is better (explore). Commit too early and you lock into a channel that only looked good on thin data; keep testing forever and you spend the budget learning instead of earning. The multi-armed bandit is the formal model of that tradeoff.
How does Thompson sampling decide where the budget goes?
It keeps a probability distribution over each channel's true return and updates it as results come in. To place the next dollar, it draws one sample from each channel's distribution and backs the highest. A channel that is probably best gets funded most of the time; an uncertain channel still gets funded occasionally, because its wide distribution sometimes samples high. Exploration falls out of the uncertainty and shrinks as the data sharpens.
Why not just build a marketing mix model instead?
A marketing mix model needs years of spend variation across many channels to separate their contributions. A business spending a few thousand dollars a month does not have that history and will not for a long time. Bandit allocation is the tool for that regime: it answers "where should the next dollar go given what I have learned so far" and keeps answering it as evidence accumulates. It does not replace an MMM; it works at the data scale below one.
Does the platform's automatic budget tuning already do this?
Platform auto-tuning is itself a kind of bandit, but a black-box one. It tunes itself against the platform's own reported outcomes, which field experiments show can overstate real causal return, toward the platform's objective, which is not guaranteed to be your incremental profit. The better method is not to abandon it but to run a transparent loop around it: feed the allocation a causal reward and reallocate as the evidence, and the auction, move.
What is the single most important input to a bandit budget?
The reward signal. A bandit chases whatever number you feed it, so if the reward is the platform's last-click return, the allocator will steer toward a measurement artifact. Activity bias and the eBay branded-keyword null result both show platform-reported value diverging from causal value. Feeding the loop a causal reward, read through holdout or geo-based tests where budget allows, is what keeps the allocator measuring reality instead of efficiently chasing a flattering number.
Why does any of this argue against a set-and-forget budget?
Because four documented facts stack against it: the split is a decision under uncertainty, the principled method is an explore/exploit loop, that loop needs a causal reward rather than a last-click one, and the channels are non-stationary because their auctions keep moving. A set-and-forget budget treats a moving problem as a fixed one. Active, measured reallocation is what the theory says the work is.
Provenance
Sources
- Xia, Y., et al., "Thompson Sampling for Budgeted Multi-armed Bandits", IJCAI 2015 (established)ijcai.org
- "Adaptive Budget Optimization for Multichannel Advertising Using Combinatorial Bandits", arXiv:2502.02920, 2026 (emerging application, extending the established budgeted-bandit line)arxiv.org
- Google, "Meridian is now available to everyone", 2025; Zhang et al., "Marketing Mix Model Calibration With Bayesian Priors", Google, 2024 (established)
- Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 155-174, 2015 (established, single-firm)doi.org
- Lewis, R. A., Rao, J. M. & Reiley, D. H., "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW 2011 (established)
- Vaver, J. & Koehler, J., "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011 (established)research.google
- Johnson, G. A., Lewis, R. A. & Nubbemeyer, E. I., "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 867-884, 2017 (established)doi.org
- Edelman, B., Ostrovsky, M. & Schwarz, M., "Internet Advertising and the Generalized Second-Price Auction", American Economic Review, 97(1), 242-259, 2007; Varian, H. R., "Position Auctions", International Journal of Industrial Organization, 25(6), 1163-1178, 2007 (established)
- Industry practitioner analysis on incremental versus platform-reported return, 2025-2026 (contested / directional, not independently audited)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.