HOLDOUT · Methodology
How HOLDOUT measures
Effective 20 July 2026
01 · What this document is
This is the method behind the figures on this site, written so a reader can check them rather than trust them.
One illustrative specimen store carries every worked example. It is a specimen, not a customer. HOLDOUT is pre-launch — public launch is Q4 2026 — so nothing here is a customer result or a claim about what your store will earn. Where a specimen figure was chosen rather than derived, this page says so where it appears.
Method version v4. When the method changes, this page changes with it and takes a new effective date.
02 · What a holdout is, and what it is not
A holdout is a random slice of sessions that receives no interventions at all: no promotion, no offer, no variant. Assignment happens before exposure and holds for the epoch. It is not a control variant — a control variant receives a different treatment; a holdout receives none.
It is not a segment, not a lookalike, not a modelled counterfactual. Randomization at assignment is the entire mechanism: the unexposed arm is a real population of real sessions in your own orders, and the effect is a difference between two observed rates.
It is not an attribution model: no view-through windows, no probabilistic credit, no last-click. There is nothing to argue with, because there is no model.
03 · Holdout against A/B split
An A/B split answers which of two variants performs better. A holdout answers whether any of it was worth doing. Only the second question is incrementality.
A split has no unexposed arm. Run variant A against variant B of the same welcome offer and both arms carry the giveaway, so the dollars handed to buyers who were already going to purchase sit inside both numbers and cancel out of the comparison. A promotion can win its split convincingly while losing money against not running it.
Always-on matters for a second reason: the baseline exists before you decide to test anything, so every discount, offer and price change running today is already being measured against it. You did not have to predict which one would leak.
The slice is 10–15%, and it has a price stated rather than hidden: that share of your traffic sees no promotion for as long as the holdout runs, converting at whatever your store does unaided. The range is a product specification, not a derived optimum — wide enough that the held-out arm accumulates usable sample inside a reasonable window, narrow enough that the cost stays bounded. The specimen store’s slice is 12.1%, inside the band.
04 · Sequential statistics and the peeking problem
A fixed-horizon test asks one question at one moment: after n sessions per arm, is the difference real? Its 95% interval is valid only if you look exactly once, at n. Look on day 4, day 7 and day 10, stopping the moment the interval clears zero, and you have run three tests rather than one. Each look is another opportunity for noise to cross the line, so the stated 5% error rate is no longer 5% — and every false positive arrives wearing a confidence interval.
That is the peeking problem. It is not a rule against curiosity; it is arithmetic — a threshold calibrated for one look does not hold for many. Operators peek, and the engine does not ask for discipline it will not get. Sequential inference is designed so the interval is valid at every look rather than only the last: read the number on any of those days and the 95% interval means the same thing each time.
The guarantee is not free. Holding error at 5% across an open-ended series of looks requires a wider interval, and a wider interval requires more sessions. That surcharge is the leading factor of the power equation, and section 06 states what is and is not known about it.
05 · Sample ratio mismatch, and the gate it holds
Before any effect is computed, the readout checks that traffic split the way assignment said it did.
Specimen, day 10: assignment is 12.1%. The epoch carries 11,940 held-out and 86,730 exposed sessions, 98,670 in total. Expected held-out at 12.1% is 11,939.1; observed is 11,940. That gives χ² ≈ 0.00008, p ≈ 0.99 — the split is what it claimed to be, and the readout is stamped SRM passed.
It matters because the failure it catches is silent. A bot filter, a redirect rule, a cache that treats one arm differently — any of these breaks randomization while every dashboard keeps rendering, and the arms stop being comparable populations. Every number computed from them is then wrong in a direction nobody can recover.
So it is a gate, not a caveat. Fail it and no verdict renders — not a hedged verdict, not a warning banner. A broken assignment does not produce a slightly wrong answer; it produces a confident answer to a different question. It is also what makes the Ledger’s dollars auditable: a booked amount asserts its arms were comparable, and this check is the evidence.
06 · Why the engine refuses under-powered tests
The refusal card prints its own equation, which is the only reason its verdict can be checked:
nseq = 1.35 × 2σ²(z0.975 + z0.8)² ÷ δ² · σ = $4.10
Term by term. σ is dispersion — how noisy contribution margin per session is in this store; $4.10 is a stated specimen input, printed on the card. δ is the effect being looked for, in dollars per session, and is also a stated input. z₀.₉₇₅ = 1.959964 and z₀.₈ = 0.841621 fix 95% confidence and 80% power. The leading 1.35 is the sequential surcharge from section 04.
Worked: 2σ² = 33.62. (1.959964 + 0.841621)² = 7.849. 1.35 × 33.62 × 7.849 = 356.24 — one constant, so the card is a single curve: sample equals 356.24 divided by δ². For the refused test — a product page badge copy change, expected effect 3.0% against an anonymized beauty-vertical prior at N=37 — that returns 99,000 sessions per arm. The store has 620 eligible sessions per day per arm, and 99,000 ÷ 620 = 160 days against an honest window of 21.
δ’s value is not printed, and the omission is deliberate. Expressing the effect in dollars per session requires a per-session baseline for the specimen store, and publishing that would make a second specimen fact out of the first — a figure about a store that does not exist, asserted as though it had been measured. δ is an input; the derivation stops there.
The verdict is refusal, with the arithmetic attached. What follows it is a reformulation the same store can power: move the free-shipping threshold from $50 to $65, same margin hypothesis, expected effect 9%, n = 11,000 per arm, 18 days — inside the window. The honest answer to “can I measure this” is sometimes no, and a tool that never says no is not measuring.
On the 1.35, which has no source at any tier. The noun is wrong in the literature’s own terms: every source defining a sequential inflation factor defines it as the ratio of the maximum sample a design may reach to the fixed-design sample, not the sample required — and such a design typically finishes faster than the fixed test, because it can stop early. Using a ceiling as a feasibility gate is defensible; calling it the sample required is a different claim, and the card makes the second. There is also no finite factor for continuous monitoring: in the Pocock family it increases without bound in the number of looks — roughly 1.23 at five analyses, 1.30 at ten, 1.43 at fifty. And 1.35 is not reachable in any published table at a design anyone runs.
The correct factor depends on a monitoring design HOLDOUT has not declared, and the candidates are not close together:
| Monitoring design | Factor | Sample / arm |
|---|---|---|
| As printed on the card — held | 1.35 | 99,000 |
| Pocock, five analyses | 1.23 | 90,000 |
| O’Brien–Fleming / Lan–DeMets | 1.03 | 75,000 |
Samples are rounded, from the same equation with only the factor changed. Guessing a design in order to source a number would be the same error as inventing a statistic, committed while correcting one — so the factor is held as printed and the gap is stated here. It is the one open item on this page, and it closes when the design is declared.
07 · The Proof Moment bridge, worked in full
The specimen day-10 readout publishes 41% — a cannibalization rate with a 95% interval of [34%, 48%] — and below it, order rates of 2.21% exposed and 1.79% held out. A reader reconciling those finds 1.79 ÷ 2.21 = 81%, or 19% incremental, and no route to 41.
Both are correct and neither is reachable from the other, because they are in different units. 41% is a share of redemptions; 2.21% and 1.79% are rates per session. What connects them is the number of redemptions, and the card did not print it.
It prints it now — 621 pop-up redemptions. The held-out arm made 214 orders in 11,940 sessions, a rate of 1.79229%. The exposed arm made 1,921 orders in 86,730 sessions. At the holdout’s own rate those sessions would have produced 86,730 × 0.0179229 = 1,554.5 orders with no promotion at all, so 1,921 − 1,554.5 = 366.5 rounds to 367 orders the promotion explains. The redemptions it does not explain are the rest: (621 − 367) ÷ 621 = 41%.
Read the direction of that chain carefully, because the direction is the disclosure. 41% is the specimen’s stated cannibalization rate. 621 was selected for coherence with it — chosen so the card’s rates and its cannibalization rate describe one consistent store. It is an input of the same class as 11,940, 214, 86,730 and 1,921, and nothing on this site claims that 41% falls out of it.
08 · Measured, projected, and how the Ledger recomputes
A measured figure is a difference observed in your own orders against the holdout, inside a stated interval. A projected figure extrapolates a measured effect forward under assumptions that have not been tested. The two never dress alike: projected figures carry the 45° hatch or a dashed underline, everywhere, and never share a total. It is a law rather than a preference because the two get confused in exactly one direction — the vendor’s.
On the specimen Ledger the annualization is recomputable from the surface: the figcaption states annualized = measured month × 12, so $31,204 × 12 = $374,448, printed as $374k per year. The verified total, $214,380, is $70,240 + $75,960 + $68,180 — the three entries carrying a holdout-verified amount, excluding the unverified $31,204 and the projected $374k. One entry is inconclusive and carries no amount, because a test that did not resolve has nothing to book.
The row amounts are stated specimen inputs, chosen as illustrative magnitudes. No decomposition into orders, basket size and margin rate will be published for them: it would be built backwards from a total that already existed and then presented as where the total came from. What the card demonstrates is coherence — every relationship it asserts between its own numbers holds — which is what a specimen owes.
In the product an entry stores its inputs rather than its conclusion — the epoch, the assignment, the raw order rows for both arms, the margin model version (v3) and the method version (v4) — so the figure is re-derived on demand rather than read from a cached total, and reconciles against the order export line by line.
09 · Glossary
- Holdout test
- A randomized slice of traffic held out of every intervention for an epoch, used as the unexposed baseline. It is distinct from a control variant, which receives a different treatment rather than none at all.
- Incrementality
- The difference in outcome between exposed traffic and held-out traffic. It is measured by subtraction against a randomized baseline, never by attributing credit across touchpoints.
- Discount cannibalization
- The share of promotional redemptions taken by buyers who would have purchased at full price. No published work reports it as a constant: it moves per offer and per store, which is why it has to be measured rather than looked up.
- Contribution margin per session
- Revenue less variable cost, divided by the sessions eligible to see a change. It is the readout unit because it is the only one that cannot be improved by discounting your way to more orders.
- Sample ratio mismatch
- A disagreement between the traffic split an experiment assigned and the split its data shows. It is evidence that randomization broke, and it invalidates every downstream figure in a direction that cannot be recovered afterwards.
- Statistical power
- The probability that a test detects an effect of a given size, if that effect is real. A test that cannot reach the stated power is not a weaker test; its result carries no information about the effect it was run to find.
10 · Sources
Every figure on this site that comes from outside the product is listed here with the document behind it. A primary or reputable-independent source sets what the surface says; an unsourceable figure is cut. Three of the four macro figures this site carried did not survive that pass.
| Figure | Source | Tier |
|---|---|---|
| Applied US apparel tariffs, 14.7% → 21.6% | Sheng Lu, University of Delaware — applied tariff rates of US apparel imports, series through May 2026, from USITC customs data | Reputable-independent |
| Meta average price per ad, +12% | Meta Platforms, Inc., Form 10-Q for the quarterly period ended 31 March 2026 | Primary |
Tariffs. Verified firsthand against the published dataset. World-average applied rate, HS chapters 61–62: 14.70% in December 2024, a peak of 31.53% in December 2025, and 21.6% in May 2026, the latest published month. The caption discloses the peak rather than publishing it as the live figure, because it has substantially unwound. Nobody has reproduced the series from primary customs data; that check is outstanding.
Meta. Quoted from the filing: average price per ad rose 12% year over year. It is deliberately not a cost per acquisition — the filing defines the metric as advertising revenue divided by ads delivered, “regardless of their desired objective.” No primary source publishes a Meta cost per acquisition and none can, because acquisition cost is an advertiser-side outcome. Caveat carried from the filing: part of the rise is attributed to a favourable currency effect, and no constant-currency figure exists.
Retail net margins — cut. No source at any tier measures the net margins of Shopify merchants as a population. The nearest primary dataset — IRS Statistics of Income, sole-proprietorship returns in retail trade, tax year 2023 — gives 4.1% on the wrong population, struck before the owner is paid, netting loss-makers against profitable firms whose counterpart is 14.0%. The site now says nothing about retail margins.
Discount cannibalization — no published band. This site previously carried a range for the share of blanket-discount redemptions taken by buyers who would have purchased anyway, introduced as published work. It is removed from every surface: the source it was credited to contains no share-of-redemptions figure at all, and the only traceable origin was a statistic whose denominator is promotions that failed to lift, not redemptions. “Published” was the specific falsehood. Nothing replaces it, and the absence is structural — peer-reviewed work declines to divide the aggregate effect of a promotion by its redemptions, because classifying one redemption as incremental needs a counterfactual for that buyer. The nearest rigorous evidence, domains named:
| Study | What it establishes |
|---|---|
| Sahni, Zou & Chintagunta, Management Science (2017), 70 randomized experiments | Most of a discount’s gain does not flow through redemption at all: ninety percent of the measured gains are not through redemption. |
| Reimers & Xie, Do Coupons Expand or Cannibalize Revenue?, Management Science 65(1):286–300 (2019) | Cannibalization of full-price customers is real and measured in a retail e-coupon setting. The authors explicitly decline to report a cannibalization level. |
| Boomhower & Davis, Journal of Public Economics 113 (2014) 67–79 | A regression discontinuity finds 43–54% of participants non-additional — the right quantity in the right unit, but measured on a Mexican government appliance-rebate program, not retail discount codes, which is why it is named here and not carried onto the surface. |
| Simester (2015), MIT survey of 61 marketing field experiments | The share of redemptions that is non-incremental appears nowhere as a reported statistic. |
The 1.35 factor. Section 06’s account rests on Wang & Tsiatis (1987), whose factors give 1.23 for a Pocock design at five analyses, two-sided α = 0.05, 80% power; on Lakens, Improving Your Statistical Inferences, for the definition of the factor as a ratio of maximum to fixed sample; and on Johari, Koomen, Pekelis & Walsh, Always Valid Inference, Operations Research 70(3):1806–1821, which publishes no inflation multiplier for continuous monitoring at all. None supports 1.35.
11 · What this document does not claim
Every specimen figure describes one illustrative store, used consistently across the site. No customers, no results, no ratings and no operating history are asserted anywhere.
The product-specification figures — the 10–15% slice, 95% confidence, 80% power, the 21-day honest window, the 15k monthly-session floor — are parameters HOLDOUT sets about its own engine. They are falsifiable against the engine and nothing else, and they become public commitments by appearing here: when one moves, this page moves in the same commit. How the three kinds of figure on this site are to be read is set out in section 04 of the Terms of Service.
Effective 20 July 2026