Design Validation (Placebo Tests)
Before a test launches, ProofPod stress-tests the design itself. Some tests can't buy statistical power with a longer duration—a pricing change might only be able to run for exactly one month—so the design has to be validated up front. That's what the Design Validation card in the location-selection step does: it runs placebo tests against your chosen treatment/control groups and tells you, before you commit, how often a design like yours "detects" effects that aren't there.
Everything the card shows is advisory. A warning never blocks you from launching—but it is recorded, permanently, alongside the design you launched.
The two placebo families
In-time (backdated) placebos
ProofPod re-runs your exact design on historical windows of the same length as your planned test, pretending the intervention happened back then. Nothing actually happened in those windows—so any "effect" the design measures is noise.
Each backdated window reports the effect it would have measured and whether that effect would have cleared the significance threshold. Two variants are computed:
- Frozen weights (drives the verdict) — uses the exact donor weights you're about to launch with. This validates the design you committed to, which is what matters when someone with audit rights asks.
- Refit weights (stored for reference) — re-runs the weight-fitting procedure inside each window. This tests the stability of the selection procedure.
The headline number is the effect ratio: the average placebo effect size relative to the minimum detectable effect at your test's duration. Below 0.5× your design is quiet; above 1× the design's noise is as large as the effect you're hoping to detect. The count of flagged windows ("2 of 6") is shown for context but doesn't drive the badge—with a handful of windows, a rate would swing on a single window.
In-space (permutation) placebos
Each control location takes a turn as a fake treatment location on a held-out window, with weights refit from the remaining donors. If your real treatment group's held-out effect ranks as extreme among these placebo units, the design already singles out your treatment locations with no treatment applied—a bias warning. This is the standard permutation inference for synthetic control designs.
Small chains get an honest answer instead of a fake number: with fewer than 8 donor units, a permutation rank is too coarse to badge, so the card reports the pool size and a warning rather than a p-value. Chain-wide tests (no untreated locations at all) are reported as having no valid donor pool.
Calendar-month alignment
Tests whose mechanism works on a day-of-month cycle (renewals, billing dates) can't be validated with windows that straddle month boundaries—the straddle smears exactly the effect being checked. When your test starts on the 1st and ends on the last day of a month, placebo windows automatically align to calendar months: each backdated window covers the same day-of-month range, shifted back whole months.
Duration-honest statistics
Every MDE (minimum detectable effect) quoted in the wizard is computed at your actual planned test duration. A 4-week test is quoted a 4-week MDE. The placebo thresholds, the effect ratio, and the analysis-phase success criteria all use the same duration-aware number.
All sources also estimate noise on the same basis: the trailing 365 days of available data, anchored to when your data ends—not to when the test is planned, so planning further ahead never changes the validation basis. Weekly noise buckets are anchored the same way in every path. Older history is excluded from diagnostics, donor-weight fitting, and placebo baselines alike, so the wizard's quoted MDE and the placebo card's MDE agree, and years-old noise can't inflate (or deflate) what you're told about the design. The stored design snapshot records the actual data span the validation ran on—if your data ends before your test starts, the record says so.
MDEs are quoted for a two-sided test (detect a lift or a drop) at the product-wide defaults of 80% confidence and 80% power (one home: app/analysis/stats_defaults.py — analysis-phase confidence intervals and Bayesian credible intervals use the same level). The AR(1) autocorrelation adjustment is clamped at φ ≥ 0 everywhere—the design-phase MDE quote and the analysis-phase confidence intervals share the same clamp: a negative weekly autocorrelation estimate on ~52 points is bucket-phase noise, and crediting it would quote an MDE (or narrow a readout CI) better than independent data supports—optimistic in exactly the direction that hurts a launch decision.
Validation metrics and falsification outcomes
By default the suite validates your primary metric. The card's Add validation metric control lets you add:
- Secondary — another outcome your test should move (e.g., a repricing test moves both join rate and revenue-per-join, in opposite directions). It gets its own full placebo suite and contributes to the overall verdict.
- Falsification — an outcome your test cannot move (existing members, ineligible segments). Its placebo run establishes a noise band that is stored with the design. If that outcome moves during your real test, you're looking at a local shock, not your intervention. Falsification results never affect the overall verdict—noisy placebos on a can't-move outcome say nothing about your design.
What gets stored (the audit trail)
Placebo results are computed and persisted server-side at design time, in an append-only record that includes: the treatment locations and donor weights, the pre-period used, the planned dates and duration, the window alignment, every per-window and per-unit placebo result for every metric, the verdict thresholds in force (versioned), and the overall verdict. When you create the test, the record is linked to it; if the test is later deleted, the record survives.
Re-running matching produces a new record rather than overwriting the old one, so losing candidate designs remain visible. Placebo runs and test creation are also written to the audit log.
The stored snapshot is visible on the test detail page for the life of the test (and after), exactly as it appeared when the design was accepted.
Known limitations
- Survival-shaped outcomes (e.g., 90-day retention) can't be expressed in the weekly panel the engine uses. Tests with survival-shaped primary outcomes get placebo validation only on proxy aggregates. This limitation is recorded inside every stored snapshot.
- In-time windows require at least 26 weeks of history before each pseudo-treatment start, so shorter histories produce fewer windows (or none, reported as insufficient history).
Under the hood
Per backdated window, the seasonal model (per-location means + chain-level Fourier) is re-fit on the fit period only and then applied to the pseudo-post weeks. Fitting on the full panel would absorb a pseudo-post shock into the seasonal term and under-detect—anti-conservative, the wrong direction for an audit defense. Fit periods shorter than 52 weeks drop to a single Fourier harmonic (a 52-week cycle is half-observed at 26 weeks).
Significance per window: |mean of demeaned difference over the window| > z·s_D/√T_eff, with s_D and the AR(1) φ estimated from the fit period and T_eff = L(1−φ)/(1+φ).
The in-space permutation is computed symmetrically: the treatment's held-out effect uses weights refit on the fit period (same procedure as the placebo units), because the accepted weights were optimized on data that includes the holdout and would rank artificially well. p = (1 + count of placebo |effects| ≥ treatment |effect|) / (n + 1).
Verdict thresholds (v1): in-time effect ratio pass ≤ 0.5, warn ≤ 1.0; in-space p pass ≥ 0.30, warn ≥ 0.10. Overall = worst component across primary + secondary metrics. Thresholds are stored inside each snapshot so old verdicts remain interpretable if defaults change.