DriftScope audit — verdict on the stream:
POSITIVE CONTROL (euron / BOCPD, full-stream): reject=True; top CP: 2015-01-23 (p=0.47), 2014-11-28 (p=0.41), 2022-03-29 (p=0.40)
NEGATIVE CONTROL (main / 3 pillars, per regime):
R1 (n=133): 0/3 (no signal); [h1=ok mmd=ok cooccurrence=ok]
R2 (n=389): 1/3 (single-pillar signal, requires DriftSim power context); [h1=ok mmd=ok cooccurrence=reject]
R3 (n=436): 0/3 (no signal); [h1=ok mmd=ok cooccurrence=ok]
Family B (per-number FDR, benjamini_yekutieli): 0/150 rejections
WATCHLIST (DoD-5): None (honest null)
DriftScope — Stationarity Audit Report
EuroJackpot discrete stream: a positive/negative control demonstration
📄 In a hurry? Read the one-page executive summary instead (print-friendly — Ctrl+P → a clean single-page PDF).
1. The finding
DriftScope audits discrete streams that are uniform by design for non-stationarity. EuroJackpot is the flagship case study with a known ground truth: the pool of euro numbers changed its rules twice (2014, 2022), while the main 1–50 pool never did.
The framework delivers an unambiguous verdict: it confirms the known changes and stays silent where there is no change.
2. Method — three independent pillars
The verdict rests on three mutually non-redundant families of detectors (the Disagreement Protocol, preregistration v6 §6.5). Each sees a different class of deviation:
| Pillar | Family | What it detects | What it is blind to |
|---|---|---|---|
| H1 (BOCPD) | temporal / global | change-points in the symbol distribution | pair structure |
| MMD | distributional | windowed frequencies departing from uniform (freq_shift, trend, autocorr) | pair structure |
| Co-occurrence | joint | over-represented number pairs under uniform margins (pair_corr) |
marginal signal |
The key property: pair_corr is visible only to co-occurrence — chi²/MMD are provably blind to it when the margins are preserved (power ≈ FPR). The three pillars do not overlap; a signal is classified by agreement: 3/3 · 2/3 · 1/3 · 0/3.
Design decision: the H1 pillar is represented by BOCPD (calibrated per field, FPR≈0.05). ADF/KPSS/Welch/ACF play a diagnostic role — they do not vote (avoiding OR-inflation of the FPR from correlated sub-tests).
3. Result on real EuroJackpot
Positive control (euron). BOCPD detects change-points covering both known ground truths — 2014-11-28 (the first draw containing the number 9, 51 days after the rule change) and 2022-03-29 (the first draw containing the number 11). The 1–50 negative control remains untouched.
Interactive view. The same BOCPD curve, interactive — hover any draw for its exact posterior, zoom into either change-point, or pan across the 958-draw stream. The red markers carry the draw date (2014-11-28, 2022-03-29):
Negative control (main 1–50). Three independent pillars, evaluated within each regime (H0 §1 is stationarity inside a regime; the 1–50 pool is invariant across regimes, so each regime is an independent negative control):
| Regime | n | Pillar | p-value | reject H0 |
|---|---|---|---|---|
| R1 | 133 | h1 | 0.792 | False |
| R1 | 133 | mmd | 0.480 | False |
| R1 | 133 | cooccurrence | 0.325 | False |
| R2 | 389 | h1 | 0.552 | False |
| R2 | 389 | mmd | 0.975 | False |
| R2 | 389 | cooccurrence | 0.010 | True |
| R3 | 436 | h1 | 0.558 | False |
| R3 | 436 | mmd | 0.655 | False |
| R3 | 436 | cooccurrence | 0.585 | False |
Disagreement per regime: R1 0/3, R2 1/3, R3 0/3.
Supplementary sequential lens (IT / Lempel-Ziv). Beyond the three Disagreement-Protocol pillars, the information-theoretic Lempel-Ziv 1976 complexity test (reporting/information_theory.py, order-shuffle null over draw blocks) reads each regime along a fourth, independent axis — the ordering between draws. It conditions on both the marginal and the within-draw joint, so it is orthogonal to the three families and fires only on sequential structure (period / autocorrelation). It is a supplement, not a fourth Disagreement-Protocol pillar — that set stays three-way (DoD-4 = 3/3). Reported here to confirm the negative control is clear along the sequential axis too:
| Regime | n | IT (LZ) p | reject H0 |
|---|---|---|---|
| R1 | 133 | 0.535 | False |
| R2 | 389 | 0.855 | False |
| R3 | 436 | 0.620 | False |
Every regime reads clear (high p, no rejection): the main 1–50 pool carries no sequential structure within any regime — the honest null confirmed along a fourth, independent axis, without touching the three-pillar verdict.
4. Rigor — FDR and the honest watchlist
The marginal per-number signal is controlled by the Family B family: per-number exact-binomial tests (count_k ~ Binomial(n, 5/50)) computed per regime — 50 numbers × 3 regimes = 150 hypotheses — pooled into one family and corrected with Benjamini-Yekutieli (valid under arbitrary dependence structure — the 5/50 counts are negatively correlated). The omnibus detectors (chi², gap, co-occurrence) are not per-number and are reported separately as complementary families (preregistration_v7 §5).
- Family size: 150 hypotheses (per-number).
- Method: benjamini_yekutieli, α = 0.05.
- Rejections (q ≤ α): 0 / 150.
- Smallest q-value: 1.000.
Honest watchlist (DoD-5). The adaptive layer returns patterns “worth monitoring” only after passing both gates: FDR (q ≤ α) and convergence ≥1 pillar. When none passes, it returns None (an honest null), not an empty list or an extrapolation:
Watchlist: None — honest null. None of the 0 candidate(s) cleared the gate (FDR q ≤ 0.05 AND convergence ≥ 1 pillar). No watchlist — the methodology produced no validated signal.
The negative control is not naively all-zero: in R2 a single pillar (co-occurrence) flags one pair on the main pool. The Disagreement Protocol does not suppress it — it classifies it as 1/3 (“single-pillar, requires power context”), not a finding. Promotion to the watchlist requires both the FDR gate (Family B, per-number) and convergence; a lone single-pillar flag on a clean main pool clears neither. So no pattern clears the rigor bar → the framework proposes nothing. This is deliberate: a raw signal is reported as raw, absence of convergent evidence as absence of evidence.
Concretely, in R2 the co-occurrence pillar localizes the flag to a single pair (10, 25) (max-pair z ≈ 4.7). The tell is that this is a purely joint excess: across 389 draws number 10 appears 28 times and number 25 appears 38 times (uniform expectation 39 each) — both marginally unremarkable, yet they land together far more often than independence allows. That is the clean cell: the per-number Family B sees nothing on either number, so the marginal FDR gate cannot corroborate the pair and it stays 1/3. One flag across three independent regime tests is also the most likely non-zero outcome under the null (P ≈ 1 − 0.95³ ≈ 14%).
A note on the 1/3 cell — and why we do not hard-gate on ≥2/3. Some departures are, by construction, visible to only one family: a pure pairwise correlation leaves every marginal and every windowed frequency exactly uniform, so neither Family B nor MMD can see it — only co-occurrence can (the “clean cell” of the complementarity proof, §6.5). Such a genuine signal would surface as 1/3, not 3/3. A naive “≥2/3 = real” rule would therefore be structurally blind to an entire class of real defects. We do not use that rule: the watchlist’s primary gate is the per-number FDR (Family B), with convergence only required at ≥1 pillar — so a true single-family signal that also clears FDR can still surface, while a lone flag without FDR support (the R2 co-occurrence pair above) does not. The 1/3 classification is thus a routing decision (“requires power context”), not a dismissal: the cost of demanding convergence is named, and the safeguard against it is explicit.
5. Reusability — the same audit on PRNGs
A null result on EuroJackpot (“we found nothing”) is only worth trusting if the instrument is sensitive. So we point the exact same battery (Family B per-number FDR + MMD + co-occurrence) at streams with a known ground truth: two well-behaved generators (MT19937, Xorshift64), two cryptographic ones (ChaCha20, AES-CTR-DRBG), the same MT19937 with two deliberately injected defects of different kinds — a marginal bias (one number over-represented in 15% of draws) and a period-truncation (a short 50-draw cycle that repeats, freezing the frequencies) — and real EuroJackpot. The framework is detector-agnostic — feeding it a PRNG stream instead of a lottery is just a different DrawRecord source (ingestion/rng_streams.py).
| Source | Class | n | Family B (reject/size) | MMD p | Co-occ p | IT (LZ) p | Verdict |
|---|---|---|---|---|---|---|---|
| MT19937 | good | 1500 | 0/50 | 0.595 | 0.055 | 0.780 | clear |
| Xorshift64 | good | 1500 | 0/50 | 0.700 | 0.320 | 0.315 | clear |
| ChaCha20 | crypto | 1500 | 0/50 | 0.140 | 0.490 | 0.635 | clear |
| AES-CTR-DRBG | crypto | 1500 | 0/50 | 0.740 | 0.225 | 0.710 | clear |
| MT19937+bias(7) | DEFECT | 1500 | 1/50 | 0.005 | 0.465 | 0.970 | FLAG |
| MT19937+period(50) | DEFECT | 1500 | 27/50 | 0.005 | 0.005 | 0.005 | FLAG |
| drand | beacon | 1500 | 0/50 | 0.915 | 0.550 | 0.260 | clear |
| NIST-Beacon | beacon | 1500 | 0/50 | 0.565 | 0.985 | 0.465 | clear |
| EuroJackpot | real | 958 | 0/50 | 0.885 | 0.940 | 0.700 | clear |
Sensitivity (defect → FLAG): ✓ · Specificity (good / crypto / real → clear): ✓.
The two defects fire differently, and that contrast is the showcase. The marginal bias is caught narrowly — Family B (the over-represented number breaks its per-number binomial) and MMD (the windowed frequency departs from uniform) — but not co-occurrence, which targets pair structure, not marginals. The period-truncation is caught broadly across all three pillars: the frozen cycle distorts marginals (Family B rejects many numbers at once), the windowed distribution (MMD), and pair structure (co-occurrence) simultaneously. The framework therefore reports not just whether a stream is defective but what kind of defect it is. The two good PRNGs, both cryptographic PRNGs, and real EuroJackpot all come back clear — specificity holds across two distinct crypto primitives (stream vs block cipher). The Verdict column applies the same Disagreement logic as the main pipeline: FLAG requires ≥2 of the three core families (marginal / distributional / joint) to reject — a lone family is not a finding (it clears), which is what keeps the specificity claim robust to a single borderline detector. Both defects clear that bar (bias: Family B + MMD; period: all three); good / crypto / real fire none.
One caveat on the EuroJackpot row: here Family B runs on the full stream (50 numbers), for parity with the synthetic sources — a PRNG stream has no calendar regimes, so the exact same battery is applied identically to every source, which is precisely the reusability claim this section makes. The regime-split treatment (50 × 3 regimes = 150 hypotheses, §4) is the canonical EuroJackpot headline; both readings return clear, since the 1–50 main pool is invariant across the 2014/2022 change-points.
A fourth, supplementary lens — an information-theoretic Lempel-Ziv 1976 complexity test (reporting/information_theory.py, order-shuffle null over draw blocks) — adds a sequential view orthogonal to the three core families: it conditions on both the marginal and the within-draw joint, firing only when structure lives in the ordering between draws. It is deliberately blind to the marginal bias (the IT (LZ) p column stays high for +bias) yet flags the period-truncation defect outright (low p) — and, like the core battery, reads real EuroJackpot as incompressible / clear. It is a supplement, not a Disagreement-Protocol pillar (that stays three-way; DoD-4 = 3/3).
This is the point of the whole framework in one table: the same detector that stays silent on EuroJackpot’s main pool lights up on a planted defect and clears a crypto-grade RNG. The honest null is not blindness — it is a calibrated instrument reporting the absence of a signal. It also cashes the reusability claim directly: DriftScope is a general non-stationarity / uniformity auditor (NIST RNG, cryptographic PRNG, financial random-walk are natural targets), not a lottery-specific script.
6. Second real-world case study — Multi Multi (20-of-80)
The PRNG benchmark proves sensitivity on synthetic streams with a known ground truth. The final question is whether the framework generalizes to a second real game of a different shape — without touching detector code. Multi Multi draws 20 numbers from a pool of 80 (vs EuroJackpot’s 5-of-50). Every detector derives its pool and draw size from the DrawRecord itself, so the same battery runs unchanged; only the data source differs (16,827 draws, 1996–2026; the audit runs on the most recent 2,000, since BOCPD is O(T²)).
Calibration first. A null on a new pool is only trustworthy if the detectors are calibrated there. We re-measured the MMD false-positive rate at pool=80 on an honest uniform null (200 trials, scripts/calibrate_mmd_pool.py): window=25 → FPR = 0.035, well within Monte-Carlo error of the nominal 0.05 — the detector is not miscalibrated at the larger pool. The BOCPD reject threshold was likewise re-derived per pool (80 → 0.34, 95ᵗʰ percentile of the null).
| Detector | value | reject H0 |
|---|---|---|
| BOCPD (main, temporal) | max=0.304 (thr 0.34) | False |
| Family B (per-number FDR) | 0/80 rejected | False |
| MMD (distributional) | p=0.030 | True |
| Co-occurrence (joint) | p=0.680 | False |
| IT / LZ (sequential, supplement) | p=0.090 | False |
Verdict: clear, with a textbook caveat. BOCPD (below the pool-80 threshold), Family B (0/80), co-occurrence and the IT/LZ supplement all read clean. A single pillar — MMD — rejects at p≈0.03. Read through the Disagreement Protocol this is 1/3 (one of the three core pillars) — the same classification as the lone R2 co-occurrence flag in §4. A single-pillar rejection at α=0.05 is expected (≈1 test in 20; across five families P(≥1)≈23%) and is not a finding: promotion requires convergence, which a lone flag does not clear. The framework neither hides it nor over-reads it — which is precisely why the verdict rests on agreement between independent families, not on any single test.
This closes the reusability loop on real data. DriftScope is not a EuroJackpot script: the same calibrated instrument audits a structurally different real game (4× the pool, 4× the draw size) and behaves exactly as designed — quiet where there is no convergent signal.
7. World Lottery Audit — Powerball, Mega Millions & UK Lotto
Multi Multi answers “does the instrument stay quiet on a clean second game?”. The World Lottery Audit asks the sharper question: does it blindly recover documented rule changes in games it has never seen? Powerball (5-of-69) and Mega Millions (5-of-75) histories come from the official NY Open Data portal (data.ny.gov); UK Lotto (6-of-59) adds the longest stream of all, reaching back to the game’s first ever draw in 1994. Their matrix changes are public record — five ground-truth change-points across 7,689 draws (1994–2026), including a pool shrink (Mega Millions 75→70 in 2017), a harder target than any expansion.
UK Lotto earns its place by being a near-perfect natural control for Powerball: the same kind of event — a pool expansion — three days apart (UK 49→59 on 2015-10-10, Powerball 59→69 on 2015-10-07), in two unrelated games on two continents. Whatever separates their outcomes below cannot be attributed to the calendar or to the kind of change.
Its history is also the one case here without a machine-readable official feed: the operator’s site is closed to non-browser clients and its export window is capped at roughly 180 days, so the stream comes from a mirror. Rather than trust it, we cross-validated it against a 64-record sample taken straight from the official export — main numbers agree on 64/64 records (scripts/validate_uk_mirror.py, reproducible). The single discrepancy found is a crossed pair of bonus balls on one date; the bonus stream is consequently not used at all.
Two reusability-layer refinements were fixed before reading any results (both documented in reporting/lottery_audit.py): a coverage warm-up of ⌈(N/k)·H_N⌉ draws — for small-k pools the default N∕∕K warm-up grossly under-covers the coupon-collector transient, whose first-appearance spikes otherwise inflate the null (mean null max 0.43 → 0.19 after the fix; thresholds re-calibrated at this warm-up: pool 69 → 0.49, pool 75 → 0.39, pool 59 → 0.53, FPR≈0.05) — and onset-based localization, because a pool expansion smears its signature over months of new-symbol debuts (measured: +81 days for Powerball, +189 days for Mega Millions).
| Documented change | BOCPD onset (blind) | Family B contrast | detected |
|---|---|---|---|
| white 59->69 (2015 matrix change) | below threshold | appeared {60..69} | yes |
| white 52->56 (2005 matrix change) | below threshold | appeared {53,54,55,56} | yes |
| white 56->75 (2013 matrix change) | 2013-10-22 (Δ +0 d) | appeared {57..75} | yes |
| white 75->70 shrink (2017 matrix change) | below threshold | vanished {71,72,73,74,75} | yes |
| main 49->59 (2015 matrix change) | 2015-10-10 (Δ +0 d) | appeared {50..59} | yes |
5/5 documented changes detected; 0 spurious onsets across 7,689 draws.
Four results deserve emphasis. First, two of the five changes are localized blind to the exact day — UK Lotto 2015 (onset 2015-10-10, Δ = 0) and Mega Millions 2013 (onset 2013-10-22, Δ = 0) — and the Family B pre/post contrast recovers the exact matrix delta symbol-by-symbol in all five events: {50..59} appearing in UK Lotto 2015, {60..69} in Powerball 2015, {71..75} vanishing in the 2017 shrink. Not one spurious onset appears across 7,689 draws and 32 years.
Second, the twin events are the sharpest evidence in this report, precisely because they disagree. UK Lotto 2015 is caught by both pillars at day zero; Powerball 2015, three days earlier and structurally identical, is a near-miss for BOCPD — its change peak reaches max cp_prob = 0.474 against the pool-69 null threshold of 0.49, an empirical p = 0.065, formally not significant at α = 0.05. The framework reports it as such, and Family B carries the detection instead. Two matched events, one caught twice and one caught once: that is what a calibrated instrument looks like, as opposed to one tuned until everything fires.
Third, the shrink case exposes a structural detector asymmetry: BOCPD reacts instantly to a new symbol (a zero-count category appearing is a likelihood shock) but is nearly blind to symbol retirement, where evidence accumulates only from absence. Family B’s two-sided binomial has no such asymmetry. Fourth, taken together these are the Disagreement Protocol’s complementarity argument — previously demonstrated on synthetic planted signals (§4) — replicating on real, documented ground truth: every change is caught, but no single pillar catches all of them.
8. Beacon withholding — a bespoke observable for a battery-invisible attack
The reusability battery (§5) reads two live public randomness beacons — drand and the NIST Beacon — as clear. But a uniformity battery is the wrong instrument for the sharpest attack on a proposer-driven beacon like Ethereum’s RANDAO, and reporting clear there would be dishonest: the battery has provably zero power against it.
An epoch’s final RANDAO mix fixes proposer duties two epochs later, and a validator can only influence it by not publishing — withholding a block. Crucially, the attacker selects among candidate mixes that are each still uniform, so the mix’s marginal distribution is unchanged; MMD, Family B and co-occurrence all see nothing. The trace withholding actually leaves is a missed slot at the tail of an epoch, because prediction is only possible when no unpredictable contribution follows — at the epoch’s last slot (position 31), extending backward only across a contiguous tail run the attacker owns. So the attack’s signature is an excess of misses at position 31, decaying toward the front of the epoch.
This turns a benign confound into an asset. Position 0 carries a structurally elevated miss rate (the epoch transition — justification, finalization, shuffling — is computed there and slow nodes miss it), but position 0 is the least profitable withholding slot, with 31 unpredictable contributions still to come. Benign and adversarial explanations point at opposite ends of the epoch, which is exactly what makes the tail test clean. Position 0 is a pre-specified exclusion from the null reference (identified in a scouting probe before the test was designed) and reported separately rather than quietly dropped. The full test — primary binomial on position 31, secondary on the tail set 28–31, omnibus Family B over all 32 positions — was committed before the scan produced any numbers (606dee3); the git history is the pre-registration.
Scanned 96,000 slots (3,000 epochs), 358 missed (0.373%). Null reference: 252 misses at positions 1–31 (position 0 excluded — 106 misses, 13× the reference mean, the epoch-transition confound).
| Test | Observed | Expected | p (one-sided) |
|---|---|---|---|
| primary — position 31 (last slot) | 9 | 8.1 | 0.426 |
| secondary — tail 28–31 | 34 | 32.5 | 0.418 |
| omnibus — Family B (32 positions) | rejects: pos_0 | — | — |
Verdict: clear. At 80% power this scan would have caught an attacker withholding 1 block per 332 epochs (9 withholds, 3.6% of all misses). A null result bounds the attack — it does not certify its absence.
The tail is quiet: the last slot sits within a whisker of its expectation and the primary test does not reject. The only Family B rejection is position 0 — the framework surfaces the benign confound explicitly rather than hiding it, and the confound sits at the opposite end of the epoch from any withholding signature, so it cannot masquerade as one. This is the discipline of the whole project applied to a live adversarial setting: the right observable chosen before the data, a clear verdict reported with the power bound that makes it readable, and a known confound named out loud instead of swept under the null.
9. Reproducibility (DoD-6)
- Determinism: every detector is a pure function of the stream; the RNG is seeded from the data contents (⊕
BASE_SEED), independent of call order. - Manifest:
scripts/archive.pyproduces a deterministic SHA-256 manifest of the seed CSV + artifacts (a cold-machine re-run yields a bit-identical result). - Pre-registration: all methodological choices are frozen in
methodology/preregistration_v7.md; revisions carry arevision_reason, split into clean / data-informed (the §0 discipline).
Generated by driftscope.pipeline.run_audit on the seed CSV (data/seed/eurojackpot_history.csv, 958 draws 2012–2026). The chunk code is the framework’s public API, covered by the tests in tests/test_pipeline.py.