Pre-registered trading backtests – 27 kills, 2 passes, 0 funded

Strawburry1 pts0 comments

The Falsification Kit — tools that kill bad backtests<br>🪦'>

doc: FALSIFICATION_KIT.md<br>status: pre-registered · checkout live 2026-07-27<br>performance promises: 0

case file · originBacktesting tools with a body count.×27 FAILED

I have run 31 pre-registered backtests of retail trading strategies.<br>The pass/fail criteria were written down before each test, and the goalposts<br>stayed welded in place after. Twenty-seven died, among them the strongest raw edge this<br>pipeline ever measured, killed by friction alone. Two passed and still didn't get my<br>money, because a tiebreaker written in advance said they hadn't earned it. Every<br>strategy died in the pipeline before it touched a live account: total tuition paid to<br>the market, $0 .

This kit is that pipeline. It will not make you money.<br>It stops your backtest from lying to you about making money , which,<br>if you trade, is the same thing wearing work clothes.

×01

×10

×16

✓17

✓18

◐19

×20

×21

×25

◐26

×27

×31

gates 17–25 pass · benched · split · 6 kills, 5 in one day<br>gates 27–31 the collector's-side sweep · 5 kills in 30 hours

31 gates · 27 kills ·<br>1 pass · 2 certified, then benched by their own<br>pre-registered tiebreaker · 1 split verdict (Japan) · $0 lost live

counted from the public gate files; every number on this page is in the repo

exhibit a · the killsEvery family, measured against the gate it pre-registered.

Each of these looked plausible. Several looked great, right<br>up until honest fills, split-adjusted data, and a de-survivorship universe were<br>applied. The picture first, then the ledger; full writeups are free in the repo.

DISTANCE FROM THE GATE — best honest profit factor per family

fast-strategy gate, pre-registered: PF ≥ 1.30 at the stated cost · every dot is a documented result

Mean reversion (RSI-2) @ 0.25%/side<br>0.77

ORB long+short "Sharpe 2.4" paper · @ 0.10%/side<br>0.85

190-name mover universe @ zero cost<br>0.91

Momentum scalp @ zero cost, best window<br>0.97

Penny stocks @ a fantasy 0.1% cost<br>1.00

Trend (Donchian, broad) @ 0.25%/side<br>1.05

Pairs stat-arb (GGR) gross 2.23! · @ 0.10%/side/leg<br>1.06

Insider clusters (Form 4) 60d hold · @ 0.25%/side<br>1.17

Earnings drift (real 8-Ks) @ 0.25%/side<br>1.11

Catalyst gaps primary → untouched holdout

1.23 primary<br>0.93 holdout

0.70<br>1.00breakeven<br>1.30the gate

Not one family reached the line. The catalyst pair is the<br>expensive lesson: a passing primary sample (hollow dot) collapsing to negative<br>expectancy on the holdout it never touched. The pairs dot is the newest and<br>sharpest: PF 2.23 at zero cost — still the strongest gross edge in all 31 gates —<br>landing at 1.06 once each hedged round trip pays the spread four times. Exact<br>figures, verdicts, and causes of death: table below.

strategy familybest honest resultverdict<br>Momentum scalping (213-trade window)49% win, ±1.18%: a fair coinFAIL<br>Opening-range breakout, long+short ("Sharpe 2.4" paper)PF 0.82–0.85 @ 0.10%/sideFAIL<br>Catalyst gaps (≥5% + volume)PF 1.23 primary, 0.93 holdout FAIL<br>Earnings drift (real SEC 8-K dates)PF 1.11, worse than random gaps (1.29)FAIL<br>Mean reversion (RSI-2)PF 0.77 @ 0.25%/sideFAIL<br>News sentiment (19.5k articles, LLM-scored)made every config worseFAIL<br>Crypto funding carryreal mechanism, ~0.4% net in 2026FAIL<br>Cross-asset trend, monthly (Faber, zero tuned params)8.1% CAGR, −1.1% in 2022 · reader-forced 2008 extension: 5.9% over 18.4y, −1.5% in 2008 — pass holds, repricedPASS<br>Same rule, sampled daily7.31%; whipsaw ate itFAIL<br>Options premium selling (5 pro income ETFs + CBOE's own indices)every fund's Sharpe below its own underlyingFAIL<br>Turn-of-month window (T-4..T+3)Sharpe 0.44 vs SPY 0.73; 42% of returns in 29% of daysFAIL<br>Factor ETFs (MTUM/VLUE/QUAL/USMV, EW + 12-1 rotation)Sharpe 0.87 / 0.84 vs SPY 0.90FAIL<br>Pairs stat-arb (GGR distance, liquid, long-short)gross PF 2.23 → 1.06 @ 0.10%/side/legFAIL<br>Insider-cluster buying (2+ insiders, $200k+, SEC Form 4)PF 1.17 @ 0.25%/side — vs 1.23 gross: friction wasn't the murdererFAIL<br>Candlestick patterns (engulfing / hammer / piercing, 43,624 signals)pooled PF 0.95 @ 0.25%/side; the hammer is a literal coin flip (1.00)FAIL<br>Cash-merger arbitrage (789 deals hand-built from SEC filings, breaks paid in full)4.03% CAGR vs a cash+3 bar of 4.32%: clears it only at exactly zero costFAIL<br>Prediction-market cross-venue arb (17 identical-resolution Kalshi/Polymarket pairs)+2.2%/yr on deployed capital, less than T-bills (3.62%)FAIL<br>Paid order flow / maker rebates (receipts only, no simulated fills)net −0.17 to −0.23¢/share at retail tier; first positive payment starts at $10M/monthFAIL<br>Unified everything-rotation (1,365 stocks + 13 ETFs, one momentum pool)7.14% vs SPY's 11.53%; ETFs entered the top-10 in 0 of 55 monthsFAIL<br>Trend overlay on a momentum index fundhalved the drawdown, cost 3.4pp/yr vs a 2pp capFAIL

PF = profit factor. The pattern: nearly every DIRECTIONAL price-derived signal is ~PF 1.05–1.10 gross, and retail friction eats it whole. The market-neutral exception measured 2.23 gross and...

side registered cost gross fail gate

Related Articles