Phase Space Research · Quantfund

Signals

High-confidence conclusions from the quantfund research programme · updated 2026-07-31 16:15 UTC · build 1ce1b8d9
The objective function. Outperformance of the trades actually taken versus the full set that triggered that same day. Daily returns cluster hard, so an absolute target would mostly measure which days you traded; indexing to the day's own triggered set strips the regime out and measures selection skill. A signal that does not improve selection against the day's own triggered set is not a signal, however interesting it looks in isolation.

Every figure below was re-derived from the raw archive by code in the repository. Where it disagrees with the previous edition of this page, both numbers are shown.

Live — actionable

892 sessions LIVEYesterday predicts today

rho +0.408 · top third beats bottom third by +73.9pp

Of 25 day-level indicators tested against the equal-weighted terminal return of a session's alerts, the prior day's outcome is the only causal one that survives. Bootstrap CI [+58.9, +88.3] on 892 sessions.

So what: This is the day filter the autocorrelation implies. Yesterday's result is known before today's first alert fires, so unlike every other day-quality candidate tested it can actually be acted on.

Caveat: Measured on positions that can span sessions, so some of the autocorrelation is mechanical rather than informational — the same artefact the previous edition of this page flagged and never resolved. It is NOT yet re-measured on same-day-closed trades only. Treat the effect size as an upper bound until that is done.

894 sessions LIVEDays cluster, and most of them lose

64.3% of sessions lose money · top 20% of days carry 90.4% of gains

Top 5% of days carry 46.2%. Lag-1 autocorrelation of the daily mean is +0.397. All four figures reproduce the previous edition within 0.8pp.

So what: Selection is a day decision before it is an alert decision.

Caveat: Same overlapping-position caveat as above. The clustering is not in doubt; its magnitude is.

20,018 trades LIVEUpside is concentrated in a handful of trades

top 5% of trades hold 39.2% of all upside

Top 1% hold 15.8%, top 10% hold 54.7%, top 20% hold 73.4%. Independently re-derived and matching the previous edition to within 1.2pp on every bucket.

So what: Any rule that caps the winners destroys the payoff. This is the arithmetic reason scaling out and trailing stops both fail here.

Caveat: Computed on maximum favourable excursion, which is an opportunity measure — it is what the contract offered, not what a position captured.

Watch — real, not yet actionable

20,018 trades WATCHSeveral published figures are not currently reproducible

32 checks re-derived: 15 match, 4 close, 13 do not

The scripts behind the previous edition of this page were lost with a deleted worktree, so its numbers could not be reproduced by anyone. Re-deriving them from the raw archive: structural findings reproduce cleanly — day clustering on all five figures, upside concentration on all four, the exit-variant ordering, the 0DTE share. Absolute levels do not. Ever-green comes back 81.5% against 85.7% published, and median MFE 51% against 64%, a bias consistent across every channel.

So what: The conclusions survive an independent rebuild; several of the levels do not. Figures on this page are now the re-derived ones, and the code behind them is in the repository.

Caveat: A uniform downward bias points at a different entry-price convention in the original, not at an error in it. Until that is pinned down, treat level figures as this implementation's, and the previous edition's as unverified rather than wrong.

Rejected — tested and dead

Published as prominently as the accepted list. Knowing which plausible ideas are dead is worth as much to trade selection as knowing which are live, and it stops the same hypotheses being re-tested indefinitely.
893 sessions DEADMarket state at the open says nothing about the day

19 indicators · not one clears zero

First-30-minute return, overnight gap, opening range and volume for SPY, QQQ, IWM, TLT and UVXY — 19 features in all. Every bootstrap CI on the top-third-minus-bottom-third spread includes zero, and every rank correlation sits between −0.06 and +0.04.

IndicatorrhoTop−bottom third95% CI
SPY first-30min return−0.013−3.7%[−18.7, +12.6]
SPY overnight gap−0.059−4.8%[−18.7, +10.0]
QQQ first-30min return+0.002−0.5%[−15.7, +15.5]
UVXY first-30min return−0.022−2.4%[−16.2, +12.4]
IWM first-30min return−0.040−11.7%[−27.4, +3.2]
time of first alert+0.024+0.2%[−14.9, +15.2]

So what: Do not build a day filter from the tape. Whatever makes a good alert day good is not visible in index price action at the open — which also means a macro or volatility overlay is not the missing piece.

Caveat: Tests the first 30 minutes only, by design: an indicator needing the full session describes a day rather than filtering it. A slower-moving regime variable could still matter and is not covered by this rejection.

886 sessions DEADThe day's first alerts do not predict the rest of the day

sign flips between interleaved halves

Days whose first 2–3 alerts are up 15 minutes later go on to return +7.4% versus −1.8% on the remaining alerts. Pooled CI [−5.3, +22.7] includes zero; the chronological first half gives −0.6% and the second +16.1%; interleaved halves disagree in sign. 2 of 20 cells clear zero against ~1 expected by chance.

So what: The day-level autocorrelation is a property of realised days. It is not detectable early enough, from the alert stream alone, to skip the rest of a bad day.

Caveat: A sign flip between interleaved halves is the signature of noise, not of regime change — a genuine regime shift fails only the chronological split.

5,090 holds DEADHolding a trigger overnight does not pay — least of all after a good day

median −21.4% overnight · 29.6% win rate

Carrying an alerted contract from the day-1 closing bid into the next session loses money on median across 5,090 holds. Splitting by day-1 quality inverts the intuition: trades from good sessions return a median −37.4% overnight with a 20.4% win rate, against 0.0% and 47.4% from bad sessions. Winsorised difference −55.9%, CI [−63.0, −48.9].

So what: "Good days follow good days" is about the next day's new alerts. It does not transfer to the positions themselves. A contract that already ran on day 1 gives it back overnight; one that went nowhere has less to give back.

Caveat: The eligibility filter is not independent of day quality. Requiring a real position at the day-1 close (at least $0.10, and at least a fifth of the entry) admits 3,354 good-day trades against 1,736 bad-day ones, so the two buckets are differently selected. The direction — overnight holding is negative — is solid; the good-versus-bad gap is suggestive, not established. Without that filter the median day-1 close is $0.01 and the ratios become meaningless.

20,018 trades DEADRepeat-alerted tickers are not hotter

flat across streak length once SPX is separated

SPX led 8 of the last 8 sessions and MU appeared in 6 of 8, so persistence in the alert stream is real. It carries no return information. Excluding SPX, sessions-in-a-row buckets of 0 / 1 / 2–4 / 5+ return −9.4%, −10.1%, −6.5% and −8.5%, all CIs overlapping. Within SPX, a fresh streak (+12.8%) and a 5+ streak (+11.9%) are indistinguishable.

So what: A repeat-name filter is not worth building. The apparent pattern is the SPX-versus-single-names split already known, wearing a disguise.

Caveat: Streak barely varies for SPX, which is alerted nearly every session, so the SPX column is a weak test by construction.

Standing method

Every claim needs an effect size, an n, a confidence interval and a stated control — same-ticker, same-day or time-matched. Chronological and interleaved splits are the minimum bar: a curve-fit fails both, a genuine regime change fails only the chronological one. A sealed holdout of 4,016 trades remains unopened and is spent once, on the single best survivor. With many hypotheses in flight the multiple-testing problem is real and is stated on each card rather than assumed away.