2026-07-30 to 2026-08-28 · 22 sessions · six index and ETF instruments
Dealer gamma regime does not separate our highest-conviction reads
A widely-held idea says directional signals should be trusted less when dealer positioning is in its dampening state. Across a month of live reads the dampened group performed slightly better, not worse.
In simple words
What happened
Two of the filter’s three own conditions fail, and the one that decides the matter points the wrong way. When it was written the regime reading was stuck on a single value for every read, so the filter carried no information at all; a later fix to how dealer gamma is signed made both states reachable for the first time, which is what finally made this testable. The filter is currently switched on, and it has downgraded 136 reads.
Who would pay?
The theory was that when dealer positioning is in its "dampening" state, big directional calls should be trusted less, so the filter strips them of top billing. That is a reasonable-sounding story. It is also testable, and the filter was written with its own explicit re-arm conditions: the regime has to actually vary, the dampened group has to be materially worse, and there have to be at least 30 independent episodes on each side.
How it was tested
Every max-conviction read was graded on the move of the thing it actually names — the index itself — from the price at the moment of the call to the closing bell, and again at a 60-minute horizon. The read names no strike, no expiry and no exit rule, so grading a specific option contract would have smuggled three unstated choices into the answer. Repeats were collapsed to one observation per name per day per direction, because the engine re-emits a read every cycle and counting rows instead of episodes is the fastest way to reach a confident wrong conclusion.
What would prove it wrong?
The filter only ever downgrades — it never creates a signal and never touches the other regime. So the entire question is whether the downgraded group is worse. It is not. On the wider of two possible groupings — the one deliberately built to flatter the filter, because it keeps downgraded reads that other rules later removed while the comparison group only holds reads that survived everything — the downgraded cohort still won more often (45.0% vs 37.4%) with an identical average (-0.067% vs -0.067%).
Leader-collapsed evidence
What each setup actually did
510 max-conviction reads carrying a dealer-regime reading; 119 independent episodes after collapsing repeats within a name, day, direction and regime
| Group | Episodes | Win rate | Lower bound | Mean to close | Median to close | Decision |
|---|---|---|---|---|---|---|
| DOWNGRADEDReads the filter demotes (dampening regime) | 120 over 16 sessions | -0.067% | -0.046% | 45.0% | Lower bound 36.4% | NEGATIVE EVIDENCE |
| UNTOUCHEDReads the filter never touches (other regime) | 99 | -0.067% | -0.069% | 37.4% | Lower bound 28.5% | NEGATIVE EVIDENCE |
| STRICTStrictly-comparable subset (before the filter was armed) | 20 over 2 sessions | -0.130% | +0.017% | 50.0% | vs 37.4% untouched | UNRUNNABLE |
What we keep
Learnings
- 01
A filter that acts on a signal changes the population you can measure it against. Because this one removes top billing, the reads it downgrades stop appearing in the group you would compare — so the comparison has to be built from a marker recorded BEFORE the downgrade, not from the outcome. Record the decision you did not take, at the moment you take the other one, or the question stops being answerable.
- 02
A filter that reads as sensible can still be measuring nothing. Before this study, the regime reading was the same value on 100% of reads — a setting that fires on everything carries exactly as much information as one that fires on nothing.
- 03
Grade the instrument the signal actually names. This read names an index, not a contract, so it is graded on the index. Reaching for an option would have added a contract choice, an entry basis and an exit rule that the signal never specified, and any of the three can flip the sign of the answer.
- 04
When you deliberately build the comparison to favour the thing you are testing and it still fails, that is a stronger result than a fair test that fails narrowly.