2026-07-30 to 2026-08-28 · 13 sessions with a disagreement
Repointing our flow read at the raw tape would have made it worse
One vivid morning suggested the raw aggressor tape was the truer source. Across fifty-three independent disagreements the proposed replacement is the weaker of the two, so the swap is refused.
In simple words
What happened
The repoint is refused: it would swap a coin flip for a slightly worse coin flip. The wider result is more useful than the narrow one — on the 306 agreement episodes both signals score 51% at an hour and 49% to the close. In this window the flow input is not carrying directional information at all, which reframes the question from "which source" to "why is this an input".
Who would pay?
On one memorable morning the fused flow vote leaned strongly bullish off a small set of curated rows while the raw tape carried $30M of calls trading on the bid — the opposite reading. If the tape is the truer source, pointing our flow input at it should improve the read.
How it was tested
Both numbers — the fused flow vote the read actually acted on, and the signed, aggression-weighted tape over the same window — were recorded side by side at the moment of every read, specifically so this question would survive the short retention that had made a retrospective version impossible. Each is scored on whether its sign matched the index’s subsequent move, at 60 minutes and to the close. Repeats within a name, day and direction are collapsed.
What would prove it wrong?
The decisive rows are the ones where the two disagree, because everywhere else the choice makes no difference. On those 53 episodes the tape is right 47.2% of the time at 60 minutes against the flow vote’s 52.8% — the proposed replacement is worse.
Leader-collapsed evidence
What each setup actually did
8,940 graded reads; 175 raw disagreements collapsing to 53 independent episodes, against an agreement control of 306
| Cohort | Episodes | Flow vote correct | Tape correct | Flow lower bound | Tape lower bound | Decision |
|---|---|---|---|---|---|---|
| CLASH60Disagreement rows, 60-minute horizon | 53 over 13 sessions | 52.8% | 47.2% | 39.7% | 34.4% | REFUSED |
| CLASHCLDisagreement rows, to the close | 53 | 54.7% | 45.3% | 41.5% | 32.7% | REFUSED |
| AGREEAgreement rows (control), 60-minute horizon | 306 | 51.0% | 51.0% | 45.4% | 45.4% | REFUSED |
What we keep
Learnings
- 01
One vivid disagreement is an anecdote. It was a real, correctly-described disagreement, and it pointed at the wrong fix — the source it favoured turned out to be the weaker sign.
- 02
Score the alternative on exactly the rows where the choice matters. Pooling in the cases where both signals agree would have buried a 5.6-point difference under 306 rows of noise.
- 03
Include the agreement rows as a control anyway. They are what turned "we picked the wrong source" into the more important "neither source predicts here" — and only the second one is actionable.
- 04
Recording both numbers at the moment of decision is what made this answerable. A later reconstruction had already been tried and collapsed from 130 candidate episodes to 5 usable ones because the raw inputs had aged out.
- 05
The remaining option — removing the input entirely — is supported by this evidence but is a behaviour change, and is being taken as its own decision rather than smuggled in on the back of this result.