Research library

2026-07-29 to 2026-08-28 · 22 sessions

A rate-of-change filter for large prints selected the losing half

Volume thresholds on a cumulative total are really a statement about the time of day, so a rate-of-change test looks like the obvious upgrade. Run silently for a month alongside the live rule, it picked the worse prints.

VerdictREFUSED · stays in shadow

In simple words

What happened

Switching the rate test on would have removed the better half of the lane. A day-block resampling of the difference gives a 95% range of -55 to +9 percentage points, so this is not proof the filter is actively harmful — but it is nowhere near clearing a bar to adopt it, and the estimate sits squarely on the wrong side of zero.

Who would pay?

The original trigger tests cumulative volume on a strike against a fixed floor. Cumulative volume only goes up, so a fixed floor is a first-crossing test — it fires on whichever strike happens to cross first, which is a statement about the time of day, not about urgency. Measuring the scan-to-scan change instead should identify a genuine burst.

How it was tested

For a month the rate test was computed and recorded on every qualifying print without ever being allowed to change what fired. That makes the comparison exact rather than reconstructed: switching it on would keep the prints it marked as qualifying and drop the ones it measured and rejected, so the two groups are simply read off the record.

What would prove it wrong?

The prints the rate test would have DROPPED were the better ones, on both measures: they hit more often (70.0% vs 47.4%) and their median peak was more than twice as large (+135% vs +56%).

Leader-collapsed evidence

What each setup actually did

89 shadow-marked prints where a rate could actually be measured; 39 independent episodes in the affected lane after collapsing by name, day and direction

GroupEpisodesHit rateLower boundMedian peakWhat switching on doesDecision
KEPTPrints the rate test marks as qualifying1947.4%+55.9%27.3%These would be keptREFUSED
DROPPEDPrints it measured and rejected2070.0%+135.2%48.1%These would be discardedREFUSED

What we keep

Learnings

  1. 01

    The original criticism was correct and the replacement still failed. Diagnosing a trigger as a clock does not mean the obvious alternative is better; the burst you can measure is not necessarily the burst that matters.

  2. 02

    Shadow-running earns its keep here. The rule was written, believed, and would have been switched on after two weeks on the strength of the argument alone. Four weeks of recording it without letting it act cost nothing and produced the opposite answer.

  3. 03

    Publish the honest range, not just the point estimate. The difference is -22.6 percentage points but the resampled range crosses zero — the correct summary is "does not clear the bar", not "proven harmful".

  4. 04

    This is the second time a tightened filter in this system has selected for the losing side. A narrower cohort feels like higher quality and often is not.

The rate test keeps running in shadow — computed and recorded on every print, never allowed to change what fires. It costs nothing and the record keeps growing, so the decision can be revisited on new data rather than re-argued from the same month.
Back to all studies