Build a screen, test it against 19,051 stock-dates, and see what the pipeline's own stages are worth.
Screen lab
Drag the weights and every one of 19,051 stock-dates is re-scored, re-ranked and re-tested in the browser — correlation, confidence interval, decile spread and the names it would have bought, all recomputed on the frame. Nothing is precomputed and nothing is hidden; this is the same data and the same statistics the rest of this page reports.
Start with Momentum only, then try Equal weight. Adding six more factors makes it worse, which is the single most useful thing this project learned.
1 = highest scored10 = lowest
2023-06-152026-05-15
Pick as-of dates in the past. Truncate every price series at that date and refuse to run if anything newer leaks in — news included, verified per article, not trusted to the API.
Run the thing under test on only what was knowable then. Same stocks, same dates, same context for every variant, so any difference is the variant and not the sample.
Label with realised 1 and 3-month returns, excess of SPY. Excess, because a rising market makes everything look clever.
Resample as-of dates, not stocks — names within one month move together. Compare variants paired on identical data, which is far more sensitive than comparing two separate intervals.
The harness is validated on cases with known answers before it is trusted on unknown ones: an arm that just echoes the input score must measure exactly zero new information, and it does — to three decimal places.
Stage A found nothing on 50 mega-caps over 12 dates. The diagnosis was that the universe, the signals and the sample were all wrong. This tests that on 19,051 pairs, 533 tickers and 36 monthly dates - 32x the sample - with sector-neutral z-scores instead of hand-set absolute thresholds.
| Score | 3M rank IC | |
|---|---|---|
| RocketScore (shipped) | +0.024[-0.022, +0.068] | inside the noise |
| 12-1 momentum, sector-neutral | +0.081[+0.043, +0.119] | separates |
| Difference (paired) | +0.057[+0.024, +0.092] | separates |
| Signal | 3M rank IC | Kind |
|---|---|---|
| 12-1 momentum | +0.081[+0.044, +0.118] | single factor (exploratory) |
| 1-month reversal | +0.005[-0.038, +0.052] | single factor (exploratory) |
| Volume surge | +0.004[-0.015, +0.023] | single factor (exploratory) |
| Low volatility | -0.043[-0.096, +0.011] | single factor (exploratory) |
| Trend slope | +0.011[-0.033, +0.052] | single factor (exploratory) |
| Near 52w high | +0.018[-0.025, +0.061] | single factor (exploratory) |
| Illiquidity | -0.032[-0.055, -0.006] | single factor (exploratory) |
| All seven, equal weight | +0.013[-0.023, +0.047] | unfitted composite |
| Fitted model (walk-forward) | +0.042[-0.008, +0.092] | fitted, purged walk-forward (out of sample) |
| Random | -0.007[-0.021, +0.006] | floor |
Three stages, each measured against its own baseline. The pipeline runs left to right: every stage's input is the previous one's output, so a stage that adds nothing passes its input through unchanged.
Does RocketScore rank forward returns?
600 pairs, 12 as-of dates. Rank correlation with forward excess return.
Does it beat one call, and beat the screen?
Paired difference in rank correlation, 7,200 API calls.
Does it beat dividing by N?
Forward total return, covariance fitted only on data before the as-of date.
Paired differences in rank correlation with forward excess return. The same stocks, the same dates, and one bootstrap resample plan applied to both arms, so the interval is on the difference rather than on two separately estimated levels. Hatched intervals cross zero.
Cost sits beside effect because that is the trade being evaluated. “New info beyond the screen” residualises each arm's score on the RocketScore it was handed and correlates the residual with forward return, isolating what the LLM knew that the screen did not.
RocketScore ranks the universe before any LLM runs. Running this first is deliberate: if the screen carries no signal, the debate is being asked to add value on top of noise. Costs nothing, so it is never constrained by budget: 600 pairs over 12 as-of dates.
inside the noise| Score | 1M rank correlation | 3M rank correlation |
|---|---|---|
| rocket | +0.026[-0.096, +0.146] | +0.013[-0.070, +0.101] |
| technical | +0.018[-0.103, +0.135] | +0.006[-0.074, +0.087] |
| volume | +0.048[-0.033, +0.130] | +0.035[-0.038, +0.109] |
| quality | -- | -- |
| macro | +0.028[-0.085, +0.131] | +0.004[-0.134, +0.126] |
| Component | Advertised | Actual influence |
|---|---|---|
| technical | 45% | 75.4% |
| volume | 25% | 20.1% |
| quality | 20% | 0.0% |
| macro | 10% | 4.5% |
The tag bonus moves rank correlation by -0.000[-0.007, +0.007] — a tight interval around zero, not an inconclusive one.
The only stage that spends money. RocketScore only is free and is the baseline that matters: the debate is handed the RocketScore and its rank inside its own context, so if it cannot beat “use the number you were given”, the four extra calls are decoration.
| Arm | 1M corr. | Per seed | 3M corr. | New info vs screen (3M) | Calls | $/decision | Latency |
|---|---|---|---|---|---|---|---|
| Full debateBull, bear, regime and value agents in parallel, then a judge. Five calls. | -0.010±0.016 | -0.001±0.005 | -0.016[-0.053, +0.024] | 5 | $0.00253 | 11.7s | |
| Single callOne call carrying the same four lenses and the same decision rule. | +0.017±0.018 | +0.001±0.004 | -0.013[-0.073, +0.039] | 1 | $0.00036 | 4.8s | |
| RocketScore onlyThe deterministic screen, which the debate is handed inside its own context. | +0.026 | +0.013 | +0.000[-0.000, +0.000] | 0 | free | 0.0s | |
| RandomUniform random ranking. The noise floor, over eight seeds. | +0.003±0.040 | +0.007±0.033 | +0.009[-0.010, +0.029] | 0 | free | 0.0s |
top 12 by RocketScore, covariance fitted only on data ending at the as-of date, evaluated on realised forward returns. The shipped backtest does none of that — it replays the same window the selection and the covariance both came from.
inside the noise| Horizon | Optimiser | Equal weight | Difference |
|---|---|---|---|
| 1M | +3.20%[+1.45%, +5.04%] | +3.18%[+1.17%, +5.34%] | +0.02%[-0.80%, +0.94%] |
| 3M | +6.31%[+3.70%, +9.14%] | +6.15%[+3.60%, +9.00%] | +0.16%[-1.29%, +1.44%] |
The honest reading is not “the debate is worthless” but “it is not measurably better, and the burden of proof was on it”.