RocketShip

The screen lab

Build a screen, test it against 19,051 stock-dates, and see what the pipeline's own stages are worth.

Pairs 600As-of dates 12Universe 533 tickersPanel 19,051 stock-datesLabels fwd return vs SPY

Screen lab

Build a screen. Break it.

Drag the weights and every one of 19,051 stock-dates is re-scored, re-ranked and re-tested in the browser — correlation, confidence interval, decile spread and the names it would have bought, all recomputed on the frame. Nothing is precomputed and nothing is hidden; this is the same data and the same statistics the rest of this page reports.

Start with Momentum only, then try Equal weight. Adding six more factors makes it worse, which is the single most useful thing this project learned.

Weights

+100

12-month return skipping the most recent month

0

1-month return, sign flipped

0

10d/60d average volume

0

60d return volatility, sign flipped

0

annualised slope of log price over 60d

0

distance from the 52-week high (negative)

0

log median dollar volume, sign flipped

Start from
Rank correlation with forward return+0.08095% CI [+0.047, +0.118]separates from zero19,051 stock-dates · 36 months · recomputed live

Mean forward return by score deciletop decile minus bottom: +7.22%

1
2
3
4
5
6
7
8
9
10

1 = highest scored10 = lowest

Month by montha signal that works only once is not a signal

2023-06-152026-05-15

What it would have boughttop 10 on 2026-05-15

  1. ECHO-37.38%
  2. CIEN-24.52%
  3. LITE-4.98%
  4. WDC+6.42%
  5. STX+20.34%
  6. SNDK+22.14%
  7. MU+34.84%
  8. TER+26.40%
  9. COHR-12.96%
  10. ALB-30.32%

How the evaluation works

Freeze

Pick as-of dates in the past. Truncate every price series at that date and refuse to run if anything newer leaks in — news included, verified per article, not trusted to the API.

Score

Run the thing under test on only what was knowable then. Same stocks, same dates, same context for every variant, so any difference is the variant and not the sample.

Wait

Label with realised 1 and 3-month returns, excess of SPY. Excess, because a rising market makes everything look clever.

Doubt

Resample as-of dates, not stocks — names within one month move together. Compare variants paired on identical data, which is far more sensitive than comparing two separate intervals.

The harness is validated on cases with known answers before it is trusted on unknown ones: an arm that just echoes the input score must measure exactly zero new information, and it does — to three decimal places.

Stage A2

Rebuilding the screen: the one thing that worked

Stage A found nothing on 50 mega-caps over 12 dates. The diagnosis was that the universe, the signals and the sample were all wrong. This tests that on 19,051 pairs, 533 tickers and 36 monthly dates - 32x the sample - with sector-neutral z-scores instead of hand-set absolute thresholds.

Score3M rank IC
RocketScore (shipped)+0.024[-0.022, +0.068]inside the noise
12-1 momentum, sector-neutral+0.081[+0.043, +0.119]separates
Difference (paired)+0.057[+0.024, +0.092]separates
The simplest thing winsA fitted seven-factor walk-forward model scores worse than the single momentum factor, and does not separate from zero. Averaging all seven equally is worse still. Adding factors diluted the one that works - the value is in specifying one signal correctly, not in the fitting.
Signal3M rank ICKind
12-1 momentum+0.081[+0.044, +0.118]single factor (exploratory)
1-month reversal+0.005[-0.038, +0.052]single factor (exploratory)
Volume surge+0.004[-0.015, +0.023]single factor (exploratory)
Low volatility-0.043[-0.096, +0.011]single factor (exploratory)
Trend slope+0.011[-0.033, +0.052]single factor (exploratory)
Near 52w high+0.018[-0.025, +0.061]single factor (exploratory)
Illiquidity-0.032[-0.055, -0.006]single factor (exploratory)
All seven, equal weight+0.013[-0.023, +0.047]unfitted composite
Fitted model (walk-forward)+0.042[-0.008, +0.092]fitted, purged walk-forward (out of sample)
Random-0.007[-0.021, +0.006]floor
What I do not believe yetThe universe is current index membership, so delisted names are absent and the survivors are disproportionately the ones that went up - which is exactly what momentum measures. The window is 36 months of a single trending regime, and momentum is documented to crash on reversals. Published momentum ICs sit near 0.02-0.05; getting 0.08 is more consistent with those two biases than with a discovery.

Where value is created

Three stages, each measured against its own baseline. The pipeline runs left to right: every stage's input is the previous one's output, so a stage that adds nothing passes its input through unchanged.

  1. Stage Afree

    The screen

    Does RocketScore rank forward returns?

    +0.013[-0.070, +0.101]
    vs random ranking
    No measurable effect

    600 pairs, 12 as-of dates. Rank correlation with forward excess return.

  2. Stage B$3.44

    The debate

    Does it beat one call, and beat the screen?

    -0.003[-0.067, +0.057]
    vs a single LLM call
    No measurable effect

    Paired difference in rank correlation, 7,200 API calls.

  3. Stage Cfree

    The optimiser

    Does it beat dividing by N?

    +0.16%[-1.29%, +1.44%]
    vs equal weight
    No measurable effect

    Forward total return, covariance fitted only on data before the as-of date.

Every comparison, against zero

1M horizon

Debate vs one call
-0.027[-0.079, +0.024]
inside the noise
Debate vs the screen
-0.034[-0.099, +0.033]
inside the noise
Debate vs random
-0.011[-0.104, +0.093]
inside the noise

3M horizon

Debate vs one call
-0.003[-0.067, +0.057]
inside the noise
Debate vs the screen
-0.013[-0.066, +0.040]
inside the noise
Debate vs random
-0.007[-0.074, +0.061]
inside the noise
-0.128-0.006+0.116

Paired differences in rank correlation with forward excess return. The same stocks, the same dates, and one bootstrap resample plan applied to both arms, so the interval is on the difference rather than on two separately estimated levels. Hatched intervals cross zero.

What each arm costs, and what it knows

Cost sits beside effect because that is the trade being evaluated. “New info beyond the screen” residualises each arm's score on the RocketScore it was handed and correlates the residual with forward return, isolating what the LLM knew that the screen did not.

ArmCost per decisionNew info beyond the screen
Full debate$0.002537.0x-0.016[-0.053, +0.024]
Single call$0.00036-0.013[-0.073, +0.039]
RocketScore onlyfree+0.000[-0.000, +0.000]
Randomfree+0.009[-0.010, +0.029]
Stage A

Does the deterministic screen rank anything?

RocketScore ranks the universe before any LLM runs. Running this first is deliberate: if the screen carries no signal, the debate is being asked to add value on top of noise. Costs nothing, so it is never constrained by budget: 600 pairs over 12 as-of dates.

inside the noise
Score1M rank correlation3M rank correlation
rocket+0.026[-0.096, +0.146]+0.013[-0.070, +0.101]
technical+0.018[-0.103, +0.135]+0.006[-0.074, +0.087]
volume+0.048[-0.033, +0.130]+0.035[-0.038, +0.109]
quality----
macro+0.028[-0.085, +0.131]+0.004[-0.134, +0.126]
The weights are not the weightsA component with no cross-sectional variance cannot move a ranking, whatever weight the config assigns it.
ComponentAdvertisedActual influence
technical45%75.4%
volume25%20.1%
quality20%0.0%
macro10%4.5%

The tag bonus moves rank correlation by -0.000[-0.007, +0.007] — a tight interval around zero, not an inconclusive one.

Stage B

Does the debate beat one call?

The only stage that spends money. RocketScore only is free and is the baseline that matters: the debate is handed the RocketScore and its rank inside its own context, so if it cannot beat “use the number you were given”, the four extra calls are decoration.

Arm1M corr.Per seed3M corr.New info vs screen (3M)Calls$/decisionLatency
Full debateBull, bear, regime and value agents in parallel, then a judge. Five calls.-0.010±0.016
-0.001±0.005-0.016[-0.053, +0.024]5$0.0025311.7s
Single callOne call carrying the same four lenses and the same decision rule.+0.017±0.018
+0.001±0.004-0.013[-0.073, +0.039]1$0.000364.8s
RocketScore onlyThe deterministic screen, which the debate is handed inside its own context.+0.026+0.013+0.000[-0.000, +0.000]0free0.0s
RandomUniform random ranking. The noise floor, over eight seeds.+0.003±0.040
+0.007±0.033+0.009[-0.010, +0.029]0free0.0s
What this actually tells youThe debate is handed the score in its own context, so the question is not whether it ranks — it is whether it ranks better than the number it started with. Residualising its output on that score isolates what the LLM contributed. It comes out at zero, and that points somewhere specific: the agents were given a table of floats and asked to forecast returns, which is the task language models are worst at. The fix is not a better debate, it is a different job — reading filings and earnings calls into features a model can rank. That is the next experiment, and this harness is what will judge it.
Stage C

Does the optimiser beat dividing by N?

top 12 by RocketScore, covariance fitted only on data ending at the as-of date, evaluated on realised forward returns. The shipped backtest does none of that — it replays the same window the selection and the covariance both came from.

inside the noise
HorizonOptimiserEqual weightDifference
1M+3.20%[+1.45%, +5.04%]+3.18%[+1.17%, +5.34%]+0.02%[-0.80%, +0.94%]
3M+6.31%[+3.70%, +9.14%]+6.15%[+3.60%, +9.00%]+0.16%[-1.29%, +1.44%]
The look-ahead premiumThe product's own in-sample framing reports Sharpe +2.12. The honest forward Sharpe on identical weights is +1.48. The gap is what backtesting on your own selection window buys you.

What would change my mind

  • Twelve as-of dates is still few. Going from four dates to twelve flipped the sign of two separate estimates in this project, including the one that had looked closest to significant. A longer history is the highest-value next step.
  • The screen it builds on has no signal either. The debate is being asked to add value on top of noise, over fifty mega-caps where edge is hard to find by construction.
  • Training-data contamination biases every arm upward, not down. The model has seen outcomes for dates this recent. These are ceilings.
  • Fundamentals are not point-in-time, so the quality component is pinned to neutral, removing a fifth of the score's nominal weight. Identical across arms, so comparisons hold; absolute scores do not match production.

The honest reading is not “the debate is worthless” but “it is not measurably better, and the burden of proof was on it”.