Can LLM agents do quantitative finance research? Two capabilities, measured
separately on private data: execution (replicate a result from a written method
spec + raw data) and judgment (predict returns from financial text). 20+ frontier
models, one identical harness, one inference platform, script-graded.
29frontier models, one identical harness, 220 graded agent runs
23/29models replicate a momentum study perfectly from spec — execution is commoditized
IC 0.61best model reading post-cutoff analyst articles; withheld-text controls collapse to ≈0
Key findings
Execution is commoditized. Given a precise method spec and raw data,
nearly every model — including sub-cent-per-run ones — perfectly replicates
standard quantitative studies (momentum portfolios, event studies, an
option-overlay backtest, an 11GB full-market study). See T1–T4.
The frontier appears at paper scale. Replicating a full
machine-learning asset-pricing study end-to-end (T6: train four model classes,
annual refits, out-of-sample portfolio evaluation) splits the field roughly in
half. Failures are precision failures — every model produces plausible
magnitudes; only some hold twenty-plus spec details simultaneously.
Reading long financial text yields real, measurable judgment. On
analyst articles published after training cutoffs, flagship models reach
Spearman ICs around 0.4–0.5 vs ≈0.2 for small models; with the text withheld,
every model collapses to ≈0. See B2.
On historical data, apparent skill is often memorization. Some
models predict historical event outcomes better with the news withheld —
they recall two decades of price history from (ticker, date) alone. The honest
metric is ΔIC = IC(text) − IC(withheld). See B1.
Models rank better than they call direction. Several models show
significantly positive ICs with below-50% hit rates: a systematic long bias.
Use these signals cross-sectionally, not for market timing.
Reliability, quantified. Re-running identical tasks: ~2–3% of runs
silently produce wrong answers and ~6% die on provider infrastructure —
consistent with what production-agent practitioners report as their top
challenge.
Published ML alpha decays. Re-running a canonical ML asset-pricing
design on 2015–2021 (after the original out-of-sample period), decile
long-short Sharpe falls from 2+ to ≈0.2–0.3 and linear signals flip negative;
the nonlinear-beats-linear ranking survives.
Methods in brief
One harness, one platform. Every model runs the same two-tool agent
loop (bash sandbox + answer submission, step-capped) through one inference
platform; provider-specific API quirks are normalized at the gateway layer and
documented.
Script-graded, private ground truth. Track A answers are graded by
tolerance bands (loose ±5%, strict ±1%) against reference implementations run
on private data extracts — published numbers can't be recalled from training.
Controls everywhere. Every text-prediction track has a
withheld-text control run per model, bounding memorization; Track B2
additionally uses only articles published after model training cutoffs, with
per-model cutoff verification.
Replicates. Reliability numbers come from repeated identical runs
(up to 5 seeds); leaderboard metrics are being extended with multi-seed error
bars.
Track A — Execution: replicate from spec
The agent gets a task sheet (method spelled out, results withheld), read-only data, a bash sandbox, and a 40-step budget. Grading is scripted: each submitted statistic vs ground truth (loose ±5% / strict ±1%). Samples are private extracts, so published numbers can't be recalled from training.
T1 · 20-stock momentum
Model
Strict
Loose
Consistency (strict across seeds)
Best steps
Best cost
Claude Haiku 4.5
1.0
1.0
—
6
$0.0668
Claude Opus 4.8
1.0
1.0
—
2
$0.0615
Claude Opus 5
1.0
1.0
5/5
2
$0.0553
Claude Sonnet 5
1.0
1.0
5/5
1
$0.0188
DeepSeek V3.2
1.0
1.0
—
15
$0.0422
DeepSeek V4 Flash
1.0
1.0
5/5
4
$0.0038
DeepSeek V4 Pro
1.0
1.0
5/5
3
$0.0297
GLM 5.2
1.0
1.0
4/5
2
$0.0154
GPT-5.4
1.0
1.0
5/5
1
$0.0193
GPT-5.5
1.0
1.0
—
3
$0.1473
GPT-5.6 Sol
1.0
1.0
5/5
1
$0.0269
GPT-5.6 Terra
1.0
1.0
—
1
$0.0155
Gemini 3.1 Pro
1.0
1.0
5/5
2
$0.0775
Gemini 3.6 Flash
1.0
1.0
—
28
$1.0086
Gemma 4 31B
1.0
1.0
5/5
1
$0.0010
Grok 4.5
1.0
1.0
5/5
2
$0.0204
Hunyuan 3
1.0
1.0
—
2
$0.0015
Kimi K2.6
1.0
1.0
—
2
$0.0132
Kimi K3
1.0
1.0
5/5
2
$0.0270
MiMo V2.5
1.0
1.0
5/5
2
$0.0013
MiniMax M3
1.0
1.0
—
6
$0.0264
MiniMaxAI/MiniMax-M2.7
1.0
1.0
—
7
$0.0182
Qwen3.8 Max
1.0
1.0
4/5
3
$0.0294
Nemotron 3 Ultra
0.0
0.0
—
—
—
Qwen/Qwen3.6-Plus
0.0
0.0
—
—
—
bytedance/seed-2.0-mini
0.0
0.0
—
—
—
kwaipilot/kat-coder-pro-v2
0.0
0.0
—
—
—
moonshotai/Kimi-K2.5
0.0
0.0
—
—
—
zai-org/GLM-4.7-FP8
0.0
0.0
—
—
—
T2 · News event study
Model
Strict
Loose
Consistency (strict across seeds)
Best steps
Best cost
Claude Opus 5
1.0
1.0
5/5
2
$0.0607
Claude Sonnet 5
1.0
1.0
5/5
1
$0.0180
DeepSeek V4 Flash
1.0
1.0
5/5
5
$0.0037
DeepSeek V4 Pro
1.0
1.0
5/5
4
$0.0470
GLM 5.2
1.0
1.0
5/5
3
$0.0205
GPT-5.4
1.0
1.0
5/5
1
$0.0190
GPT-5.6 Sol
1.0
1.0
5/5
1
$0.0285
Gemini 3.1 Pro
1.0
1.0
5/5
2
$0.0553
Gemma 4 31B
1.0
1.0
5/5
1
$0.0008
Grok 4.5
1.0
1.0
5/5
2
$0.0220
Kimi K3
1.0
1.0
5/5
2
$0.0330
MiMo V2.5
1.0
1.0
4/5
2
$0.0011
Qwen3.8 Max
1.0
1.0
5/5
2
$0.0339
T3 · Covered-call option backtest
Model
Strict
Loose
Consistency (strict across seeds)
Best steps
Best cost
Claude Opus 5
1.0
1.0
—
2
$0.1268
Claude Sonnet 5
1.0
1.0
—
10
$0.2537
DeepSeek V4 Flash
1.0
1.0
—
32
$0.1155
DeepSeek V4 Pro
1.0
1.0
—
21
$0.7153
GLM 5.2
1.0
1.0
—
14
$0.3479
GPT-5.4
1.0
1.0
—
7
$0.1813
Gemini 3.1 Pro
1.0
1.0
—
16
$0.7649
Gemma 4 31B
1.0
1.0
—
6
$0.0109
Grok 4.5
1.0
1.0
—
5
$0.1129
Kimi K3
1.0
1.0
—
4
$0.2244
MiMo V2.5
1.0
1.0
—
11
$0.0226
Qwen3.8 Max
1.0
1.0
—
13
$0.3382
GPT-5.6 Sol
0.6
0.7
—
2
$0.0718
T4 · Full-market momentum (11GB)
Model
Strict
Loose
Consistency (strict across seeds)
Best steps
Best cost
Claude Opus 5
1.0
1.0
—
5
$0.1968
Claude Sonnet 5
1.0
1.0
—
11
$0.2056
DeepSeek V4 Pro
1.0
1.0
—
13
$0.3799
GLM 5.2
1.0
1.0
—
26
$0.3970
Gemma 4 31B
1.0
1.0
—
5
$0.0054
Grok 4.5
1.0
1.0
—
3
$0.0381
Kimi K3
1.0
1.0
—
11
$0.2920
MiMo V2.5
1.0
1.0
—
5
$0.0073
Qwen3.8 Max
1.0
1.0
—
19
$0.7046
GPT-5.4
0.8
1.0
—
1
$0.0300
GPT-5.6 Sol
0.8
0.8
—
5
$0.1065
Gemini 3.1 Pro
0.4
1.0
—
30
$1.0432
DeepSeek V4 Flash
0.2
0.4
—
20
$0.0372
T6 · Replicate an ML asset-pricing study (RFS-style, 4 models trained)
Model
Strict
Loose
Consistency (strict across seeds)
Best steps
Best cost
Claude Opus 5
1.0
1.0
1/3
7
$0.5712
Claude Sonnet 5
1.0
1.0
3/3
5
$0.1292
DeepSeek V3.2
1.0
1.0
—
45
$0.2838
DeepSeek V4 Flash
1.0
1.0
1/3
28
$0.0669
DeepSeek V4 Pro
1.0
1.0
2/3
9
$0.1135
GLM 5.2
1.0
1.0
3/3
12
$0.1919
Gemini 3.6 Flash
1.0
1.0
—
22
$0.9978
Grok 4.5
1.0
1.0
2/3
5
$0.1123
Kimi K3
1.0
1.0
3/3
13
$0.5373
MiMo V2.5
1.0
1.0
1/3
15
$0.0271
MiniMax M3
1.0
1.0
—
14
$0.1622
Qwen3.8 Max
1.0
1.0
3/3
15
$0.4760
GPT-5.6 Sol
0.5
0.8
0/3
10
$0.3244
Claude Haiku 4.5
0.4
0.5
—
8
$0.0991
Claude Opus 4.8
0.4
0.5
—
9
$0.5695
Gemini 3.1 Pro
0.4
0.5
0/3
12
$0.3600
GPT-5.6 Terra
0.4
0.4
—
10
$0.2201
Hunyuan 3
0.1
0.5
—
16
$0.0406
GPT-5.4
0.1
0.2
0/3
5
$0.1480
Gemma 4 31B
0.0
0.4
0/3
9
$0.0139
GPT-5.5
0.0
0.1
—
11
$1.4065
Kimi K2.6
0.0
0.0
—
—
—
Track B — Judgment: predict returns from text
B1 · News headlines
1,008 stratified news events on 20 mega-caps (2004–2024) from an institutional news-event dataset. Model sees headline + ticker + date + trailing stats; predicts 5-day beta-adjusted return. Control: headline withheld (memorization baseline).
Model
n
Hit rate
Spearman IC
L/S spread (bps)
Control IC
ΔIC (text − control)
Claude Opus 5
1008
0.6713
0.522
973.7
0.5326
-0.0106
Kimi K3
971
0.5637
0.2413
593.1
0.2157
0.0256
Claude Opus 4.8
1008
0.5733
0.2299
550.6
0.1626
0.0673
Gemini 3.1 Pro
1008
0.5712
0.1991
548.6
0.2281
-0.029
Qwen3.8 Max
1000
0.5612
0.1658
344.2
0.179
-0.0132
Claude Sonnet 5
1008
0.5481
0.1435
425.9
0.1187
0.0248
GLM 5.2
919
0.5531
0.1368
345.2
0.0438
0.093
GPT-5.6 Sol
1008
0.5327
0.1324
271.3
0.1643
-0.0319
DeepSeek V4 Pro
632
0.5768
0.0992
117.9
0.0445
0.0547
DeepSeek V4 Flash
986
0.528
0.052
103.8
0.0236
0.0284
Grok 4.5
1008
0.5257
0.0442
136.6
0.001
0.0432
GPT-5.4
1008
0.5051
0.039
72.5
0.0298
0.0092
Gemma 4 31B
1008
0.5154
0.0377
43.6
0.0362
0.0015
MiMo V2.5
1006
0.5082
0.0145
77.6
0.0324
-0.0179
B2 · Analyst articles (post-cutoff)
Full-text investment research articles published Apr–Aug 2026 — after the training cutoff of the models under test. Model predicts 20-day market-adjusted return. Control: article withheld.
Model
n
Hit rate
Spearman IC
IC (Jun–Jul only)
Cutoff
L/S spread (bps)
Control IC
Claude Opus 5
297
0.5118
0.610 ±0.028 (2)
0.6271
2026-05 ⚠️
4447.3
-0.0261
GPT-5.5
292
0.476
0.449 ±0.027 (2)
0.562
UNDISCLOSED ⚠️
2759.0
0.0376
Qwen3.8 Max
296
0.4881
0.446 ±0.025 (3)
0.6027
UNDISCLOSED ⚠️
2687.2
Kimi K3
279
0.482
0.417 ±0.049 (2)
0.481
UNDISCLOSED ⚠️
2773.4
0.2397
Claude Sonnet 5
297
0.4678
0.415 ±0.024 (3)
0.4161
2026-01
2198.1
-0.1583
google_gemini-3.5-flash
296
0.449
0.368
0.4266
UNDISCLOSED ⚠️
2865.3
0.0491
google_gemini-3-flash-preview
297
0.4561
0.3425
0.3345
UNDISCLOSED ⚠️
1737.3
0.0345
GPT-5.6 Sol
297
0.4411
0.3251
0.3876
2026-02
2558.4
-0.1052
anthropic_claude-sonnet-4.6
297
0.4646
0.3174
0.3319
UNDISCLOSED ⚠️
1917.4
-0.0876
Claude Opus 4.6
297
0.431
0.2846
0.3942
UNDISCLOSED ⚠️
1775.9
0.1571
Gemini 3.6 Flash
297
0.4527
0.2839
0.2685
UNDISCLOSED ⚠️
1845.6
0.1131
Claude Opus 4.8
297
0.4257
0.2825
0.3092
UNDISCLOSED ⚠️
2864.3
Claude Opus 4.7
297
0.4291
0.2745
0.2531
UNDISCLOSED ⚠️
2355.2
-0.1365
Qwen_Qwen3.7-Max
297
0.4414
0.2657
0.2784
UNDISCLOSED ⚠️
2239.2
Claude Opus 4.5
297
0.4242
0.2592
0.3221
UNDISCLOSED ⚠️
1254.0
0.1184
MiniMax M3
293
0.4418
0.258
0.2526
UNDISCLOSED ⚠️
2048.9
0.0151
GPT-5.4
297
0.4276
0.2579
0.2797
2025-08
1873.3
-0.0174
GPT-5.6 Terra
297
0.4291
0.2414
0.2072
UNDISCLOSED ⚠️
1813.3
0.2342
anthropic_claude-sonnet-4.5
297
0.4426
0.217
0.1823
UNDISCLOSED ⚠️
1589.6
Grok 4.5
297
0.4257
0.207 ±0.011 (2)
0.2718
2026-02
1096.4
MiniMaxAI_MiniMax-M2.7
266
0.4264
0.2026
0.2001
UNDISCLOSED ⚠️
920.1
-0.1788
Kimi K2.6
264
0.4275
0.2022
0.163
UNDISCLOSED ⚠️
2285.2
0.0483
Gemini 3.1 Pro
297
0.4305
0.1941
0.1713
2025-01
1701.2
MiMo V2.5
297
0.4276
0.192 ±0.007 (2)
0.2055
2024-12
1072.5
0.0408
zai-org_GLM-5-FP8
276
0.4197
0.1832
0.2132
UNDISCLOSED ⚠️
1260.9
-0.0697
DeepSeek V4 Pro
142
0.4892
0.1776
0.2043
UNDISCLOSED ⚠️
222.8
-0.0639
DeepSeek V4 Flash
292
0.4316
0.168 ±0.053 (2)
0.2357
UNDISCLOSED ⚠️
1416.4
-0.0293
moonshotai_Kimi-K2.5
297
0.4242
0.1644
0.2038
UNDISCLOSED ⚠️
548.6
-0.1331
google_gemini-3.5-flash-lite
297
0.4122
0.1632
0.2079
UNDISCLOSED ⚠️
1006.8
GLM 5.2
282
0.4107
0.155 ±0.047 (2)
0.2317
UNDISCLOSED ⚠️
1648.6
0.0238
deepseek-ai_DeepSeek-V3-0324
297
0.4175
0.1536
0.1641
UNDISCLOSED ⚠️
1048.3
-0.0123
Claude Haiku 4.5
297
0.4377
0.149
0.1444
UNDISCLOSED ⚠️
782.2
Hunyuan 3
278
0.417
0.1423
0.1164
UNDISCLOSED ⚠️
1044.3
0.0067
deepseek-ai_DeepSeek-R1-0528
283
0.4413
0.1338
0.123
UNDISCLOSED ⚠️
1098.4
-0.2132
DeepSeek V3.2
279
0.3943
0.1292
0.1567
UNDISCLOSED ⚠️
966.7
0.0534
Gemma 4 31B
297
0.415
0.115 ±0.024 (2)
0.1251
2025-01
747.9
B4 · Earnings press releases (post-cutoff)
8-K item-2.02 earnings press releases filed with the SEC Feb–Jul 2026 (after most models' training cutoffs), fetched from EDGAR. Model reads the release and predicts the 5-day post-filing market-adjusted return. Control: text withheld.
Model
n
Hit rate
Spearman IC
L/S spread (bps)
Control IC
Claude Sonnet 5
96
0.5495
0.119 ±0.075 (2)
116.1
0.0407
Claude Opus 5
100
0.52
0.093 ±0.020 (2)
-197.2
0.2066
Kimi K2.6
97
0.5224
0.0419
-46.1
0.0906
Claude Opus 4.8
99
0.4941
0.0417
-94.9
0.0571
Kimi K3
96
0.5806
0.0382
-601.4
-0.0904
DeepSeek V3.2
97
0.4839
0.0038
564.2
0.1773
GPT-5.6 Sol
100
0.5231
0.0006
-334.0
-0.171
GPT-5.4
100
0.5102
-0.0218
-25.5
0.0675
MiniMax M3
100
0.4615
-0.0479
-124.0
0.1595
DeepSeek V4 Flash
96
0.46
-0.051 ±0.150 (2)
-691.3
0.053
GPT-5.5
100
0.4821
-0.0825
-170.8
-0.0039
Gemini 3.6 Flash
100
0.459
-0.085
-350.3
-0.0441
MiMo V2.5
99
0.4833
-0.0855
-458.3
-0.0056
Qwen3.8 Max
100
0.4783
-0.1047
-492.6
Grok 4.5
100
0.4259
-0.105 ±0.054 (2)
-716.5
0.0052
GPT-5.6 Terra
100
0.4286
-0.1058
187.8
0.0235
Gemini 3.1 Pro
100
0.46
-0.1114
-817.8
-0.1402
DeepSeek V4 Pro
78
0.4643
-0.1173
-544.6
-0.0912
Claude Haiku 4.5
99
0.4699
-0.1238
-417.5
GLM 5.2
88
0.4375
-0.1358
-594.6
-0.0791
Hunyuan 3
75
0.4565
-0.1439
-828.4
Gemma 4 31B
100
0.3913
-0.1916
-1049.5
-0.1419
B6 · Earnings-call transcripts (post-cutoff)
Full earnings-call transcripts (Feb–Jul 2026). Model reads management remarks + Q&A and predicts the 5-day post-call market-adjusted return. Control: transcript withheld.
Model
n
Hit rate
Spearman IC
L/S spread (bps)
Control IC
Claude Opus 5
99
0.5152
0.0033
-261.5
0.1585
Claude Sonnet 5
99
0.5253
0.0016
329.8
-0.2382
GLM 5.2
89
0.5169
0.0003
-272.3
-0.0163
Kimi K3
99
0.5051
-0.0267
-71.5
GPT-5.5
86
0.5233
-0.0353
118.2
0.0047
Qwen3.8 Max
99
0.4949
-0.0391
-96.1
Gemma 4 31B
99
0.5152
-0.0532
-194.2
-0.05
GPT-5.4
99
0.5051
-0.1049
56.6
-0.136
Grok 4.5
99
0.5152
-0.1134
-34.0
Gemini 3.1 Pro
99
0.4848
-0.1169
-109.5
0.1236
MiMo V2.5
98
0.5408
-0.1394
37.4
-0.2147
DeepSeek V4 Flash
95
0.5158
-0.1568
-215.3
-0.0972
GPT-5.6 Sol
99
0.4848
-0.1632
-361.0
-0.0823
B5 · 13F smart-money adds (post-cutoff, small n)
Top new/increased positions from Q1-2026 13F filings of 11 prominent funds (filed mid-May 2026). Model sees who bought, size, and portfolio weight; predicts the 20-day post-filing market-adjusted return. Control: same stock/date without the 13F context. n≈38 — indicative only.
Model
n
Hit rate
Spearman IC
L/S spread (bps)
Control IC
GPT-5.4
38
0.6053
0.5802
2012.7
0.1344
DeepSeek V4 Pro
20
0.5882
0.5457
2321.5
0.0552
GLM 5.2
28
0.6071
0.4478
1945.9
0.0279
Claude Opus 5
38
0.6316
0.3672
1573.2
-0.0491
GPT-5.6 Sol
38
0.5833
0.2691
1521.1
-0.331
Qwen3.8 Max
38
0.6216
0.268
767.3
-0.0141
DeepSeek V4 Flash
36
0.6538
0.231
882.8
-0.0354
Grok 4.5
38
0.6111
0.1674
1373.8
-0.013
Gemma 4 31B
38
0.5526
0.1501
-6.8
-0.1553
Kimi K3
37
0.5833
0.1328
949.4
-0.0483
Gemini 3.1 Pro
38
0.6053
0.1211
-175.6
0.0059
Claude Sonnet 5
36
0.5625
0.0431
999.5
-0.2808
MiMo V2.5
38
0.6053
-0.1569
-835.4
-0.0061
B7 · FOMC statements & minutes (macro)
42 Federal Reserve releases (2024-01..2026-07, 8 post-cutoff). Model reads the release and predicts SPY's close-to-close move over the next 5 trading days (TLT also collected). Control: text withheld — the model knows only that a Fed release happened that day.
Model
n
Hit rate
Spearman IC
L/S spread (bps)
Control IC
Claude Opus 5
42
0.881
0.7083
305.7
0.7129
Kimi K3
41
0.8293
0.6001
257.2
0.4589
Qwen3.8 Max
42
0.6774
0.2925
80.6
0.54
GPT-5.5
33
0.7333
0.2433
63.5
0.4513
Gemini 3.1 Pro
42
0.7059
0.1715
4.4
0.451
GLM 5.2
41
0.5833
0.1477
122.0
0.2338
MiMo V2.5
42
0.6
0.11
23.1
-0.1731
Claude Sonnet 5
42
0.6667
0.0537
-7.3
0.0634
GPT-5.6 Sol
42
0.575
-0.0161
0.8
0.1097
Gemma 4 31B
42
0.5909
-0.0608
-124.9
-0.0162
GPT-5.4
42
0.5897
-0.2047
-102.3
0.09
Grok 4.5
42
0.4667
-0.2838
-162.5
0.0288
DeepSeek V4 Flash
27
0.45
-0.292
-170.0
0.2004
B3 · SEC filing MD&A
Management's Discussion & Analysis sections from 10-K/20-F filings of 33 AI-infrastructure companies (2017–2024); predict 20-day post-filing market-adjusted drift. Control: text withheld.
Model
n
Hit rate
Spearman IC
L/S spread (bps)
Control IC
GPT-5.5
204
0.6765
0.6098
1983.3
0.6215
Claude Opus 5
204
0.6897
0.5466
2007.4
0.4427
DeepSeek V4 Pro
53
0.6154
0.5358
2220.7
-0.1678
Kimi K3
150
0.6533
0.4475
1339.3
0.3815
Gemini 3.6 Flash
204
0.5196
0.3109
1811.4
0.5076
Qwen3.8 Max
204
0.532
0.281 ±0.046 (2)
1003.7
0.3856
Gemini 3.1 Pro
204
0.549
0.2748
1536.7
0.4834
GPT-5.6 Sol
204
0.5147
0.1555
694.7
0.2055
Kimi K2.6
192
0.4688
0.1515
547.3
0.1058
GLM 5.2
179
0.5028
0.140 ±0.011 (2)
786.6
0.2724
Claude Opus 4.8
200
0.545
0.1151
392.5
0.2136
Claude Sonnet 5
204
0.4745
0.105 ±0.105 (2)
10.2
0.2044
Claude Haiku 4.5
204
0.5343
0.0867
317.2
-0.0905
GPT-5.6 Terra
204
0.4804
0.0861
578.7
0.0988
DeepSeek V4 Flash
187
0.5135
0.077 ±0.031 (2)
327.1
0.1499
MiniMax M3
204
0.5147
0.0716
342.4
0.096
DeepSeek V3.2
204
0.5616
0.0646
459.5
0.1981
Gemma 4 31B
204
0.5588
0.061 ±0.018 (2)
304.3
0.1181
Hunyuan 3
204
0.4706
0.0293
399.6
0.1292
GPT-5.4
204
0.4559
0.011
312.3
0.1737
Grok 4.5
204
0.4412
-0.0082
179.6
0.0263
MiMo V2.5
203
0.4455
-0.050 ±0.010 (2)
-195.2
0.0551
Judgment across model generations
Running each family's successive releases on the same post-cutoff article sample asks: is financial judgment improving generation over generation? Broadly yes — most families climb — though unevenly, and several legacy serving pools were no longer available to test (oldest generations excluded where marked on the platform).
Anthropic Claude Opus
Generation
Spearman IC
Hit rate
Control IC
Claude Opus 4.5
0.2592
0.4242
0.1184
Claude Opus 4.6
0.2846
0.431
0.1571
Claude Opus 4.7
0.2745
0.4291
-0.1365
Claude Opus 4.8
0.2825
0.4257
Claude Opus 5
0.610 ±0.028 (2)
0.5118
-0.0261
OpenAI GPT
Generation
Spearman IC
Hit rate
Control IC
GPT-5.4
0.2579
0.4276
-0.0174
GPT-5.5
0.449 ±0.027 (2)
0.476
0.0376
GPT-5.6 Sol
0.3251
0.4411
-0.1052
GPT-5.6 Terra
0.2414
0.4291
0.2342
Moonshot Kimi
Generation
Spearman IC
Hit rate
Control IC
Kimi K2.6
0.2022
0.4275
0.0483
Kimi K3
0.417 ±0.049 (2)
0.482
0.2397
DeepSeek
Generation
Spearman IC
Hit rate
Control IC
DeepSeek V3.2
0.1292
0.3943
0.0534
DeepSeek V4 Flash
0.168 ±0.053 (2)
0.4316
-0.0293
DeepSeek V4 Pro
0.1776
0.4892
-0.0639
Cost-optimal routing
Because every model runs the same tasks in the same harness, the results double as a routing table: the cheapest model that solves each difficulty tier perfectly. Routine replication does not need a flagship.
Task tier
Cheapest perfect solver
Cost per run
T1 · 20-stock momentum
Gemma 4 31B
$0.0010
T2 · News event study
Gemma 4 31B
$0.0008
T3 · Covered-call option backtest
Gemma 4 31B
$0.0109
T4 · Full-market momentum (11GB)
Gemma 4 31B
$0.0054
T6 · Replicate an ML asset-pricing study (RFS-style, 4 models trained)
MiMo V2.5
$0.0271
Agent efficacy — from ranking skill to P&L
Three falsifiable checks on the strongest signal (B2). 1) Tradability: weekly long/short tercile portfolios from each model's predictions, 20-day overlapping holds, 10bps one-way costs — all configurations finished positive over the (short, 3-month) sample. 2) Stability: re-running the same model on the same articles gives prediction rank-correlation ≈ 0.9 and near-identical ICs (Qwen3.8-Max 0.451→0.468, Sonnet-5 0.389→0.437 across seeds). 3) Pre-registration: predictions for brand-new articles are committed to this repository before outcomes exist — see the locked file; check back a month later.
Configuration
Cohorts
Ann. ret
Ann. vol
Sharpe
Max DD
Opus 5 (cutoff 2026-05)
10
134.89%
48.96%
2.76
-13.08%
GPT-5.5
10
80.0%
43.76%
1.83
-10.5%
Qwen3.8 Max
10
135.54%
50.65%
2.68
-13.71%
Kimi K3
10
108.53%
48.77%
2.23
-17.48%
Sonnet 5 (clean)
10
151.77%
45.15%
3.36
-10.54%
GPT-5.6 Sol (clean)
10
6.73%
40.07%
0.17
-22.7%
GPT-5.4 (clean)
10
32.85%
32.24%
1.02
-12.35%
Grok 4.5 (clean)
10
58.65%
16.16%
3.63
-6.25%
ENSEMBLE (4 clean models)
10
74.33%
40.51%
1.83
-10.54%
ENSEMBLE (all 8)
10
138.57%
45.1%
3.07
-10.54%
67 trading days, Apr–Aug 2026, AI-heavy universe, overlapping cohorts — treat annualized figures as directional, not expected returns.
Which financial text carries alpha?
The same models, the same harness, four post-cutoff text sources — very different outcomes. Analyst opinion articles (B2) support strong cross-sectional ranking (best models IC ≈ 0.4–0.6): opinions diffuse slowly. Earnings press releases (B4) show near-zero or negative IC for most models: hard numbers are priced within minutes, so reading them the next day adds nothing — models that naively map "good quarter → buy" get systematically caught by post-announcement reversals. 13F disclosures (B5) show positive signal on a small sample. The lesson: text selection dominates model choice.
Side finding — ML asset-pricing alpha after publication
Building T6's ground truth required re-running a scoped version of a canonical machine-learning asset-pricing study (94 firm characteristics, expanding-window refits) on 2015–2021 — entirely after the original paper's out-of-sample period. The equal-weighted decile long-short portfolios that earned Sharpe ratios above 2 in the original sample largely vanish:
Model
OOS R² (%)
Decile L/S Sharpe 2015–2021
OLS-3 (size/value/momentum)
0.267
-0.101
Elastic net
0.197
-0.312
Gradient-boosted trees
0.102
0.196
Neural net (3 layers)
-0.175
0.320
Nonlinear models still beat linear ones — the paper's qualitative ranking survives — but linear signals flip negative and the economic magnitude is a fraction of the published era. (Scope deviations: characteristics only, equal-weighted deciles, training history starts 2004.)
Honest caveats
Track B2's "post-cutoff" claim is per-model: models released mid-2026 may have
training data extending into the article window; a per-model cutoff table is in
progress. Withheld-text controls bound the contamination for every track.
Article/filing forward windows overlap in calendar time, so cross-sectional
correlation inflates naive significance; treat Track B as a ranking across
models on a common sample, not a tradable alpha estimate.
Small samples where noted (B2 n<500); error bars matter and seeds are being added.
Underlying market and text datasets are licensed; this page publishes
aggregates only.
Scores are harness-dependent: the same model can score differently under a
different runtime (a point model vendors themselves acknowledge). All numbers
here come from ONE neutral harness on one serving platform — comparable to each
other, not to vendor-reported benchmarks.
Costs are computed at full input-token list price; provider-side prefix
caching (not metered in early runs) would reduce flagship agent costs somewhat.