FinAgentBench

📖 Plain-language report · 中文报告 · 📄 Working paper (PDF) · 🔒 Pre-registered predictions

Can LLM agents do quantitative finance research? Two capabilities, measured separately on private data: execution (replicate a result from a written method spec + raw data) and judgment (predict returns from financial text). 20+ frontier models, one identical harness, one inference platform, script-graded.

Source & docs on GitHub →

29frontier models, one identical harness, 220 graded agent runs
23/29models replicate a momentum study perfectly from spec — execution is commoditized
IC 0.61best model reading post-cutoff analyst articles; withheld-text controls collapse to ≈0

Key findings

  1. Execution is commoditized. Given a precise method spec and raw data, nearly every model — including sub-cent-per-run ones — perfectly replicates standard quantitative studies (momentum portfolios, event studies, an option-overlay backtest, an 11GB full-market study). See T1–T4.
  2. The frontier appears at paper scale. Replicating a full machine-learning asset-pricing study end-to-end (T6: train four model classes, annual refits, out-of-sample portfolio evaluation) splits the field roughly in half. Failures are precision failures — every model produces plausible magnitudes; only some hold twenty-plus spec details simultaneously.
  3. Reading long financial text yields real, measurable judgment. On analyst articles published after training cutoffs, flagship models reach Spearman ICs around 0.4–0.5 vs ≈0.2 for small models; with the text withheld, every model collapses to ≈0. See B2.
  4. On historical data, apparent skill is often memorization. Some models predict historical event outcomes better with the news withheld — they recall two decades of price history from (ticker, date) alone. The honest metric is ΔIC = IC(text) − IC(withheld). See B1.
  5. Models rank better than they call direction. Several models show significantly positive ICs with below-50% hit rates: a systematic long bias. Use these signals cross-sectionally, not for market timing.
  6. Reliability, quantified. Re-running identical tasks: ~2–3% of runs silently produce wrong answers and ~6% die on provider infrastructure — consistent with what production-agent practitioners report as their top challenge.
  7. Published ML alpha decays. Re-running a canonical ML asset-pricing design on 2015–2021 (after the original out-of-sample period), decile long-short Sharpe falls from 2+ to ≈0.2–0.3 and linear signals flip negative; the nonlinear-beats-linear ranking survives.

Methods in brief

Track A — Execution: replicate from spec

The agent gets a task sheet (method spelled out, results withheld), read-only data, a bash sandbox, and a 40-step budget. Grading is scripted: each submitted statistic vs ground truth (loose ±5% / strict ±1%). Samples are private extracts, so published numbers can't be recalled from training.

T1 · 20-stock momentum

ModelStrictLooseConsistency (strict across seeds)Best stepsBest cost
Claude Haiku 4.51.01.06$0.0668
Claude Opus 4.81.01.02$0.0615
Claude Opus 51.01.05/52$0.0553
Claude Sonnet 51.01.05/51$0.0188
DeepSeek V3.21.01.015$0.0422
DeepSeek V4 Flash1.01.05/54$0.0038
DeepSeek V4 Pro1.01.05/53$0.0297
GLM 5.21.01.04/52$0.0154
GPT-5.41.01.05/51$0.0193
GPT-5.51.01.03$0.1473
GPT-5.6 Sol1.01.05/51$0.0269
GPT-5.6 Terra1.01.01$0.0155
Gemini 3.1 Pro1.01.05/52$0.0775
Gemini 3.6 Flash1.01.028$1.0086
Gemma 4 31B1.01.05/51$0.0010
Grok 4.51.01.05/52$0.0204
Hunyuan 31.01.02$0.0015
Kimi K2.61.01.02$0.0132
Kimi K31.01.05/52$0.0270
MiMo V2.51.01.05/52$0.0013
MiniMax M31.01.06$0.0264
MiniMaxAI/MiniMax-M2.71.01.07$0.0182
Qwen3.8 Max1.01.04/53$0.0294
Nemotron 3 Ultra0.00.0
Qwen/Qwen3.6-Plus0.00.0
bytedance/seed-2.0-mini0.00.0
kwaipilot/kat-coder-pro-v20.00.0
moonshotai/Kimi-K2.50.00.0
zai-org/GLM-4.7-FP80.00.0

T2 · News event study

ModelStrictLooseConsistency (strict across seeds)Best stepsBest cost
Claude Opus 51.01.05/52$0.0607
Claude Sonnet 51.01.05/51$0.0180
DeepSeek V4 Flash1.01.05/55$0.0037
DeepSeek V4 Pro1.01.05/54$0.0470
GLM 5.21.01.05/53$0.0205
GPT-5.41.01.05/51$0.0190
GPT-5.6 Sol1.01.05/51$0.0285
Gemini 3.1 Pro1.01.05/52$0.0553
Gemma 4 31B1.01.05/51$0.0008
Grok 4.51.01.05/52$0.0220
Kimi K31.01.05/52$0.0330
MiMo V2.51.01.04/52$0.0011
Qwen3.8 Max1.01.05/52$0.0339

T3 · Covered-call option backtest

ModelStrictLooseConsistency (strict across seeds)Best stepsBest cost
Claude Opus 51.01.02$0.1268
Claude Sonnet 51.01.010$0.2537
DeepSeek V4 Flash1.01.032$0.1155
DeepSeek V4 Pro1.01.021$0.7153
GLM 5.21.01.014$0.3479
GPT-5.41.01.07$0.1813
Gemini 3.1 Pro1.01.016$0.7649
Gemma 4 31B1.01.06$0.0109
Grok 4.51.01.05$0.1129
Kimi K31.01.04$0.2244
MiMo V2.51.01.011$0.0226
Qwen3.8 Max1.01.013$0.3382
GPT-5.6 Sol0.60.72$0.0718

T4 · Full-market momentum (11GB)

ModelStrictLooseConsistency (strict across seeds)Best stepsBest cost
Claude Opus 51.01.05$0.1968
Claude Sonnet 51.01.011$0.2056
DeepSeek V4 Pro1.01.013$0.3799
GLM 5.21.01.026$0.3970
Gemma 4 31B1.01.05$0.0054
Grok 4.51.01.03$0.0381
Kimi K31.01.011$0.2920
MiMo V2.51.01.05$0.0073
Qwen3.8 Max1.01.019$0.7046
GPT-5.40.81.01$0.0300
GPT-5.6 Sol0.80.85$0.1065
Gemini 3.1 Pro0.41.030$1.0432
DeepSeek V4 Flash0.20.420$0.0372

T6 · Replicate an ML asset-pricing study (RFS-style, 4 models trained)

ModelStrictLooseConsistency (strict across seeds)Best stepsBest cost
Claude Opus 51.01.01/37$0.5712
Claude Sonnet 51.01.03/35$0.1292
DeepSeek V3.21.01.045$0.2838
DeepSeek V4 Flash1.01.01/328$0.0669
DeepSeek V4 Pro1.01.02/39$0.1135
GLM 5.21.01.03/312$0.1919
Gemini 3.6 Flash1.01.022$0.9978
Grok 4.51.01.02/35$0.1123
Kimi K31.01.03/313$0.5373
MiMo V2.51.01.01/315$0.0271
MiniMax M31.01.014$0.1622
Qwen3.8 Max1.01.03/315$0.4760
GPT-5.6 Sol0.50.80/310$0.3244
Claude Haiku 4.50.40.58$0.0991
Claude Opus 4.80.40.59$0.5695
Gemini 3.1 Pro0.40.50/312$0.3600
GPT-5.6 Terra0.40.410$0.2201
Hunyuan 30.10.516$0.0406
GPT-5.40.10.20/35$0.1480
Gemma 4 31B0.00.40/39$0.0139
GPT-5.50.00.111$1.4065
Kimi K2.60.00.0

Track B — Judgment: predict returns from text

B1 · News headlines

1,008 stratified news events on 20 mega-caps (2004–2024) from an institutional news-event dataset. Model sees headline + ticker + date + trailing stats; predicts 5-day beta-adjusted return. Control: headline withheld (memorization baseline).

ModelnHit rateSpearman ICL/S spread (bps)Control ICΔIC (text − control)
Claude Opus 510080.67130.522973.70.5326-0.0106
Kimi K39710.56370.2413593.10.21570.0256
Claude Opus 4.810080.57330.2299550.60.16260.0673
Gemini 3.1 Pro10080.57120.1991548.60.2281-0.029
Qwen3.8 Max10000.56120.1658344.20.179-0.0132
Claude Sonnet 510080.54810.1435425.90.11870.0248
GLM 5.29190.55310.1368345.20.04380.093
GPT-5.6 Sol10080.53270.1324271.30.1643-0.0319
DeepSeek V4 Pro6320.57680.0992117.90.04450.0547
DeepSeek V4 Flash9860.5280.052103.80.02360.0284
Grok 4.510080.52570.0442136.60.0010.0432
GPT-5.410080.50510.03972.50.02980.0092
Gemma 4 31B10080.51540.037743.60.03620.0015
MiMo V2.510060.50820.014577.60.0324-0.0179

B2 · Analyst articles (post-cutoff)

Full-text investment research articles published Apr–Aug 2026 — after the training cutoff of the models under test. Model predicts 20-day market-adjusted return. Control: article withheld.

ModelnHit rateSpearman ICIC (Jun–Jul only)CutoffL/S spread (bps)Control IC
Claude Opus 52970.51180.610 ±0.028 (2)0.62712026-05 ⚠️4447.3-0.0261
GPT-5.52920.4760.449 ±0.027 (2)0.562UNDISCLOSED ⚠️2759.00.0376
Qwen3.8 Max2960.48810.446 ±0.025 (3)0.6027UNDISCLOSED ⚠️2687.2
Kimi K32790.4820.417 ±0.049 (2)0.481UNDISCLOSED ⚠️2773.40.2397
Claude Sonnet 52970.46780.415 ±0.024 (3)0.41612026-012198.1-0.1583
google_gemini-3.5-flash2960.4490.3680.4266UNDISCLOSED ⚠️2865.30.0491
google_gemini-3-flash-preview2970.45610.34250.3345UNDISCLOSED ⚠️1737.30.0345
GPT-5.6 Sol2970.44110.32510.38762026-022558.4-0.1052
anthropic_claude-sonnet-4.62970.46460.31740.3319UNDISCLOSED ⚠️1917.4-0.0876
Claude Opus 4.62970.4310.28460.3942UNDISCLOSED ⚠️1775.90.1571
Gemini 3.6 Flash2970.45270.28390.2685UNDISCLOSED ⚠️1845.60.1131
Claude Opus 4.82970.42570.28250.3092UNDISCLOSED ⚠️2864.3
Claude Opus 4.72970.42910.27450.2531UNDISCLOSED ⚠️2355.2-0.1365
Qwen_Qwen3.7-Max2970.44140.26570.2784UNDISCLOSED ⚠️2239.2
Claude Opus 4.52970.42420.25920.3221UNDISCLOSED ⚠️1254.00.1184
MiniMax M32930.44180.2580.2526UNDISCLOSED ⚠️2048.90.0151
GPT-5.42970.42760.25790.27972025-081873.3-0.0174
GPT-5.6 Terra2970.42910.24140.2072UNDISCLOSED ⚠️1813.30.2342
anthropic_claude-sonnet-4.52970.44260.2170.1823UNDISCLOSED ⚠️1589.6
Grok 4.52970.42570.207 ±0.011 (2)0.27182026-021096.4
MiniMaxAI_MiniMax-M2.72660.42640.20260.2001UNDISCLOSED ⚠️920.1-0.1788
Kimi K2.62640.42750.20220.163UNDISCLOSED ⚠️2285.20.0483
Gemini 3.1 Pro2970.43050.19410.17132025-011701.2
MiMo V2.52970.42760.192 ±0.007 (2)0.20552024-121072.50.0408
zai-org_GLM-5-FP82760.41970.18320.2132UNDISCLOSED ⚠️1260.9-0.0697
DeepSeek V4 Pro1420.48920.17760.2043UNDISCLOSED ⚠️222.8-0.0639
DeepSeek V4 Flash2920.43160.168 ±0.053 (2)0.2357UNDISCLOSED ⚠️1416.4-0.0293
moonshotai_Kimi-K2.52970.42420.16440.2038UNDISCLOSED ⚠️548.6-0.1331
google_gemini-3.5-flash-lite2970.41220.16320.2079UNDISCLOSED ⚠️1006.8
GLM 5.22820.41070.155 ±0.047 (2)0.2317UNDISCLOSED ⚠️1648.60.0238
deepseek-ai_DeepSeek-V3-03242970.41750.15360.1641UNDISCLOSED ⚠️1048.3-0.0123
Claude Haiku 4.52970.43770.1490.1444UNDISCLOSED ⚠️782.2
Hunyuan 32780.4170.14230.1164UNDISCLOSED ⚠️1044.30.0067
deepseek-ai_DeepSeek-R1-05282830.44130.13380.123UNDISCLOSED ⚠️1098.4-0.2132
DeepSeek V3.22790.39430.12920.1567UNDISCLOSED ⚠️966.70.0534
Gemma 4 31B2970.4150.115 ±0.024 (2)0.12512025-01747.9

B4 · Earnings press releases (post-cutoff)

8-K item-2.02 earnings press releases filed with the SEC Feb–Jul 2026 (after most models' training cutoffs), fetched from EDGAR. Model reads the release and predicts the 5-day post-filing market-adjusted return. Control: text withheld.

ModelnHit rateSpearman ICL/S spread (bps)Control IC
Claude Sonnet 5960.54950.119 ±0.075 (2)116.10.0407
Claude Opus 51000.520.093 ±0.020 (2)-197.20.2066
Kimi K2.6970.52240.0419-46.10.0906
Claude Opus 4.8990.49410.0417-94.90.0571
Kimi K3960.58060.0382-601.4-0.0904
DeepSeek V3.2970.48390.0038564.20.1773
GPT-5.6 Sol1000.52310.0006-334.0-0.171
GPT-5.41000.5102-0.0218-25.50.0675
MiniMax M31000.4615-0.0479-124.00.1595
DeepSeek V4 Flash960.46-0.051 ±0.150 (2)-691.30.053
GPT-5.51000.4821-0.0825-170.8-0.0039
Gemini 3.6 Flash1000.459-0.085-350.3-0.0441
MiMo V2.5990.4833-0.0855-458.3-0.0056
Qwen3.8 Max1000.4783-0.1047-492.6
Grok 4.51000.4259-0.105 ±0.054 (2)-716.50.0052
GPT-5.6 Terra1000.4286-0.1058187.80.0235
Gemini 3.1 Pro1000.46-0.1114-817.8-0.1402
DeepSeek V4 Pro780.4643-0.1173-544.6-0.0912
Claude Haiku 4.5990.4699-0.1238-417.5
GLM 5.2880.4375-0.1358-594.6-0.0791
Hunyuan 3750.4565-0.1439-828.4
Gemma 4 31B1000.3913-0.1916-1049.5-0.1419

B6 · Earnings-call transcripts (post-cutoff)

Full earnings-call transcripts (Feb–Jul 2026). Model reads management remarks + Q&A and predicts the 5-day post-call market-adjusted return. Control: transcript withheld.

ModelnHit rateSpearman ICL/S spread (bps)Control IC
Claude Opus 5990.51520.0033-261.50.1585
Claude Sonnet 5990.52530.0016329.8-0.2382
GLM 5.2890.51690.0003-272.3-0.0163
Kimi K3990.5051-0.0267-71.5
GPT-5.5860.5233-0.0353118.20.0047
Qwen3.8 Max990.4949-0.0391-96.1
Gemma 4 31B990.5152-0.0532-194.2-0.05
GPT-5.4990.5051-0.104956.6-0.136
Grok 4.5990.5152-0.1134-34.0
Gemini 3.1 Pro990.4848-0.1169-109.50.1236
MiMo V2.5980.5408-0.139437.4-0.2147
DeepSeek V4 Flash950.5158-0.1568-215.3-0.0972
GPT-5.6 Sol990.4848-0.1632-361.0-0.0823

B5 · 13F smart-money adds (post-cutoff, small n)

Top new/increased positions from Q1-2026 13F filings of 11 prominent funds (filed mid-May 2026). Model sees who bought, size, and portfolio weight; predicts the 20-day post-filing market-adjusted return. Control: same stock/date without the 13F context. n≈38 — indicative only.

ModelnHit rateSpearman ICL/S spread (bps)Control IC
GPT-5.4380.60530.58022012.70.1344
DeepSeek V4 Pro200.58820.54572321.50.0552
GLM 5.2280.60710.44781945.90.0279
Claude Opus 5380.63160.36721573.2-0.0491
GPT-5.6 Sol380.58330.26911521.1-0.331
Qwen3.8 Max380.62160.268767.3-0.0141
DeepSeek V4 Flash360.65380.231882.8-0.0354
Grok 4.5380.61110.16741373.8-0.013
Gemma 4 31B380.55260.1501-6.8-0.1553
Kimi K3370.58330.1328949.4-0.0483
Gemini 3.1 Pro380.60530.1211-175.60.0059
Claude Sonnet 5360.56250.0431999.5-0.2808
MiMo V2.5380.6053-0.1569-835.4-0.0061

B7 · FOMC statements & minutes (macro)

42 Federal Reserve releases (2024-01..2026-07, 8 post-cutoff). Model reads the release and predicts SPY's close-to-close move over the next 5 trading days (TLT also collected). Control: text withheld — the model knows only that a Fed release happened that day.

ModelnHit rateSpearman ICL/S spread (bps)Control IC
Claude Opus 5420.8810.7083305.70.7129
Kimi K3410.82930.6001257.20.4589
Qwen3.8 Max420.67740.292580.60.54
GPT-5.5330.73330.243363.50.4513
Gemini 3.1 Pro420.70590.17154.40.451
GLM 5.2410.58330.1477122.00.2338
MiMo V2.5420.60.1123.1-0.1731
Claude Sonnet 5420.66670.0537-7.30.0634
GPT-5.6 Sol420.575-0.01610.80.1097
Gemma 4 31B420.5909-0.0608-124.9-0.0162
GPT-5.4420.5897-0.2047-102.30.09
Grok 4.5420.4667-0.2838-162.50.0288
DeepSeek V4 Flash270.45-0.292-170.00.2004

B3 · SEC filing MD&A

Management's Discussion & Analysis sections from 10-K/20-F filings of 33 AI-infrastructure companies (2017–2024); predict 20-day post-filing market-adjusted drift. Control: text withheld.

ModelnHit rateSpearman ICL/S spread (bps)Control IC
GPT-5.52040.67650.60981983.30.6215
Claude Opus 52040.68970.54662007.40.4427
DeepSeek V4 Pro530.61540.53582220.7-0.1678
Kimi K31500.65330.44751339.30.3815
Gemini 3.6 Flash2040.51960.31091811.40.5076
Qwen3.8 Max2040.5320.281 ±0.046 (2)1003.70.3856
Gemini 3.1 Pro2040.5490.27481536.70.4834
GPT-5.6 Sol2040.51470.1555694.70.2055
Kimi K2.61920.46880.1515547.30.1058
GLM 5.21790.50280.140 ±0.011 (2)786.60.2724
Claude Opus 4.82000.5450.1151392.50.2136
Claude Sonnet 52040.47450.105 ±0.105 (2)10.20.2044
Claude Haiku 4.52040.53430.0867317.2-0.0905
GPT-5.6 Terra2040.48040.0861578.70.0988
DeepSeek V4 Flash1870.51350.077 ±0.031 (2)327.10.1499
MiniMax M32040.51470.0716342.40.096
DeepSeek V3.22040.56160.0646459.50.1981
Gemma 4 31B2040.55880.061 ±0.018 (2)304.30.1181
Hunyuan 32040.47060.0293399.60.1292
GPT-5.42040.45590.011312.30.1737
Grok 4.52040.4412-0.0082179.60.0263
MiMo V2.52030.4455-0.050 ±0.010 (2)-195.20.0551

Judgment across model generations

Running each family's successive releases on the same post-cutoff article sample asks: is financial judgment improving generation over generation? Broadly yes — most families climb — though unevenly, and several legacy serving pools were no longer available to test (oldest generations excluded where marked on the platform).

Anthropic Claude Opus

GenerationSpearman ICHit rateControl IC
Claude Opus 4.50.25920.42420.1184
Claude Opus 4.60.28460.4310.1571
Claude Opus 4.70.27450.4291-0.1365
Claude Opus 4.80.28250.4257
Claude Opus 50.610 ±0.028 (2)0.5118-0.0261

OpenAI GPT

GenerationSpearman ICHit rateControl IC
GPT-5.40.25790.4276-0.0174
GPT-5.50.449 ±0.027 (2)0.4760.0376
GPT-5.6 Sol0.32510.4411-0.1052
GPT-5.6 Terra0.24140.42910.2342

Moonshot Kimi

GenerationSpearman ICHit rateControl IC
Kimi K2.60.20220.42750.0483
Kimi K30.417 ±0.049 (2)0.4820.2397

DeepSeek

GenerationSpearman ICHit rateControl IC
DeepSeek V3.20.12920.39430.0534
DeepSeek V4 Flash0.168 ±0.053 (2)0.4316-0.0293
DeepSeek V4 Pro0.17760.4892-0.0639

Cost-optimal routing

Because every model runs the same tasks in the same harness, the results double as a routing table: the cheapest model that solves each difficulty tier perfectly. Routine replication does not need a flagship.

Task tierCheapest perfect solverCost per run
T1 · 20-stock momentumGemma 4 31B$0.0010
T2 · News event studyGemma 4 31B$0.0008
T3 · Covered-call option backtestGemma 4 31B$0.0109
T4 · Full-market momentum (11GB)Gemma 4 31B$0.0054
T6 · Replicate an ML asset-pricing study (RFS-style, 4 models trained)MiMo V2.5$0.0271

Agent efficacy — from ranking skill to P&L

Three falsifiable checks on the strongest signal (B2). 1) Tradability: weekly long/short tercile portfolios from each model's predictions, 20-day overlapping holds, 10bps one-way costs — all configurations finished positive over the (short, 3-month) sample. 2) Stability: re-running the same model on the same articles gives prediction rank-correlation ≈ 0.9 and near-identical ICs (Qwen3.8-Max 0.451→0.468, Sonnet-5 0.389→0.437 across seeds). 3) Pre-registration: predictions for brand-new articles are committed to this repository before outcomes exist — see the locked file; check back a month later.

ConfigurationCohortsAnn. retAnn. volSharpeMax DD
Opus 5 (cutoff 2026-05)10134.89%48.96%2.76-13.08%
GPT-5.51080.0%43.76%1.83-10.5%
Qwen3.8 Max10135.54%50.65%2.68-13.71%
Kimi K310108.53%48.77%2.23-17.48%
Sonnet 5 (clean)10151.77%45.15%3.36-10.54%
GPT-5.6 Sol (clean)106.73%40.07%0.17-22.7%
GPT-5.4 (clean)1032.85%32.24%1.02-12.35%
Grok 4.5 (clean)1058.65%16.16%3.63-6.25%
ENSEMBLE (4 clean models)1074.33%40.51%1.83-10.54%
ENSEMBLE (all 8)10138.57%45.1%3.07-10.54%

67 trading days, Apr–Aug 2026, AI-heavy universe, overlapping cohorts — treat annualized figures as directional, not expected returns.

Which financial text carries alpha?

The same models, the same harness, four post-cutoff text sources — very different outcomes. Analyst opinion articles (B2) support strong cross-sectional ranking (best models IC ≈ 0.4–0.6): opinions diffuse slowly. Earnings press releases (B4) show near-zero or negative IC for most models: hard numbers are priced within minutes, so reading them the next day adds nothing — models that naively map "good quarter → buy" get systematically caught by post-announcement reversals. 13F disclosures (B5) show positive signal on a small sample. The lesson: text selection dominates model choice.

Side finding — ML asset-pricing alpha after publication

Building T6's ground truth required re-running a scoped version of a canonical machine-learning asset-pricing study (94 firm characteristics, expanding-window refits) on 2015–2021 — entirely after the original paper's out-of-sample period. The equal-weighted decile long-short portfolios that earned Sharpe ratios above 2 in the original sample largely vanish:

ModelOOS R² (%)Decile L/S Sharpe 2015–2021
OLS-3 (size/value/momentum)0.267-0.101
Elastic net0.197-0.312
Gradient-boosted trees0.1020.196
Neural net (3 layers)-0.1750.320

Nonlinear models still beat linear ones — the paper's qualitative ranking survives — but linear signals flip negative and the economic magnitude is a fraction of the published era. (Scope deviations: characteristics only, equal-weighted deciles, training history starts 2004.)

Honest caveats

  • Track B2's "post-cutoff" claim is per-model: models released mid-2026 may have training data extending into the article window; a per-model cutoff table is in progress. Withheld-text controls bound the contamination for every track.
  • Article/filing forward windows overlap in calendar time, so cross-sectional correlation inflates naive significance; treat Track B as a ranking across models on a common sample, not a tradable alpha estimate.
  • Small samples where noted (B2 n<500); error bars matter and seeds are being added.
  • Underlying market and text datasets are licensed; this page publishes aggregates only.
  • Scores are harness-dependent: the same model can score differently under a different runtime (a point model vendors themselves acknowledge). All numbers here come from ONE neutral harness on one serving platform — comparable to each other, not to vendor-reported benchmarks.
  • Costs are computed at full input-token list price; provider-side prefix caching (not metered in early runs) would reduce flagship agent costs somewhat.