Financial AI evaluations

Scoreboard

The public reference point for the investment performance and characteristics of financial AI agents.

How to interpret results

Investment performance is strategy-dependent and multi-faceted:

Strategy and books

systematic stockrate

Each agent is independently deployed to rate a ticker on each decision date. The ratings are then aggregated into a cross-sectional signal that we transform into holdings in the following ways:

  • MVO. Mean-variance-optimized portfolio based on the agent rating and stock covariance.
  • Naive tilt. Over/underweight the benchmark based on the agent rating, constrained to zero net exposure.
  • SW Long >0.0. Long every ticker scored above zero; higher score, larger weight.

fundamental stockpick

Each agent is responsible for identifying the most attractive tickers across the universe. They output holdings and their respective ratings.

  • SW Long >0.0. Long every ticker scored above zero; higher score, larger weight.

Mission

Both harnesses receive this rubric. The score is a rating, not a trade.

You are a portfolio manager. You will be asked, independently for one company at a time, to rate that company's fundamental attractiveness and the trajectory of its business on a continuous scale from -1 to +1.

## What the score combines

Judge three pillars, in this order of importance:

1. **Fundamental health & trajectory** (primary) — earnings path and durability, margins, cash generation, balance-sheet resilience, franchise / competitive position.
2. **Valuation vs opportunity** (co-primary) — mainly vs the company's *own* recent history (is the market paying more or less for the same earnings power?).
3. **Near-term rerating chance** (secondary boost) — a dated, fundamental catalyst that could re-rate the stock over the coming quarter-to-year (guidance, earnings event, product/regulatory clearance, capital-allocation shift, leadership change with a clear strategy implication). This pillar exists because assessment is quarterly: a sound fundamental view should be slightly stronger when a credible re-rating path is in the pack.

Pillar 3 is a *modifier*, not a thesis. It can raise or lower the score within a band once pillars 1–2 are set; it cannot turn a weak or deteriorating business into a high score, and price momentum alone is never a rerating catalyst.

## Score anchors (use the continuous range; these are landmarks)

**Long side**


| Score    | Meaning                                                                                                                                                                                                                                                             |
| -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **+1.0** | Extremely attractive: excellent fundamental health *and* clear positive trajectory, clearly undervalued vs own history (and peers if shown), *and* a credible near-term fundamental catalyst that makes positive re-rating likely. All three pillars aligned; rare. |
| **+0.5** | Strongly attractive: excellent health + trajectory *and* clearly undervalued vs own history. Catalyst helpful but not required. Multi-factor, well-cited.                                                                                                           |
| **+0.2** | Mildly attractive: good / solid fundamentals and trajectory; valuation fair-to-slightly-cheap vs own history (OK, not a screaming bargain). Optional small bump (toward ~+0.3) if a credible near-term catalyst is present without stretching the fundamental case. |
| **0.0**  | Neutral: mixed or fully priced — nothing clearly mis-set in fundamentals or valuation, or pillars cancel.                                                                                                                                                           |


**Short side** (mirror the long side)


| Score    | Meaning                                                                                                                           |
| -------- | --------------------------------------------------------------------------------------------------------------------------------- |
| **−0.2** | Mildly unattractive: soft or deteriorating fundamentals and/or somewhat rich vs own history; no offsetting catalyst.              |
| **−0.5** | Strongly unattractive: poor / worsening health or trajectory *and* expensive vs own history.                                      |
| **−1.0** | Extremely unattractive: bad/worsening fundamentals, clearly expensive, *and* near-term downside catalyst or de-rating risk. Rare. |


In-between values (e.g. +0.35, −0.15) are encouraged when the case sits between anchors. Prefer milder scores when evidence is thin, gapped, or conflicting. Reserve **|score| ≥ 0.5** for multi-factor, well-cited cases; reserve **|score| ≥ 0.8** for near-full alignment of all three pillars.

## How to combine (discipline)

- Start from pillar 1 (fundamentals). If health/trajectory is merely OK, you should rarely exceed ~+0.2 even if valuation looks cheap.
- Use pillar 2 to move within / across bands: cheap vs history can lift a good franchise toward +0.5; rich vs history should cap or flip an otherwise decent story.
- Apply pillar 3 last, as a modest adjuster (think roughly ±0.1 to ±0.2 on the continuous scale), and only when the catalyst is dated and fundamental — not a headline spike or "stock has been going up."
- Cross-lane weight when synthesizing specialist reports: quantitative fundamentals and own-history valuation dominate; qualitative franchise / structural web context supports; recent news/updates and macro/sentiment mainly feed pillar 3 (and risk offsets), they do not replace pillars 1–2.



## Anti-momentum (still binding)

Do not momentum-chase. A rising or falling price is not itself a reason to rate a company attractive or unattractive — price action is at most a check on whether the market has already priced in the fundamentals you see. A good quarter or a headline does not make a bad business attractive, and a bad quarter does not undo a durable franchise. If the only evidence for a view is that the stock has been going up, say so and score down your conviction.

## Output hygiene

The score is a rating, not a trade instruction — do not discuss position sizing, stop losses, or entry/exit timing; that is handled downstream, deterministically, from your rating alone.

Use only the point-in-time evidence you are given through your tools. You will not be told the current date beyond your decision date, and you have no access to information published after it. Cite the specific evidence (a filing, a price level, a news item, a macro reading) behind every material claim in your rationale — do not rely on general knowledge about the company that isn't grounded in what you were actually shown. Quote or omit: do not upgrade vague text into precise figures, and do not treat an incomplete or non-contiguous quarterly history as trailing twelve months.

When you have formed a view, submit it via the tool made available to you for that purpose.

Data they could access

Same point-in-time surface for both harnesses. Nothing published after the decision date.

  • Prices. Massive daily prices, 365-day lookback.
  • Fundamentals. Massive fundamentals, 540-day lookback.
  • Valuation ratios. Own-history and peer multiples, 365-day window.
  • Trailing returns. Simple returns vs DJIA at 1d / 2w / 1m / 3m / 6m / 12m, plus path, vol, and beta.
  • Macro. FRED regime pack — rates, curve, vol, credit, FX, oil, breakeven, labour; 90-day lookback.
  • News. Recent articles only, 14-day window.
  • Web search. Structural queries (business model, competition, risk, strategy) over 30 days, plus an updates query over 7 days. Age-clamped.

Harnesses

fintel_GFA

A structured specialist pipeline. A quantitative specialist reads a pre-built pack (prices, fundamentals, valuation ratios, trailing returns, macro). A qualitative specialist reads news and web context. The two run in parallel. A portfolio-manager call then synthesizes one −1 to +1 score and must submit through a fixed schema. The research path is the same every cell — the agent does not choose tools or extra steps.

OpenClaw

A tool-calling ReAct harness. The same point-in-time tools are available, but the agent decides what to pull and how many reasoning steps to take before submitting a score. That can surface more evidence; it also adds tokens, retries, and a 2–4× cost multiple versus GFA on the same models.

Twins

A twin is the same model run on both harnesses. That is how we hold the model fixed and attribute remaining differences in the book — sector tilt, residual skill, cost — to the harness.

Metrics

Ann. return

Compound growth of the scoreboard book, annualized over the window: (1 + total)^(1 / years) − 1. This is a book result, not a rating-quality result.

Spearman IC

Each decision date, rank-correlate the 30 scores with next-quarter returns (h=1). The scoreboard reports the mean of those correlations (16 periods). It asks whether higher-rated names subsequently returned more. Empty for stock-pick agents that do not rate the full universe.

IC t-stat

One-sample t of mean IC versus 0: (mean IC / sample std) × √n. We treat t > 3 as a conventional bar showing that ranking skill is distinguishable from noise over this window.

Residual IC

Each date, residualize the scores on point-in-time FF6 loadings (Mkt, SMB, HML, RMW, CMA, Mom): s = a + Bγ + u. Residual IC is Spearman corr(u, next-period excess return) — skill left after stripping common-factor bets. Empty for stock-pick agents.

Residual t

The same t-stat construction as IC t, applied to the residual-IC series. Residual t > 3 is the bar we use for idiosyncratic skill after neutralization.

Sharpe

Mean period return of the scoreboard book divided by period volatility, then annualized (× √ppy). Does not subtract the DJIA.

Information ratio

Mean active return versus the price-weighted DJIA, divided by tracking error, then annualized. Measures the assumed holding book versus the price-weighted index.

Max drawdown

Peak-to-trough decline of net NAV on the scoreboard book. Shallower is better.

Eval cost

USD charged for the 510 rating cells — model inference, not portfolio trading cost. OpenClaw is typically a 2–4× multiple of GFA on the same model.

Periods

Number of h=1 IC observations (16 in this window). One fewer than the 17 decision dates, because the last date has no next-quarter return yet.

Cumulative return
-4.0%30.7%65.5%100.2%135.0%2022-092022-112023-012023-032023-052023-072023-092023-112024-012024-032024-052024-072024-092024-112025-012025-032025-052025-072025-092025-112026-012026-032026-052026-072026-09
OC GLM 5.3OC Muse 1.3GFA GLM 5.3GFA MiMo 2.6GFA Grok 4.6GFA Muse 1.3OC MiMo 2.5GFA DSV 4.1OC Muse 1.3 pickGFA MiMo 2.5OC GLM 5.3 pickDJIA PWDJIA MW

Cumulative return of a hypothetical portfolio that proportionately holds every name scored above 0 (attractive) by the respective agent. The universe is the DJIA, shown price-weighted and market-cap-weighted.

Click a row for the resume

Agentic system

fintel commentary

Actionable insights from our evaluators.

01

MiMo 2.6 and Muse 1.3 leading on total return.

MiMo 2.6 and Muse 1.3 led on total return. Rank IC (how correlated the agent scores are with next-period returns) is positive for all eleven agents, but only four clear a significant threshold: GFA GLM 5.3 (t=4.3), GFA Grok 4.6 (t=3.1), OC GLM 5.3 (t=3.1), and GFA MiMo 2.6 (t=3.0). After factor neutralization, only GFA GLM 5.3 remains (residual t=3.5) — suggesting idiosyncratic insight beyond common-factor loadings. GFA MiMo 2.6 has the lowest drawdown at -1.2%.

Choosing the right model depends on your investment objective.

02

Harness affects model performance and characteristics.

On headline metrics the ranking is preserved across both harnesses: total return Muse > GLM > MiMo, IC GLM > Muse > MiMo. On more granular dimensions the harness is the larger driver — especially sector bias and residual IC.

Don’t evaluate the model in isolation. Harness and data also shape investment performance and characteristics.

same model, different harness — correlation
0.000.250.500.751.00GLM 5.3 · XS ratings: 0.62GLM 5.3 · residual IC: 0.42GLM 5.3 · naive-tilt factor: 0.79GLM 5.3 · naive-tilt sector: 0.61GLM 5.3MiMo 2.5 · XS ratings: 0.58MiMo 2.5 · residual IC: 0.47MiMo 2.5 · naive-tilt factor: 0.78MiMo 2.5 · naive-tilt sector: 0.56MiMo 2.5Muse 1.3 · XS ratings: 0.55Muse 1.3 · residual IC: 0.54Muse 1.3 · naive-tilt factor: 0.75Muse 1.3 · naive-tilt sector: 0.67Muse 1.3
XS ratingsresidual ICnaive-tilt factornaive-tilt sector

03

Model capability is just one driver of performance.

Model capability (intelligence) is just one driver of performance. It is also heavily driven by reliability and hallucination (Omniscience), and by the harness of choice.

Picking the right model often means balancing intelligence and reliability (frontier models tend to have higher hallucination rates). Suitability to the harness should also be assessed.

IR vs AA Intelligence Index

OLS y = 0.133 + 0.027 x · R² 0.62 · n=6

0.821.041.261.481.6924.4030.2036.0041.8047.60IRAA Intelligence IndexGrok 4.6: AA Intelligence Index 43.00, IR 1.27Grok 4.6GLM 5.3: AA Intelligence Index 45.00, IR 1.63GLM 5.3MiMo 2.5: AA Intelligence Index 26.00, IR 0.88MiMo 2.5Muse 1.3: AA Intelligence Index 45.00, IR 1.22Muse 1.3DeepSeek v4.1: AA Intelligence Index 40.00, IR 1.05DeepSeek v4.1MiMo 2.6: AA Intelligence Index 46.00, IR 1.37MiMo 2.6
IR vs AA-Omniscience Index

OLS y = 1.131 + 0.009 x · R² 0.19 · n=6

0.821.041.261.481.69-7.641.9311.5021.0730.64IRAA-Omniscience IndexGrok 4.6: AA-Omniscience Index 28.00, IR 1.27Grok 4.6GLM 5.3: AA-Omniscience Index 14.00, IR 1.63GLM 5.3MiMo 2.5: AA-Omniscience Index 3.00, IR 0.88MiMo 2.5Muse 1.3: AA-Omniscience Index 23.00, IR 1.22Muse 1.3DeepSeek v4.1: AA-Omniscience Index -5.00, IR 1.05DeepSeek v4.1MiMo 2.6: AA-Omniscience Index 8.00, IR 1.37MiMo 2.6

04

Agents are better deployed to systematically rate the universe than to pick names out of it.

Holding the model and the OpenClaw harness fixed, systematic stockrate beats fundamental stockpick on information ratio, volatility, and drawdown for both GLM 5.3 and Muse 1.3. Picking a handful stocks concentrates the book and exposes hallucinations and errors.

Adopt implementations that could control for AI-specific risks.

IR
-0.090.420.921.431.94GLM 5.3 · systematic stockrate: 1.84GLM 5.3 · fundamental stockpick: 0.54GLM 5.3Muse 1.3 · systematic stockrate: 1.68Muse 1.3 · fundamental stockpick: 0.95Muse 1.3
systematic stockratefundamental stockpick
vol
-0.8%3.5%7.9%12.2%16.5%GLM 5.3 · systematic stockrate: 9.4%GLM 5.3 · fundamental stockpick: 15.7%GLM 5.3Muse 1.3 · systematic stockrate: 9.0%Muse 1.3 · fundamental stockpick: 10.1%Muse 1.3
systematic stockratefundamental stockpick
max dd
-7.5%-5.6%-3.6%-1.6%0.4%GLM 5.3 · systematic stockrate: -3.9%GLM 5.3 · fundamental stockpick: -7.2%GLM 5.3Muse 1.3 · systematic stockrate: -2.2%Muse 1.3 · fundamental stockpick: -6.3%Muse 1.3
systematic stockratefundamental stockpick

Subscribe for new eval insights

We'll email you when we publish.