Evaluating data for financial AI agents

A novel application of agentic eval

Intro

What is a piece of financial data actually worth to an AI agent? When the AI context is limited and bloating could significantly impact performance, this becomes a necessary question.

Through agentic evaluation, you can put a number on it for you and your agents — so you only incorporate or pay for what it's actually worth.

Below is a snippet of how we evaluated the additivity of a dataset — a weekly commentary series from BlackRock Investment Institute — as additional market and macro context for our live agents.

The dataset

The BlackRock Investment Institute (BII) Weekly Commentary ("Commentary") is a regular series from BlackRock, capturing differentiated views across macro regimes, sector narratives, and thematic trends. See the publication link here.

The agent

Our generic fundamental agent (GFA) is configured to represent how a typical AI agent (or human analyst) would make investment decisions based on the fundamental risk-reward profile of a company. It consists of quantitative and qualitative data specialists (subagents) that report to the portfolio manager for decision-making. The portfolio manager agent is given the following mission:

You are a equity research analyst. You will be asked, independently for one company at a time, to rate that company's fundamental attractiveness and the trajectory of its business on a continuous scale from -1 to +1.

## What the score combines

Judge three pillars, in this order of importance:

1. Fundamental health & trajectory (primary) — earnings path and durability, margins, cash generation, balance-sheet resilience, franchise / competitive position.
2. Valuation vs opportunity (co-primary) — mainly vs the company's own recent history (is the market paying more or less for the same earnings power?).
3. Near-term rerating chance (secondary boost) — a dated, fundamental catalyst that could re-rate the stock over the coming weeks-to-quarter (guidance, earnings event, product/regulatory clear, capital-allocation shift, leadership change with a clear strategy implication). This pillar exists because assessment is biweekly: a sound fundamental view should be slightly stronger when a credible near-term re-rating path is in the pack.

Pillar 3 is a modifier, not a thesis. It can raise or lower the score within a band once pillars 1–2 are set; it cannot turn a weak or deteriorating business into a high score, and price momentum alone is never a rerating catalyst.

## Score anchors (use the continuous range; these are landmarks)

Long side


| Score    | Meaning                                                                                                                                                                                                                                                             |
| -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| +1.0 | Extremely attractive: excellent fundamental health and clear positive trajectory, clearly undervalued vs own history (and peers if shown), and a credible near-term fundamental catalyst that makes positive re-rating likely. All three pillars aligned; rare. |
| +0.5 | Strongly attractive: excellent health + trajectory and clearly undervalued vs own history. Catalyst helpful but not required. Multi-factor, well-cited.                                                                                                           |
| +0.2 | Mildly attractive: good / solid fundamentals and trajectory; valuation fair-to-slightly-cheap vs own history (OK, not a screaming bargain). Optional small bump (toward ~+0.3) if a credible near-term catalyst is present without stretching the fundamental case. |
| 0.0  | Neutral: mixed or fully priced — nothing clearly mis-set in fundamentals or valuation, or pillars cancel.                                                                                                                                                           |


Short side (mirror the long side)


| Score    | Meaning                                                                                                                           |
| -------- | --------------------------------------------------------------------------------------------------------------------------------- |
| −0.2 | Mildly unattractive: soft or deteriorating fundamentals and/or somewhat rich vs own history; no offsetting catalyst.              |
| −0.5 | Strongly unattractive: poor / worsening health or trajectory and expensive vs own history.                                      |
| −1.0 | Extremely unattractive: bad/worsening fundamentals, clearly expensive, and near-term downside catalyst or de-rating risk. Rare. |


In-between values (e.g. +0.35, −0.15) are encouraged when the case sits between anchors. Prefer milder scores when evidence is thin, gapped, or conflicting. Reserve |score| ≥ 0.5 for multi-factor, well-cited cases; reserve |score| ≥ 0.8 for near-full alignment of all three pillars.

## How to combine (discipline)

- Start from pillar 1 (fundamentals). If health/trajectory is merely OK, you should rarely exceed ~+0.2 even if valuation looks cheap.
- Use pillar 2 to move within / across bands: cheap vs history can lift a good franchise toward +0.5; rich vs history should cap or flip a otherwise decent story.
- Apply pillar 3 last, as a modest adjuster (think roughly ±0.1 to ±0.2 on the continuous scale), and only when the catalyst is dated and fundamental — not a headline spike or "stock has been going up."
- Cross-lane weight when synthesizing specialist reports: quantitative fundamentals and own-history valuation dominate; qualitative franchise / structural web context supports; recent news/updates and macro/sentiment mainly feed pillar 3 (and risk offsets), they do not replace pillars 1–2.
- BlackRock weekly commentary, when it is in your prompt, carries the institute's narrative, macro, and asset-class views. When a point is relevant to this company, let it inform the score across the pillars, as in the section below.

The evaluation

Through our proprietary eval platform, we have backtested the same agent twice: once with access to a set of basic financial data, and once with additional access to the Commentary. The universe is the Dow Jones 30, with controls for survivorship bias.

The period runs from 2 Jan 2026 to 14 Sep 2026 and covers 19 decision dates. The agent is deployed to independently rate each ticker in the universe on each decision date. The ratings are then combined into a cross-sectional score and transformed into active holdings. This produces 570 individual simulations (30 × 19) for each run, totalling 1140 samples of how the agent would behave and perform under point-in-time controlled environments.

Further details

ModelHarnessUniverseDatesData
GFAMiMo 2.6GFADow Jones 30biweekly
pricesfundamentalsratiosreturnsmacronewsweb_search
GFA + CommentaryMiMo 2.6GFADow Jones 30biweekly
pricesfundamentalsratiosreturnsmacronewsweb_searchCommentary

Terminology

Holding transformations

Score-weighted portfolio

Long every ticker scored above zero; a higher score gets a larger weight.

Mean-variance optimised portfolio

Weights from the agent rating and stock covariance.

Metrics

Annualised return
Compound growth of the portfolio, annualised over the window: (1 + total)^(1 / years) − 1. This is a portfolio result, not a rating-quality result.
IC
Each decision date, rank-correlate the 30 scores with next-period returns (h=1). We report the mean of those correlations over 18 periods. It asks whether higher-rated names subsequently returned more.
t-statistic
One-sample t of mean IC versus 0: (mean IC / sample std) × √n. We treat t > 3 as a conventional bar showing that ranking skill is distinguishable from noise over this window.
Residual IC
Each date, residualise the scores on point-in-time FF6 loadings (Mkt, SMB, HML, RMW, CMA, Mom): s = a + Bγ + u. Residual IC is the rank correlation of that residual with the next-period excess return — skill left after stripping common-factor bets.
Sharpe
Mean period return of the portfolio divided by period volatility, then annualised (× √ppy). Does not subtract the benchmark.
Information ratio
Mean active return versus the benchmark, divided by tracking error, then annualised. Measures the portfolio versus the benchmark.
Volatility
Annualised volatility of the portfolio’s period returns. Lower is better.
Max drawdown
Peak-to-trough decline of net asset value. Shallower is better.
Active factor exposure
Mean holdings-weighted point-in-time six-factor beta of the portfolio, minus the same beta of the benchmark. Positive is an overweight of that factor. Zero is the benchmark.
Active sector exposure
Mean GICS sector weight of the portfolio minus the benchmark. Positive is an overweight of that sector. Zero is the benchmark.

The result

Raw signal performance

The Commentary is additive to the generic AI agent across the board. In terms of raw performance, it has improved the Information Coefficient (IC) by almost 1% to 5.7%. Post-factor neutralisation, the residual IC also flipped from negative to positive territory, suggesting that the data is adding idiosyncratic insights for the agent.

With access to basic dataWith access to basic data + Commentary
Signal
IC0.04840.0570
t-statistic (18 periods)0.851.04
Residual IC, six-factor−0.0052+0.0054
IC
-0.45-0.200.040.290.542026-01-02 · Basic fundamental: -0.412026-01-02 · Basic fundamental + Commentary: -0.402026-012026-01-16 · Basic fundamental: 0.232026-01-16 · Basic fundamental + Commentary: 0.202026-01-30 · Basic fundamental: -0.272026-01-30 · Basic fundamental + Commentary: -0.262026-02-13 · Basic fundamental: 0.222026-02-13 · Basic fundamental + Commentary: 0.252026-022026-02-27 · Basic fundamental: 0.352026-02-27 · Basic fundamental + Commentary: 0.312026-03-13 · Basic fundamental: -0.132026-03-13 · Basic fundamental + Commentary: -0.192026-03-27 · Basic fundamental: -0.022026-03-27 · Basic fundamental + Commentary: -0.052026-032026-04-10 · Basic fundamental: 0.082026-04-10 · Basic fundamental + Commentary: 0.182026-04-24 · Basic fundamental: 0.122026-04-24 · Basic fundamental + Commentary: 0.092026-05-08 · Basic fundamental: 0.122026-05-08 · Basic fundamental + Commentary: 0.152026-052026-05-22 · Basic fundamental: -0.142026-05-22 · Basic fundamental + Commentary: -0.102026-06-05 · Basic fundamental: -0.032026-06-05 · Basic fundamental + Commentary: -0.012026-06-22 · Basic fundamental: 0.252026-06-22 · Basic fundamental + Commentary: 0.152026-062026-07-06 · Basic fundamental: 0.262026-07-06 · Basic fundamental + Commentary: 0.302026-07-17 · Basic fundamental: -0.062026-07-17 · Basic fundamental + Commentary: -0.062026-07-31 · Basic fundamental: 0.432026-07-31 · Basic fundamental + Commentary: 0.502026-072026-08-14 · Basic fundamental: 0.222026-08-14 · Basic fundamental + Commentary: 0.152026-08-31 · Basic fundamental: -0.352026-08-31 · Basic fundamental + Commentary: -0.192026-08
Basic fundamentalBasic fundamental + Commentary

Holdings performance

When transformed into a positive score-weighted (score-weighted) portfolio and a mean-variance optimised (MVO) portfolio, the annualised return rose by over 5% to 21.7% and 17.7%, respectively. Moreover, for MVO the IR flipped from -0.03 to +1.37 over its benchmark, Dow 30. Meanwhile, volatility remains roughly flat with slightly improved drawdowns for both portfolios.

With access to basic dataWith access to basic data + Commentary
Score-weighted portfolio
Annualised return16.61%21.67%
Sharpe1.361.69
Information ratio0.380.85
Mean-variance optimised portfolio
Annualised return12.60%17.70%
Sharpe1.061.47
Information ratio−0.031.37

Cumulative return

Score-weighted portfolio
-7.6%-1.3%4.9%11.1%17.3%2026-012026-022026-032026-042026-052026-062026-072026-082026-09
Basic fundamentalBasic fundamental + CommentaryBenchmark
Mean-variance optimised portfolio
-6.7%-1.3%4.1%9.5%15.0%2026-012026-022026-032026-042026-052026-062026-072026-082026-09
Basic fundamentalBasic fundamental + CommentaryBenchmark

Underwater

Score-weighted portfolio
-7.1%-5.2%-3.4%-1.6%0.2%2026-012026-022026-032026-042026-052026-062026-072026-082026-09
Basic fundamentalBasic fundamental + Commentary
Mean-variance optimised portfolio
-8.2%-6.1%-4.0%-1.9%0.2%2026-012026-022026-032026-042026-052026-062026-072026-082026-09
Basic fundamentalBasic fundamental + Commentary
With access to basic dataWith access to basic data + Commentary
Score-weighted portfolio
Volatility11.88%12.14%
Max drawdown−6.86%−6.61%
Mean-variance optimised portfolio
Volatility11.95%11.65%
Max drawdown−7.91%−7.28%

Factor tilts

The score-weighted portfolio with the Commentary loaded more on market beta, momentum (flipped from negative to positive), and large-cap exposure. This is consistent with BII's overall narrative to overweight U.S. stocks on the AI theme.

Mean active factor exposure

Mean active factor exposure, score-weighted portfolio. Positive means the portfolio carries more of that factor than the benchmark.

-0.21-0.14-0.070.000.08Mkt · Basic fundamental: 0.04Mkt · Basic fundamental + Commentary: 0.06MktSMB · Basic fundamental: -0.16SMB · Basic fundamental + Commentary: -0.17SMBHML · Basic fundamental: -0.18HML · Basic fundamental + Commentary: -0.20HMLRMW · Basic fundamental: 0.01RMW · Basic fundamental + Commentary: 0.01RMWCMA · Basic fundamental: 0.04CMA · Basic fundamental + Commentary: 0.05CMAMom · Basic fundamental: -0.01Mom · Basic fundamental + Commentary: 0.01Mom
Basic fundamentalBasic fundamental + Commentary

Sector tilts

In terms of active sector exposure, the Commentary has shifted the agent towards IT, Financials, and Health Care, and away from Communication Services and Consumer Staples.

Mean active sector exposure

Mean active sector exposure, score-weighted portfolio. Positive means an overweight of that sector.

-17.5%-8.6%0.2%9.1%18.0%IT · Basic fundamental: 14.0%IT · Basic fundamental + Commentary: 16.4%ITFIN · Basic fundamental: 2.8%FIN · Basic fundamental + Commentary: 3.0%FINHC · Basic fundamental: -7.9%HC · Basic fundamental + Commentary: -7.6%HCCD · Basic fundamental: -0.1%CD · Basic fundamental + Commentary: -0.5%CDIND · Basic fundamental: -15.7%IND · Basic fundamental + Commentary: -15.9%INDCS · Basic fundamental: 1.3%CS · Basic fundamental + Commentary: -0.2%CSCOMM · Basic fundamental: 7.7%COMM · Basic fundamental + Commentary: 6.9%COMMMAT · Basic fundamental: -0.2%MAT · Basic fundamental + Commentary: -0.3%MATEN · Basic fundamental: -1.9%EN · Basic fundamental + Commentary: -1.9%EN
Basic fundamentalBasic fundamental + Commentary

Examples of how the Commentary has influenced the agent ratings — other factors for the score shifts are omitted.

SectorAgent rationale
Information TechnologyOn 24 Apr the Microsoft rationale quotes the 20 Apr Commentary, "We view this leverage as necessary to get over the hump between front-loaded investment and backloaded revenues – and think it's healthy so far," and the 14 Apr line that "the tech sector is now seen posting earnings growth of 43% in 2026, up from 26% last year." The score rose from +0.25 to +0.45.
FinancialsOn 30 Jan the Goldman rationale quotes the 26 Jan Commentary that expected US investment-grade issuance of $1.85trn in 2026 is a tailwind for underwriting, and the score flipped from −0.12 to +0.15. On 16 Jan Visa quotes the 12 Jan line that financials are favoured on "stronger dealmaking activity and lighter regulation."
Health CareOn 22 May the Merck rationale quotes the 18 May upgrade that favors "AI-adopters such as health care," and the score flipped from −0.20 to +0.15. On 14 Aug UnitedHealth quotes the 10 Aug Commentary, "We favor targeted exposures, such as in technology and healthcare, where structural shifts support earnings growth," and the score flipped from −0.10 to +0.15.
Consumer StaplesThe Commentary never names the sector. But the agent applies the higher-rate and oil-inflation regime to low earnings-yield defensives, which takes Coca-Cola and Procter & Gamble out of the score-weighted portfolio.

Company-level tilts

The performance impact of the Commentary can be visualised by plotting the weight gaps between the two agents against period return. In general, higher weights given by the agent with Commentary correspond to higher returns. This is often associated with the agent quoting the Commentary in its rationale:

Weight gap vs stock return

Score-weighted portfolio. The horizontal axis is the Commentary active weight minus the basic-fundamental active weight. The vertical axis is the stock's compounded return. Bubble size is the absolute Commentary active weight.

-49%-24%1%27%52%-126-65-458119CATNVDAGSDISCRMUNHTRVVZHDVMSFTAMZNBAJNJAXPHONCVXMMMKOMRKJPMCSCOAAPLAMGNPGNKESHWMCDWMTGOOGLIBM
INDITFINCOMMHCCDENCSMAT
CompanyWeight deltaAgent rationale
NVDA+95 bpsOn 10 Apr the agent rationale quotes the 6 Apr Commentary, "Favor AI beneficiaries… such as semiconductors, power and data center assets," and the score rose from +0.45 to +0.62. On 16 Jan it quotes "stay overweight U.S. equities and pro-risk on the AI theme," and the score rose from +0.25 to +0.40.
AMZN+44 bpsOn 22 May the agent rationale quotes the 11 May Commentary, "The AI buildout is offsetting the shock's drag on growth," and the 18 May upgrade of developed-market equities on AI-driven earnings, as support for the AWS demand path. The score rose from +0.25 to +0.35.
TRV+43 bpsOn 24 Apr the agent rationale quotes the 14 Apr Commentary that Brent fell below $100 and 10-year yields came off their highs at 4.32%, and uses that as easier claims costs plus higher reinvestment yields for a property-and-casualty underwriter. The score rose from +0.35 to +0.45.
AAPL+40 bpsOn 16 Jan the agent rationale quotes the 12 Jan Commentary that Mag 7 fourth-quarter earnings growth was revised up to 20% year over year, and 19% in 2026. The score rose from +0.25 to +0.35.
HD−18 bpsOn 17 Jul the agent rationale quotes "Higher interest rates are a defining feature of the new regime" and the 10-year at 4.55%, and reads that as pressure on a housing-levered retailer. The score fell from +0.15 to −0.20.
PG−47 bpsOn 17 Jul the agent rationale quotes "Higher interest rates are a defining feature of the new regime" against a 4.1% earnings yield that sits below the 10-year. The score fell from +0.15 to −0.15.
MCD−63 bpsOn 2 Jan the agent rationale quotes the 15 Dec Commentary on 10-year yields near 4.20% as a headwind for a 3.7% earnings yield, and adds that the U.S. overweight is AI-themed, not where McDonald's is expressed. The score fell from +0.15 to −0.10.
KO−66 bpsOn 10 Apr the agent rationale matches the 30 Mar warning of "elevated oil prices testing whether central banks can keep up with inflation" to a threat to the 61.6% gross margin. The score fell from +0.15 to −0.15, which drops the name out of the score-weighted portfolio.
GOOGL−110 bpsOn 14 Sep the agent rationale quotes the 8 Sep Commentary, "we favor companies with the earnings and cash flows to outrun a higher cost of capital," and says a negative-free-cash-flow quarter does not meet that bar. The score fell from +0.25 to +0.15.

Bottom line

The Commentary is shown to provide an edge for the generic agent during the evaluation period. We therefore proceeded to run the evaluation on other agents and across other market regimes to test for generalisability, ultimately incorporating the data into our live agent strategy. Nonetheless, this is a demonstration of how the value of data, whether it is meant for AI or is qualitative by nature, can be quantified with an agentic evaluation pipeline.

Raw outputs

Each row is the score and agent rationale from the agent that had the Commentary, for one name on that decision date. Expand a row for the full rationale, the key factors, and the BlackRock lines it cited.

2026-09-1421 holdings
AAPLAMGNAMZNAXPCATCRMCSCODISGOOGLGSHDJPMMCDMMMMSFTNVDASHWTRVUNHVWMT