Finance agents stall near ~50% on hard tasks. How much is the data layer?
Give an AI agent the right SEC-data tool and multi-step financial research gets ~2× faster and ~2× cheaper, with no measurable accuracy trade-off — and on building a valuation from filings, 11% → 56%. Every number clicks back to the 10-K.
Point your AI agent at the right SEC-data tool and the slow, error-prone part of financial research — finding the exact number in the right filing — gets dramatically faster and cheaper. In our test, a Claude Sonnet 4.6 agent ran vals.ai's Finance-Agent v2 questions ~2× faster and ~2× cheaper, with no measurable hit to accuracy — and on the single hardest task, building a valuation from a company's filings, it climbed from 11% to 56%. Every figure is cited to the exact line in the 10-K, so you can check the work. Don't take it on faith — watch a real run or connect it free.
You've probably seen the finance score Anthropic now ships with its models — it's measured on an independent benchmark, vals.ai's Finance-Agent. Version 2 is the hard, current one: the multi-step work a research analyst actually does — pull the exact figures out of 10-Ks, build the model, compare the peers — where even the best models clear only about 46% of questions cleanly. That gap is mostly not the model. It's the data layer underneath it. So we held the model fixed (Claude Sonnet 4.6) and toggled one thing: a SEC-data MCP, on and off.
Why would a data tool move the score at all? Because on the hard questions the bottleneck isn't reasoning — it's the data. An agent left to web search and raw EDGAR has to find the right filing, open it, and read the exact figure out of dense HTML, and that's where it slips: it pulls a number from the wrong period, misses a later restatement, reads a press-release summary instead of the audited statement, or can't show where a figure came from at all. MetricDuck hands the agent the structured number with its source attached, so the model spends its budget reasoning, not foraging.
What the benchmark actually asks
These are real questions from the public Finance-Agent v2 set — the same kind our agent ran. They look simple and turn out not to be: each one hinges on the exact figure, from the right period, in the right filing.
| Question (verbatim, abbreviated) | Type | What it takes |
|---|---|---|
| "Who is the CFO of Airbnb (NASDAQ: ABNB) as of April 07, 2025?" | Simple retrieval | A point-in-time fact, current as of a date |
| "How has Netflix's (NFLX) Average Revenue Per Paying User changed from 2019 to 2024?" | Trends | A clean six-year series, one definition |
| "How many basis points did MU beat or miss its Q3 2024 GAAP gross-margin guidance?" | Beat or miss | Guidance vs. actual, to the basis point |
| "Of AMZN, META, or GOOG, who plans to spend the most on capex in 2025?" | Complex retrieval | Latest guidance across three filers |
| "What was the total consideration cost TKO paid for the Endeavor assets…?" | Simple retrieval | One number from a specific transaction |
| "It is March 11, 2025… will TSM beat or miss Q1 2025 guidance, and by how much? Show your work." | Financial modeling | Build a projection from monthly data |
Questions from the vals.ai Finance-Agent v2 public validation set (CC-BY-4.0). Want to watch the agent work one of these end-to-end, every number clickable to the filing? See a real run →
| Metric | Without (web + EDGAR) | With MetricDuck | Difference |
|---|---|---|---|
| Accuracy (all-pass) | 53% | 57% | ≈ same |
| Cost per test | $1.65 | $0.77 | 2.1× cheaper |
| Latency | 495s | 237s | 2.1× faster |
| Agent steps† | 88 | 28 | fewer |
Same model (Claude Sonnet 4.6), same 27 questions from the vals.ai Finance-Agent v2 set, the MCP toggled on and off (N=3 seeds). Accuracy is comparable — its lift confidence interval spans 0, so we say same, not higher. Cost per correct answer: $0.60 with vs $1.21 without.
† “Agent steps” is a harness-side iteration count we're still re-confirming (it conflates the agent's text and tool-call events) — read cost and latency as the robust efficiency signals.
The accuracy numbers sit close on purpose: many questions ask for a fact the model already knows, so a tool can't change the answer and the arms tie. MetricDuck's edge shows up where the work is genuinely hard.
The hardest test: building a valuation from the filings
vals' single hardest category is Financial Modeling — assembling a company's valuation from its filings, where the agent has to pull the exact numbers (the restated figure, the right period, the right segment) and put them together. A free web-and-EDGAR agent gets 11% of these right. The same agent, with MetricDuck, gets 56% — a 5× jump — because the data arrives pre-structured and every figure is cited to the line it came from.
That's the shape of the whole result: on easy questions a free agent already wins, so the arms tie; MetricDuck pulls ahead exactly where the job is pulling and assembling many exact figures. We lose one — Precedents (finding comparable past deals), where the free agent wins by 33 points — and we show you that too. See all nine categories →
Where this fits
That a data layer makes finance agents better is settled — well-funded players are proving it. The builder's question isn't whether to add one, it's which. The deep, human-verified, global options are enterprise-priced and sales-gated — right for an institution. MetricDuck is the other end: fully automated, US-SEC-focused, free and self-serve, and verifiable in the open. Not deeper or more accurate than the institutional vendors — the one you can try yourself, right now, at no cost, and check on every number.
How we measured it — and how to read it
MetricDuck isn't a model; it's an MCP over SEC filings, transcripts, and IR data. So we measured a system — same model, same questions, the MCP toggled on and off — and the gap is MetricDuck's contribution. Three things to keep us honest:
- A delta, not a leaderboard score. Every number is with-vs-without on one fixed model (Claude Sonnet 4.6), against our own free web+EDGAR baseline — so it isn't comparable to vals.ai's hosted ranking.
- Accuracy is comparable, not proven equal. With 27 questions × 3 seeds the accuracy difference wasn't statistically distinguishable — the lift's confidence interval spans 0. We read that as no accuracy trade-off we could measure, not as proof the two are identical, and not "more accurate" (or "more reliable" — reliability is mixed). The robust wins are cost, latency, and the hard-category capability above.
- Checked, not asserted. A different-family judge (Gemini) re-graded every answer (κ = 0.92), the answer is pulled from the agent's response by rule, and the harness is public so you can re-run it.
See it for yourself: watch a real agent run · full method & numbers · connect MetricDuck to your agent (free tier, no card).
MetricDuck Research
SEC filing analysis and structured context for AI agents in finance