Financial research. Measured against expert judgment.
FrontierFinance tests how AI systems investigate open-ended investment questions: connecting evidence across sources, comparing figures, and synthesizing findings against expert-written criteria.
An open benchmark for frontier financial intelligence, measuring how well AI systems support the full investment workflow, from screening to earnings analysis.
Queries
220
Largest open-source dataset for financial workflows
Use cases
6
Diverse expert-crafted queries spanning the entire investment workflow
FrontierFinance covers the full investment workflow
A professional investor first discovers and screens securities. Then researches to understand them better, builds financial models to evaluate their potential, and finally monitors events that might impact the investment. FrontierFinance reflects this end-to-end workflow.
Active investment workflow
220 queries across the jobs an analyst actually does
1 - Discover
2 - Research
3 - Evaluate
4 - Monitor
Company Research
Analyze a company's business model, competitive position, financial health, and risks.
Query
Annotated 2025-04-09
How does Marqeta make money, and what are the unit economics of the business? Include details from the S-1 document at the time of their IPO if needed.
Expert rubric
44 rubric items · 13 must have
Provides information for Marqeta, Inc. (MQ) stating that the company provides a single, global, cloud-based, open API Platform for modern card issuing and transaction processing.
States that Marqeta, Inc. (MQ) employs a usage-based revenue model tied to processing volume.
Provides information for Marqeta, Inc. (MQ) stating that the company derives the majority of its revenue from Interchange Fees generated by card transactions through its Platform.
Provides information for Marqeta, Inc. (MQ) stating that the company also generates revenue from other processing services, including monthly platform access, ATM fees, fraud monitoring, and tokenization services.
Provides information for Marqeta, Inc. (MQ) stating that the Total Processing Volume (TPV) on the Marqeta Platform for full year 2024 was USD 291.1 billion.
Provides information for Marqeta, Inc. (MQ) stating that the net revenue for full year 2024 was USD 506.99 million.
Provides information for Marqeta, Inc. (MQ) stating that the gross margin for full year 2024 was 69%.
Provides information for Marqeta, Inc. (MQ) stating that the net profit margin for full year 2024 was 5%.
Financial data extraction is the easiest use case to annotate and measure. Every existing benchmark focuses on it. FrontierFinance also covers other use cases that matter just as much and are harder to evaluate.
This benchmark
FrontierFinance
n=220
Comparison benchmarks (Open Source)
FinanceBench
n=150
BigFinanceBench
n=50
vals.ai Finance Agent v2
n=27
Financial Data ExtractionCompany ResearchEarnings & EventsSector, Industry & MacroCoverage & Catalyst MonitoringScreening & DiscoveryOther
Each dot is one example · grey marks examples outside the six use cases shown.
FrontierFinance is harder
FrontierFinance covers a broader range of difficulty and has a higher mean hardness than existing finance benchmarks. This follows directly from its coverage of diverse use cases, which demand a higher level of agent capability.
Interquartile range (25–75%)MedianMean10–90th percentile
How hardness is measured
We pooled all examples, both queries and rubrics, from existing finance benchmarks into a single set, then asked an AI agent to compare them pairwise, judging which of two examples is more difficult to answer and fully satisfy. From these pairwise judgments we fit Bradley-Terry scores, giving each example a relative hardness rating on a shared scale across all benchmarks.
System Performance
With FrontierFinance, we benchmark model performance across three harnesses: the Web Search only harness (●), the open-source Finance Agent v2 harness (◆), and the proprietary Samaya harness (★). For each, the charts below plot overall rubrics qualification rate against cost and latency per query.
Sets the chart’s x-axis
Updated September 28, 2026
★ Samaya System◆ Finance Agent v2● Web Search
Only the top 12 systems are shown in the charts. See the leaderboard below for all systems.
Hover a row to spotlight it in the chart above.
FrontierFinance leaderboard · September 28, 2026
#
System
Reasoning effort (API default)
Score
Avg. latency
Avg. cost
1
★Samaya System (high effort)
–
56.0%
278s
$1.81
2
◆Claude Fable 5.1 · Finance Agent v2
high
55.9%
268s
$6.89
3
★Samaya System
–
52.9%
214s
$0.93
4
◆Gemini 3.8 Flash · Finance Agent v2
medium
50.0%
247s
$1.21
5
◆Claude Fable 5 · Finance Agent v2
high
49.2%
165s
$4.06
6
◆GPT 5.6 Sol · Finance Agent v2
medium
46.8%
171s
$3.03
7
◆Ling 3.0 Flash Fin · Finance Agent v2
–
46.6%
318s
$0.05
8
◆Kimi K3 · Finance Agent v2
max
46.4%
336s
$0.90
9
◆Gemini 3.6 Flash · Finance Agent v2
medium
46.3%
164s
$2.41
10
◆Claude Opus 4.8 · Finance Agent v2
high
45.0%
156s
$2.61
11
◆GPT 5.5 · Finance Agent v2
medium
43.5%
233s
$2.80
12
◆GLM 5.2 · Finance Agent v2
max
42.8%
297s
$0.63
13
◆DeepSeek V4 Pro · Finance Agent v2
high
40.5%
203s
$0.68
14
●Claude Opus 4.8 · Web Search
high
33.0%
54s
$1.47
15
◆Kimi K2.6 · Finance Agent v2
on; level not supported
32.3%
201s
$0.95
16
●Gemini 3.1 Pro · Web Search
high
30.7%
91s
$0.24
17
◆Gemini 3.1 Pro · Finance Agent v2
high
30.5%
222s
$1.72
18
●GPT 5.5 · Web Search
medium
20.7%
192s
$0.55
Notes
For the Finance Agent v2 harness, we adapted the open-source implementation with an added tool call limit of 200 calls and a 300-second tool call time limit for all systems. The same limits were applied to Samaya’s system for a fair comparison.
For every web search harness, we evaluated against the official web-search-grounded APIs from the corresponding model providers.
For all models compared, we evaluated with their thinking mode turned on and the thinking or reasoning level set to their default settings.
All cost calculations count only the agentic model cost and exclude related costs such as the search API and data-index pricing, since precise estimates for those are hard to obtain.
For “GPT 5.5 · Web Search,” the low rubrics score was driven by a substantial share of queries (64 of 220) triggering an AI guardrail and returning a null response.
For “Gemini 3.1 Pro · Web Search,” the cost per query is an underestimate, because the API reports token usage only for the initial input prompt and the final answer tokens.
Compare systems
3 selected
Samaya System (high effort)Claude Fable 5 · Finance Agent v2GPT 5.6 Sol · Finance Agent v2
Performance Breakdown by Use Cases
Macro-averaged rubric qualification rate (%) calculated over all queries for each individual use case.
Samaya System (high effort)Claude Fable 5 · Finance Agent v2GPT 5.6 Sol · Finance Agent v2
Performance Breakdown by Rubric Categories
Micro-averaged rubric qualification rate (%) calculated over all rubrics for each rubric category.
Samaya System (high effort)Claude Fable 5 · Finance Agent v2GPT 5.6 Sol · Finance Agent v2
Queries
220
Rubric items
11,543
Avg / Query
52.5
Loading the question explorer…
Research and resources
Research blog
Follow the team’s work on financial research, reasoning, and evaluation.