Financial research. Measured against expert judgment.

FrontierFinance tests how AI systems investigate open-ended investment questions: connecting evidence across sources, comparing figures, and synthesizing findings against expert-written criteria.

An open benchmark for frontier financial intelligence, measuring how well AI systems support the full investment workflow, from screening to earnings analysis.

Queries
220
Largest open-source dataset for financial workflows
Use cases
6
Diverse expert-crafted queries spanning the entire investment workflow
Rubrics / Query
52.5
Comprehensive expert-written annotations
Best system
56.0%
Hard dataset with a large headroom

System quality vs. cost comparison

Desirable Region45%50%55%$0$1$2$3$4$5$6$7Cost / Query (decreasing →)Rubrics Qualification RateSamaya System (high effort): 56.0%, $1.81, 278sSamaya System (high effort)Claude Fable 5.1 · Finance Agent v2: 55.9%, $6.89, 268sClaude Fable 5.1 · FASamaya System: 52.9%, $0.93, 214sSamaya SystemGemini 3.8 Flash · Finance Agent v2: 50.0%, $1.21, 247sGemini 3.8 Flash · FAClaude Fable 5 · Finance Agent v2: 49.2%, $4.06, 165sClaude Fable 5 · FAGPT 5.6 Sol · Finance Agent v2: 46.8%, $3.03, 171sGPT 5.6 Sol · FALing 3.0 Flash Fin · Finance Agent v2: 46.6%, $0.05, 318sLing 3.0 Flash Fin · FAKimi K3 · Finance Agent v2: 46.4%, $0.90, 336sKimi K3 · FAGemini 3.6 Flash · Finance Agent v2: 46.3%, $2.41, 164sGemini 3.6 Flash · FAClaude Opus 4.8 · Finance Agent v2: 45.0%, $2.61, 156sClaude Opus 4.8 · FAGPT 5.5 · Finance Agent v2: 43.5%, $2.80, 233sGPT 5.5 · FAGLM 5.2 · Finance Agent v2: 42.8%, $0.63, 297sGLM 5.2 · FA
Read more about system performance

FrontierFinance covers the full investment workflow

A professional investor first discovers and screens securities. Then researches to understand them better, builds financial models to evaluate their potential, and finally monitors events that might impact the investment. FrontierFinance reflects this end-to-end workflow.

Active investment workflow

220 queries across the jobs an analyst actually does

  1. 1 - Discover
  2. 2 - Research
  3. 3 - Evaluate
  4. 4 - Monitor

Company Research

Analyze a company's business model, competitive position, financial health, and risks.

Query

Annotated 2025-04-09

How does Marqeta make money, and what are the unit economics of the business? Include details from the S-1 document at the time of their IPO if needed.

Expert rubric

44 rubric items · 13 must have

  • Provides information for Marqeta, Inc. (MQ) stating that the company provides a single, global, cloud-based, open API Platform for modern card issuing and transaction processing.
  • States that Marqeta, Inc. (MQ) employs a usage-based revenue model tied to processing volume.
  • Provides information for Marqeta, Inc. (MQ) stating that the company derives the majority of its revenue from Interchange Fees generated by card transactions through its Platform.
  • Provides information for Marqeta, Inc. (MQ) stating that the company also generates revenue from other processing services, including monthly platform access, ATM fees, fraud monitoring, and tokenization services.
  • Provides information for Marqeta, Inc. (MQ) stating that the Total Processing Volume (TPV) on the Marqeta Platform for full year 2024 was USD 291.1 billion.
  • Provides information for Marqeta, Inc. (MQ) stating that the net revenue for full year 2024 was USD 506.99 million.
  • Provides information for Marqeta, Inc. (MQ) stating that the gross margin for full year 2024 was 69%.
  • Provides information for Marqeta, Inc. (MQ) stating that the net profit margin for full year 2024 was 5%.
+36 more rubric items

FrontierFinance is more diverse

Financial data extraction is the easiest use case to annotate and measure. Every existing benchmark focuses on it. FrontierFinance also covers other use cases that matter just as much and are harder to evaluate.

This benchmark

FrontierFinance

n=220

Comparison benchmarks (Open Source)

FinanceBench

n=150

BigFinanceBench

n=50

vals.ai Finance Agent v2

n=27
Financial Data ExtractionCompany ResearchEarnings & EventsSector, Industry & MacroCoverage & Catalyst MonitoringScreening & DiscoveryOther

Each dot is one example · grey marks examples outside the six use cases shown.

FrontierFinance is harder

FrontierFinance covers a broader range of difficulty and has a higher mean hardness than existing finance benchmarks. This follows directly from its coverage of diverse use cases, which demand a higher level of agent capability.

-8-6-4-2024681012FinanceBenchIslam et al., 2023n=150BigFinanceBenchpublicn=50Finance Agent v2publicn=27FrontierFinancen=2203.7Hardness score (higher is harder)
Interquartile range (25–75%)MedianMean10–90th percentile

How hardness is measured

We pooled all examples, both queries and rubrics, from existing finance benchmarks into a single set, then asked an AI agent to compare them pairwise, judging which of two examples is more difficult to answer and fully satisfy. From these pairwise judgments we fit Bradley-Terry scores, giving each example a relative hardness rating on a shared scale across all benchmarks.

System Performance

With FrontierFinance, we benchmark model performance across three harnesses: the Web Search only harness (●), the open-source Finance Agent v2 harness (◆), and the proprietary Samaya harness (★). For each, the charts below plot overall rubrics qualification rate against cost and latency per query.

Sets the chart’s x-axis
Updated September 28, 2026
Desirable Region45%50%55%$0$1$2$3$4$5$6$7Cost / Query (decreasing →)Rubrics Qualification RateSamaya System (high effort): 56.0%, $1.81, 278sSamaya System (high effort)Claude Fable 5.1 · Finance Agent v2: 55.9%, $6.89, 268sClaude Fable 5.1 · FASamaya System: 52.9%, $0.93, 214sSamaya SystemGemini 3.8 Flash · Finance Agent v2: 50.0%, $1.21, 247sGemini 3.8 Flash · FAClaude Fable 5 · Finance Agent v2: 49.2%, $4.06, 165sClaude Fable 5 · FAGPT 5.6 Sol · Finance Agent v2: 46.8%, $3.03, 171sGPT 5.6 Sol · FALing 3.0 Flash Fin · Finance Agent v2: 46.6%, $0.05, 318sLing 3.0 Flash Fin · FAKimi K3 · Finance Agent v2: 46.4%, $0.90, 336sKimi K3 · FAGemini 3.6 Flash · Finance Agent v2: 46.3%, $2.41, 164sGemini 3.6 Flash · FAClaude Opus 4.8 · Finance Agent v2: 45.0%, $2.61, 156sClaude Opus 4.8 · FAGPT 5.5 · Finance Agent v2: 43.5%, $2.80, 233sGPT 5.5 · FAGLM 5.2 · Finance Agent v2: 42.8%, $0.63, 297sGLM 5.2 · FA
★ Samaya System◆ Finance Agent v2● Web Search

Only the top 12 systems are shown in the charts. See the leaderboard below for all systems.

Hover a row to spotlight it in the chart above.
FrontierFinance leaderboard · September 28, 2026
#SystemReasoning effort (API default)ScoreAvg. latencyAvg. cost
1Samaya System (high effort)–56.0%278s$1.81
2Claude Fable 5.1 · Finance Agent v2high55.9%268s$6.89
3Samaya System–52.9%214s$0.93
4Gemini 3.8 Flash · Finance Agent v2medium50.0%247s$1.21
5Claude Fable 5 · Finance Agent v2high49.2%165s$4.06
6GPT 5.6 Sol · Finance Agent v2medium46.8%171s$3.03
7Ling 3.0 Flash Fin · Finance Agent v2–46.6%318s$0.05
8Kimi K3 · Finance Agent v2max46.4%336s$0.90
9Gemini 3.6 Flash · Finance Agent v2medium46.3%164s$2.41
10Claude Opus 4.8 · Finance Agent v2high45.0%156s$2.61
11GPT 5.5 · Finance Agent v2medium43.5%233s$2.80
12GLM 5.2 · Finance Agent v2max42.8%297s$0.63
13DeepSeek V4 Pro · Finance Agent v2high40.5%203s$0.68
14Claude Opus 4.8 · Web Searchhigh33.0%54s$1.47
15Kimi K2.6 · Finance Agent v2on; level not supported32.3%201s$0.95
16Gemini 3.1 Pro · Web Searchhigh30.7%91s$0.24
17Gemini 3.1 Pro · Finance Agent v2high30.5%222s$1.72
18GPT 5.5 · Web Searchmedium20.7%192s$0.55

Notes

  1. For the Finance Agent v2 harness, we adapted the open-source implementation with an added tool call limit of 200 calls and a 300-second tool call time limit for all systems. The same limits were applied to Samaya’s system for a fair comparison.
  2. For every web search harness, we evaluated against the official web-search-grounded APIs from the corresponding model providers.
  3. For all models compared, we evaluated with their thinking mode turned on and the thinking or reasoning level set to their default settings.
  4. All cost calculations count only the agentic model cost and exclude related costs such as the search API and data-index pricing, since precise estimates for those are hard to obtain.
  5. For “GPT 5.5 · Web Search,” the low rubrics score was driven by a substantial share of queries (64 of 220) triggering an AI guardrail and returning a null response.
  6. For “Gemini 3.1 Pro · Web Search,” the cost per query is an underestimate, because the API reports token usage only for the initial input prompt and the final answer tokens.

Compare systems

3 selected
Samaya System (high effort)Claude Fable 5 · Finance Agent v2GPT 5.6 Sol · Finance Agent v2

Performance Breakdown by Use Cases

Macro-averaged rubric qualification rate (%) calculated over all queries for each individual use case.

20%40%60%80%Financial DataExtractionEarnings &EventsCompanyResearchCoverage &Catalyst MonitoringSector, Industry& MacroScreening &DiscoverySamaya System (high effort) · Financial Data Extraction: 59.51%Samaya System (high effort) · Earnings & Events: 75.13%Samaya System (high effort) · Company Research: 56.45%Samaya System (high effort) · Coverage & Catalyst Monitoring: 58.27%Samaya System (high effort) · Sector, Industry & Macro: 38.50%Samaya System (high effort) · Screening & Discovery: 36.24%Claude Fable 5 · Finance Agent v2 · Financial Data Extraction: 55.56%Claude Fable 5 · Finance Agent v2 · Earnings & Events: 63.91%Claude Fable 5 · Finance Agent v2 · Company Research: 40.54%Claude Fable 5 · Finance Agent v2 · Coverage & Catalyst Monitoring: 49.01%Claude Fable 5 · Finance Agent v2 · Sector, Industry & Macro: 38.07%Claude Fable 5 · Finance Agent v2 · Screening & Discovery: 33.28%GPT 5.6 Sol · Finance Agent v2 · Financial Data Extraction: 48.95%GPT 5.6 Sol · Finance Agent v2 · Earnings & Events: 58.03%GPT 5.6 Sol · Finance Agent v2 · Company Research: 42.15%GPT 5.6 Sol · Finance Agent v2 · Coverage & Catalyst Monitoring: 46.73%GPT 5.6 Sol · Finance Agent v2 · Sector, Industry & Macro: 40.94%GPT 5.6 Sol · Finance Agent v2 · Screening & Discovery: 36.52%
Samaya System (high effort)Claude Fable 5 · Finance Agent v2GPT 5.6 Sol · Finance Agent v2

Performance Breakdown by Rubric Categories

Micro-averaged rubric qualification rate (%) calculated over all rubrics for each rubric category.

25%50%75%100%Factual DataExtractionQualitative &ContextualInformationAnalysis &InterpretationComparativeAnalysisForward-LookingInformationFormat &PresentationSamaya System (high effort) · Factual Data Extraction: 46.58%Samaya System (high effort) · Qualitative & Contextual Information: 48.60%Samaya System (high effort) · Analysis & Interpretation: 54.66%Samaya System (high effort) · Comparative Analysis: 48.86%Samaya System (high effort) · Forward-Looking Information: 51.47%Samaya System (high effort) · Format & Presentation: 75.57%Claude Fable 5 · Finance Agent v2 · Factual Data Extraction: 43.15%Claude Fable 5 · Finance Agent v2 · Qualitative & Contextual Information: 41.24%Claude Fable 5 · Finance Agent v2 · Analysis & Interpretation: 49.91%Claude Fable 5 · Finance Agent v2 · Comparative Analysis: 50.19%Claude Fable 5 · Finance Agent v2 · Forward-Looking Information: 48.02%Claude Fable 5 · Finance Agent v2 · Format & Presentation: 74.42%GPT 5.6 Sol · Finance Agent v2 · Factual Data Extraction: 35.70%GPT 5.6 Sol · Finance Agent v2 · Qualitative & Contextual Information: 44.18%GPT 5.6 Sol · Finance Agent v2 · Analysis & Interpretation: 52.51%GPT 5.6 Sol · Finance Agent v2 · Comparative Analysis: 53.03%GPT 5.6 Sol · Finance Agent v2 · Forward-Looking Information: 43.37%GPT 5.6 Sol · Finance Agent v2 · Format & Presentation: 79.39%
Samaya System (high effort)Claude Fable 5 · Finance Agent v2GPT 5.6 Sol · Finance Agent v2
Queries
220
Rubric items
11,543
Avg / Query
52.5

Loading the question explorer…

Research and resources

Research blog

Follow the team’s work on financial research, reasoning, and evaluation.

Research papers

Read the papers behind Samaya’s research and FrontierFinance.

FrontierFinance dataset

Inspect the benchmark’s queries and expert-written rubrics on Hugging Face.