Finance · Agent-style number retrieval

FinRetrieval

8 models · 2 composite metrics

8.4 pts
Claude Opus 4.7's lead over #2
15 pts
spread across all nine models
Overview

What this dataset measures

150 financial number retrieval questions (daloopa/finretrieval subset)

Results

Composite ranking across 8 models

Composite score is the unweighted mean of 2 active metrics: Numeric match, LLM judge accuracy.

Claude Opus 4.7
14.7%
GPT-5 mini
6.3%
Claude Opus 4.8
5.0%
Groq Llama 4 Scout
3.3%
Groq Llama 3.3 70B
1.3%
Claude Haiku 4.5
0.7%
Claude Sonnet 4.6
0.7%
GPT-5
0.0%
Metric breakdown

How models differ across scoring dimensions

Scores below are the metrics that feed the composite for this dataset.

Numeric match
Claude Opus 4.74.0%
GPT-5 mini2.0%
Claude Opus 4.83.3%
Groq Llama 4 Scout2.7%
Groq Llama 3.3 70B1.3%
Claude Haiku 4.50.7%
Claude Sonnet 4.61.3%
GPT-50.0%

Unweighted mean across test rows. Measures whether extracted numbers match the reference answer.

LLM judge accuracy
Claude Opus 4.725.3%
GPT-5 mini10.7%
Claude Opus 4.86.7%
Groq Llama 4 Scout4.0%
Groq Llama 3.3 70B1.3%
Claude Haiku 4.50.7%
Claude Sonnet 4.60.0%
GPT-50.0%

Unweighted mean across test rows. LLM judge scores answer correctness against the reference.

Failure analysis

Where models go wrong

Representative examples of where models struggle on this benchmark. Each pattern shows a common error type drawn from the evaluation set.

Takeaways

What this means for deployment

Interactive charts

Cost & workflow analysis