Overview
What this dataset measures
150 financial number retrieval questions (daloopa/finretrieval subset)
Results
Composite ranking across 8 models
Composite score is the unweighted mean of 2 active metrics: Numeric match, LLM judge accuracy.
Claude Opus 4.714.7%
GPT-5 mini6.3%
Claude Opus 4.85.0%
Groq Llama 4 Scout3.3%
Groq Llama 3.3 70B1.3%
Claude Haiku 4.50.7%
Claude Sonnet 4.60.7%
GPT-50.0%
Metric breakdown
How models differ across scoring dimensions
Scores below are the metrics that feed the composite for this dataset.
Numeric match
| Claude Opus 4.7 | 4.0% |
| GPT-5 mini | 2.0% |
| Claude Opus 4.8 | 3.3% |
| Groq Llama 4 Scout | 2.7% |
| Groq Llama 3.3 70B | 1.3% |
| Claude Haiku 4.5 | 0.7% |
| Claude Sonnet 4.6 | 1.3% |
| GPT-5 | 0.0% |
Unweighted mean across test rows. Measures whether extracted numbers match the reference answer.
LLM judge accuracy
| Claude Opus 4.7 | 25.3% |
| GPT-5 mini | 10.7% |
| Claude Opus 4.8 | 6.7% |
| Groq Llama 4 Scout | 4.0% |
| Groq Llama 3.3 70B | 1.3% |
| Claude Haiku 4.5 | 0.7% |
| Claude Sonnet 4.6 | 0.0% |
| GPT-5 | 0.0% |
Unweighted mean across test rows. LLM judge scores answer correctness against the reference.
Failure analysis
Where models go wrong
Representative examples of where models struggle on this benchmark. Each pattern shows a common error type drawn from the evaluation set.
Takeaways