These are the publicly available benchmarks the market uses to evaluate AI. They are a starting point. The best evaluations are built on custom datasets grounded in specific workflows. That is what Caliper Lab specialises in.
01
Financial documents and reasoning
7 datasets
FinanceBench
150 examples
Best used to test
Whether a model can answer specific questions from SEC filings using the source document as context.
Open-book QA on real 10-K, 10-Q, and 8-K filings. Designed as a minimum performance standard for financial LLMs. Built by PatronusAI with annotated evidence strings.
Whether a model can handle scanned or image-based financial documents, the format most real-world trade and legacy documents arrive in.
Document visual QA extension of TAT-QA. Questions over page images of financial reports. Relevant for vendors processing scanned or low-quality documents.
Whether an AI agent can retrieve specific numeric values from structured financial databases versus searching unstructured sources.
Agent retrieval benchmark built by Daloopa across 14 agent configurations. Key finding: structured API access outperforms web search by 71 percentage points. Full tool call traces released.
Whether a model correctly classifies sentiment in financial news text, the baseline for any research or monitoring product making tone or signal claims.
Sentiment classification for financial news sentences. Annotated by 16 people with finance backgrounds across three classes. Most widely used financial NLP baseline.
Whether a model can identify and extract specific clause types from commercial contracts, the core task in M&A diligence and contract review.
Contract Understanding Atticus Dataset. 13,000+ expert annotations across 41 clause types. Built over a year by legal experts. Gold standard for contract clause extraction.
Whether a model can classify whether a contract clause supports, contradicts, or is silent on a specific obligation, the core NDA and compliance review task.
Natural language inference over contracts. Three-way classification across NDA text. Used across multiple legal benchmarks as a foundational test.
Whether a model retrieves the right passage from a legal contract before answering, testing retrieval quality separately from generation quality.
Retrieval-augmented generation benchmark specifically for legal contracts. Built by ZeroEntropy to isolate whether the retrieval layer finds the right clause.
Whether a model correctly identifies legal risk in contract clauses, including detecting laziness where models avoid flagging risk rather than analysing it.
Clause-level legal risk identification built on CUAD. Measures correctness (F1), output effectiveness, and laziness rates across 19 models. April 2025.
These are the benchmarks frontier models are evaluated against. Useful as a general capability baseline. Do not mistake high scores here for task-specific performance on professional workflows.
MMLU
14,000+ questions
Best used to test
Broad academic knowledge across 57 subjects including finance, law, and economics. The standard measure of frontier model breadth.
Measuring Massive Multitask Language Understanding. Most frontier models now score above 85%. Useful as a general baseline, not a workflow performance measure.
Multi-dimensional model evaluation across accuracy, calibration, robustness, fairness, and efficiency. The methodology reference for rigorous evaluation design.
Holistic Evaluation of Language Models from Stanford. The gold standard for multi-dimension evaluation methodology.
Whether a model can analyse spreadsheets, financial statements, and business reports across QA, data visualisation, and file generation tasks.
AI Data Analytics Benchmark evaluating LLM agents on document analysis spanning spreadsheets, financial statements, and business reports across three task dimensions.
A single reference point for the full financial AI evaluation landscape across information extraction, QA, forecasting, risk management, and decision-making.
HuggingFace-hosted leaderboard covering 42 financial datasets across 7 categories. Updated continuously by the FINOS Foundation.
Consulting, advisory, and complex multi-document workflows have no meaningful public benchmarks. The best evaluations are built on custom datasets grounded in real tasks. Caliper Lab designs and builds evaluation datasets for AI teams that need to measure what actually matters in their domain.