Real workflows, not toy tasks
We build test sets from the actual work: contracts pulled from live deal processes, filings from active coverage universes, workflows described by practising lawyers and analysts. Combined with established academic benchmarks, this means our evaluations reflect what tools and models actually face in professional use, not what makes them look capable in a demo.







