The index is a weighted average of four independent benchmarks: Vals AI's Excel Modeling Benchmark at 35%, LiveBench Data Analysis at 30%, Vals AI's Terminal-Bench 2.1 at 20%, and LiveBench Agentic Coding at 15%. The two analytics components carry 65% between them because that is what this page is about; the agentic components are included because warehouse work is mostly code an agent has to run and correct, not code it writes once.
Only third-party measurements are used. Vendor-reported figures were collected and deliberately excluded — a model's own announcement is not evidence on the same footing as a leaderboard the vendor does not control. Where a vendor published a conflicting figure for the same benchmark, that is noted rather than averaged.
Where a model is missing a component, the index is renormalised over the components it does have, and the row says so. A missing benchmark renders as an em dash, never as zero. DeepSeek V4 is the one model here scored on two components rather than four, so its index is not directly comparable to the rest.
Cost per task is the average USD the agent actually spent completing one Excel Modeling Benchmark task, as measured by Vals AI — not a token-price estimate. It is the real cost of the run, including whatever reasoning the model chose to do. Three other boards publish cost figures over different task sets; they are not interchangeable, so only one is used here.
There is no text-to-SQL component, and that is a finding rather than an omission. Spider 2.0, BIRD-CRITIC, DABstep, LiveSQLBench and the BIRD Data Intelligence Index were all checked: not one of them lists a single model on this page. Those boards take months to validate a submission and DABstep has closed validation entirely, so they currently stop a generation or more behind the models people are actually deploying. The Text-to-SQL Leaderboard page carries that evidence for the models that do appear on it.
Model names are printed as the board prints them. Provider names are normalised to the vendor's own name, so the Vals label 'SpaceXAI' appears here as xAI. Release dates come from each provider's own documentation; where a provider publishes none, the field is left blank rather than inferred.
List prices are the published base tier. Long-context requests are billed higher by OpenAI, Google, xAI and Alibaba once a threshold is crossed, and an analytics agent working a wide schema will cross it routinely — so treat every price here as a floor. DeepSeek figures are peak-rate; off-peak is exactly half.