Evaluation

SyneHQ Analytics Agent Index

v0.2Updated

How well do agents actually do analyst work — and what does it cost per task?

A curated comparison of current frontier models on independent analytics and agentic benchmarks, weighted into a single index and set against the measured cost of running each one.

The index combines 4 evaluations: Vals EMB (Excel Modeling), LiveBench Data Analysis, Terminal-Bench 2.1, and LiveBench Agentic Coding.

Claude Fable 5.1 leads the index at 77.8, ahead of GPT-6 Astra at 76.0 — a 1.8-point gap.

Models compared
11
Component benchmarks
4
Lowest cost per task
$0.80

Ranked results

Every model, ranked

Higher is better

  1. 1Claude Fable 5.177.8
  2. 2GPT-6 Astra76.0
  3. 3GPT-5.6 Sol74.8
  4. 4Claude Opus 574.8
  5. 5Kimi K372.4
  6. 6Claude Sonnet 568.5
  7. 7Grok 4.668.3
  8. 8Claude Opus 4.866.0
  9. 9GLM 5.265.0
  10. 10GLM 5.364.2
  11. 11DeepSeek V451.1
Analytics Agent Index — bar length is the raw value; order follows the metric's better direction. Figures are curated public data, not SyneHQ measurements.

Analytics Agent Index vs Cost per Task

Capability per dollar

X axis
BEST VALUE — HIGH ANALYTICS AGENT INDEX, LOW COST PER TASK51.157.864.571.277.8$0.80$1.69$3.57$7.54$15.93Cost per Task (USD) — lower is better · log scaleAnalytics Agent Index — higher is betterClaude Fable 5.1GPT-6 AstraGPT-5.6 SolClaude Opus 5Kimi K3Claude Sonnet 5Grok 4.6Claude Opus 4.8GLM 5.2GLM 5.3DeepSeek V4
Analytics Agent Index against Cost per Task for every plotted model.
ModelProviderAnalytics Agent IndexCost per TaskProvenance
Claude Fable 5.1Anthropic77.8$15.93public-benchmark
GPT-6 AstraOpenAI76.0$5.83public-benchmark
GPT-5.6 SolOpenAI74.8$6.01public-benchmark
Claude Opus 5Anthropic74.8$6.07public-benchmark
Kimi K3Moonshot AI72.4$3.48public-benchmark
Claude Sonnet 5Anthropic68.5$15.44public-benchmark
Grok 4.6xAI68.3$3.06public-benchmark
Claude Opus 4.8Anthropic66.0$12.06public-benchmark
GLM 5.2Zhipu AI65.0$4.47public-benchmark
GLM 5.3Zhipu AI64.2$3.79public-benchmark
DeepSeek V4DeepSeek51.1$0.80public-benchmark

Figures are curated from public sources and vendor announcements, not SyneHQ measurements.

Leaderboard · v0.2

Full results table

Sort by any column. Filter by provider or model name.

SyneHQ Analytics Agent Index v0.2 — ranked results, sortable by column.
#Source
1
Claude Fable 5.1Anthropic
77.876.7%80.3%85.0%66.1%$15.93public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
2
GPT-6 AstraOpenAI
76.071.7%83.0%87.3%57.3%$5.83public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
3
GPT-5.6 SolOpenAI
74.872.3%79.8%85.8%56.2%$6.01public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
4
Claude Opus 5Anthropic
74.873.6%74.6%84.6%65.2%$6.07public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
5
Kimi K3Moonshot AI
72.466.4%78.7%80.9%62.2%$3.48public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
6
Claude Sonnet 5Anthropic
68.566.3%71.7%74.5%59.4%$15.44public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
7
Grok 4.6xAI
68.362.7%73.9%78.3%57.0%$3.06public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
8
Claude Opus 4.8Anthropic
66.069.4%66.0%71.9%50.5%$12.06public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
9
GLM 5.2Zhipu AI
65.061.5%73.7%67.8%51.8%$4.47public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
10
GLM 5.3Zhipu AI
64.256.3%70.2%71.5%60.9%$3.79public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1LiveBench (2026-06-25 release)
11
DeepSeek V4DeepSeek
51.151.6%50.2%$0.80public-benchmarkVals AI — Excel Modeling BenchmarkVals AI — Terminal-Bench 2.1

Every number on this page is read from an independent third-party leaderboard, never from a vendor's own announcement. These are not SyneHQ measurements. A dash (—) means the figure is not published, not zero.

Score by benchmark

What the SyneHQ Analytics Agent Index is made of

The index score is a weighted average of the component benchmarks below. Every component is a public leaderboard or paper result, so a model only appears on a benchmark its authors actually published — a model missing from a panel was never run on it, and is not scored zero.

Vals EMB (Excel Modeling)

11 of 11 published
  • Claude Fable 5.176.7
  • Claude Opus 573.6
  • GPT-5.6 Sol72.3
  • GPT-6 Astra71.7
  • Claude Opus 4.869.4
  • Kimi K366.4
  • Claude Sonnet 566.3
  • Grok 4.662.7
  • GLM 5.261.5
  • GLM 5.356.3
  • DeepSeek V451.6

0–100 scale · 35.0% of the index

LiveBench Data Analysis

10 of 11 published
  • GPT-6 Astra83.0
  • Claude Fable 5.180.3
  • GPT-5.6 Sol79.8
  • Kimi K378.7
  • Claude Opus 574.6
  • Grok 4.673.9
  • GLM 5.273.7
  • Claude Sonnet 571.7
  • GLM 5.370.2
  • Claude Opus 4.866.0

0–100 scale · 30.0% of the index

Terminal-Bench 2.1

11 of 11 published
  • GPT-6 Astra87.3
  • GPT-5.6 Sol85.8
  • Claude Fable 5.185.0
  • Claude Opus 584.6
  • Kimi K380.9
  • Grok 4.678.3
  • Claude Sonnet 574.5
  • Claude Opus 4.871.9
  • GLM 5.371.5
  • GLM 5.267.8
  • DeepSeek V450.2

0–100 scale · 20.0% of the index

LiveBench Agentic Coding

10 of 11 published
  • Claude Fable 5.166.1
  • Claude Opus 565.2
  • Kimi K362.2
  • GLM 5.360.9
  • Claude Sonnet 559.4
  • GPT-6 Astra57.3
  • Grok 4.657.0
  • GPT-5.6 Sol56.2
  • GLM 5.251.8
  • Claude Opus 4.850.5

0–100 scale · 15.0% of the index

Components and weights

Vals EMB (Excel Modeling)

35.0% of the index

The agent is handed a real financial-modelling task in a spreadsheet and must build or extend the model itself. Template tasks are graded by exact cell match, Scratch tasks against a partial-credit rubric. This is the closest public benchmark to the analyst work an agent actually gets asked to do.

Tasks
0
Weight
0.35
Vals AI — Excel Modeling Benchmark

LiveBench Data Analysis

30.0% of the index

LiveBench's data-analysis category: table reformatting, join prediction and column-type inference over tables drawn from recent Kaggle and Socrata datasets. Questions are refreshed from sources published after the models trained, so a score here is hard to reach by memorisation.

Tasks
0
Weight
0.3
LiveBench — Data Analysis category

Terminal-Bench 2.1

20.0% of the index

89 end-to-end terminal tasks run under the Terminus 2 harness, scored pass@1 with the model required to pass every pytest for a task. Measures whether an agent can actually drive tooling to completion rather than describe how it would.

Tasks
89
Weight
0.2
Vals AI — Terminal-Bench 2.1

LiveBench Agentic Coding

15.0% of the index

LiveBench's agentic coding category, where the model works a task through a tool loop instead of emitting one shot of code. Included at a low weight because warehouse work is mostly code the agent has to run, check and correct.

Tasks
0
Weight
0.15
LiveBench — Agentic Coding category

Index score = 35.0% × Vals EMB (Excel Modeling) + 30.0% × LiveBench Data Analysis + 20.0% × Terminal-Bench 2.1 + 15.0% × LiveBench Agentic Coding. Raw weights (0.35 : 0.3 : 0.2 : 0.15) are shown as percentages of their total, so the 4 components add up to 100% of the index.

Figures are curated public data — vendor announcements and public benchmark leaderboards. They are not SyneHQ measurements.

Reading the index

Methodology

The index is a weighted average of four independent benchmarks: Vals AI's Excel Modeling Benchmark at 35%, LiveBench Data Analysis at 30%, Vals AI's Terminal-Bench 2.1 at 20%, and LiveBench Agentic Coding at 15%. The two analytics components carry 65% between them because that is what this page is about; the agentic components are included because warehouse work is mostly code an agent has to run and correct, not code it writes once.

Only third-party measurements are used. Vendor-reported figures were collected and deliberately excluded — a model's own announcement is not evidence on the same footing as a leaderboard the vendor does not control. Where a vendor published a conflicting figure for the same benchmark, that is noted rather than averaged.

Where a model is missing a component, the index is renormalised over the components it does have, and the row says so. A missing benchmark renders as an em dash, never as zero. DeepSeek V4 is the one model here scored on two components rather than four, so its index is not directly comparable to the rest.

Cost per task is the average USD the agent actually spent completing one Excel Modeling Benchmark task, as measured by Vals AI — not a token-price estimate. It is the real cost of the run, including whatever reasoning the model chose to do. Three other boards publish cost figures over different task sets; they are not interchangeable, so only one is used here.

There is no text-to-SQL component, and that is a finding rather than an omission. Spider 2.0, BIRD-CRITIC, DABstep, LiveSQLBench and the BIRD Data Intelligence Index were all checked: not one of them lists a single model on this page. Those boards take months to validate a submission and DABstep has closed validation entirely, so they currently stop a generation or more behind the models people are actually deploying. The Text-to-SQL Leaderboard page carries that evidence for the models that do appear on it.

Model names are printed as the board prints them. Provider names are normalised to the vendor's own name, so the Vals label 'SpaceXAI' appears here as xAI. Release dates come from each provider's own documentation; where a provider publishes none, the field is left blank rather than inferred.

List prices are the published base tier. Long-context requests are billed higher by OpenAI, Google, xAI and Alibaba once a threshold is crossed, and an analytics agent working a wide schema will cross it routinely — so treat every price here as a floor. DeepSeek figures are peak-rate; off-peak is exactly half.

Frequently asked

Why is there no text-to-SQL benchmark on a page about analytics agents?

Because none of the public text-to-SQL leaderboards list any of these models. Spider 2.0, BIRD-CRITIC, LiveSQLBench and the BIRD Data Intelligence Index were each read in full; their newest entries are a generation or more behind, and DABstep has closed validation pending a v2. Publishing a SQL column here would have meant inventing numbers or silently mixing model generations. The Text-to-SQL Leaderboard page shows what those boards do measure.

Are these SyneHQ's own measurements?

No. Every figure is read from an independent public leaderboard and cited per row. SyneHQ ran none of these evaluations and has no submission on any of these boards.

Why exclude the numbers vendors publish themselves?

They were collected — 163 cited vendor figures across Anthropic, OpenAI, DeepSeek, Moonshot and Zhipu — and then left out of the index on purpose. Vendors choose which benchmarks to report and under which harness, and several published mutually inconsistent figures for their own models across their own pages. A third-party board the vendor does not control is stronger evidence, and mixing the two into one ranking would imply they are equally rigorous.

Does a high index score mean the model will be good on my warehouse?

Not directly. These benchmarks use public spreadsheets, tables and terminal tasks, not your schema, your dialect or your data volumes. The index is a prior, not a prediction. The clearest warning on the page is GPT-6 Astra: first on Terminal-Bench and first on LiveBench Data Analysis, sixth on the Excel modelling benchmark.

Why is cost per task so much higher than the token price suggests?

Because it is measured, not estimated. It is the actual spend to complete one benchmark task, including whatever reasoning the model chose to do. That is why Claude Sonnet 5 costs more per task than Claude Opus 5 despite a list price less than half as high — Opus 5 finishes tasks in fewer tokens.

How often does this page change?

It is regenerated from the source boards rather than hand-edited, so it moves when they do. The Vals boards were read on 2026-09-05 and the LiveBench release used here is dated 2026-06-25. Leaderboards restate scores as evaluation suites are corrected, so treat this as a snapshot with the date attached.

Why do two boards disagree about the same model?

Different task sets, different harnesses, different scoring. GLM 5.2 beats GLM 5.3 on both data-analysis components and loses on both agentic ones — neither board is wrong; they measure different things. That disagreement is a large part of why the index blends four boards rather than trusting one.

What is missing from this page?

Latency and output-token counts, which none of these boards publish per task in a form worth reproducing. Also any model too new or too small to have been picked up by an independent evaluator — absence here means unmeasured, not bad.

Provenance

Every number on this page is read from an independent third-party leaderboard, never from a vendor's own announcement. These are not SyneHQ measurements.

SyneHQ Analytics Agent Index v0.2 · last updated 2026-09-09. All evaluations

Bring the question, the work, and the answer into one governed workspace.