Skip to content

Blog /

Cheap Tokens Still Need Reviewed Evidence

SyneHQ

Picture a Monday variance review. An analytics agent is asked why gross margin dropped two points in August. It runs forty queries, writes a tidy page, and cites three drivers. The model bill for the whole session is smaller than the coffee on the table. Then a finance lead asks which of the forty queries produced the second driver, whether refunds were netted before or after the currency conversion, and why the category totals do not match last month's close. Nobody can answer in under an hour.

The tokens were cheap. The answer was not.

That is the gap we want to look at this week. Jyn's essay "tokens too cheap to meter" made the rounds on Hacker News, and it is worth reading in full. We mostly agree with it. We also think analytics teams should draw a specific conclusion from it: when model spend stops being the constraint, verification becomes the constraint. Plan for that now.

What the essay actually argues

The essay collects evidence that the cost of machine intelligence is falling fast on several fronts at once. It points to GPUs getting more power-efficient, inference engines like vLLM getting more efficient per token, mixture-of-experts and hybrid architectures cutting compute and memory for a given quality, and specialized yes/no classifiers that are far cheaper than generative models. Adding it up, the author estimates roughly two and a half orders of magnitude of cost decline in a year, driven mostly by models becoming about 100x more cost-efficient per task.

Two points in it matter for data teams. First, the author is careful to say that the cost per token of frontier models is not reliably falling; the cost per task is. Second, the essay predicts that the limiting factor for AI use will soon be quality and access, not the sheer number of tokens. It even does a Fermi estimate showing a model turn is already within a few orders of magnitude of running grep or a build on a laptop, and suggests that once models are cheaper than tools, people will put models inside tools.

The essay also lists, among its guesses about what cheap intelligence changes, that the hard part of software moves toward requirements, testing, and interface design, with a possible resurgence in QA roles. That is the line we want to pull on, because in analytics, "testing" means reviewing evidence.

Commenters on Hacker News pushed back on the extrapolation. Several argued that efficiency curves bend and that deterministic tools like grep benefit from the same hardware gains. Fair points. Our argument holds whether token prices fall 100x or 10x.

The tool call in analytics is a warehouse scan

The essay's cost table compares a model turn to tools running on a laptop: grep, parsing HTML, cargo build. In that setting the tool is nearly free and the model is the expensive part.

Analytics flips that. The "tool" an agent calls is a warehouse query, and warehouses bill for work. BigQuery's on-demand pricing is $6.25 per TiB scanned in us-central1, after the first free TiB each month. The essay's own estimate for a GPT-5.6 Luna turn is about 10k tokens at roughly 30 cents per million, a third of a cent.

Now a hypothetical: an agent runs an unfiltered SELECT over a 200 GiB fact table because it forgot the partition column. That one scan costs about $1.22 at on-demand rates, the same as several hundred model turns at the essay's estimate. Do it forty times in an exploratory session, which is not unusual when agents search instead of plan, and the warehouse bill dominates before a single token is metered.

So even if the essay is right about tokens, the agent's most expensive action in analytics is already the query, not the thought. And the query is not even the scarcest resource. The reviewer is.

What falls and what does not

Cost Falls with cheap tokens? Why
Model spend per investigation Yes, sharply This is the essay's core claim, and the evidence is strong
Drafting SQL, prose, and charts Yes Generation is exactly what gets cheaper
Retries and self-critique passes Yes A second or fifth attempt costs little in tokens
Warehouse bytes scanned No Billed by the warehouse per TiB or slot, independent of model price
Query latency and concurrency No Large scans still take seconds to minutes and compete for slots
Metric definitions and data contracts No Someone has to decide and sign what "net revenue" means
Human review of the evidence No, and it gets worse More output per hour means more to check per hour
Cost of a wrong number in a board deck No Trust, restatements, and rework are priced in people, not tokens

The bottom half of the table is where analytics work actually stalls. Cheap tokens make the top half almost free, which means the natural failure is to produce more of it: more drafts, more narratives, more "insights." Unreviewed output is not an asset. It is a queue for someone senior.

The bottleneck moves to verification

Back to the Monday review. With cheap tokens, the agent explored forty paths and summarized the ones it liked. The summary looked better. The review got harder, because the reviewer now has to reconstruct which paths were tried, which were discarded, and why.

This is the practical reading of "quality and access become the limiting factor." In analytics, quality is not the model's eloquence. It is whether a second person can confirm the number from the evidence left behind. If the evidence is a chat transcript that was compacted at 11 a.m., the quality ceiling is set by memory, not by the model.

We made a related argument about routing work by token cost: cheap models for governed lookups, stronger models for multi-step investigations. That still holds as prices fall. What changes is the next line of the budget. When the model line shrinks, the review line becomes the one that decides whether you ship.

What to spend the savings on

If your model bill drops by an order of magnitude, do not bank the savings as "more agent sessions." Reinvest them in checks that make each session cheaper to review.

  1. Run the reconciliation automatically. For every metric the agent reports, have it also run the check against the system of record: the close total, the dashboard tile, the prior month. Cheap tokens make it trivial to add a verification pass that compares the answer to a known number and flags any gap above a threshold you set.
  2. Pay for adversarial passes, not polish. Spend tokens on a second pass whose only job is to find the join fan-out, the double-counted refund, or the timezone mismatch. Do not spend them on polishing tone.
  3. Log every statement, including the rejected ones. Keep the SQL text, parameters, row counts, and failures in order. The discarded join is often the evidence a reviewer needs.
  4. Cap what the agent can scan. Enforce partition filters, row limits on exploratory reads, and byte budgets per session. Token savings mean nothing if the warehouse bill doubles.
  5. Return projections, not dumps. Aggregates and samples go to the model; the full frame stays where a human can inspect it. Long contexts still degrade on distractors, even when they are cheap.
  6. Put a person where the blast radius is real. Reads can be cheap and logged. Writes, schema changes, and anything that persists for the team should go through explicit approval.
  7. Measure review time, not answer time. Track how long it takes a second analyst to confirm an agent's result from its artifacts. If that number rises as your token spend falls, you are producing faster than you can verify.

A rule of thumb (ours, not the essay's): for every dollar the model line saves, ask what check it could buy. A check that saves one senior reviewer ten minutes is almost always worth more than another draft.

Where SyneHQ fits

We built SyneHQ around the bottom half of that table. Quantum Lab is an ordered notebook of SQL, Python, and charts, so an investigation leaves blocks a teammate can rerun rather than a paragraph they have to trust. Kole explores shared business data in plain language and works inside that surface. Saved queries and run history let people return to what ran, and audit trails, coming next on Pro, will record query activity through Syne. The approval workflow routes consequential actions, such as data modifications, schema changes, and agent-proposed notebook edits, through explicit review before they run.

None of that makes review free. It makes review possible at the volume cheap tokens will produce.

The essay closes by suggesting we plan for cheap intelligence, not just cheap compute. For analytics teams, planning for it means deciding now what counts as evidence, where it is kept, and who signs off. The tokens will take care of themselves.

Put it to work

Bring the next question into your workflow.

See how Kole brings questions, source data, and review into a shared workflow.