Blog /
Keep the Analysis Log. Do Not Compact the Evidence
Harsh Vardhan Goswami

On this page
At 11 a.m. the agent tries an inner join, drops 4% of orders, and “fixes” it with a left join. At 11:04 it compacts the session because the window is full. At 5 p.m. a VP asks why the number moved. The summary says “we aligned the join.” The dropped keys are gone. You cannot tell whether 4% was test orders or real revenue. You rerun from memory. Memory is wrong.
That is not a model personality issue. It is a storage issue. Production agents hit a context ceiling and compact: older turns become a paragraph so the next call fits. Reasonable for chat. Fatal for analysis. The thing you need later is often the path you rejected—the filter, the timezone, the row count that did not match the dashboard. Summaries drop those details because they looked like noise at compression time.
Academic work in 2026 (including Scroll, arXiv:2608.21690) argues long-horizon agents need addressable history, not a lossy recap. OpenAI’s GPT-6 Astra notes for Codex—searchable prior windows, notes instead of only compaction—are the vendor rhyme. You do not need their harness. You need to stop treating the investigation as a disposable prompt.
Leave with two stores, not one. If you only remember “bigger windows help,” you will still compact the proof.
Two memories, two jobs
| Store | Holds | May compact? | Who it serves |
|---|---|---|---|
| Working view | Instructions, current schema snippet, last tool output | Yes. Evict chatter and huge dumps | The model’s next token |
| Analysis log | SQL text, parameters, row counts, charts, definition notes, rejected paths | No. Append. Point back | The human’s next argument |
The working view is a scratchpad. The log is the court record. Mixing them is how “we already checked refunds” disappears.
Here is what the compacted paragraph usually contains: “Tried inner join, switched to left join, numbers now match the dashboard.” Here is what it usually omits: orders.customer_id was null on 412 rows; those rows were test accounts except 19 that were real; the left join reintroduced the 19 and also reintroduced duplicates from a fan-out on order_items. The dashboard “match” was coincidental on the rounded thousands. The log would have shown row counts at each step. The paragraph cannot.
A bigger context window reduces how often you compact. It does not make compaction safe. Long-context tests still apply: needles move, distractors exist. A log you can grep is cheaper than hoping token 800,000 still attends to grain.
What belongs in the log (and what is vanity)
Keep:
- The question in the user’s words and in SQL (window, timezone, grain).
- Each statement that ran, including failures, timeouts, and “empty result.”
- The definition used, or the explicit sentence “no signed definition.”
- Outputs a reviewer might challenge: aggregates, exception samples, charts.
- Steering: “ignore test accounts” applied at 14:02, not buried in prose.
- The discarded inner join and its 4% drop. Especially that.
Do not keep as if it were evidence:
- The agent’s self-congratulations.
- Full
EXPLAINdumps unless you are debugging one query. - Ten copies of the same schema.
- Entire fact tables in the prompt. Hold frames in the notebook; send projections to the model.
Notebook discipline you can start Monday
You do not need a sandboxed research kernel on day one.
- Treat the notebook as the log. Ordered SQL / Python / markdown blocks are the Event Log if you refuse to delete intermediate blocks that still explain the result.
- Do not overwrite SQL. Duplicate the cell and change it. The discarded statement is evidence. Editors who “clean up” before sharing are destroying the court record.
- Cap what returns to the model. A DataFrame can be large for the human and small in the tool payload (
head, aggregates, samples). - If the product must compact, compact dialogue. Keep a sidecar: query text, hash, row count, timestamp.
- Name the kept query. “Final” is not a name.
net_revenue_aug_v3_left_joinis a name.
If your agent UI only keeps the last answer, you do not have an analysis tool. You have a search box with a confident tone.
A reconstruction test
Before you trust a session product, run this:
- Finish a messy investigation (at least one failed join).
- Wait two hours. Open only the artifact the tool persisted.
- Can a colleague rerun the winning SQL and see the rejected SQL without asking you?
If they cannot see the rejected path, you will lose the 5 p.m. argument. That test fails most chat UIs. It passes a notebook that nobody “tidied.”
Astra’s notes-across-windows feature is interesting because it admits compaction was the wrong default for long work. Copy the admission. You do not need to copy the brand.
Product shape
Quantum Lab is an ordered notebook: the investigation sits beside the result so a teammate can rerun and challenge it. Kole writes into that notebook rather than into a vanishing chat bubble. Keep the blocks. Compact the chatter. Approval workflow still sits on writes; the log is how you explain what you almost ran.
Related: tools, not warehouse paste, agents query at machine volume.
Take this home: append evidence, compact chatter. If you start duplicating SQL cells instead of editing in place, you already upgraded your team. If you want that discipline to be the default surface instead of a personal habit, Quantum Lab is the surface.
Put it to work
Bring the next question into your workflow.
See how Kole brings questions, source data, and review into a shared workflow.
