Skip to content
syneHQ

Blog /

A Semantic Layer an SMB Can Actually Keep

Harsh Vardhan Goswami

Two dashboards disagree by 8%. Both say “revenue.” The agent, asked to explain, averages them. Finance believes neither. Engineering believes the warehouse. Sales believes the CRM export. Nobody is lying. Nobody named the grain.

Databricks, Hex, dbt, and Salesforce spent 2026 repeating a true sentence: agents need a semantic contract or they will invent metrics. The false sentence that often follows is that your first artifact must be a metrics platform. A ten-person company does not have a metrics-platform group. It has Slack arguments. The fix that actually ships is a short, owned document next to the queries—not a second career.

Leave with a card you can fill this afternoon. If “revenue” still has two owners at 5 p.m., you found the real bug. The agent was only amplifying it.

Three industry words, one SMB job

Primers split context, semantics, and ontology. Collapse them until you have evidence you need the split:

  • Context: which connection, which environment, which tables are in bounds. Staging is not prod.
  • Semantics: how to compute the named metrics the business already fights about.
  • Ontology: shared entities across teams. Skip until two groups share “customer” and still mean different keys.

MotherDuck’s September 2026 primer is explicit: a human still decides what is true. Keep that. Automation that drafts a layer from chat logs can be a draft. It is not true until an owner signs.

The metric card (copy this)

One card per metric. Five to fifteen cards is a layer. Fifty is a job posting.

Name (meeting word): net_revenue
Owner (human): 
Grain: order | line | daily_account | invoice
Clock: created_at | recognized_at
Timezone:
Population excludes: is_test, deleted, internal SKUs
Measure: column + status filter
Joins allowed:
Joins forbidden:
Known lies: columns that look like this metric and are not
Last signed: date
SQL of record: link or saved query id

If you cannot name the owner, you do not have a metric. You have a query someone ran once. Agents will run it forever.

Worked failure: net versus gross

Gross is sum(amount) for rows that existed. Net subtracts refunds, or uses a status set, or uses a recognition date. All three appear in startups. All three get called “revenue” in a board deck.

An agent without a card does the statistically common thing: it finds a column named amount and a table named orders and it sums. If refunds live in order_adjustments with a delay, August “revenue” includes July sales and misses August refunds—or the reverse. The narrative will be “seasonality.” The controller will find the adjustments table in November.

With a card, the agent is not asked to be wise. It is asked to run the SQL of record or to stop. Stopping is a successful session. Inventing is a failed one.

Write the negatives in the same card. They are cheaper than clever recovery:

  • Do not join users.id to events.user_uuid without user_id_map.
  • Do not mix UTC dashboards with company-local “day.”
  • Do not use staging.orders in customer-facing numbers.
  • Do not treat amount_ex_tax as net revenue.

Forbidden joins beat a 40-page happy-path spec that nobody updates.

Where the card must live

A wiki-only definition will not be in the prompt. A prompt-only definition will drift from the warehouse. Put the canonical text next to the queries:

  • Comment on the saved query finance actually runs.
  • A markdown block at the top of the investigation notebook.
  • A short connection guide the agent is allowed to read.

Hex’s agent docs say the same thing in warehouse language: curate what the agent can see; sync descriptions; do not dump every schema. SMB translation: hide staging, document five facts, refresh when DDL lands.

Data Explorer is where people inspect grain and keys before they bless a card. Look at distinct status values. Look at whether customer_id is nullable. Look at a handful of refund rows. Then sign. Kole should inherit those names rather than inventing synonyms. If it cannot find net_revenue, it should ask. Asking is alignment. Guessing is a product bug you will feel in a meeting.

A 90-minute workshop that actually finishes

Do not “kick off a semantic-layer initiative.” Do this:

  1. List the last ten numbers that caused an argument. Those are your candidate metrics. Not a brainstorm of 80.
  2. Fill cards until you stall. Stall is information: two owners, or no grain.
  3. Pick one SQL of record per card. Save it. Run it. Put the row count in the card.
  4. Hide one staging schema from the agent role.
  5. Put the forbidden-join list in the connection guide.

You now have a layer. It is embarrassing how small it is. Small is how it stays true.

When to buy a catalog product

Graduate to dbt MetricFlow, Cube, or a warehouse semantic product when multiple surfaces—BI, notebooks, agents, reverse ETL—must share one calculation and the markdown is already a merge-conflict generator. Until then, a second system is how definitions fork: the wiki says net, the BI tool says gross, the agent averages.

Agents do not make a semantic layer optional. They make a short, owned, negative-aware layer urgent. That is still a document. It does not start as a platform.

If you filled even one card while reading, you already did the work most teams postpone until after the bad board meeting. If you want the SQL of record, the notebook, and the agent to share that card without copy-paste, SyneHQ is built around saved queries, Explorer, and Kole in one workspace. The document still comes first. The product is where it stops being a PDF.

Put it to work

Bring the next question into your workflow.

Keep the query, result, and explanation together when the analysis becomes team work.