Skip to content
~/reflective_data

blog / rd platform

Why a Semantic Layer Is Essential for Giving AI Access to Your Data

By Jason Dolan · Oct 8, 2026 · 6 min read

RD PlatformTechnical

In short: Ask five people at your company for last month’s revenue and you may get five different numbers. An AI agent will happily give you a sixth. A semantic layer fixes this by defining every metric once, in code, so people, dashboards and AI agents all get the same answer. If you want AI to work with your data, this is the layer to build first.

The problem: AI does not know what your numbers mean

Large language models are very good at writing SQL. What they cannot know is what your business means by “revenue”, “active customer” or “qualified lead”. Is revenue gross or net of refunds? Does it include tax? Which orders count as completed? Is a customer someone who signed up, or someone who paid?

Those answers live in people’s heads, in old dashboards and in scattered queries. When an AI agent is pointed at raw tables, it has to guess. Sometimes it guesses right. Often it produces a confident, well-formatted answer that is quietly wrong, and nobody notices until the number reaches a board slide.

What a semantic layer is

A semantic layer is a governed description of your business on top of your data models. It defines:

  • Metrics: how each measure is calculated, such as revenue, conversion rate or churn
  • Dimensions: the ways you slice them, such as date, channel, country or product
  • Entities and relationships: how customers, orders and sessions connect
  • Ownership and permissions: who is responsible for a definition and who may see the data

Here is what a metric definition can look like in code (illustrative):

metric: revenue
type: simple
measure: sum(order_total)
filter: status = 'completed'
owner: finance
used_by: [dashboards, notebooks, agents]

It is written once, reviewed in git like any other code, and used everywhere. Common ways to build one include dbt with MetricFlow, LookML, SQLMesh and Omni, and we work with all of them.

Why AI specifically needs it

Consistent answers

Without a semantic layer, the same question phrased two ways can produce two different queries and two different numbers. With one, the agent does not write the calculation. It asks for the metric by name and the layer produces the correct query. The answer matches your dashboards because it comes from the same definition.

Fewer hallucinations

A model choosing from a short list of defined metrics and dimensions has far less room to invent a join, pick the wrong column or misread a field. The space of possible mistakes shrinks from “anything in the warehouse” to “something the business has already agreed on”.

Trust you can verify

When an executive asks “where does this number come from?”, the answer should not be “the AI said so”. With a semantic layer, every figure traces back to a named metric, an owner and a tested model. Agents can show their working, and people can check it.

Security and access control

Giving an agent raw warehouse access also gives it everything in the warehouse, including data it should not show. A semantic layer is a natural control point: you expose only the metrics and dimensions the user is allowed to see, and the same rules apply whether the request comes from a person, a dashboard or an agent.

Speed and cost

Agents that explore raw event tables scan far more data than they need, which is slow and expensive. Queries generated from a semantic layer run against modelled, optimised tables, so answers come back faster and cost less.

One layer for every tool

Your analysts use BI, your marketers use spreadsheets, your product team uses notebooks and now your AI assistants use chat. A semantic layer gives all of them the same definitions. Add a new tool or a new AI model later and it inherits the existing logic instead of starting from scratch. This is also why the Reflective Data agent and MCP server read from a semantic layer and not from raw sources.

What goes wrong without one

We see the same patterns again and again when teams connect AI to unmodelled data:

  • Revenue, customers and conversion rate calculated differently in every dashboard, and now a different way again in the chat assistant
  • Agents that quietly include test orders, bot traffic or duplicate events
  • Answers that change from one day to the next because the model made different assumptions
  • No way to audit or explain a number once it has left the tool
  • A pilot that impressed everyone in the demo and was switched off after the first wrong answer in front of leadership

In practice: how we built it for Barbora

Barbora is one of the largest e-commerce businesses in the Baltics. They wanted their teams and their AI tools to work from the same trusted numbers, so we built the foundation described above, step by step, on the Reflective Data platform:

  1. A dbt setup on our platform. We built their dbt project and semantic layer on our hosted dbt, so their business logic and metric definitions are written once, tested and version-controlled.
  2. Interactive dashboards in Explore. On top of that layer we built a set of interactive dashboards using the platform’s Explore feature, so teams can look at the same governed metrics themselves.
  3. An MCP server for their data. We then built an MCP for them that gives AI tools access to the same governed data, through the semantic layer.

Today the whole company uses it where they already work. People ask questions in Slack and directly in Claude, and the same layer powers their weekly stakeholder emails. Because everything reads from one set of definitions, the answer in Slack, the number in a dashboard and the figure in the email all match.

That is the real payoff of a semantic layer. You do not build one integration per tool. You build the foundation once, and every new way of using your data inherits it.

How to build a semantic layer that AI can use

  1. Start with clean models. Raw data is staged, tested and documented in dbt, so the layer sits on a foundation you trust.
  2. Pick the metrics that matter. Begin with the dozen or so numbers leadership and teams use every week. Do not try to model everything on day one.
  3. Agree on definitions with the business. Finance, marketing and product often disagree. Settling it once, in writing, is most of the value.
  4. Define them in code. Metrics, dimensions and relationships live in version control, with review, tests and an owner.
  5. Add descriptions for humans and machines. Clear names and descriptions help AI choose the right metric. Good documentation improves the answers.
  6. Connect your tools. BI, notebooks and AI agents all read from the same layer, for example through an MCP server.
  7. Monitor and improve. Track which questions are asked, where answers are wrong, and extend the layer accordingly.

How we can help

Building this layer is what our data build and semantic layer service is for. We design and build the models, work out the metric definitions with your stakeholders, and connect the layer to your BI tools and AI agents. If you already have a semantic layer in one tool, we can improve it or migrate it. If you would rather not run it yourself, our platform includes hosted dbt and a semantic layer, and our team can run it for you as a managed service.

Once the foundation is in place, our AI services help you put it to work: a company brain, custom agents, bots for your website or Slack, and training for your team.

Want to see how this would look on your data? Request a demo or talk to one of our experts.

newsletter.sh

Join 7,900+ people building with data and AI

Field notes from our experts on tracking, data pipelines, experimentation and what AI is changing in data work. One or two emails a month, no fluff.

// unsubscribe any time
author.md

written by

Jason Dolan

Jason is an expert in digital analytics, data pipelines and data warehouses. Since 2009, he has helped numerous companies become more data-driven and gain full control over their data flows.

We love AI, but this article was written by one of our experts.

discussion

0 comments

Leave a comment

Your email address will not be published.

put it into practice

Run it yourself, or have our experts do it

platform.md

the platform

Pipelines, dbt and exploration in one place

Reflective Data is our data platform: managed pipelines from hundreds of sources into your warehouse, hosted dbt with a semantic layer, and an AI data agent. Use all of it or just one part.

services.md

expert services

Senior specialists for your data stack

Analytics implementation, audits, A/B testing, BI and AI, delivered by the experts who write these articles. Vendor-neutral, for startups and enterprises.