Skip to main content

Key Concepts

Feature Availability
DataHub Core (OSS)
DataHub Cloud

This page explains the ideas the guide relies on. You can follow the steps without it, but it's a good place to start if the terms are new.

ConceptWhere it appears in the guide
The Context LayerStep 2: Ingest context, Step 3: Generate context
Validating Context Through EvalsStep 1: Define evals
Validating Context Changes (Human in the Loop)Step 3: Generate context
Global vs. Domain-Scoped Data AgentsStep 4: Activate context
Maintaining Context Through FeedbackStep 5: Context feedback & improvement

The Context Layer​

DataHub gives your agents a single context layer with two parts: the technical context that inventories your data, and the business context that explains what it means. Domains organize both into smaller units.

Technical Context​

Technical context is the foundation. When you connect your warehouse, BI tools, and pipelines, DataHub builds an inventory of your data landscape:

  • Assets: tables, columns, dashboards, charts, pipelines, and the queries that run against them
  • Lineage: where each asset comes from and what depends on it
  • Ownership: who is responsible for each asset
  • Usage: how often each asset is used, and by how many people
  • Past queries: SQL that has already run against each table
  • Status: when each asset last changed, and whether it's deprecated

An agent uses these as trust signals, the way an experienced analyst would. They help it tell a curated table from the raw data behind it, prefer the table people actually use over an abandoned copy, reuse proven joins and filters, and cite its sources.

DataHub also uses these signals to rank what your agent sees. When an agent searches, context that's widely used and recently updated ranks higher, deprecated assets rank lower, and published documents rank above drafts. The more clearly your technical context separates good data from bad, the more reliably your agent finds the right answer.

Business Context​

Business context is the semantic knowledge that explains what your data means and how to use it. It includes:

  • Context documents: guides, recipes, runbooks, and examples an agent can follow to give the canonical answer to a business question. A good one reads like a map: "To calculate net revenue retention, start with these tables, join them like this, and exclude trial accounts." Context documents also capture governance and process ("how to request access to customer PII") and guardrails for specific tables ("amounts are in cents"). Relate a document to the tables it describes, and your agent sees it whenever it looks those tables up.
  • Descriptions: plain-language explanations of tables and columns
  • Glossary terms: shared definitions of business concepts, such as "active customer," linked to the columns that implement them
  • Semantic models and metrics: precise metric definitions from tools like dbt, Snowflake, Looker, and Cube

Context documents come from two sources:

  • Your team writes them, or imports them from tools like Notion, Confluence, and GitHub (step 2).
  • Context Generation writes them for you, by capturing the patterns your team already uses to answer important business questions in its query history, BI tools, and semantic models (step 3). These cover the long tail no one has time to document by hand.

Context documents are searched by meaning, so an agent's question finds the right document even when its wording differs. If you've worked with agent skills, the idea is similar: focused know-how that an agent loads only when a question calls for it. The difference is that context documents are written, generated, and reviewed by the people who know the data.

Domains​

A domain is an optional way to organize your data and context into smaller units, such as everything that belongs to Finance. You can assign almost everything to a domain:

WhatHow it joins the domain
Tables, metrics, and semantic modelsAdd them from the domain's page
Context documentsSet the domain yourself, or have Context Generation assign it
EvalsSet the eval's domain
OwnersThe domain's owners, and the owners of its tables, review its context

Domains let you manage context one area at a time, and they make domain-scoped data agents possible. There are three ways to expose a domain to an agent:

  • A custom agent in DataHub, scoped to the domain
  • A scoped MCP server, limited to the domain with a View
  • Instructions that teach your agent how to navigate your domain hierarchy, such as which domain to search for which kinds of questions

The first two limit what the agent sees; the third guides where it looks.

Validating Context Through Evals​

An eval is a real business question paired with the answer an expert would give, usually as SQL. DataHub asks an agent the question, and an AI judge checks whether the agent's answer is equivalent to the expected one.

Together, your evals form a test suite for your context. They give you a baseline before you add any context, show whether each change helps, and catch regressions the day after they happen. That's why the guide starts with them: every later step is measured against your evals. See Context Evals.

Validating Context Changes (Human in the Loop)​

Agents act on what they read, so the people who know the data decide what goes in. DataHub lets anyone propose a change to context, whether it comes from a person, DataHub's AI, your own agents, or Context Generation. The owners of the context involved review it before agents see it.

Reviewers can run evals against a proposed change before approving it, so they learn whether it helps before your agent does. You choose where review is required: for example, generated documents can publish automatically and be verified by evals, or require approval for sensitive domains. See Reviewing Context Changes.

Global vs. Domain-Scoped Data Agents​

The most consequential decision is what kind of agent you're building. The guide calls them global agents and domain agents for short.

Global data agentDomain-scoped data agent
Who uses itAnalysts, data engineers, and data scientistsA business team, such as Finance, Sales Ops, or Marketing
Typical question"Where do we keep clickstream data, and who owns it?""What was net revenue retention for enterprise accounts last quarter?"
What it can seeAll the data and context the person asking can seeOne domain's trusted tables, metrics, and documents

We recommend rolling out domain-scoped agents one domain at a time. A global agent for your data team is easy to add alongside them. You can run many domain-scoped agents side by side, each with its own context, evals, and owners. See Step 4 for how to set up each kind.

Maintaining Context Through Feedback​

Your data and your questions change constantly, so context needs ongoing maintenance. Feedback is how you keep it current.

No eval suite anticipates every question. Once your agent is in real use, it reports the gaps it runs into: context that's missing, incorrect, or conflicting, and answers a user has corrected. These notes collect in one place, so your team can fix each gap once, and add an eval so it stays fixed. Scheduled Context Generation and custom agent tasks can help with upkeep too.

Feedback closes the loop: real usage improves your context, and your evals confirm each improvement. See Context Feedback.