One theme stood out in the talks at this year’s dbt Summit: everyone is assembling context for data agents. This context takes different forms: semantic definitions, documentation, business rules, and vendor-specific configuration.
At Cassis, our goal is to organize this knowledge in an open, tool-agnostic structure that different agents can use. One requirement drives many of our design choices: context should be debuggable.
When an agent gives a wrong answer, where do you fix it? What if the rule it should have applied was already documented, but the agent never saw it?
We distinguish four kinds of failure:
-
Missing information: the context doesn’t contain a rule or definition needed to answer.
-
Incorrect information: a definition is wrong, outdated, or contradicts another.
-
Retrieval failure: the information exists, but the agent didn’t find it.
-
Model error: the agent retrieved the right information but interpreted it incorrectly or generated the wrong query.
Each calls for a different fix. Missing or incorrect information requires updating the context. Model errors require inspecting how the available information was used: clearer instructions, an example, or a more capable model may help.
Retrieval failures raise a different question: how should we organize context so that, when an agent misses something, we know where to investigate?
Giving the agent a path through the context
Several patterns exist to present context to agents. Vector search retrieves chunks based on semantic similarity. Keyword search finds exact terms in files. Both are useful, and their results can be inspected. But when an agent misses a business rule, understanding why it was not retrieved does not necessarily tell the maintainer where to make a fix: should they change the rule’s wording, the surrounding documentation, or the agent’s search instructions? Neither gives the maintainer an explicit decision point where they can act on whether a certain piece of context is retrieved or not.
Data context has its own characteristics as well. Code is well suited to keyword searches: searching a function’s name is the standard way to look for call sites. It is easy to know where to start looking and, more importantly, where to stop. Data context is not as structured: several rules may refer to the same thing while wording it slightly differently. Moreover, it is very hard to know when to stop exploring: how can you know you’re not missing an important rule about how a certain system works, that would change the interpretation of the numbers you are querying in a specific context?
We address this by organizing context into a tree of business domains and requiring the agent to explore only along the branches of that tree. It starts at the root. Each domain it opens contains its own context in Markdown and short descriptions of its children, which help the agent decide where to explore next.
This is essentially recursive progressive disclosure. The general pattern is established: systems such as PageIndex also use trees for retrieval, but while they build a tree over documents, we build a tree of business knowledge alongside the tables and metrics it explains.

The tree constrains the path the agent takes, and it gives us a concrete way to investigate a missed rule: follow the agent’s path and find where it diverged from the relevant domain. When that happens, update the domain description to ensure future queries retrieve it.
For example, imagine a revenue question whose answer depends on a rule in a billing subdomain. If the agent saw the billing description but skipped it, we can inspect whether that description makes the connection clear.
The agent can still take a wrong turn. But we have a visible decision point and a place for the fix in the description, resulting in more repeatable behavior.
Keeping each piece of knowledge in a clear place
Constraining the context discovery path to make it debuggable and fixable is one thing, but there’s still one problem. When the context has missing information, where is it inserted? The context structure should also include a set of rules that specify where to position each piece of context, so that it can’t be accidentally duplicated and drift into diverging rules.
We approach it this way: domains hold the organizational knowledge needed to interpret the data. Each rule lives in the deepest domain that still covers everything it applies to. Table objects, attached to domains, hold descriptions of columns, grain, and joins. Metrics carry definitions of computations and required filters.
This hierarchy lets broad rules live in parent domains. When the agent descends through the tree, it encounters those rules before reaching more specific context. Rules specific to how a table works are positioned at the table level, and retrieved only if the table in question is useful for answering the question.
This gives each fact one home, at the most specific level where it applies. It makes context creation and maintenance straightforward, and it also ensures that we include as little unnecessary information in the context as possible.

Cross-domain joins introduce a complication: they can lead to a useful table without its parent context being present. When this happens, we automatically include the context along the path from the root to that table that’s not already been read, preserving the principle that every table is read within the context of its parent domain, and every domain within its own parent’s context, all the way to the root.
Scaling the context without overloading the agent
As the context grows, it can be refactored into subdomains. The agent sees descriptions of the available branches at each level and opens the ones relevant to the question. Entire unrelated branches stay outside its context window. This pruning enables the agent to remain accurate even on warehouses containing thousands of tables and hundreds of thousands of columns.
For our larger clients, the switch to this domain tree structure had a massive impact: 50% to 80% fewer context tokens loaded on the same questions, with better answer quality and accuracy. Since those tokens depend on the context discovery path, they are not easily cached across conversations.
There is a maintenance tradeoff in domain size: domains that are too small create unnecessary navigation overhead. Domains that are too large fill the agent’s context window with unnecessary information, and are also harder to work with for maintainers. The sweet spot depends on the model, but our recommendation is currently around 10,000 tokens per domain, including all that is loaded with the domain: business rules in Markdown, metrics, table descriptions, and subdomain descriptions.
What about the semantic layer?
The context structure described here works with or without a separate semantic layer.
If a customer has one, its established measures and metrics remain authoritative, and the agent queries through it. Otherwise, those definitions live in the context, and the agent generates SQL against the database directly.
In either setup, the hard part is discovering which definitions apply, understanding the surrounding business rules, and using them correctly. The domain tree organizes that discovery.
Curious to see how it works in practice? Get in touch with us!