These patterns are for the engineers and architects who have to make agents work on real lakehouse data. Each one is a step-by-step build with code, a test that proves it, and a "done when" checklist. They fit together: governed views are what the policies filter, the guardrails run on what the compiler emits, and the MCP server is how the agent reaches all of it.
Agent-safe governed views
Raw and governed schemas, PII left out of the column list, joins written once, and an engine that cannot read files.
2Row and column policies for agents
Roles in reviewed configuration, row filters baked into views, denied metrics and dimensions, and tests that try to escape.
3Query guardrails for agents
Name and relation allowlists, row limits with truncation, timeouts, Iceberg metadata cost caps, session budgets, and an audit log.
4Exposing a semantic layer to agents over MCP
Metric tools instead of a SQL tool, definitions an agent can choose from, errors it can act on, and tests through a real MCP client.
5An MCP server over an Iceberg catalog
PyIceberg catalog and partitioned tables, reading into DuckDB, metadata-driven decisions, and the swap to a REST catalog.
Suggested order
If you are starting from nothing, read them in the order above and run the companion code as you go. If you already have an MCP server in front of your lakehouse, start with query guardrails and check your server against its "done when" list; it is the fastest way to find gaps.
Where they sit in the architecture
The reference architecture has eight labeled layers from object storage to agents. These patterns cover the middle of it: the catalog and table format (pattern 5), the semantic layer (pattern 4), governance (patterns 1 and 2), and the agent interface (patterns 3 and 4).