Scoping Agent Permissions to Pre-Modeled Metrics

Restrict agent access to pre-modeled datasets instead of raw database credentials.

Contributing Editor · · 10 min read
Cover illustration for “Scoping Agent Permissions to Pre-Modeled Metrics”
Agent-Safe Data Access · October 10, 2026 · 10 min read · 2,234 words

An AI agent that can query a production database directly is dangerous for a reason that has little to do with sloppy settings. The danger comes from how agents inherit permissions from the credentials they use: give an agent a connection string with broad read access, and the agent now has broad read access, whether or not its actual task ever requires more than a handful of columns. Security researchers call this excessive agency, the condition in which an agent holds more authority than the job in front of it demands. That gap between granted permission and needed permission doesn't sit quietly. It waits. A prompt injection, a malformed query, a chained tool call the agent wasn't meant to make: any of these can turn dormant over-permissioning into an actual data exposure, and the agent did not need to be compromised in some dramatic way for that to happen. It just needed to be asked the wrong question at the wrong moment.

Most organizations build their agent infrastructure backward. CData's 2026 governance guide frames this ordering as the mistake: governance, in its account, is a prerequisite that has to exist at the source before anything connects to it, not a layer applied retroactively once an agent is already live. The fix is not a better settings panel or a stricter credential rotation policy. The fix is to stop handing agents database credentials in the first place and to ask instead what the specific agent needs to see.

How MCP Accelerates Agent-to-Data Connectivity

The Model Context Protocol has become the standard way agents reach external data and tools, and that standardization is why its security properties matter at scale. The result is a protocol that behaves more like the rest of the web: cacheable, easier to scale, friendlier to the kind of infrastructure teams already run. That shift is a large part of why MCP adoption keeps accelerating.

Faster adoption does not mean safer adoption. Token passthrough of this kind is a named, recurring pattern that any serious MCP security model has to account for directly. The architecture that research proposes in response is a multi-protocol gateway bridging the Agent Chat Protocol, A2A, and MCP, routing requests through a layer that can at least enforce identity and scope consistently.

A gateway like that solves routing. It does not solve meaning. An agent can be authenticated correctly, scoped correctly, and routed through a gateway that enforces every policy it's configured to enforce, and still land on a raw production table with no definition attached to any of its columns. Fast, correctly authorized access to a table nobody has defined is still unsafe, just unsafe in a different register: the agent now has clean, policy-approved access to data it cannot interpret correctly. That's the semantic gap a transport-layer fix can't close on its own.

A Governed Dataset Is a Permission Decision

A pre-modeled dataset is a permission boundary for analysts, enforced by the simple fact that an agent can only query what the model has explicitly exposed. There is no separate step where someone configures what the agent can see, because the model itself already did that. The blast radius of anything that goes wrong shrinks to exactly the surface the model defines, which is a much smaller and much more legible thing than a production schema.

Research on automating data access permissions makes the architectural case directly: governing agents at the source, before they ever connect, is the principle that reliably contains excessive agency. A structural boundary doesn't depend on anyone remembering to enforce it correctly every time.

The semantic benefit runs alongside the access benefit, and the two reinforce each other. Research on normative infrastructure for the agentic web contrasts access control stapled on after the fact, which depends on every downstream tool respecting rules it could, in principle, ignore, with governance built directly into the data layer, which removes the option to ignore them.

This has a direct consequence for how much damage a compromised agent can do. The root cause of excessive agency is architectural, an agent granted database credentials inherits everything those credentials permit, regardless of what its task actually calls for. Governing at the data layer, by restricting agents to pre-modeled metrics before they ever connect, is how teams using Dreambase avoid the problem before something leaks.

The strongest objection to this approach is that pre-modeling everything is itself a bottleneck. No model can anticipate every question a business will eventually ask, and plenty of the context that matters lives outside any predefined layer, in the judgment calls and edge cases nobody wrote down. That's a fair complaint, but it compares the wrong two costs. An incomplete metrics layer produces a bounded, improvable gap: a question that can't be answered today gets modeled tomorrow. Full table access for agents produces a cost that only grows as agent autonomy grows, because an agent that can read anything can also leak anything, and that kind of exposure doesn't get undone after the fact.

The Architecture in Practice: Modeled Parquet, DuckDB, and a Single MCP Server

Turning that argument into a working system takes three pieces: pre-calculated datasets stored as governed Parquet files, a fast query engine that reads them, and a single MCP server that sits as the only thing agents are ever allowed to touch.

Parquet is the artifact where governance actually lives. Agents never open a connection to the live database. That single fact protects two things at once: production performance, since nothing an agent does can touch the systems serving real traffic, and production schema, since there's no path from an agent's query back into the structure of the tables it was never shown.

DuckDB is the engine that makes querying those files fast enough to be useful. The agent is, in a real sense, querying a snapshot, not a system.

The MCP server is where the permission boundary becomes concrete. Dreambase implements exactly this pattern: agents query modeled datasets through a single MCP server, never the live Supabase database behind it, with the governed datasets as the only interface available to query. There is no second path in. Research on the Redpanda Agentic Data Plane describes a version of the same idea at the infrastructure level, using out-of-band metadata channels to carry security context, policy signals, and audit trails entirely outside the agent's own read and write path, so governance gets enforced invisibly rather than surfacing as something the agent has to navigate around. It's the same gateway-plus-semantic-layer pattern, applied one level down in the stack.

How Supabase's Row Level Security Fits Into This Architecture

Row Level Security in Supabase is a real and necessary control. It is not, by itself, a sufficient permission boundary for agents, and teams that treat it as one are leaving a gap they may not discover until it's already been exploited.

RLS pushes access policy down into PostgreSQL itself. Every query gets evaluated against SQL policies the database enforces directly, so there's no application layer to route around and no middleware to misconfigure. For governing human users who need different visibility into the same tables, that's the correct model, and it works precisely because the enforcement happens at the lowest possible level.

Agents introduce a failure mode that RLS was never built to handle. The service_role key in Supabase bypasses Row Level Security. An agent running under service_role credentials operates with none of the row-level policies in effect, and because that bypass is silent, nothing about the agent's behavior necessarily looks wrong until the data it exposed is already out. RLS was designed around a model of access, logged-in users with row-level differentiation, that doesn't map cleanly onto a service account running autonomous queries at machine speed.

Performance is a second, more mundane concern at agent query volumes. A policy that was invisible under normal dashboard traffic can become a real bottleneck once an autonomous agent is hitting it on a loop.

RLS also has nothing to say about meaning. It controls which rows an agent can see, but it doesn't define what those rows represent. An agent with RLS-scoped access to a raw events table still has to decide for itself what counts as an "active user," a definition a metrics layer exists to fix consistently. The two layers are not competing solutions to the same problem. RLS governs which production rows are allowed to feed into the pre-modeling pipeline in the first place; the metrics layer governs what agents are actually allowed to query once that pipeline has run. Each does a job the other can't.

What scoped agents can actually do once the permission boundary is in place

Once an agent is scoped to a governed metrics layer, its output becomes trustworthy in a specific and checkable way: the numbers it reports are already certified, and for the agent to get one wrong, it has to actively contradict a value sitting in its own context.

Consider the reporting workflow this makes possible. An agent reads from the metrics layer, runs queries against pre-calculated datasets, assembles the structure of a report, and hands the numbers to an LLM that writes the narrative around them. The LLM retrieves and cites metric values from the governed layer rather than generating them itself; a hallucinated number requires the model to contradict a source it already has in front of it, a much rarer failure than generating a number from nothing.

Anomaly detection benefits from the same stability. Because the metrics layer has already fixed what "Revenue" and "MAU" mean, an agent watching for anomalies is checking against a definition that doesn't shift from one run to the next. Dreambase reflects this division of labor directly: KPIs get calculated on a schedule and pushed out to inbox, Slack, and an iOS app, with the agent responsible for delivery and narrative rather than for deciding what the metric is in the first place. That split, metric construction handled upstream, delivery and framing handled by the agent, is what makes the output something a person can actually rely on.

This also makes agents auditable in a way raw-query agents aren't. Auditability here is a property of the governance structure, not a capability the agent happens to have.

None of this removes the value of a human checking high-stakes output before it goes out the door. Board reports and investor updates still benefit from a human approval step. What changes is how fast that review can happen: a reviewer checking a metrics-layer-grounded report is checking the narrative, not reconstructing the underlying numbers from scratch, because those numbers were already correct when the agent picked them up.

The objection that pre-modeling is a bottleneck, and why the trade-off still favors it

The complaint that pre-modeling constrains what agents can answer on demand is legitimate, and dismissing it would understate the real cost of the approach. A metrics layer only covers the KPIs someone thought to define. Research on automating data access permissions acknowledges this tension directly: the more complete the metrics layer, the more useful an agent built on top of it becomes, but keeping a layer complete takes ongoing curation, and curation is its own recurring cost.

The two costs being traded off here are not the same kind of cost. Research on agent permission interfaces frames this as a trade-off between task capability and safety constraints, and the engineering choice that trade-off points toward is to widen the metrics layer deliberately over time, not to grant raw access as a shortcut around the work of modeling something properly.

That difference appears most clearly in how each kind of change gets made. A new query pattern run by an agent with raw table access leaves no such record, and nobody finds out about it until something has already gone wrong.

Implementing This Architecture Without a Dedicated Data Team

Diagram: Three-Stage Architecture: From Raw Database to Governed Agent Access. Visualizes: Show a sequential three-stage implementation path that transforms how agents connect to data.

None of this requires an enterprise data organization to put into practice. A small team already running a Supabase project can build toward this architecture in three stages, starting from the Postgres database they already have.

The first stage is defining the metrics that actually matter. The moment those definitions exist in code rather than in someone's head, the recurring argument over which number is correct disappears for both the humans checking the dashboard and any agent querying the same layer.

The second stage is exposing that layer to agents through a single MCP server, in place of database credentials. Agents connect to that server. They never see the production database behind it.

The third stage is incremental: as new agent workflows come up, the same pattern applies every time. Define the metric in the governed layer first, then expose it. Never grant an agent raw table access as a shortcut to save a modeling step, because that shortcut is exactly the decision that reintroduces the unbounded risk the first two stages were built to remove.

The payoff compounds as the layer grows. Every investor update, every board report, every agent-generated digest draws from the same pre-calculated source, so the team and its agents work from one set of numbers. The permission boundary holds because the structure makes misconfiguration hard to do by accident, not because someone remembered to configure each new agent correctly. As the metrics layer expands, trust in the numbers it produces expands with it, for the people reading the reports and for the agents generating them.

Sources

  1. Towards Automating Data Access Permissions in AI Agents
  2. The Agentic Web Requires New Normative Infrastructure
  3. How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
  4. Infrastructure for the Agentic Web: Gap Analysis and Architecture from the Agentverse Platform
  5. The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane
  6. The 2026-07-28 MCP Specification Release Candidate

More in Agent-Safe Data Access