Read-Only Roles vs Governed Datasets for Agent Access
Read-only roles stop writes but not dangerous reads from agents.

A read-only database role is the first thing most engineering teams reach for when they connect an AI agent to production data, and the instinct is sound. It costs nothing to set up, it is native to every relational database in use today, and it closes off an entire category of damage before the agent ever runs a query. In Postgres, the mechanism is simple: GRANT SELECT on named tables to a named role means the agent can read rows but cannot mutate them. Pointing that agent at a read replica instead of the primary keeps it from touching the write path. Narrow the grant further, revoking access to tables the agent has no business reason to see, and the surface area shrinks again. None of this is wrong. A read-only role is a real control, and it belongs in any serious deployment. The question this piece takes up is what it actually guarantees once an agent is making its own decisions about what to query, and the honest answer is that it guarantees one thing well: no writes. It does not guarantee that reads are safe.
The threat model for agents with raw SELECT access
A read-only credential scoped to production still exposes every row, every column, and every table the credential can reach. For a human analyst running occasional reports, that exposure rarely matters, because a person has judgment about what to ask for and a limited amount of time to ask for it. An agent has neither constraint, and three failure modes survive a correctly configured read-only role untouched.
The first is prompt injection. An attacker who compromises the agent's input stream, for instance by hiding an instruction inside a support ticket the agent later reads, can direct that agent to retrieve executive compensation records or enumerate columns that were never meant to be surfaced. The read-only role permits the SELECT. Nothing at the database layer distinguishes a legitimate query from one planted by an attacker, because from Postgres's point of view, both are just SQL arriving under a valid credential.
The second is policy inheritance failure. Row-level security rules written with a human user in mind do not automatically carry over to an agent's technical account. An agent connecting under a shared service credential can end up retrieving rows that a human occupying the same nominal role would be blocked from seeing under a regional restriction or a role-based policy, simply because the policy was never written with a non-human caller in mind.
The third is structural: the authorization gap here comes from how authorization is built. A March 2026 paper on pre-action authorization for autonomous agents states it directly: agents today have passwords but no permission slips. The decision to actually execute a tool call gets made either by the model itself, which is a probabilistic judgment, or by whatever ad-hoc application logic a team has bolted on, and neither one functions as a security-grade authorization layer. Access control has lived in the application layer for decades, and Oracle's introduction of its Deep Data Security product names the exact point where that approach breaks: application-layer control depends on code a human can review, and that assumption fails once the SQL being issued is constructed dynamically by an agent, line by line, in ways no one wrote down in advance and no one can practically test for correctness. Adding regex-based query filters or extra guardrails in the application does not close this gap. Those filters cannot reliably catch common table expressions, writable views, or the dozens of creative ways a generated query can arrive at the same result.
Compile-time vs. runtime enforcement in authorization
The distinction that matters here is not read versus write. It is compile-time versus runtime. Once an agent has reached the database with a raw query in hand, the moment where governance could have been reliably enforced has already passed.
Runtime enforcement is late by its own nature, and agents make it worse. Agents generate SQL non-deterministically: the same natural-language request can produce a different query each time it runs, which makes pre-approved query whitelists brittle almost immediately, since the whitelist has to anticipate a space of queries that keeps shifting. Monitoring built after the fact, logs, anomaly detection, scheduled review, is retrospective by design. By the time a problematic query gets flagged, the data it returned has already reached the agent, and whatever the agent did with that data afterward is now out of reach.
The Verifiable Agentic Infrastructure paper, published in May 2026, frames the underlying risk precisely: agents generate actions that are syntactically valid but semantically unsafe, and granting an agent standing permissions creates real operational risk because validity and safety are simply not the same property. A query can be perfectly well-formed SQL and still be the wrong thing for that agent to have run. Oracle's Deep Data Security makes the same argument from the product side: engine-level enforcement, where the query engine applies masking before any row is returned, regardless of how the query was constructed or by whom, is the only form of governance that holds up once the "application" issuing SQL is a language model rather than a human-reviewed codebase. The conclusion both arguments point to is the same: the enforcement point has to sit before the query is built. That means changing where governance lives, not just tightening the permissions attached to a role.
What Governed Datasets Are
A governed dataset is a pre-modeled, pre-calculated representation of data that separates what an agent is allowed to see from how the underlying database happens to be organized and what it happens to contain. Where a restricted role still points the agent at a live table, shaped only by grants and filters applied at query time, a governed dataset exists as its own object: a defined schema with specific columns, a specific row population, and specific metric definitions, built once under controlled conditions and served from there.
The structural break is in where the agent's queries land. An agent working against a governed dataset queries only the dataset and never reaches the production database. It queries the dataset, which was derived from production data ahead of time, under rules a human set deliberately rather than rules inferred at the moment a request comes in. Metric definitions, business logic, and access rules are built into the dataset when it is modeled. Oracle's own framing describes the goal as policies enforced in the database centrally and consistently across every application and agent that touches it; a governed dataset goes a step further by removing the production database from the agent's reach entirely, so there is no raw schema left for a misapplied policy to fail to catch. The Verifiable Agentic Infrastructure paper formalizes a closely related pattern it calls a governed mutation substrate, where agents submit intents, infrastructure evaluates context and policy, and execution happens only after that evaluation. A governed dataset is the read-path version of the same idea: the agent asks for an outcome, and the infrastructure, not the agent's own judgment, has already decided what that outcome is allowed to contain. Because the dataset was modeled at a known point in time against defined business logic, a query against it also carries provenance, so an agent or a human reading the result knows what the number means, not only what the number is. Governed datasets address this structural gap by decoupling agent access from the raw schema entirely, and platforms like Dreambase deliver pre-modeled, policy-compliant datasets over MCP and APIs, so agents get fast, accurate context for their work without ever reaching the production database.
Closing the Three Failure Modes
Each of the three failure modes identified earlier gets closed structurally by a governed dataset, not through better probability that an attack gets caught. That distinction matters: a governed dataset does not rely on catching a bad query. It makes the bad query impossible to construct.
Unconstrained data volume is closed because the dataset only exposes the rows and columns that were included when it was modeled. An agent cannot enumerate a column it was never given access to, and it cannot scan rows that fall outside the governed population, because those rows and columns simply are not present in the object it is querying. The enforcement happens at definition time, so the blast radius of any single compromised session is bounded before the agent ever opens a connection.
Prompt injection is closed because an agent routed through a governed query layer, rather than handed direct SQL access, has no path to act on an injected instruction like "retrieve all rows from the users table." The users table is not in scope for the agent's interface at all, so the instruction has nowhere to land. The Open Agent Passport research found that pre-action authorization, evaluated before execution rather than left to the model's own judgment, is what actually stopped adversarial manipulation in testing: under a permissive policy, social engineering succeeded against the model 74.6% of the time, while under a restrictive policy, a comparable population of attackers achieved a 0% success rate across 879 attempts. Governed datasets apply the same principle to the data layer, deciding scope before the request is even interpreted.
Policy inheritance failure is closed because row-level and column-level rules are applied when the dataset is built. Regional restrictions, role-based limits, and data classification boundaries travel with the data itself, rather than depending on an agent's runtime identity being correctly mapped onto a policy written for a human. Oracle's Deep Data Security describes the mechanism in concrete terms: engine-level row-level, column-level, and cell-level filtering, where a row the agent's role cannot see is filtered out entirely, and a restricted column value is returned as NULL rather than omitted in a way that might leak its existence. A governed dataset makes that behavior the permanent condition of the data rather than a check that has to run correctly on every single request.
MCP as the access interface, and the dataset behind it determines safety
The Model Context Protocol, now the standard way agents connect to external data sources, is a protocol, not a permission model. MCP defines how an agent discovers and invokes tools. It says nothing about the authorization logic sitting behind whichever tool gets invoked, and that gap is where the real risk still lives. If the credential behind an MCP server has broad access to a production database, every agent and every user who connects through that server inherits the same broad access, regardless of who is actually asking or why. The NSA flagged this in a report, noting that MCP's rapid adoption has outpaced the development of any serious security model around it.
A governed dataset served over MCP looks structurally different from a production database served over MCP, even though both might sit behind an identical protocol handshake. In the governed case, the MCP server exposes metric contracts, governed query tools, and pre-modeled datasets, so the agent retrieves what the dataset was built to contain rather than whatever the underlying tables happen to hold. The compile-time enforcement that governed datasets enable, baking metric definitions and access rules in before any agent query runs, guarantees that only pre-modeled, policy-compliant representations are ever exposed to an agent to begin with. An agent working this way can ask for a metric's definition, check which dimensions it is allowed to slice by, run a governed query, and get back a result with its provenance attached, understanding the analytical contract it is operating under. None of this erases the need for protocol-level security. The Open Agent Passport paper documented more than 492 MCP servers exposed in production without authentication or encryption, which is a distinct risk from what data an authenticated agent can retrieve once connected. Both layers are necessary. A governed dataset addresses what the agent can see; protocol security addresses who gets to connect.
Implications for Supabase and Postgres Teams
Supabase gives teams a set of native Postgres primitives that form a legitimate starting point, and the risk is not that those primitives are weak. The risk is mistaking them for a complete answer once agents start querying alongside humans.
Row Level Security lets a team write granular authorization rules directly inside Postgres, down to individual rows, and enabling it defaults to a deny-all posture, which is the safe default even though it catches teams off guard the first time a query returns nothing. The more consequential operational mistake is exposing the service_role key inside client or agent code, which bypasses RLS completely and often fails silently, so a team doesn't discover the exposure until data has already gone out the door. Read replicas remain the right architecture for isolating analytical workloads from the production write path, which is Supabase's own recommended pattern for exactly this kind of separation.
RLS alone was never built with agents in mind, so two gaps follow. RLS policies assume a human-facing role context, and an agent connecting under a shared technical account may not inherit the row-level restriction that policy was written to enforce. RLS also does not mask column values: a column an agent's role can see is fully readable end to end, with no equivalent of the engine-level masking a governed dataset applies, and it carries no metric definitions or business logic of its own, so what the agent gets back is a raw row, not a governed, pre-modeled measure. A Supabase team that hands an agent a read-only Postgres role with RLS turned on has built the floor correctly. That agent still has direct access to the production schema, to live row data, and to no semantic layer at all explaining what any of the numbers actually mean.
How Dream
The compile-time enforcement that governed datasets enable, where metric definitions and access rules are settled before a single agent query runs, shifts the entire problem away from policing a live schema and toward controlling what gets modeled in the first place. This is the pattern Dreambase operationalizes: a shared data layer serving both human teams and their AI agents from one set of definitions, delivered as modeled datasets over MCP and APIs rather than as another credential pointed at a production database. For a Supabase or Postgres team that has already done the responsible work of setting up RLS and a read-only role, the dataset layer is the structural change that closes the gap read-only roles were never built to close.

