Move credentials out of callers

Radia is not a secret store and does not replace one. It changes which processes need a secret at all: callers hold an identity, a few narrow workers hold the dangerous credentials, and authority becomes permission to request an operation rather than possession of a token. This guide shows how to make that change incrementally, and what it costs.

Credentials become requests

In most deployments a secret manager provisions a downstream credential to every process that needs the system behind it. Each process is then trusted with whatever that credential can do, and compromise of the process is compromise of the credential.

One secret manager provisioning a downstream credential to every process A secret manager hands a database credential to two services, a payment key to a third, a source-control token to a fourth and an internal API key to a fifth. Each process holds whatever the credential can do. Vault one credential per process service A database credential service B database credential service C payment key service D source-control token service E internal API key
Compromise of any process is compromise of the credential it was deployed with. The blast radius of a process equals the authority of its secret.

The target is not fewer secrets. It is fewer holders, and a different way for application code to express authority.

Callers hold an identity; narrow workers hold the credentials Applications and agents send records to Radia. Radia authorizes each request against grants and hands the work to a narrow capability worker, which holds the one credential its system needs and reads it from the secret manager. application A application B agent C agent D Radia grants decide who may request what, for which content no downstream credential here database worker holds one secret database Vault source-control worker holds one secret forge Vault
The number of secrets is unchanged. The number of processes that hold one falls.
Credentials migrate from callers to capability providers. The caller keeps one identity that authorizes requests. The credential that can act on the external system lives in a worker that does one job.

Nothing below requires a rewrite. Each transition is independently useful and can leave the previous path working while it is evaluated. Stop where the next boundary does not justify its operational or data cost.

Before trying this

Start with one survivable path and one external operation. Do not begin with a production administrative credential merely because it has the largest theoretical payoff.

Inventory the authority

Name one credential, every process that can retrieve it and the concrete operations each caller actually needs.

Design the records

Define one request kind and one result kind. Declare every field that must participate in routing or authorization as an indexed path.

Price the data boundary

Choose retention before request bodies enter the space. Decide which fields may remain visible to matching, operators and observers.

Bound external effects

Confirm that the downstream operation accepts an idempotency key or has another durable deduplication mechanism.

Begin with a staged workload. Radia has no independent deployment history yet. The first path should be real enough to measure and safe enough to survive a runtime, worker or integration failure.

Start beside the existing path

Deploy Radia next to the current architecture and change no existing credential. Selected processes gain a Radia identity as an additional interface.

Starting beside the existing path adds an identity without removing anything The secret manager keeps provisioning the existing database credential to the service and adds a Radia bootstrap credential. The service exchanges the bootstrap credential for a short-lived run token and keeps its old path working. Vault DB credential Radia definition service A unchanged path plus an identity exchange run token 15 minutes, renewable to 12 hours
The bootstrap credential can mint runs and nothing else. A request presenting it is refused with a message saying to mint a run first, so the durable half is never the acting credential.

The process exchanges its durable bootstrap credential for a short-lived run token. The durable credential can mint runs but cannot perform ordinary operations, so it is never the acting credential. Still, minting is a single call, so a leaked bootstrap credential is one hop from the agent's full authority. What the split buys is a short-lived acting credential, one minting handle to revoke, and a record of every mint.

At this point no downstream credential has been removed, so containment has not improved. What Radia establishes is the foundation later transitions need: principals, record kinds and their indexed paths, grants, and the claim and lease behaviour your workers will rely on. Decide where the bootstrap credential comes from now rather than later.

You gain evidence, not containment. Continue when identities, kinds, grants, lease recovery and the expected request volume have been exercised without changing the existing service path.

Add a record-based request path

Take one service that holds a credential and give it a second way to be asked. The existing HTTP path stays exactly as it is.

The same service, reachable two ways Before, the invoice agent calls the ERP directly with an ERP token. During evaluation, it retains that token for rollback but uses the record path: Radia checks the grant, and the existing invoice service claims the work and calls the ERP with its token. before invoice agent POST /erp/invoices with ERP_TOKEN ERP after invoice agent token retained for rollback record Radia grant checked claim invoice service holds ERP_TOKEN ERP
Both paths are live at once: the record path can carry one caller while every other caller keeps using the direct call.

The service still reads ERP_TOKEN from Vault. The selected caller uses the record path but retains its token for rollback during evaluation. Removing it is the next transition.

A legacy system does not have to learn anything. A bridge worker translates a record into whatever the system already speaks, which is how Radia spreads from the edges inward rather than from a rewrite.

A bridge worker translates a record into whatever the system already speaks A customer lookup record is claimed by a bridge worker, which holds the legacy API token and issues an ordinary authenticated HTTP request on the caller's behalf. crm.customer.lookup tenant: acme customer_id: 123 claim bridge worker holds the legacy token legacy REST API GET /customers/123 Bearer token
The legacy system learns nothing about Radia. Later, the implementation behind the record kind can change without touching a caller.

A bridge worker claims one operation, calls the existing system and writes a result record. A failed call is an answer, not a refusal to settle: a downstream rejection that nacks the claim is redelivered, while a failure result becomes one durable record that consumers handle with the normal delivery semantics. See the worked examples →

You gain an independently routable request path. Continue when its latency, result shape, retry behavior and duplicate handling have been observed under representative load. The caller still holding its old credential means containment has not improved yet.

Move the credential to a narrow worker

After this transition, compromise of the caller yields only what its grants permit, not the ERP token.

One caller with three secrets becomes one identity and three narrow workers The invoice agent stops holding the ERP token, the object-store token and the database password. Each moves into a worker that does one job and holds one credential. before invoice agent ERP_TOKEN S3_TOKEN CUSTOMER_DB_PASSWORD after invoice agent one identity erp worker ERP_TOKEN document worker S3_TOKEN customer worker CUSTOMER_DB_PASSWORD
Compromise of the caller now yields whatever its grants permit it to write, claim and read, rather than three external systems.
Keep the workers boring. One worker holding the ERP token, the CRM credential, cloud keys and a database password is a privileged middleware service with extra steps. Prefer narrow workers with narrow secret-manager policies: customer-reader, invoice-writer, deploy-worker. The runtime, its storage and any credential-holding host stay trusted components, so the goal is to reduce their number and authority rather than to assume Radia makes them harmless.
You gain containment. Continue when every legitimate caller operation is expressible through a concrete request kind and the old credential has actually been removed from the caller's deployment, diagnostics and recovery path.

Express caller authority as grants

A secret-manager policy answers one question: may this workload retrieve this credential. Retrieval then implies everything that credential can do.

Possession of a credential becomes permission to request one operation The secret manager policy says which worker may read the ERP secret. The Radia grant says which caller may request an invoice creation, and for which tenant. Deletion, export and administrative operations are not granted at all. secret manager policy the ERP worker may read secret/erp/prod Radia grant the invoice agent may put erp.invoice.create for tenant acme not granted erp.invoice.delete erp.customer.export erp.admin.*
Secret-manager policy limits who holds the downstream credential; Radia grants limit what callers may ask that holder to do.

A grant names a principal, a concrete kind and the operations allowed on it. Wildcard kinds are refused, so authority is enumerated rather than implied. Grant changes remain visible in the record history.

the caller's whole authority over the ERP
{ "principal": "agent:invoice-agent",
  "kind": "erp.invoice.create",
  "operations": ["put"] }

Advertising a capability is not authorization. A worker may publish what it handles, and only an operator or authorized supervisor can write the grant that permits it. See how grants are enforced →

You gain enumerable request authority. Continue when effective permissions answer who may request each operation without relying on a downstream token shared across callers.

Enforce tenant scope before execution

The database worker holds one credential that can read every tenant. Without scoping, each caller's code is what keeps tenants apart, which means every caller is part of the security boundary.

the scope moves out of the caller and into the grant
{ "principal": "agent:acme-assistant",
  "kind": "customer.query",
  "operations": ["put"],
  "pattern": { "tenant": "acme" } }

A pattern bounds both reads and writes. A broad read is narrowed to permitted content, and a write outside the pattern is refused before a worker receives it.

Design scope into the record kind. A grant can scope only fields declared as indexed paths. If tenant is the isolation boundary, declare it when the kind is created rather than burying it inside an opaque payload.

What each principal can actually do is a read rather than a belief: radia permissions agent:acme-assistant answers from the same path the enforcement uses. See effective permissions →

You gain a boundary the caller cannot omit. Continue when the indexed fields encode the real isolation decision and tests show that broad reads and writes are narrowed or refused before the worker receives them.

Bound shared workers by the caller

A shared worker serving many callers is the point where a broker usually becomes a way to borrow the worker's authority. Radia's answer is a delegated run whose authority is the intersection of the caller's grants and the worker's.

A delegated run carries the intersection of the caller's grants and the worker's The caller is scoped to one tenant and the shared worker can reach all of them. The delegated run the worker acts under carries only the overlap, so a read on the caller's behalf stays inside the caller's tenant. caller tenant = acme shared worker all tenants delegated run tenant = acme reads bounded by both
The caller reaches one tenant and the worker reaches all of them; the run minted for work on the caller's behalf carries only the overlap.

The worker uses its own run for its own work and a delegated run for work performed on a caller's behalf. Radia derives that caller from the authenticated run, not from a claim in the request body.

An intersection narrows; it never grants. A delegated run cannot contain a capability the caller lacks, so the caller needs its own grant on the kind being read. See delegated authority →
Use delegation only for genuine on-behalf-of operations. A worker acting only through its own grants needs no delegated run. Continue when the application has a shared worker whose reads or writes must remain inside each caller's existing authority.

Keep credentials outside generated code

A container that holds cloud keys, a source-control token, a database URL and model-written Python is the combination this transition exists to end.

Generated code proposes operations and holds no credential A generated process sends proposals to a trusted host over a private channel. The host holds the run token and supplies the identity, parents, labels and idempotency keys the child must not control, then Radia applies the agent's grants. generated process no secret-manager token no Radia token no database password no cloud credential proposal trusted host holds the run token stamps identity, parents labels, idempotency authorized Radia applies the agent's grants
There is no credential inside the sandbox to steal, replay or exfiltrate, which is a different property from injecting a short-lived one.

Generated code receives a narrow proposal interface rather than a credential. The trusted host supplies identity and provenance, then Radia applies the agent's ordinary grants. This is stronger than injecting a short-lived token because the sandbox has nothing to steal or replay. See credentialless execution →

You gain credentialless execution, not a harmless process. The trusted host, runtime and storage remain inside the security boundary, and the selected sandbox must prove the isolation properties it advertises.

Where the bootstrap credential comes from

The built bootstrap chain keeps a durable credential in the secret manager and is available on any deployment.

A durable bootstrap credential, stored like any other secret An operator creates the agent definition and assigns grants. The credential is stored in the secret manager, read by the workload using its platform identity, and exchanged by the trusted host for a short-lived run token that performs the ordinary operations. operator creates + grants Vault retrieved through workload identity trusted host reads it, exchanges it run token 15 minutes, 12 hours maximum ordinary operations under the run
The durable credential only bootstraps short-lived runs; application work happens under those runs. A run renews itself while it works and cannot outlive twelve hours from its mint, so a long-running host re-exchanges the durable credential rather than renewing forever.

Store the durable definition credential like any other infrastructure secret and expose it only to the trusted host, as a file with restrictive permissions rather than an environment variable, which tends to reach diagnostics, crash dumps and deployment manifests. The host exchanges it for renewable, short-lived runs rather than using it for application work.

Radia can also mint a run from a verified OIDC identity, resolving the principal through an explicit mapping. A compatible workload-identity platform can therefore avoid storing a durable Radia credential. The space has to be started with the issuer and the audience it accepts, and that audience is a single value per space, so one space accepts tokens minted for one audience. Ready-made Kubernetes and SPIFFE profiles are not shipped today, so treat this as an integration to validate, not a turnkey deployment path.

A workload identity token minting a run with no stored secret A platform-issued workload identity token is presented to Radia, which verifies it against the issuer and reads the identity mapping record to decide which principal it is. The result is an ordinary short-lived run token. workload identity platform-issued token presented once Radia verifies the signature reads the identity mapping run token agent grants no durable Radia credential is stored
An explicit mapping decides which Radia principal the external identity becomes.
Create the mapping before first use. An unmapped identity receives no grants; retiring a mapping prevents it from being used again.
Revocation and containment are separate. Revoking a definition prevents new runs and deliberately leaves live ones alone, so a compromised process keeps acting until its run expires. Stopping those runs is a second action (radia runs --for <principal> --stop), and it covers both the principal's own runs and delegated runs held by workers on its behalf. An incident response does both.

Give each environment its own principal and its own stored credential. Compromise of development should not yield a minting credential for production, and separate spaces are a stronger boundary than separate grants inside one space.

Effects that leave the space

Delivery is at-least-once. A worker can perform an external effect and fail before its settle lands, after which the work is redelivered. Any bridge worker that charges a card, sends mail or files a document needs deduplication at the external boundary, designed in from the first one rather than added after the first duplicate.

The record id is the deduplication key at the external boundary A payment request record is claimed by a worker that may be redelivered the same record after a failure. The external API call carries the record id as its idempotency key, so a repeat delivery is one charge rather than two. payment.request rec_7f912... claim payment worker may be redelivered payment API Idempotency-Key rec_7f912... the record id, not the attempt
A worker that lost its lease may still be running, so the external key has to identify the request rather than the attempt.

See what happens when a worker dies →

What changes when requests become records

Moving credentials out of callers also moves request data into a shared store. Price that trade before adding the record path, not after relying on it.

What the migration improves, and what it costs On one side, fewer holders of each credential, authority that can be enumerated and audited, and tenant scoping enforced before a worker runs. On the other, request bodies living in a shared store, records kept unless retention was declared, the space sitting in the synchronous path, and shared storage for several runtime instances. what improves fewer holders of each downstream credential authority is enumerable and auditable tenant scoping is enforced before the worker runs requests and results are one audit trail what it costs request bodies live in a shared store records are kept unless retention was declared the space is in the synchronous path several instances need shared storage
Both columns move together. An evaluation that counts only the left one will be surprised by the right one at the first compliance review.
ConcernWho owns it
Grant intersection, lease fencing and write authorizationRadia enforces
Storage and rotation of the downstream credentialSecret manager and capability-worker host
External API deduplicationCapability worker and downstream system
Sandbox isolationSelected backend and trusted host
Request-body retention and routable plaintext fieldsApplication and kind design
TLS, request limits, backup and shared deployment storageDeployment infrastructure

Record bodies are stored as plaintext JSON, not only the fields that route: matching needs the routed ones, and the rest sit beside them. Operators and principals with observation authority can read those bodies, so sensitive payloads and retention need deliberate design. Application-level encryption is available for fields that do not participate in routing.

Records are immutable and remain unless the writer or kind declares retention. Large or erasable content belongs in artifacts rather than record bodies.

Claiming, leasing and settling add more work than a direct HTTP call, and several runtime instances need shared storage. Radia has not yet established production readiness through independent adoption. Pilot it on a real but survivable path rather than replacing a mature authorization boundary at once. See what it is not →

Choose transitions by risk and need

Start with credentials whose risk justifies a new boundary, while keeping the first path survivable. The worst combination is broad authority, many holders, agent-controlled code and a high-value target; the best first pilot is often one step below it.

Credentials ranked by adoption priority Production database administrative credentials reachable by agents and cloud administrative credentials rank very high. Organization-wide source-control tokens and production CRM write credentials rank high. Narrow read-only tokens are medium and internal low-risk service credentials are low. very high production database admin credential reachable by agents very high cloud administrative credential high organization-wide source-control token high production CRM write credential medium narrow read-only API token low internal low-risk service credential
These priorities are illustrative. Rank your own credentials by authority, number of holders, exposure to agent-controlled code and the value of the target.
A map of independently useful architecture transitions A reference path from a caller holding its own downstream credential to a retired direct path, passing through an added identity, a request record, a narrow worker holding the credential, grants, pattern scoping, delegation and credentialless generated code. 0 caller holds the downstream credential and calls the system directly 1 a Radia identity is added; nothing is taken away 2 the caller writes a record; the existing service claims it and calls the system 3 the downstream credential leaves the caller and lives in a narrow worker 4 the caller's authority becomes a grant on a kind, not a token it holds 5 the grant carries a pattern, so tenant scoping is enforced before the worker runs 6 a shared worker acts under caller grants intersected with its own 7 generated code holds nothing and proposes through a trusted host 8 the old direct path is retired
This is a reference path, not a completion checklist. Stop after the transition that establishes the boundary the application needs; removing the caller credential already produces a concrete result.

The objective is not zero credentials. It is that very few processes hold one that can act on an external system, and that possession stops being how application code expresses authority. Two numbers say whether that is happening.

How many processes can exfiltrate each downstream credential, and how many principals can read each sensitive body? The first should fall as credentials move. Track the second from the moment requests become records, and keep it small with narrow kinds and pattern-scoped grants.

A production database credential reachable by thirty-seven workloads today, and by three capability workers after the first transitions, is a real result even if the secret manager never moves.