The repository includes a coordination pipeline (TypeScript and Python), a model team that
writes and reviews music, a multi-process chat application, a content-keyed analysis
application, a shared world, a scripted marketplace that calls no model at all, and a
load generator. Only live model calls require an API key; the automated suites use
deterministic providers. To put agent harnesses
you already run on a space instead of running an example, see
radia team: it writes no application
code, and radia team up runs those harnesses as workers from one directory of
prompts.
Run deno task demo to start a planner, two workers and an aggregator. The planner
emits pieces, each worker claims the kind it handles and the aggregator writes a summary after
all results arrive.
The four programs do not import or address one another. Their integration consists of the kinds and patterns they publish to the shared space.
The same pipeline runs on the Python SDKs: python3 examples/pipeline-py/demo.py
against a space started with --ext. Its workers also advertise their ops as
presence-policed capabilities, so a worker that dies takes its tool out of discovery when
its beat ages out.
Open the console while it runs and you can watch records appear and change colour as they are claimed and answered.
Song Creator turns one request into a brief, several independently written parts, an assembled draft and two blind reviews. A deterministic producer sends rejected work around another round or renders every draft into a playable workspace.
Models decide the arrangement, write the parts and judge musical shape. Programs perform fan-in, validate the notation, count measurable faults, enforce revision limits and synthesize the WAV. The output has a measurable quality signal, so the model-free test requires the review loop to reduce the fault count rather than trusting a prose claim that it improved.
Run deno task test:chat for the no-key suites, or deno task chat with a
configured provider. The application separates the client, model routing, inference, tool
execution and code execution into processes with distinct credentials and host permissions.
The terminal client writes a request record and renders result records. Workers perform model and tool calls independently, so a terminal restart does not terminate an active turn.
a message // what you typed
a request // no model named: the chat does not choose one
a request // a router picked a model and asked again
some chunks // the reply, streaming in
a reply
a tool call // whichever worker offers that tool claims it
some progress
a tool result
a verdict // written by the runner, not by the model
a reply // carrying the provider's own token and dollar figures
Conversations persist as record threads and can be resumed by name. Code checks are written by the runner under a kind the assistant cannot write, keeping model claims separate from execution verdicts.
Tool workers advertise schemas and descriptions as records. The assistant watches those records to build the model's tool list, so a newly deployed worker is available on the next turn without client configuration.
Replies retain provider-reported token and cost fields for later queries. Model tiers are also records, allowing routing and escalation workers to read the same active registry.
With --encrypt, message prose, tool arguments and tool output are encrypted while
routing fields remain readable to the runtime. Conversation keys live in shreddable artifacts,
allowing the prose to be erased while records, lineage and the event chain remain.
Run deno task analysis to start a web application with sign-in, CSV upload and
staged processing. The application relays the browser's token and holds no service credential.
Each stage is keyed by dataset, input digest and code digest. Completed work is found by query. Changing code or input creates a new key for that stage and its descendants; unchanged stages reuse their existing artifacts.
The code that runs is a record too. Each stage lives in a workspace tree named by its digest, and that digest is the same value the work request carries and a promotion grant pins. So a stage can only ever be handed work naming its promoted tree, and a result claiming any other digest is refused at the moment it is written. The question every cached answer rests on, which code produced this, went from a field a worker reported to a fact the grant table enforces in both directions. The code itself runs in a jail holding no credential: a generic host fetches its input under the stage's own permission, and stamps its output with the person the work was for, so that person's scope reaches bytes an agent computed on their behalf.
And the pipeline's shape is itself data. The list of stages is a registry the planner and the page both read, so adding a stage is a deployment: store the tree, declare where it goes in the order, promote its digest, bind it. The example's own test does exactly that, deploying a fourth stage into the live pipeline while it runs, and only the new work computes; nothing restarts and no process that was already running learns anything. A tree authored anywhere qualifies, including one an agent wrote in conversation, because deployment names a digest and nothing else.
planner.ts. Deciding it depends on what you consider an
input. Everything else (the DAG, the routing, the leases, the audit, and now which code may
answer) is already there.
Run deno task test:mud for the headless suite, or
deno task mud -- --player alice to walk around. It needs no API key and calls no
model. Rooms, players and non-player characters share one space: a player writes what they
typed, a narrator claims it and decides what happened, and the characters standing there
react.
Each character is a principal, not a branch in a game loop. The gatekeeper's permission to speak is pinned to the room she stands in and to the name she may speak under. A gatekeeper that tried to speak in the tavern, or to put words in a player's mouth, is refused when the record is written. No code in the example checks for either, and the test plants both writes to show the refusal comes from the runtime.
a command // "north", claimed by the narrator under a lease
an event // "alice goes north." in the room she left
a presence // where she is now, which is the authority on it
an event // "alice arrives from the south." in the room she entered
a cue // one per character standing where it happened
an event // the character's line, written as the character
A player writes one kind of record and the client reads two. Everything else they see was written by somebody else, under a permission the player does not hold: they cannot narrate, cannot move themselves by declaring where they are, and cannot type as another player. What a room contains is a question about the newest record per actor rather than a query for records mentioning the room, because the log keeps every place everyone has ever stood.
Characters also act when nobody is watching, and no process holds a timer to make that happen. Each one answers its cue with the next cue, marked to become claimable half a minute later, so a wandering guard is a chain of records. Code that Radia runs in a jail is handed one record and then exits, so a timer is not available to it at all, while a deferred record is.
Start a space and there is a web console at the same address. It is built on the public API, so it can only show you what any other client could ask for.
Overview ranks what deserves attention, and every finding links to the records it rests on. Feed is a live view of everything happening. Graph draws how records connect, a conversation and its messages, a job and its pieces, with a waterfall view that reads like a trace: time as the axis, so you can see which tool the seconds went to. Kinds draws who is listening for what right now, mined from live declarations rather than configuration. Space arranges every record by what it is rather than by what it links to, so a fleet doing something odd shows up as a cluster that should not exist. Flows renders the recurring shapes with their evidence. Auth shows who exists and what they may do.
Whatever you are looking at is in the address bar, so you can send someone a link to it, and it survives your session expiring while you are in the middle of looking at something.
This design creates a problem for itself. If nothing declares how work flows, there is no diagram to look at. New people ask where the workflow is, and no file holds one.
So it gets worked out from what happened. Take everything that is causally connected, reduce it to the sequence of steps involved, and count how often each shape occurs.
180x 100% n= 2 1.3s a request → a reply
104x 100% n= 7 7.6s conversation ⇒ request → request + message → reply
73x 100% n= 4 810ms conversation ⇒ tool call → message + result
3x 100% n=10 229ms job → pieces ×4-7 → results ×4-7 → summary
The bottom line is the pipeline from earlier, found without anybody describing it. The columns are how often it happened, how often it finished, how many records were involved and how long it took.
Shapes that start and rarely finish show up next to the ones that work. A job nobody ever picks up and a job whose worker keeps crashing look identical in a log, and they need completely different fixes.