A run that goes wrong on a shared space often does not throw. Nothing crashes, no stack trace appears, and the work simply does not finish. This page is how to find out why, taken from building a team of five models and two programs that writes music.
A claim is by content, not by address. Any worker holding the right grant can take any matching record, so the failures are ones a single process cannot see: a leftover from an earlier run claimed beside the new one, two copies of a worker racing, a handler deciding on evidence the run has already moved past.
Every one of those looks identical from inside a process. It claimed a record, it did the work, it wrote a result. The mistake is only visible from outside.
So the first rule is: read the space as well as the log. The log says what one process printed. The space preserves the records, claim transitions and relationships committed across the run, which is the evidence no one process holds.
That evidence has a boundary. A read leaves no event, a take that found nothing committed nothing, and an external side effect may live outside Radia. Absence from the space does not prove that an agent never looked, tried or acted elsewhere.
Which read to use is decided by the question, not by habit. A read that answers a different question returns a clean result, which reads as evidence that nothing is wrong.
| The question | The read |
|---|---|
| What does this record say? | radia query <kind> --match '{…}', radia get <id> |
| Is anything still claimable, leased, or given up on? | radia doctor for the summary, GET /v0/ops/records?state= for the records |
| Which live runs advertise an authorized interest? | client.dryRun(kind), or POST /v0/ops/dry-run |
| When did it happen, and in what order? | Each record's server-assigned creation time, radia activity, radia events |
A record query returns a record whatever its claim state, so it cannot tell you that an old task is still sitting there waiting to be claimed. Only the envelope knows that.
The third row matters because a grant says who may claim a kind, while an interest says what a run intends to claim. Interest publication is best-effort, and a live run does not prove its process is still running.
The fourth carries a trap: the database clock is the only clock every agent shares. Three harnesses on three machines each report their own, so a timeline built from anything else is a picture of the clocks rather than of the run.
A team of five models and two programs kept producing two songs at once. The log showed a clean start every time: the members came up, the work was seeded, nothing errored.
The envelope query showed four records from a previous run still available to claim. The team routes its own record kinds rather than generic tasks, and the sweep that clears leftovers only knew about tasks, so it had swept nothing and reported success.
Fixing that surfaced the next layer. A run stopped part way leaves its records leased, not available, and a lease lapses only when someone next tries to claim. So a sweep of available records still reported a clean space and handed the old song out seconds later.
Then a warning fired saying another copy of the team might be running. It was wrong. A worker mints a short-lived run, exiting does not stop that run, and an interest is live for as long as its run is. Two dead processes still looked like listeners.
Three defects, each invisible in the process logs and each exposed by a different read. The space supplied the record ids and state needed to locate the corresponding lifecycle and ordering mistakes in the code.
Provoke it in a space of its own first. A port stolen while a server boots, a service that ignores a polite shutdown, a deadline reached with nothing usable to hand back: each was reproduced against a throwaway space before anything was changed. Diagnosing against the space you are also using confuses your own writes with the fault.
Make the guard fail first. Write the test, revert the fix, and watch it go red. Twice in one day that step showed the test was wrong rather than the code, which is a cheap thing to learn then and an expensive one to learn later.
A flake is a finding. One intermittent failure turned out to be a real ordering bug: two reviews of the same round could be claimed after the next round already existed, and the handler was deciding on evidence that had been superseded. Re-running until it passed would have shipped it.
Five model turns per round makes the distance between a change and evidence about it minutes long and dollars wide. That is slow enough to change how you work, and not in a good direction: you start guessing instead of checking.
The answer is a second version of the same pipeline with a scripted stand-in for every model turn, running the real workers, the real records and the real grants against a real space. It runs in seconds, costs nothing, and asserts the properties that must hold.
Then the paid path is spent only on the thing the cheap one cannot produce, which is what models actually write. Every defect above was found by running, not by reading, and most were then reproduced for free.