Records who owns each Run and Attempt and how each ended, as atomic commands on the command journal; it starts no process, measures no time and judges nothing.
Status: "Accepted for trusted local supervision" (build review): durable cancellation shares the Run transaction; stale completion, retry and replay cannot bypass it; permanent commit-edge crash and competing-controller tests. Receipts: build receipt (independent execution review by gpt-6-astra at xhigh; commit c4111d6; its final_sources hash is for the earlier source, before cancellation requests existed) and continuation receipt (durable cancellation, 31 focused cancellation checks). Only the continuation receipt's files hashes for the source, both test files and the helper below match the current files. This is local scope: no epic is complete, and all 27 remain open.
Source: src/execution.ts. Tests: test/execution.test.ts (24), test/execution-cancellation-boundaries.test.ts (2, cross-process commit boundaries, with test/helpers/execution-cancellation-child.ts).
Intelligence: none — Run and Attempt ownership, Budget and cancellation are invariants checked on every read and write; there is no judgement to delegate.
What it hides
- One Run per journal aggregate. Each Run is stored under
run:<runId>, and every command is oneCommandJournal.executecall. State, receipt, events and outbox entries commit together, and each command adds one version. - Ownership epochs. A claim opens Attempt
<runId>/<epoch>, where the epoch is the Attempt's position from 1. An Attempt ends only by its owner's report naming that epoch, or by a trusted cleanup confirmation for that epoch. There is no takeover, lease or timeout. - Durable cancellation in the Run record. A pending
cancelRequestedis seen by every command in the same transaction, whatever version the caller read. - Whole-state validation. Every state
readreturns, every non-null state a command reads, and every state about to be written passescheckRun(execution.ts:575). Every change also passescheckTransition(execution.ts:628).
Public interface
EXECUTION_LIMITS (execution.ts:49), frozen:
| Limit | Value | Applies to |
|---|---|---|
maxRunIdLength | 64 | runId: /^[A-Za-z0-9][A-Za-z0-9._:-]{0,63}$/ (no /) |
maxAttempts | 20 | budget.maxAttempts and every epoch (1–20) |
maxAttemptTimeoutMs | 86_400_000 | budget.attemptTimeoutMs (1 to 86,400,000) |
maxEvidencePerAttempt | 128 | evidence array length |
maxTextLength | 500 | reason, cleanup.method (UTF-16 units; not blank, well-formed, no control characters) |
The other rules are internal. commandId, worker, recipeRevision, scenarioRevision, confirmedBy and each Evidence reference must be 1–128 characters of A-Z a-z 0-9 . _ : / -, starting with a letter or digit. candidateDigest is sha256: plus 64 lowercase hex characters. expectedVersion is an integer from 0 to Number.MAX_SAFE_INTEGER. Evidence references must not repeat. Every command input, and each record inside it, is a plain data object with exactly the documented keys. Unknown keys, accessors, class instances and holes are refused.
class Execution (execution.ts:315). constructor(journal: CommandJournal) throws InvalidInputError for anything that is not a CommandJournal.
| Method | Accepts when | Leaves the Run | Events |
|---|---|---|---|
submit(input: SubmitRun): CommandResult | no Run with this ID | queued | run.submitted |
claim(input: ClaimRun): ClaimResult | queued, nothing pending | running; new epoch = attempts.length + 1 | attempt.claimed |
complete(input: CompleteAttempt): CommandResult | running, epoch is current, nothing pending | completed | attempt.completed, run.completed |
fail(input: FailAttempt): CommandResult | as complete | queued if attempts.length < maxAttempts, else exhausted | attempt.failed [+ run.exhausted] |
interrupt(input: InterruptAttempt): CommandResult | running, cleanup.epoch is current (pending allowed) | as fail; cancelled with the pending reason if one is pending | attempt.interrupted [+ run.exhausted or run.cancelled] |
requestCancel(input: RequestCancel): CommandResult | open, nothing pending | queued becomes cancelled; running stays running with cancelRequested: {reason} | run.cancelled or run.cancel-requested |
cancel(input: CancelRun): CommandResult | queued without cleanup; running with cleanup for the current epoch; if pending, reason must equal the pending reason | cancelled; a running Attempt becomes interrupted | [attempt.interrupted +] run.cancelled |
Reads: read(runId: string): Run | undefined returns undefined if the Run was never submitted. history(runId: string, page?: PageOptions): readonly JournalEvent[] returns events in commit order, or [] for a Run never submitted; page is checked by the journal, not as a strict record: afterSequence (default 0) is a safe integer from 0, limit (default 100) an integer from 1 to 1,000, and other keys are ignored. Both check the runId (InvalidInputError) and both work over a read-only journal.
Events, in the order a command records them (each data also carries runId): run.submitted {pins, budget} · attempt.claimed {attemptId, epoch, worker} · attempt.completed {epoch, evidence} · run.completed {} · attempt.failed {epoch, reason, evidence} · attempt.interrupted {epoch, cleanup} · run.exhausted {attempts} · run.cancel-requested {epoch, reason} · run.cancelled {reason}.
Inputs. submit takes commandId and runId; every other command takes those plus expectedVersion:
SubmitRun { commandId; runId; pins: RunPins; budget: Budget }ClaimRun { …; worker }CompleteAttempt { …; epoch; evidence: readonly string[] }FailAttempt { …; epoch; reason; evidence }InterruptAttempt { …; cleanup: CleanupConfirmation }RequestCancel { …; reason }CancelRun { …; reason; cleanup? }. Omitcleanuprather than passingundefined.
Types:
RunPins { candidateDigest; recipeRevision; scenarioRevision }Budget { maxAttempts; attemptTimeoutMs }. It has no cost field, so amaxCostUsdfield is refused.CleanupConfirmation { epoch; confirmedBy; method }: a trusted supervisor's statement that the Attempt's protected work ended or was irreversibly fenced against resuming, with resources released or quarantined. It need not prove OS process exit. Execution records who made it and the method, and cannot check it.RunningAttemptandAttempt. Each Attempt hasattemptId,epochandworker, plus one of:running;completedwithevidence;failedwithreasonandevidence; orinterruptedwithcleanup.CancelRequest { reason }RunStatus = 'queued'|'running'|'completed'|'exhausted'|'cancelled'andTerminalStatus = 'completed'|'exhausted'|'cancelled'Run: the Run's state plusversion, deeply frozen.cancelledcarriescancelReason;runningmay carrycancelRequested.CommandResult { run: Run; replayed: boolean }andClaimResult extends CommandResult { attempt: RunningAttempt }
Errors: ExecutionErrorCode, ExecutionError and its subclasses (below).
Invariants and guarantees
runId,pinsandbudgetnever change aftersubmit, and settled Attempts are never rewritten.checkTransitionenforces this, and the test "pins and Budget are fixed at submission…" covers it.attempts[i]has epochi + 1and ID${runId}/${i + 1}. Only a claim adds an Attempt, one at a time and only when none is running.attempts.length <= budget.maxAttempts. Enforced bycheckAttemptandcheckTransition; covered by the Astra budget matrix, which tests every ceiling from 1 to 20.- Only the last Attempt may be
runningorcompleted; every earlier one isfailedorinterrupted. - Status matches the Attempts.
queuedhas no Attempt running and at least one left.runningmeans the last Attempt is running, andcompletedthat it completed.exhaustedmeans every budgeted Attempt was used and none completed.cancelledmeans none is running or completed.cancelReasonis present exactly when the Run iscancelled, andcancelRequestedonly when it isrunning. completed,exhaustedandcancelledare terminal: no later transition is accepted. A command that reaches the decision is refused withRunTerminalError; a matching repeat replays; a stale version, a conflicting command ID and a freshsubmitare refused earlier withVersionConflictError,CommandConflictErrorandRunExistsError.- A running Attempt ends only by its owner's report for its epoch, or by a cleanup confirmation for its epoch. The ending keeps the same
worker. Elapsed time, leases and heartbeats end nothing. The tests refuseforce,replacesandlastHeartbeatAt. - A pending cancellation is never cleared, replaced or bypassed. It ends only as
cancelledwith its own reason, and never asqueuedorexhausted, so no retry follows. Adding it changes no Attempt. - Each command is atomic: one version, and state, receipt, events and outbox together. A SIGKILL before commit leaves nothing; one after commit leaves everything (boundary tests). The stored Run always equals the fold of its own events (test "each Run is the fold…").
- Input is validated before storage is touched. A stored record that breaks any invariant raises
CorruptRunErrorand is never interpreted. One exception: a persisted JSONnullunderrun:<runId>isCorruptRunErroronread, but a command treats it as a missing Run (RunNotFoundErrorat its matching version, execution.ts:490). Nothing is written either way. - Execution has no Verdict vocabulary. Inputs carrying
verdictorstatusare refused.completedmeans the owner ran the Recipe to its end; Evidence references are stored as given and never interpreted.
Failure semantics
Refusal order: input validation (InvalidInputError, nothing touched), then the journal (JournalClosedError or JournalReadOnlyError before any transaction; then replay or CommandConflictError; then VersionConflictError), then the decision (CorruptRunError, RunNotFoundError, RunTerminalError, CancelPendingError, then AttemptActiveError or StaleAttemptError). A refusal records nothing new. A command ID that was unused stays unused, so the same ID can succeed later; an ID already committed stays bound to its original command and only replays with matching inputs.
Error (code) | Meaning | Extra fields |
|---|---|---|
RunNotFoundError (RUN_NOT_FOUND) | never submitted, so reachable with expectedVersion: 0; also a persisted null state at its matching version (invariant 9) | — |
RunExistsError (RUN_EXISTS) | another command submitted this Run ID. submit maps the journal's VersionConflictError to this | — |
RunTerminalError (RUN_TERMINAL) | Run is terminal | status |
AttemptActiveError (ATTEMPT_ACTIVE) | claim while running, or cancel on a running Run without cleanup | epoch |
StaleAttemptError (STALE_ATTEMPT) | the named epoch is not the running Attempt | epoch, activeEpoch (null if none) |
CancelPendingError (CANCEL_PENDING) | claim, complete, fail, a second requestCancel, or cancel with another reason, while pending | epoch, reason (first) |
CorruptRunError (CORRUPT_RUN) | stored record is not a valid Run; a command treats a persisted null as absent instead (invariant 9) | cause |
All carry code and runId. Errors from the journal pass through unchanged:
InvalidInputError(INVALID_INPUT)VersionConflictError(VERSION_CONFLICT): re-read, then decide againCommandConflictError(COMMAND_CONFLICT,mismatched): the same ID with a different payload,expectedVersionor Run. The operation is part of the payload.StorageError(STORAGE),JournalClosedError(CLOSED)JournalReadOnlyError(READ_ONLY): every command on a read-only journal, even a replayDecisionError(DECISION_FAILED, the thrown error ascause): a transition broke an invariant, socheckRunorcheckTransitionthrew. This is a bug, and nothing is recorded.
Idempotency and unknown outcomes. Command IDs are journal-wide keys. Repeating a command with the same Run, expectedVersion and payload returns replayed: true with the Run as that command left it (version expectedVersion + 1). A replay is not current state. If a caller crashed and does not know whether its command committed, it repeats the command: a replay means it committed, and a fresh result means it did not. Never launch work from a replayed claim; re-read ownership instead.
Failed, interrupted and unknown are different. failed is the owner's own report. interrupted is a supervisor's cleanup confirmation. A silent owner leaves the Attempt running: its outcome is unknown, never treated as failed. fail and interrupt retry by re-queueing while the Budget lasts. Execution does not judge whether a retry is worthwhile; a caller that decides against one calls cancel.
Trust scope
Established, in local scope:
- atomic, idempotent Run and Attempt commands over a local SQLite journal;
- one owner under racing claims from separate processes;
- stale reports refused after replacement;
- cancellation that commit order alone decides, surviving SIGKILL on either side of its commit.
Not established:
- Authentication. Callers are trusted local code. An epoch identifies an owner but is not a secret.
confirmedByis attribution, not authentication, and Execution cannot check a cleanup statement. - Protection from other writers. Anyone who can write the journal can store a well-formed forged Run. Only invariant-breaking records are detected.
- Process control or timing. Execution does not start, stop or time anything.
attemptTimeoutMsis recorded for the worker adapter to enforce. A hung Attempt staysrunninguntil a trusted confirmation arrives. - Proof that work stopped. A pending request only says work should stop. A cleanup confirmation says protected work ended or was irreversibly fenced, with resources released or quarantined; it need not prove OS process exit. No check-and-spawn is atomic across processes, so a claim committed first may already have a launch in flight, which its owner must stop and reconcile.
- Evidence. Execution never checks that Evidence references exist. A completed Run is not a Verdict.
- Outbox delivery. Events are pending in the outbox; the journal does not deliver them yet.
- Durability beyond process crashes. Power-loss durability, other platforms, production reliability, release and customer value are not established.
Composition
Depends on:
src/journal.ts:CommandJournal,DecisionError,InvalidInputError,VersionConflictErrorand types.src/canonical-json.ts:canonicalJson,jsonArrayItems,jsonObjectEntries.
Used by:
src/local-factory.ts:withStateconstructsnew Execution(journal)over a writable journal, or over a read-only one for inspection;submitRunsubmits with command IDverify:<runId>:submit;existingandreadLegacyRunread;legacySettlecompletes, fails or cancels withverify:<runId>:settle:<epoch>; its cancelcleanupis confirmed by the Attempt'sworker;legacyRequestCancelrequests cancellation withverify:<runId>:cancel:<uuid>.
src/guarded-verification.ts:dispatchOneclaims withverify:<runId>:claim:<token>;settleFromReportcompletes, fails or cancels withverify:<runId>:settle:<epoch>; its cancelcleanupis confirmed by the Attempt'sworker, from the stored report;reconcileRunninginterrupts withverify:<runId>:recover:<epoch>, after it records a fence proof;confirmedByisFENCE_PROTOCOL(sf-attempt-fence/1);requestCancelDurablyrequests cancellation withverify:<runId>:cancel:<uuid>;readBoundRunreads;buildGuardedRequestandparseReportuseEXECUTION_LIMITS.maxAttempts.
src/factory-support.ts: the typesBudget,ExecutionandRun.
Command IDs are journal-wide, so these prefixes must stay distinct from any other module's. Mirrored identifiers are not imported: Evidence pins (PIN_FIELDS in src/evidence.ts) carry runId, attemptId and epoch, and a fence's FenceOwner (src/attempt-fence.ts) carries runId, epoch and worker.
Changing it safely
Run
npm run check(typecheck and all tests).test/guarded-verification.test.tsand thetest/local-factory*.test.tssuites exercise Execution through composition;test/attempt-fence.test.tschecks thatMAX_FENCE_EPOCHequalsEXECUTION_LIMITS.maxAttempts;test/helpers/recovery-observer-fault.tspatchesExecution.prototype.readto inject a fault, so renamingreadbreaks it.Know which tests prove what. In
test/execution.test.ts:- ownership: the claim-race, replacement and killed-worker tests;
- idempotency: the replay and reused-ID tests;
- Budget: the ceiling and matrix tests;
- terminal refusal, input validation, fold/outbox consistency and corrupt-record refusal: one test each;
- durable cancellation: the tests under the
// --- Durable cancellationdivider.
test/execution-cancellation-boundaries.test.tscovers SIGKILL on both sides of the commit and both commit orders across processes.Keep stored records readable. Every read re-validates old records. A new stored field must be optional, as
cancelRequestedis, with a test in the style of "records without a pending field stay readable…".Change mirrored rules together.
EXECUTION_LIMITS.maxAttempts(20) is mirrored byMAX_FENCE_EPOCHinsrc/attempt-fence.tsandMAX_EPOCHinsrc/attempt-ledger.ts. The Run ID rule is copied asRUN_IDinsrc/local-factory.ts,src/attempt-fence.tsandsrc/attempt-ledger.ts, and the reference rule asREFinsrc/attempt-fence.ts. The 500-unit text rule is copied incancelFixture(src/local-factory.ts) andasReason(src/factory-support.ts).maxRunIdLengthkeepsrun:<runId>and<runId>/<epoch>within the journal's 128-character ID limit.Treat events and command payloads as durable history. Outbox consumers and the test
foldread the events, so keep old types and shapes readable. Execution itself reads stored state, not events: an old journal stays readable while its stored Runs still passcheckRun. Replay compares the canonical command payload ({ op, … }), so reshaping a payload turns a repeat of an earlier command intoCommandConflictErrorinstead of a replay.Reviewers reject:
- endings based on time, leases or heartbeats;
- takeover switches;
- Verdict fields;
- any path that clears or bypasses a pending cancellation;
- a changed refusal order.
Record receipts for any change. A change to
src/execution.tsor its tests invalidates the hashes in the continuation receipt. It needs a fresh independent review and a new receipt, as later increments recorded theirs. Keep existing receipts as historical evidence. Update the build review row only from that new evidence.
