Toolchain
- Node.js 26.8.1 or later and npm 11. TypeScript runs directly through Node's built-in type stripping, so there is no build step.
- The only dependencies are exact dev dependencies:
typescript(type checking) and@types/node. SQLite comes fromnode:sqlite, and nothing needs a network service.
npm ci # install the locked dev dependencies
npm test # node:test suite; every test uses real SQLite files in a temporary directory
npm run typecheck # tsc against tsconfig.json: strict settings, no emit
npm run check # both
The dashboard is a separate package with its own lockfile: npm --prefix apps/dashboard ci, then npm --prefix apps/dashboard run check (typecheck, build, rendered tests). The root check never needs its dependencies.
Source code has to use erasable TypeScript syntax only (no enums, namespaces or parameter properties), and relative imports include the .ts extension.
Command journal (src/journal.ts)
A local, single-host journal backed by one SQLite file in WAL mode. It records commands against aggregates. Each command changes the aggregate's state, emits events and returns a result, and it can be retried safely. Runs, attempts, leases and outbox delivery are out of scope; they are separate assignments.
import { CommandJournal, type Decide } from './src/journal.ts';
const add: Decide = (state, payload) => {
const count = (state === null ? 0 : (state as { count: number }).count) + (payload as { amount: number }).amount;
return { state: { count }, result: { count }, events: [{ type: 'counter.added', data: payload }] };
};
const journal = CommandJournal.open('data/journal.sqlite');
journal.execute('cmd-1', { amount: 2 }, 0, 'counter-1', add); // { version: 1, result: { count: 2 }, replayed: false, ... }
journal.execute('cmd-1', { amount: 2 }, 0, 'counter-1', add); // same result, replayed: true; add does not run
journal.readAggregate('counter-1'); // { aggregateId, version: 1, state: { count: 2 } }
journal.readHistory('counter-1'); // events in commit order, paged by { afterSequence, limit }
journal.readPendingOutbox(); // every event not yet delivered, oldest first
journal.close();
const reader = CommandJournal.openReadOnly('data/journal.sqlite'); // existing current journal only
reader.readAggregate('counter-1'); // the same reads, each seeing the latest commit
reader.execute('cmd-2', { amount: 1 }, 1, 'counter-1', add); // throws JournalReadOnlyError
Guarantees
One atomic transaction per command.
execute(commandId, payload, expectedVersion, aggregateId, decide)runs insideBEGIN IMMEDIATE. The new state, the version increment, the command receipt, the events and their outbox entries commit together, or none of them do. SQLite usessynchronous = FULL.Explicit version check. The decision runs only when the aggregate is at
expectedVersion(0 for a new aggregate); otherwise the call throwsVersionConflictErrorwith the actual version. Each successful command adds exactly 1 to the version and records at least one event.Idempotent retries. A command ID's receipt stores its aggregate, expected version and canonical payload. Repeating the same command returns the recorded, deeply frozen result without running the decision. Changing any of those inputs under the same command ID throws
CommandConflictError.Decisions are trusted, pure and synchronous. They receive frozen copies of the state and payload. If a decision throws, returns a promise or returns a malformed decision, the call throws
DecisionErrorand rolls back. Reads and replays never call back into application code.Fails closed. The file is marked with a SQLite
application_idanduser_version(schema 1).openrefuses an unknown schema version or a database that is not a journal (UnsupportedSchemaError), and it does so before changing that file.Read-only opening.
openReadOnly(path, options?)opens with SQLite's native read-only mode and requires an existing journal of the current schema. It creates, migrates, switches to WAL and checkpoints nothing:- a missing file or an empty database throws
JournalNotFoundError(NOT_FOUND) and stays missing or empty; - a foreign or unsupported database throws
UnsupportedSchemaErrorand is left unchanged; executethrowsJournalReadOnlyError(READ_ONLY) before any transaction or decision, even for a command that would replay.
Reads share the writable journal's readers and parsers. Each read sees the latest commit, including a live writer's commit after the reader opened, and a crashed writer's committed WAL. This is not a snapshot across several reads.
Read-only is not directory-byte stability. The main file and the WAL keep their bytes, but SQLite may create or update the
-shmindex, and an empty-wal, to coordinate readers with writers.immutable=1and copying a live database are not used: either could miss committed WAL frames.- a missing file or an empty database throws
Input rules (violations throw InvalidInputError before storage is touched):
- IDs and event types are 1–128 characters of
A–Z a–z 0–9 . _ : / -and start with a letter or digit. Versions and page options are non-negative safe integers; a page holds at most 1000 events and defaults to 100. - JSON documents (payload, state, result, event data) may contain only plain objects, dense arrays, finite numbers, well-formed strings, booleans and null. Each document is at most 1 MiB and 64 levels deep. The keys
__proto__,constructorandprototypeare rejected, as are symbol keys, class instances, cycles,undefined,NaNand infinities. - Anything
JSON.stringifywould silently skip is also rejected: non-enumerable properties, and array properties other than the elements andlength. Values are read from property descriptors, so a getter is rejected without being called. The decision's return value and itseventslist follow the same rules, so a list with holes cannot commit. This is strict input checking for trusted callers, not a sandbox againstProxyobjects. - Canonical form sorts keys by UTF-16 code units, as in RFC 8785, and uses ECMAScript number formatting (
-0is written as0). Receipts compare this form, so{ "a": 1, "b": 2 }and{ "b": 2, "a": 1 }are the same payload.
All deliberate failures extend JournalError and carry a code. SQLite failures, including a lock not acquired within busyTimeoutMs (default 5000), surface as StorageError, with the original error as cause. Several processes may open and create the same new journal file concurrently. SQLite does not call its busy handler for the switch to WAL mode, so open retries only that switch, only on SQLITE_BUSY and only within busyTimeoutMs. The schema check reads the application ID, schema version and object count in one statement, so a concurrent creator is never mistaken for a foreign database.
Current limitations
- There is no outbox delivery or acknowledgement yet, so every event stays pending.
- Schema migrations don't exist: any schema change needs a new version and an explicit migration.
- Stored JSON is trusted when it is read back; it is not re-validated.
- All connections must be on one host: WAL mode requires a local file system, not a network share.
- Retention and compaction are not implemented; receipts and history grow without bound.
Evidence and Assurance
EvidenceStore in src/evidence.ts stores a single content-addressed record containing bytes and metadata. put(metadata, bytes) returns its identity; get(id) verifies its content and provenance format. Records pin the Candidate digest, Recipe and scenario revisions, Run, Attempt, epoch, environment and collector. Concurrent identical writes converge; corrupt or non-regular objects fail closed. POSIX local filesystems are required.
gatherEvidence(store, refs) and pure judge(expectation, gathered) in src/assurance.ts produce a frozen local-fixture Verdict. Every required check needs matching observations for Verified. An unexpected exit is Failed; missing, corrupt, mismatched, timed-out or lost observations cannot produce Verified. The trusted collector records process observations; a producer's passed flag is not proof.
These are local trust boundaries, not a sandbox or authenticated external collector. The composition must own the full evidence-reference set and persist its issued Verdict; the pure judge cannot detect deliberately omitted evidence. Process-crash tests do not establish power-loss durability.
Execution
Execution(journal) owns Run and Attempt transitions: submit, claim, complete, fail, interrupt, requestCancel, cancel, read and history. Candidate/Recipe/scenario pins and Budget cannot change. Claims increment ownership epochs; stale reports and terminal resurrection are rejected. Replayed commands return their historical result; a worker must re-read current ownership and must never launch from a replayed claim.
A trusted supervisor must confirm that protected work ended or was irreversibly fenced against resuming, with resources released or quarantined, before interruption, replacement or active cancellation. Fencing need not prove process exit. Execution records the actual proof; it does not authenticate it, stop processes or enforce elapsed time. The worker adapter enforces its recorded timeout. A completed Run is recipe completion, never an Assurance Verdict.
Cancellation. Cancellation is part of the Run record itself, so every command sees it in the same transaction as its own change. requestCancel({commandId, runId, expectedVersion, reason}) works as follows:
- Queued Run: it becomes
cancelledat once. No Attempt is used and no cleanup is recorded. - Running Run: the Run stays
runningwith the same Attempt, epoch and Budget, and gains a pendingcancelRequested: {reason}. The version rises by one and arun.cancel-requestedevent is recorded. - While cancellation is pending,
claim,completeandfailare refused withCancelPendingError(CANCEL_PENDING), even for a caller that has just re-read the Run. So are a second request and acancelgiving a different reason. The first reason is never replaced or cleared. Repeating the same command replays its receipt with no new version or event; a changed payload under the same command ID is aCommandConflictError. - Settling: cleanup confirmed for the current epoch, through
interruptor throughcancelwith the pending reason, moves the Run straight tocancelledwith that reason. It is never queued or exhausted, so no retry is admitted and the Budget is untouched. Cleanup without the current epoch is refused, and so iscancelwithout cleanup. - Terminal Runs refuse every request.
- Old records: a record without
cancelRequestedhas no pending request and is read as stored. A pending field on a non-running Run, or a malformed one, isCorruptRunError.
A pending request says the work should stop, not that it has stopped; only the trusted cleanup confirmation settles it. A fresh claim is dispatch admission, not a process start:
- A request committed before a claim prevents admission.
- A claim committed first may already have one launch in flight. Its owner should check for a pending request right before the physical launch and stop and reconcile the work, but no check-and-spawn is atomic across processes.
- No later epoch is admitted once a request is pending.
- A replayed claim never grants launch authority.
Execution starts and stops no processes and judges nothing.
Attempt fence (src/attempt-fence.ts)
A per-Attempt ownership gate: a process performs an Attempt's work only while it holds the fence, and recovery proves that nothing holds it before closing it permanently. The guarded fixture worker and the CLI's guarded Recipe /2 use it.
import { createFence, holdFence, tryFence } from './src/attempt-fence.ts';
const identity = createFence('/abs/state/fences/run-1.1.sqlite', { runId: 'run-1', epoch: 1, worker: 'local-cli/…' });
const held = holdFence(identity); // { status: 'holding', hold } | { status: 'fenced' } | { status: 'busy' }
if (held.status === 'holding') { held.hold.assertActive(); /* work */ held.hold.release(); }
tryFence(identity); // { status: 'fenced', alreadyFenced } | { status: 'held' }; never waits
What a fence is. A fence is one small SQLite file in rollback (DELETE) journal mode. It has one row: protocol
sf-attempt-fence/1, Run, epoch, worker, a random nonce and a state ofopenorfenced.createFencereturns a frozen, JSON-safe identity that pins that row together with the file's canonical path, device and inode. It writes a fully built, fsynced temporary database and publishes it with a non-replacinglink(). An existing file is never replaced (exists), and only the call's own temporary name is ever removed.Holding. A hold is a read-only connection that runs
BEGINand then an actualSELECT. It keeps SQLite's SHARED lock, a kernel-released fcntl lock, untilrelease()or process exit. Many processes may hold a fence at once, and a hold never writes. A dropped handle stays reachable, so garbage collection never ends a hold. A hold is frozen, andvalidatedHoldIdentity(hold)returns the identity its own read validated, orundefinedfor anythingholdFencedid not return. Consumers use that function, never a caller-supplied label.busyTimeoutMs(0–10000, default 1000) bounds the wait while a fence attempt is in progress; after that the result isbusy.Fencing.
tryFencerunsBEGIN EXCLUSIVEwith a zero busy timeout. Any holder givesheld, and the probe connection closes at once, so it never keeps a legitimate reader out. On success it re-checks the file identity under exclusion, then commitsopen→fenced. Triggers in the file refuse every other change, so a fence never reopens and every later hold returnsfenced.What
fencedproves. At that moment no holder remained, and none can join later. A process that works only while it holds the fence can do no more of that work. It says nothing about what the work did, and nothing about processes that never held the fence, so never record it as an exit or any other observation.Fails closed (
unavailable, never recreated or reinitialised):- a missing or replaced file, a symbolic or hard link, or another user's file;
- a non-canonical or group/other-writable directory;
- a corrupt, foreign, WAL-mode or unsupported database;
- any row field that differs from the pinned identity.
Input is checked before anything is touched (
invalid-input).POSIX rule. Closing any descriptor for an inode drops this process's fcntl locks on it. While a process holds a fence it must not open, read, hash or copy that file by any other means;
lstatis safe. A second hold or atryFenceon an inode this process already has open is refused within-use. Use fences from one thread per process.Scope. One host, on a local POSIX filesystem with working fcntl locks; tested on macOS with Node 26.8.1.
node:sqliteignores Node's--permissionfile flags, so the guarantee rests on the trusted program holding the fence. It is not a sandbox and not a defence against a writer with access to the fence directory.
Tests: test/attempt-fence.test.ts runs holders, fencers and creators as real subprocesses (test/helpers/fence-child.ts), ordered by barrier files. Crashes and file swaps are injected by instrumenting node:fs in the child.
Local fixture worker
LocalFixtureWorker({store, workspaceRoot?}).run(request, signal?) runs the fixed pass, fail, hang or noisy fixture. It allows no executable, shell, environment or working-directory input. Each child receives a minimal environment, a private workspace, output/time bounds and its own watchdog. Node permission restrictions add protection; they are not an arbitrary-code sandbox.
Results distinguish an observed process, failed Evidence storage and no process started. Read the termination and workspace fields: stored Evidence alone does not establish safe cleanup. Unconfirmed termination or unexpected workspace contents are quarantined. Controller death stops the fixed child, but leaves its workspace for later recovery; this module does not reclaim it or kill a saved PID. Tested on macOS/Node 26.8.1.
Guarded runs (runGuarded, src/fixture-gate.ts, src/attempt-workspace.ts, src/fixture-manifest.ts)
A guarded run ties one fixture process to its Attempt fence, so a later recovery can prove that no guarded work can continue. The legacy run() path and its sf-local-fixture-worker/1 collector keep their behaviour. src/fixture-child.ts gained one start-callback call that does nothing without the gate. It is part of the legacy /1 Candidate source, so Runs pinned before this change now honestly report source drift.
const fence = createFence(`${state}/fences/${attemptKey}/fence.sqlite`, { runId, epoch, worker }); // one directory per fence
const held = holdFence(fence); // the controller's hold: keep it through report and settlement
const workspace = provisionWorkspace(`${state}/work`, fence, held.hold); // once, while holding; persist before claim
const executionDigest = guardedFixtureManifest().digest; // pin before the first Run
const result = await worker.runGuarded(request, { fence, hold: held.hold, workspace, executionDigest }, signal);
Inputs. Pass the complete fence identity, the caller's own live hold on it, the workspace ref provisioned for that fence and the pinned manifest digest. The request pins must name the fence's Run, its epoch and the Execution Attempt ID
<runId>/<epoch>, and use collectorsf-guarded-fixture-worker/1. If the parts don't belong together, the call throwsTypeErrorbefore anything is touched.Ownership checks. The worker never releases the hold. It checks it:
- when the call starts;
- right before spawning;
- after the process exits, before it removes the workspace or stores Evidence;
- before returning.
Only a hold that
holdFencereturned in this process counts. Its fence is the identity its own read validated (validatedHoldIdentity), which must equal the pinned fence exactly. The pinned file must also still be the held inode; a relabelled identity on the same file, or a look-alike handle, is aTypeError. These are checks, not a borrow. The caller must keep its hold for the whole awaited call and through its own report and settlement; a check cannot undo a write that was already in progress when a caller broke that rule.- If the hold is lost before spawning:
not-started/ownership-lost. - If it is lost afterwards:
ownership-lostwith the termination summary. Nothing further is removed or stored, the workspace is left for recovery, andevidenceIdis set only if Evidence was stored.
Other refusals before launch.
drift(nothing starts): the listed sources no longer hash toexecutionDigest.setup-failed: just before spawning, with nothing awaited in between, the workspace is not exactly the provisioned, empty, private directory, or the fence directory holds anything besides that fence.
The child. It runs
node --permission --allow-fs-read=<fixture-child.ts, fixture-gate.ts, attempt-fence.ts, the fence directory> --import fixture-gate.ts?enter fixture-child.ts <fixture> <lifetime>in the workspace. It gets the fixed environment plus the fence identity inSF_FIXTURE_GATE_FENCE, and a control pipe on descriptor 3.- The gate removes that variable and holds the exact pinned fence until the process exits. It then writes
sf-fixture-gate/1 admitted <nonce>to the control pipe and installs a start callback. - The fixture program calls that callback as the first statement of
main, which runs only once the program has compiled. The callback writessf-fixture-gate/1 payload <nonce>and closes the pipe. - If the fence is closed, missing, replaced, foreign or doesn't match, stays locked beyond 2 s, or the control pipe is missing, the gate writes one
sf-fixture-gate: refused:line and exits 123. No fixture code runs. - The fixture's exit status counts as a fixture outcome only when both control records arrived, exactly. Otherwise it is a collector error, so Assurance gives Inconclusive, never Verified or Failed, while the real exit status and output are kept. That covers a refusal, a death before admission (for example while loading the gate), an admitted program that never started (missing or failing to compile), a timeout before admission and malformed records.
- The records (at most 256 bytes) travel outside the output budget and never appear in Evidence output.
- Minimal environment, no shell, time and output bounds, the child's own watchdog, stdin end-of-file on parent loss and signalling only through the live
ChildProcesshandle all work as before.
- The gate removes that variable and holds the exact pinned fence until the process exits. It then writes
Workspace.
provisionWorkspace(root, fence, hold)needs the caller's own live hold on exactly that fence, so nothing can be prepared for a closed fence. It never adopts an existing directory. One workspace per Attempt is still the caller's protocol: while holding, the same owner could re-create a disposed path with a new inode, which the old ref never matches. So provision once, persist the ref durably before claiming, and never provision again or adopt another ref for an accepted Attempt.disposeWorkspaceremoves only the exact pinned directory when it is empty. It returnsremovedwithdurable: falseif the parent directory's fsync failed afterwards. Anything else isquarantined(orabsent) and left untouched. The worker disposes of the workspace only after the process it started is confirmed gone while the hold is still live, and itsremovedmeans gone at that moment. If no process started, the workspace is left for the caller.Manifest.
guardedFixtureManifest()hashes an explicit, test-verified list of the guarded execution's sources:attempt-fence.ts,attempt-workspace.ts,evidence.ts,fixture-child.ts,fixture-gate.ts,fixture-manifest.tsandlocal-worker.ts. The list is closed under relative imports, and the digest also covers the protocol and the collector. Changing a dependency alone changes the digest.parseExecutionManifest(value)validates a stored manifest exactly and recomputes its digest rather than trusting it. The runtime is pinned separately, in the scenario.Limits.
- A closed fence and a held fence prove admission and exclusion, not that a process exited; never record them as an exit.
node:sqliteignores the--permissionfile flags, and a directory grant covers everything inside that directory. These are fixed trusted programs cooperating, not confinement of arbitrary code, and not protection against a hostile concurrent writer in the state directory.- No power-loss qualification. Single host, local POSIX filesystem; tested on macOS with Node 26.8.1.
Tests: test/guarded-worker.test.ts covers the cases below, using test/helpers/guarded-worker-child.ts (a killable controller with start-up fault injection, and a pending writer), test/helpers/broken-entry-preload.mjs and test/helpers/json-lines.ts. test/attempt-workspace.test.ts and test/fixture-manifest.test.ts cover the workspace contract, including hold-bound provisioning, and dependency drift in a temporary source copy.
- healthy pass and genuine fail;
- invalid or mismatched guards, including another Attempt ID;
- lost or borrowed holds;
- a child stopped before its gate while the fence closes;
- a replaced or removed fence;
- direct gate refusals, including a missing control pipe;
- a gate waiting behind a pending writer;
- a bootstrap death before admission, a missing entry and a compile failure after admission;
- a timeout before admission;
- controller death;
- workspace swaps;
- cancellation.
Local factory CLI (src/cli.ts, src/local-factory.ts, src/guarded-verification.ts, src/attempt-ledger.ts)
The runnable composition: verification of a built-in fixture under Attempt ownership, with autonomous recovery from controller death within a pinned Budget.
npm run factory -- verify pass --run-id demo-pass --state-dir .factory # also fail, hang, noisy
npm run factory -- verify hang --run-id demo-hang --timeout-ms 500 --max-attempts 2
npm run factory -- inspect --run-id demo-pass --state-dir .factory # read-only journal
npm run factory -- cancel --run-id demo-hang --reason "no longer needed" # durable request
npm run factory -- recover --run-id demo-hang # reconcile; never dispatches
node src/cli.ts verify pass --run-id demo-pass # the same, without npm's header on stdout
State directory.
--state-dirdefaults to.factory(ignored by git), relative to the working directory; npm runs scripts from the repository root. It holdsjournal.sqlite,evidence/,fences/<owner token>/fence.sqliteandworkspaces/. Its canonical path is used throughout, and it must belong to this user and not be writable by others.Bounds.
--timeout-msis 1–60000 (default 10000),--max-output-bytesis 1–1048576 (default 65536) and--max-attemptsis 1–20 (default 3). Only plain decimal integers are accepted.Validation. All input is validated before anything is written. That covers the flags, the exact
verifyFixtureoptions (fixture,runId, optionaltimeoutMs,maxOutputBytes,maxAttempts) and theAbortSignal.Recipe
/2, pinned before the first Attempt:- the fixture;
- the complete guarded execution and collector source closure (
guardedFixtureManifest(): every listed file's SHA-256 and the digest over them); - the recipe, check and collector (
sf-guarded-fixture-worker/1); - the runtime (
node-<version> <platform>-<arch>), timeout and output bound; - the Budget.
The Candidate digest is the SHA-256 of
{"candidate":"sf-local-fixture/2","fixture":...,"execution":<closure digest>}. Reusing a Run ID with another fixture, bound or Budget is a conflict. The sources and runtime are checked against the pins before every dispatch instead. A drifted checkout dispatches nothing (unresolved), and a completed Run's current validity fails; it never silently re-pins. This includes a dependency-only change with the fixture entry untouched.Output. stdout is one JSON document, and errors are also summarised on stderr. It contains:
outcome,exitCode,detailandaction;historicalVerdict: as issued, even if its Evidence no longer backs it;currentValidity: the report and Evidence re-judged now, plus the source closure checked;- the last Attempt's
evidence,cleanupandtermination; attempts: per epoch, its status, whether its report isvalid,rejectedormissing, and its resources;steps: what this call did;pinnedandrun.
Resources vocabulary (each Attempt's
resources, andcleanupfor the last one):confirmed: its owner saw the process reaped with its pipes closed and removed its workspace;fenced: a recoverer closed its fence with no holder left, so no protected work can resume, and removed its empty workspace or found it gone. This is not a process exit;quarantined: something was left in place;no-process: no fixture process was started;unknown: nothing stored says.
Exit status:
- 0 verified (the only success);
- 1 failed, 2 inconclusive;
- 3 invalidated (a completed Run whose proof no longer holds);
- 4 cancelled;
- 5 not-completed (every budgeted Attempt used);
- 6 unresolved;
- 64 usage, 65 conflict, 66 not found, 70 error (including records that contradict each other).
action:executed: this call started a fixture process;recovered: it settled or fenced an earlier Attempt;recorded: it claimed, prepared or cancelled without starting a process;replayedandinspected: it changed no record, though SQLite may still coordinate its own journal files; only these two setreplayed: true.
Inspection.
inspectopens the journal withopenReadOnly. It makes no application or schema change, dispatches no fixture, probes no fence, and settles or records nothing. Normal SQLite reader coordination may touch the journal's-shm(and an empty-wal); the main file and a crashed controller's committed WAL keep their bytes, so inspect never checkpoints. A missing or empty journal isnot-foundand is not created. A foreign or unsupported one is anerror, and it is left unchanged.verify,cancelandrecoveropen it writable as before, and they create and checkpoint the journal as needed. Directory-hash equality is not the inspection invariant; unchanged records are. Earlier captures that saw the first inspect of a crashed controller's state checkpoint its WAL remain valid observations of the previous opening.
Dispatching an Attempt.
- Before the claim: a controller creates a fresh owner token, its own fence directory and fence, and holds the fence. It provisions the workspace under that hold and records the preparation durably in the attempt ledger: token, epoch, full FenceIdentity and WorkspaceRef.
- The claim: it then claims the queued Run as
local-cli/<token>. A claim that is definitely refused leaves nothing claimed, so the controller closes its own fence and removes its own empty workspace. Anything uncertain is left in place. A crash before the claim leaves only inert, preserved resources. - Before launch: right before the physical launch it rereads the Run, and does not launch if cancellation is pending or the sources or runtime drifted. A fresh claim is dispatch admission, not a spawn, and a replayed claim never admits anything.
- Running: the worker runs with the controller's own
AbortController. - Reporting and settling: the controller stores its immutable report under
report:<runId>/<epoch>/<token>while it still holds the fence, settles the Attempt from that report, and only then releases the hold.
Recovery. verify and recover bring a running Attempt they do not own to an end:
- With a valid report: a stored report that re-validates and confirms cleanup settles that same Attempt without re-running it. Re-validation covers exact identity including the token, the pinned expectation, the sources around the run, Evidence re-judged to its Verdict, and termination consistency.
- Otherwise, the fence decides:
- The Attempt's preparation is found from its worker name, and its fence is probed with zero wait.
- If the fence is held (a live owner, or its fixture), the verifier waits at most the pinned timeout plus 3 s, then reports
unresolved. No elapsed time, PID or lease takes over. - A missing, corrupt, replaced or WAL-mode fence, a damaged preparation, or a stored proof that is damaged or not this Attempt's own is
unresolved. Nothing is interrupted or recreated. - Once the fence is closed under exclusion, the workspace is removed only if it is exactly the provisioned empty directory (otherwise it is quarantined). The proof is recorded under
recovery:<runId>/<epoch>, and the Attempt is interrupted.
- What happens next: with cancellation pending the Run ends cancelled; otherwise it is queued for a fresh epoch, or exhausted by its Budget.
verifythen dispatches the next epoch. - Idempotent steps: a recoverer that dies after fencing, after recording the proof, or after interrupting is completed by the next call, and competing recoverers converge on one proof and one fresh epoch.
- Consistent records: every read requires the Run's pins and Budget to equal its stored request's; a contradiction is an error, never repaired.
Cancellation. It is recorded only through Execution.requestCancel, in the Run itself.
- Signals: SIGINT or SIGTERM, or an already-aborted
AbortSignal, is recorded durably first, and only then stops this call's own worker. If recording fails, nothing acts on it, and the failure is reported. - Other processes: a request from another process (
cancel) reaches a live owner, which polls for it while its worker runs. - Queued Runs are cancelled at once and never dispatched.
- Running Runs: a pending request forbids completion and any further epoch. It is settled
cancelledwith the first reason, from the owner's clean report or from the fence proof, so a controller that dies after the request never leads to a retry.
Legacy Recipe /1 Runs keep their narrower contract and are never upgraded, re-pinned or dispatched again:
- A completed
/1Run is rechecked against its pinned fixture source;fixture-child.tschanged for the gate, so/1Runs recorded before5c5addcreportinvalidated. - A
/1Attempt with a stored report that re-validates is settled from it. - A
/1Attempt without a report ran without a fence, so it staysunresolved. - A queued
/1Run is reported as retired.
Limits. One host, the fixed trusted fixtures, local POSIX filesystem, macOS with Node 26.8.1 tested, no power-loss qualification. The journal, Evidence store and state directory are trusted local storage: contradictions are detected where they break validation, but a writer with full access is not defended against. This is a local pilot, not arbitrary-code isolation.
Tests. All suites drive the real CLI or public API in subprocesses. test/helpers/factory-harness.ts injects faults by instrumenting CommandJournal.execute, EvidenceStore.put and node:diagnostics_channel: kills before or after a command, child stop/kill at spawn, kill or stop at child exit, barriers and source drift. Drift is tested only in a temporary copy of src/.
test/local-factory.test.ts: the/2contract.test/local-factory-recovery.test.ts: the recovery fault matrix, cancellation, Budgets 1, 2 and 20, and damaged proof.test/local-factory-legacy.test.ts: real/1state written by the last/1composition (f0dbdf6, extracted from git).test/local-factory-inspection.test.ts: inspect of missing, empty, foreign and unsupported journals, of one removed between the existence check and the open, and of a killed controller's committed WAL with cancellation pending. The main file and WAL are compared byte for byte, and only-shmmay differ. This is process-fault evidence, not power-loss qualification.test/journal-read-only.test.tscovers the opening itself.test/attempt-ledger.test.ts: the ledger's own boundary.
Confined executor (src/confined-executor.ts)
Runs one bounded, ordinary Node.js program against read-only snapshots inside a macOS kernel sandbox, and reports what the trusted host observed. Nothing composes it yet: no Recipe, CLI, repository or provider uses it. It observes; it never judges, stores Evidence or interprets what the program prints.
const runtime = describeRuntime(); // pin runtime.digest (node + library closure + launch tools + OS build)
const profile = confinedProfile(); // pin profile.digest (sf-confined-node-stateless/1)
const files = [{ path: 'check.mjs', content: bytes }];
const executor = new ConfinedExecutor({ root: '/abs/private/dir', maxRetained: 8 });
const result = await executor.run({
runtime: runtime.digest, profile: profile.digest,
program: { files, digest: snapshotDigest(files), entry: 'check.mjs', args: [] },
input: { files: inputFiles, digest: snapshotDigest(inputFiles) }, // becomes the read-only working directory
timeoutMs: 10_000, maxOutputBytes: 65_536, cpuSeconds: 10, maxDescriptors: 256,
}, signal);
Results. Invalid input throws
TypeErrorbefore anything is created. Otherwise:not-started: nothing ran. The reason isaborted,unsupported-host,runtime-mismatch,profile-mismatch,snapshot-mismatch,retained-limitorsetup-failed.observed: the process was reaped and its pipes closed.attributableis true only when all of these hold: the program was admitted under the verified limits, and the runtime and the namespace were unchanged afterwards.problemssays why otherwise.processholdsadmission, the measuredlimits,exit({code, signal}),stopped(null,timeout,cancelled,output-limit,limits-mismatchorcontrol-malformed),killSent, boundedstdout/stderr,outputBytesanddurationMs.identitiesholds the program, input, runtime and profile digests (template and instantiated) and the bounds.unresolved: the process was not confirmed gone, or its pipes stayed open.exitis set only if it was reaped, and the namespace is quarantined.
A stop reason is a fact, not a verdict. A program that fails to compile is an admitted, attributable failure.
Snapshots. At most 1024 files per tree and 16 MiB in total. Paths are relative, made of segments of
A-Z a-z 0-9 . _ -, and exclude.and... Case-insensitive duplicates and file/directory clashes are rejected. The executor writes only regular files: there are no links, modes or special files. Each run claimsroot/slot-<n>with an exclusivemkdir(n <maxRetained) and writes a freshsf-exec-<random>inside it, holdingprogram/,input/,home/and a hostcanary. The whole namespace is re-verified by device, inode, link count and content immediately before the synchronous spawn and again after collection.Launch.
/bin/dashruns a fixed script, and every value is a positional parameter. It sets CPU and descriptor limits (soft = hard), core 0 and file size 0, reports the effective values on descriptor 3, and execs/usr/bin/env -iwith exactlyLC_ALL=C TZ=UTC HOME=TMPDIR=<ns>/home OPENSSL_CONF=/dev/null. That runs/usr/bin/sandbox-exec -p <profile>, which runs the pinned node with adata:preload.- The preload checks that the canary is unreadable, then waits for the host's go byte on stdin. The host sends it only if the measured limits equal the request.
- The preload then writes
sf-confined/1 admittedand closes descriptor 3 before the program's entry loads. Only descriptors 0–3 are mapped, and stdin is at end-of-file for the program. - Output, markers and exit codes printed by the program carry no authority.
Profile
sf-confined-node-stateless/1. This is the tested deny-by-default profile from<local evidence file>:(deny process-info*)with a self-only allowance, and oneprocess-exec(the real node path);- read-only
program/andinput/; - named runtime library-directory subtrees and runtime files,
/System/Library,/usr/lib,/and/dev/{null,random,urandom}. The hashed static non-system library closure is derived from Mach-O load commands and independently matches anotoolderivation; it does not identify every readable public runtime file; - ancestor metadata, and the same twelve sysctls.
It grants no writable path, network, Mach service, signal or fork. It adds one tested denial:
file-map-executableon the snapshot, because a snapshot.dylib's constructor otherwise ran under the tested profile. Profile template, launch script, preload or tool-path changes alterprofile.digest; launcher-byte and OS-build changes alterruntime.digest. Unsupported hosts (only darwin/arm64 with Node v26.8.1 is proven), missing tools and pin mismatches start nothing. There is no fallback.
Tested denials (test/confined-executor-boundaries.test.ts, judged from host-side evidence with negative controls):
- Filesystem and SQLite. Reads, writes, stat, rename, truncate, chmod, symlinks, hard links, traversal and
/varaliases on a protected secret, another Product's files, a real Evidence store and another namespace are denied. Nativenode:sqliteread, write and create outside memory are denied. Worker threads are denied the same. All of those trees are byte-for-byte and inode/time-identical afterwards. - Network. TCP (IPv4 and IPv6 loopback, and a documentation address), UDP, a local DNS canary, Unix sockets, and TCP or Unix listening are denied. The host received zero connections and datagrams; an unconfined control shows the probe reaches them.
- Processes. Spawn, detached spawn, fork and shells are denied, including from a worker thread, and no process mentions the namespace afterwards. Signals and probes to a sentinel, the parent or launchd are denied; the sentinel received only the host's control signal.
- Native code. A snapshot library cannot be mapped. A C probe uses the instantiated policy plus explicit read/execute grants for that test executable; it is denied
fork,posix_spawn, a SecurityServer lookup and another process'sKERN_PROCARGS2. Without the explicit process-info rules that probe reads the canary, so the rules stay. It also measures effective limits and denied hard-limit increases through an independent launcher. The separate Node-FFI regression measures those limits and denied increases through the actual executor API and its protected startup path. - Descriptors and environment. A controller's secret descriptor and variables (
NODE_OPTIONS,DYLD_*, a token) never reach the program, while a deliberately mapped descriptor is readable in the negative control. - Broken setup. A broken profile, a permissive profile (the canary refuses it), a missing
sandbox-execand a limit mismatch never run the program.
Contract tests (test/confined-executor.test.ts) cover:
- pins, validation and root/alias refusal;
- the retained limit, including a slot claimed between count and claim;
- abort before and during launch (including an abort delivered between spawn and collection);
- timeout, output overflow, CPU and descriptor limits;
- forged control records and reopening descriptor 3;
- swaps of the program, namespace or input, and hard links.
Limits (not established):
- CPU time is not a termination bound.
RLIMIT_CPUsends SIGXCPU. A program with the default action dies, but one that handles SIGXCPU keeps running, sotimeoutMs(host wall time and SIGKILL) is the enforced bound. - No hard memory or aggregate-disk limit.
RLIMIT_AS/RLIMIT_DATAcould not be set on this host. The program has no writable path. Snapshot content is bounded to 16 MiB per reservation; filesystem overhead and the small control files are additional, so this is not a hard disk quota. - Controller loss. Automatic termination is not established after the controller dies. Its slot and namespace stay, are never reported as exited, and count against
maxRetained. The executor never reclaims uncertain resources; independently proved cleanup and recovery are still needed. - Scope of the program. One process only: no build tools, subprocesses, package managers or writable scratch.
- Host information. The program may read public system and runtime files and the listed machine sysctls; this is not zero host-information disclosure.
- Runtime files. Homebrew runtime files are user-writable. They are hashed before and after each run, but not held immutable in between.
- Tested scope. One host, local POSIX filesystem; tested on macOS (Darwin 27, arm64) with Node 26.8.1. Controllers sharing a root must use the same
maxRetained.
Repository (src/repository.ts)
Protected local source and sealed, immutable Candidates in a factory-owned bare Git repository. It validates bounded text-file proposals, seals a prospective commit on a pinned base, and exports the exact sealed files for the confined executor. It also performs the one local integration effect: a guarded move of the target ref from a Candidate's base B to its commit M, at most once per prepared Action occurrence. Nothing composes it yet. It never checks out, merges, rebases, contacts a remote, calls a provider, writes the journal, or decides authority. Grants, revocation, dispatch admission and owner replacement belong to the later Product gate.
const identity = await createRepository({path: '/abs/private/product.git', productId: 'product-a', targetRef: 'refs/heads/main', files: [{path: 'src/a.txt', text: 'alpha\n'}]});
const repo = openRepository(identity); // persist `identity` (JSON-safe); reopen it after a restart
const base = await repo.target(); // read-only: the commit the target ref names
const source = await repo.snapshot(base); // {commit, tree, files: [{path, mode, blob, sha256, bytes, text}]}
const candidate = await repo.seal(
{base, slice: {id: 'slice-1', revision: 1}, scope: ['src'], acceptance, recipe, environment}, // "sha256:<hex>" pins
{edits: [{path: 'src/a.txt', prior: source.files[0].sha256, text: 'ALPHA\n'}, {path: 'src/b.txt', prior: null, text: 'new\n'}]},
); // {digest, manifest, files, snapshotDigest}
const copy = await repo.exportCandidate(candidate.digest); // re-verified from protected objects
await executor.run({...request, input: {files: copy.files, digest: copy.snapshotDigest}});
const prepared = await repo.prepareIntegration({
action: {id: 'action-1', occurrence: 1}, product: 'product-a', slice: {id: 'slice-1', revision: 1},
repository: identity, ref: identity.targetRef, base, commit: candidate.manifest.commit,
candidate: candidate.digest, proof, owner: {id: 'owner-1', generation: 1},
}); // {protocol, digest, intent, payload}; persist it (JSON-safe)
const result = await repo.invokeOnce(prepared, finalRead); // only after the Product gate admitted this dispatch
// finalRead is optional: (claim) => {proceed: true} | {proceed: false, reason}, awaited just before the move
const later = await repo.inspectIntegration(prepared); // read-only reconciliation, e.g. after a restart
Creation.
createRepositoryrefuses an existing path (exists) and a parent that is not a canonical directory of this user or is writable by others. It lays out the bare repository itself, in a directory with mode 0700: fixedconfig,HEAD,objects/{info,pack}andrefs/{heads,tags}, with no hooks, templates,info/or description. It writes the initial files as one deterministic commit on the target ref (refs/heads/<name>), then an ownership markersf-repository.jsonwith a random nonce, last. The identity pins Product, canonical path, device, inode, nonce, object format (sha1), target ref and initial commit. Until the marker exists, every Git process re-checks the provisional root. A failed create leaves an unmarked directory that is never opened or adopted.Opening and ownership.
openRepository(identity, {gitTimeoutMs?})parses the identity exactly. It then requires:- the same canonical, private (0700) directory of this user, with the same device and inode;
- exactly the five top-level entries. The one exception is Git's own
HEAD.lock, held while a concurrent trusted writer updates the target ref that HEAD names. It is tolerated only as an empty regular file of this user with at most one link, or if it disappears during the check. A lock released mid-lookup can still be seen with no links; a hard link (two or more) is refused. It is never opened, removed or waited for. A target update that meets the held lock fails within Git's bounds, and after the claim that failure isunknown. Any other lock name or shape is an unexpected entry; - byte-identical marker,
configandHEAD(regular, singly linked files of this user); - empty
objects/infoandobjects/pack, so no alternates, grafts or foreign packs; - real directories of this user for every loose-object directory, the target ref's ancestors,
refs/sf/candidatesandrefs/sf/integrations, never links elsewhere.
Anything else is
unavailable. A missing or replaced repository is never recreated. OnlyopenRepositorymakes aRepository, and every Git process it runs repeats the check immediately before it starts and after it ends. The paths that process traverses (the ref it reads or creates, and the loose objects it writes) must also be real entries of this user. So no effect or result is accepted across a root replaced mid-operation (REP-1), and a linked internal directory cannot redirect Git's writes (REP-2).Supported source. Regular files of mode
100644holding valid UTF-8 without NUL bytes that are not LFS pointer files. Byte-order marks and CRLF are kept exactly. Paths are relative and portable: 1–240 characters and at most 16 segments ofA-Z a-z 0-9 . _ -, excluding.and.., and no segment may start with.git(case-insensitive:.git,.gitattributes,.gitmodules,.github). A tree holds at most 256 files (REPOSITORY_LIMITS.files), 128 KiB each and 1 MiB in total. Names that differ only in case, and paths that are both a file and a directory, are rejected. Links, executables, submodules and other modes are refused wherever they appear. These are a strict subset of the executor's snapshot rules.Snapshots.
snapshot(commit)takes a full 40-hex ID. It accepts only this repository's initial commit or one of its Candidate commits, in the exact fixed format (author, date and message). Seal commits, foreign commits and other Products' commits arecorrupt, and an unknown ID ismissing. Every tree is parsed and re-encoded canonically, and every object is re-hashed here rather than trusted from Git.Proposals.
{edits: [{path, prior, text}]}holds 1–64 edits and at most 256 KiB of paths plus replacement text:prior: nullcreates a file that must not exist;- otherwise
priormust equal the current file'ssha256; text: nulldeletes the file.
Edits must fall within the frozen
scope(a listed path or anything below it). The whole proposal is copied synchronously and checked against the base in memory, and the resulting tree is checked against every limit. Any of these fails withrejectedbefore any object, ref or file is written:- a malformed shape, including an accessor, holes, extra fields or a claimed
passed; - a bad path;
- a duplicate edit (compared ignoring case);
- a wrong or missing prior;
- binary, invalid or over-limit text;
- an unchanged file;
- an out-of-scope edit;
- a collision.
Malformed frozen inputs (
base,slice {id, revision ≥ 1}, a non-empty duplicate-freescope, andsha256:pins for acceptance, recipe and environment) areinvalid-input.Sealing. Every object is built here from validated bytes:
- the files, trees and the Candidate commit M, whose sole parent is the base B;
- the canonical manifest (
sf-candidate/1); - a seal commit S, whose tree is
{manifest.json}and whose sole parent is M.
They go as one pack to
git unpack-objects --strict, which also checks that every link resolves, including B. Everything is then re-read and verified, and only then is the single create-only refrefs/sf/candidates/<manifest sha256>→ S published. The Candidate digest is the SHA-256 of the canonical manifest bytes. It is distinct from the Git SHA-1 IDs.The manifest records:
- the repository identity, Slice, scope, and acceptance, recipe and environment pins;
- B and its tree, M and its tree;
- every file's path, mode, blob, SHA-256 and size;
- the proposal digest and each edit's prior/next digest.
M's message carries the Product, Slice and those pins, plus the scope and proposal digests, never the manifest. The manifest names M without a cycle. Repository location and ownership identity also affect the manifest digest; they need not change M when its source, parent and delivery inputs are identical.
Identical inputs return the same Candidate and write nothing. Concurrent identical seals converge on one ref. A process killed after writing objects leaves unreachable objects, which never qualify, and a later seal completes it.
Export.
exportCandidate(digest)reads the ref and verifies, independently of Git's own checks:- S and its manifest, parsed strictly and canonically;
- that the manifest's repository equals this handle's identity;
- M and every tree, rebuilt from the manifest and compared byte for byte;
- every blob against the manifest's size and SHA-256;
- that B is a factory commit with the recorded tree.
It returns fresh copies of the files and their executor
snapshotDigest. Mutating those copies changes nothing stored, and the executor refuses a mutated copy against the pinned digest. A missing ref ismissing. Any damage iscorrupt: missing, swapped or bit-flipped objects, a ref moved to another seal, a forged manifest, or objects copied from another Product. Nothing is ever repaired.Integration records. All of them are protected commits in the factory's fixed format, each named by its own create-only ref under
refs/sf/integrations/<key>/. The<key>is the SHA-256 of the canonical Action occurrence, so Action IDs never shape ref names. The top-level layout is unchanged.intent→ I: a root commit whose tree is{intent.json}, holding the canonical payload (at most 4 KiB).dispatch→ D: parent I, the empty tree and a fresh random nonce. This is the exclusive claim, written before the target is touched.receipt→ R: parent D. It records that Git's update exited 0 and that the readback named exactly M.
There is no negative receipt. This module never moves or deletes these refs.
Preparation.
prepareIntegration(payload)parses the payload exactly:action {id, occurrence ≥ 1},product,slice {id, revision},repository(a full identity),ref,base,commit,candidate,proof(sha256:) andowner {id, generation ≥ 1}. Any other field isinvalid-input, including a claimedpassedorverified. It then re-exports the Candidate withexportCandidate, so M is never taken from the caller. The payload's Product, repository identity, target ref, B, M and Slice must equal this repository's and the manifest's, or it isrejected; an unpublished Candidate ismissing.Only then is the intent retained. The same payload again returns the same
PreparedIntegrationand writes nothing. Any other payload for the same Action occurrence isconflict: a changed proof, owner, generation or Candidate needs a new occurrence. Preparation moves no target and grants nothing.proofis bound into the digest and retained, but it is not authenticated here; issuing and checking the protected proof belong to composition.One invocation.
invokeOnce(prepared)re-checkspreparedagainst its payload and the retained intent, and re-verifies the Candidate. It then decides in this order:- If a dispatch claim already exists, it is a replay: it returns the inspection and starts nothing.
- If the target no longer names B, it refuses to launch and makes no claim.
- Otherwise it writes a fresh D and runs
update-ref <dispatch> D <zero>. Only the call whose ownupdate-refexits 0 holds the claim. A claim held by another call is a replay. A claim with no acknowledgement throwsgit-failedand starts nothing, even if the ref turns out to name its own D. An existing identical ref is never taken as ownership. - Given a
guard(step 1, increment 3b), the claim holder awaits it once, now, immediately before the update. Only exactly{proceed: true}continues. A refusal{proceed: false, reason}(1–256 printable ASCII characters), a throw or rejection, any other answer, or no answer withingitTimeoutMsstarts no update (synchronous caller code cannot be interrupted, so an allowance also counts as late when the monotonic clock, read after the answer is validated, shows the bound has passed) and returnsinvocation: "withheld"withdispatched: false, the claim D, the target as read before the claim andproblem"withheld before submission: <reason>". The claim stays, so the occurrence is consumed and every later read reports itunknown. Without a guard this step does not exist. - The claim holder runs the single
update-ref --no-deref <target> M B, reads the target back, and writes R only if Git exited 0 and the readback named M. There is no retry, dereferencing, rebase, merge, amend or checkout.
So at most one target update is ever started per intent in this repository, whether calls repeat, overlap, run in other processes or follow a restart or failure.
invokeOncethrows only when it started no target update. Onceupdate-refhas been spawned, it returns an observation withdispatched: trueand leaves unresolved factsunknown, with the reason inproblem. That covers a failed exit, a signal, a timeout, a failed readback, a root replaced during the update and a receipt that could not be written. A failed exit or a readback other than M never becomes "not applied" (INT-1): Git can fail after the rename, and another writer may have moved the ref since.The caller must call
invokeOnceonly for a dispatch the Product gate has admitted, after its final current-state read, and should pass a guard that repeats that read. The guard narrows the window but does not close it: a change after it answers and before Git's rename is not seen. It must never treatunknownor a missing receipt as permission to prepare the same effect under a new occurrence while the earlier one may still be in flight.Inspection.
inspectIntegration(prepared)is read-only: no object, ref, retry or repair. It verifies the intent and Candidate, then reads the target, the receipt and the claim, in that order. Because a receipt is always written after its claim, and a claim before any target update, each read is consistent with the ones before it. It reports two kinds of fact separately:- Present:
target, andpostcondition. The postcondition isdesired-state-observedwhen the target names M and the intent and Candidate verified in the same call; otherwisenot-observed, orunknownwhen the target could not be read. - History:
dispatch,receipt, andinvocation. The invocation isnot-dispatched(no claim, so nothing was started),confirmed(a receipt), orunknown(a claim without a receipt).
Each observation has a random
observation.idand anobservation.attime. The postcondition does not say who wrote M, does not prove the earlier sender has stopped, and grants no further authority. A confirmed history survives later drift of the target. An unknown history is never resolved by M appearing or disappearing. Mismatched records arecorruptand dispatch nothing: a claim or receipt of another intent, a receipt without its claim, a receipt that names anything but M, or an intent ref naming another occurrence's intent. A forged intent for the same occurrence isconflict.- Present:
Git surface. Only
GIT_EXECUTABLE(/Applications/Xcode.app/Contents/Developer/usr/bin/git, the real binary rather than the/usr/bin/gitxcrun shim) runs, and only when it is root-owned and writable only by root. It runs with no shell and a fixed argument list per operation:unpack-objects -q --strict,cat-file --batch,update-ref --no-deref <ref> <id> <zero>(create only),update-ref --no-deref <target> <M> <B>(the one guarded move) andfor-each-refon one exact ref. There is no--stdinbatch, no symbolic-ref write and no other ref deletion or move. Every variable argument is a validated ID or ref.The environment is explicit:
GIT_DIR;GIT_CONFIG_NOSYSTEM=1andGIT_CONFIG_GLOBAL=/dev/null;GIT_ATTR_NOSYSTEM,GIT_NO_REPLACE_OBJECTS,GIT_NO_LAZY_FETCHandGIT_TERMINAL_PROMPT=0;HOME=/var/empty,LC_ALL=C,TZ=UTCandPATH=/usr/bin:/bin.
Hooks (
core.hooksPath=/dev/null), fsmonitor, reflogs, gc and maintenance are disabled both in the owned config and with-con every command. Each process has a time bound (gitTimeoutMs, 100–120000, default 30000), after which it gets SIGKILL andgit-failed. It also has a stdout bound derived from the source limits, which failscorrupt, and a 16 KiB stderr bound. Repository content is never executed.Errors.
RepositoryError.codeis one of:invalid-input: the caller's own arguments;rejected: the proposal, or an integration payload that does not match this repository or its Candidate;exists;unavailable: not the owned repository;missing: no such commit, Candidate or prepared integration;conflict: another payload is already retained for the same Action occurrence;corrupt: stored data failed verification;git-failed: Git could not run, failed or timed out.
Limits (not established):
- Trusted storage. The repository directory is trusted local storage. Damage and relabelling fail closed, but a writer with full access could forge a self-consistent Candidate by rebuilding M, its manifest and S, or forge integration records. Other trusted writers of the target are expected and reported, never prevented. Checks bracket each Git process; this is not atomic against a concurrent hostile writer, since Node has no directory-relative (
openat) operations. - Several refs are not one transaction. The claim, the target move and the receipt are separate single-ref
update-refruns on the files backend. A crash can leave a claim with no target change, or a target at M with no receipt. Those states stayunknown, and no crash atomicity across refs is claimed.not-dispatchedrelies on the claim being written, and never deleted, before any target update. That holds for process faults on trusted storage, not across power loss:core.fsync=committeddoes not fsync refs. - No external fencing. The claim makes this module start at most one update per intent. It cannot stop a sender that composition has already admitted, and it cannot stop another writer. A race loser that claimed before the winner moved the target stays
unknown, and composition must reconcile it. - Durability. No power-loss qualification. The owned config sets
core.fsync=committed; the marker and layout are fsynced. - Storage. Bounds apply to each source snapshot, proposal and process output. They are not a total repository disk quota; retention and aggregate storage budgeting remain unimplemented.
- Tested scope. Fresh owned repositories only, with no import of existing repositories, one direct target ref, SHA-1 object format and the files ref backend. No active product checkout is an integration target. Tested on macOS (Darwin 27, arm64) with Node 26.8.1 and Apple Git 2.50.1.
Tests:
test/repository.test.ts:- round trip with independent
gitreads andfsck --strict; - replay, concurrency and changed-input identities;
- a restart in a fresh process;
- 46 malformed or inapplicable proposals with byte-identical storage afterwards;
- invalid input;
- exact limits;
- one real confined-executor run on an export.
- round trip with independent
test/repository-boundaries.test.ts:- forged, moved, linked and replaced identities;
- REP-1 root swaps during seal and create;
- REP-2 linked
refs/sf,refs/sf/candidates,refs/heads, loose-object directories and object paths, with outside sentinels unchanged; - layout, marker and config tampering;
- cross-Product copies;
- object corruption, deletion and swaps;
- a forged manifest;
- unsupported stored modes and commits;
- hostile ambient Git config, hooks, templates and environment, with a negative control proving the sentinels are live;
- a kill after objects are written;
- a FIFO object hitting the time bound, an oversized object hitting the output bound, and an unwritable object store.
test/repository-integration.test.ts(child processes viatest/helpers/repository-integration-child.ts):- exact B→M with independent readback of the target, D, R and
fsck --strict; - preparation replay; changed-payload conflicts; cross-Product, mismatched, unpublished and malformed payloads with byte-identical storage;
- three rounds of two processes racing B→M1/M2, where exactly one wins and is preserved;
- three processes invoking one intent, where exactly one dispatches;
- a stale base refused before any claim; a failed guarded update after the claim staying
unknown; Astra's INT-1 case (real update, then another writer, then a failed close) stayingunknownafter reopening; - restart inspection and replay in fresh processes writing nothing;
- processes killed after the claim and after the acknowledgement;
- another writer establishing M, then drift;
- confirmed history surviving drift; a target at M with no claim;
- forged and mismatched claims, receipts and intents;
- a corrupt Candidate or intent;
- linked
refs/sf/integrations, Action ref directory, intent object directory andrefs/heads(after the claim, before spawn), with outside sentinels unchanged; - the root replaced before and after the update was spawned.
- exact B→M with independent readback of the target, D, R and
- The hostile-configuration regression in
test/repository-boundaries.test.tsalso prepares and invokes an integration, so the claim, move and receiptupdate-refruns are covered by the live-hook negative control.
Producers (src/producer.ts, src/claude-producer.ts, src/openai-producer.ts)
Bounded, proposal-only model transports. A producer turns one validated request into at most one provider exchange and returns an untrusted Repository Proposal for repository.seal, or an explicit non-success. It never applies, writes, integrates, verdicts or grants anything. The module is accepted under independent implementation review (Astra round 9 PASS, 2026-09-28, <local evidence archive>); that is not a live transport qualification. Nothing composes it yet, and neither transport is commissioned: until a person commits a qualification, every production produce is unavailable not-commissioned before anything is spawned or sent. Agent reference: producers.
const producer = new ClaudeProducer({root: '/abs/private/dir', gate, commission}); // commission: a parsed ProducerCommission, or null
const result = await producer.produce({
protocol: 'sf-producer-request/1', product: 'product-a', slice: {id: 'slice-1', revision: 1},
attempt: {runId: 'run-1', epoch: 1, attemptId: 'run-1/1', owner: 'worker-a'},
base: {commit, tree}, files: snapshotFiles, scope: ['src'], brief: {digest, text},
provider: 'claude', model: 'claude-opus-5-5', effort: 'high', allowance: 'allowance/slice-1',
bounds: {inputBytes, outputBytes, diagnosticBytes, deadlineMs, gateMs, maxEdits, maxEditBytes, maxFiles, maxOutputTokens, maxCostCents},
hardBounds: ['deadlineMs', 'inputBytes'], // bounds the Slice needs enforced, sorted
}, signal);
// {status: 'proposal', proposal: {edits: [{path, prior, text}]}, digest, bytes, receipt, allowance}
// | {status: 'unavailable' | 'failed' | 'interrupted', reason, message, receipt, allowance}
await repo.seal(frozen, result.proposal); // the repository re-decides everything
- Request.
parseProducerRequestcopies and validates exactly (descriptors, no extra keys, nothing defaults): Product, Slice, the Attempt identity with its owner (the ownership token the gate checks), the pinned base, 0 tomaxFilescontext files in path order whosesha256andblobare recomputed from their text (1 MiB together, the repository's complete layout rule), a sorted scope, a brief without control, separator or bidirectional characters, provider, exact model, effort, allowance, every bound withinPRODUCER_LIMITS, andhardBounds.requestDigestis the SHA-256 of the canonical request with each file reduced to its identity; the prompt digest covers the texts. - Capabilities and hard bounds. Each transport publishes
capabilitiesbefore any call: which bounds it enforces (host,provider,provider-estimate,none), whether its request count is exact, andprocessContainment: "none". A request naming a bound the transport only estimates or does not enforce inhardBoundsisunavailable unenforceable-boundbefore contact: Claude cannot enforcemaxOutputTokensand only estimatesmaxCostCents(--max-budget-usd); OpenAI enforcesmaxOutputTokensand has no spend bound. - Commission and gate. A
ProducerCommission(written by the qualification below and committed by a person) binds provider, model, the transport-policy digest (CLAUDE_POLICY_DIGEST/OPENAI_POLICY_DIGEST, which change with every reviewed rule change), the observed managed-policy state, the completeness audit and the qualification receipt. Any difference isunavailable qualification-mismatch. Step 1 supplies theDispatchGate:admit(dispatch)is called once, after all preparation and immediately before spawn/send, with the dispatch ID, request digest, transport digest (and the Claude call directory), attempt owner and capabilities; a denial, an exception or no answer withingateMson the monotonic clock (a late answer, even one delivered after the gate blocked the event loop, is no answer) isinterrupted not-admittedand nothing is sent.record(consumption)is called once afterwards: a thrownGateRejectionis a definiteno, anything elseunknown, never retried and never changing the status. - Wire format. The model returns exactly
{"edits":[{"path","baseSha256","content"}]};parseProposalTextchecks the string before any encoding (well-formed Unicode, withinoutputBytes, one JSON object, no fence or commentary), then paths, scope, case-insensitive duplicates, digests against the supplied files, content, the repository's byte accounting and the layout of the visible result, and maps each edit toFileEdit {path, prior: "sha256:" + baseSha256 | null, text: content}in path order.malformed-proposalis a shape or Unicode fault,rejected-proposala well-formed edit that breaks a rule. Messages name edits by index, never by the provider's path. - Results.
unavailable(not-commissioned, qualification-mismatch, unenforceable-bound, executable-missing/-unreadable/-mismatch, no-platform-credential, authentication),failed(input-overflow, spawn, exit, missing-init, malformed-event, unexpected-event, unexpected-field, capability-inventory, wrong-model, tool-use, multiple-turns, inconsistent, refusal, truncated, incomplete, cli-error, provider-error, output-overflow, malformed-proposal, rejected-proposal) andinterrupted(deadline, cancelled, not-admitted).remoteisnonewhen nothing was sent,completedonly for a validated terminal response (for OpenAI: aresponseobject with aresp_id, the exact model and a status whoseerrorandincomplete_detailsare consistent with it; for Claude: a successful,terminal_reason: "completed"result with valid numbers and exactmodelUsageidentity, from a CLI that exited on its own, closed its pipes before the deadline and wrote nothing after it) or a definitive 401/403 rejection, andunknownotherwise (a Claude CLI error result included).localcarries the call directory disposition and process facts into the consumption record, so step 1 can block re-dispatch onretained, a non-empty group orremote: unknownuntil it has reconciled. - Retention. Provider values enter receipts, consumption records, messages and
sf:producerchannel notifications only as fixed classifications, grammar-checked tokens (GRAMMARS), checked numbers, byte counts or digests. stderr and error bodies are counted and digested, never kept; inventories are compared with the allowlist and recorded as the policy value or as{key, kind, count}deviations with the digest of the raw line or body. Every list has aRECEIPT_LIMITScap, so receipts always canonicalise. - Claude transport. Only on darwin/arm64, only the pinned 2.1.283 executable (hashed through an
O_NOFOLLOWdescriptor immediately before every spawn), onlyclaude-opus-5-5. The root must be canonical, private and outside any directory holding.git,.claude,CLAUDE.md,AGENTS.mdor.mcp.json; each call gets<root>/call-<random>/{cwd,tmp}. Argv is fixed (--print --verbose --output-format stream-json --model --effort --tools "" --safe-mode --restricted --strict-mcp-config --mcp-config '{"mcpServers":{}}' --setting-sources "" --settings '{}' --disable-slash-commands --no-chrome --no-session-persistence --permission-mode dontAsk --permission-prompts none --max-budget-usd --system-prompt), with no--bare, resume, fallback model or--json-schema; the prompt goes on stdin. The environment is exactlyHOME(the real one, so the existing sign-in works),PATH=/usr/bin:/bin,LANG=en_US.UTF-8andTMPDIR; descriptors 0-2 only; its own process group. Every managed-policy source the binary names (/Library/Application Support/ClaudeCoderecursively, and the system and per-usercom.anthropic.claudecodeMDM plists) is observed before every spawn and must equal the commission's record; links, deeper directories, more than 16 files or 1 MiB, or ill-named entries are refused outright. - Stream rules. The first event must be
system/init; its key set, version, model, cwd and inventory must match (CLAUDE_INERT_INVENTORY: no tools, MCP servers, skills, commands, agents, plugins or memory paths;permissionMode: dontAsk; only the observed protocol capabilities). Then onlysystem/thinking_tokens,rate_limit_event, one assistant message (text, thinking and redacted thinking blocks; any tool block stops the run) and one successful result whose text equals the assistant text, with one turn, no permission denials,end_turn,terminal_reason: completedand exact model usage. A protocol violation, overflow, deadline or cancellation kills the group at once while the process is unreaped; a reaped process is never signalled. The deadline is checked on the monotonic clock whenever output, the exit or the pipes' close is delivered, and its timer expires any exchange not yet ended, even after the exit: output, an exit or a close delivered after the deadline, including after an exit observed in time, isinterrupted deadlinewithremote: unknown, never a proposal. A proposal also needs exit 0, the whole prompt written and closed pipes. - Cleanup. Only an unreaped process is ever signalled destructively, while its PID and group number cannot be reused. After reaping,
kill(-pgid, 0)only observes the group, repeated for up toCLAUDE_KILL_CONFIRM_MSuntil it answers ESRCH (empty); otherwise the last answer decides. Signal 0 shows existence, not ownership, so a present group (a survivor or a reused number) isaliveand is never signalled; EPERM or another error isunknown. EPERM is observed through first: a member still being torn down after the group SIGKILL answers EPERM for a few milliseconds before ESRCH (seen on macOS under CPU load). The qualification script settles its discovery control the same way. The call directory is removed entry by entry (links unlinked, never followed) only when the process was reaped, its pipes closed and its group empty; otherwise it is retained with the reason and reported inlocal. A never-spawned directory isremoved-unsent. - OpenAI transport (design candidate).
buildOpenAiRequestmakes the exact body (tools: [],tool_choice: "none",store: false,stream: false,truncation: "disabled",background: false, the strict proposal schema) and a credential-free manifest whose digest the gate admits. The module-private send boundary rebuilds and compares both byte for byte before it touches the credential, sends once withredirect: "error"and no retries, rechecks the deadline on the monotonic clock and cancellation after the headers and every settled body read (an overdue or cancelled exchange isinterrupted,remote: unknown), caps the body, and treats any body containing the credential asprovider-errorwithout parsing it. Responses are parsed strictly (known keys, exact model, no advertised tools, exactly one completed assistant message).OpenAiCredentialkeeps its value in a module-private record: it never serialises, prints or inspects to the value. - Seams.
unsafeClaudeTransportPins,unsafeOpenAiEndpoint(loopback HTTP only),unsafeTestCredential(sk-test-only),unsafeOpenAiSendBoundaryandunsafeQualificationRunare the only ways to reach test or qualification paths; a test proves production code never calls them. Pairing is decided from the credential's private record: a production credential never goes to a test endpoint and a test credential never leaves loopback. - Qualification.
node scripts/qualify-claude-producer.ts --archive <new absolute dir> --auditor <name> [--completeness-audit <file>] [--test-summary <file>]spends: a discovery control with sentinels and an empty HOME (no prompt), one legitimate edit, and (only if that met its required outcome in full) one adversarial call, never more than three spawns. It writesdocs/claude-producer-receipt.json,docs/openai-producer-receipt.jsonand, only foraccepted-scoped,docs/claude-producer-commission.json; committing them is the person's act.--openai-onlyspawns nothing and records the OpenAI transport as unavailable without a Platform credential. - Qualification result (2026-09-28). The qualification receipt is for rules revision 4; the current rules revision 5 requires re-qualification. The receipt is kept byte-unchanged as the record of that run (policy digest
sha256:4685bfc7…5246); the rule changes made in review since (deadline classification,remoteevidence, observation-only cleanup after reaping) are revision 5, and a commission binding the revision-4 digest is refused asqualification-mismatch. Claude:unavailableand no commission (receipt; archive<local evidence archive>). The discovery control exited 1 after 4.4 s with no init event, so it isnot-established, although its SessionStart hook and MCP sentinel both ran. Call 1 made one spawn. Its init named version 2.1.283 and modelclaude-opus-5-5, with no tools, MCP servers, skills, commands or memory paths, but with 5 agents and 2 plugins. It therefore endedfailed capability-inventory: the group was killed 188 ms after spawn, before any assistant or result event.remote: unknown, spend unknown, group empty, call directory removed; sentinels, canary and both repositories were unchanged.pssaw one new, unattributedmdworker_sharedprocess, which the rule counts. Call 2 did not run. Authentication under the narrowed settings was not observed. Per plan step 3 the CLI transport is unavailable: the built-in agent and plugin metadata has no established inert meaning, and no completeness audit exists. The narrower alternative is a separately authorised Messages API adapter withtools: []. OpenAI:unavailable no-platform-credential(receipt). 80/80 offline tests pass; no live request was made. The OpenAI receipt is for rules revision 4; revision 5 requires re-qualification when a Platform credential exists. The receipt is kept byte-unchanged as that record (policy digestsha256:31243da3…115d); the review changes to itsremoteand deadline rules are revision 5, and a commission binding the revision-4 digest is refused asqualification-mismatchwith nothing sent.
Limits (not established):
- Containment. No containment of the Claude CLI's process tree exists: a descendant that changes session is invisible to the group check, so
quiescenceis alwaysnot-established. The transport is a trusted host component that runs no product code; a Slice for which that is unacceptable does not use it. - Provider side. The CLI's internal request count and retries are not observable; a kill or an aborted request does not prove remote work stopped; no billed cost is known.
- Inventory completeness. Matching the allowlist proves conformity to the audited policy, not that the init inventory lists every route; that is a person's completeness audit.
- Pins. The executable is checked by path immediately before spawn; a replacement between the hash and execve is not excluded. The CLI's own state under HOME is neither confined nor inspected.
- OpenAI. No live request has been made; the response shape, effort echo and snapshot naming are unverified, and observed
xhighexecution is never claimed from the manifest.
Tests:
test/producer.test.ts: every request rule, request and commission digests, the prompt fixture, every parser rule, consumption and acknowledgement classes, and receipt bounds.test/claude-producer.test.tsandtest/claude-producer-boundaries.test.ts(a fake CLI run by node, pinned by node's real SHA-256): exact argv and environment, commission and managed-policy binding, hard bounds, admission, init inventory, identity, stream and result rules, exit precedence, stops, pins, cleanup, canary scans, process groups, construction refusals and adversarial context, with negative controls for the environment and descriptor canaries.test/openai-producer.test.tsandtest/openai-producer-boundaries.test.ts(a loopback fake server): manifest exactness, response rules, credential echo, statuses, deadlines, redirects and send-boundary tampering.test/claude-producer-seam.test.ts: seams, policy digests and credential pairing.
