Factory docs, home
Page navigation

Verifies one built-in fixture under exclusive Attempt ownership (Recipe sf-local-fixture-verify/2). When a controller dies, it brings the orphaned Attempt to an end from its stored report or a closed fence, within the Run's pinned Budget.

Status: "Accepted for trusted fixtures" (build review): durable preparation, immutable reports, exact ownership, bounded replacement and original legacy contracts; independent 23-case recovery review plus actual CLI crash recovery. Receipts: continuation receipt (recovery, cancellation and ledger review by gpt-6-astra at xhigh; its hashes for src/attempt-ledger.ts and for this module's tests and helpers still match the files at 574e70c) and inspection receipt (read-only inspection; its hashes for guarded-verification.ts, local-factory.ts, cli.ts and factory-support.ts match the files at 574e70c, and supersede the continuation receipt's hashes for those four files). This is local synthetic-fixture infrastructure, not a pilot. No epic is complete; all 27 remain open. Human page: guide.

Source: src/guarded-verification.ts (policy), src/attempt-ledger.ts (durable Attempt records). Composition: src/local-factory.ts, src/cli.ts, src/factory-support.ts. Tests: guarded-verification, attempt-ledger, local-factory, local-factory-recovery, local-factory-legacy, local-factory-inspection, local-factory-observer-boundary, inspection-boundaries.

Intelligence: none — Picks the next step from durable facts under fixed rules, so any controller reproduces the same recovery; a model would make it unrepeatable.

What it hides

  • The dispatch protocol. A fresh owner token, fence and workspace; a durable preparation; the claim; a recheck right before launch; the guarded run; the report written under the hold; then settlement. Callers see four operations and one FactoryResult.
  • Recovery decisions. For a running Attempt this call does not own, it chooses one path: settle it from a report that re-validates; wait on a held fence; close the fence and interrupt; or return unresolved.
  • Cancellation bridging. A caller's AbortSignal becomes a durable Execution.requestCancel before it stops anything. Requests from other processes are polled.
  • Presentation. The historical Verdict, current validity, and each Attempt's report and resource states, all derived from stored records only.

Public interface

src/guarded-verification.ts constants: GUARDED_RECIPE_REVISION = "sf-local-fixture-verify/2", DEFAULT_MAX_ATTEMPTS = 3, RECOVERY_WAIT_MARGIN_MS = 3_000 (a held fence is waited on for scenario.timeoutMs + 3_000), DEFAULT_CANCEL_REASON = "cancellation was requested; the Run records no Verdict". Private values: check "fixture-exit-status", Candidate kind "sf-local-fixture/2", fence poll every 50 ms, cancellation poll every 100 ms, 8 settle tries.

Every operation takes ctx: Context (factory-support.ts): { runId, stateDir, paths: StatePaths, journal: CommandJournal, execution: Execution, store: EvidenceStore }, where StatePaths = { root, journal, evidence, fences, workspaces }.

Request (pure, except readBoundRun):

  • buildGuardedRequest(inputs: GuardedInputs): GuardedRequest. GuardedInputs = { runId: string; fixture, timeoutMs, maxOutputBytes, maxAttempts, execution, environment: unknown }.
    • Refuses with FactoryError("invalid-input"): a fixture outside FIXTURES (pass, fail, hang, noisy); timeoutMs outside 1..MAX_TIMEOUT_MS (60 000); maxOutputBytes outside 1..MAX_OUTPUT_BYTES (1 048 576); maxAttempts outside 1..EXECUTION_LIMITS.maxAttempts (20).
    • A non-string environment is a plain Error. execution goes through parseExecutionManifest, which recomputes the digest and throws TypeError on any mismatch.
    • Derives candidate.digest = digestOf({ candidate: "sf-local-fixture/2", fixture, execution: execution.digest }).
    • Derives recipe = { revision, check: "fixture-exit-status", expectedExitCode: 0, collector: "sf-guarded-fixture-worker/1" }.
    • Derives scenario.revision = "local-fixture-scenario/" + digestOf({ environment, timeoutMs, maxOutputBytes }) and budget = { maxAttempts, attemptTimeoutMs: timeoutMs }.
  • storedGuardedRequest(value: unknown, runId: string): GuardedRequest. It rebuilds the request from the stored inputs and requires canonical equality; otherwise it throws "…is not consistent with its own pins".
  • isGuardedRecord(value: unknown): boolean. True when recipe.revision === GUARDED_RECIPE_REVISION.
  • runPinsOf(request: GuardedRequest): RunPins. Returns { candidateDigest, recipeRevision, scenarioRevision }.
  • readBoundRun(ctx: Context, request: GuardedRequest): Run | undefined (:195). Throws unless request.runId === ctx.runId and the Run's [pins, budget] equal the request's. It returns undefined for a Run that was never submitted.

Operations (each returns Promise<FactoryResult>):

  • verifyGuarded(ctx: Context, request: GuardedRequest, environment: string, signal: AbortSignal | undefined) (:707).
    • Binds first, then creates paths.fences and paths.workspaces (mode 0700), then records an already-aborted signal.
    • Then loops for at most 4 × maxAttempts + 4 rounds: a running Run is reconciled; a queued Run with an abort wanted is cancelled; any other queued Run is checked for source or runtime drift (drift blocks) and then dispatched (dispatchOne, :580).
    • Stops at a terminal Run or the first blocker. action is executed, else recovered, else recorded, else replayed.
  • recoverGuarded(ctx, request). Reconciles only, for at most 2 × maxAttempts + 2 rounds, and never dispatches. action: recovered or replayed.
  • cancelGuarded(ctx, request, reason: string). Records the request durably; if it is pending, reconciles once with zero wait. action: recorded or replayed. A Run that was never submitted is a plain Error here; the composition refuses it as not-found first.
  • inspectGuarded(ctx, request). Presentation only: no fence probe, no record change, no process. It still reads the Evidence store and hashes the sources for currentValidity. action: inspected.
  • interface GuardedReport { runId, attemptId, epoch, worker, token, expectation, execution: { beforeLaunch, afterRun }, cancelRequested, result }.
    • The execution digests are each sha256:… or null.
    • result.status is observed (evidence, termination, verdict | null), storage-failed (message, termination) or not-started.
    • not-started carries a reason (aborted, setup-failed or drift), a message and the owner's WorkspaceDisposal. Stored reports are parsed with exact keys.

Result (factory-support.ts): FactoryResult = { command: "verify" | "inspect" | "cancel" | "recover"; runId; outcome; detail; action; replayed (true only for replayedandinspected); historicalVerdict: Verdict | null; currentValidity: { state: "valid" | "invalid" | "not-applicable"; verdict; reasons }; evidence (the last Attempt's); cleanup: Cleanup (the last Attempt's resources); termination: TerminationSummary | null; attempts: AttemptView[]; steps; pinned: PinnedRequest; run: Run; stateDir }. PinnedRequest = LegacyRequest | GuardedRequest; the guarded operations always give the /2 branch. AttemptView = { epoch, worker, status, report: "valid" | "rejected" | "missing", resources: Cleanup }; Cleanup.state is confirmed, fenced, quarantined, no-process or unknown.

outcome (statusOf, :885): a completed Run gives verified, failed or inconclusive from its historical Verdict when recheck finds it still valid, otherwise invalidated; cancelled gives cancelled (detail is the cancelReason); exhausted gives not-completed; queued and running give unresolved, with the blocker as detail when there is one.

src/attempt-ledger.ts:

  • TOKEN_PATTERN (a lowercase UUID). workerFor(token): string returns local-cli/<token> and throws LedgerError for a bad token. tokenOf(worker): string | undefined.
  • new AttemptLedger(journal: CommandJournal, runId: string). It follows Execution's Run ID rule ^[A-Za-z0-9][A-Za-z0-9._:-]{0,63}$ and accepts epochs 1..20. Every record is create-once (expectedVersion 0, :137), and each write emits one event: attempt.prepared, attempt.reported or attempt.fenced.
MethodAggregate · command IDNotes
recordPreparation(p: Preparation): voidattempt:<runId>/<token> · verify:<runId>:prepare:<token>worker must equal workerFor(token); fence and workspace must name the same Run, epoch, worker and fence nonce
readPreparation(token): Preparation | undefinedsameRe-validated on read; a malformed record throws LedgerError
recordReport(epoch, token, report: JsonValue): voidreport:<runId>/<epoch>/<token> · verify:<runId>:report:<epoch>:<token>The caller validates the content
readReport(epoch, token): JsonValue | nullsameFound only under its own epoch and token
recordFenceProof(proof: FenceProof): FenceProofrecovery:<runId>/<epoch> · verify:<runId>:fence-proof:<epoch>Returns the proof that stands: the first one
readFenceProof(epoch): FenceProof | undefinedsameRe-validated on read

Preparation = { runId, token, worker, epoch, fence: FenceIdentity, workspace: WorkspaceRef }. FenceProof = { runId, epoch, token, fenceNonce (32 hex), workspace }, where workspace is removed { durable }, absent or quarantined { path, reason }. class LedgerError extends Error.

Composition (src/local-factory.ts, src/cli.ts):

  • verifyFixture(stateDir, options: VerifyOptions, signal?), inspectFixture(stateDir, runId), cancelFixture(stateDir, runId, reason = DEFAULT_CANCEL_REASON) and recoverFixture(stateDir, runId). Each Run ID must follow the rule above, and a cancel reason must be well-formed text of at most 500 UTF-16 units with non-whitespace content and no control characters; otherwise invalid-input.
  • VerifyOptions = { fixture, runId, timeoutMs? (10_000), maxOutputBytes? (65_536), maxAttempts? (3) }. Options must be a plain object; any other field is refused; defaults apply only to absent fields, so an explicit undefined or null is invalid-input; a supplied signal must be an AbortSignal. The runtime is pinned as node-<version> <platform>-<arch>.
  • Also exported: DEFAULT_TIMEOUT_MS, DEFAULT_MAX_OUTPUT_BYTES, RECIPE_REVISION = "sf-local-fixture-verify/1" (legacy Runs are read, settled from a valid report or cancelled, and never dispatched or upgraded), and re-exports of DEFAULT_MAX_ATTEMPTS, GUARDED_RECIPE_REVISION and the result types.
  • CLI: npm run factory -- verify|inspect|cancel|recover, with main(argv, signal?): Promise<CliOutput> ({ exitCode, stdout, stderr }) and runCli(argv): Promise<number>.
  • EXIT_CODES: verified 0, failed 1, inconclusive 2, invalidated 3, cancelled 4, not-completed 5, unresolved 6, usage 64, conflict 65, not-found 66, error 70. DEFAULT_STATE_DIR = ".factory".
  • State: journal.sqlite, evidence/, fences/<token>/fence.sqlite and workspaces/, all under the canonical (realpathSync.native) state path. The request is stored under verification:<runId> (command verify:<runId>:request) and the Run is submitted with verify:<runId>:submit.

Invariants and guarantees

  1. Bound before effect. Every public operation calls readBoundRun first. A request for another Run, or a Run whose pins or Budget differ from its stored request, throws and changes nothing. Such records are never repaired or re-pinned (guarded-verification.test.ts; the recovery test "records that contradict each other").

  2. Self-consistent pins. storedGuardedRequest rebuilds the request and compares canonically. Reusing a Run ID with another fixture, bound or Budget is a FactoryError("conflict") (guardedDifferences in local-factory.ts). Sources and runtime are checked before every dispatch instead.

  3. Preparation before claim. The token, epoch, full FenceIdentity and WorkspaceRef are durable before Execution.claim runs as local-cli/<token>. The worker name of any accepted claim therefore leads to its fence.

  4. Launch admission. runGuarded runs only after a fresh, non-replayed claim that is still the current ownership, with no cancellation observed at the final admission check, and with guardedFixtureManifest().digest and the runtime equal to the pins. A request committed after that check can meet one launch in flight, which the owner's poll then stops. Right before the claim the Run is read again: a Run no longer queued at the version read, or a caller abort already wanted, means no claim. A claim that is definitely refused closes only this call's own fence and disposes only its own empty workspace; anything uncertain stays in place.

  5. Record cancellation, then stop. A caller's abort is stored through Execution.requestCancel before own.abort(). If storing fails, the worker is not stopped on that account. The failure becomes the blocker of an unresolved result while the Run is still queued or running; if the dispatch completes meanwhile, the terminal outcome is returned and the failure appears in neither steps nor detail (:726, :763, :894). The worker never receives the caller's signal.

  6. Report under hold. hold.assertActive() runs before and after the report is built (:648, :665). The report is keyed by exact epoch and token, and settlement from it is attempted under the same hold. If that settlement stays unresolved (cleanup not confirmed), the owner releases its hold and takes the fence route like any recoverer; the Attempt may stay running.

  7. Verdicts only for pinned sources. verdictOf returns null unless the source digest before launch and after the run both equal the pin. A null Verdict fails the Attempt; it never completes it.

  8. Settle only from proof. judgeReport (:259) accepts a report only when all of these hold:

    • exact identity: Run, Attempt, epoch, worker and token;
    • the expectation rebuilt from the pins;
    • for an observed result, source digests that agree with whether a Verdict is present, and exactly one Evidence record; with a Verdict present, that record must judge again to exactly the stored Verdict, agree with the termination summary and stay within the pinned output bound (observationProblem).

    A valid report settles the Attempt only when cleanup is confirmed (unsafeCleanup, or no quarantine for not-started). Anything else takes the fence route.

  9. No takeover by time. A held fence is waited on for at most timeoutMs + RECOVERY_WAIT_MARGIN_MS (zero under cancel), then the result is unresolved. No elapsed time, PID or lease takes over, and a fence is never recreated.

  10. Interrupt only with its own proof. After tryFence closes the fence, the workspace is removed only if it is exactly the provisioned empty directory; otherwise it is quarantined. The proof is recorded (the first proof stands) and must match the preparation's token, fence nonce and quarantine path. Execution.interrupt then records confirmedBy: "sf-attempt-fence/1" with a method that says "not a process exit".

  11. Cancellation wins. With a request pending, the Attempt settles through Execution.cancel with cleanup, or interrupt ends the Run cancelled. A report with cancelRequested never completes a Run, and Execution refuses any further epoch.

  12. Idempotent convergence. Command IDs are deterministic: verify:<runId>:claim:<token>, …:settle:<epoch> and …:recover:<epoch>. Only a cancellation request uses a random ID (verify:<runId>:cancel:<uuid>); the Run's version and pending state guard it. Ledger records are create-once, so concurrent or crashed recoverers converge on one proof and one fresh epoch.

  13. Historical versus current. historicalVerdict is shown as issued. currentValidity (recheck) judges the Evidence again and hashes the sources again now. invalidated never rewrites the Run.

  14. Bounded work. Verify runs at most 4 × maxAttempts + 4 rounds and recover at most 2 × maxAttempts + 2. Settlement and cancellation recording each try at most 8 times. reconcileRunning itself has no counter: it ends when the Attempt changes, a blocker appears or the fence wait runs out.

Failure semantics

  • Refused. FactoryError codes invalid-input (CLI 64 usage), conflict (65) and not-found (66). Invalid input is rejected before anything is written. A journal missing at the existence check is not-found before any open (inspect, cancel, recover); inspect also returns not-found during its read-only open for an empty journal or one removed after that check, without recreating it; a missing request or Run in an existing journal is not-found after the open (writable for cancel and recover, which may initialise an empty file, recreate one removed after the check or checkpoint it). conflict is raised after verify's writable open, which may create or checkpoint the journal. A refused call dispatched, claimed and settled nothing.
  • Contradiction or corruption. A plain Error: bound-Run mismatch, an inconsistent stored request, a malformed verification:<runId> record, or a Run that "kept changing while cancellation was requested". The CLI reports these as 70 error, and nothing is repaired. A LedgerError from a stored record is caught. A damaged preparation makes reconciliation unresolved even when a report exists; a damaged fence proof blocks only the fence route, so a valid, clean report still settles. Presentation with an unreadable preparation or proof falls back to what the report says, and shows resources.state: "unknown" only when no valid report exists. Only a malformed selector, which validated input prevents, would surface as 70.
  • unresolved (6) is a returned outcome, not an exception. detail names the blocker:
    • the worker name is not local-cli/<uuid>, so nothing can prove it stopped;
    • the preparation is missing, damaged or for another epoch;
    • the fence is not usable proof (missing, replaced, WAL-mode, corrupt or open in this process, so tryFence throws);
    • the fence is still held when the wait ends;
    • the fence proof is damaged or belongs to another Attempt;
    • the Attempt kept changing across 8 settle tries;
    • the claim was not the current ownership afterwards;
    • ownership was lost before or after launch;
    • the sources or runtime drifted while the Run was queued;
    • the cancellation could not be recorded, while the Run is still queued or running (invariant 5).
  • Drift after the claim. Drift found before launch, by the owner's recheck or the worker's own, yields a not-started/drift report. Drift found around an observed run (a digest before launch or after the run that differs from the pin) yields an observed report with verdict: null. With confirmed cleanup and no cancellation pending, either fails that Attempt. While Budget remains, the next round's drift check blocks with unresolved; with none left the Run is exhausted (not-completed).
  • Unknown is not failed. resources.state: "unknown" means nothing stored says what happened. report: "missing" or "rejected" shows an absent or unproven account. Neither is ever presented as a failed fixture.
  • Retries inside the module. VersionConflictError, CancelPendingError and CommandConflictError retry settlement. StaleAttemptError and RunTerminalError mean another caller ended the Attempt. A refused claim (VersionConflictError, AttemptActiveError, RunTerminalError or CancelPendingError) abandons only this call's own resources. A replayed claim never launches.
  • Leftovers before the claim. A controller that dies after provisioning but before its claim leaves an inert fence directory and workspace, and a preparation record if it was committed. Recovery follows only a running Attempt's worker name, so nothing reclaims them later. They consume no Budget, which bounds claimed Attempts only.
  • Replacement. After a failed or interrupted Attempt, Execution queues the Run again while its Budget allows, and otherwise marks it exhausted, reported as not-completed (5). Only verify dispatches the fresh epoch. The Budget is never extended.
  • Replays. A repeated verify of a finished Run is replayed. An exact ledger repeat replays. A different preparation or report under the same key raises CommandConflictError; recordFenceProof swallows that conflict and returns the first proof instead.

Trust scope

Established (local, fixed trusted fixtures):

  • Recovery from controller death at every protected boundary: before and after preparation, after the claim, while the fixture runs, after exit, after Evidence and after the report. Each recovery uses one fresh payload per epoch. Budgets of 1 and 2 are used up exactly by controller deaths, a Budget of 20 is preserved through one death and completes at epoch 2, a terminal Run never resurrects, and competing verifiers converge.
  • A recoverer killed after closing the fence, after recording its proof or after interrupting is finished by the next call. A fixture child stopped before its gate and orphaned by controller death is fenced out and adds no Evidence.
  • Actual CLI crash recovery (recovery_cli_demo in the continuation receipt):
    • SIGKILL after the claim commit, then verify: verified, with Attempts [interrupted, completed].
    • SIGKILL after the report commit, then recover: verified without a second run.
  • Durable cancellation survives controller death and forbids any retry. Read-only inspection is covered by the inspection receipt.
  • Legacy /1 Runs keep their original pins and Budget. Their missing-report Attempts stay unresolved, and they are never dispatched again.

Not established:

  • Isolation of arbitrary code.
  • Defence against a hostile writer of the journal, Evidence store or state directory. These are trusted local storage.
  • Process exit: a closed fence proves only that no protected work can resume.
  • Reliability after power loss, or on other platforms. Testing covered macOS with Node 26.8.1, on one host and a local POSIX filesystem.
  • Cleanup of the fence, workspace and preparation a controller leaves when it dies before its claim.
  • Background recovery: reconciliation runs only inside verify, recover and, without waiting on a held fence, cancel; only verify dispatches a replacement.
  • Byte-identical state under inspect: SQLite may touch -shm and an empty -wal.
  • An atomic snapshot under inspect: each journal read sees its own committed state, so a live writer's commit can appear between the reads of one call.
  • Repository or provider delivery, release and customer value. The optional broad mutation sweep was not completed.

Composition

Depends on:

  • execution.ts (Execution.read, claim, complete, fail, cancel, interrupt and requestCancel; EXECUTION_LIMITS; its error classes).
  • attempt-fence.ts (createFence, holdFence, tryFence, FenceError, FENCE_PROTOCOL; the ledger uses parseFenceIdentity) and attempt-workspace.ts (provisionWorkspace, disposeWorkspace; the ledger uses parseWorkspaceRef).
  • local-worker.ts (LocalFixtureWorker.runGuarded, FIXTURES and the limits), fixture-child.ts (isFixtureName) and fixture-manifest.ts (guardedFixtureManifest, parseExecutionManifest, GUARDED_FIXTURE_COLLECTOR).
  • assurance.ts (gatherEvidence, judge); ctx.store is an EvidenceStore.
  • journal.ts (CommandJournal for the ledger, VersionConflictError, CommandConflictError).
  • factory-support.ts (observationProblem, unsafeCleanup, isTermination and the shared result types, which the /1 path also uses).

Used by:

  • local-factory.ts:
    • verifyFixture builds the request (buildGuardedRequest), records it, re-reads it with storedGuardedRequest, compares with guardedDifferences, submits the Run (runPinsOf) and calls verifyGuarded.
    • inspectFixture calls inspectGuarded over a read-only journal.
    • cancelFixture calls cancelGuarded, and recoverFixture calls recoverGuarded.
    • existing() routes stored records by isGuardedRecord.
  • cli.ts: main() calls those four functions. runCli() turns the first SIGINT or SIGTERM into an abort of its own AbortController.
  • Nothing else imports it; the dashboard, portfolio and repository modules do not.

Changing it safely

  • Which tests prove what:
    • local-factory.test.ts: the /2 contract through the real CLI and public API. It covers pinning, conflicts, concurrent identical requests, corrupt Evidence, signals, poisoned reports, a quarantined workspace and dependency drift.
    • local-factory-recovery.test.ts: the crash-boundary matrix, convergence, cancellation, Budgets, damaged or foreign proof, contradictory records and the ordering of aborts.
    • guarded-verification.test.ts: binding before effect.
    • attempt-ledger.test.ts: create-once records and selectors.
    • local-factory-legacy.test.ts: the /1 contract, on real /1 state written by commit f0dbdf6.
    • local-factory-inspection.test.ts and inspection-boundaries.test.ts: read-only inspection.
    • local-factory-observer-boundary.test.ts: a fault in the cancellation poll stays contained.
    • Run npm run check.
  • Source drift is deliberate. GUARDED_FIXTURE_SOURCES lists seven files: attempt-fence, attempt-workspace, evidence, fixture-child, fixture-gate, fixture-manifest and local-worker. Editing any of them changes the pinned execution digest. Existing Runs then dispatch nothing, and completed ones report invalidated. This module's own files are not in that list. Test drift only in a temporary copy of src/.
  • Stored shapes are exact. GuardedRequest, GuardedReport, Preparation and FenceProof are parsed with exact keys, so a shape change makes existing records unreadable. The precedent is /1 to /2: a new Recipe revision, with old Runs read and settled but never upgraded.
  • Receipts. Any runtime change leaves the hashes in both receipts stale. It needs independent review, a new receipt, and updates to the build review row and the "Local factory CLI" section of local development.
  • Reviewers check: invariants 1, 4, 5, 8, 9 and 10 above; that no path records a fence as a process exit; and that unresolved is never turned into failed or verified.

Source: docs/agents/guarded-verification.md