Factory docs, home
Page navigation

Stores attributable, content-addressed process observations immutably, and judges them purely against a pinned expectation to issue a local-fixture Verdict.

Status: "Accepted for trusted local fixtures" (build review): immutable content identity, corruption, concurrent publication, crash durability, bounded special-file reads and strict provenance. Receipts: build receipt (independent evidence review by gpt-6-astra at xhigh; module commit d832c69; its final_sources hashes for the four files below matched the current files on 28 September 2026) and review receipt (the docs clarity revision, document assurance only and not a code review of this module; its final_sources hash for docs/design.md also matched on that date, and that glossary defines the Evidence and Verdict terms used here). Local development only: no release, production or customer-value claim, and all 27 epics remain open.

Source: src/evidence.ts, src/assurance.ts. Tests: test/evidence.test.ts (12 tests), test/assurance.test.ts (13 tests).

Intelligence: none — A Verdict is a pure fold over trusted collector observations; a model's opinion is not evidence, only a producer 'passed' flag by another name.

What it hides

  • Record format and identity. One self-contained file per observation: sf-evidence/1\n<metadata byte length>\n<canonical metadata JSON><observed bytes>. Its name is the SHA-256 hex of the whole record. Canonical JSON has sorted keys and no whitespace, so key order never changes identity.
  • Atomic, never-replacing publication. Each record is created exclusively as <root>/tmp/<id>.<uuid>.tmp with mode 0o444, written and fsynced, then hard-linked to <root>/objects/<id>, and objects/ is fsynced. The link never replaces an existing name.
  • Fail-closed reads. get re-derives everything from the bytes on disk: id format, file type, size, digest, header, length line, strict metadata and canonical re-encoding.
  • Verdict rules. A three-valued fold (Verified / Failed / Inconclusive) over per-check outcomes, with a provenance qualification that callers cannot satisfy with hand-built objects.

Public interface

src/evidence.ts (all fields readonly; unmarked fields are string):

  • MAX_EVIDENCE_BYTES = 16 * 1024 * 1024 (observed bytes) and MAX_METADATA_BYTES = 16 * 1024 (canonical metadata JSON). Unexported limits: text fields 1–256 printable ASCII characters; collectorError a well-formed Unicode string of 1–1024 UTF-16 code units (length); a stored object at most 16,793,620 bytes (MAX_RECORD_BYTES, evidence.ts:32).
  • interface EvidencePins { candidateDigest; scenarioRevision; recipeRevision; runId; attemptId; epoch: number; environment; collector }. PIN_FIELDS is the frozen tuple of those eight names in that order (evidence.ts:56).
  • interface ProcessObservation { exitCode: number | null; signal: string | null; timedOut: boolean; collectorError: string | null }. It deliberately has no "passed" flag.
  • interface EvidenceMetadata { pins: EvidencePins; check: string; collectedAt: string; observation: ProcessObservation }
  • interface EvidenceRecord { id: string; metadata: EvidenceMetadata; bytes: Uint8Array }
  • type EvidenceErrorCode = "invalid-input" | "invalid-id" | "missing" | "corrupt" | "conflict" and class EvidenceError extends Error { readonly code: EvidenceErrorCode; constructor(code, message) } with name "EvidenceError".
  • class EvidenceStore (evidence.ts:109):
    • constructor(root: string): resolves the path and does no I/O; a non-string or empty root throws invalid-input.
    • put(metadata: EvidenceMetadata, bytes: Uint8Array): Promise<string>: returns the 64-hex id.
    • get(id: string): Promise<EvidenceRecord>
  • isStoreRecord(value: unknown): value is EvidenceRecord: true only for objects returned by EvidenceStore.get in this process.
  • parsePins(value: unknown): EvidencePins and parseCheckName(value: unknown): string: validate and throw EvidenceError invalid-input; parsePins returns a frozen copy.

src/assurance.ts (same conventions):

  • MAX_REQUIRED_CHECKS = 256, MAX_EVIDENCE_REFERENCES = 1024.
  • interface Expectation { pins: EvidencePins; requiredChecks: readonly string[]; expectedExitCode: number }
  • type Gathered = { ref; record: EvidenceRecord } | { ref; fault: EvidenceErrorCode; reason: string }
  • type CheckOutcome = "passed" | "failed" | "inconclusive"; interface CheckJudgement { check; outcome: CheckOutcome; evidence: readonly string[]; reason }; interface RejectedEvidence { ref; reason }
  • interface Verdict { result: "Verified" | "Failed" | "Inconclusive"; scope: "local-fixture"; expectation: Expectation; checks: readonly CheckJudgement[]; rejected: readonly RejectedEvidence[]; reason: string }
  • gatherEvidence(store: EvidenceStore, refs: readonly string[]): Promise<readonly Gathered[]>: reads the refs sequentially through store.get (assurance.ts:66).
  • judge(expectation: Expectation, gathered: readonly Gathered[]): Verdict: pure (assurance.ts:83).

Invariants and guarantees

  1. The id is sha256(record) and the record is canonical: identical metadata and bytes give an identical id in any store, whatever the key order (test "identity is the SHA-256…").
  2. put validates and copies its inputs before its first await, so later caller mutation cannot change what is stored (test "the store copies inputs before its first await").
  3. An existing name is never replaced. EEXIST with identical bytes returns the same id; different or unreadable existing bytes throw conflict and are left untouched (tests "duplicate and concurrent…", "conflicting contents…").
  4. Identical puts, concurrent or repeated, converge on one object and, in the tests, leave tmp/ empty. Temp-file cleanup is best effort: put tries to unlink its own temp file and ignores any error from doing so.
  5. An id acknowledged by put survives the writer being killed with SIGKILL. Readers never consult tmp/, and leftovers there are never readable as objects (tests "a process killed mid-write…", "missing evidence and crash leftovers under tmp…").
  6. get checks the id against /^[0-9a-f]{64}$/ before any filesystem access, so no path traversal is possible (test "ids that are not plain lowercase…").
  7. get opens with O_RDONLY | O_NOFOLLOW | O_NONBLOCK and requires a regular file of at most MAX_RECORD_BYTES. A symbolic link, directory or FIFO is corrupt; the FIFO case is tested to fail promptly rather than block the open (test "a FIFO in place of an object…").
  8. get fails closed on a digest mismatch, a wrong header, a missing, malformed or leading-zero length line, truncated metadata, invalid metadata, oversized bytes, a non-canonical encoding, or growth after stat. A valid record placed under another record's name is corrupt (tests "corrupt, truncated…", "content matching its name but not in canonical record form…").
  9. Returned records, their metadata, pins and observation are frozen. bytes is a fresh copy on every get, so mutating it never affects the store.
  10. Metadata is strict. metadata, pins and observation must be plain objects that own every listed field and have no other enumerable string-keyed field, so an extra field such as passed is rejected; non-enumerable and symbol-keyed properties are ignored. Text is 1–256 printable ASCII characters with no surrounding spaces. candidateDigest is sha256: plus 64 lowercase hex characters; epoch is a non-negative safe integer; collectedAt must match YYYY-MM-DDTHH:MM:SS.mmmZ and round-trip exactly through toISOString; exitCode is null or a safe integer; signal is null or /^SIG[A-Z0-9]{1,16}$/; timedOut is a boolean; collectorError is null or a well-formed string of 1–1024 UTF-16 code units. Invalid input writes nothing: objects/ is not even created (test "oversized, malformed or self-asserted metadata…").
  11. judge is pure: repeated calls on the same inputs are deep-equal. Its only process-local dependency is the isStoreRecord WeakSet.
  12. A gathered entry qualifies only if it is an object without an own fault property whose record passes isStoreRecord, has record.id === ref, matches all eight PIN_FIELDS exactly (!==) and names a required check. Anything else goes to rejected with a reason (an entry whose ref is not a string is reported as <invalid reference>). Hand-built, spread-copied, relabelled or JSON round-tripped records therefore never qualify (test "hand-built or relabelled records are not proof").
  13. Each qualifying observation is classified in this order: collectorError set → inconclusive; timedOut → inconclusive; any signal → inconclusive, because a signal may come from the environment; exitCode null → inconclusive; exitCode !== expectedExitCode → failed; otherwise passed. A check with no qualifying observation is inconclusive with reason "no qualifying evidence". With several observations of one check, failed wins over inconclusive, which wins over passed, and the reason is prefixed with the decisive ref; a repeated ref counts once.
  14. The Verdict is Failed if any required check failed, even when others are missing or rejected; otherwise Inconclusive if any check is inconclusive or any supplied reference was rejected; otherwise Verified. Verified therefore needs every supplied reference to qualify.
  15. The Verdict and all its parts are frozen, and its expectation is the validated, frozen copy. checks follows requiredChecks order and rejected follows gathered order. reason strings are deterministic.

Failure semantics

  • EvidenceError codes:
    • invalid-input: an invalid root, metadata or bytes, bytes over MAX_EVIDENCE_BYTES, or canonical metadata over MAX_METADATA_BYTES.
    • invalid-id: a malformed id passed to get.
    • missing: ENOENT.
    • corrupt: any integrity or format fault, including a symbolic link (ELOOP) or a non-regular file.
    • conflict: the name exists with different or unreadable contents.
  • Other filesystem errors, such as EACCES, ENOSPC and EIO, propagate raw from put and get and are not EvidenceErrors, with two exceptions: an existing publication target that cannot be read becomes conflict, and any error from unlinking the temp file is ignored.
  • put is idempotent and safe to retry after any failure, including a lost acknowledgement: a retry re-links or finds the identical object and fsyncs objects/ again. Directory preparation is memoised and reset on failure.
  • gatherEvidence turns only EvidenceErrors into { ref, fault, reason } entries (a non-string ref becomes fault: "invalid-id" with ref: String(ref)); these are inputs to judge, not errors. Any other error rejects the whole call. A non-array or more than 1024 refs rejects with TypeError.
  • judge throws TypeError "invalid expectation: …" for an invalid expectation. The expectation must be a non-null object whose sorted own enumerable string keys, joined with commas, equal expectedExitCode,pins,requiredChecks (assurance.ts:138); this comparison does not enforce an exact key set, since comma-containing keys can collide and the three field values can be inherited. The values are then validated: 1–256 unique valid check names with no array holes, a safe-integer exit code and valid pins. judge also throws a TypeError when gathered is not an array or has more than 1024 entries. Ordinary evidence problems never throw: they land in rejected, and the Verdict is then at best Inconclusive. judge is not an exception-isolating boundary for arbitrary objects, though: a caller-supplied getter, proxy or toString that throws propagates its own error.
  • Unknown is never a pass and never a known failure. A timeout, signal or collector error makes that observation inconclusive; each check then folds its qualifying observations with failed > inconclusive > passed, so a timed-out rerun beside a failed one still gives failed. A missing or unreadable reference is rejected and touches no check's outcome. Any failed required check makes the Verdict Failed; otherwise an inconclusive check or a rejected reference makes it Inconclusive.

Trust scope

Established, for trusted local fixtures on a POSIX local filesystem with hard links; tested on macOS with Node 26.8.1:

  • detection of corruption, truncation, misplaced objects, special files and overwrite attempts;
  • never-replacing publication that converges under concurrent writers;
  • acknowledged objects that survive a process kill;
  • strict provenance pins;
  • Verdicts that producer flags and hand-built records cannot forge.

Not established:

  • It is not a sandbox, an access-control boundary or a cryptographic authority. Anything that can write the store directory can write well-formed evidence, and collector attributes an observation without authenticating it (evidence.ts:3).
  • Power-loss durability. Process-crash tests do not establish it.
  • Completeness. judge cannot detect deliberately omitted references; the composition must own the full reference set and persist the Verdict it issues (local development).
  • Anything beyond scope: "local-fixture". A Verdict is never a production, release, public-assurance or customer-value claim, and it is only as trustworthy as the collector and the store.
  • Time. collectedAt is caller-supplied and is not checked against a clock.
  • Housekeeping. There is no delete, retention or garbage-collection API. Each put only tries to unlink its own temp file; crash leftovers under tmp/ stay until removed by hand, which is safe only while no put is running.

Composition

  • Depends on: evidence.ts uses only node:crypto, node:fs, node:fs/promises and node:path. assurance.ts imports only evidence.ts.
  • src/local-worker.ts: LocalFixtureWorker requires an EvidenceStore instance. run() and runGuarded() call store.put(...) once for the process they observed: an observed result references that one record, and a failed put gives storage-failed in run(). runGuarded() skips storage when ownership ended after the process exited, and re-checks ownership after the put succeeds or fails: lost ownership wins and returns ownership-lost, with evidenceId set only if the record was stored; storage-failed is returned only while ownership still holds. parseRequest uses parsePins and parseCheckName, requires pins.collector to equal LOCAL_FIXTURE_COLLECTOR in run() or GUARDED_FIXTURE_COLLECTOR in runGuarded(), and rewraps EvidenceError as TypeError. The worker observes and never judges.
  • src/factory-support.ts: observationProblem(store, expectation, maxOutputBytes, result) requires exactly one evidence reference, re-gathers and re-judges it, compares the canonical JSON of the new Verdict with the stored one, and checks the termination fields against the stored observation.
  • src/guarded-verification.ts: expectationFor(...) builds one-check expectations (requiredChecks: [recipe.check], expectedExitCode: recipe.expectedExitCode); verdictOf(...) judges only when the source digests before and after the run equal the pinned digest, and otherwise returns null; recheck(...) re-judges attempt.evidence for current validity; judgeReport(...) calls observationProblem; dispatchOne(...) passes ctx.store to the worker.
  • src/local-factory.ts: withState(...) constructs new EvidenceStore(<state-dir>/evidence); legacyRecheck(...) re-judges and compares canonically; legacyReportProblem(...) calls observationProblem.
  • src/repository-assurance.ts (Repository Assurance): imports PIN_FIELDS, isStoreRecord and the EvidencePins, EvidenceRecord and ProcessObservation types from evidence.ts, and only types from assurance.ts. It never calls judge, and neither judge accepts the other's records. The Product gate and delivery records import assurance.ts's Gathered type only.
  • src/repository-collector.ts (Repository collector, step 1, increment 3: accepted, not composed): requires an EvidenceStore instance and calls store.put(...) once per repository check with sf-repository-observation/1 bytes; a failed put is storage-failed and stops the collection. It never calls judge or get.
  • Tests elsewhere: test/helpers/factory-harness.ts monkey-patches EvidenceStore.prototype.put for fault injection; test/confined-executor-boundaries.test.ts uses a real store as a protected target.

Changing it safely

  • Run node --test test/evidence.test.ts test/assurance.test.ts and npm run typecheck, then the dependent suites: test/local-worker*.test.ts, test/guarded-*.test.ts, test/local-factory*.test.ts, test/fixture-manifest.test.ts and test/confined-executor-boundaries.test.ts. npm run check runs everything.
  • Editing evidence.ts changes the guarded fixture manifest digest, because it is listed in GUARDED_FIXTURE_SOURCES in src/fixture-manifest.ts; Runs pinned to the old digest report source drift. assurance.ts is not in that list.
  • Record format or canonicalisation changes break existing records. An id is the hash of the bytes as stored, so existing records keep their ids, but get re-encodes and compares, so it rejects them as corrupt unless the sf-evidence/1 decoding and canonical validation are kept exactly; the same content would also get a different id under the new encoding. Add a new format version rather than changing this one. Never relax a corrupt check.
  • Verdict shape, ordering and reason wording are compared canonically against stored Verdicts by observationProblem, recheck and legacyRecheck. Changing them makes earlier issued Verdicts fail re-validation.
  • Keep put's signature: the factory harness patches it.
  • Reviewers check that no path lets a non-store record, a passed flag, a pin mismatch, a signal, a timeout or a missing check produce Verified; that failures stay Failed and unknowns stay Inconclusive; and that validation still happens before the first await.
  • After a change, get a fresh independent review, issue a new successor receipt (the build receipt and its hashes stay unchanged), and update the build review row. The rows in module studies and local development must stay true.

Source: docs/agents/evidence-assurance.md