A per-Attempt ownership gate (a one-row SQLite file held through a kernel-released read lock) plus one fresh private workspace directory bound to that gate.
Status: "Accepted for trusted local POSIX fixtures" (build review). Its evidence summary ends: "A fence proves protected work cannot resume, not process death." Receipt: continuation receipt. The SHA-256 hashes it pins for both source files, the three test files and both fence helpers match the files at 574e70c. This is local acceptance only. It is not a release or customer claim, and all 27 epics remain open. Narrative docs: local development, sections "Attempt fence" and "Guarded runs". Human page: guide.
- Source:
src/attempt-fence.ts,src/attempt-workspace.ts - Tests:
test/attempt-fence.test.ts(18),test/attempt-fence-crash.test.ts(1),test/attempt-workspace.test.ts(5). Helpers:test/helpers/fence-child.ts,test/helpers/fence-commit-crash-child.ts.
Intelligence: none — Ownership proven by kernel locks and a private directory bound to the fence; there is nothing to reason about and nothing a model could prove.
What it hides
- SQLite lock mechanics. A hold is a read-only connection that runs
BEGINand then an actualSELECT. It keeps a SHARED lock, which is an fcntl lock the kernel releases on any process exit.BEGINalone takes no lock, and WAL mode would let writers pass readers, so both are ruled out. - Closing for good.
tryFencerunsBEGIN EXCLUSIVEwith a zero busy timeout. The probe is closed at once on BUSY, so it never lingers as a PENDING writer. On success it commitsopen→fenced. Triggersfence_keptandfence_one_wayrefuse any other change. - Identity pinning. Two checks. The file check (
checkPinnedFile,:417) requires a canonical parent directory that this user owns and no other user can write, and at the pinned path a regular file with the pinned device and inode,nlink1 and this user's uid. It runs before the open, again under exclusion beforetryFencewrites, and after the read or commit. The read inside the transaction (readPinnedState,:383) runs once and requires journal modedelete, the schema compared verbatim withsqlite_schema,application_id,user_version, a cleanquick_check, and exactly one row whose protocol, Run, epoch, worker and nonce equal the pinned identity. - Workspace lifecycle. A workspace is
mkdir 0700at a nonce-derived name under a trusted root; the root is then fsynced, and the directory is pinned by device and inode. Disposal checks the pinned directory's identity and emptiness, then runs a non-recursivermdirby pathname; anything the checks catch is quarantined or reported absent. The checks and the removal are not atomic, so callers must not rename or replace the path during disposal.
Public interface
src/attempt-fence.ts imports only node: built-ins.
| Constant | Value |
|---|---|
FENCE_PROTOCOL | "sf-attempt-fence/1" |
MAX_FENCE_EPOCH | 20 (equals Execution's EXECUTION_LIMITS.maxAttempts) |
MAX_HOLD_WAIT_MS | 10_000 (the default wait is 1_000, not exported) |
WORKSPACE_PROTOCOL | "sf-attempt-workspace/1" |
Types (all readonly):
FenceOwner:{ runId: string; epoch: number; worker: string }FenceIdentity:FenceOwnerplus{ protocol: typeof FENCE_PROTOCOL; path: string; nonce: string; dev: string; ino: string }.createFenceandparseFenceIdentityreturn it frozen, in the key orderprotocol, path, runId, epoch, worker, nonce, dev, ino. It is JSON-safe, so it can be journaled.FenceHold:{ identity: FenceIdentity; active: boolean; assertActive(): void; release(): void }HoldResult:{ status: "holding"; hold: FenceHold } | { status: "fenced" } | { status: "busy" }FenceResult:{ status: "fenced"; alreadyFenced: boolean } | { status: "held" }FenceErrorCode:"invalid-input" | "exists" | "in-use" | "unavailable" | "storage" | "released", carried byclass FenceError extends Error { readonly code: FenceErrorCode }WorkspaceRef:{ protocol: typeof WORKSPACE_PROTOCOL; path; runId; epoch; worker; fenceNonce; dev; ino }.provisionWorkspaceandparseWorkspaceRefreturn it frozen; it is JSON-safe.WorkspaceDisposal:{ state: "removed"; durable: boolean } | { state: "absent" } | { state: "quarantined"; path: string; reason: string }WorkspaceErrorCode:"invalid-input" | "exists" | "unavailable" | "not-held", carried byclass WorkspaceError extends Error { readonly code: WorkspaceErrorCode }
Fence lifecycle:
createFence(path: string, owner: FenceOwner): FenceIdentity(attempt-fence.ts:179)holdFence(identity: FenceIdentity, options: { readonly busyTimeoutMs?: number } = {}): HoldResult(:231)tryFence(identity: FenceIdentity): FenceResult(:250). It never waits.
Validating stored or passed values:
parseFenceIdentity(value: unknown): FenceIdentity(:285). It accepts an exact plain object with data properties only, and returns a frozen copy.validatedHoldIdentity(hold: unknown): FenceIdentity | undefined(:280). It returns the identity that the hold's own read validated, and only for holds thatholdFencecreated in this module instance. It says nothing about liveness.parseWorkspaceRef(value: unknown): WorkspaceRef(attempt-workspace.ts:173). The path must equaljoin(dirname(path), "<runId>.<epoch>.<fenceNonce>").
Workspace:
provisionWorkspace(root: string, fence: FenceIdentity, hold: FenceHold): WorkspaceRef(:89). It creates<root>/<runId>.<epoch>.<nonce>with mode0700.checkWorkspace(ref: WorkspaceRef, fence: FenceIdentity): WorkspaceRef(:123)isWorkspaceEmpty(ref: WorkspaceRef): boolean(:135). It only parses the ref and runsreaddirSync(ref.path): no device, inode, owner, mode or symbolic-link check, so it can follow a substituted symbolic link. CallcheckWorkspace(ref, fence)first, as the worker does; the two calls are not atomic against a concurrent replacement.disposeWorkspace(ref: WorkspaceRef): WorkspaceDisposal(:145)controllerHoldProblem(fence: FenceIdentity, hold: FenceHold): string | undefined(:222). It returnsundefinedonly if this process holds exactlyfenceright now:holdmust be oneholdFencereturned in this process,JSON.stringify(validatedHoldIdentity(hold))must equalJSON.stringify(fence),hold.activemust be true, and a probeholdFence(fence, { busyTimeoutMs: 0 })must throwin-use. A probe that holds instead is released at once and reported as a problem. Because the comparison is byJSON.stringify, pass an identity in the canonical key order above, ascreateFenceorparseFenceIdentityreturns it;provisionWorkspaceand the worker parse first, but a direct caller must too. The answer holds only at the moment of the call.
Input limits:
| Input | Rule |
|---|---|
runId | ^[A-Za-z0-9][A-Za-z0-9._:-]{0,63}$ |
worker | ^[A-Za-z0-9][A-Za-z0-9._:\/-]{0,127}$ |
epoch | Safe integer from 1 to 20 |
| Fence path | Absolute and normalised, no NUL, with a basename matching ^[A-Za-z0-9][A-Za-z0-9._-]{0,120}\.sqlite$ |
nonce / fenceNonce | 32 lowercase hex characters |
dev / ino | Decimal strings of at most 20 digits, no leading zeros (^(0|[1-9][0-9]{0,19})$) |
busyTimeoutMs | Integer from 0 to 10000 |
Workspace root | Absolute and normalised, with no trailing /, no NUL, and at most 1024 characters |
| Object shape | owner has exactly runId, epoch and worker; options has at most busyTimeoutMs; identities and refs have exactly their fields. Each must be a plain object (Object.prototype or null prototype) with own data properties. Extra keys, symbol keys, arrays, class instances and accessors are invalid-input. A FenceIdentity is refused as a FenceOwner: pick out { runId, epoch, worker } first. |
Invariants and guarantees
Holding never writes. A hold lasts until
release()or process exit, and garbage collection never ends it: the module-levelconnectionsmap keeps the connection reachable. Tests: "creates a durable…", "a hold whose handle is dropped…".Many holders, no fencing. Any number of processes may hold at once.
tryFencereturnsheldwhile any holder remains, and SIGKILL of the last holder frees the fence. Test: "live readers in several processes coexist…".No lingering writer. A failed
tryFenceleaves no pending writer: a zero-wait reader joins straight afterwards. Test: "a failed fence attempt leaves no pending writer…" (25 rounds).fencedis one-way and permanent.- Later holds return
fenced, and repeatedtryFencereturnsalreadyFenced: true. - The file's own triggers refuse
DELETEand everyUPDATEother thanopen→fencedwith all other columns unchanged, even from a direct writer. The test triesUPDATE fence SET state = 'open'andDELETE FROM fenceon a fenced file. - Of six concurrent fencers, exactly one sees
alreadyFenced: false. - A holder delayed before its first read gets
fenced. - Tests: "simultaneous fencers converge…", "a holder delayed before its first read…".
- Later holds return
Commit is crash-atomic. SIGKILL before the commit leaves the fence
open; SIGKILL after it leaves itfenced. Test:attempt-fence-crash.test.ts.Identity fails closed. All of these are
unavailable, and the file is never recreated or rewritten:- a missing or replaced file (new inode), a hard or symbolic link;
- a corrupt, WAL-mode, tampered-schema, wrong-
user_versionor foreign database; - any row field that differs from the pinned identity;
- a swap between the check and the open;
- a symbolic-link or other-writable parent directory (tested for
createFence;holdFenceandtryFencerun the same check); - a file or directory owned by another user (source only: the tests run as one user, so this case is not exercised).
Tests: "invalid input is refused…", "a pinned identity that differs…", "missing, replaced, aliased…", "a fence replaced between…".
createFencenever replaces a file. It publishes a fsynced temporary database with a non-replacinglink().- Of six concurrent creators, exactly one wins; the rest get
exists. - A crash before the link publishes nothing.
- A crash after the link but before the temporary name is unlinked (the tested point,
:204–:210) leaves a file withnlink2, which every check refuses while that second name remains. A crash after a successful unlink, or a failure of the parent fsync that follows it (:210–:211), can leave a single-link fence whose identity was never returned. Other post-publicationstoragefailures (:212–:220) can leave both names, or a missing or changed target; the error does not establish the resulting file state. A surviving valid single-link fence accepts a reconstructed identity, so callers must never reconstruct or adopt one. - Tests: "simultaneous creators…", "creation faults…".
- Of six concurrent creators, exactly one wins; the rest get
One connection per fence inode per process. A second
holdFenceortryFenceon it throwsin-use, so a repeat never drops the live lock. Test: "one process cannot open a fence it holds a second time…".Holds cannot be relabelled. Holds are frozen, and
validatedHoldIdentityrejects look-alikes andObject.create(hold). Test: "a hold reports the identity its own read validated…".Provisioning needs a live hold. A workspace is created only with the caller's own live hold on exactly
fence, so a closed fence can never get one.- A foreign hold, a look-alike or a relabelled identity gives
invalid-input. - An ended hold, a closed fence or an unusable pinned file gives
not-held. - An existing directory is never adopted (
exists). - Tests: the "provisioning needs…" and "provisions one fresh…" workspace tests.
- A foreign hold, a look-alike or a relabelled identity gives
checkWorkspaceis strict. It requires the owner and nonce to match the fence, a canonical trusted parent, a directory that is not a symbolic link, the same device and inode, this user's uid, andmode & 0o077 == 0.disposeWorkspaceremoves only its own directory. It checks that the path is the exact pinned directory (device, inode, owner, mode, not a link) and empty, then removes it by pathname with a non-recursivermdir(attempt-workspace.ts:147–159). Anything the checks catch isquarantined(orabsent) and left in place, and links are never followed. The checks and the removal are not atomic: a concurrent replacement with another empty directory between them would be removed. The test covers substitutions made before disposal, not during it; callers must prevent renaming or replacement of the path while disposal runs. Test: "disposal removes only the exact empty directory…".Execution alignment.
MAX_FENCE_EPOCH === EXECUTION_LIMITS.maxAttempts, andattempt-fence.tsimports onlynode:modules. Test: "the fence module depends only on node: built-ins…".
Failure semantics
| Result or error | Meaning | What the caller does |
|---|---|---|
invalid-input (either module) | The argument was malformed. It is checked before anything is touched. | Fix the caller; do not retry. |
FenceError "exists" | A file is already at the path and was left unchanged. | Use a fresh path; never adopt the file. |
FenceError "storage" | Creation failed. If the fence was published but not confirmed, it is left in place and "must not be used". | Treat it as unknown; do not recreate at that path. |
FenceError "unavailable" | The fence proves nothing. A non-FenceError failure while withConnection opens the connection or runs the work maps here. A failure of db.close() in its finally, or of the ROLLBACK or close inside release(), propagates as a raw error instead. | Never recreate. Recovery reports it unresolved and replaces nothing. |
FenceError "in-use" | This process already has that inode open. | Do not open it again; keep the existing hold. |
FenceError "released" | assertActive() was called on a hold that has ended. | Stop all protected work. |
holdFence → busy | A fence attempt kept the file locked for the whole wait. Nothing is held. | Retry later, or treat it as not admitted. |
holdFence → fenced | The fence is closed. | The Attempt's work must not start or continue. |
tryFence → held | A holder, or another fence attempt, exists. This is not a failure: work may still be running. | Poll; dispatch no replacement meanwhile. |
WorkspaceError "not-held" | Ownership could not be established: every controllerHoldProblem answer maps here. The hold has ended, the fence is closed, the pinned file is unusable (missing, replaced or in an untrusted directory, even while hold.active is true), or the probe found this process not holding it. Nothing was created. It does not by itself prove a release, a closed fence or a process exit. | Do not provision. |
WorkspaceError "exists" | The workspace name is taken, and it is never reused. | Do not retry for this fence. |
WorkspaceError "unavailable" | The root or ref failed identity or trust checks, mkdir failed for a reason other than EEXIST, or isWorkspaceEmpty could not read the directory. | Leave everything in place. |
removed + durable: false | The directory is gone now, but the fsync of its parent failed. | Record that durability is unconfirmed. |
absent | Nothing is at the path. It says nothing about who removed it. | Do not claim a removal. |
quarantined | Left in place, with reason. | Leave it for inspection. |
Retries and idempotency:
tryFenceis idempotent (alreadyFenced).release()returns at once on a repeat call; its first call sets the hold inactive and frees the inode slot even ifROLLBACKorclosethrows a raw error, in which case treat the lock state as unknown.createFenceandprovisionWorkspaceare create-once: a replacement Attempt needs a new fence, which gets a new nonce and so a new workspace path.- In
provisionWorkspace, the root fsync andlstataftermkdir(attempt-workspace.ts:105–106) are not wrapped. If they throw, the error is a raw Node error, not aWorkspaceError, and the directory exists but no ref is returned.
Trust scope
Established (tests on macOS with Node 26.8.1):
- trusted local code on one host;
- a local POSIX filesystem with working fcntl locks;
- real multi-process holders, fencers and creators;
- SIGKILL of holders and at commit edges;
- file swaps, aliases and corruption;
- non-recursive cleanup.
Not established:
- Process exit.
fencedproves that no holder remains and none can join. It does not prove that any process exited, what the work did, or anything about processes that never held the fence. Never record it as an exit status, output orexited: true. Likewise,removedsays nothing about process exit. - Confinement. This is not a sandbox.
node:sqliteignores Node's--permissionfile flags. The identity checks catch accidents, not a writer with access to the fence directory or workspace root. - Power loss. Neither
durablenor the fsyncs are a claim about surviving power loss. - One workspace per Attempt. This is the caller's protocol, not something the module enforces. While the fence is held, a disposed path can be created again with a new inode, which the old ref never matches.
- Threads. Bookkeeping is per module instance, so use fences from one thread per process.
- Other environments. Other operating systems, network filesystems and Node versions have not been tested.
Composition
Depends on:
attempt-fence.tsusesnode:crypto,node:fs,node:path,node:sqliteandnode:urlonly.attempt-workspace.tsusesnode:fsandnode:path, and importsFENCE_PROTOCOL,FenceError,holdFence,parseFenceIdentityandvalidatedHoldIdentityfromattempt-fence.ts.
Used by:
- Guarded verification / recovery, in
src/guarded-verification.ts:dispatchOnemakesmkdir <fences>/<token>(0700, one directory per fence), then callscreateFence(<dir>/fence.sqlite),holdFenceandprovisionWorkspace. It records the preparation beforeclaim, callshold.assertActive()before the report, and releases the hold infinally. If settlement is not confirmed it releases its own hold and runsreconcileRunningitself.reconcileRunningpollstryFenceevery 50 ms (POLL_MS) up to a deadline. Onfencedit callsdisposeWorkspace, thenrecordFenceProof, theninterruptwithcleanup.confirmedBy: FENCE_PROTOCOL. Any error thrown bytryFence(:425–:428) makes the Attempt unresolved; nothing is recreated or replaced.abandon, for a preparation whose claim was refused, callstryFenceand thendisposeWorkspace, best effort.
- Repository collector (step 1, increment 3: accepted, not composed), in
src/repository-collector.ts:collectrequiresvalidatedHoldIdentity(hold)with the pinned Product as itsrunId(elseTypeError), and callshold.assertActive()before every executor process and every Evidence write, so once the hold ends nothing further is started or stored andcollectrejects withFenceErrorreleased. - Local fixture worker, in
src/local-worker.ts:runGuarded(request, guard, signal?)takesguard: { fence, hold, workspace, executionDigest }.parseGuardusesparseFenceIdentity,parseWorkspaceRefandvalidatedHoldIdentity, and throwsTypeErroron mismatch.ownershipProblem→controllerHoldProblemruns at the start, just before spawning, after exit and before returning. The worker never releases the hold.checkWorkspace,isWorkspaceEmptyandcheckDedicatedFenceDirectoryrun just before spawning. The fence's directory must hold nothing but the fence and its-journal, because the child gets a read grant on that whole directory; anything else isnot-started/setup-failed.disposeWorkspaceruns viadispositionOf, only after a process it started is confirmed gone.- The child runs with
--permissionand read grants forfixture-child.ts,fixture-gate.ts,attempt-fence.tsand the fence directory.
enter()insrc/fixture-gate.tsruns in the fixture child. It callsholdFence(identity, { busyTimeoutMs: GATE_WAIT_MS })(2000 ms) and keeps the hold until the process exits.parsePreparationinsrc/attempt-ledger.tsre-parses both values and requiresworkspace.fenceNonce === fence.nonce.src/fixture-manifest.tslists both files inGUARDED_FIXTURE_SOURCES.- Product gate, in
src/product-gate.ts, importsparseFenceIdentityand theFenceIdentitytype to validate a dispatch owner's fence identity ({runId: productId, epoch: 1, worker: delivery/<token>}). It never opens or probes a fence: a successor presents the result of its owntryFenceas proof, and the gate checks it against the current generation and fence nonce. Only tests call it.
Changing it safely
- Focused tests:
node --test test/attempt-fence.test.ts test/attempt-fence-crash.test.ts test/attempt-workspace.test.ts. - Downstream tests:
test/guarded-worker.test.tstest/guarded-worker-boundaries.test.tstest/fixture-manifest.test.tstest/attempt-ledger.test.tstest/local-factory*.test.ts: the recovery tests calltryFenceto check what recovery did to a fence, the inspection tests call it to prove thatinspectGuardednever probes or closes one (it reads stored records only,guarded-verification.ts:802), and "dependency-only source drift" editsattempt-workspace.ts.
- Every byte change is guarded-source drift. It changes
guardedFixtureManifest().digest. A completed guarded Run pinned to the old digest reportsinvalidatedon recheck while its historical Verdict is kept (guarded-verification.ts:894–907). For an unfinished Run, drift blocks further dispatch (runGuardedreturnsnot-started/drift), and a queued or running Run reportsunresolved, notinvalidated(:918–933; both outcomes are tested intest/local-factory.test.ts, "dependency-only source drift…"). That is intended; do not work around it. - The fence file format is compared verbatim (
SCHEMA,APPLICATION_ID = 0x53464631,user_version = 1), so any format change makes existing fencesunavailable. Treat such a change as a protocol revision (FENCE_PROTOCOL). - Identity and ref fields are closed.
parseFenceIdentityandparseWorkspaceRefaccept exactlyIDENTITY_FIELDSandREF_FIELDS, andparsePreparationre-parses both from the journal. Adding or renaming a field makes every stored preparation unreadable, so it is a protocol revision too (FENCE_PROTOCOL,WORKSPACE_PROTOCOL). - Keep
attempt-fence.tsfree of relative imports. The gate's read grant and its closed source list depend on it, andattempt-fence.tsmust stay the only relative import offixture-gate.ts(which also importsnode:fs);test/fixture-manifest.test.tschecks the relative imports. - Reviewer checklist:
- holds come from an actual read, never
BEGINalone; - DELETE journal mode, never WAL;
tryFencestays zero-wait and closes its probe on BUSY;- identity is re-checked under exclusion before the write;
- no recreation of a missing fence;
- no recursive delete;
- nothing opens, reads, hashes or copies a held fence file except
lstat; - no
fenced-as-exit wording anywhere.
- holds come from an actual read, never
- Re-review. Any source or test change voids the pinned hashes in the continuation receipt. It needs an independent review, a new receipt and an updated build-review row before it is claimed as accepted.
