Factory docs, home
Page navigation

Occurrence identity, the discovery-window capacity ledger and deferral for the proposed discovery design (lifecycle §2, §4, §5), as one journal aggregate (discovery/schedule) driven by an injected clock. It decides when a discovery occurrence is due, admits, defers, resumes and settles it exactly once, admits each provider call before contact, and never runs one: no model, provider, sign-in, file, process or LaunchAgent is touched.

Status: Implemented as lifecycle increment 1 (spec §9, row 1) on branch factory/lifecycle-schedule, from 217887b. Tests first; an independent adversarial pass then added 12 failing tests, and all 12 defects they found are fixed (dispositions: <local evidence archive>). Independent Astra review of the fixed revision: rounds 1 to 3 FAIL (four, two and two findings, all fixed), round 4 PASS (<local evidence archive>). Every log is under <local evidence archive>. No separate reference review (docs-astra) or composition exists yet, and nothing calls it outside its tests. Every commission value is a fixture or a proposal (spec §11 defaults), not the owner's signed S1. Local development only: no scheduled run, release or customer claim, and all 27 epics remain open. Human page: guide.

  • Source: src/lifecycle-schedule.ts
  • Tests: test/lifecycle-schedule-time.test.ts (12: zone arithmetic, ISO weeks, the commission, day, week and DST boundaries, the LaunchAgent text, the module's boundary), test/lifecycle-schedule.test.ts (36: transitions, coalescing, limits, overage, the ceiling, replay, the clock, corrupt state, a three-week soak, and the rules the adversarial and review fixes added), test/lifecycle-schedule-processes.test.ts (2, real processes), test/lifecycle-schedule-adversarial.test.ts (12, written by the independent adversary)
  • Helpers: test/helpers/schedule.ts (the fake clock and harness), test/helpers/schedule-child.ts

Intelligence: none — Occurrence identity, the capacity ledger and deferral are rules over the clock and counters; a model would make scheduled runs unrepeatable.

What it hides

  • One aggregate, one decision. Every command is one CommandJournal.execute on discovery/schedule, whose pure decision re-parses the stored state (CORRUPT on any defect, never repaired), applies its rules, and emits its events. So a transition sees the lane, the occurrence, the week's ledger and the account records at one version. A method reads, plans and executes at the version it read; another writer's commit makes it re-plan, at most maxRounds (16) times (CONTENDED after that).
  • Time outside the payload. Each method reads the injected Clock when it plans, and the decision reads it again inside the journal's transaction, so no decision uses a reading taken before it waited for the lock (a plan the new reading no longer supports re-plans). Readings reach decisions as arguments, never the payload, so an identical retry at any later time has the same payload and replays. The module reads no host clock: Date.UTC and ICU's zone data through Intl.DateTimeFormat are its only calendar sources.
  • Command IDs, all journal-wide and distinct from other modules' prefixes: discovery/commission/<revision>, discovery/account/<claude|chatgpt>/<sequence>, and under each occurrence <id>/admit, /call/<k>, /call/<k>/end, /defer/<n>, /reconcile/<n>, /resume/<n> and /settle (n is the segment). Each occurrence records every transition's suffix and the version it produced; a retry of a recorded one is re-executed at that version so the journal compares it with its receipt: an identical one replays (replayed: true), any other is CONFLICT. A retry whose record has left the state (an earlier commission revision or account record, or a transition of an occurrence settled and pruned) finds that version in the journal's history, read in pages of 1 000 events up to the version just read, so it replays too; a pruned occurrence's record comes from its settlement event. A deferral's segment is the one the request names; without one, the segment an identical earlier request (same runner, reason and bound) deferred, so its retry replays even after the occurrence resumed; otherwise the running segment.
  • Occurrence IDs (spec §5): discovery/<product>/basic/<local date>, discovery/<product>/full/<ISO week> and discovery/portfolio/<audit|synthesis|bundle>/<ISO week>, in the commission's zone. portfolio is reserved and refused as a Product ID.
  • Zone arithmetic. A wall time in a gap moves forward by the gap, and one that happens twice takes its earlier instant (Temporal's "compatible" rule), assuming at most one offset change within a day either side. Europe/London never changes its clocks inside 02:00-05:00, so the tests also use Europe/Berlin (both changes fall inside it) and Asia/Tokyo (the local date runs a day ahead of the UTC one).
  • Lanes. Each lane of a revision is in the state from the revision that names it: since (when it joined; a revision that names a lane again rejoins it then), watermark (the latest occurrence admitted and its local date), lastWeek (the ISO week of the latest week-keyed occurrence the lane settled by admitting or coalescing it), open and recent. An occurrence is due only on a day at or after effectiveFrom whose window closes after the lane joined, after the watermark, and, if week-keyed, of a week after lastWeek.
  • State pruning. Each lane keeps its open occurrence and its two latest settled ones; each settlement event carries the whole record, so older ones leave the state but never the history. The ledger keeps its eight latest weeks and any week a call in flight debits. A revision that removes a Product keeps its lane's watermark and any open occurrence and drops its settled records.

Public interface

src/lifecycle-schedule.ts. Runtime imports: Buffer (node:buffer), createHash (node:crypto), canonicalJson (./canonical-json.ts), and CommandJournal and its conflict and decision errors (./journal.ts). Nothing else under src/ imports it.

  • Constants: SCHEDULE_PROTOCOL = "sf-lifecycle-schedule/1", COMMISSION_PROTOCOL = "sf-schedule-commission/1", SCHEDULE_AGGREGATE = "discovery/schedule". SCHEDULE_LIMITS (frozen): maxProducts 9, maxLanes 48, maxRevision 999 999, maxSegments 8, maxTasksPerRhythm 4, maxCallsPerTask 16, maxCallsPerOccurrence 24, maxCallBoundTokens 2 000 000, maxWeeklyTokens and maxUsageTokens 10 000 000 000, maxDeadlineMinutes 180, stagger 5-60 minutes, maxRecordAgeDays 30, maxResetHorizonMs seven days, recentPerLane 2, ledgerWeeks 8, maxRounds 16, maxStateBytes 786 432.
  • class LifecycleSchedule(journal: CommandJournal, clock: Clock):
    • recordCommission(commission) records the next revision (REVISION otherwise) and recordAccount({account, sequence, signIn, overage, recordedBy}) the owner's next record for claude or chatgpt, dated by the clock. Both return CommandResult {commandId, version, replayed, occurrence}.
    • start({lane, runner}) is one tick of a lane, not a transition request: a repeated tick acts on the state it finds, and the admission or resumption it recorded replays to the same runner while that segment runs. Its results: admitted, resumed or settled (with the occurrence, replayed and commandId), running, or waiting with until (null when no time will do).
    • admitCall({occurrenceId, runner, task, work, bound, signIn, attempt?}) returns {admitted: true, call, week, replayed, commandId, dispatch, inFlight} or a refusal {admitted: false, reason, detail} that records nothing. A request is identified by its runner, task, work, bound and attempt (default 1): an identical one replays the recorded call's receipt in any segment, after it ended or the occurrence was pruned, and a runner that re-sends refused work under the same token asks with the next attempt. Only dispatch: true, which a fresh admission alone carries, licenses one contact; a replay licenses none, even while the call is in flight (inFlight), because the journal cannot tell a call never dispatched from one whose result was lost: a call whose dispatch is uncertain is ended as lost (endCall, or reconcile), unknown and never sent again. Reasons, in the order checked: limited (this segment's closed window), call-open, deadline, call-bound, unknown-effect, duplicate-work, then the account's limited (its window closed) or sign-in-needed (a refusal since its current record), overage-unverified, budget-exhausted.
    • endCall({occurrenceId, runner, call, report}) returns {commandId, replayed, outcome, debit, occurrence}; defer({occurrenceId, runner, reason, bound?, segment?}), reconcile({occurrenceId, confirmation}) and settle({occurrenceId, settledBy, status, reason}) return a CommandResult.
    • read() (ScheduleView, or undefined before the first commission), plan() (one LanePlan per lane of the current commission: admit, resume, settle, running or wait, with its instant) and status(occurrenceId) (OccurrenceStatusView {id, state, status, coalescedInto}, where state is scheduled, due, missed, running, interrupted, deferred, settled, coalesced, pruned or not-commissioned).
  • Pure helpers: proposedCommission({products, effectiveFrom, weeklyTokens, revision?}), ceilingFromShadow(p95), parseCommission, commissionDigest, lanesOf, laneOccurrenceOn(commission, lane, date), parseOccurrenceId, zonedInstant(zone, date, time), localDateOf(zone, instant), isoWeekOf(date), launchIntervals(commission, hostZone) and renderLaunchAgentPlist(spec).
  • Types (spec §4 names): Rhythm (basic, full, audit, synthesis, bundle), OccurrenceId, Occurrence, OccurrenceStatus (completed, partial, deferred, unavailable), DeferralReason (usage-limit, budget-exhausted, sign-in-needed, volume-unmounted, logged-out), ScheduleCommission, UsageLimitObservation {status, rateLimitType, resetsAt (seconds), isUsingOverage}. The spec's CapacityLedger and LedgerEntry are LedgerWeek {week, ceiling, debited} and each occurrence's CallRecords. Also Clock, CallReport (result, usage-limit, sign-in-needed, lost), CallOutcome, InterruptionConfirmation, AccountRecord, Segment, Deferral (with the budget deferral's bound), Settlement, Coalesced, LaneState {lane, since, watermark, lastWeek, open, recent}, AccountState {account, record, version, overageObservedAt, windowClosedUntil, signInRefusedAt}, CallRefusal, SettlementReason, ScheduleError {code}.
  • ScheduleErrorCode: INVALID CORRUPT CONFLICT CONTENDED LIMIT NOT_COMMISSIONED REVISION NOT_A_LANE UNKNOWN_OCCURRENCE UNKNOWN_CALL NOT_RUNNING NOT_OWNER CALL_OPEN SETTLED STALE RULE.

Invariants and guarantees

Each item names the test that holds it ("time" is lifecycle-schedule-time, "processes" is lifecycle-schedule-processes, the rest are lifecycle-schedule).

  1. Day boundary. A basic occurrence takes its local date and is admitted only from its slot until its window closes; a late start's deadline stops at the close; a night admits nothing twice. Tests (time): "day boundary: a basic occurrence is admitted at its local slot, …" and "day boundary in Asia/Tokyo: …".
  2. Week boundary. Full, audit, synthesis and bundle occurrences take their ISO week, across the 53-week year 2026 (Thursday 31 December to Sunday 3 January are 2026-W53; Monday 4 January is 2027-W01), and the ledger opens a new week on Monday. Tests (time): "occurrence identity is stable: …", "isoWeekOf names the ISO 8601 week, …" and "week boundary: …".
  3. DST boundaries. One occurrence per local night at its local slot, across London's 23- and 25-hour days (the naive offset's instant is refused as not yet due), a Berlin slot in the gap moved forward by the gap, and Berlin's repeated hour admitting once. Tests (time): "zonedInstant resolves a local wall time …", "DST boundaries in Europe/London: …" and "DST boundaries in Europe/Berlin: …".
  4. A restart gives exactly one occurrence. A lane holds at most one open occurrence. Another runner finds it running and admits nothing; the same runner's retry replays the journal's receipt; across the next schedule boundary a running segment is never taken over by time, and it resumes as the same occurrence only after a trusted interruption confirmation (reconcile). Reconciling never skips what a deferral would enforce: a segment a usage limit or closed window had stopped resumes only at the lane's first slot after the reset, one a sign-in refusal had stopped only after a newer owner record, and none the same night after that night's deadline. Six real processes racing to start one lane admit one occurrence, and a runner killed by SIGKILL mid-call leaves one logical occurrence whose call is never re-sent. Tests: "a restart gives exactly one occurrence: …", "a running segment answers only to its runner: …"; (processes) "runners racing in separate processes …" and "a runner killed mid-call leaves one logical occurrence: …"; (adversarial) "a runner that dies after a usage limit, once reconciled, …", "a runner that dies after a 401, once reconciled, …" and "reconciling a segment past its deadline …".
  5. Coalescing. A tick admits the lane's current occurrence; every earlier occurrence of the lane that was due and never admitted, whether no runner fired or an open occurrence held the lane, coalesces into it, recorded as {count, first, last}, and each reads as coalesced, status unavailable, naming the one admitted. None is admitted later, and no backlog is replayed. Nothing is due before its lane joined the commission, so a Product a later revision adds has nothing missed before it (it reads not-commissioned). A week-keyed occurrence has one terminal record: once admitted or coalesced it is never admitted again, so a revision that moves a full night later in the same ISO week leaves that night with nothing due; one that ran is not counted in a later range (the one day a range can skip), and once pruned it reads pruned, not coalesced; one that was coalesced reads coalesced into the occurrence that holds it. Tests: "missed occurrences coalesce into the next one admitted, …", "a revision that moves a full night later in its ISO week …", "a weekly occurrence that was coalesced is never admitted later, …", and the usage-limit and segment-limit tests; (adversarial) "a revision moving a full night later …" and "a Product added by a later revision …".
  6. Usage limits (F1-F3). A call the provider refused for its usage window is usage-limit; a lost call that saw a limit event is unknown. Either stops the segment: further calls are refused limited, and the occurrence may not settle while running; it is deferred for usage-limit only, to the lane's first slot at or after the checked resetsAt (after the observation, within seven days), and resumes there as itself under resume/<n>. A performed call whose usage-window observation has any status but allowed or allowed_warning (an unrecognised one fails closed) also closes the segment until its checked reset: no further call, and a deferral must be usage-limit; its completed work may still settle. A closed window belongs to the account, not the occurrence (the owner's sessions and every Product share it): calls of any occurrence on that account are refused limited until the latest checked reset any occurrence observed on it (the account's stops only move forward, whatever order the reports land in), even after the observing occurrence settles or leaves the state, and an occurrence may defer usage-limit for it without contacting the provider, to its lane's first slot after that reset. A refusal the provider performed none of (no tokens reported used) leaves its work unsent: it is sent again under a new call and uses no call allowance. A refusal that reports tokens used went out in part: its effect is unknown, it counts against the allowance, and like any unknown work it is never sent again and keeps the occurrence from completing. The occurrence settles once. F3's other half, never switching model, provider or tier, is the Intelligence selector's and the transport's: the schedule names tasks, never models. Tests: "an injected usage limit defers the occurrence and later resumes the same one …", "a usage limit that refused the call leaves its work unsent: …", "a usage limit's reset is checked: …", "a usage-window observation closes the segment unless …", "a usage window belongs to the account: …", "a reconciled segment a usage limit closed resumes only at the lane's first slot …" and "a decision reads the clock inside the journal transaction: …"; (adversarial) "a usage-limit report that shows the provider consumed tokens …", "a sign-in refusal that shows the provider consumed tokens …" and "a call reported with a rejected usage-limit observation …".
  7. Sign-in (F5). A 401 or 403 is sign-in-needed: deferred with no resume time (or reconciled after its runner died), it waits until the owner records a newer sign-in record, then resumes at the next slot, or tonight within tonight's deadline. The record must postdate the refusal (the account's latest), not the deferral, so a sign-in repaired before the runner deferred releases it too. The refusal stops the account: every occurrence's calls on it are refused sign-in-needed until that newer record, and any of them may defer sign-in-needed for it without contact; a report that lands late never moves the stop back before a newer one. Tests: "a sign-in failure is an owner blocker: …", "a sign-in refusal belongs to the account: …", "an account stop only moves forward: …", "an owner record made between a sign-in refusal and its deferral …" and "the rhythm's deadline is per night: …".
  8. Paid overage fails closed (F4). A call is admitted only with the owner's current record of that account (at most recordMaxAgeDays old, not dated after the clock), overage off, the same sign-in the runner observes, and no overage observed since the record. Any report with isUsingOverage: true is overage, whatever else it says, an unusable reset included: it stops that account, not the other, for every Product until a newer record (the stop never moves back), and its occurrence cannot defer or complete. Its observation closes the usage window as well only when its status is closed and its reset usable; then a newer owner record lifts the overage stop but not the window. Tests: "paid overage fails closed: …", "paid overage takes precedence over its reset: …", "paid overage whose observation also closes the usage window …", "an account stop only moves forward: …" and "the owner's record is current for its maximum age and no longer, …".
  9. The ceiling refuses and is never extended. A call is admitted only if its bound fits the week's remaining allowance; a refused call records nothing, and the runner defers budget-exhausted naming the bound that did not fit (refused if it fits), to the lane's first slot of the next ISO week. A week's ceiling is set when its first call opens it; a later revision lowers it at once and never raises it. A known usage replaces a call's debit, above its bound too (flagged overBound); an unknown call stays debited at its bound. Tests: "the discovery-window ceiling refuses a call it cannot fit, …" and "a usage above its bound is debited as observed and flagged, …".
  10. Journal-backed transitions. Each transition is its own command; an identical retry replays, another request under a recorded ID is CONFLICT, revisions and records must follow in order, and replays and refusals write nothing. The command a request maps to depends on the request, not on later state: a deferral retried after the occurrence resumed, under the same runner token or another, replays its receipt; so does a call admission retried while its call is in flight (licensing no second contact), after it ended or after the occurrence resumed, and a retry of an earlier commission revision or account record, or of a transition of an occurrence settled and pruned, while anything new on a pruned occurrence is SETTLED. Tests: "each transition is one journal command: …", "a deferral is for one segment: …", "a call admission is identified by its request: …", "an admission replay licenses no contact, …" and "a retry of a transition whose occurrence has left the state …"; (adversarial) the two "an identical retry of a deferral, …" tests and "an identical retry of an earlier commission revision, …".
  11. Segments and settlement. A running segment answers only to its runner (NOT_OWNER), holds one call in flight, admits calls until its deadline (the rhythm's, capped at the window's close, and per night: a segment resumed on a night another segment of the occurrence started keeps that night's deadline, and after it the occurrence waits for the next night) and ends by defer, settle or reconcile. completed needs a running occurrence with no call whose effect is unknown (lost, overage, or a refusal reporting tokens used) or whose output was malformed, and no work left unsent; unavailable is refused once work completed. At most eight segments: the tick that would start a ninth settles the occurrence segment-limit instead. Tests: "a running segment answers only to its runner: …", "a deferral for the volume or a log-out waits for the next slot, …" and "a deferred occurrence can be settled by any named caller, …".
  12. Corrupt state. Any stored defect (protocol, keys, an open occurrence missing or settled, a segment or call out of order, a debit other than its bound or known usage, an ID that is not its lane and date, a watermark whose date is not its occurrence's or that is not the lane's latest admission, a week-keyed admission after the watermark, a lane of the current commission missing, a ledger week that debits less than its recorded calls, a deferral out of order or naming a bound other than a budget deferral's, an account observation that is not an instant, a commission digest mismatch, an unknown value) is CORRUPT on every read and command, and nothing is written over it; so is a transition the state records but the journal holds no receipt for, which is never replayed or repeated. Every decision also re-reads the state it is about to store and refuses one its reader would call corrupt. Tests: "stored state that breaks a rule is corrupt: …" and "a transition the state records but the journal holds no receipt for is corrupt, …".
  13. The clock never runs backwards in the record. A transition whose clock reads earlier than its occurrence's last recorded instant is INVALID and writes nothing; it succeeds once the clock is past it again. Test: "a clock that reads earlier than an occurrence's last transition is refused, …".
  14. Bounded state. Over three weeks of every lane of six Products across the autumn clock change (135 occurrences, each settled once), the state stays at or below 96 KiB (38 761 bytes measured); a worst-case full occurrence (eight segments, 16 calls, 48 commands, every token, key, observation and confirmation at its longest) measured 14 207 bytes (<local evidence archive>, script fix/measure.ts). Growing commands refuse LIMIT above 768 KiB, so recording commands can always settle what was admitted. A retry answered from history reads at most the history up to the version just read, in pages of 1 000. Test: "state stays bounded over three weeks …".
  15. Boundary. The module imports only the four listed; its source names no host clock, process, environment, file, child process, launchctl, network, timer or randomness, and every reading goes through the injected clock. The LaunchAgent text is returned, never written or loaded. Tests (time): "the module reads time only from the injected clock …" and "the LaunchAgent renderer is output only: …".

Failure semantics

  • Input defects, and a clock reading earlier than the occurrence's last transition, are INVALID; nothing is written. Rule failures throw their code and record nothing; admitCall's rule refusals are results, not errors. Stored defects are CORRUPT. CommandConflictError on a recorded command is CONFLICT; VersionConflictError re-plans; journal storage errors pass through.
  • A missing commission is NOT_COMMISSIONED; a lane the current commission does not name is NOT_A_LANE; an occurrence neither the state holds nor the history records as settled is UNKNOWN_OCCURRENCE for transitions (status still answers for any occurrence of the current commission); a pruned one replays recorded transitions and refuses others SETTLED. A deferral naming a segment that is not the running one, and not one it already deferred, is STALE.

Trust scope

  • Established locally (macOS arm64, Node 26.8.1 with tz data 2026a, real SQLite journals in temporary directories, a fake clock): the rules above, including a six-process race and a SIGKILL mid-call.
  • Proposals, not decisions (spec §11; the owner's S1 is unsigned): Europe/London, the 02:00-05:00 window, the 20-minute stagger, two full nights a night Thursday to Saturday, the portfolio times and §5's call bounds and deadlines (decisions 5 and 6). Choices the spec leaves open, taken here and open to review: deadlines of 30, 60 and 10 minutes for the audit, synthesis and bundle and an 08:00 close for the bundle; a "slot" is the lane's own start time each day, so a limit that resets inside tonight's window still waits for the next night's slot; the rhythm's deadline counts once per night; allowed and allowed_warning as the only open usage-window statuses (only allowed appears in the fake event; any other status closes the segment); a refusal reporting more than zero tokens as partly performed; F2, F3 and F5 read as account-wide (a closed window or a sign-in refusal stops every occurrence on that account) and an overage observation closing the window only with a usable reset; a lane's occurrence due only when its window closes after the revision naming the lane was recorded; a reused runner token naming the segment to defer a later one; tokens as the ledger's unit, with the week's ceiling set at its first call; a work key counted per task; at most two settled occurrences kept per lane; the eight-segment limit; the seven-day reset horizon. A revision that changes the zone re-plans dates from then on; recorded occurrences keep theirs. The ceiling has no default: ceilingFromShadow (p95 x 1.25, rounded up) waits for the shadow run (L8).
  • Not established: any live runner, LaunchAgent (S6, L5) or sign-in from launchd (A1); that the live limit event matches the fake one (A2), which statuses it uses, or that resetsAt is in seconds; how the runner observes a sign-in identity; that a call's bound is a true upper bound (it rests on the transport's enforced bounds, not built for these tasks); the owner's S1 record or that paid overage can be turned off (A13); authentication (records and confirmations are attribution: anyone who can write the journal can forge them); whether the Mac's zone matches the commission's (launchIntervals refuses a mismatch it is told about); power-loss durability, other hosts, and any Signal, Capture or customer value.

Composition

  • Depends on: Command journal and its canonical JSON, unchanged. Its journal is meant to be <local discovery journal> (spec A9); tests use disposable ones.
  • Used by: nothing yet. Intelligence declares it none in MODULE_INTELLIGENCE with its one file and does not import it.
  • Intended caller contract (spec §8 Runner, increments 2-8, not built): a LaunchAgent fires one Factory command at each lane time (launchIntervals); the command calls start for the lanes due; the running segment's runner holds an Attempt fence and calls admitCall before each provider contact, dispatches only on an admission with dispatch: true, reports it with endCall (as lost when it cannot tell whether it dispatched); it defers on a limit, a sign-in failure, an exhausted week, a missing volume or a log-out, and settles at the end. A supervisor that finds a runner gone confirms it with reconcile and never re-sends an unknown call.

Changing it safely

  • Run node --test test/lifecycle-schedule-time.test.ts test/lifecycle-schedule.test.ts test/lifecycle-schedule-processes.test.ts test/lifecycle-schedule-adversarial.test.ts test/intelligence-lifecycle.test.ts, then npm run check.
  • A change to a command ID, an occurrence ID, a stored shape or a rule's order changes recorded history: stored schedules must still parse, or add a versioned reader. Keep the clock out of payloads and every rule inside the decision.
  • Reviewers check: no I/O and no host clock; refusals record nothing; a replay is history, never a second effect, and the command a request maps to never depends on later state; no rule ends a running segment by time, and no path back to running (resume after a deferral or a reconciliation) skips a reset, an owner record or the night's deadline; nothing raises a week's ceiling; every list stays bounded.

Source: docs/agents/lifecycle-schedule.md