Factory docs, home
Page navigation

A read-only watch on each Product's default branch on GitHub (the Factory bet's module M1, quick win H1). Every poll derives one state per Product (red, unknown or green) from the check runs and commit statuses on the branch head and from each workflow's runs, state and schedule. It keeps one episode per failing, disabled, overdue or waiting spell of a workflow, mirrored as one Portfolio blocker fact.

It reads through gh api --method GET only. It never writes to GitHub, never asks gh for a token, never reruns, pauses or re-enables anything, and records an observed signature for each episode, never a cause.

Status: Accepted local implementation: independent Astra reviews and combined checks passed. No live polling, scheduling or delivery. Acceptance and evidence.

  • Source: src/branch-health.ts (derivation, episodes, the journal record, the Portfolio mirror) and src/branch-health-github.ts (the gh reader, response shapes, the read plan, trigger analysis and cron schedules)
  • Tests: test/branch-health.test.ts , test/branch-health-github.test.ts , test/branch-health-replay.test.ts and the adversary's test/branch-health-adversarial.test.ts . What each covers: Tests and helpers.
  • Helpers: test/helpers/fake-gh.mjs, test/helpers/branch-health-recording.ts, test/helpers/branch-health-portfolio.ts; fixtures in test/fixtures/branch-health/ (replay.json says how the copied capture is read)
  • Design: <local evidence archive> §2 (KR2's episode, onset and closing rules), §3 (M1) and §4 (H1's acceptance)
  • Human page: guide

Intelligence: none — Health, episodes and signatures are fixed rules over what GitHub reports; a model would guess at causes where only observed signatures are allowed.

What it hides

  • Reading GitHub. ghCliReader spawns a named gh executable once per GET, with no shell, a timeout and a byte bound. Its arguments are the fixed GH_ARGUMENTS (api --method GET --include, a JSON Accept, a pinned API version, --) and the path. Its environment is explicit: ghEnvironment passes PATH, HOME, GH_CONFIG_DIR, GH_TOKEN, GITHUB_TOKEN and GH_HOST only, and the reader refuses GH_DEBUG and DEBUG. It parses the status line, headers and JSON body; a Link header only says whether a next page exists, and the next path is built here. Standard error is drained and dropped, so nothing gh prints there is kept.
  • The read plan. readProduct reads one Product in a fixed order. First come the branch head (commits/<branch>), the head's check runs (up to 3 pages) and its combined status, then the workflows (up to 2 pages). For each workflow not disabled by a person it reads the file's schedule (contents/<path>) and the runs on the branch, page after page while the triggers read so far are not answered or, the runs have not reached the persisted workflow scan boundary or any open episode's readBackTo, up to 5 pages. A listing cut short by that choice or by its bound may hide a trigger whose latest run is older than every run read, so each branch event it did not reach is then read on its own (event=, one page each). Last come the jobs (up to 3 pages each) of the red runs not yet accounted for (for an open episode, its own red runs it has not read), newest first and 20 a poll, the rest at later polls; then of up to two successful runs on the head that might close an episode, retaining candidate reads so older useful runs are eventually examined: at most 22 runs a workflow.
  • Identity and totals. A check run or combined status naming another commit than the head, a run naming another workflow (or, on an event's page, another event), jobs naming another run or commit, or a workflow file at another path make that read malformed: evidence about something else is never taken for this head or run. A listing whose pages declare a total that the items read do not match is cut short as malformed, never taken as whole.
  • Trigger analysis. Runs count only on the branch and from a branch event (push, schedule, workflow_run, workflow_dispatch, repository_dispatch, dynamic); merge-queue and release runs are left out, since they test a queue branch or a tag, not the branch. A GitHub-managed workflow (dynamic/…) is read for dynamic runs only, the one event its runs have in the recorded capture. A run is red when completed as failure, timed_out or startup_failure, green when success; anything else (skipped, cancelled, neutral, running) is no evidence either way. Each event is analysed by completion time, with attempt order distinguishing retries. New failure spells during the watch are retained, including failures followed by success between polls; spells already ended before that workflow's first scan are baseline history. A cancelled or unfinished run on the current head remains unknown until newer conclusive evidence supersedes it.
  • Check identity. The latest run of each check on the head stands. For GitHub Actions a check is a job name within one workflow: each check run's suite is traced to its workflow through the runs read, and an untraced one is a job name within its suite. A check run that names no suite is never superseded. Another app's check is its name, as GitHub reads a required check.
  • Cron schedules. schedulesOf reads cron: lines of a workflow file and refuses to guess: a file that mentions a schedule without a cron it can parse is unreadable. parseCron takes five numeric fields with ranges, lists and steps (no names or macros); nextFiring finds the next minute in UTC within 400 days, either day field matching when both are restricted, as cron does. A dynamic workflow has no file; its schedule is GitHub's and is not checked.
  • The record. One journal aggregate per Product, branch-health:product:<productId>, holds the last poll, the health derived by the last poll that changed it, the episodes (every open one and the 16 most recently closed) and, per workflow, its scan boundary and newest accounted red run. Its stored shape is sf-branch-health/2; older records are refused, not migrated implicitly. A snapshot of what a poll read is never kept; its SHA-256 over a key-sorted serialisation is.

Public interface

src/branch-health.ts:

NameWhat
BranchHealthnew BranchHealth(journal, {reader, watch, now, portfolio?}): watch lists 1–16 WatchTargets (no Product or repository twice); now gives whole milliseconds; portfolio defaults to a Portfolio over the same journal, and null turns the mirror off. poll(): Promise<PollReport>
PollReport, ProductPollReport{pollId, at, products}; each Product's outcome is recorded, unchanged, replayed, conflict, refused (the result broke a bound) or corrupt (its stored record fails validation and was left as it is), with state, the episode IDs opened and closed, a reason when not recorded, and the mirror's result (mirrored, skipped or failed with a reason)
readBranchHealth(journal, productId){version, record} or null before the first poll; re-validates the stored record (CORRUPT on any defect); reads only
openedSince(journal, productId, afterSequence, limit?)the Product's journal events after a sequence, each with its poll and the episodes it opened; what an exception batch reads
deriveProduct(snapshot, stored, pollId), flagSharedSteps(products), blockerObservation(productId, repository, episode)the pure steps of a poll: one Product's health and episodes, the shared-step flags across Products, and one episode's Portfolio observation
displayText(value, max?), distinctLabel(value, max?), runUrl(repository, runId)a name from GitHub made printable, single-spaced and bounded (an ellipsis past max, "?" when nothing is left); a label two names never share (the name itself when displayText leaves it unchanged, else its display text cut shorter, then # and 12 hex digits of the SHA-256 of the whole name), used for failing jobs' and failed steps' names; a run's page on GitHub
aggregateOf(productId), EPISODES_PER_EVENT (50), HEALTH_REASONS, BRANCH_HEALTH_PROTOCOL, BRANCH_HEALTH_COLLECTOR (branch-health)identities of the record, its events and the mirror, and the reasons in their order
BRANCH_HEALTH_LIMITS16 Products; 200 open episodes a Product (several may belong to one workflow) and 16 closed kept (more for one poll when more close at it; 216 in all); 256 workflow scans and 256 accounted red runs; 100 failing jobs an episode, each known by the SHA-256 of its whole name (the first 16 with 8 failed steps each), and 1,000 red runs kept as read or unread; 8 shared steps with 8 partners; 8 failing check names; a runaway at 6 bot-dispatched failures on one head within 24 hours; waiting over 1 hour; 1 hour's grace on a schedule; a blocker refreshed at least every 12 hours; texts of 280 and names of 120 characters; 786,432 bytes of state
TypesEpisode, RunRef, FailingJob, Onset, Runaway, SharedStep, Closure, Disposition, AccountedRed, WorkflowScan, ProductHealth, Coverage, Tally, HealthReason, Signature, NoSignalCause, Condition, ProductState, BranchHealthRecord, DerivedProduct, EpisodeOpened; BranchHealthError with code INVALID, CONFLICT, CORRUPT or LIMIT

src/branch-health-github.ts:

NameWhat
GitHubReader{get(path): Promise<ReadResult>}; ReadResult is {kind: "response", status, next, rateLimitExhausted, body} or {kind: "failed", reason}
ghCliReader({executable, argvPrefix?, env, timeoutMs?, maxResponseBytes?}), ghEnvironment(source), GH_ARGUMENTS, hasNextPage(link)the gh reader: an absolute executable, at most 120,000 ms (default 30,000) and 8 MiB (default and ceiling) a response
readProduct(reader, target, {now, open, accounted?, scans?})one Product's ProductSnapshot; open lists the stored open episodes (OpenEpisodeHint, each with its readBackTo; a hint without one is a TypeError), so their runs are read back to where the last poll left off and their closing runs' jobs are read; accounted gives each workflow's accounted red run; scans preserves history continuity even without an open episode
githubPaths, checkWatchTarget, SHA, REPOSITORYevery request path this module builds, and the target check (owner/name, a branch name, a Product ID)
triggerStates, branchRuns, latestRed, closingSince, closingCandidates, possibleEvents, uniqueRuns, workflowRunsOf, isAfter, RunKey, RunPosition, atOrBefore, reachesBack, inEpisode, redRunKey, isRed, isGreen, isBot, isRetired, newestFirst, RED_CONCLUSIONS, WAITING_STATUSES, BRANCH_EVENTSthe run classification the plan and the derivation share
parseCron, nextFiring, schedulesOfcron schedules, as above
GITHUB_READ_LIMITS, UNKNOWN_REASONS100 items a page; 5 pages of runs, 3 of check runs, 2 of workflows, 3 of a run's jobs; 20 red runs' jobs and 22 job reads in all a workflow a poll; 8 MiB a response; 30 s by default, 120 s at most; 256 KiB a workflow file; 16 crons; text of 512 characters, 32 labels, 256 steps. Reasons: unauthenticated, forbidden, rate-limited, not-found, server-error, unexpected-status, unavailable, timeout, too-large, malformed

Stored layout. Commands branch-health/poll/<productId>/<pollId> (the poll ID is the poll's start instant), one a Product a poll: a full payload of the health, every kept episode, workflow scans and the accounted red runs, or, when nothing but the time changed, an unchanged payload with the content digest. Events: branch-health.polled; branch-health.episodes-opened (each episode with its workflow, signature, onset, detection, repository, latest red run and first failed step) and branch-health.episodes-closed, 50 episodes each at most; branch-health.state-changed. Mirror: collector branch-health, stable blocker ID the episode ID, branch-health/w<workflow>/<head, 12 hex>/r<first red run>.<attempt> for a red spell or …/<cause>-<detection instant> for no signal. blockerSource maps each logical version to a Portfolio source revision of 1–1,000; later generations use an /source/<generation> record suffix while retaining the blocker ID. A refused episode mirror does not prevent unrelated mirrors, and a repeated poll retries it.

Invariants and guarantees

  1. GET only, and no token is printed. Every request is gh with GH_ARGUMENTS and a path built by githubPaths from checked parts; ghCliReader refuses any other path (an absolute URL, a leading dash, a dot segment, a path outside repos/) before spawning. Nothing asks gh for a token, standard error is never kept, and a failure is recorded as a reason code. Tests: the fake gh refuses every other method, any field or input, --paginate, auth token and --show-token. The replay's GETs are all accepted, with only allowed environment names, and a token set in the environment appears in no journal byte or report.
  2. Bounded, refused whole. Pages, bytes and time are bounded (GITHUB_READ_LIMITS); a response over its byte bound is refused whole (too-large), never cut short. Tests: the oversized, hung and malformed cases; a poll that reads over 1 MiB records only its digest, and five pages bound each listing.
  3. Unknown is never green. A request refused (401 unauthenticated, 403 forbidden, 403 with the rate limit spent or 429 rate-limited), failed, timed out or malformed is unknown, never an empty listing; a listing whose next page was not read is incomplete. A later page of runs that fails, or a trigger the runs listing did not reach whose own page could not be read, is a gap (runs-unread), whatever was read before it. A red run's jobs cut at their page bound are unknown, never "no failing job". An event's own page that was refused, failed or malformed, or that names a next page while holding no run, is a gap too, never an event without runs; so is a run listing cut short that holds no run. A page's declared total that the items read do not match, or evidence that names another commit, workflow, event or run, is malformed. Each gap becomes a HealthReason, and any reason or open episode keeps the Product from green. Tests: truncated check runs, 401 on check runs, 403 on workflows, a spent rate limit, 403 on runs and on a later page of runs, 401 and no answer on the head; the adversary's cut jobs listing; totals that do not match for check runs, workflows, runs and jobs; a check run, a status and jobs naming another commit or run; an event's own page malformed (empty or not), refused, or empty with a next page.
  4. Green only on complete evidence. Green needs a readable head carrying at least one passing check or status (neutral and skipped are no pass), nothing failed, pending, cancelled, stale or awaiting action (checks-unsettled), every listing read to its end (for a workflow's runs: every branch event's latest run read), every schedule read and run at least once, and no open episode. A head with no check (no-check-on-head) or only skipped ones (no-passing-check) is unknown. A same-named job of another workflow never hides a failed check. A red check or status on the head makes the Product red whatever else passed. Tests: the positive control; four ways to stay unknown; Family Watch red despite its passing status; the adversary's same-named checks, action_required and stale, and a trigger beyond the first page; a re-run of the same workflow's job superseding its failure, and a cancelled one or one naming no suite that does not.
  5. Signatures within one trigger. A red workflow's episode begins at the earliest spell start among its red events. Compared with that event's last green: new-head-red when the green was on another head, same-head-red on the same head, unclassed when there was never a green in a whole listing. When a red event's last green is out of reach, the signature is unclassed and the onset unknown. A signature is kept as observed at detection and is never a cause. Tests: the three signatures, a green of another event ignored, and out of reach; Pasta's Alpha within its workflow_run event.
  6. Onset and detection. A red spell's onset is its first failed run's completion (failed-run). An overdue schedule's is its expected run plus 1 hour (overdue): the next cron firing after the last scheduled run. A schedule that has never run on the branch in a whole history has no date to count from (GitHub dates the workflow, not its cron line), so it opens no episode and keeps the Product unknown (schedule-never-ran). A waiting run's onset is its creation plus 1 hour (waiting). A workflow disabled for inactivity has an unknown onset (GitHub does not date it), as has a spell reaching past the runs in reach; notAfter bounds it. Detection counts from the first poll (detectedAt) and is never earlier than a known onset: a run that failed while the poll was reading is detected at its failure.
  7. One episode per spell. A workflow may have several open spells when a later failure begins before older failing jobs have proven recovery. Each keeps its identity and observed onset. spellEnd records the first following success; it is not proof that the failing jobs passed. Failures are owned by completion chronology, including retries of older runs. An episode retains its read and unread job obligations. Accounted frontiers advance only after continuous history has been examined; they never discard another open episode. Workflow scan boundaries persist across closure, move only after a complete enough read, and move back when an older run is pending. An owner-disposed failure remains unknown if re-enabled without a newer run (reenabled-red).
  8. Closing. A red episode closes only on a successful run on the current head, after its latest red run (or a retry of that run), whose jobs were read to their end, in which every failing job it recorded succeeded and at least one job succeeded. A job is known by its whole name (the SHA-256 of the name GitHub gives, FailingJob.key), never by its label, and a name that more than one job of the run holds passes only when each of them succeeded. Its failing jobs must be complete: every red run of it (each retained red attempt belonging to its failure period) read to its end, its spell's start in reach, no more than 100 failing jobs and 1,000 red runs, and its workflow's runs read back, at this poll, to its readBackTo. That is the newest run the last poll that reached back read or, when older, the oldest run that was still running then; for a red spell whose first red run the listing did not reach (only that event's own page), that first red run. Until a poll's listing reaches it, a run between two polls may be unread: readBackTo stays, the failing jobs are incomplete, and an episode of any kind closes only if its workflow is disabled by a person or removed; once more than five pages of runs lie between, no poll ever reaches it. An unread red attempt remains an obligation even when its run is retried; its own attempt's jobs are read separately. That run's own retry closes it as passed-on-retry ("cause unresolved"); otherwise it is recovered. A no-signal episode closes once its condition ended: a disabled one on a successful run on the head after its detection; a waiting one only when nothing waits over an hour, on a successful run completed after its detection; an overdue one only on a successful scheduled run on the head after its detection. A workflow disabled by a person or deleted, one missing from a whole workflow listing, or an overdue schedule removed from its file closes as owner-disposition, apart from recovery. A green run that skipped the failing jobs, a run on an old head, or a push run during an overdue spell closes nothing. Tests: each case, and the adversary's; 17 failing jobs; a job that failed only in an earlier red run; 25 red runs read over two polls; jobs alike in their first 120 characters, in spacing or but for a tab, and one name held by two jobs; a run that failed between two polls behind 100 successful ones; a failed second page, then 520 runs past the page bound; a run still running at one poll that failed behind 100 more.
  9. Flags, never merges. Runaway: at least 6 red runs by a bot (…[bot]) on one head completing within the 24 hours before the poll, unknown when the runs in reach cover less than that and fewer are seen. A failing step shared by two open episodes, in any Products, is flagged on each with its partners; the episodes stay separate. Tests: 6, 5 and 6 by a person; both reconciles.
  10. One command a Product a poll. A poll is recorded only if later than the recorded one (else conflict), and a poll at the instant already recorded is replayed only when its repository and branch match: nothing is read or derived again, and its recorded episodes are mirrored. A changed repository or branch returns conflict, with no fetch or mirror for that Product; other Products can still proceed. A result unchanged but for the time records only that the poll happened, keeping the health of the poll that last changed it (and moving lastGreen while green). The decision re-validates the payload: new episodes start at revision 1 detected at this poll or, later, at their onset; open episodes are never dropped (every episode that closes at a poll is kept with it, however many; earlier closures fill what is left of 16), closed ones never reopen, revisions and accounted red runs never go back.
  11. The mirror. After the command, every kept episode's current blocker revision is written to the Portfolio, one episode at a time so a refusal cannot block unrelated episodes; an identical revision is a duplicate there, so repeating it every poll costs nothing and repairs a lost mirror. A revision moves when the blocker's content changes, and an open one is republished at least every 12 hours so the Portfolio never shows it stale. The fact names no one responsible (responsible: null), so it is never owner-required; its affected IDs carry the workflow, head, signature, onset and flags, and nextAction states the signature as a suspicion and links the latest red run's log with its first failed step.
  12. Stored state is checked, never repaired. Every read and decision re-validates the record; a Product whose record fails is reported corrupt and left as it is, and the others are still polled.

Failure semantics

WhereWhat happens
A readNever throws for GitHub's answer: a reason code in the snapshot (UNKNOWN_REASONS), so the Product is unknown or keeps its episodes as they were. A path that is not the module's own is a TypeError before any process starts
new BranchHealthBranchHealthError INVALID for a bad journal, reader, clock or watch list
pollINVALID for a clock outside 2000–2999. Per Product, never thrown: conflict (a racing writer, an older poll, or a changed repository or branch on same-instant replay), refused (a bound broken, such as the state size), corrupt (stored record fails validation)
The mirrorfailed with the Portfolio's code, such as NOT_FOUND for a Product not registered; the branch-health record stands. Any other error propagates after the command was recorded; the same poll again replays it and mirrors
readBranchHealth, openedSinceCORRUPT for a stored record or event without its exact shape; INVALID for a bad Product ID

Journal errors other than version and command conflicts propagate unchanged.

Trust scope

Established, on this Mac with Node 26.8.1, in tests only:

  • The replay of the recorded capture of six repositories (29 September 2026, read through a fake gh) gives the states the bet recorded:
    • Family Watch red despite a passing status, and Pasta's Alpha new-head red and runaway;
    • both reconciles same-head red with their shared step, and Tally's smoke tests no signal with the onset unknown;
    • Focus Sentinel and the Factory unknown, none green, and one blocker fact per episode.
  • A read-only reader against a fake gh that refuses anything but a GET and never prints a token; unknown for truncation, 401, 403, rate limits, timeouts, oversize and malformed answers.
  • Closing only on the failing jobs passing on the current head, matched by their whole names, with the runs read back without a gap to where the last poll left off, and with retry and owner disposition apart; one episode per spell across closures, re-enabling and a moving head; per-Product isolation of corrupt records; bounded events and mirror batches.

Not established:

  • Live GitHub. The module has never called gh against GitHub: no live read, no LaunchAgent, no schedule, and the bet's live check (the first owner-launched poll agreeing with a manual read in the same hour) has not run. How real gh prints --include output and exit codes is taken from its documented behaviour and matched by the fake, not observed.
  • The capture's gaps. It holds no workflow files, no other runs' jobs, no later pages, no event-filtered pages and no check suite IDs; the replay reads what is missing as unknown, so its cut run listings are runs-unread and no check run on its heads supersedes another. Its head SHAs and some workflow IDs are inferred, and Focus Sentinel's head is a stand-in (replay.json).
  • Runners. Local runner services are not read; a stopped runner shows only as a run waiting over an hour.
  • Detection time. Detection within 60 minutes of onset (KR2) needs polls every 30 minutes, which nothing runs.
  • Schedules. Cron lines are read line by line, not by a YAML parser; a cron inside a script block could be misread. Dynamic workflows' schedules are not checked.
  • Spells beyond reach. A red episode whose spell began before the runs in reach, or with more than 100 failing jobs or 1,000 red runs, never has complete failing jobs, so only the owner's disposition (the workflow disabled by a person or removed) closes it. The same holds for one whose red run's jobs are no longer readable, and for an episode of any kind whose workflow ran more than five pages of runs between two polls that reached back.
  • Reruns and long runs. A new attempt of a run keeps its run's place in the listing, so a rerun of a run older than the runs a poll reads is not seen; nor is a run created before the runs read when an episode was detected and still running then. Either could fail a job the episode never records.
  • Never-run workflows. A workflow with no run at all (other than a schedule that never fired) opens no episode; only its head's checks speak for it.
  • Causes. A signature is a suspicion: nothing here establishes why a build is red, and nothing restores one.
  • Scale. At most five pages of runs are read, so a long spell's start can be out of reach; the Portfolio keeps at most 2,000 source records a Product; retained closed episodes and source generations can eventually exhaust that bound. A refused mirror is reported and never cited as saved.

Composition

  • Depends on: the Command journal (execute, readAggregate, readHistory, canonical JSON); the Portfolio (Portfolio.recordObservations, PORTFOLIO_LIMITS); Node's child_process, crypto and buffer.
  • Used by: the Owner digest, which reads readBranchHealth, openedSince, HEALTH_REASONS, runUrl and the types; its blocker facts reach Required actions and blockers and the Dashboard through the Portfolio. Nothing in src/ calls poll.

Changing it safely

What must stay true

  • A change to the argument list, the environment or the path rules changes what reaches gh: every request must stay a GET with no field, body or token request. Keep the fake gh's refusals in step with any new argument.
  • The episode ID, the stored shape and the event types are what the owner digest and the Portfolio key on: a change to any of them is a protocol revision, not an edit in place.
  • Keep accounted frontiers, per-workflow scans and retained episode obligations together: no one of these proves that all earlier failures were read.
  • Workflow scans and an open episode's readBackTo keep failures between polls from going unread: it moves only at a poll whose listing reached it, and a closure needs that poll to have reached it.
  • The fixtures under test/fixtures/branch-health/ are the recorded capture byte for byte (their SHA-256 are in replay.json); never edit them, and record a new capture beside them instead.

Tests and helpers

Focused tests: node --test test/branch-health.test.ts test/branch-health-github.test.ts test/branch-health-replay.test.ts test/branch-health-adversarial.test.ts, then test/owner-digest.test.ts and test/intelligence.test.ts. Every change also needs the steps in AGENTS.md's documentation obligations.

FileCoverage
test/branch-health.test.tsHealth, failure ownership, closure, history continuity, bounds and record validation
test/branch-health-github.test.tsGET-only reader, environment and path restrictions, pagination, failure responses, schedules and triggers
test/branch-health-replay.test.tsRecorded six-Product capture, exact reader requests and Portfolio mirrors
test/branch-health-adversarial.test.tsIndependent failure cases and continuation regressions: retries, intervening spells, read gaps, cross-trigger failures, mirror rollover and digest claims
test/helpers/fake-gh.mjsRecorded GET responses; refuses other commands
test/helpers/branch-health-recording.tsImmutable capture adapter and synthetic repositories
test/helpers/branch-health-portfolio.tsDisposable six-Product Portfolio

Run the local commands. Live use remains separately commissioned.

Source: docs/agents/branch-health.md