Factory docs, home
Page navigation

The deterministic collector of the proposed lifecycle (§2 discovery-sources, §4, §8 "Collector"; increment 2): it fetches a Product's commissioned public Sources into Captures, records a coverage status for every Source, and makes and re-verifies Citations. No model takes part, nothing on a page is acted on, and only hashes, counts, item keys and exact dates are kept durably.

Status: Implemented as lifecycle increment 2 on branch factory/lifecycle-sources (from 217887b), by Opus 5.5; accepted locally on independent implementation review (Astra round 5 PASS, 29 September 2026, <local evidence archive>). An independent adversarial pass wrote 17 failing tests (test/discovery-sources-adversarial.test.ts, unchanged since); all 17 defects are confirmed and fixed, with the class each belongs to closed (dispositions: <local evidence archive>). Independent review by Astra (gpt-6-astra, high effort): rounds 1 to 4 FAIL (five, three, one and one findings), all fixed (same dispositions), and round 5 PASS on the fixed revision; the highest-numbered astra-r*.md in that directory holds the current verdict. This reference has no separate reference-review report yet. Tested only against a loopback HTTP corpus the tests serve: no live web request, model call, sign-in or schedule. Nothing composes it: no schedule, signal-intake or gate calls it yet. Its receipt records the implementation, the adversarial pass, the checks and the five review rounds; the build review row records it as accepted locally and not composed. All 27 epics remain open. Owner decisions stay open (lifecycle §11): the collector follows decision 13's default (public pages only, 15-word quotes, a 35-day cache) as a proposal, not a signed commission, and takes any valid brief as confirmed, since no brief or allow-list is commissioned yet (decision 12, S5).

  • Source: src/discovery-sources.ts (brief and occurrence parsing, transports, cache, collector, records, Citations, drift) and src/discovery-text.ts (pure text operations: extraction, normalisation, SimHash, robots.txt, the quote rules)
  • Tests: test/discovery-sources.test.ts (40, one with three subtests), test/discovery-text.test.ts (19), test/discovery-sources-adversarial.test.ts (17, the adversary's); loopback corpus test/helpers/discovery-corpus.ts. Test-first: the red run, green runs and checks are logged under <local evidence archive>.
  • Design: lifecycle §2 (module row), §4 (ResearchBrief, Source, SourceStatus, Capture, CaptureNovelty, Citation as sf-citation/1, CoverageRecord), §5 (occurrence IDs and Budgets), §8 (Collector) and §9 row 2 (acceptance).

Intelligence: none — Fetching commissioned Sources, hashing, de-duplication and quote checks are exact byte operations; a model could only blur what must match.

What it hides

  • Fetching. One HTTPS GET at a time, in the brief's order, through a Transport. Every request carries exactly Host, User-Agent: SoftwareFactoryDiscovery/1, a fixed Accept, Accept-Encoding: identity and Connection: close, plus If-None-Match and If-Modified-Since when a prior fetch can answer a 304. Nothing else: no cookie (a Set-Cookie is ignored), credential, contact address, From header or query of its own. No redirect is followed. robots.txt is read first on each host, once per occurrence.
  • One Budget per occurrence. Just before its first request, an attempt claims the occurrence, create-once (aggregate discovery-sources:claim:<occurrenceId>, sf-collection-claim/1: the brief digest, the attempt's start and the deadline it sets; command discovery-sources/claim/<occurrenceId>, one discovery-sources.collection-claimed event). Any other attempt, racing in this process or another, or retrying after a failure, finds the claim: before its deadline it is refused CONFLICT without a request; after it, with no record yet, it records the occurrence partial, every due Source exhausted interrupted, no Capture, and requests and bytes null (unknown), still without a request. So the occurrence's request, byte and time Budget is spent at most once and never renewed; what an interrupted attempt fetched is unknown, never zero.
  • Text. Markup becomes one normalised form: every invisible character removed (each format character, Unicode Cf, such as zero-width characters, bidirectional controls including U+061C, the soft hyphen and Unicode tag characters, and each default-ignorable code point, such as variation selectors and Hangul fillers) with the control characters, then NFC, then single spaces; removal comes before composition, so the form is idempotent. A page therefore cannot carry an instruction or an address its reader cannot see. Scripts and styles are skipped unread; head, template, SVG, MathML, noscript, iframe and object content is not text; comments are dropped. An unclosed head ends as an HTML parser ends it, at the first start tag or text that cannot be head content, so a page that omits </head> and <body> keeps its body. A text/plain page is read whole as text, never as markup, and only by the page extractor (with any other it is refused-content content-type: it has no items, and zero would be a false exact count). Every other character a page carries is kept as text or dropped by a fixed rule. Links, forms, meta refreshes, base URLs and Link, Refresh and Content-Location headers are never followed.
  • Items. A Source's extractor splits its page: page (one item, the page's text), article (each top-level <article>) or feed-entry (each Atom <entry> or RSS <item>, whose escaped or CDATA markup is stripped once more). A feed is read as XML, with no raw text (a self-closing <script/> or <style/> is empty, and script and style content is excluded by its element): element names keep their case and every name character (a prefix such as a.b or été is whole, and A and a are distinct), and each name resolves in its namespace scope, default (xmlns) and prefixed (xmlns:<prefix>) alike, a declaration applying to its own element and what it contains. The root fixes the dialect: an Atom feed in the Atom namespace (entries and ids in it) or in no namespace (entries and ids in none), an RSS 2.0 rss in no namespace (item, guid), or an RSS 1.0 or 0.90 rdf:RDF (items in their namespace); any other element, including one of the same local name in another namespace or case, is an extension, never an entry. A feed-entry page whose root starts none of these dialects, or with any element whose prefix no declaration in scope binds, is refused-content feed-syntax, even when some entries were read: unknown, never a partial or false exact count. Each extractor reads only what it can split: article only HTML or XHTML, feed-entry only an Atom, RSS or XML media type, page any accepted type (plain text whole); any other pairing is refused-content content-type. HTML tag names, like XML ones, run to whitespace, / or >, so <article.x> is not an article. An unknown entity reference stays as written: only an own entry of the fixed table is decoded. Each item's key is its HTML id (when valid, unique and not of a reserved form), entry- and 16 hex digits of the SHA-256 of a feed entry's id or guid, or item-<position>.
  • Novelty. Each Capture is compared with the Captures fetched in the 35 days before the occurrence started, from every recorded collection of the Product that started in the 36 days before it (found through the start lists below, by when each started, so an occurrence resumed long after the day its ID names is still found), in the order they became known, and with those already made in this occurrence: unchanged (same URL, same SHA-256), repeat (the same SHA-256 at another URL, else a SimHash within Hamming distance 3 of one, with duplicateOf naming the earliest such Capture, its host, the distance and whether the host differs), changed (same URL, other hash), otherwise new.
  • Storage. Capture text lives only in a cache (below). The journal keeps one create-once aggregate per occurrence, discovery-sources:collection:<occurrenceId>, written by command discovery-sources/record/<occurrenceId> with one discovery-sources.collection-recorded event: the sf-collection/1 record of coverage, counts, Capture hashes, item keys, novelty and exact dates, never page text. Before any request it also registers the occurrence's start in the Product's list for the UTC day it started, aggregate discovery-sources:starts:<productId>:<day> (sf-collection-starts/1: the sorted occurrence IDs, at most 64, one version each), by command discovery-sources/start/<productId>/<day>/<version> with one discovery-sources.collection-started event; a racing registration of another occurrence conflicts and is re-read (at most 16 attempts), one of the same occurrence replays.
  • Validators. A response's ETag and Last-Modified are page-controlled, so the record keeps neither as text: the ETag (of RFC 9110 form, at most 128 characters) only as the SHA-256 of its value, whose value goes to the cache under that fetch like a Capture's text; Last-Modified only when it is an exact IMF-fixdate naming a real instant (Tue, 29 Sep 2026 01:00:00 GMT), a date and nothing else. A conditional GET sends If-None-Match only while that value is still in the cache.
  • Provenance. A Capture's id is a public hash of its fields, not a signature, so a Capture, Citation or later collection handed in is trusted only as far as this journal's record holds it: text, cite, verify and drift look each up in its occurrence's record.
  • Cache. <root>/<productId>/<UTC day of fetchedAt>/<sha256>.txt, mode 0600 in 0700 directories, holding Capture texts and ETag values; default root <local cache> (spec §8), configurable. Written through a temporary file and a hard link, so an existing file is never replaced; read with every byte re-hashed. Every collection first prunes it (all Products): a day directory goes once every text in it has expired, 36 days after the day began.

Public interface

Constants (src/discovery-sources.ts:70-113): kinds sf-research-brief/1, sf-capture/1, sf-citation/1, sf-collection/1, sf-drift/1, sf-collection-starts/1, sf-collection-claim/1; ROBOTS_TOKEN = "SoftwareFactoryDiscovery"; USER_AGENT = "SoftwareFactoryDiscovery/1"; DEFAULT_CACHE_ROOT; DISCOVERY_LIMITS (frozen):

BoundValueSource
maxFetchBytes2 MiB per fetch, larger refusedspec §8
maxCaptureBytes4 096 UTF-8 bytes of text per Capture, larger refusedspec §8
maxQuoteWords15, counted strictly (below); a quote is also at most MAX_QUOTE_LENGTH 1 024 UTF-16 code unitsspec §8, decision 13; the length is this increment's bound, the one a Citation record holds
cacheDays35spec §8
simhashDistance3spec A7 (a proposal)
basic40 requests, 8 MiB, 10 minutesspec §5
full150 requests, 45 minutes; 32 MiBspec §5; the 32 MiB is this increment's proposal (§5 sets no byte bound for a full occurrence)
maxSources 64, maxDomains 32, maxUrlLength 256, maxSourceIdLength 64, maxKeyLength 64, maxValidatorLength 128, maxItemsPerPage 100, maxItemsPerOccurrence 1 000, maxCapturesPerOccurrence 600this increment's proposals, so one occurrence's record always fits the journal's 1 MiB (tested with a worst case)
maxStartsPerDay 64occurrences of one Product that may start collecting on one UTC day; the next is refused INVALIDthis increment's proposal
fetchTimeoutMs 30 000 (at most maxFetchTimeoutMs 120 000)per exchange, body included, never past the occurrence's deadlinethis increment's proposal

Parsing (each throws DiscoveryError INVALID and returns a frozen copy):

  • parseResearchBrief(value) → ResearchBrief {kind, productId, revision ≥ 1, allowedDomains (1–32 lower-case DNS names of two or more labels, so never an IP address or localhost), sources (1–64 of {id, url, rhythm: "basic" | "full", extractor})}. A Source URL must be https, stated exactly as its WHATWG href, with no credentials, port or fragment, at most 256 characters, and its host must be on the allow-list exactly (a subdomain is another host). Ids and URLs are unique. briefDigest(brief) is sha256: and the hex SHA-256 of its canonical JSON.
  • parseOccurrenceId(id) → {id, productId, rhythm, period} for discovery/<productId>/basic/<YYYY-MM-DD> (a real date) or discovery/<productId>/full/<YYYY-Www> (a week its ISO year has), the schedule's grammar (spec §5).
  • parseCollection(value) and parseCitation(value): exact shapes, bounds and cross-checks (each Capture's id recomputed; Captures only for fetched or not-modified Sources, at their page's time and URL, counted exactly by coverage; status partial exactly when a Source is exhausted; requests and bytes null exactly when a Source is interrupted, and then no Capture and every due Source interrupted).

class DiscoverySources (:922), new DiscoverySources({journal, cache, transport, now?, fetchTimeoutMs?}):

  • collect(occurrenceId, brief): Promise<{collection, replayed}> (:950). Order: validate both and check the brief is the occurrence's Product's; answer a recorded occurrence from the journal (CONFLICT if it was recorded under another brief digest); cache.check(); prune the cache; register the start; load prior collections; claim the occurrence (refused CONFLICT while another attempt's claim runs, recorded as interrupted after it); then each Source in order; validate the built record with parseCollection; record it.
  • collection(occurrenceId): Collection | null (:994), CORRUPT when the stored record fails validation or is not that occurrence's.
  • text(capture): {state: "present", text} | {state: "unverifiable", reason} (:1012); INVALID for a Capture its occurrence's record does not hold field for field.
  • cite(capture, quote): {ok: true, citation} | {ok: false, refused} (:1025), refused one of malformed, too-long, personal-data, not-verbatim, inconsistent (fetched after now), not-recorded or unverifiable. Every Citation it makes, verify at the same clock verifies.
  • verify(citation): {state: "verified"} | {state: "rejected", reason} | {state: "unverifiable", reason} (:1051): rejected malformed, inconsistent, not-recorded, too-long, personal-data or not-verbatim; unverifiable expired, bytes-missing, bytes-corrupt, cache-unavailable or record-corrupt.
  • drift(citation, later): DriftObservation (:1080); INVALID for a rejected Citation, a later that is not this journal's record, or one that did not start after the Citation's fetch.

class CaptureCache (:608): new CaptureCache(root = DEFAULT_CACHE_ROOT) checks only the path (absolute, normalised, not /, no trailing slash); path(productId, fetchedAt, sha256), check(access = "write") ("read" needs only read access), put(productId, fetchedAt, text) (a file-system failure is CACHE_UNAVAILABLE), read(productId, fetchedAt, sha256) (present, missing, corrupt, or unavailable when the root cannot be read), prune(now) → {removedFiles, removedDirectories, kept}.

Transports (:404-584): interface Transport { open({url, headers, timeoutMs}) → Promise<{status, headers, body, close()}> }; TransportError {reason: "timeout" | "network" | "non-public-address" | "refused-url"}; httpsTransport() (port 443, TLS verified, minimum TLS 1.2, a fresh agent per request so no keep-alive, proxy or shared state, and a lookup that refuses a host when any address it resolves to is not public); loopbackTransport(port) (for recorded corpora and tests: plain HTTP to 127.0.0.1:<port> with the URL's host in Host, so one local server stands in for several domains); isPublicAddress(address): IPv4 outside the special-purpose blocks (this network, private, shared, loopback, link-local, IETF protocol assignments, documentation, 6to4 relay, benchmarking, multicast, reserved); IPv6 only inside global unicast 2000::/3 and outside 2001::/23, 2001:db8::/32, 2002::/16 and 3fff::/20, so every other IPv6 address (unspecified, loopback, IPv4-mapped and IPv4-compatible ::/96, NAT64, discard, unique-local, link-local, site-local fec0::/10, multicast and all reserved space) is refused.

Types: Source, ResearchBrief, Occurrence, SourceStatus, CoverageRecord {sourceId, url, extractor, status, reason, httpStatus, fetchedAt, bodySha256, bodyBytes, validators, captures, emptyItems, refusedItems, basis}, Validators {etagSha256, lastModified}, CaptureNovelty, Duplicate {id, host, distance, crossDomain}, Capture {kind, id, occurrenceId, sourceId, url, fetchedAt, sha256, simhash, bytes, novelty, previous, duplicateOf}, Collection {kind, occurrenceId, productId, rhythm, brief {revision, digest}, startedAt, finishedAt, status, requests, bytes, coverage, captures} (requests and bytes null exactly when the occurrence was interrupted), Citation {kind, captureId, occurrenceId, url, fetchedAt, sha256, quote}, CitationCheck, CiteOutcome, CaptureText, Unverifiable, DriftObservation {kind, captureId, url, occurrenceId, state, reason, current}, CollectOutcome, DiscoveryErrorCode.

Text operations (src/discovery-text.ts): extractItems(markup, extractor, xhtml = false) (the collector passes xhtml for application/xhtml+xml, where a trailing slash closes any element; in HTML it closes only void elements and <svg> and <math>, and a self-closing feed entry is an empty item), feedSyntaxProblem(markup), normaliseText, utf8Length, simhash (64-bit, over word 3-shingles after NFKC and lower-casing, so case, punctuation, spacing and markup change nothing), hammingDistance, parseRobots(text, token) and robotsAllows(policy, path), personalDataIn, quoteWords(quote), quoteProblem(text, quote, maxWords), EXTRACTORS, MAX_KEY_LENGTH, MAX_QUOTE_LENGTH.

Coverage statuses and their fixed reasons: fetched; not-modified (with the basis occurrence); not-due (a full Source in a basic occurrence); refused-size (fetch-over-limit, items-over-limit, occurrence-over-limit); refused-content (content-type, charset, content-encoding, invalid-utf-8, feed-syntax); robots-denied; robots-unknown (robots-status-<code>, robots-oversized, robots-timeout, robots-network, robots-address, robots-content-encoding); redirect-refused; denied (401, 403, 407, 451); unavailable (status-<code>, timeout, network, non-public-address, refused-url, unexpected-304); exhausted (requests, bytes, deadline, interrupted).

Invariants and guarantees

  1. Only allow-listed URLs. A brief is refused unless each Source is an exactly stated https URL on its allow-list; the collector requests those URLs and /robots.txt on their hosts, and nothing else. Test "a research brief is refused unless every Source is an https URL on its allow-list, stated exactly"; the injected-instructions and owner-email tests list every request the corpus saw.
  2. Fixed headers; no cookies, credentials or scripts. Test "the collector sends only its fixed headers: no cookies, credentials or contact data, and no retry on 401": exact header names on every request, including a later occurrence after a Set-Cookie; a 401 is denied and never retried with anything.
  3. robots.txt honoured. Read once per host per occurrence before any page there; the collector's own group wins over *, the longest match decides and Allow wins a tie (RFC 9309). Rules and paths compare in one canonical form (RFC 9309 §2.2.2): percent-encoded unreserved characters decoded, other percent-encodings in upper case, reserved characters kept apart from their encodings, anything else percent-encoded as UTF-8, and a path's literal * and $ compared as %2A and %24 (§2.2.3), so /%61dmin cannot step around Disallow: /admin. A 404 or 410 means no rules. Anything else that is not a readable 200 (5xx, 401, 403, a redirect, over 2 MiB, a timeout) is robots-unknown and nothing on that host is fetched: stricter than RFC 9309, which reads other 4xx answers as no rules. Wildcard matching cannot backtrack exponentially. Tests "robots.txt is honoured before any page request …", the robots tests in discovery-text.test.ts and the adversarial "robots.txt rules match percent-encoded paths …".
  4. Refuse, never truncate. A body over 2 MiB is refused-size fetch-over-limit, before reading when its length is declared, otherwise within one network chunk of the limit, with nothing kept. Exactly 2 MiB is read. A page of more than 100 items is refused whole; an item over 4 096 UTF-8 bytes is listed in refusedItems with its key and size, and 4 096 bytes is kept. Tests "a fetch over 2 MiB is refused …" and "Capture text over 4 KiB is refused item by item …".
  5. Only UTF-8 text of known types, with no content coding; plain text read whole; each extractor only on what it can split. Tests "only UTF-8 text/html, XML feeds and plain text with no content coding are read", "each extractor reads only what it can split: …", "a prefixed Atom feed is read like a default one, …", "feeds are read as XML in the collector: …" (a dotted and an upper-case prefix, with an unchanged entry's Citation still holding, and a foreign default namespace refused), "in a feed, read as XML, a self-closing script or style is empty: …", the feed tests in discovery-text.test.ts, "validators keep no page text: … a plain-text page is read only whole", and the adversarial "a text/plain Source is captured whole …" and "a page that omits the optional </head> and <body> tags …".
  6. Redirects never followed. Test "redirects are never followed …".
  7. Identical Captures on replay. The same corpus and clock give byte-identical collections and cached texts in two fresh stores; a recorded occurrence answers from the journal with no request, also after a restart; of collectors racing on one occurrence (two instances, a second journal connection, the same instant, another brief) exactly one claims and records it within one Budget, the others are refused before any request, and later calls answer with the record; an attempt that failed after its claim leaves the occurrence refused until the claim's deadline and then recorded as interrupted, without a request. Tests "identical Captures on replay: …", "collectors racing on one occurrence: …" and "an attempt that fails after claiming its occurrence …".
  8. Novelty and duplicates. Tests "novelty: new, unchanged, changed and repeat, with a cross-domain duplicate detected" (an exact copy on another domain, and a reformatted copy found by SimHash at distance 0), "a conditional GET that returns 304 re-emits the prior Captures unchanged from their recorded texts" and the adversarial "novelty compares with every Capture recorded in the 35 days before: an occurrence collected after its named day is not forgotten".
  9. Page text is inert data. A page's visible instructions are kept verbatim as text and change nothing: no request, allow-list, novelty or record follows from them; invisible ones (tag characters, bidirectional controls) are removed. Tests "injected instructions in page text are carried as inert data …" and the adversarial "normalised text drops invisible format characters …".
  10. No owner data leaves; no page text is kept. The collector reads no identity from its environment and sends no address; a quote carrying an address or handle (an @ before a word character, combining marks allowed between) or a phone number (nine or more digits of any script) is refused, and no invisible character can hide one; the journal holds no Capture text and no response header's text (validators above). Tests "owner-email bait: …" (the owner's address in EMAIL, GIT_AUTHOR_EMAIL, GIT_COMMITTER_EMAIL and USER_EMAIL; every byte the corpus received, and the journal files), "validators keep no page text: …" and the adversarial "page-controlled validators put no text in the journal …" and "the personal-data screen refuses an e-mail address hidden by an invisible soft hyphen".
  11. Citations. A quote is at most 15 words and 1 024 code units, normalised text, with no personal data, and occurs in the Capture's recorded text at word boundaries. Words are counted strictly (quoteWords): the larger of the space-separated parts and the runs of word characters, where only an apostrophe inside a word joins it (sync/export is two words), and each character of a script written without spaces (Han, kana, Thai and the like) is a word. cite takes only a Capture its occurrence's record holds field for field, fetched no later than now. verify rejects a forged or altered quote, a Citation whose Capture id does not bind its occurrence, URL, fetch time and hash, one fetched after now, and one whose occurrence's record holds no such Capture, so a real quote re-labelled to another URL, domain, fetch time or occurrence never verifies, whatever id it carries. Every Citation cite makes, verify verifies at the same clock. Tests "a Citation quotes at most 15 words verbatim; …", "a Capture, Citation or later collection is trusted only as this journal recorded it …", the quote tests in discovery-text.test.ts, and the adversarial Citation tests (re-labelled, re-dated, fabricated Capture, unspaced words, cite-then-verify).
  12. Unknown, never failed. From 35 days after its fetch, a Citation is unverifiable expired, even if its bytes remain; missing, altered or unreachable bytes, or a corrupt record of its occurrence, are unverifiable too, never verified. Tests "a Citation re-verifies from the recorded bytes until 35 days, then is unverifiable, never failed; …" and "a Capture, Citation or later collection is trusted only as this journal recorded it …".
  13. Drift is marked, never rewritten. drift reports holds, changed (new text still holds the quote), drifted, gone (no text of the item, which is not among the page's refused items: it is no longer on the page, or it is now empty, since empty items are counted, not keyed) or unknown (with the later Source's status; item-over-limit when the item is still there but over 4 KiB; extractor-changed when the later record split the page with another extractor, each record keeping its own) against a later collection: one this journal recorded for the same Product that started after the Citation's fetch; any other is INVALID. The Citation still verifies against its own bytes. Tests "drift: …", "drift is unknown, never gone, when the later occurrence split the page with another extractor" and the two adversarial drift tests.
  14. Budgets end an occurrence partial. The first request, byte or deadline exhaustion ends the occurrence: every later Source is exhausted, and the collection is partial. A page whose body ran the byte Budget out, or whose exchange the deadline cut off (its per-exchange timeout is the time left before the deadline), keeps its request time and bytes read and is exhausted, not unavailable; only a per-fetch timeout well before the deadline is the Source's unavailable timeout. A chunk that overruns the byte Budget is checked for that before the fetch limit, so it ends the occurrence partial even when it also makes its page oversized; a body of undeclared length can overrun the Budget by at most one network chunk (bytes already received cannot be refused), and the record states the bytes actually read. Tests "budgets: …" (three subtests), "a request the occurrence's deadline cuts off, in its headers or its body, …", "a chunk that overruns both the fetch limit and the occurrence's byte Budget …" and "a hanging server is cut off at the fetch timeout …".
  15. Every Source has a status. One coverage record per Source of the brief, in its order. Test "every Source has a status: …".
  16. Fail closed; never repair. A missing, linked or non-directory cache root, or one that cannot be pruned, is CACHE_UNAVAILABLE before any request and is never created; an existing cache file with other bytes is CACHE_CORRUPT, left as found, and nothing is recorded; a stored record, list of starts or claim that fails validation is CORRUPT for its own reads and stops any later occurrence of the Product before a request. Tests "the cache root must already exist …", "the cache never overwrites …", "a stored collection that fails validation is corrupt …", "a stored record is re-read strictly …", "each occurrence registers its start before any request; …" and "each collection first prunes expired texts; …".
  17. Prune, at every collection. collect prunes the cache before it registers or fetches anything; prune removes a day directory once every text in it has expired (day D once now reaches D + 36 days), and in it only the file names the cache writes; links are never followed and anything else is kept and reported. So a text is never read after 35 days, and stays on disk at most until its day is 36 days old or, if no occurrence runs then, until the next one. Tests "prune removes only …" and the adversarial "no Capture text outlives the cache window …".
  18. No model, provider or process. Imports are only Node built-ins for hashing, files and HTTP, the journal and canonical JSON; no dynamic import or evaluator. Test "the module reaches no model, provider or process: …".
  19. Starts are registered. Before any request, every occurrence that collects is registered in its Product's list for the UTC day it started; a day holds at most 64, and prior collections are found only through these lists. Test "each occurrence registers its start before any request; …".
  20. Public addresses only. httpsTransport refuses a host when any address it resolves to is not public (isPublicAddress above). Tests "the https transport refuses a host that resolves to a non-public address before it connects" and the adversarial "isPublicAddress refuses reserved and documentation IPv6 space …".

Failure semantics

DiscoveryError.codeRaised whenWritten?
INVALIDa malformed brief, occurrence ID, Capture, Citation or option; a brief of another Product; text of a Capture the journal did not record; drift on a rejected Citation or on a collection that is not recorded or not later; a 65th start on one dayNo
CONFLICTa recorded occurrence replayed under another brief digest, or recorded meanwhile under one; an occurrence another attempt has claimed, before that claim's deadline or under another brief; a start that could not be registered in 16 attemptsNo
CORRUPTa stored collection, list of starts or claim that fails validation, or a collection that does not read back as writtenNo
CACHE_UNAVAILABLEthe cache root is missing, a link, not a directory or not writable; it cannot be pruned; a directory cannot be createdNo; nothing fetched when found at the start
CACHE_CORRUPTa cache file already holds other bytes, or a cache directory is a linkNo record; texts already written stay
  • A Source's failure is a coverage status, never an error: the occurrence continues.
  • Unknown is not zero. Only a fetched or not-modified Source's counts are exact; under every other status (not-due, refused-size, refused-content, robots-denied, robots-unknown, redirect-refused, denied, unavailable, exhausted) the Source's content is unknown, never zero, and an interrupted occurrence's use is unknown too.
  • SQLite and the network are never one transaction. The start is registered first, then the claim, then fetches and cache writes, then the record. A failure or crash after the claim leaves a claim, perhaps texts, and no record (a registered occurrence with no record is skipped as a prior); the Budget is never spent again: until the claim's deadline another call is refused CONFLICT, and after it the occurrence is recorded as interrupted without a request. After a lost acknowledgement of the record, call collect again with the same occurrence and brief: the committed record answers.
  • Retries: none. A failed Source waits for the next occurrence. Repair: none.

Trust scope

Established locally (76 offline tests, 17 of them the adversary's, on one macOS host with Node 26.8.1, against a loopback corpus; independently reviewed, Astra round 5 PASS):

  • Every behaviour in the invariants above, over real HTTP on 127.0.0.1, real SQLite journals and real cache files in disposable directories under the worktree's ignored tmp/.

Not established:

  • The web. httpsTransport has never fetched a real page: its refusal of a non-public address is tested (for localhost), but TLS, real DNS, real sites and their terms are not. Terms of use are honoured only by the owner's commission of the allow-list (S5); code cannot read them.
  • Authority and identity. A brief is taken as confirmed; nothing checks who confirmed it (S5 is an owner signature this increment does not have). The records are unauthenticated: anyone who can write the journal or the cache can forge consistent ones, and a Capture id is a hash, not a signature.
  • Duplicate detection beyond its tested cases. A domain is a host (no public-suffix list), so two subdomains of one site count as two domains. SimHash at distance 3 on short texts catches copies that differ in case, punctuation and spacing, not reliably a reworded one; the threshold is a proposal (A7).
  • Item identity. Items without a stable id are keyed by position, so an insertion shifts later keys and shows as changed, never as lost.
  • An article page with no articles. An HTML page with no <article> records zero items for an article Source: exact for that page, though it can also mean the site changed its markup. The coverage record shows it (fetched, no Capture, no empty or refused item); the spec's rule that proposes removing a Source after 8 weeks at zero (§2) is what is to catch it, and is not built.
  • The personal-data screen. It refuses an @ before a word character and nine or more digits with phone separators; an obfuscated address (name [at] example) passes, and a long number that is not a phone number is refused.
  • Item keys are the page's own. A key is the item's HTML id when valid (at most 64 characters of letters, digits and _ . : -, starting with a letter, so never an address or a sentence), and it is kept durably in Capture URLs and refused items: the only page-controlled strings the journal holds besides hashes and dates. An id could still name a person (a username), which nothing screens.
  • Word counts for unspaced scripts. The module segments no words, so a quote in a script written without spaces is at most 15 characters: stricter than 15 words, by design.
  • Retention on disk. No text is read after 35 days, but a day directory is removed only once its day is 36 days old and only when a collection runs (or a caller prunes); between occurrences, expired texts stay on disk.
  • Scale, politeness and durability. No crawl delay or Crawl-delay; one request at a time. The cache under <local volume> may be cleared by the host's clean-up; that makes Citations unverifiable, which is the designed outcome, not a failure. No retention of journal records, no migration.
  • Composition and value. No schedule (increment 1), triage (increment 3) or gate (increment 4) uses it; the gate's "valuable" measure (Captures cited by admitted Opportunities, and proposing removal after 8 weeks at zero) needs those and is not built. No Verdict, release or customer value.

Composition

  • Depends on: src/journal.ts (Command journal): CommandJournal (execute, readAggregate), CommandConflictError, VersionConflictError, canonicalJson; src/canonical-json.ts: jsonObjectEntries, jsonArrayItems; Node built-ins node:buffer, node:crypto, node:dns, node:fs, node:http, node:https, node:net, node:path.
  • Used by: nothing yet. The designed callers are the lifecycle-schedule occurrence (increment 1), which is to pass an occurrence ID and the confirmed brief, and signal-intake (increment 3), which is to read new and changed Captures' text, turn triage quote spans into Citations with cite, and have the gate re-verify them (G2).
  • Journal file: the caller opens it; the spec's <local discovery journal> (A9) is an assumption, and no Portfolio collector exists yet.

Changing it safely

  • Run: focused node --test test/discovery-sources.test.ts test/discovery-text.test.ts test/discovery-sources-adversarial.test.ts test/intelligence.test.ts test/intelligence-lifecycle.test.ts; npx tsc; npm run check before acceptance; after a docs change rm -rf site && npm run docs and node --test test/docs-site.test.ts.
  • A record shape change is a new kind (sf-collection/2), never an edit of stored records; keep parseCollection exact.
  • Keep every page effect out: any new header, redirect, link or retry would need its own acceptance. A new request header breaks the owner-email and header tests by design.
  • The test corpus is loopback only; never point a test at the web.

Source: docs/agents/discovery-sources.md