No study yet: this module has no plate position.
What it does for you: It is the part of the proposed discovery and post-development design that reads the public pages you commission for a project: its forums, reviews, changelogs and competitors' pages. It fetches only the addresses on your confirmed list, asks each site's robots.txt first, and keeps each post or entry as a short piece of text with a fingerprint, so it can tell what is new, what changed, what repeats and what is copied from another site. It lets a quote of at most 15 words be cited only when it appears word for word in what was fetched, and checks it again later. No AI model takes part, and nothing a page says is ever acted on.
How it links: It records what it fetched in the Command journal: fingerprints and counts only, never the text, which stays in a separate cache for 35 days. The design's next parts will use it: the nightly and weekly schedule will run it, and the sorting step will read its new and changed texts and turn the quotes it finds into citations. None of them is built yet, so nothing uses it.
Honest limits:
- It has only been tried against a small test web server on this computer: it has never fetched a real page. Your confirmed brief and list of sites (commission S5) do not exist yet, and nothing schedules it.
- It refuses rather than cuts short: a page over 2 MiB, a page of more than 100 items or an item over 4 KiB of text is refused and counted, never trimmed. When a night's request, size or time allowance runs out, the rest of that night's pages are marked as not reached: unknown, never zero. Each night's allowance is spent once: if a run stops part-way, another run of the same night is turned away until that night's time is up, and the night is then recorded as unknown without fetching again.
- It sends nothing about you: no cookies, sign-in, e-mail address or contact details, whatever a page asks for. Quotes that contain an e-mail address, a handle or a phone number are refused, and it strips the invisible characters a page could use to hide one, or to hide instructions from you. It cannot read a site's terms of use; your approval of each site is what covers those.
- It keeps no page text in its records: only fingerprints, counts, dates and each post's own label from the page (its HTML id), which could occasionally be someone's user name. A quote can only be cited from a page it actually recorded, so a quote moved to another site or date is rejected.
- After 35 days a quote can no longer be checked against the text it came from, so it counts for nothing: unknown, not failed. If the cache is cleared sooner, the same happens sooner. Each run clears expired text first, so old text stays on disk only until the next run after its 36th day.
- A page that is not what its source should be (a feed that is not a feed, or a web page where a feed should be) counts as unknown, never as empty.
- It counts quote words strictly: in languages written without spaces, such as Japanese or Thai, each character counts as a word, so such a quote is at most 15 characters.
- It spots copies that differ only in capitals, punctuation or spacing, not reliably a reworded one, and it treats two parts of one website (such as a forum and its blog on separate addresses) as two sites. Its limits and thresholds are proposals, not settings you have signed.
Status: Accepted as lifecycle increment 2 on independent review, 29 September 2026: an independent adversarial pass found 17 faults and four rounds of independent review found ten more, all fixed, and the fixed code passed the fifth round (build review). Nothing uses it yet, and it has never fetched a real page. This is local development only: not a release or a customer benefit, and all 27 epics remain open.
