Factory docs, home
Page navigation

Reference document, shown as written except that local paths appear as placeholders. Where it describes the Factory as intended, read it as design, not current state: only local infrastructure is accepted, no release, customer value or scheduled automation is established, and all 27 customer-value epics remain open. Current state: Factory model.

The factory and its products are maintained through the same Slice lifecycle. These are required operating behaviours; no schedules or runtime are active yet.

Ownership and cadence

A Capability is an observable ability with an owner and executable checks. Planning owns which abilities are required; Assurance owns their independent checks. Keep the baseline versioned, including infrequent recovery and fallback paths. Track required/retiring/retired disposition separately from healthy/degraded/unknown health. Low usage alone does not prove an ability is expendable.

RhythmWork and completion
Continuous healingOperations detects a failure, contains impact, restores safe required or replacement behaviour and verifies it. Delivery then fixes the cause through a Slice.
Nightly researchDelivery researches relevant tools, models and engineering methods against observed weaknesses. Planning selects bounded experiments; Assurance compares against the current baseline. Promote useful improvements incrementally through normal release gates. A night may conclude with no worthwhile change.
Weekly defragDelivery inspects dependencies, module boundaries, duplication, dead code, stale flags, prompts, documentation and abandoned work. Planning commissions small simplification Slices; Assurance proves preserved capabilities.

Research is input, never executable authority. Each domain revises its operating rules within delegation; wider authority or undelegated commitments require the owner. Keep source links, a falsifiable improvement claim, a baseline comparison and a rejection reason when useful. Measure correctness, reliability, latency, user experience, cost and complexity. A smaller diff or newer model is not evidence of improvement. Use a fixed regression suite plus fresh independent scenarios so the factory cannot optimise only for known tests.

Remove complexity without losing capability

Inventory affected capabilities and consumers before removal. Replace behaviour where necessary; remove implementation only after checks pass. Keep rollback artefacts and compatible state for the product's recovery window. Verify again after release. If required behaviour regresses, restore safe behaviour or its approved replacement, add the missing regression check and reassess the cleanup. Do not cancel a valid retirement or revive harmful or Retired behaviour. Never erase the failed check to justify removal. Correct a wrong check in a separate independently evaluated change that names the behaviour still protected. Intentional retirement is a separate Planning decision under existing authority, not disguised cleanup. Retention never overrides privacy deletion requirements.

Recovery must survive its own failure

Keep minimal Operations recovery outside the deployment being changed, protecting product availability as well as the factory. It uses pinned known-good artefacts, health checks and restricted rollback access; it does not depend on the failed planner or model provider. Test factory-down recovery, state compatibility, Action reconciliation, current authority and privacy deletion obligations before autonomous self-updates. Recovery uses independently established healthy artefacts, bounded controls and flap protection. Recover service before investigating causes. Repair the controller through a separate validated update; never replace both recovery and its target in one rollout.

Durable schedules and bounded work

Nightly and weekly Recipes use stable IDs for each scheduled occurrence, surviving restarts without duplicate Runs. Coalesce missed research/defrag windows into one current pass; never replay a backlog of speculative changes. Incidents take priority. Reserve bounded capacity for improvement without starving product delivery; when necessary, persist a wait and resume. Exact times, timezone, budgets and capacity shares remain deployment settings to select before activation.

Evidence required before activation

  • Inject a cleanup regression: detect, restore and independently re-prove the lost Capability.
  • Stop the factory during self-update: the separate recovery controller restores operation without its planner or model.
  • Restart across a schedule boundary: one logical occurrence, with no duplicate external effect.
  • Reject an attractive research change that worsens capability checks or the stated Outcome.
  • Demonstrate a weekly simplification that removes complexity while preserving the Capability baseline and the recovery path.

These extend the existing failure checklist. Research, cleanup and repair create ordinary Slices and Events; they do not introduce another service, event bus or approval chain.

Source: docs/self-improvement.md