Revision r4 · acceptance proposal, not implemented. Design owns the architecture; this page owns the first proof checklist.
This proof belongs inside E02 in the implementation guideline, alongside repair of the real pilot problem. It is not a separate platform-building phase. E01–E05 complete the first customer-value experiment.
Outcome and composition
Submit a pinned repository revision and verification Recipe once. The factory creates an isolated checkout, runs independent checks, stores Evidence and publishes the resulting Verdict as a GitHub check in a controlled repository. Interruption or duplicate requests must not lose work, invent success or blindly repeat an uncertain write.
Planning supplies the Slice contract → Execution starts a Run and bounded Attempt → Assurance checks the immutable Candidate → Evidence store preserves observations → Assurance records its Verdict → the Action gateway publishes it and confirms a Receipt. The Recipe specifies how failed/inconclusive Verdicts map to check conclusions; a completed verification Run can legitimately report failing product behaviour.
Start with known passing and failing fixtures, a real worker process and the chosen workflow engine. Use a fake external adapter for exhaustive faults, then the actual GitHub API and an existing product's check. No dashboard or release credentials are needed. An independent agent assesses requirement coverage; it cannot turn missing Evidence into a pass.
External write contract
Assurance allocates the verification identity from the Candidate, scenario revision and logical occurrence. Retries and replacement Runs reuse it. Persist the Action before dispatch; serialise ownership by Action key. GitHub's external_id is a correlation reference, not a uniqueness constraint. On confirmation, retain check_run_id and update that object.
After an uncertain creation, reconcile paginated check listings for the revision, app and verification identity, then fetch the matching result. Zero matches is not proof of absence. Multiple matches are an explicit conflict to reconcile, never silently claimed as exactly one. Leave ambiguous results unknown; use durable wakeups and continue unrelated work. Only positive evidence of non-application permits a guarded retry. Claim no exactly-once guarantee from the external API. See GitHub's check-run contract.
Failure evidence required
| Injected condition | Required observation |
|---|---|
| Duplicate submission | Same command returns its result; changed payload under its ID is rejected. |
| Two Runs request the same verification publication | Both attach to one persisted Action. Neither mints a new occurrence. |
| Worker crash or hang | Replacement resumes; stale epoch cannot submit results or dispatch Actions. |
| Old worker still owns a resource | Terminate or quarantine it before exclusive reuse. |
| Lost acknowledgement after domain commit | Retry returns the recorded command result. |
| Lost acknowledgement after GitHub creation | Reconcile the actual check; absence alone never triggers another create. |
| Known failing behaviour or missing/corrupt Evidence | Failed or inconclusive Verdict; never a fabricated pass. |
| Candidate/scenario revision changes | Prior Verdict cannot qualify the new inputs. |
| Cancellation during dispatch | Stop new work; reconcile in-flight Actions even after the Run ends. |
| Required Action stays unknown | No successful Run; durable reconciliation and honest status. |
| Authority revoked before dispatch | No external write despite earlier permission. |
| Conflicting writers or duplicate outbox delivery | Version checks and message deduplication preserve one authoritative decision. |
| Recovery ceiling exhausted | No runaway loop; persist a wait/escalation and continue eligible work. |
| Workflow lost or terminal unexpectedly | Reconcile against public Run state; never resurrect a terminal Run. |
| Projection rebuilt from history | Zero commands or external writes. |
Inject faults before and after persistence and acknowledgements. Reproduce failures with fixed seeds; add generated transition sequences for cancellation and stale Attempts. Record input revisions, runtime versions, expected/actual observations and Receipts. A fixture pass proves only that fixture.
Following slices
One repair adds a real model adapter; a stub proves outage handling before a second provider is compared. One delivery adds merge-queue Candidates, alpha/stable health gates, instrumentation and recovery. Backup restoration must be proven before production reliance. Before autonomous self-updates, prove Capability restoration and factory-down recovery. Nightly research and weekly defrag then use the same gates. Customer feedback closes the Outcome loop; a second product proves reuse and isolation. Introduce only the contracts each working slice needs.
Unsettled choices
The stack, controlled pilot repository, operating Budget and numerical recovery targets remain proposals. A crash/retry demonstration should test the proposed TypeScript/Temporal/PostgreSQL choice before broad adoption. No production reliability claim is made yet.
