# pgpipe Repair Approval Operator Runbook

This runbook is the production operating procedure for the proof-bearing,
source-authoritative repair workflow. It covers rollout, approval, monitoring,
failure handling, restart, and recovery. For general security boundaries, see
[SECURITY.md](SECURITY.md).
For incident diagnosis, see
[TROUBLESHOOTING.md](TROUBLESHOOTING.md).

## Safety contract

A repair means exactly this:

> For the approved primary-key rows or bounded primary-key range, make the
> destination match the current transformed source.

The source always wins. pgpipe never copies a destination value back to the
source and never asks the operator to choose which side wins. Inserts add rows
missing from the destination, updates overwrite differing destination rows,
and deletes remove destination rows that are not present in the approved source
selection.

The workflow is intentionally fail-closed:

- `repair.execution_enabled` defaults to `false`.
- Verification remains read-only. A preview also performs no destination
  writes.
- Authentication and CSRF origin protection are required for every v2 repair
  mutation. Actor identity comes only from the authenticated session; do not
  send `requested_by`, `approved_by`, or another actor field.
- Every approval is bound to the exact proof-bearing preview, current mapping,
  source and destination relation identities, one-time token, server-generated
  confirmation, reason, and required acknowledgements.
- Approval returns `202 Accepted`. A pipeline-owned worker executes
  asynchronously; closing the browser does not cancel it.
- Cancellation is allowed only before a worker claim. It does not interrupt or
  roll back active work.
- A partial or ambiguous outcome is never retried automatically. It becomes an
  offline recovery event.

Do not edit pgpipe state, repair fences, or chunk-ledger rows to bypass a
refusal. The clean v6 boundary removes the entire old unversioned repair route
family; clients must use `/api/v2/repair`.

## Capability and rollout state

The first production approval-workflow release must keep execution off by
default:

```yaml
config_version: 6

repair:
  execution_enabled: false
  approved_queue_lifetime: 15m
  max_execution_duration: 30m
  connection_acquire_timeout: 10s
  statement_timeout: 30s
  lock_timeout: 5s
  recovery_cleanup_timeout: 30s
```

Released `v5.2.0` does not contain the guided repair provisioning command or
the dashboard activation workflow described below. They are current
development-tree behavior for a later release and must not be advertised as a
`v5.2.0` capability.

Phase 7 documentation is present in the development tree, but release
certification is not complete. The unchanged candidate still needs its clean
installation, complete browser workflow, destructive soak/recovery,
normal-replication performance, vulnerability, and final human-review evidence.
A skipped, noisy, stale, expiring, or different-commit result is not approval to
enable production repair. The canonical pending/pass register is in the
[Greenfield Destination Safety Contract](DESTINATION_SAFETY_CONTRACT.md#15-phase-7-documentation-and-release-certification),
and executable release gates are described in [CI.md](CI.md).

This is an API/backend capability, not just a hidden dashboard button. When it
is false, token issuance, approval, new worker claims, and destination repair
writes remain disabled. Verification, mismatch reporting, proof previews, and
ordinary replication continue. The always-on read-only destination guard check
still blocks startup if an older unresolved repair fence exists; disabling the
feature is never a recovery bypass.

Check the live process, not only the YAML file:

```bash
curl --disable --fail --silent --show-error \
  --cacert /etc/pgpipe-tls/ca.crt \
  https://pgpipe.example.com/api/v2/repair/capabilities
```

The bounded response reports `execution_enabled`, `approval_available`, and
`readiness`. Approval is safe to expose only when all three conditions are
`true`, `true`, and `ready`. Other readiness values are `disabled`, `starting`,
`setup_required` (server authentication/CSRF or required state, source, and
destination setup is incomplete), `not_streaming`, `reconciling`, and
`recovery_required`; use the dashboard and sanitized service logs for the next
step rather than guessing from the label.

### Install, prepare, enable, and operate

The development tree contains the protected-storage kernel, runtime integration,
and the native `pgpipe init` bootstrap UI. There is no supported repair-specific setup
command or legacy migration workaround in the normal fresh-install flow.

The greenfield operator flow is:

1. Run the token-gated `pgpipe init` browser wizard.
2. On the final **One-time destination preparation** page, review the locked
   destination route/TLS/runtime identity and enter a temporary destination
   administrator username and password.
3. Choose **Prepare installation**. Initialization returns an asynchronous operation,
   proves both connections reach the same database, creates the exact protected
   manifest, validates the runtime role, discards the administrator credential,
   and publishes repair disabled. Closing the browser does not cancel the
   server-owned operation.
4. Start pgpipe and confirm ordinary replication is healthy.
5. Optionally run `pgpipe repair check -c <config>` with the runtime identity;
   it is database-enforced read-only and never admits a repair.
6. After release gates pass, an authenticated dashboard administrator enables
   repair, reviews the locked limits, saves, and performs the required restart.
7. Confirm **Repair enabled — ready** before materializing or approving work.
8. Review the exact proof-bearing action plan, enter a non-sensitive reason and
   acknowledgements, then approve.
9. Monitor fenced asynchronous execution and require final verification to
   report `matched`.
10. Retain preview, approval, execution, verification, and actor audit history.

Installation and activation remain separate: successful initialization always
publishes `repair.execution_enabled: false`. Day-to-day enable/disable and
repair operation never receive the DBA credential.

Do not start pgpipe unless the completion page reports all three assurances:

- **Repair safety: Prepared**
- **Repair execution: Disabled**
- **Administrator credential: Not persisted**

An authentication/authorization failure clears the password and allows retry
against the retained reviewed candidate. If destination preparation committed
but local file publication failed, run the same `pgpipe init` command again; it
validates and resumes the exact journaled attempt rather than repeating DDL.
Every other unsuccessful attempt publishes neither canonical YAML nor a start
instruction.

### Greenfield configuration and restart behavior

This safety redesign supports only a fresh pgpipe control-plane installation.
It does not migrate an older safety schema, owner, grant set, or repair state.
Existing development/demo objects are unsupported and disposable.

All `repair.*` values are restart-required. A dashboard save changes desired
configuration only; `SIGHUP` does not activate them. Before enabling repair,
confirm that initialization completed, ordinary replication is healthy, and
the read-only runtime check passes. After the controlled restart, require the
capability response to report `execution_enabled: true`,
`approval_available: true`, and `readiness: ready`.

To disable new execution, first ensure there is no queued or active repair,
set `execution_enabled: false`, validate, and perform a controlled restart.
Never edit or discard an unresolved claim, cleanup state, or destination
fence: startup reconciliation may correctly keep streaming blocked even while
new repair execution is disabled.

## Pre-enablement checklist

Complete and record every item for the pipeline:

Complete release/auth/TLS/backup/alert/write-freeze/ownership prerequisites
before activation. Complete initialization and live-readiness checks before
issuing any approval token.

- [ ] Phases 1-6 correctness, backend-parity, browser, failure-injection, and
  performance gates pass for the exact unchanged release candidate.
- [ ] The complete Phase 7 register is satisfied: clean-host Docker/package/raw
  installation, same-candidate security and vulnerability evidence,
  destructive soak and an operator recovery drill, final normal-replication
  performance, zero unresolved high-severity findings, and named security,
  correctness/recovery, browser UX/accessibility, operations, and performance
  approvals are recorded in the authoritative release evidence. For a product
  release, confirm the five protected GitHub Environment deployment records and
  the signed `pgpipe-<version>-release-gate-certification.tar.gz` receipt bundle
  against the authoritative GitHub policy and approval-history API evidence
  documented in
  [CI.md](CI.md#required-protected-github-environments). Each environment must
  use directly named individual reviewers; team-only review and configuration-
  marker variables are not accepted as policy evidence. No single approval
  substitutes for another domain, and the dedicated read-only
  `PGPIPE_RELEASE_POLICY_AUDIT_TOKEN` must be available for the environment-
  policy and immutable-release API checks. Each protected review job validates
  its own GitHub approval-history record separately.
- [ ] Dashboard authentication and CSRF origin protection are enabled.
- [ ] HTTPS or a trusted TLS-terminating proxy protects the dashboard.
- [ ] The configured destination role is a dedicated, least-privileged runtime
  role; the one-time initialization used a separate reviewed destination DBA.
- [ ] Source-table recreation protection and destination ownership fencing are
  healthy.
- [ ] The exact `pgpipe_safety` manifest is owned by
  `pgpipe_safety_owner`; `pgpipe repair check` and startup preflight pass.
- [ ] Backup/PITR restores of destination data and pgpipe state have been
  rehearsed as one coordinated recovery point.
- [ ] Alerts, metrics scraping, and operator-audit retention have been reviewed.
- [ ] Application owners can freeze direct destination writes to selected rows
  for the entire preview, approval, execution, and final-verification window.
- [ ] A named incident owner and database administrator are available for the
  maintenance window.

### Isolated demo and staging soak

The ordinary/default `make demo-up` path keeps repair execution off, but it is
still a disposable development demo—not a production deployment. The isolated
repair overlay requires Docker Compose 2.24.4 or newer for `!override`. Confirm
`docker compose version` first, then use it only on a development host:

```bash
make repair-demo-up
make repair-demo-test
make repair-demo-down
```

To exercise the prepare-disabled, enable-later lifecycle in the ordinary
disposable Docker demo, reset it and let `demo-up` perform the demo-internal
bootstrap automatically:

```bash
make demo-reset  # destructive: permanently removes the disposable demo volumes
make demo-up
# Confirm ordinary replication while repair is still disabled.
# In Configuration -> Repair Execution, enable and save the reviewed limits.
docker compose restart pgpipe
```

`make demo-up` bootstraps the exact greenfield manifest with demo-only
credentials before the service starts and leaves the checked-in default
`execution_enabled: false` unchanged. A network-isolated handoff receives the
fixed demo-only administrator and runtime DSNs as Compose secrets. The
bootstrap consumes private copies, and the setup containers, host files, and
credential handoff volume are removed afterward; the normal runtime container
never receives the administrator DSN. It refuses stale pre-greenfield volumes
with `repair_safety_reset_required` and never deletes them
automatically. Run `make demo-reset` explicitly if the disposable data may be
permanently erased. Confirm ordinary replication, then save the dashboard
setting and restart. Confirm the dashboard reports
**Repair enabled — ready** and the capability values described above before
creating a repair. This path uses demo credentials and disposable local
containers only; it is never a production provisioning procedure.

Use `make demo-down` for an ordinary stop that preserves all demo database,
configuration, and state volumes. It does not reset safety storage.

The same reset selects the immutable destination replication profile. Origin
is the default. To test the replica profile, use:

```bash
make demo-reset
PGPIPE_DEMO_SAFETY_PROFILE=replica make demo-up
```

The one-shot bootstrap and generated runtime configuration receive the same
profile. The normal dashboard displays it read-only. Changing an existing
destination between origin and replica requires another full disposable reset
and bootstrap; an in-place UI, API, or YAML switch is refused because it would
make protected destination identity disagree with runtime behavior.

The smoke test authenticates, creates one source-authoritative repair containing
exactly two destination updates, requires the complete approval contract, and
proves `completed / matched` plus durable audit history after restart.
`repair-demo-up` deletes the generated owner-only setup DSN file and its private
setup-only handoff volume after success, failure, or interruption. A successful
lifecycle fails closed if that cleanup cannot be proven and prints the exact
resource that must be removed. `repair-demo-down` removes only the
`pgpipe-repair-demo` Compose project and its volumes, and also performs a
fallback cleanup of the host setup file even when Compose teardown returns an
error. Those resources are disposable by design. This is not a production
enablement procedure.

First validate the soak harness without Docker mutation:

```bash
make repair-soak-contract-test
```

Run the repeated destructive/failure soak only in an explicitly approved,
disposable Docker environment:

```bash
PGPIPE_REPAIR_SOAK_ALLOW_DESTRUCTIVE=yes-delete-disposable-repair-soak-data \
REPAIR_SOAK_ITERATIONS=3 \
REPAIR_SOAK_IMPACT_RUNS=3 \
make repair-soak-test
```

The guarded harness creates and deletes only uniquely named harness resources,
repeats race/state/PostgreSQL failure tests, runs paired repair-impact trials,
sanitizes retained logs, records source-tree/build/environment identity, and
writes checksummed evidence under `benchmark-results/repair-soak-*` unless
`REPAIR_SOAK_ARTIFACT_DIR` selects an absolute, physical path whose final leaf
does not yet exist. Its parent must already exist, be owned by the invoking
user without group/world write access (or be a root-owned sticky shared
directory), and contain no symlink components. The harness atomically creates
the evidence leaf mode `0700`; it never reuses even an empty directory. A
source-tree change, inconclusive trial, reused project identity, or failed
stage fails the soak. Keep the capability default-off until reviewers accept
this evidence and the separate final performance certification.

Evidence review is explicit, not just a green command exit:

- require `run-status.tsv` to say `PASS` and every row in `validation.tsv` to
  say `PASS`;
- from the artifact root, verify every entry in `SHA256SUMS` with
  `sha256sum -c SHA256SUMS` (or `shasum -a 256 -c SHA256SUMS`);
- match the outer source-tree hash, immutable pgpipe/PostgreSQL image IDs, and
  pgpipe binary hash to the exact release candidate;
- confirm both required test inventories are exact and the JSON-derived run
  counts equal the requested repetitions with no failure or skip;
- require every nested impact run to have a unique pair identity, matching
  outer source/image/binary identities, a valid nested checksum manifest, and
  a passing secret scan;
- treat paired active-repair impact as a correctness/overlap diagnostic, not
  the steady-state replication performance gate; certify normal replication
  separately on the exact release artifact.

Gitignored workstation evidence and the workflow's 90-day CI artifact are
temporary review inputs, not durable release records. Retain only the
sanitized, checksum-verified bundle in the access-controlled release record.

## Destination initialization and privilege matrix

Destination safety storage is a greenfield installation artifact. The native
`pgpipe init` browser flow performs one bounded, destination-global
bootstrap transaction with a reviewed DBA connection, validates the runtime
connection against the same database, discards the DBA credential, and publishes
repair disabled. The setup origin is random and loopback-only; use the printed
SSH-tunnel instructions for a remote browser. The normal dashboard never
accepts the DBA credential, and the fresh-install path has no separate repair
bootstrap command.

A fresh disposable Docker demo performs the same bootstrap internally during
`make demo-up`. Do not copy its credentials into a real deployment. Old
development objects are unsupported and require a destructive demo reset.

### Managed PostgreSQL that prohibits role creation

If the administrator offered by the database provider cannot execute `CREATE
ROLE`, initialization fails with the stable role-creation-prohibited error. It
must never fall back to `postgres`, the runtime writer, or an
administrator-owned manifest. Ask the provider or highest-authority DBA to
create only the fixed cluster-global owner before rerunning `pgpipe init`:

```sql
CREATE ROLE pgpipe_safety_owner
  NOLOGIN NOINHERIT NOSUPERUSER NOCREATEDB NOCREATEROLE
  NOREPLICATION NOBYPASSRLS
  CONNECTION LIMIT -1
  PASSWORD NULL
  VALID UNTIL 'infinity';

COMMENT ON ROLE pgpipe_safety_owner IS
  'pgpipe protected destination safety owner v1';

GRANT SET ON PARAMETER session_replication_role
  TO pgpipe_safety_owner;
```

The provider/DBA must leave this role with no direct or indirect memberships,
role settings, login path, other owned objects in the destination database, or
grant/admin path from the runtime role. Do not pre-create either pgpipe schema
or any protected table. The browser finalization then validates the role and
creates the exact database-local manifest; any extra authority fails closed.
PostgreSQL 16 and newer may record an automatic `ADMIN` membership for the role
creator, so the highest-authority administrator must remove that edge before
initialization. The complete validation query and provider notes are in the
[fresh-install guide](README.md#managed-postgresql-that-cannot-create-roles).

Do not treat similar names as one credential. The installation has seven
separate identities:

| Identity | Where it exists | Purpose and ownership boundary |
| --- | --- | --- |
| Source replication role | Source PostgreSQL | Streams WAL and reads selected source rows for snapshot/proof work; has no destination or dashboard authority. |
| Destination application-table owner | Destination PostgreSQL | Owns customer schemas/tables in the hardened profile or explicitly delegates a bounded customer-schema DDL profile; owns no safety object. |
| Temporary bootstrap administrator | Destination PostgreSQL, only during `pgpipe init` | Creates/validates protected storage and exact grants in the already reviewed database; the credential is discarded and is never accepted by the running dashboard. |
| `pgpipe_safety_owner` | Destination PostgreSQL | Permanent non-login owner of `pgpipe_safety`, its exact manifest, and the replica-session activator; it cannot be used as a service login or reached with `SET ROLE`. |
| `pgpipe_writer` (configured destination runtime role) | Destination PostgreSQL | Applies destination changes and receives only the required protected-object DML/activator privileges; it never owns protected storage. |
| Dashboard administrator | pgpipe authentication | Enables/disables repair, reviews and approves exact jobs, and audits results; it receives no PostgreSQL DBA power. |
| Linux service account (`pgpipe`) | Operating system | Runs the process and owns local config/state files; it is not automatically a database role or dashboard approver. |

One person may administer several realms, but these remain distinct authority
boundaries. In particular, the temporary DBA, application-table owner,
dashboard administrator, and Linux service account do not become
`pgpipe_safety_owner` or `pgpipe_writer` by sharing a human operator.

The protected owner is the fixed `pgpipe_safety_owner` role. It is `NOLOGIN`,
`NOINHERIT`, `NOSUPERUSER`, `NOCREATEDB`, `NOCREATEROLE`, `NOREPLICATION`,
`NOBYPASSRLS`, has no password, and has no role memberships. It owns
`pgpipe_safety` and the exact protected manifest; the configured runtime login
owns none of them and cannot create in that schema. `PUBLIC` has no access.

This owner is `NOLOGIN` because PostgreSQL ownership itself includes powers
that grants cannot take away: an owner can alter or drop its objects and change
their ACLs. If the service authenticated as the owner, a leaked runtime
credential or defective DDL path could bypass the safety boundary. `NOLOGIN`
must be combined with the required empty, non-administrable membership graph;
otherwise a login could still reach the same powers with `SET ROLE`.

| Protected object | Direct, non-grantable runtime privileges |
| --- | --- |
| `pgpipe_safety.destination_identity` | `SELECT` |
| `pgpipe_safety.pipeline_fence` | `SELECT, INSERT, UPDATE` |
| `pgpipe_safety.pipeline_fence_audit` | `SELECT, INSERT` |
| `pgpipe_safety.apply_ledger` | `SELECT, INSERT, UPDATE, DELETE` |
| `pgpipe_safety.table_rebuild_ledger` | `SELECT, INSERT, UPDATE` |
| `pgpipe_safety.repair_execution_fence` | `SELECT, INSERT, UPDATE` |
| `pgpipe_safety.repair_chunk_ledger` | `SELECT, INSERT` |
| `pgpipe_safety.activate_replica_session_v1()` | `EXECUTE` only when the immutable destination profile is `replica`; no runtime grant when it is `origin` |

The runtime role receives `USAGE` on `pgpipe_safety` without `CREATE` or grant
options. Runtime-managed heartbeat objects live in the separate
`pgpipe_runtime` schema. Both schema names are reserved against table
selection, mapping, transformation, dynamic addition, and DDL replay.

Application-table authority is a separate installation choice:

| Profile | Customer-object authority | What remains available |
| --- | --- | --- |
| **Hardened DML-only** | The application role retains ownership. Give the runtime schema `USAGE` and only the reviewed table/sequence privileges required for the configured pipeline, normally `SELECT, INSERT, UPDATE, DELETE` on selected tables. | Streaming, verification, and repair against pre-created compatible tables. Automatic table creation, reset/truncate, structured destination DDL, and shadow rebuild remain unavailable unless their additional privileges are deliberately granted. |
| **pgpipe-managed DDL** | In a dedicated application schema, deliberately grant the runtime `USAGE, CREATE` plus only the table/sequence ownership or operation privileges needed by enabled snapshot, supported DDL, reset, and rebuild workflows. | The explicitly enabled creation/DDL/rebuild features, with a larger customer-data and schema blast radius. |

Neither profile grants database-level `CREATE`, ownership of `pgpipe_safety`,
or any extra protected-object privilege. This is not a dashboard switch. If
preflight rejects a missing customer privilege, either enable the intended
managed capability deliberately or keep the hardened profile and perform that
schema operation through the application owner; never broaden safety grants to
silence the error.

Runtime startup validates this exact owner, role graph, ACL, relation shape,
constraints, indexes, comments, executable behavior, dependencies, and
destination identity in a database-enforced read-only transaction. It never
creates, alters, upgrades, grants on, or repairs protected storage. The same
validation runs immediately before repair admission, and physical-relation
plus fence checks remain inside every repair chunk transaction. None of the
complete catalog checks run per replicated event.

After initialization, validate through the configured runtime role:

```bash
sudo -u pgpipe -- /usr/bin/pgpipe repair check -c /etc/pgpipe/pgpipe.yaml
```

This command is read-only, accepts no DBA credential, performs no DML probe,
and never enables or admits repair. A successful check does not replace live
startup and pre-admission readiness checks.

The runtime destination role also needs ordinary DML on destination
application tables and, when hybrid writing is enabled, `TEMPORARY` on the
destination database. The source role retains its logical-replication and
read/proof privileges. Final repair verification needs:

```sql
GRANT EXECUTE ON FUNCTION pg_catalog.pg_logical_emit_message(boolean, text, text)
TO pgpipe_source;
```

The normative object, role, and stable-error contract is the
[Greenfield Destination Safety Contract](DESTINATION_SAFETY_CONTRACT.md).

### Future protected-schema upgrades

The current runtime, normal dashboard, and DEB/RPM package scripts never create
or upgrade `pgpipe_safety`. A future release that changes its manifest must
ship an explicit release-specific upgrade workflow. Expect it to require a
maintenance window, every old process stopped, no mixed old/new binaries, a
coordinated destination/state/global-role backup, and one newly supplied
temporary DBA credential for a bounded identity-checked upgrade. It must then
discard that credential, validate through `pgpipe_writer`, and retain a
sanitized receipt before restart.

Until such a workflow is shipped for the exact future version, a safety-schema
version mismatch is a fail-closed condition. Do not rerun today's greenfield
`pgpipe init` over an existing config, edit protected tables, grant ownership to
the writer, or expect a package upgrade to change the database.

## Understand `session_replication_role: replica`

This setting applies PostgreSQL's replication-session behavior on destination
connections. It normally suppresses ordinary destination row triggers and
foreign-key rechecks for both streaming and repair writes. That can improve
throughput and prevent a destination trigger from duplicating source-side
effects, but it changes the destination integrity boundary:

- application audit/outbox triggers do not run unless explicitly configured to
  fire for replica sessions with `ALTER TABLE ... ENABLE REPLICA TRIGGER ...`
  or for every session with `ENABLE ALWAYS TRIGGER`;
- foreign keys are not rechecked for each applied row;
- destination-only trigger side effects are not produced;
- pgpipe still validates repair proofs, relation identities, fences, and its
  internal chunk ledger; the setting does not bypass those controls;
- sequences, materialized views, aggregates, and external side effects are not
  automatically reconciled by a targeted row repair.

Use it only when the source is authoritative and source integrity is trusted.
After a repair, verify application invariants that would otherwise have been
enforced by destination triggers or constraints. Keep the setting empty when a
destination audit trigger or other downstream consumer must observe each row
operation.

## Backup, PITR, and destination-write freeze

Before approval:

1. Freeze direct application, administrator, ETL, and maintenance writes to the
   selected destination rows. For a range repair, freeze the whole range.
2. Keep source application writes quiet when an audit-quality preview and final
   verification are required. At minimum, understand that separate source and
   destination snapshots can make a concurrent-write result inconclusive.
3. Confirm the destination backup/PITR point and the matching pgpipe state
   backup. A plain live copy of a BoltDB or SQLite file is invalid: briefly
   stop pgpipe and verify it is quiescent before an offline file-backend
   capture, or use a documented backend-native consistent snapshot. For
   PostgreSQL state, use the database backup system's transactionally
   consistent method. Record timestamps, database timeline/LSN evidence,
   state backend, pipeline identity, build version, and job ID.
4. Preserve WAL, logs, alerts, audit data, and destination ledger evidence
   through the maintenance window.

Destination data and pgpipe state are one recovery cohort. Restoring only the
destination or only the state can create a false authority boundary and must
not be used to clear a repair blocker. For coordinated PITR, restore both to a
proven compatible point, validate destination fencing and state lineage, then
run verification before writes resume.

Database-only backups commonly do not capture PostgreSQL cluster-global role
metadata. The recovery plan must separately preserve or recreate the exact
`pgpipe_safety_owner` attributes, empty membership graph, password-null
attestation, and `SET` privilege on `session_replication_role`, plus the runtime
role boundary. A dump restored with owner/ACL suppression is not ready merely
because the seven tables exist. After restoring or pre-creating those globals,
run `pgpipe repair check`, startup preflight, destination/state identity checks,
and content verification before admitting repair or resuming application
writes. A clone or restore into a different database identity is a new
lifecycle unless the documented lineage recovery proves otherwise.

The dashboard's “destination writes stopped” acknowledgement is an operator
assertion, not a database lock over every outside client. pgpipe cannot stop an
administrator or application using another connection.

## Server-owned time limits

Requests cannot override these settings:

| Setting | Default | Accepted range | Meaning |
| --- | ---: | ---: | --- |
| `repair.approved_queue_lifetime` | 15m | 1m-1h | Maximum time an admitted, unclaimed job may wait. |
| `repair.max_execution_duration` | 30m | 1m-2h | Absolute wall-clock deadline persisted when a worker claims the job. |
| `repair.connection_acquire_timeout` | 10s | 1s-30s | Bound for obtaining each repair database connection. |
| `repair.statement_timeout` | 30s | 1s-5m | PostgreSQL deadline for an individual repair statement. |
| `repair.lock_timeout` | 5s | 100ms-30s | PostgreSQL lock wait; it must not exceed `statement_timeout`. |
| `repair.recovery_cleanup_timeout` | 30s | 1s-2m | No-new-writes evidence reconciliation and proven pre-write cleanup budget. |

Every sub-deadline is clamped to the remaining whole-execution time. Shutdown
stops new claims and chunks immediately; an already open destination
transaction remains bounded by the existing statement, lock, connection, and
absolute execution deadlines.

Timeout classifications are evidence-based:

| Situation | Durable classification/action |
| --- | --- |
| Queue lifetime expires before claim | `expired_before_claim`; no worker starts. |
| Whole execution deadline expires | reason `repair_execution_timed_out`; classify by proven write effect. |
| PostgreSQL statement limit fires first | reason `repair_statement_timed_out`; classify by proven write effect. |
| PostgreSQL lock limit fires first | reason `repair_lock_timed_out`; classify by proven write effect. |
| Timeout before any write is positively proven | `interrupted_prewrite`; fenced cleanup must complete before guard release. |
| Some commit is proven | `interrupted_partial`; offline recovery is required. |
| COMMIT result or evidence is uncertain | `interrupted_ambiguous`; offline recovery is required. |
| Writes completed but verification did not prove the result | `interrupted_postwrite`; offline recovery is required. |

A timeout is never labeled canceled, auto-resumed, or automatically retried.

## Repair-control rate limits

The first release uses fixed one-minute windows. Both the authenticated actor
bucket and pipeline bucket must have capacity:

| Mutation | Per actor/minute | Per pipeline/minute |
| --- | ---: | ---: |
| Create preview | 10 | 30 |
| Materialize Verify candidate | 10 | 30 |
| Issue/rotate approval token | 10 | 30 |
| Approve | 5 | 15 |
| Cancel before claim | 10 | 30 |
| Rerun verification | 5 | 15 |

Exhaustion returns `429 repair_rate_limited` and an integer `Retry-After`
header. Wait at least that many seconds. For preview, materialization, approval,
or cancellation, retry the exact request with the same `Idempotency-Key`; do
not generate a new key merely because the response was lost. Token issuance
rotates the one-time token, so use only the latest successful response. A
rerun-verification request is not a repair retry and has no idempotency key.

The limiter is control-plane only and adds no work to the replication hot path.
Its buckets are memory owned by one running API process and reset on process
restart. Active/passive or horizontally exposed instances therefore need a
shared reverse-proxy/API-gateway limit if operators require one aggregate
cross-instance policy; pgpipe's durable idempotency and fencing still remain
the execution authority.

### HTTP remediation

Every v2 error uses a bounded structured body with a stable `code`, safe
message, remediation, and retryable flag. Handle the status before deciding
whether to create a new token or idempotency key:

| HTTP | Meaning | Safe response |
| ---: | --- | --- |
| `401` | Login is missing or expired. | Sign in again, reload job detail, and obtain a fresh token before approval. Never add an actor field. |
| `403` | Authentication authority is disabled/forbidden, or CSRF origin was rejected. | Correct dashboard auth/authorization or the exact `Origin`/`server.csrf.trusted_origins`; do not bypass middleware. Restart if configuration changed. |
| `409` | Current phase, token/revision, idempotency content, claim/cancel race, or recovery state conflicts. | Reload detail first. Poll an already admitted job; create a fresh preview/token only when the stable error remediation says so. Never assume a conflict means zero writes. |
| `422` | Materialization/fresh preview, eligibility, reason, or acknowledgement is incomplete. | Fix the stated precondition and review the resulting current job again. No preview-drift override exists. |
| `429` | Per-actor or per-pipeline control rate was exceeded. | Wait for the integer `Retry-After`; retry the exact idempotent operation with the same key. Use only the latest successful token issuance. |
| `503` | Execution is disabled or pipeline/state/destination readiness is unavailable. | Restore the prerequisite, re-read capabilities and detail, then retry the exact still-current idempotent request. If its preview/token expired, create fresh authority and a new key rather than changing content under the old key. |

## Preferred dashboard procedure

Use the dashboard unless a reviewed automation client is required:

1. Run Verify & Repair for the exact table and inspect the mismatch.
2. Select **Review repair**. A Verify candidate is materialized into a fresh,
   proof-bearing job; an older candidate is superseded, not rewritten.
3. Confirm the selected keys/range and the exact insert, update, and delete
   counts. For example, `0 insert / 1 update / 0 delete` means one destination
   row will be overwritten from the source.
4. Freeze direct destination writes and record backup/PITR evidence.
5. Enter a non-sensitive reason, satisfy the destructive/delete/large-repair
   acknowledgements shown by the server, acknowledge the write freeze, and type
   the exact current confirmation phrase.
6. Approve once. The accepted job enters `approved_queued`; use **View
   progress** rather than resubmitting.
7. Keep the write freeze until the terminal result and final verification have
   been reviewed.

If the page was refreshed, the token response was lost, or the dialog became
stale, reload job detail and request a new token. Never expect a GET response to
recover a plaintext approval token.

## Safe v2 API example

The dashboard is preferred because it handles fresh detail, one-time token
rotation, acknowledgements, and confirmation safely. The following example is
for a short-lived, access-controlled Bash shell with `curl`, `jq`, `python3`,
`cmp`, `find`, and standard core utilities, plus the correct CA file. Do not
enable shell tracing or paste secrets into tickets. Remove temporary session
files afterward; retain any unreconciled owner-only replay bundle until its
exact outcome is proven.

Set the endpoint and authenticate:

```bash
set -euo pipefail
set +x
export PGPIPE_URL='https://pgpipe.example.com'
export PGPIPE_ORIGIN='https://pgpipe.example.com'
export PGPIPE_CA='/etc/pgpipe-tls/ca.crt'
read -r -p 'Dashboard username: ' PGPIPE_DASHBOARD_USERNAME
umask 077
PGPIPE_PRIVATE_ROOT="${TMPDIR:-/tmp}"
readonly PGPIPE_PRIVATE_ROOT
PGPIPE_PRIVATE_DIR="$(mktemp -d "$PGPIPE_PRIVATE_ROOT/pgpipe-repair.XXXXXX")" || exit 1
case "$PGPIPE_PRIVATE_DIR" in
  /*) ;;
  *) echo 'Private working directory is not absolute.' >&2; exit 1 ;;
esac
[ -d "$PGPIPE_PRIVATE_DIR" ] && [ ! -L "$PGPIPE_PRIVATE_DIR" ] &&
  [ -O "$PGPIPE_PRIVATE_DIR" ] || exit 1
chmod 0700 -- "$PGPIPE_PRIVATE_DIR"
readonly PGPIPE_PRIVATE_DIR
PGPIPE_AUTH_CONFIG="$PGPIPE_PRIVATE_DIR/curl-auth.conf"
cleanup_pgpipe_repair_session() {
  case "$PGPIPE_PRIVATE_DIR" in
    "$PGPIPE_PRIVATE_ROOT"/pgpipe-repair.*) ;;
    *) return 1 ;;
  esac
  [ -d "$PGPIPE_PRIVATE_DIR" ] && [ ! -L "$PGPIPE_PRIVATE_DIR" ] &&
    [ -O "$PGPIPE_PRIVATE_DIR" ] || return 1
  for file in "$PGPIPE_PRIVATE_DIR"/*; do
    [ -e "$file" ] || continue
    [ -f "$file" ] && [ ! -L "$file" ] || return 1
    rm -f -- "$file"
  done
  rmdir -- "$PGPIPE_PRIVATE_DIR"
}
trap cleanup_pgpipe_repair_session EXIT

read -r -s -p 'Dashboard password: ' PGPIPE_DASHBOARD_PASSWORD
printf '\n'
printf '%s' "$PGPIPE_DASHBOARD_PASSWORD" \
  >"$PGPIPE_PRIVATE_DIR/dashboard-password.txt"
unset PGPIPE_DASHBOARD_PASSWORD
jq -n --arg username "$PGPIPE_DASHBOARD_USERNAME" \
  --rawfile password "$PGPIPE_PRIVATE_DIR/dashboard-password.txt" \
  '{username:$username,password:$password}' \
  >"$PGPIPE_PRIVATE_DIR/login-request.json"
rm -f -- "$PGPIPE_PRIVATE_DIR/dashboard-password.txt"

LOGIN_HTTP_STATUS="$(curl --disable --silent --show-error --cacert "$PGPIPE_CA" \
    -H "Origin: $PGPIPE_ORIGIN" \
    -H 'Content-Type: application/json' \
    --data-binary "@$PGPIPE_PRIVATE_DIR/login-request.json" \
    -o "$PGPIPE_PRIVATE_DIR/login-response.json" -w '%{http_code}' \
    "$PGPIPE_URL/api/auth/login")" || exit 1
[ "$LOGIN_HTTP_STATUS" = 200 ] || {
  echo "Login failed with HTTP $LOGIN_HTTP_STATUS." >&2
  exit 1
}
unset LOGIN_HTTP_STATUS
rm -f -- "$PGPIPE_PRIVATE_DIR/login-request.json"

PGPIPE_API_TOKEN="$(jq -er '.token' "$PGPIPE_PRIVATE_DIR/login-response.json")"
printf 'header = "Authorization: Bearer %s"\n' "$PGPIPE_API_TOKEN" \
  >"$PGPIPE_AUTH_CONFIG"
unset PGPIPE_API_TOKEN
PGPIPE_AUTHENTICATED_ACTOR="$PGPIPE_DASHBOARD_USERNAME"
unset PGPIPE_DASHBOARD_USERNAME
rm -f -- "$PGPIPE_PRIVATE_DIR/login-response.json"

PGPIPE_STATUS="$PGPIPE_PRIVATE_DIR/status.json"
curl --disable --config "$PGPIPE_AUTH_CONFIG" \
  --fail --silent --show-error --cacert "$PGPIPE_CA" \
  -o "$PGPIPE_STATUS" "$PGPIPE_URL/api/status"
PGPIPE_PIPELINE_NAME="$(jq -er '.pipeline_name | select(type == "string" and length > 0)' \
  "$PGPIPE_STATUS")"
PGPIPE_PIPELINE_ID="$(printf '%s' "$PGPIPE_PIPELINE_NAME" | python3 -c \
  'import hashlib,sys; print("pgpipe-" + hashlib.sha256(b"pgpipe-pipeline\0" + sys.stdin.buffer.read()).hexdigest()[:32])')"
case "$PGPIPE_PIPELINE_ID" in
  pgpipe-*) pipeline_id_suffix=${PGPIPE_PIPELINE_ID#pgpipe-} ;;
  *) echo 'Current pipeline identity could not be derived safely.' >&2; exit 1 ;;
esac
case "$pipeline_id_suffix" in
  ''|*[!0-9a-f]*) echo 'Current pipeline identity is invalid.' >&2; exit 1 ;;
esac
[ "${#pipeline_id_suffix}" -eq 32 ] || exit 1
rm -f -- "$PGPIPE_STATUS"
unset PGPIPE_PIPELINE_NAME pipeline_id_suffix
```

The owner-only directory keeps the password, bearer value, approval token, and
request bodies out of child-process command arguments. Keep the shell and
directory private; the trap removes the complete directory at session exit.

Preview and materialization do not write destination rows, but both accept an
idempotency key and materialization changes durable repair-control state. A
lost response must therefore remain exactly retryable. For each request, use a
different existing empty owner-only directory on durable incident storage,
outside `PGPIPE_PRIVATE_DIR`. Create it with mode `0700` beneath a parent chain
owned by the current user or root and not writable by another user. A physical
root-owned sticky shared directory such as `/tmp` on Linux or `/private/tmp` on
macOS is accepted, but no symlink component is accepted. Preserved bundle files
must already be owner-only mode `0600`; the helpers fail closed instead of
changing permissions after authority may have leaked:

If a preserved bundle later fails either check, do not `chmod` it and replay.
Preserve it as incident evidence, treat its authority as potentially exposed,
and reconcile the durable job outcome before issuing any fresh authority.

```bash
require_private_replay_path() {
  local replay_path=$1 replay_kind=$2
  python3 - "$replay_path" "$replay_kind" <<'PY'
import os
import stat
import sys

path, kind = sys.argv[1:]
if kind not in ("directory", "file"):
    raise SystemExit(1)
if (not path.startswith("/") or path == "/" or path.endswith("/") or
        "//" in path or any(ord(character) < 32 or ord(character) == 127
                            for character in path)):
    raise SystemExit(1)
parts = path.split("/")[1:]
if not parts or any(part in ("", ".", "..") for part in parts):
    raise SystemExit(1)

if (not hasattr(os, "O_NOFOLLOW") or not hasattr(os, "O_DIRECTORY") or
        not hasattr(os, "O_NONBLOCK")):
    raise SystemExit(1)
if (os.stat not in os.supports_dir_fd or
        os.stat not in os.supports_follow_symlinks):
    raise SystemExit(1)
directory_flags = os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW
if hasattr(os, "O_CLOEXEC"):
    directory_flags |= os.O_CLOEXEC
file_flags = os.O_RDONLY | os.O_NOFOLLOW | os.O_NONBLOCK
if hasattr(os, "O_CLOEXEC"):
    file_flags |= os.O_CLOEXEC

effective_uid = os.geteuid()
try:
    current_fd = os.open("/", directory_flags)
except OSError:
    raise SystemExit(1)
try:
    for index, component in enumerate(parts):
        final = index == len(parts) - 1
        flags = file_flags if final and kind == "file" else directory_flags
        named_details = None
        if final and kind == "file":
            try:
                named_details = os.stat(
                    component, dir_fd=current_fd, follow_symlinks=False
                )
            except OSError:
                raise SystemExit(1)
            if not stat.S_ISREG(named_details.st_mode):
                raise SystemExit(1)
        try:
            next_fd = os.open(component, flags, dir_fd=current_fd)
        except OSError:
            raise SystemExit(1)
        details = os.fstat(next_fd)
        if (named_details is not None and
                (details.st_dev, details.st_ino) !=
                (named_details.st_dev, named_details.st_ino)):
            os.close(next_fd)
            raise SystemExit(1)
        mode = stat.S_IMODE(details.st_mode)
        if final:
            expected_type = stat.S_ISREG if kind == "file" else stat.S_ISDIR
            expected_mode = 0o600 if kind == "file" else 0o700
            if (not expected_type(details.st_mode) or
                    details.st_uid != effective_uid or mode != expected_mode):
                os.close(next_fd)
                raise SystemExit(1)
        else:
            if not stat.S_ISDIR(details.st_mode):
                os.close(next_fd)
                raise SystemExit(1)
            if details.st_uid not in (0, effective_uid):
                os.close(next_fd)
                raise SystemExit(1)
            writable_by_others = mode & 0o022
            root_sticky_shared = details.st_uid == 0 and mode & stat.S_ISVTX
            if writable_by_others and not root_sticky_shared:
                os.close(next_fd)
                raise SystemExit(1)
        os.close(current_fd)
        current_fd = next_fd
finally:
    os.close(current_fd)
PY
}

require_empty_owner_only_replay_dir() {
  local replay_dir=$1 first_entry
  if ! require_private_replay_path "$replay_dir" directory; then
    echo 'Replay directory and its parent chain are not private physical paths.' >&2
    return 1
  fi
  if ! first_entry="$(find "$replay_dir" -mindepth 1 -maxdepth 1 -print -quit)"; then
    echo 'Replay directory contents could not be inspected safely.' >&2
    return 1
  fi
  if [ -n "$first_entry" ]; then
    echo 'Replay directory must be empty.' >&2
    return 1
  fi
}

require_owner_only_replay_bundle() {
  local replay_dir=$1 manifest recorded_url recorded_origin recorded_actor recorded_pipeline
  if ! require_private_replay_path "$replay_dir" directory; then
    echo 'Replay bundle and its parent chain are not private physical paths.' >&2
    return 1
  fi
  manifest="$replay_dir/manifest.tsv"
  require_private_replay_path "$manifest" file || return 1
  awk -F '\t' 'NF != 2 { exit 1 }' "$manifest" || return 1
  recorded_url="$(awk -F '\t' '
    $1 == "api_base_url" { if (NF != 2) exit 2; value = $2; count++ }
    END { if (count == 1) print value; else exit 1 }
  ' "$manifest")" || return 1
  recorded_origin="$(awk -F '\t' '
    $1 == "api_origin" { if (NF != 2) exit 2; value = $2; count++ }
    END { if (count == 1) print value; else exit 1 }
  ' "$manifest")" || return 1
  recorded_actor="$(awk -F '\t' '
    $1 == "actor" { if (NF != 2) exit 2; value = $2; count++ }
    END { if (count == 1) print value; else exit 1 }
  ' "$manifest")" || return 1
  recorded_pipeline="$(awk -F '\t' '
    $1 == "pipeline_id" { if (NF != 2) exit 2; value = $2; count++ }
    END { if (count == 1) print value; else exit 1 }
  ' "$manifest")" || return 1
  if [ "$recorded_url" != "$PGPIPE_URL" ] ||
     [ "$recorded_origin" != "$PGPIPE_ORIGIN" ] ||
     [ "$recorded_actor" != "$PGPIPE_AUTHENTICATED_ACTOR" ] ||
     [ "$recorded_pipeline" != "$PGPIPE_PIPELINE_ID" ]; then
    echo 'Current API, actor, or pipeline differs from the saved replay authority.' >&2
    return 1
  fi
}

require_replay_bundle_files() {
  local replay_dir=$1 path name allowed_file allowed
  shift
  require_private_replay_path "$replay_dir" directory || return 1
  for path in "$replay_dir"/* "$replay_dir"/.[!.]* "$replay_dir"/..?*; do
    if [ ! -e "$path" ] && [ ! -L "$path" ]; then
      continue
    fi
    require_private_replay_path "$path" file || return 1
    name=${path##*/}
    allowed=false
    for allowed_file in "$@"; do
      if [ "$name" = "$allowed_file" ]; then
        allowed=true
        break
      fi
    done
    [ "$allowed" = true ] || {
      echo "Replay bundle contains unexpected file: $name" >&2
      return 1
    }
  done
}

require_replay_manifest_value() {
  local replay_dir=$1 field=$2 expected=$3 actual
  actual="$(awk -F '\t' -v field="$field" '
    $1 == field { if (NF != 2) exit 2; value = $2; count++ }
    END { if (count == 1) print value; else exit 1 }
  ' "$replay_dir/manifest.tsv")" || return 1
  [ "$actual" = "$expected" ] || {
    echo "Replay manifest field differs from the exact request: $field" >&2
    return 1
  }
}

require_uuid_idempotency_pair() {
  local key_file=$1 config_file=$2 path key key_bytes expected actual config_bytes
  for path in "$key_file" "$config_file"; do
    require_private_replay_path "$path" file || return 1
  done
  key_bytes="$(wc -c <"$key_file" | tr -d '[:space:]')" || return 1
  [ "$key_bytes" = 36 ] || return 1
  key="$(<"$key_file")"
  case "$key" in
    *[!0-9a-f-]*|'') return 1 ;;
  esac
  [ "${key:8:1}" = - ] && [ "${key:13:1}" = - ] &&
    [ "${key:18:1}" = - ] && [ "${key:23:1}" = - ] || return 1
  expected="header = \"Idempotency-Key: $key\""
  actual="$(<"$config_file")"
  config_bytes="$(wc -c <"$config_file" | tr -d '[:space:]')" || return 1
  [ "$actual" = "$expected" ] &&
    [ "$config_bytes" = "$((${#expected} + 1))" ] || return 1
}

preview_response_job_id() {
  jq -er '
    select(type == "object" and (.id | type) == "string" and
      (.id | test("^[A-Za-z0-9._:-]+$")) and .phase == "previewed")
    | .id
  ' "$1"
}

materialized_response_job_id() {
  jq -er --arg candidate "$2" '
    select(type == "object" and (.id | type) == "string" and
      (.id | test("^[A-Za-z0-9._:-]+$")) and .phase == "previewed" and
      .supersedes == $candidate)
    | .id
  ' "$1"
}
```

Create a bounded manual preview for key `id=42`:

```bash
PREVIEW_REPLAY_DIR='/absolute/path/to/empty-preview-replay-directory'
require_empty_owner_only_replay_dir "$PREVIEW_REPLAY_DIR"
PREVIEW_KEY="$(python3 -c 'import uuid; print(uuid.uuid4())')"
PREVIEW_REQUEST="$PREVIEW_REPLAY_DIR/preview-request.json"
PREVIEW_RESPONSE="$PREVIEW_REPLAY_DIR/preview-response.json"
PREVIEW_KEY_CONFIG="$PREVIEW_REPLAY_DIR/preview-idempotency.conf"
PREVIEW_KEY_FILE="$PREVIEW_REPLAY_DIR/preview-idempotency-key.txt"
PREVIEW_STATUS_FILE="$PREVIEW_REPLAY_DIR/preview-http-status.txt"
PREVIEW_MANIFEST="$PREVIEW_REPLAY_DIR/manifest.tsv"
printf '%s' "$PREVIEW_KEY" >"$PREVIEW_KEY_FILE"
printf 'header = "Idempotency-Key: %s"\n' "$PREVIEW_KEY" >"$PREVIEW_KEY_CONFIG"
unset PREVIEW_KEY
jq -n '{
  schema:"public",
  table:"users",
  selector:{key_columns:["id"],keys:[{id:"42"}]}
}' >"$PREVIEW_REQUEST"
{
  printf 'contract_revision\t1\n'
  printf 'api_base_url\t%s\n' "$PGPIPE_URL"
  printf 'api_origin\t%s\n' "$PGPIPE_ORIGIN"
  printf 'actor\t%s\n' "$PGPIPE_AUTHENTICATED_ACTOR"
  printf 'pipeline_id\t%s\n' "$PGPIPE_PIPELINE_ID"
  printf 'preview_endpoint\t/api/v2/repair/preview\n'
} >"$PREVIEW_MANIFEST"
chmod 0600 -- "$PREVIEW_REQUEST" "$PREVIEW_KEY_FILE" "$PREVIEW_KEY_CONFIG" \
  "$PREVIEW_MANIFEST"
require_uuid_idempotency_pair "$PREVIEW_KEY_FILE" "$PREVIEW_KEY_CONFIG"
if PREVIEW_HTTP_STATUS="$(curl --disable --config "$PGPIPE_AUTH_CONFIG" \
    --config "$PREVIEW_KEY_CONFIG" --silent --show-error --cacert "$PGPIPE_CA" \
    -H "Origin: $PGPIPE_ORIGIN" \
    -H 'Content-Type: application/json' \
    --data-binary "@$PREVIEW_REQUEST" \
    -o "$PREVIEW_RESPONSE" -w '%{http_code}' \
    "$PGPIPE_URL/api/v2/repair/preview")"; then
  printf '%s' "$PREVIEW_HTTP_STATUS" >"$PREVIEW_STATUS_FILE"
else
  printf 'transport_error' >"$PREVIEW_STATUS_FILE"
  echo "Preview result is unknown; preserve $PREVIEW_REPLAY_DIR and retry exactly." >&2
  exit 1
fi
[ "$PREVIEW_HTTP_STATUS" = 201 ] || {
  echo "Preview was not proven; preserve $PREVIEW_REPLAY_DIR and reconcile it." >&2
  exit 1
}
if ! REPAIR_ID="$(preview_response_job_id "$PREVIEW_RESPONSE")"; then
  echo "Preview response was malformed; preserve $PREVIEW_REPLAY_DIR." >&2
  exit 1
fi
case "$REPAIR_ID" in
  ''|*[!A-Za-z0-9._:-]*) echo 'Server returned an invalid repair ID.' >&2; exit 1 ;;
esac
rm -f -- "$PREVIEW_REQUEST" "$PREVIEW_RESPONSE" "$PREVIEW_KEY_FILE" \
  "$PREVIEW_KEY_CONFIG" "$PREVIEW_STATUS_FILE" "$PREVIEW_MANIFEST"
rmdir -- "$PREVIEW_REPLAY_DIR"
unset PREVIEW_HTTP_STATUS
```

If starting from a Verify-created candidate instead, materialize it once and
use the returned new job ID. Materialization changes repair-control state but
does not write destination rows. Preserve the exact candidate, endpoint, and
idempotency key until HTTP `201` and a valid successor job are proven:

```bash
CANDIDATE_ID='verify-candidate-id-from-job-detail'
case "$CANDIDATE_ID" in
  ''|*[!A-Za-z0-9._:-]*) echo 'Candidate ID is invalid.' >&2; exit 1 ;;
esac
MATERIALIZE_REPLAY_DIR='/absolute/path/to/empty-materialization-replay-directory'
require_empty_owner_only_replay_dir "$MATERIALIZE_REPLAY_DIR"
MATERIALIZE_KEY="$(python3 -c 'import uuid; print(uuid.uuid4())')"
MATERIALIZE_RESPONSE="$MATERIALIZE_REPLAY_DIR/materialize-response.json"
MATERIALIZE_KEY_CONFIG="$MATERIALIZE_REPLAY_DIR/materialize-idempotency.conf"
MATERIALIZE_KEY_FILE="$MATERIALIZE_REPLAY_DIR/materialize-idempotency-key.txt"
MATERIALIZE_STATUS_FILE="$MATERIALIZE_REPLAY_DIR/materialize-http-status.txt"
MATERIALIZE_CANDIDATE_FILE="$MATERIALIZE_REPLAY_DIR/candidate-id.txt"
MATERIALIZE_MANIFEST="$MATERIALIZE_REPLAY_DIR/manifest.tsv"
printf '%s' "$MATERIALIZE_KEY" >"$MATERIALIZE_KEY_FILE"
printf 'header = "Idempotency-Key: %s"\n' "$MATERIALIZE_KEY" >"$MATERIALIZE_KEY_CONFIG"
unset MATERIALIZE_KEY
printf '%s' "$CANDIDATE_ID" >"$MATERIALIZE_CANDIDATE_FILE"
{
  printf 'contract_revision\t1\n'
  printf 'api_base_url\t%s\n' "$PGPIPE_URL"
  printf 'api_origin\t%s\n' "$PGPIPE_ORIGIN"
  printf 'actor\t%s\n' "$PGPIPE_AUTHENTICATED_ACTOR"
  printf 'pipeline_id\t%s\n' "$PGPIPE_PIPELINE_ID"
  printf 'materialize_endpoint\t/api/v2/repair/%s/materialize\n' "$CANDIDATE_ID"
} >"$MATERIALIZE_MANIFEST"
chmod 0600 -- "$MATERIALIZE_KEY_FILE" "$MATERIALIZE_KEY_CONFIG" \
  "$MATERIALIZE_CANDIDATE_FILE" "$MATERIALIZE_MANIFEST"
require_uuid_idempotency_pair "$MATERIALIZE_KEY_FILE" "$MATERIALIZE_KEY_CONFIG"
if MATERIALIZE_HTTP_STATUS="$(curl --disable --config "$PGPIPE_AUTH_CONFIG" \
    --config "$MATERIALIZE_KEY_CONFIG" --silent --show-error \
    --cacert "$PGPIPE_CA" -X POST -H "Origin: $PGPIPE_ORIGIN" \
    -o "$MATERIALIZE_RESPONSE" -w '%{http_code}' \
    "$PGPIPE_URL/api/v2/repair/$CANDIDATE_ID/materialize")"; then
  printf '%s' "$MATERIALIZE_HTTP_STATUS" >"$MATERIALIZE_STATUS_FILE"
else
  printf 'transport_error' >"$MATERIALIZE_STATUS_FILE"
  echo "Materialization result is unknown; preserve $MATERIALIZE_REPLAY_DIR and retry exactly." >&2
  exit 1
fi
[ "$MATERIALIZE_HTTP_STATUS" = 201 ] || {
  echo "Materialization was not proven; preserve $MATERIALIZE_REPLAY_DIR and reconcile it." >&2
  exit 1
}
if ! REPAIR_ID="$(materialized_response_job_id "$MATERIALIZE_RESPONSE" "$CANDIDATE_ID")"; then
  echo "Materialization response was malformed; preserve $MATERIALIZE_REPLAY_DIR." >&2
  exit 1
fi
case "$REPAIR_ID" in
  ''|*[!A-Za-z0-9._:-]*) echo 'Server returned an invalid repair ID.' >&2; exit 1 ;;
esac
rm -f -- "$MATERIALIZE_RESPONSE" "$MATERIALIZE_KEY_FILE" \
  "$MATERIALIZE_KEY_CONFIG" "$MATERIALIZE_STATUS_FILE" \
  "$MATERIALIZE_CANDIDATE_FILE" "$MATERIALIZE_MANIFEST"
rmdir -- "$MATERIALIZE_REPLAY_DIR"
unset MATERIALIZE_HTTP_STATUS
```

On a timeout, disconnect, or malformed response, do not create a new key. Keep
the replay directory, re-authenticate, and repeat the exact endpoint/body/key.
Before replaying a preview, run `require_owner_only_replay_bundle`, allow only
its manifest, request, key/config, response, and status filenames through
`require_replay_bundle_files`, then run `require_uuid_idempotency_pair`. Also
require manifest revision `1` and the exact `/api/v2/repair/preview` endpoint
with `require_replay_manifest_value`. For materialization, first inspect the
candidate and job list: if durable detail proves the linked successor, use it
and remove the bundle; otherwise replay the saved candidate and key. For
example:

```bash
MATERIALIZE_REPLAY_DIR='/absolute/path/to/preserved-materialization-replay-directory'
require_owner_only_replay_bundle "$MATERIALIZE_REPLAY_DIR"
require_replay_bundle_files "$MATERIALIZE_REPLAY_DIR" \
  manifest.tsv candidate-id.txt materialize-idempotency-key.txt \
  materialize-idempotency.conf materialize-response.json \
  materialize-http-status.txt
require_uuid_idempotency_pair \
  "$MATERIALIZE_REPLAY_DIR/materialize-idempotency-key.txt" \
  "$MATERIALIZE_REPLAY_DIR/materialize-idempotency.conf"
CANDIDATE_ID="$(<"$MATERIALIZE_REPLAY_DIR/candidate-id.txt")"
case "$CANDIDATE_ID" in
  ''|*[!A-Za-z0-9._:-]*) echo 'Saved candidate ID is invalid.' >&2; exit 1 ;;
esac
require_replay_manifest_value "$MATERIALIZE_REPLAY_DIR" contract_revision 1
require_replay_manifest_value "$MATERIALIZE_REPLAY_DIR" materialize_endpoint \
  "/api/v2/repair/$CANDIDATE_ID/materialize"
if MATERIALIZE_HTTP_STATUS="$(curl --disable --config "$PGPIPE_AUTH_CONFIG" \
    --config "$MATERIALIZE_REPLAY_DIR/materialize-idempotency.conf" \
    --silent --show-error --cacert "$PGPIPE_CA" -X POST \
    -H "Origin: $PGPIPE_ORIGIN" \
    -o "$MATERIALIZE_REPLAY_DIR/materialize-response.json" -w '%{http_code}' \
    "$PGPIPE_URL/api/v2/repair/$CANDIDATE_ID/materialize")"; then
  printf '%s' "$MATERIALIZE_HTTP_STATUS" \
    >"$MATERIALIZE_REPLAY_DIR/materialize-http-status.txt"
else
  printf 'transport_error' \
    >"$MATERIALIZE_REPLAY_DIR/materialize-http-status.txt"
  echo 'Materialization replay remains ambiguous; preserve the bundle.' >&2
  exit 1
fi
if [ "$MATERIALIZE_HTTP_STATUS" = 201 ]; then
  materialized_response_job_id \
    "$MATERIALIZE_REPLAY_DIR/materialize-response.json" "$CANDIDATE_ID" >/dev/null || {
    echo 'Materialization replay returned a malformed success; preserve the bundle.' >&2
    exit 1
  }
fi
```

Delete either replay bundle only after a valid HTTP `201` response or an
authoritative detail/list reconciliation proves its resulting job. Treat the
bundle as sensitive control authority; never attach it to logs or tickets.

Fetch fresh review detail. Inspect the selector, counts, acknowledgement flags,
and `pipeline_id` before proceeding:

```bash
DETAIL="$PGPIPE_PRIVATE_DIR/repair-detail.json"
curl --disable --config "$PGPIPE_AUTH_CONFIG" \
  --fail --silent --show-error --cacert "$PGPIPE_CA" \
  -o "$DETAIL" \
  "$PGPIPE_URL/api/v2/repair/$REPAIR_ID"
jq '{id,phase,pipeline_id,schema,table,preview,selector,
     requires_destructive_acknowledgement,
     requires_delete_acknowledgement,
     requires_large_repair_acknowledgement}' "$DETAIL"
```

Immediately before approval, use the function below. It issues the one-time
token, verifies the typed confirmation, and writes the exact approval request
plus idempotency key to a dedicated owner-only replay directory before the
network submission. The directory must already exist, be empty, and be on
durable incident storage; it is intentionally outside the temporary directory
so an SSH/browser disconnect cannot erase replay authority.

```bash
approve_pgpipe_repair() {
  local replay_dir token_response typed_file expected_file reason_file
  local approval_key approval_request approval_key_file approval_key_config
  local approval_response approval_status_file repair_id_file replay_manifest
  local approval_http_status typed_confirmation approval_reason acknowledgement
  local approve_destructive=false approve_deletes=false approve_large=false

  read -r -p 'Existing empty owner-only replay directory: ' replay_dir
  require_empty_owner_only_replay_dir "$replay_dir" || return 1

  token_response="$PGPIPE_PRIVATE_DIR/approval-token.json"
  if ! curl --disable --config "$PGPIPE_AUTH_CONFIG" \
      --fail --silent --show-error --cacert "$PGPIPE_CA" \
      -X POST -H "Origin: $PGPIPE_ORIGIN" \
      -o "$token_response" \
      "$PGPIPE_URL/api/v2/repair/$REPAIR_ID/approval-token"; then
    echo 'Token issuance failed; no approval was submitted.' >&2
    return 1
  fi
  jq -r '"Type this exact server phrase after reviewing it:\n\(.confirmation)"' \
    "$token_response" || return 1

  read -r -p 'Type the exact confirmation: ' typed_confirmation
  typed_file="$PGPIPE_PRIVATE_DIR/typed-confirmation.txt"
  expected_file="$PGPIPE_PRIVATE_DIR/expected-confirmation.txt"
  printf '%s' "$typed_confirmation" >"$typed_file"
  unset typed_confirmation
  jq -er '.confirmation' "$token_response" | tr -d '\n' >"$expected_file" || return 1
  if ! cmp -s "$typed_file" "$expected_file"; then
    echo 'Confirmation did not match; request a fresh token.' >&2
    return 1
  fi

  read -r -p 'Type DESTINATION WRITES ARE STOPPED: ' acknowledgement
  if [ "$acknowledgement" != 'DESTINATION WRITES ARE STOPPED' ]; then
    echo 'Destination write-freeze acknowledgement was not provided.' >&2
    return 1
  fi
  if jq -e '.requires_destructive_acknowledgement == true' "$DETAIL" >/dev/null; then
    read -r -p 'Type I APPROVE DESTRUCTIVE REPAIR: ' acknowledgement
    [ "$acknowledgement" = 'I APPROVE DESTRUCTIVE REPAIR' ] || return 1
    approve_destructive=true
  fi
  if jq -e '.requires_delete_acknowledgement == true' "$DETAIL" >/dev/null; then
    read -r -p 'Type I APPROVE DELETE ACTIONS: ' acknowledgement
    [ "$acknowledgement" = 'I APPROVE DELETE ACTIONS' ] || return 1
    approve_deletes=true
  fi
  if jq -e '.requires_large_repair_acknowledgement == true' "$DETAIL" >/dev/null; then
    read -r -p 'Type I APPROVE LARGE REPAIR: ' acknowledgement
    [ "$acknowledgement" = 'I APPROVE LARGE REPAIR' ] || return 1
    approve_large=true
  fi
  unset acknowledgement

  read -r -p 'Approval reason/ticket (protected audit data): ' approval_reason
  reason_file="$PGPIPE_PRIVATE_DIR/approval-reason.txt"
  printf '%s' "$approval_reason" >"$reason_file"
  unset approval_reason
  approval_key="$(python3 -c 'import uuid; print(uuid.uuid4())')" || return 1
  approval_request="$replay_dir/approval-request.json"
  approval_key_file="$replay_dir/approval-idempotency-key.txt"
  approval_key_config="$replay_dir/approval-idempotency.conf"
  approval_response="$replay_dir/approval-response.json"
  approval_status_file="$replay_dir/approval-http-status.txt"
  repair_id_file="$replay_dir/repair-id.txt"
  replay_manifest="$replay_dir/manifest.tsv"

  jq -n \
    --rawfile reason "$reason_file" \
    --slurpfile detail "$DETAIL" \
    --slurpfile authority "$token_response" \
    --rawfile confirmation "$typed_file" \
    --argjson destructive "$approve_destructive" \
    --argjson deletes "$approve_deletes" \
    --argjson large "$approve_large" \
    '($detail[0]) as $detail | ($authority[0]) as $authority | {
      reason:$reason,
      expected_pipeline_id:$detail.pipeline_id,
      expected_approval_token:$authority.approval_token,
      expected_approval_token_revision:$authority.approval_token_revision,
      expected_preview_version:$authority.preview_version,
      approve_destructive:$destructive,
      approve_deletes:$deletes,
      approve_large_repair:$large,
      destination_writes_stopped:true,
      confirmation:$confirmation
    }' >"$approval_request" || return 1
  printf '%s' "$approval_key" >"$approval_key_file"
  printf 'header = "Idempotency-Key: %s"\n' "$approval_key" >"$approval_key_config"
  printf '%s' "$REPAIR_ID" >"$repair_id_file"
  {
    printf 'contract_revision\t2\n'
    printf 'api_base_url\t%s\n' "$PGPIPE_URL"
    printf 'api_origin\t%s\n' "$PGPIPE_ORIGIN"
    printf 'actor\t%s\n' "$PGPIPE_AUTHENTICATED_ACTOR"
    printf 'pipeline_id\t%s\n' "$PGPIPE_PIPELINE_ID"
    printf 'approval_endpoint\t/api/v2/repair/%s/approve\n' "$REPAIR_ID"
  } >"$replay_manifest"
  chmod 0600 -- "$approval_request" "$approval_key_file" "$approval_key_config" \
    "$repair_id_file" "$replay_manifest" || return 1
  require_uuid_idempotency_pair "$approval_key_file" "$approval_key_config" || return 1
  rm -f -- "$token_response" "$typed_file" "$expected_file" "$reason_file"

  if ! approval_http_status="$(curl --disable --config "$PGPIPE_AUTH_CONFIG" \
      --config "$approval_key_config" \
      --silent --show-error --cacert "$PGPIPE_CA" \
      -o "$approval_response" -w '%{http_code}' \
      -H "Origin: $PGPIPE_ORIGIN" \
      -H 'Content-Type: application/json' \
      --data-binary "@$approval_request" \
      "$PGPIPE_URL/api/v2/repair/$REPAIR_ID/approve")"; then
    printf '%s' "${approval_http_status:-transport_error}" >"$approval_status_file"
    echo "Approval result is unknown; preserve $replay_dir and reconcile detail." >&2
    return 1
  fi
  printf '%s' "$approval_http_status" >"$approval_status_file"
  if [ "$approval_http_status" = 202 ] &&
     jq -e --arg id "$REPAIR_ID" \
       '.accepted == true and .id == $id and .phase == "approved_queued"' \
       "$approval_response" >/dev/null; then
    rm -f -- "$approval_request" "$approval_key_file" "$approval_key_config" \
      "$approval_response" "$approval_status_file" "$repair_id_file" "$replay_manifest"
    rmdir -- "$replay_dir" || return 1
    echo 'Repair admitted as approved_queued.'
    return 0
  fi
  echo "Approval was not proven; preserve $replay_dir and reconcile detail." >&2
  return 1
}

approve_pgpipe_repair
```

On success, the server returns HTTP `202`, `accepted: true`, the same job ID,
and phase `approved_queued`; only then does the function remove the replay
bundle. On timeout, disconnect, non-202, or an invalid body, first inspect job
detail. If it already proves `approved_queued`, claimed, or a later phase,
monitor that job. Otherwise retry the exact saved endpoint, body, and
idempotency key; do not wait for a GET to prove that an in-flight request did
*not* happen. The server's idempotency/CAS boundary must resolve the exact
replay. A definitive structured stale/expired-token rejection means the old
authority can be discarded and a fresh preview/token/key created. An ambiguous
transport or malformed response does not: preserve and retry the exact bundle.
For example, after re-authenticating:

```bash
REPLAY_DIR='/absolute/path/to/the-preserved-replay-directory'
require_owner_only_replay_bundle "$REPLAY_DIR"
require_replay_bundle_files "$REPLAY_DIR" manifest.tsv repair-id.txt \
  approval-idempotency-key.txt approval-idempotency.conf approval-request.json \
  approval-response.json approval-http-status.txt
require_uuid_idempotency_pair "$REPLAY_DIR/approval-idempotency-key.txt" \
  "$REPLAY_DIR/approval-idempotency.conf"
REPAIR_ID="$(<"$REPLAY_DIR/repair-id.txt")"
case "$REPAIR_ID" in
  ''|*[!A-Za-z0-9._:-]*) echo 'Saved repair ID is invalid.' >&2; exit 1 ;;
esac
require_replay_manifest_value "$REPLAY_DIR" contract_revision 2
require_replay_manifest_value "$REPLAY_DIR" approval_endpoint \
  "/api/v2/repair/$REPAIR_ID/approve"
if approval_http_status="$(curl --disable --config "$PGPIPE_AUTH_CONFIG" \
    --config "$REPLAY_DIR/approval-idempotency.conf" \
    --silent --show-error --cacert "$PGPIPE_CA" \
    -o "$REPLAY_DIR/approval-response.json" -w '%{http_code}' \
    -H "Origin: $PGPIPE_ORIGIN" -H 'Content-Type: application/json' \
    --data-binary "@$REPLAY_DIR/approval-request.json" \
    "$PGPIPE_URL/api/v2/repair/$REPAIR_ID/approve")"; then
  printf '%s' "$approval_http_status" \
    >"$REPLAY_DIR/approval-http-status.txt"
else
  printf 'transport_error' >"$REPLAY_DIR/approval-http-status.txt"
  echo 'Approval replay remains ambiguous; preserve the bundle.' >&2
  exit 1
fi
if [ "$approval_http_status" = 202 ]; then
  jq -e --arg id "$REPAIR_ID" \
    '.accepted == true and .id == $id and .phase == "approved_queued"' \
    "$REPLAY_DIR/approval-response.json" >/dev/null || {
    echo 'Approval replay returned a malformed success; preserve the bundle.' >&2
    exit 1
  }
fi
```

Validate `202`, `accepted:true`, the saved repair ID, and
`phase:"approved_queued"` before deleting the bundle. Delete it only after
accepted admission or documented incident reconciliation is proven.
Treat the entire replay directory as highly sensitive plaintext approval
authority: restrict it to the operator, never attach it to tickets or logs,
and retain it only for the shortest reconciliation window.

```bash
curl --disable --config "$PGPIPE_AUTH_CONFIG" \
  --fail --silent --show-error --cacert "$PGPIPE_CA" \
  "$PGPIPE_URL/api/v2/repair/$REPAIR_ID" | jq .
```

To cancel a preview or queued/unclaimed job, provide a required reason and its
own idempotency key. A `409 repair_cancel_conflict` means the worker already won
the claim race; monitor it instead. Cancellation has the same ambiguous-network
rule as approval, so preserve its exact request/key outside the temporary
session until HTTP `200` plus the expected job and `canceled_before_claim` body
is proven:

```bash
cancel_pgpipe_repair() {
  local replay_dir reason_file cancel_request cancel_key cancel_key_file cancel_key_config
  local cancel_response cancel_status_file cancel_http_status

  read -r -p 'Existing empty owner-only cancellation replay directory: ' replay_dir
  require_empty_owner_only_replay_dir "$replay_dir" || return 1

  reason_file="$PGPIPE_PRIVATE_DIR/cancel-reason.txt"
  cancel_request="$replay_dir/cancel-request.json"
  cancel_key_file="$replay_dir/cancel-idempotency-key.txt"
  cancel_key_config="$replay_dir/cancel-idempotency.conf"
  cancel_response="$replay_dir/cancel-response.json"
  cancel_status_file="$replay_dir/cancel-http-status.txt"
  cancel_key="$(python3 -c 'import uuid; print(uuid.uuid4())')" || return 1
  read -r -p 'Cancellation reason/ticket (protected audit data): ' CANCEL_REASON
  printf '%s' "$CANCEL_REASON" >"$reason_file"
  unset CANCEL_REASON
  jq -n --rawfile reason "$reason_file" '{reason:$reason}' \
    >"$cancel_request" || return 1
  printf '%s' "$cancel_key" >"$cancel_key_file"
  printf 'header = "Idempotency-Key: %s"\n' "$cancel_key" >"$cancel_key_config"
  printf '%s' "$REPAIR_ID" >"$replay_dir/repair-id.txt"
  {
    printf 'contract_revision\t2\n'
    printf 'api_base_url\t%s\n' "$PGPIPE_URL"
    printf 'api_origin\t%s\n' "$PGPIPE_ORIGIN"
    printf 'actor\t%s\n' "$PGPIPE_AUTHENTICATED_ACTOR"
    printf 'pipeline_id\t%s\n' "$PGPIPE_PIPELINE_ID"
    printf 'cancel_endpoint\t/api/v2/repair/%s/cancel\n' "$REPAIR_ID"
  } >"$replay_dir/manifest.tsv"
  chmod 0600 -- "$cancel_request" "$cancel_key_file" "$cancel_key_config" \
    "$replay_dir/repair-id.txt" "$replay_dir/manifest.tsv" || return 1
  require_uuid_idempotency_pair "$cancel_key_file" "$cancel_key_config" || return 1
  rm -f -- "$reason_file"
  unset cancel_key

  if ! cancel_http_status="$(curl --disable --config "$PGPIPE_AUTH_CONFIG" \
      --config "$cancel_key_config" --silent --show-error --cacert "$PGPIPE_CA" \
      -o "$cancel_response" -w '%{http_code}' \
      -H "Origin: $PGPIPE_ORIGIN" -H 'Content-Type: application/json' \
      --data-binary "@$cancel_request" \
      "$PGPIPE_URL/api/v2/repair/$REPAIR_ID/cancel")"; then
    printf '%s' transport_error >"$cancel_status_file"
    echo "Cancellation is ambiguous; preserve and exactly retry $replay_dir." >&2
    return 1
  fi
  printf '%s' "$cancel_http_status" >"$cancel_status_file"
  if [ "$cancel_http_status" = 200 ] &&
     jq -e --arg id "$REPAIR_ID" \
       '.id == $id and .phase == "canceled_before_claim"' \
       "$cancel_response" >/dev/null; then
    rm -f -- "$cancel_request" "$cancel_key_file" "$cancel_key_config" "$cancel_response" \
      "$cancel_status_file" "$replay_dir/repair-id.txt" "$replay_dir/manifest.tsv"
    rmdir -- "$replay_dir" || return 1
    echo 'Repair canceled before claim.'
    return 0
  fi
  echo "Cancellation was not proven; preserve $replay_dir, inspect detail, and exactly retry on ambiguity." >&2
  return 1
}

cancel_pgpipe_repair
```

For an ambiguous result, re-authenticate, set the same API URL/origin, run
`require_owner_only_replay_bundle` on the saved directory, validate its
exact allowed filenames with `require_replay_bundle_files`, and run
`require_uuid_idempotency_pair` on `cancel-idempotency-key.txt` and
`cancel-idempotency.conf` before replay. Require manifest revision `2` and an
exact `/api/v2/repair/<saved-id>/cancel` endpoint with
`require_replay_manifest_value`; never generate a new reason or key.
For a definitive structured `repair_cancel_conflict`, inspect detail and retain
the bundle until the claimed-or-later phase is documented.

Clear only session-local detail files. Any ambiguous preview, materialization,
approval, or cancellation bundle remains in its explicitly selected owner-only
directory until reconciled:

```bash
rm -f -- "$DETAIL"
```

## Monitor phases and progress

The normal durable sequence is:

```text
previewed -> approved_queued -> claimed -> cutover_captured -> copying
          -> replaying -> verifying -> completed
```

Use the dashboard, `GET /api/v2/repair/{id}`, Prometheus, service logs, and
operator audit together. Progress explicitly reports its evidence source:

- `live_backfill` is a current in-process sample;
- `durable_table_state` is persisted restart/recovery evidence;
- `repair_queue` means admitted but not yet claimed;
- `repair_job_history` is historical durable evidence, not a live counter.

Never interpret a stale historical row count as current execution progress.

### Terminal outcomes

| Result | Meaning | Operator action |
| --- | --- | --- |
| `completed`, final verification `matched` | Approved destination selection matches the source proof. | Review audit and metrics, release write freeze, close maintenance record. |
| `completed`, final verification `inconclusive` | Writes completed, but the final read could not prove an audit-quality match, commonly because source rows changed concurrently. | Keep evidence; freeze source writes if possible and use **Rerun verification**. This does not rerun repair. |
| `expired_before_claim` | Queue deadline passed before a worker claimed it. | Create and review a fresh preview if repair is still required. |
| `canceled_before_claim` | Operator canceled before claim; no active worker was interrupted. | Review the cancellation audit. Create a fresh preview for later work. |
| `interrupted_prewrite` with cleanup complete | Zero destination writes were positively proven and fenced cleanup completed. | Review evidence. The old job remains terminal; create a fresh preview if needed. |
| `cleanup_pending` or cleanup failed | A no-write interruption is still guarded while artifacts are reconciled/removed. | Keep pgpipe controlled and follow sanitized startup guidance. Do not delete artifacts. |
| `interrupted_partial`, `interrupted_ambiguous`, or `interrupted_postwrite` | Destination effects are partial, uncertain, or not verified. | Keep writes frozen, preserve all evidence, and perform offline recovery. Never retry the old job. |
| `recovery_required` | Startup/streaming remains blocked by durable safety evidence. | Escalate to the incident owner and database administrator; follow the recovery procedure below. |

For an inconclusive completed job, the authenticated mutating endpoint
`POST /api/v2/repair/{id}/rerun-verification` starts one fresh diagnostic table
verification and returns `202`. It never retries or changes the terminal repair.

```bash
curl --disable --config "$PGPIPE_AUTH_CONFIG" \
  --fail --silent --show-error --cacert "$PGPIPE_CA" \
  -X POST \
  -H "Origin: $PGPIPE_ORIGIN" \
  "$PGPIPE_URL/api/v2/repair/$REPAIR_ID/rerun-verification" | jq .
```

## Restart and recovery procedure

An ordinary restart before any job is admitted is safe. After durable admission:

- an unclaimed queued job may be claimed only while its queue deadline and all
  readiness checks remain valid;
- a claimed job is never automatically resumed after restart in the first
  release;
- state, exact local artifacts, destination fence, and chunk ledger are
  reconciled before any normal streaming or new repair can proceed;
- only positively proven pre-write cleanup may run automatically, under its
  separate no-new-writes timeout;
- partial, ambiguous, post-write, stale, foreign, or contradictory evidence
  remains blocked.

When recovery is required:

1. Stop pgpipe. Keep source and destination application writes frozen for the
   affected table/range.
2. Preserve the exact config, pgpipe binary/version, service logs, state
   backend, local repair artifacts, source WAL/slot evidence, destination
   database, alerts, audit events, and internal repair fence/ledger rows.
3. Do not run a new repair, edit a terminal job, delete a fence/ledger row, or
   repoint the same state database at another destination.
4. Choose one offline recovery with the incident owner and DBA:
   - rebuild the affected destination data from the authoritative source under
     a new, clean pipeline lifecycle; or
   - restore destination and pgpipe state from the same coordinated backup/PITR
     point.
5. Run config validation, destination/state fence inspection, startup
   preflight, and content verification before restoring application writes.
6. Retain the interrupted job and its audit/ledger evidence as immutable
   incident history.

The exact command depends on the pipeline's backup system and affected table.
There is intentionally no generic `force`, `retry interrupted repair`, or
`delete evidence` command.

## Metrics and alerts

Scrape the configured Prometheus endpoint. In the first production rollout,
monitor these bounded repair series:

| Metric | Use |
| --- | --- |
| `pgpipe_repair_jobs_total{result}` | Durable repair outcomes by bounded result. Alert on any recovery-required or unexpected interruption result. |
| `pgpipe_repair_approval_attempts_total{result}` | Accepted and rejected approval attempts. Investigate sustained rejection or rate-limit patterns. |
| `pgpipe_repair_active` | Current claimed execution gauge. Alert if it remains active near the configured whole-execution deadline. |
| `pgpipe_repair_duration_seconds` | Execution-duration histogram. Compare high percentiles with `repair.max_execution_duration`. |
| `pgpipe_repair_actions_total{operation}` | Committed insert/update/delete action counts. Reconcile with reviewed preview and destination chunk evidence. |
| `pgpipe_repair_preview_drift_total` | Preview invalidations caused by changed proof/config/relation identity. Investigate repeated drift before approving again. |
| `pgpipe_repair_final_verification_total{status}` | Matched versus inconclusive final verification. Alert on inconclusive results. |
| `pgpipe_repair_control_rate_limited_total{operation}` | Rate-limit rejections for preview, materialization, token issuance, approval, cancellation, and verification rerun. Correlate with `Retry-After` and audit/auth activity. |

Metrics are process observability, not the durable authority record. A process
failure can occur between a database commit and an in-memory counter update.
Use the job state, destination chunk ledger, and atomic operator audit as the
correctness evidence.

At minimum, route the critical durable alert type
`repair_recovery_required` and warning type
`repair_verification_inconclusive` to an attended channel. Validated alerts use
the job ID as their stable fingerprint; malformed durable evidence emits only
generic bounded context. Also alert when:

- `pgpipe_repair_active` remains `1` close to the whole-execution limit;
- approval rejections or preview drift increase unexpectedly;
- any partial, ambiguous, post-write, or cleanup-failed outcome appears, or the
  audit reports queue expiry;
- final verification is not matched;
- `/api/v2/repair/capabilities` leaves `ready` during a maintenance window.

Use `GET /api/alerts` and the dashboard alert view for durable alert status.
Resolving an alert is only an operator acknowledgement; it does not change a
repair job, release a destination fence, or prove recovery.

## Audit review

Approval admission and its audit event are atomic. Audit records use the
server-authenticated actor and required reason; they do not store raw row
images, private proofs, or plaintext approval tokens. Review the pipeline-
scoped repair events through the authenticated endpoint:

```bash
curl --disable --config "$PGPIPE_AUTH_CONFIG" \
  --fail --silent --show-error --cacert "$PGPIPE_CA" \
  "$PGPIPE_URL/api/ops/audit?type=repair&command_id=$REPAIR_ID&limit=500" | jq .
```

For each production job, retain or forward the bounded events needed to prove:

- candidate materialization or manual preview provenance;
- token issuance/rotation without plaintext token material;
- approval actor, reason, idempotent admission, and selected job ID;
- claim and execution phase transitions;
- committed action counts and final verification result;
- cancellation, queue expiry, timeout, interruption, cleanup, or recovery
  classification when applicable.

Compare audit continuity with the job detail and destination ledger evidence.
Missing or contradictory durable evidence is a blocker, not permission to infer
success. Limit audit access because actor identities and reasons are protected
operational data.

When monitoring and audit review are complete, remove the temporary bearer
config and clear the remaining shell state:

```bash
if cleanup_pgpipe_repair_session; then
  trap - EXIT
  unset PGPIPE_AUTH_CONFIG PGPIPE_URL PGPIPE_ORIGIN PGPIPE_CA REPAIR_ID
  unset PGPIPE_AUTHENTICATED_ACTOR PGPIPE_PIPELINE_ID PGPIPE_STATUS
  unset PREVIEW_KEY PREVIEW_REQUEST PREVIEW_RESPONSE PREVIEW_KEY_FILE PREVIEW_KEY_CONFIG
  unset PREVIEW_REPLAY_DIR PREVIEW_STATUS_FILE PREVIEW_MANIFEST PREVIEW_HTTP_STATUS
  unset MATERIALIZE_KEY MATERIALIZE_RESPONSE MATERIALIZE_KEY_FILE MATERIALIZE_KEY_CONFIG CANDIDATE_ID
  unset MATERIALIZE_REPLAY_DIR MATERIALIZE_STATUS_FILE MATERIALIZE_CANDIDATE_FILE
  unset MATERIALIZE_MANIFEST MATERIALIZE_HTTP_STATUS
  unset DETAIL approval_http_status REPLAY_DIR
  unset -f approve_pgpipe_repair cancel_pgpipe_repair \
    require_empty_owner_only_replay_dir require_owner_only_replay_bundle \
    require_replay_bundle_files require_replay_manifest_value \
    require_uuid_idempotency_pair \
    preview_response_job_id materialized_response_job_id cleanup_pgpipe_repair_session
else
  echo 'Private session cleanup was refused; the EXIT trap remains active.' >&2
fi
```

The readonly temporary root/directory path variables are non-secret after the
validated cleanup and remain only until this deliberately short-lived shell
exits. Do not continue ordinary administration in the secret-handling shell.

## Rollout stages and flag-removal criteria

Use a deliberate progression for the exact release artifact:

1. **Disabled production deployment:** ship with
   `repair.execution_enabled: false`; verify that streaming and Verify & Repair
   remain unaffected.
2. **Demo/integration:** enable only after Phases 1-5 pass; exercise successful,
   inconclusive, cancellation, restart, timeout, and injected partial/ambiguous
   failures on disposable databases.
3. **Staging soak:** run repeated destructive and failure-injection trials with
   the production topology, PostgreSQL versions, state backend, role model,
   proxy/auth configuration, and monitoring stack.
4. **Limited production:** select one low-risk pipeline and maintenance window;
   require backup/PITR proof, write freeze, named approver, live monitoring, and
   post-repair audit review.
5. **Broader opt-in:** expand only after security, correctness, UX, operations,
   and performance owners sign off on the accumulated evidence.

Do not enable by default or remove the flag until all of these are true:

- repeated soak shows no duplicate claim/execution, audit gap, stale-token
  admission, cross-pipeline access, or unsafe restart outcome;
- failure injection proves correct pre-write cleanup and conservative
  partial/ambiguous/post-write recovery classification;
- final performance certification of the exact build shows no unacceptable
  steady-state replication regression and quantifies active-repair impact;
- the runbook and backup/PITR recovery have been exercised by operators who did
  not build the feature;
- metrics, alerts, audit retention, privilege preflight, and support escalation
  are operating in production;
- security, correctness, UX, operations, and performance reviews all approve
  default enablement;
- a major-release migration/deprecation plan exists for removing the temporary
  capability without silently enabling execution on older configurations.

Until then, `execution_enabled: false` remains the safe production default.
