Skip to content

FEAT-024 Phase 2C: Staging Proof Runbook

Exact revision ledger ceilings

The detailed revision ledger admits at most 100,000 retained rows per project and no new materialized write while any retained row is at least 90 days old. These are provisional ProjectStatisticsLifecycleLimits, not measured production limits. Admission is checked only for a point write or replacement publication that would insert a revision row. The row count stops at the ceiling even if an older ledger is already oversized; the age check uses the IX_StatsRevision_ProjectObservedAt index. Deploy that index and this writer version on every API and project-management mutation host before treating the ceiling as a hard bound during a rolling deployment. Older writers do not enforce it.

Under pressure, an admitted source mutation uses the existing source-only route: its source change, minimal operation receipt, clocks, and required notification slot still commit, while affected statistics serve authoritative calculation. A rebuild refuses publication with DeltaLedgerCapacityExceeded and releases its candidate. Neither path deletes a pinned row or advances the replay or receipt floor. Run only the supported bounded delta-maintenance trigger when its independent gate is enabled and safety watermarks allow compaction; then use the ordinary family backfill/rebuild to return Stale rows to Fresh. A pin that prevents safe compaction must be resolved through its owning workflow, never by direct ledger deletion.

The existing materialized Writes flag and project allowlist remain the rollout gates; no new flag or automatic cleanup is introduced. Bounded-growth/soak and per-write query latency still require measurement before broader activation; these tests do not establish the HTTP or mutation-overhead acceptance targets. On an uncontended materialized point write, admission adds two MongoDB commands: a count capped at the row ceiling and an indexed, limit-one age-existence query. The command-budget regression holds the pre-existing path to 23 commands and these probes to two more; that count is not a latency or less-than-10-percent write-overhead measurement.

Purpose

Phase 2's staging proof is to "dark-write an allowlisted corpus, run shadow parity and load, force Stale/Rebuilding/Incompatible and kill-switch paths, and prove legacy response equivalence" (technical plan, Phase 2). This runbook is the operational half of that: which gates must close first, what to change, what to run, what each response means, and what evidence to keep.

It covers the project screening family only, on one staging project, with no page, SignalR or export consumer. Phase 2C is explicitly "no ordinary page or SignalR cutover", so the only reader of a materialized value during this proof is the administrator-only parity audit.

What this runbook is not

  • It is not an approval. Running it requires the gates in Preconditions to be closed, including an explicit go from the FEAT-024 programme owner.
  • It is not a production procedure. Nothing here may be pointed at production, and the soak evidence it produces is an input to a later separate production-pilot decision, not that decision.
  • It does not authorize a consumer cutover, a second project, a second family, or any legacy retirement.

Preconditions and gates

# Gate Why
1 syrf #3196 merged The parity audit endpoint lives on this pull request. As of 38a272d4c main has only backfill and rebuild; without #3196 there is no way to compare the projection against the authoritative calculation, which is the entire point of the proof.
2 syrf #3232 merged Binds capacity writes to durable-mode transactions.
2a syrf #3371 merged The administrative mode-transition surface (issue #3369). Without it no production code opens the fleet control or a project's narrow gate, so a completed backfill writes rows that can never serve and the parity audit reports Disabled for every scope — steps 3, 5 and 10 are unperformable and the proof establishes nothing.
3 syrf issue #3185 closed The README's pilot gate: until its single-transaction settings write lands, a Project.AgreementThreshold write can commit between a rebuild's pinned snapshot and its publication without advancing the control's source revision, so a freshly rebuilt row can fail the reader's row/control equality. The pilot must not be activated for any project before this closes.
4 Pilot project chosen and its GUID recorded See Choosing the pilot project.
5 Explicit go from the FEAT-024 programme owner The README holds every flag off and the allowlist empty "until the separately authorized single-project staging activation". This runbook does not grant that authorization.
6 An administrator account on staging — or, for the step-6 parity read only, the read-only statistics-parity evidence credential Every mutating route (backfill, rebuild, mode, narrow gate, fleet, maintenance) sits behind ApplicationAuthorization.BatchAdminProjectsPolicy, whose activity BatchAdminProjects is granted to the administrator application group only. The parity read (GET …/{projectId}/parity) sits behind ProjectStatisticsParityReadPolicy, which admits exactly the same administrators or a client-credentials token carrying only statistics:parity:read. An authenticated non-administrator receives 403; an anonymous caller receives 401. The evidence credential is off by default and cannot perform any step except step 6. The statistics operator identity may instead run the pending-index, fold, backfill/rebuild and parity operations (not mode, narrow gate, fleet or maintenance); it is also off by default.

Confirm the deployed staging API and project-management images actually contain #3196 and #3232 before starting. A promoted chartTag in cluster-gitops is the intent; the running pod's image digest is the fact.

Step 1 — enable the pilot in cluster-gitops

The change is prepared as a draft pull request against camaradesuk/cluster-gitops, held closed behind the gates above. It touches two files:

File Change
syrf/environments/staging/api/values.yaml statistics flags into the existing featureFlags: map; allowlist into the existing env: map
syrf/environments/staging/project-management/values.yaml the same block

Both hosts are required. The API is the serving side and hosts the administrative endpoints; project-management is the source-transaction writing side. Enabling one without the other produces a projection that is either written and never read or read and never maintained.

Flag values

Three on, nine off:

Flag Value Note
materializedProjectStatisticsWrites true Global write kill switch.
materializedProjectStatisticsServing true Global serving kill switch; the runtime catalog also requires the write gate.
materializedProjectStatisticsScreening true The project screening family.
materializedProjectStatisticsMembershipScreening false A separate Phase 4 family, not part of the screening family.
materializedProjectStatisticsAnnotation false
materializedProjectStatisticsMembershipAnnotation false
materializedProjectStatisticsQuestionAnswers false
materializedProjectStatisticsSearchPopulation false
materializedProjectStatisticsDerivedSummaries false
materializedProjectStatisticsFold false Asynchronous point fold (#3863); deployment-managed and stays off for this proof. The separate ProjectScreening fold pilot turns it on: see the fold staging pilot.
materializedProjectStatisticsPages false Consumer gate — stays off for the whole proof.
materializedProjectStatisticsSignalR false Consumer gate — stays off for the whole proof.
materializedProjectStatisticsExports false Consumer gate — stays off for the whole proof.

ProjectStatisticsFlagMap.FamilyFlagKey maps ProjectStatisticsMetricFamily.ProjectScreening to materializedProjectStatisticsScreening and to nothing else, so the screening proof needs no second family gate. A family with no gate is denied rather than defaulted on.

Allowlist

ProjectStatistics:ProjectAllowlist is deployment configuration, not a generated feature flag. ProjectStatisticsAllowlistConfiguration reads it through IConfiguration, so a non-production administrator cannot widen it from the runtime flag admin UI — which matters, because it is the thing that decides which real projects a dark projection may serve.

It has no env-mapping.yaml entry, so it is set through the chart's generic .Values.env map, the same mechanism ASPNETCORE_ENVIRONMENT already uses:

env:
  SYRF__ProjectStatistics__ProjectAllowlist: "<STAGING_PILOT_PROJECT_ID>"

_deployment-dotnet.tpl copies .Values.env verbatim into the container environment, and the hosts call AddEnvironmentVariables("SYRF__") last, which strips the prefix and turns __ into :, yielding exactly ProjectStatistics:ProjectAllowlist.

The value is a comma-separated list of project GUIDs. There is no wildcard and no permissive failure: absent, empty and malformed all admit nothing. The draft therefore ships the literal placeholder <STAGING_PILOT_PROJECT_ID>, which cannot enable any project even if merged by mistake.

Verifying it rendered — read-only

After ArgoCD syncs, confirm the running pods carry the values. The staging preflight workflow performs these reads and comparisons with a least-privilege identity and records the result; the commands below are the manual equivalent. All three are reads.

# The rendered environment on each host. Expect the three true flags, the nine false ones,
# and the allowlist GUID.
kubectl -n syrf-staging get deploy syrf-api -o json \
  | jq -r '.spec.template.spec.containers[0].env[]
           | select(.name | test("MaterializedProjectStatistics|ProjectAllowlist"))
           | "\(.name)=\(.value)"'

kubectl -n syrf-staging get deploy syrf-projectmanagement -o json \
  | jq -r '.spec.template.spec.containers[0].env[]
           | select(.name | test("MaterializedProjectStatistics|ProjectAllowlist"))
           | "\(.name)=\(.value)"'

# Prove the pods were actually replaced, rather than the Deployment spec having moved ahead
# of a stuck rollout.
kubectl -n syrf-staging get pods -l app=api-staging-syrf-api \
  -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.containerStatuses[0].imageID}{"\n"}{end}'

kubectl exec ... env is an acceptable alternative; it is still a read. Do not use kubectl set env, kubectl apply, kubectl edit or helm upgrade to correct anything found here — a discrepancy is a cluster-gitops change, per the repository's GitOps-only policy.

The clearest functional confirmation costs nothing and comes in step 4: a project outside the allowlist answers the backfill endpoint with 409 and failureReason: "NotAllowlisted". If the pilot project answers that way, the allowlist did not reach the API host — a 404 there means something else, namely that the project document does not exist.

Staging preflight

Before any credentialed step, capture an automated preflight. It needs no personal login and no administrator account. Dispatch the Statistics Staging Preflight workflow (.github/workflows/statistics-staging-preflight.yml, optional gitops_ref and project_ids). It has two jobs:

  • preflight (main only) reads the HTTP surfaces, cluster-gitops and the running cluster;
  • parity (main only, environment statistics-evidence) captures one parity report per allowlisted project with the read-only evidence credential.

The HTTP and GitOps half also runs locally:

# GitOps through the GitHub contents API (token needs read access to cluster-gitops):
GH_TOKEN=... python3 scripts/capture-statistics-preflight.py --gitops-ref main
# Without a token, either skip GitOps or read a local copy (recorded as unpinned):
python3 scripts/capture-statistics-preflight.py --skip-gitops
python3 scripts/capture-statistics-preflight.py --gitops-dir <cluster-gitops>/syrf/environments/staging
# With saved kubectl output (any identity with get on Deployments and pods):
kubectl -n syrf-staging get deploy syrf-api syrf-projectmanagement syrf-quartz -o json > deploy.json
kubectl -n syrf-staging get pods -l 'app in (api-staging-syrf-api,project-management-staging-syrf-projectmanagement,quartz-staging-syrf-quartz)' -o json > pods.json
GH_TOKEN=... python3 scripts/capture-statistics-preflight.py --gitops-ref main \
  --cluster-deployments deploy.json --cluster-pods pods.json

The preflight job reads, and only reads:

  • anonymous /health/live from the API, project-management and Quartz staging hosts;
  • anonymous GET /api/runtime-feature-flags from the API host: the revision and each flag's static, override and effective values;
  • syrf/environments/staging/{api,project-management,quartz,web}/{config,values}.yaml from cluster-gitops at one resolved commit. Only image and chart tags, replica counts, SharedReaderMode, materializedProjectStatistics* flags and SYRF__ProjectStatistics* / SYRF__FeatureFlags__MaterializedProjectStatistics* settings are kept. Setting names that look like credentials are redacted, and no other value is copied;
  • kubectl -n syrf-staging get deploy syrf-api syrf-projectmanagement syrf-quartz -o json and get pods for those Deployments' selectors, as the GCP service account syrf-stats-preflight@camarades-net.iam.gserviceaccount.com. That identity has roles/container.clusterViewer and the built-in view ClusterRole in syrf-staging only (view excludes Secrets), and GitHub OIDC can impersonate it only from repo:camaradesuk/syrf:ref:refs/heads/main (camarades-infrastructure#16, cluster-gitops#1390, syrf #3806). From that output it keeps each Deployment's image, rollout counters and the same whitelisted FEAT-024 environment, and each pod's image, imageID digest, readiness and whether its FEAT-024 environment equals the Deployment template. A value supplied through valueFrom.secretKeyRef is recorded as <from secret>: it is never resolved, and it is reported as "not captured" rather than compared. The raw kubectl JSON stays in the runner's temp directory and is deleted; it is not uploaded.

It writes one dated JSON file labelled kind: "preflight" and parity: "not captured", with UTC timestamps, the tool version and every source URL and commit, then compares:

Check Failure behaviour
Live gitInformationalVersion equals the desired image tag, per service exit 1
Each deployment-managed statistics flag (runtimeManageable: false, or the field missing) has static == effective == API GitOps value exit 1
API and project-management GitOps agree on each deployment-managed flag exit 1
Each runtime-manageable statistics flag's static value equals the API GitOps value exit 1
A whitelisted setting name appears in GitOps text but the parser could not read it (for example a list-shaped env entry) exit 1; its lifecycle job reads unknown (unparsed)
cluster-image:<service>: the Deployment's image tag equals the GitOps image tag exit 1
cluster-pods:<service>: every matching pod is Running and ready, runs the Deployment's image, reports one shared sha256 digest, and carries the Deployment template's FEAT-024 environment (a stuck or half-finished rollout fails here) exit 1
cluster-flag:<service>:<flag>: on the API and the writer host (project-management), each deployed SYRF__FeatureFlags__MaterializedProjectStatistics* value equals that host's GitOps value (absent reads as off) exit 1
cluster-setting:<service>:<name>: each deployed SYRF__ProjectStatistics* setting and SharedReaderMode equals that host's GitOps value, including ProjectAllowlist (order and case ignored) and ProjectStatistics{DailyObservations,Repair,Maintenance,DriftCheck}:Enabled exit 1
cluster-allowlist-agreement: the API and project-management Deployments carry the same allowlist exit 1
A Deployment or pod could not be read (cluster-read) exit 1; the other comparisons still run
A value comes from a Secret reported as not captured plus a warning; never a pass
ProjectStatisticsRepair:Enabled is not on in the deployed writer environment warning, exit 0 (cluster-gitops#1397 turns it on; the setting check above still fails if GitOps and the Deployment disagree)
A rollout is not settled (generation not observed, or updated/ready replicas short) warning, exit 0
A runtime override makes an effective statistics flag differ from GitOps warning, exit 0; listed

Exit 2 means an HTTP or GitOps source could not be read. The capture is still written, with an error and no comparison.

The job summary shows a cluster table (Deployment image, pod digests, ready pods) and the writer host's deployed FEAT-024 environment: the flags that are on, the allowlist, SharedReaderMode, the four lifecycle switches and any Secret-sourced names. This is the read-only form of the kubectl checks in Verifying it rendered; record its JSON as the "deployed flag values and allowlist as read from the running pods" and "deployed image digests" of step 8.

The parity job then requests one statistics:parity:read token and GETs the parity report for each project in the running API Deployment's allowlist (an empty allowlist counts as absent, so it falls back to the GitOps API allowlist; a project_ids input overrides both). When both allowlists are empty the parity job is skipped and the preflight summary says so; that is not a failure. An inconclusive report is re-requested once, per the retry rule, and every attempt is kept. It uploads two files as the statistics-parity-<run> artifact:

  • statistics-parity_staging_<stamp>.jsonl: every 200 answer in the parity JSONL envelope, readable unchanged by validate-statistics-evidence.py --parity, so each run is one soak parity observation;
  • statistics-parity_staging_<stamp>.json: per attempt, inParity, isInconclusive, readSource, fallbackReason, capacityFailure, configurationMatchesCurrentSettings, the committed projection revision and each metric's expected, actual and delta; plus the verdict per project.

The verdict follows the validator's parity rules. pass only when the validator would accept the sample. fail (exit 1) for a conclusive divergence (inParity: false or a non-zero delta with a materialized value present) or an unexpected answer (401, 403, 404, 5xx). Everything else is warn and the run's result is unproven, which is not a pass: an inconclusive audit, a disabled or fallback report (409 writes-disabled, materializedAvailable: false), a stale scope, or configurationMatchesCurrentSettings false or null. A credential or transport failure exits 2. The job runs on main only and reads its secret from the main-only statistics-evidence GitHub environment (setup); the preflight job itself has no environment, because an environment would change its OIDC subject away from the ref:refs/heads/main binding its GCP impersonation requires.

Schedule. The workflow carries a six-hourly schedule (23 */6 * * *) for periodic soak observations of both jobs. It does nothing unless the repository variable STATISTICS_PREFLIGHT_SCHEDULE is true; it is off by default and turning it on is the programme owner's decision. Scheduled runs are main runs.

The preflight does not prove:

  • runtime flag overrides on the writer host (deployment env is read from the Deployments and pods; overrides come only from the API's anonymous snapshot);
  • any value that arrives through a Secret (reported as <from secret>, never resolved);
  • that scheduled lifecycle jobs are registered or have run, or anything about the web Deployment;
  • Fresh/fallback behaviour on consumer reads, performance, or that a soak has started. The parity job's samples are soak parity evidence only once the evidence manifest, read and receipt captures and the offline validator put them in a window.

Choosing the pilot project

Record the decision and its reasoning in the evidence file. Criteria:

  • Real screening activity. The screening profiles must have something to distribute over: studies with include and exclude decisions, ideally from more than one reviewer, and ideally some studies with none. A project of all-zero counters proves that zero equals zero.
  • Small enough to backfill in minutes. The backfill is synchronous (see step 4), so its duration is a request duration. Prefer hundreds to low thousands of studies for the first run.
  • Not production-critical and not somebody's live work. Staging data is shared. Do not choose a project that a colleague is mid-review on: step 7 requires making a real screening decision on it.
  • Ideally one whose AgreementThreshold is stable. Configuration changes during the proof invalidate the control's configuration digest and complicate the report (ConfigurationMatchesCurrentSettings).

Record: project id, name, study count, reviewer count, screening decision count, and who confirmed it is safe to use. Read those counts from the project's existing UI or an authoritative read; do not write to the database to arrange them.

Step 2 — capture the "before" state

Before enabling anything, capture the project's current authoritative screening statistics through the ordinary application surface, and keep them in the evidence file. This is the legacy-response baseline that "prove legacy response equivalence" is measured against. With the consumer flags off, every page read is served authoritatively for the whole proof, so this baseline should still hold at the end for anything that did not change.

Step 3 — declare the fleet versions and open the fleet gate

POST https://api.staging.syrf.org.uk/api/admin/project-statistics/fleet/mode
Authorization: Bearer <administrator token>
Content-Type: application/json

{ "mode": "Enabled" }

Nothing serves until this runs. The flags in step 1 say what this deployment requests; the durable fleet control says what it is allowed to do, and it defaults to Disabled. Before this endpoint existed no production code created that control at all, so a completed backfill wrote rows that could never be read and the parity audit reported Disabled for every scope. Enabling is deliberately a separate act from backfilling: the rollout order is dark-write, prove parity, then serve, and a backfill that switched serving on would collapse the middle step it exists to make possible.

This one call does three things in one bounded transaction: it creates the singleton pmProjectStatisticsGlobalControl when the fleet has never had one, declares the catalogue/storage/source versions this build's writers produce, records the deployment's effective reviewer mode, and moves the durable gate to Enabled. A first declaration preserves both the write epoch and mode epoch at zero; it does not retire backfilled rows.

Before this first declaration, verify that every API and project-management replica agrees on ActiveReviewerTrackingEnabled && SignalRActive, and record that effective value in the evidence. The endpoint captures that process value once and stores it with the singleton and invalidation slot in one transaction. It does not assume the default false: an existing true deployment must retain its capacity semantics, including outside the pilot allowlist. This endpoint cannot verify other replicas' configuration. Use the coordinated static rollout; the runtime host-consistency gap remains

3360. Once a singleton exists, its reviewer mode is authoritative: a disagreeing enable returns

409 InvalidState and never silently changes the formula or resets its epoch.

The versions are not yours to choose. They are derived constants (ProjectStatisticsFleetVersions.Deployed). The reader applies two independent version equalities — fleet-versus-project-control and row-versus-control — and every published row is stamped from the deployed family writer's own constants, so a fleet declaring anything else would guarantee that no row is ever readable. Today they are catalogue 1, storage 0, source 1.

What the response reports is the durable row, not this build. recordedCatalogueVersion, recordedStorageVersion and recordedSourceVersion are read back from the control row the transition acted on — the fleet singleton here, the project's own control row for a narrow gate. A successful first declaration therefore echoes the constants above, but a transition over a row an earlier deployment enabled reports that row's numbers, so a disagreement with the constants above is real version drift rather than a reporting artefact. They are null only when the transition found no control row at all (a narrow gate refused NotFound, or a fleet transition whose singleton neither existed nor was created).

Run it before the backfill

The project control row copies the fleet's storage version verbatim when a backfill creates it, and the backfill refuses (ControlVersionMismatch) a project whose control disagrees with the fleet. Declaring first means the two agree by construction. It is also what makes the step-6 parity audit capable of reporting anything but Disabled.

Outcomes

Status outcome Means
202 Applied The gate is now Enabled. mode, writeEpoch and clientInvalidationRevision in the body are read back from the durable row.
200 AlreadyApplied The fleet was already Enabled under exactly these versions. Nothing changed and no revision was allocated.
409 InvalidState The recorded reviewer mode disagrees with this host, the fleet is serving under different versions — a version change is a disable/declare/re-enable sequence, never a silent edit under live traffic — or the requested stage is not legal from the current mode.
409 RebuildRequired This fleet control has been through a durable disable, so it re-opens only under rebuild evidence. See step 10.
409 QuarantineNotElapsed Only for a Disabled claim; see step 10.
409 InvalidationSlotUnavailable The fixed global invalidation slot could not be admitted, so the durable mode was deliberately left unchanged: a gate change nobody can be told about is worse than no gate change. Retry.
409 ConcurrencyLost A competing transition or writer won. The transaction was aborted; reload and retry.
400 — mode was not Enabled, Disabling or Disabled.

Both fleet and project transitions retry only an uncertain commit acknowledgement, at most three commit attempts on the same transaction. They never replay the transition body or rebuild sweep. If all acknowledgements remain uncertain, the error remains an outage rather than a false 409 claim that nothing committed; inspect the durable mode before resubmitting. Definitive write conflicts and duplicate-key races instead return 409 ConcurrencyLost after abort. Applied outcomes carry the committed invalidation revision; refusal responses do not allocate one.

Record the whole response body in the evidence file: mode, writeEpoch, clientInvalidationRevision and the three recorded*Version fields. Compare those three against catalogue 1, storage 0, source 1 before moving on: a difference means the durable fleet row was declared by another deployment, and every row published against it is ControlVersionMismatch for the reader.

Step 4 — backfill

POST https://api.staging.syrf.org.uk/api/admin/project-statistics/{projectId}/backfill
Authorization: Bearer <administrator token>

No request body, no query string. The route parameter is deliberately named statisticsProjectId rather than projectId in source, so that the shared authorization handler does not probe for the project and leak its existence to an authenticated non-administrator; on the wire it is just the project GUID.

Do not screen during the backfill

The backfill marks the served row Rebuilding before it calculates. A screening save on the pilot project while it runs cannot use that row as a baseline (#3831), so the save goes source-only, the backfill returns PublicationRaceLost, and the family stays Stale — pages recalculate live, so nothing wrong is shown, but the pilot is off materialized serving. Nothing retries on its own while the scheduled repair is off. If a parity audit or the fallback counters show the family Stale after a backfill, re-run this step at a quiet time, with no screening in progress, and confirm it answers 202. See saves during a backfill.

It is synchronous despite the 202

The work runs inline in the request, on the caller's cancellation token, and the response is returned only after every scope has been attempted. There is no background job, no hosted service, no Quartz schedule, no Location header, no job id and no polling endpoint. Treat 202 exactly as you would 200.

Size the client timeout for the whole rebuild. A client that times out aborts the request's cancellation token mid-run, which leaves scopes Stale and a rebuild lease held under the owner string stats.backfill.screening:admin:<32-hex-user-guid>. Use curl --max-time generously rather than a default.

202 Accepted — the success body

{
  "projectId": "...",
  "metricFamily": "ProjectScreening",
  "forced": false,
  "scopesRebuilt": 1,
  "scopesAlreadyCurrent": 0,
  "scopesStale": 0,
  "scopesAbsent": 0,
  "checkpointStatus": "Succeeded",
  "scopes": [
    {
      "scopeKind": "Project",
      "scopeKey": "project",
      "disposition": "Rebuilt",
      "status": "Succeeded",
      "failureReason": "None",
      "publishedGeneration": 1,
      "detail": null
    }
  ]
}

Screening declares exactly one scope, so scopes has one entry with scopeKind: "Project" and scopeKey: "project". disposition is one of Rebuilt, AlreadyCurrent, Stale, Absent. forced is false on /backfill and true on /rebuild.

A 202 is returned only when no scope is Stale and the bootstrap checkpoint was not refused. An Absent scope does not block a 202 — it means the authoritative calculator reported the scope does not exist and the rebuild returned it to Stale rather than publishing a fabricated zero row.

404 Not Found — empty body

One cause: the project document does not exist, so it declares no screening scope to rebuild. The service's explanatory detail is discarded, because NotFoundResult carries no body.

Note that this is not how backfill answers a project outside the allowlist — that is a 409 with failureReason: "NotAllowlisted" (below). The parity endpoint behaves differently and answers 404 for both; see step 6.

409 Conflict — not admitted

Nothing has been written; the check runs before any storage is touched. Key on the failureReason extension, not on status: Type and Instance are never set, so there is no RFC problem-type URI to match, and the Newtonsoft ProblemDetailsConverter flattens extensions to top level so status appears both as the framework's 409 and as a lifecycle enum name.

failureReason Meaning What to do
NotAllowlisted The project is outside ProjectStatistics:ProjectAllowlist. There is no implicit wildcard. Fix the cluster-gitops value, or the allowlist did not reach this host. This is the expected answer for every project today, since the allowlist is empty everywhere.
WriteDisabled materializedProjectStatisticsWrites is off. A backfill that ignored the global kill switch would be the one write the switch exists to stop. Turn the flag on through cluster-gitops.
FamilyNotServable materializedProjectStatisticsScreening is off. Turn the family flag on through cluster-gitops.
FleetVersionMismatch Either the fleet is serving but never declared its catalogue/source versions, or it declares versions the screening writer cannot produce. Every row published would be stamped at the writer's versions and rejected by the reader. Do not retry. This is a deployment-consistency problem: complete the fleet version declaration or deploy the matching family writer.
ControlVersionMismatch The project's control row carries catalogue/storage/source versions that do not match the writer's. Backfill only — the forced rebuild is exempt. Run POST .../rebuild, which reconciles the control and its current rows together inside the publication's own control compare-and-swap.

409 Conflict — partial failure or refused checkpoint

The title distinguishes three cases:

Title Means
The statistics rebuild could not publish every scope. at least one scope is Stale
The statistics rows are current, but the bootstrap checkpoint is incomplete. rows published, checkpoint refused
The statistics rebuild could not publish every scope, and the bootstrap checkpoint is incomplete. both
The statistics rebuild was interrupted by concurrent writes; nothing was published for the interrupted scope. Try again. failureReason: "RetryExhausted" (see below)

The body carries failureReason, status and a summary extension holding the full success DTO, so you can see which scopes did publish. Save it; it is evidence.

Common per-scope reasons and their responses:

failureReason Means Response
StatisticsRebuildBusy Another owner holds the per-scope rebuild lease, or the family guard's sole candidate slot is held. Wait for the lease to expire and retry. If it was your own timed-out client, wait it out rather than forcing.
FamilyFenced An active operation fence covers the family — a bulk, import or definition-rewrite operation. No visible scope can become servable while a fence covers it. Wait for the fencing operation to finish, then retry.
StatisticsInclusionRecalculationInProgress An inclusion recalculation or definition-rewrite fence is live. Wait and retry.
PublicationRaceLost A concurrent publication won the compare-and-swap. Retry.
LeaseLost The lease expired or was taken over mid-run — usually a run that took longer than expected. Retry.
RetryExhausted MongoDB aborted the scope's publication transaction on a write conflict — typically the notification-outbox dispatcher claiming the family's slot — on every one of the rebuild's bounded attempts (#3826). Nothing was published for that scope; it is left Stale, its lease and candidate released. Earlier versions answered this case with HTTP 500. Wait a moment and repeat the same request; the backfill skips scopes already current. Repeated refusals mean sustained write contention on the project — capture the response and the API log.
DigestMismatch A replayed operation identity carried different content. Do not retry blindly; capture and escalate.
ScopeAbsent The authoritative calculator says the scope does not exist. Reported as Absent, not Stale; does not fail the run.
InclusionStatusDrift MembershipScreening or ReviewerScreening (#3937): some Studies store a derived field that family reads (the inclusion status for ReviewerScreening, the agreement measure for MembershipScreening) that differs from the recalculated one, so an ordinary whole-Study save would change the family without a reviewer move. The scope is left Stale (status Blocked) and reads fall back. detail gives the count. Not retryable as it stands. Follow Before turning the reviewer flag on, item 3.
CheckpointCapacityExceeded, StatisticsScopeCapacityExceeded, StatisticsPublicationOperationCapacityExceeded A bounded capacity limit was hit. Capture; this is a Phase 1 provisional-limit finding worth reporting.

Recovering an inclusion recalculation fence

The 2026-09-23 staging incident on project 00000000-0000-0000-0000-000000000102 showed why a batch request's HTTP 200 is insufficient evidence of completion. Its update-all-study-inclusion request with replaceExisting=true updated all 32 matched Studies, then the fence completion transaction lost a MongoDB write conflict on a notification slot. The control retained InclusionRecalculationToken, and 53 operation fences stayed Active. The batch response carried a per-project failure. The source pass was complete, but no completion was committed.

A failed administrative recalculation now leaves a visible record. When the call holds the fence (its BeginAsync raised it or found its own token already held) and anything afterwards throws, the single-project and all-projects routes record an entry with Kind: "AdminInclusionRecalculationFailed" on pmProject.Errors and still rethrow. The entry carries the operation ID, token, exception type, typed failure reason and the recovery below. There is one entry per operation ID: a repeat that fails again increments FailureCount and updates the latest-failure fields, keeping FirstFailedAtUtc. A typed-busy refusal (StatisticsInclusionRecalculationInProgress because a recalculation under a different screening threshold holds the fence) raised nothing, so it propagates without an entry. The fan-out isolates each project: one project's failure is reported in Failures and leaves the others to complete. Nothing releases the fence on failure; that is deliberate.

1. Diagnose (read-only). As an application administrator, call GET /api/admin/project-statistics/{projectId}/inclusion-recalculation-fence?strandedAfterMinutes=30. It writes nothing and reports:

Field Meaning
fenceHeld, token Whether InclusionRecalculationToken is set, and its value
operationId The operation that derived the token, resolved from today's configuration or the guards' active operation IDs; null if neither matches
tokenMatchesCurrentConfiguration True when a repeat under today's threshold would resume the held token
fenceRaisedAtUtc, fenceAgeMinutes Earliest active operation-fence creation time for that operation
liveSourceJob Whether pmProject.ActiveInclusionInfoCalculationJob is open
stranded, strandedReason NoLiveSourceJob (held, no open job, and older than strandedAfterMinutes), OlderThanThreshold (open job but older than strandedAfterMinutes), or None. The threshold is 1 to 10080 minutes, default 30
activeOperationFenceCount, families[] Active operation fences for the operation, and per family the guard lifecycle, guard active-fence count and active operation ID
recordedAdminFailureCount Total FailureCount across the project's AdminInclusionRecalculationFailed entries
recovery The applicable step below, or null when not stranded

An administrative recalculation runs without the job marker, so the diagnostic cannot tell a running admin pass from an abandoned one; it therefore reports a marker-less fence as stranded only once it is older than strandedAfterMinutes, and recovery is null until then. Raise the threshold if a project's passes legitimately take longer. Serialize recovery with any other inclusion recalculation for the project.

What a repeat does while an earlier pass under the same threshold is still running: it joins rather than blocks. Both calls derive the same token, so the second BeginAsync answers AlreadyHeld without writing, and both run the three idempotent passes towards the same target state. Whichever finishes first releases the fence; the other's completion then answers NotHeld and returns normally. That is safe, because the passes converge on identical values, but it doubles the source load for nothing, which is why the diagnostic withholds the recovery until the grace has passed. A repeat under a different threshold is refused typed-busy and changes nothing. Since #3841 a repeat is also refused typed-busy, changing nothing, while liveSourceJob is true: the screening-settings change's own job then owns the token, and a repeat that joined it would clear the token while that job's passes were still running.

Two limits on reading the diagnostic during recovery:

  • fenceAgeMinutes counts from the earliest operation-fence creation, and a repeat that finds its own token (AlreadyHeld) does not reset it. A recovery repeat still running after strandedAfterMinutes is therefore reported stranded again, with recovery text that advertises another repeat. Wait for the running repeat to return, or raise strandedAfterMinutes, before acting on that report.
  • recordedAdminFailureCount is history, not a live signal: failure entries stay on the project after a successful repeat clears the fence. Judge current state from fenceHeld, token and stranded.

2. Repeat the recalculation under the unchanged threshold. Only when tokenMatchesCurrentConfiguration is true and liveSourceJob is false (with the job open, the repeat refuses with a typed 409, code StatisticsInclusionRecalculationInProgress; follow step 2a instead): with an authenticated Project.Edit principal, call POST /api/projects/{projectId}/update-study-inclusion with JSON body true. BeginAsync finds the same token (AlreadyHeld), the three source passes rerun, and CompleteAsync releases the matching token after their acknowledged completion. Already-correct Studies count as success on replay. A transient, definitively aborted completion write conflict gets at most three fresh-transaction attempts; unknown commit results retry the commit of the same transaction only. Do not change the screening threshold first: a different threshold derives a different token and fails typed-busy against the held one. If tokenMatchesCurrentConfiguration is false, the fence was raised under an earlier configuration; restore that threshold or escalate.

2a. The settings change's own job is still open (liveSourceJob true, reason OlderThanThreshold). Only that job's command can finish it. Find out where the command is before acting:

  1. In project-management logs, search for Project Agreement Threshold update failed for project {projectId}; the consumer also appends that text to the project's Errors. Its presence means the consumer ran and faulted.
  2. In the environment's RabbitMQ management UI (virtual host of the SyRF bus), inspect queue SyRF.ProjectManagement.IUpdateStudyScreeningStatsCommand (a message still waiting or unacknowledged: the job is only slow, so wait) and SyRF.ProjectManagement.IUpdateStudyScreeningStatsCommand_error (the consumer faulted). Use "Get messages" with Nack message requeue true to read without removing, and match the body's projectId.
  3. If this project's message is in _error: fix the cause shown in the fault headers, then use the management UI's Move messages (shovel plugin) to move it to SyRF.ProjectManagement.IUpdateStudyScreeningStatsCommand. The consumer resumes the same token (AlreadyHeld), reruns the three passes, closes the job and clears the token. Continue at step 3.
  4. If neither queue holds it and no consumer failure was logged, the command was never sent (a dispatch lost after the commit). Do not assume an error-queue message exists and do not publish a hand-written message. No supported path re-sends it today: escalate under #3086 / #3849. Since #3841 an aborted request and an unknown commit result no longer cause this (the commit ignores request cancellation, and an unknown result is resolved by a fresh read before dispatch); the remaining cause is a Send failure inside the post-commit domain-event dispatch.

Host flag prerequisite. Since #3841 the API raises the inclusion token on a settings change and the project-management consumer clears it. The API reads materializedProjectStatisticsWrites through runtime overrides; project management reads its deployed value. If the API sees writes on and project management sees them off, the consumer's completion is a no-op and the token stays held after the job closes (a NoLiveSourceJob stranding, recovered by the step 2 repeat). Keep the flag's effective value the same on both hosts, and prefer deployed values over a runtime override for it.

3. Verify. Call the diagnostic again: fenceHeld must be false. Inspect the returned AllSuccess and the durable control: the token must be null, all operation fences for the ID terminal, and affected families Stale. If any proof is absent, keep the fence and investigate the typed failure; do not delete fences or clear the token directly.

4. Backfill. Run the ordinary administrator backfill for ProjectScreening and any other affected family intended to serve. Confirm a Fresh publication before considering the consumer recovered. Retained history is not rewritten by this fence recovery.

Superseded Rebuilding current rows left by earlier rebuild candidates are benign: the reader selects by guard generation and never serves them. They are not reclaimed today, because ProjectStatisticsCurrentCleanup has no producer; do not delete them by hand. Reclamation is tracked in #3809.

For a staged search import or bulk Study operation, source completion and statistics-fence release are separate steps. Release now retries a definitively aborted MongoDB write conflict at most three times, each time rereading the matching operation's active fences in a new transaction. An unknown commit result retries only the original commit. If release still fails, the family remains fenced: confirm the source operation's durable completion and redeliver that same operation ID through its owning workflow so it can release its own fence. Verify the operation fences are terminal and affected families Stale before backfilling. Do not clear fence documents directly or start a different operation to take ownership.

A refused bootstrap checkpoint on a non-forced run publishes nothing — current rows must not outrun the single bootstrap point they are paired with. The detail says to retry the backfill, and, if another checkpoint occupies the identity, to run a forced rebuild (which publishes and therefore moves the projection revision) and then backfill again to record the bootstrap at the new identity. The forced rebuild is exempt from this short-circuit.

Forced rebuild of a multi-scope family

The project control holds one configuration identity — the shared configuration digest and the catalogue/source versions — and the reader serves a scope's row only while the row carries exactly that identity. A forced rebuild publishes one scope at a time, and when its authoritative calculation disagrees with the control (a threshold changed without a source write, or the digest definition was bumped in code) the first scope it publishes replaces the control's identity. From that moment every other scope of the family still carries the old identity and falls back to the authoritative path.

The rebuild service therefore finishes the job itself: a forced publication that replaced the identity immediately rebuilds every sibling scope of that family whose Fresh row carried the replaced identity, using the ordinary automatic rebuild (so a sibling whose own calculation disagrees again is refused DigestMismatch rather than replacing the identity a second time). Search population (one scope per linked search), the reviewer screening families (one per member), stage annotation (one per stage) and the membership-stage and reviewer annotation families all benefit; the single-scope project-screening family has no sibling and is unchanged. Automatic runs never replace an identity and so never cascade.

What to expect from POST .../<family>/rebuild after an identity change:

  • The route response is unchanged. The family sweep still republishes every scope itself, so its per-scope summary is the final word; a scope the cascade already rebuilt is simply republished again by the sweep (roughly twice the calculations for that one run; a forced run whose identity does not move rebuilds each scope once, as before).
  • The sibling pass runs after the triggering scope's publication has committed, so it never turns that publication into a failure. A sibling whose rebuild is refused (contended, raced, fenced) is reported Refused with its typed outcome; one whose rebuild throws (a transient storage error) is reported Failed and logged at error; if the request is cancelled mid-pass, the sibling in flight is Failed and the rest NotAttempted; and if the siblings cannot even be discovered, the pass is reported incomplete with none listed. In every such case the triggering scope still counts as a completed rebuild in the rebuilds.completed metric.
  • A scope that stays Contended, Stale or raced after the run — or that the sibling pass left Refused, Failed or NotAttempted — is still behind the old identity. Retry the family's plain backfill: once the identity has been adopted the automatic pass no longer disagrees with the control, and it republishes exactly the scopes whose rows carry the old digest. Retrying the forced rebuild would not re-run the sibling pass, because the identity no longer changes.
  • Other families on the same project keep rows stamped with the old identity and fall back until their own backfill (or rebuild) runs. Run it for every family the project serves.

500

Non-development hosts return {"error":"An unexpected server error occurred."}. Capture the timestamp and correlate with the API logs; do not retry in a loop.

Step 4b — stage-annotation backfill

Only needed when the pilot will exercise the Stage Overview annotation pie (materializedProjectStatisticsStageOverview). Run it after step 4 and before any consumer flag is opened: until a stage has a Fresh row, the pie's query falls back and opening its flag changes nothing an operator could observe.

POST https://api.staging.syrf.org.uk/api/admin/project-statistics/{projectId}/stage-annotation/backfill
Authorization: Bearer <administrator token>

Same shape as step 4 in every respect — no body, the same statisticsProjectId route token, the same BatchAdminProjects policy, synchronous despite the 202, /rebuild for the forced variant — and the same 202/404/409 tables apply unchanged, because the same service runs behind both routes. Three things differ, and they are the whole of the difference:

  • metricFamily is StageAnnotation, and the family gate this run checks is materializedProjectStatisticsAnnotation, not materializedProjectStatisticsScreening. A deployment with screening on and annotation off gets 409 FamilyNotServable here while step 4 succeeds, which is the expected answer rather than a fault.
  • One scope per stage — every stage, not only the annotation-mode ones. The legacy aggregation projects an annotation section for each of project.Stages with no review-mode filter, so a screening-mode stage has one too (all zeroes). scopes therefore has one entry per stage with scopeKind: "ProjectStage" and a scopeKey naming the stage. Stages are enumerated from the Project document, not from the projection, so a first run covers every stage. A project with no stage completes with an empty scopes array and a 202; 404 still means the project document does not exist. A stage whose project carries no agreement threshold is reported Absent — the legacy aggregation cannot run without one — rather than failing the whole run.
  • The lease owner string is stats.backfill.screening:admin:<32-hex-user-guid>, unchanged — the operation namespace is shared across families, so a stuck lease is told apart by its scope, not by its owner.

Record the response body in the evidence file alongside step 4's. Re-running it is a no-op: every scope whose published row already matches the authoritative calculation is reported AlreadyCurrent and no generation is churned.

The original screening/annotation-pie pilot did not require membership-stage annotation rows. PR #3561 adds POST api/admin/project-statistics/{projectId}/membership-stage-annotation/backfill and /rebuild; establish those baselines before enabling a consumer that requests that family. Verify the deployed version and its independently gated consumer rather than assuming a code PR has activated this additional family.

Step 5 — open the project narrow gate

POST https://api.staging.syrf.org.uk/api/admin/project-statistics/{projectId}/narrow-gate
Authorization: Bearer <administrator token>
Content-Type: application/json

{ "enabled": true }

The fleet gate from step 3 is necessary and not sufficient: each project carries its own narrow serving gate on its control row, fenced by its own ProjectWriteEpoch, and it also defaults to Disabled. Both must be Enabled before the bundle reader will serve a materialized value or the parity audit can report anything but Disabled.

Run it after the backfill, not before. The narrow gate lives on the project's control row, and only a backfill or a materialized source write creates that row. There is deliberately no create-on-enable here: a control row created by this endpoint would carry no configuration digest and no versions, and the reader would refuse every row published against it for ever. A project with no control row is answered 404.

The route parameter is named statisticsProjectId in source for the same reason as the backfill route — so the shared authorization handler does not probe for the project and leak its existence to an authenticated non-administrator. On the wire it is just the project GUID.

Outcomes

Status outcome Means
202 Applied The narrow gate is now Enabled. The body carries the resulting mode and writeEpoch (the project write epoch), the project control row's own recorded*Version fields and, on a re-enable, scopesRebuilt.
200 AlreadyApplied The gate was already Enabled. Nothing changed.
404 — The project has no statistics control row. Run step 4 first.
409 RebuildRequired The gate has previously been disabled, so it re-opens only under a complete rebuild proof, and the sweep this call ran did not complete. rebuildFailureReason carries the first scope's typed reason — FamilyFenced, StatisticsRebuildBusy, PublicationRaceLost and so on — which is the thing to clear.
409 InvalidState An unfinished Disabling quarantine may not be jumped out of. Complete the disable, then re-enable.
409 QuarantineNotElapsed Only for a { "enabled": false } call claiming Disabled; see step 10.
409 InvalidationSlotUnavailable / ConcurrencyLost As for the fleet gate. Retry.

Confirm it took

Re-run the parity audit (step 6). Before this step it reports Disabled for every scope — the audit reader deliberately keeps every durable gate, overriding only the two flag-level serving switches, so an all-Disabled report is exactly what a shut narrow gate looks like. After it, the report should carry real scopeStates and a real inParity verdict.

An all-Disabled parity report is not a passed gate. It is the report saying nothing was audited. Read fallbackReason before concluding anything from it.

Step 6 — parity audit

GET https://api.staging.syrf.org.uk/api/admin/project-statistics/{projectId}/parity
Authorization: Bearer <administrator token | statistics-parity evidence token>

This endpoint arrives with #3196 and is not on main as of 38a272d4c. It is read-only on both sides: nothing here creates a control row, family summary, guard or lease, and nothing advances a revision. It answers while serving is off, deliberately — the rollout order is dark-write, then prove parity, then serve.

Outcomes:

  • 200 OK with ProjectStatisticsParityReportDto.
  • 404 — not allowlisted, or the project does not exist. Here the two are deliberately indistinguishable, so the endpoint cannot be used to enumerate the pilot. This differs from backfill, which answers 409 NotAllowlisted for the first case and reserves 404 for the second.
  • 200 OK with fallbackReason: "Disabled" on every scope — the durable gates are shut. The audit reader overrides only the two flag-level serving switches; it deliberately keeps every durable gate, including the fleet control's mode and the project's narrow gate. Run steps 3 and 5. This is a report of nothing having been audited, not a passed gate.
  • 409, type: "urn:syrf:project-statistics:writes-disabled", title "There is no materialized projection to audit." — materializedProjectStatisticsWrites is off, so nothing is maintaining the projection and any report would describe a frozen artefact. Enable writes and backfill first.

Automated capture with the read-only evidence credential

Automation (a scheduled workflow, or an operator's terminal) does not need a human administrator account to poll parity. SyRF Identity can seed a dedicated confidential client, syrf-statistics-evidence, whose only grantable scope is statistics:parity:read, requested with the client-credentials grant.

What it can do. Read GET /api/admin/project-statistics/{projectId}/parity for a project in the pilot allowlist, and nothing else:

  • The API accepts the scope on that one route and method. On every other route — including POST …/parity, backfill, rebuild, mode, narrow gate, fleet, maintenance, pilot status, every project, study and account route and the SignalR hub — the token is not a valid API token, so the request is anonymous (401). Behind that, the parity policy is attached only to the parity action, so even a principal that got past the first gate would receive 403 from every administrator policy.
  • The allowlist still applies. An unlisted project answers 404, exactly as for an administrator, before anything is read.
  • The client cannot request syrf_api, admin:users, the notifier scopes or a refresh token, and has no redirect URI, so it cannot sign in as or on behalf of a user.
  • The audit runs under a fixed non-user caller with no application groups; no investigator is read or created for it.

What the report exposes. Counters, deltas, revisions, epochs, digests, scope/publication states, the project GUID and a timestamp — the ProjectStatisticsParityReportDto shape shown below, pinned by ProjectStatisticsParityAdminControllerTests.TheReportCarriesOnlyCountersRevisionsAndDigests. It carries no study, citation, reviewer, annotation or user content.

Threat model. A leaked secret lets its holder read parity counters for the few allowlisted pilot projects until the secret is rotated or removed; it cannot change any state, read project content or reach another route. Least privilege is enforced twice (authentication binds the scope to the route; authorization binds the route to the policy) and both halves are covered by tests. Keep the secret out of logs and chat, store it only in the Secret and the statistics-evidence GitHub environment secret named below, and remove it when the proof ends.

Revocation is not instantaneous. Removing the secret takes effect only when Identity next starts and re-seeds; until then the old registration keeps working. Rotating the secret stops new tokens from the old secret but does not revoke tokens already issued. Deleting the registration also cascades to the authorizations and tokens OpenIddict stored for it, which should make introspection report them inactive, but this has not been proved for client-credentials tokens here. Plan for the worst case: an issued access token remains usable until it expires, at most one hour. Treat a suspected leak as readable until an hour after the Identity rollout completes, and restart Identity promptly after changing the secret.

Enabling it (off by default). Nothing is seeded unless a secret is configured. Identity treats an absent, blank or CHANGEME (any case) secret as "off" and, on startup, deletes every registration whose permissions include statistics:parity:read, so removing the secret revokes the credential. The secret is trimmed; one shorter than 32 characters is treated the same way (disabled and revoked, with a warning in the Identity log) without affecting any other client. The client id is fixed, syrf-statistics-evidence, and is not configurable, so disabling can always find it. To enable, in cluster-gitops for the Identity service of the target environment (see required secrets):

  1. Create the Kubernetes Secret syrf-statistics-evidence (key clientSecret) through the environment's usual secret path (External Secrets / GCP Secret Manager).
  2. Set statisticsParityEvidence.enabled: true in the Identity values (optionally statisticsParityEvidence.clientSecret.{secretName,key}). This renders SYRF__StatisticsParityEvidence__ClientSecret.
  3. For automation, store the same secret as the environment secret SYRF_STATISTICS_EVIDENCE_CLIENT_SECRET of the main-only statistics-evidence environment (setup), not as a repository secret.

The API needs no configuration change: it already accepts the scope on the parity route, and without a seeded client no one can obtain such a token.

Capturing. The standard-library helper requests a token and appends one record in the parity JSONL envelope per run:

export SYRF_STATISTICS_EVIDENCE_CLIENT_SECRET=...   # from the secret store, never on the command line
python3 scripts/fetch-statistics-parity.py \
  --identity-url https://identity.staging.syrf.org.uk \
  --api-url https://api.staging.syrf.org.uk \
  --project-id <pilot project GUID> \
  --output /private/run/parity.jsonl
python3 scripts/test-fetch-statistics-parity.py   # offline tests of the helper

Every 200 answer is appended, including out-of-parity and inconclusive reports. A non-200 answer (404 unlisted, 409 writes disabled, 401/403) is printed to stderr, not recorded, and exits 1; a token or transport failure exits 2.

In CI, the Statistics Staging Preflight workflow's parity job does this for every allowlisted project with the statistics-evidence environment secret SYRF_STATISTICS_EVIDENCE_CLIENT_SECRET, through scripts/capture-statistics-parity.py (one token, one retry while inconclusive, a verdict per project and a job summary). See Staging preflight.

Reading the report

{
  "projectId": "...",
  "metricKey": "project-screening",
  "inParity": true,
  "isInconclusive": false,
  "sourceAdvancedDuringAudit": false,
  "materializedAvailable": true,
  "readSource": "...",
  "fallbackReason": "...",
  "capacityFailure": "...",
  "checkpoint": { "checkpointSourceRevision": 0, "checkpointProjectionRevision": 0,
                  "checkpointModeEpoch": 0, "canonical": "..." },
  "watermarks": { "globalClientInvalidationRevision": 0, "sourceInvalidationRevision": 0,
                  "committedProjectionRevision": 0, "clientInvalidationRevision": 0,
                  "modeEpoch": 0, "catalogueVersion": 0, "storageVersion": 0,
                  "sourceVersion": 0, "configurationDigest": "..." },
  "scopeStates": [ { "selection": "...", "scope": "...", "availability": "...", "state": "...",
                     "isTombstone": false, "isPublished": true,
                     "publicationGeneration": 1, "lastChangedRevision": 0 } ],
  "metrics": [ { "metricKey": "screening.sufficientlyScreened",
                 "expected": 0, "actual": 0, "delta": 0, "inParity": true } ],
  "observedAtUtc": "...",
  "controlConfigurationDigest": "...",
  "currentSettingsConfigurationDigest": "...",
  "configurationMatchesCurrentSettings": true
}

inParity is the verdict. It is true only when a materialized value exists and equals the expected one, for every metric. Crucially, a bundle that legitimately refuses to serve is not "in parity" — it is unaudited, inParity is false, materializedAvailable is false, and fallbackReason says why. Read fallbackReason before concluding anything from a false verdict.

The comparison runs over the union of both key sets with absence read as zero, in the catalogue's declaration order. That union matters both ways: a metric the projection never wrote is a legitimate zero, and a counter the projection holds that the authoritative side does not produce is a real defect the report surfaces rather than skips.

isInconclusive means a write landed underneath the comparison, so the deltas may be an artefact of the race rather than real drift. The two halves are separate reads; the project's revision is checked either side of the authoritative half, a move is retried once internally, and a second move reports the run as inconclusive rather than as a divergence an operator would chase. sourceAdvancedDuringAudit reports the same underlying observation.

Retry rule for inconclusive. Re-run the audit unchanged. If it is still inconclusive, the project is under sustained write load and cannot be audited coherently at all — that is the report telling you so, not a transient. Re-run during a quiet window, or choose a quieter pilot project. Do not re-run repeatedly until it happens to look quiet, and never record an inconclusive run as a parity pass or as a divergence. Record every attempt in the evidence file, including the inconclusive ones.

configurationMatchesCurrentSettings is separate from the verdict on purpose:

  • true — the control row is stamped under the project's current settings.
  • false — the control is still stamped under settings the project no longer has, so its counters answer a superseded question even when they agree with today's calculation. This is drift to chase separately, not a divergence in the values compared here. It is also false whenever isInconclusive is true.
  • null — the comparison could not be made. Null never means it passed. It happens when there is nothing to digest, for example a project with no agreement threshold or one that vanished mid-run.

A fresh forced rebuild adopts the digest of the project it read, so it reports a match; the field says nothing about how the digest came to be stamped.

Also capture watermarks and scopeStates verbatim — availability, state, isPublished and publicationGeneration are what distinguish Fresh from Stale, Rebuilding, Incompatible, fenced and tombstoned, and they are the record of which of those paths the proof actually exercised.

Step 7 — exercise a real screening decision

With the projection backfilled and audited, make an actual screening decision on the pilot project through the ordinary UI, as an ordinary reviewer would. This is the point of the proof: the transactional delta path must move the projection in step with the source, without a rebuild.

  1. Note the pre-decision watermarks.committedProjectionRevision and the relevant metric values.
  2. Screen one study — include or exclude — through the normal screening surface. Note the study, the reviewer, the decision and the time.
  3. Re-run the parity audit.
  4. Expect inParity: true with the counters moved by exactly the decision made, and committedProjectionRevision advanced. A false verdict here with materializedAvailable: true and a non-zero delta is the finding the proof exists to catch: capture it in full and stop.
  5. Repeat for a decision of the opposite kind, and for a study that takes a profile from one populated bucket to another, so more than one counter is proven to move.

Worth exercising deliberately while here, since they are named in the Phase 2 staging proof:

  • Kill switch. With the flags turned off in cluster-gitops, confirm reads fall back to the authoritative calculation and the parity endpoint answers 409 writes-disabled. Turn them back on and confirm the projection is still there and still in parity — the rollback preserves data.
  • Stale / Rebuilding. Observe them in scopeStates.state around a rebuild.
  • Forced rebuild. POST .../rebuild and confirm the report is still in parity afterwards, with a new publicationGeneration.

Step 8 — capture evidence

Write one JSON file per proof run to evidence/phase2c-staging/, named phase2c-staging_<YYYY-MM-DD>_<short-run-label>.json. Start from the committed template, phase2c-staging-evidence-template.json, which follows the same shape as the Phase 0 baselines.

Record, at minimum: the date, environment, project id and selection rationale, the exact deployed flag values and allowlist as read from the running pods, the deployed image digests, the backfill outcome, every parity report including inconclusive attempts, the screening decisions made, the fallback reasons observed, and notes. Store the raw response bodies rather than a summary of them — a verdict without its watermarks and scopeStates cannot be re-interpreted later.

Never put reviewer identities, study titles or other project content into the evidence file. Ids, counts and the response envelope are enough.

Evidence from a fold-mode project

On a project in async point-fold mode the source-operation receipt is written by the fold, not by the save (how a fold-mode save is counted). Three consequences for the evidence:

  • Receipts are observed at fold time. ObservedAtUtc is the fold commit, normally under a second after the save but up to the next one-minute sweep, and longer while the worker is down. A save made within that interval of a capture boundary can fall into the next capture window, so close a capture only once GET …/fold/status shows pendingStudies: 0, or accept and record the boundary difference. CommittedSourceBeforeRevision and CommittedSourceAfterRevision are still the save's own revisions, so the validator's rules (after > before >= 0, deduplicated by namespace and id) hold unchanged.
  • Overflowed and quarantined saves have no receipt. They have a quarantine record instead. Record the totals of pending_entries.overflowed and fold.quarantined (by reason) for the same window next to the validator summary's confirmedSourceOperations, so that a save missing from the receipts is accounted for rather than silently lost. scripts/validate-statistics-evidence.py does not count either; the read-only receipts export of #3952 is to supply them. On a healthy run both are 0.
  • source_operation_receipts.committed is recorded by the fold worker (project-management), once per consumed entry after its commit is confirmed, so read it on the project-management pods for a fold project, not the API pods.

Step 9 — soak

Where the soak runs. On 30 September 2026 the owner decided that the soak must not need a real user account and must run in an isolated environment. It therefore runs in the isolated local e2e stack on Bramble (seeded mock-oidc users, a synthetic project, generated data only), not on staging. The driver, its drills and the read-only receipts export are #3952; the gate it evidences is #3510, whose body still says "staging soak" and is updated with #3952. Steps 1 to 8 of this runbook (the Phase 2C proof on one staging project) still run on staging; only the 7-day volume soak moves. Run nothing in the soak against staging or production, and use no real account or credential. #3952 adds the start, stop and resume procedure and where its evidence lands; until it merges there is no soak procedure to follow.

The fold-mode soak uses the same administrative operations as the Statistics Operator workflow (pending-index build, fold-enable, backfill), against the local stack's API.

The README's soak gate and Phase 6's acceptance criteria require, before any production pilot:

  • at least seven actual days of staging soak — calendar days, not seven days' worth of traffic compressed into an afternoon;
  • 10,000 reads;
  • 1,000 relevant mutations;
  • 100% exact counter parity, zero stale materialized serves, 100% injected fallback success, no unresolved rebuild failures, and bounded delta-ledger and storage growth.

Two honest caveats about counting those numbers on a single project, which apply in the isolated stack as they did on staging.

The read and write counters now exist, but nothing scrapes them yet — and they are silent until the flags are on. Since syrf #3482 both hosts register a real Meter and record reads by source and reason, confirmed write commits, rebuild and backfill outcomes and parity results — see Telemetry for the instruments and their tags. Two limits matter for counting. The read counters record only once a consumer actually enters the adapter, which the consumer flags gate, so a dark deployment produces none of them (its startup line is the whole FEAT-024 signal). And there is still no destination: OTEL_EXPORTER_OTLP_ENDPOINT is unset in every environment, so the series live only inside each pod and must be read there with dotnet-counters, or reproduced by a driver that counts its own calls. A driver remains the more defensible option for the controlled 10,000/1,000 figures, since the plan calls them "controlled"; ambient traffic on one project will not reach them in seven days, and the isolated stack has none. Use the counters to corroborate the driver's own totals and to see the distribution — which reasons, which commit paths — that a driver cannot tell you.

With every consumer flag off, ordinary page reads do not touch the projection at all. They are served authoritatively. So "10,000 reads" against a Phase 2C deployment means 10,000 reads of the projection, which during this phase means the parity/admin surface or a driver exercising the bundle reader — not 10,000 page loads. Decide which you are counting, write it down in the evidence file, and do not let the distinction blur.

Observing fallback reasons

Every gated read that declines to serve a materialized value logs a fallback reason. Watch the distribution across the soak: a fallback that is supposed to be rare (Stale, Rebuilding, Incompatible, an epoch mismatch, an active fence) appearing steadily is the signal the soak is for. The two reason enums, the exact log message templates and the queries to count them by are in Fallback reasons and log fields below, and the counter that carries the same two reasons as tags is in Telemetry. Read both before planning the soak, and note what neither gives you before step 1: both API consumers check materializedProjectStatisticsPages (and, for the overview, ...ProjectOverview) and return the authoritative answer without entering the adapter at all, so in a fully dark deployment the read counters stay at zero and no per-read line is emitted. The only FEAT-024 series a dark pod produces is the one-per-host startup line. Once the flags are on, the counter is the more reliable of the two because it does not depend on a log level or a log pipeline.

Record the observed set — not just the counts, the set — in the evidence file. A fallback reason nobody expected to see is a finding whether or not parity held.

Step 10 — rollback

If the project is in async point-fold mode, disable the fold first. Run fold-disable (Statistics Operator workflow, confirm ticked; or POST …/fold/disable), then poll fold-status until foldMode: "Disabled": the worker finalizes it once the five-minute writer grace has passed and nothing is pending. Only then continue with step 1 below, and turn the fold flag off with the other flags in step 4. Turning the fold flag off while fold mode is still Enabled strands any pending entries (Disable and rollback). Rolling the images back past the fold follows the design's rollback guard and runbook.

Rollback closes the durable gates first, through the API, and then the flags, through cluster-gitops. Both halves matter and they are not interchangeable: the flags stop this deployment from asking to serve, while the durable gates and their write epochs stop any deployment from serving the rows — including a replica whose per-process flag cache has not caught up. The flag cache is explicitly not a correctness authority.

The flag half is a cluster-gitops change, reverted through git and synced by ArgoCD. Never kubectl set env, never helm upgrade.

1. Close the project narrow gate.

POST /api/admin/project-statistics/{projectId}/narrow-gate   { "enabled": false }

Answers 202 with mode: "Disabling". That is not a half-done call — it is the first of two stages. The project write epoch advances at this commit, so a source transaction that observed the gate as Enabled and commits afterwards writes rows carrying the superseded epoch, which are already non-servable. Reads fall back to the authoritative calculation immediately.

2. Claim Disabled once the quarantine has elapsed. Repeat the identical call. Before the deployment-verified maximum transaction lifetime has passed it answers 409 with QuarantineNotElapsed; that interval, not the mode field, is the durable barrier. Afterwards it answers 202 with mode: "Disabled".

3. Close the fleet gate, the same two stages.

POST /api/admin/project-statistics/fleet/mode   { "mode": "Disabled" }

The first call answers 202 with mode: "Disabling" and advances the fleet write epoch; the second, after the quarantine, answers 202 with mode: "Disabled".

4. Turn the flags off. Set materializedProjectStatisticsWrites, materializedProjectStatisticsServing and materializedProjectStatisticsScreening back to false on both hosts, together with materializedProjectStatisticsFold and any other family flag that was turned on. This is a cluster-gitops change for api and project-management (a git revert, synced by ArgoCD); the project-management host does not see runtime flag toggles (#3360).

5. Return SYRF__ProjectStatistics__ProjectAllowlist to the placeholder, or remove the key.

Any one of steps 1, 3, 4 and 5 alone is sufficient to stop the pilot serving — the allowlist admits nothing without a valid GUID, and the flags and the durable gates each gate every read independently — but do all of them, so the deployed state matches the intended state in one reading.

Confirm with the parity audit: after the gates are closed it reports Disabled again, and after the flags are off it answers 409 writes-disabled.

Narrow-gate requests must explicitly supply enabled: true or enabled: false; an omitted or null value returns 400 before any transition. Contention responses report the mode, epoch and invalidation revision observed together after abort. These request-validation and reporting corrections use the existing administrative gates and need no additional feature flag.

Re-enabling after a durable disable

A gate that has been through a disable does not re-open the way it was first opened. Its quarantine retired a whole epoch of rows, so:

  • The narrow gate re-opens only under a complete rebuild proof. { "enabled": true } produces that proof itself — it rebuilds every scope the project's routed families declare and presents the outcomes to the transition — and refuses with RebuildRequired plus the sweep's own rebuildFailureReason if any scope could not be published.
  • A successful re-enable still serves nothing until you rebuild again. The re-enable allocates a new project write epoch, and the sweep that produced the proof necessarily ran before that epoch existed, so its rows carry the superseded tuple and the reader answers EpochMismatch. This is the plan's stated residual — "a re-enable that serves nothing until a rebuild runs, never one that serves stale data" — not a defect. Run POST .../rebuild after the gate is open and confirm the parity report before treating the project as serving again.
  • The fleet gate answers 409 RebuildRequired and is deliberately not re-openable from this surface: a fleet re-enable's evidence would have to span the whole fleet, and no bounded transaction can establish that. Re-enabling a fleet that has been durably disabled is out of scope for the Phase 2C pilot; a pilot that needs it should be torn down and rebuilt from a fresh fleet control rather than coaxed back open.

What rollback leaves behind

Data, deliberately. "Rollback to authoritative reads is immediate and does not require data deletion", and "a flag rollback returns every read to the authoritative screening facets and preserves history". Concretely, after rollback:

  • The pmProjectStatistics* documents written during the proof — current rows, control row, delta ledger, the backfill-observed checkpoint — remain. They are simply never read.
  • Every read returns to the authoritative calculation immediately. There is no drain and no window in which a stale materialized value is served.
  • Source mutations continue committing normally. With writes off they commit the source, mark affected scopes Stale and fall back, so the projection stops tracking the source from the moment the flag goes false. This is why re-enabling requires a fresh backfill or forced rebuild, not just flipping the flags back: the rows are intact but no longer current, and the reader will correctly refuse them as Stale.
  • History is preserved. The bootstrap checkpoint keeps the identity it was captured under and is never rewritten.

Nothing needs to be deleted, and nothing should be. If the pilot is genuinely being abandoned rather than paused, removing the projection documents is a separate, separately approved cleanup — not part of rollback.

Telemetry

Since syrf #3482 the FEAT-024 instruments are real. ProjectStatisticsMetrics (SyRF.ProjectManagement.Core/Telemetry/ProjectStatisticsMetrics.cs) owns one System.Diagnostics.Metrics.Meter named SyRF.ProjectManagement.ProjectStatistics, registered as a Lamar singleton by ProjectStatisticsRegistry and named in AddMeter(...) by both hosts' Program.cs, alongside the existing SyRF.API.BffAuth meter.

Recording is not flagged. It has no behavioural effect, and the whole point is that a dark pilot should be observably dark; gating telemetry on the flags being on would leave the dark period unobservable, which is the state this section used to describe.

The instruments

Every tag value is a bounded enum member rendered snake_case (ProjectNotAllowlisted becomes project_not_allowlisted). No instrument is ever labelled with a project, membership, investigator, question or scope-key identifier — that is a plan rule, and a unit test asserts no tag value parses as a GUID. deployment.environment (development / preview / staging / production / unknown) is on every series, exactly as it is on the BFF meter.

Instrument Kind Tags beyond deployment.environment Recorded by
syrf.project_statistics.requests counter family, scope_type, request_kind, outcome (materialized / authoritative_fallback) ProjectScreeningStatisticsQueryAdapter, once per gated read
syrf.project_statistics.materialized_reads counter family, scope_type, request_kind the same adapter, when the projection served
syrf.project_statistics.materialized_read.duration histogram (ms) family, scope_type, request_kind the same adapter: the bundle read plus its mapping
syrf.project_statistics.authoritative_fallbacks counter family, scope_type, request_kind, reason, reader_reason the same adapter, when it answered authoritatively
syrf.project_statistics.authoritative_aggregations counter as above the same adapter, only when the fallback really ran an aggregation
syrf.project_statistics.authoritative_aggregation.duration histogram (ms) as above the same adapter, on the same condition
syrf.project_statistics.source_operation_receipts.committed counter outcome (point_path / source_only) the source-write owners, after the transaction's commit is confirmed
syrf.project_statistics.source_operation_receipts.deduplicated counter outcome (point_path / source_only) ResolveByReceiptAsync, when a receipt proves the operation already committed
syrf.project_statistics.source_operation_receipts.rejected counter outcome (refused), reason either coordinator entry point, on any typed ProjectStatisticsWriteRejectedException
syrf.project_statistics.rebuilds.requested counter family, operation, scope_type (scope rebuilds only) ProjectStatisticsRebuildService.RebuildScopeAsync and ProjectScreeningBackfillService.BackfillAsync
syrf.project_statistics.rebuilds.completed counter as above the same two, on Succeeded
syrf.project_statistics.rebuilds.failed counter as above plus reason the same two, on every other status
syrf.project_statistics.rebuild.duration histogram (ms) as above plus outcome (succeeded / failed) the same two
syrf.project_statistics.parity.audits counter family, outcome, reader_reason ProjectScreeningParityAuditService.AuditAsync
syrf.project_statistics.parity.mismatches counter as above the same service, only on mismatch
syrf.project_statistics.drift_check.comparisons counter family, outcome (the parity outcomes), reader_reason the #3636 periodic drift check, once per scope (never per retry); kept apart from parity.audits so the screening parity-audit soak evidence is not mixed with scheduled checks
syrf.project_statistics.parity.invalidations counter family the periodic drift check, once per drifted served row it made Stale
syrf.project_statistics.drift_check.projects counter outcome (passed / failed / unverified / continuing / inconclusive / refused / conflict / error) ProjectStatisticsDriftCheckRunner, once per project visit
syrf.project_statistics.drift_check.verifications counter verification_state (confirmed / suspect) the drift check, per batch of snapshot verification records it moved
syrf.project_statistics.stale_repair.projects counter outcome (inspected / skipped), reason (the project-level skip reason, when skipped) ProjectStatisticsStaleRepairRunner (FEAT-024 C1), once per project visit
syrf.project_statistics.stale_repair.families counter family, outcome (current / repaired / refused / configuration_refused / bootstrap_not_confirmed / skipped / failed / deferred / backed_off), reason (trigger or skip reason), failure_reason (typed backfill refusal) the same runner, once per inspected family
syrf.project_statistics.maintenance.passes counter operation (history_maintenance / delta_maintenance / receipt_maintenance), outcome (completed / refused / capacity_refused / failed / skipped), reason (typed refusal, unexpected, or the disabled switch, including scheduled_receipt_maintenance_disabled) ProjectStatisticsScheduledMaintenanceRunner, once per scheduled pass
syrf.project_statistics.maintenance.reclaimed counter operation, item (roots_retired / roots_deleted / reference_pages / observations / build_markers / delta_revisions / receipts / unchanged_days / checkpoint_verifications) the same runner, per completed pass that reclaimed anything
syrf.project_statistics.daily_observations.scopes counter family, provenance (copied / calculated) ProjectStatisticsBackfillService, once per family per ordinary daily build that reached calculation, with the option on or off (#3636)
syrf.project_statistics.pending_entries.appended counter none the fold-path save owners (screening, annotation, reservation), once per save that appended a pending entry
syrf.project_statistics.pending_entries.overflowed counter outcome (overflow_recorded / overflow_unrecorded) the same owners, once per save whose change landed as an overflow drop (32 entries or 48 KiB per Study). It has no receipt
syrf.project_statistics.fold.save_outcomes counter outcome (saved, saved_without_entry, duplicate, unknown_result_saved, digest_mismatch, unresolved, not_found, at_capacity, refused, retries_exhausted, not_on_fold_path) the screening and annotation fold saves, once per save
syrf.project_statistics.fold.reviewer_fallbacks counter reason (no_threshold / no_evidence / membership_bound) fold-path screening saves that staled the reviewer families instead of moving them
syrf.project_statistics.fold.batches / fold.entries counters none ProjectStatisticsFoldWorker, per committed fold transaction and the entries it consumed
syrf.project_statistics.fold.discarded counter family, reason the worker, per family that had moves it could not apply and staled instead (intent-only families never appear)
syrf.project_statistics.fold.quarantined counter reason (unknown_protocol, unknown_kind, undeserializable, duplicate_receipt, conflicting_receipt, duplicate_delta, retired_source_revision, quarantine_exists, overflowed, unrecorded_drops, livelocked) the worker, per quarantine record a committed fold inserted. A quarantined entry has no receipt
syrf.project_statistics.overlay.invariant_violation counter family the reader overlay, when an overlaid row broke an invariant (it then falls back)
syrf.project_statistics.fold.lag histogram (ms) none the worker, per consumed entry: from the entry's creation to the fold commit
syrf.project_statistics.fold.worker_heartbeat_age histogram (s) none the fold sweep, per run: the current-protocol heartbeat age (above two minutes means no worker)

Four conventions are worth knowing before you write a query.

  • reason and reader_reason are different enums and both matter. reason is the consumer's own ProjectScreeningStatisticsFallbackReason — including the decisions taken before any storage was touched, such as writes_disabled or project_not_allowlisted. reader_reason is the bundle reader's ProjectStatisticsFallbackReason, which stays none for those pre-snapshot decisions. reason=project_not_allowlisted means the pilot is not on; reason=bundle_fallback, reader_reason=stale means it is on and the projection is behind. Those are the two answers the soak exists to separate. The same pair appears in the adapter's Debug log line, /-separated. The parity counters publish the reader's enum under reader_reason too, so one enum is never split across two tag names.
  • authoritative_fallbacks and authoritative_aggregations are not the same number, by design. The first counts reads that fell back; the second counts authoritative calculations actually executed. ReviewController.GetFullStats runs the one 15-facet pipeline for the whole response before consulting the adapter and hands the screening section over, so its fallbacks move the first counter and not the second — and contribute no sample to authoritative_aggregation.duration, which is one of the five histograms the plan's p50/p95 gate reads. Only the screening-overview consumer, which passes a delegate that runs the query, moves both. Read the difference between the two counters as "fallbacks that cost nothing extra".
  • source_operation_receipts.committed counts confirmed commits, not attempts. The coordinator writes inside the caller's still-open transaction and a TransientTransactionError aborts and replays it, so the source-write owners record this only after ProjectStatisticsTransaction.CommitWithUnknownResultRetryAsync returns. Refusals (...rejected) are counted where they are raised, because a refusal is terminal for its transaction by contract.
  • operation separates a sweep from a scope. scope_rebuild is one scope's rebuild; backfill and forced_rebuild are the administrative sweeps, which run many scope rebuilds underneath and therefore also produce scope_rebuild rows. A sweep carries no scope_type — it spans every scope kind the family declares, so there is no honest single value — which means a query filtered on scope_type silently drops sweeps. Sum the two separately.
  • parity.audits{outcome} has four values and only one of them is a defect. in_parity (conclusive and equal), mismatch (conclusive and unequal — alert on this), inconclusive (the source moved between the two halves), and unavailable (the projection legitimately refused: Stale, fenced, rebuilding, epoch-mismatched, disabled — which is also what every audit reports before step 5 opens the narrow gate, and after rollback closes it).

Reading them during the soak

There is still no OTLP destination. AddOpenTelemetryConfig exports only when OTEL_EXPORTER_OTLP_ENDPOINT is set, and that variable appears nowhere in the charts or in cluster-gitops. There is no Prometheus scrape endpoint either — neither host calls MapPrometheusScrapingEndpoint. So the series are produced but not shipped, and the two ways to read them are:

  1. dotnet-counters, in the pod. This needs no deployment change and is the baseline. Both Dockerfiles' final stage is mcr.microsoft.com/dotnet/aspnet:10.0 — the runtime, not the SDK — so dotnet tool install is not available in the container. Copy the single-file build in instead:
# once, on your workstation
curl -sSL https://aka.ms/dotnet-counters/linux-x64 -o dotnet-counters && chmod +x dotnet-counters

POD=$(kubectl -n syrf-staging get pod -l app=api-staging-syrf-api -o name | head -1)
kubectl -n syrf-staging cp dotnet-counters "${POD#pod/}:/tmp/dotnet-counters"
kubectl -n syrf-staging exec -it "${POD#pod/}" -- /tmp/dotnet-counters ps
kubectl -n syrf-staging exec -it "${POD#pod/}" -- /tmp/dotnet-counters monitor \
  --process-id <pid from ps> \
  --counters SyRF.ProjectManagement.ProjectStatistics

Verify this end to end on one pod before the soak starts rather than discovering at day seven that it does not attach: the diagnostic IPC socket lives under /tmp and belongs to the app's user, so a container running as a different user, with a read-only root filesystem, or with DOTNET_EnableDiagnostics=0 will refuse. Substitute the project-management deployment for the write, rebuild and backfill instruments — the API host records those only for operations it initiates. Use collect --format csv --output <file> rather than monitor when you want a series to keep as evidence, and kubectl cp it back out. dotnet-counters reports per process, so a multi-replica deployment needs one session per pod and the totals summed; record the replica count in the evidence file alongside the numbers, and remember a rolling restart resets every counter.

  1. Set OTEL_EXPORTER_OTLP_ENDPOINT in the staging values and point it at a collector. That is a cluster-gitops change and a decision outside this runbook — do not make it as part of running the proof. If it has been made by the time you run the soak, everything above is already flowing and dotnet-counters becomes a cross-check rather than the primary source.

Whichever you use, put the raw counter output in the evidence file, not a summary: a fallback total without its reason/reader_reason breakdown cannot be re-interpreted later, and that breakdown is the finding.

Fallback reasons and log fields

Read this section before planning the soak: what is observable is narrower than it looks.

The metrics exist; the pipe to a dashboard does not

ProjectStatisticsTelemetry.cs used to declare the instrument names and nothing else, so this section used to read "there are no metrics". Since syrf #3482 the meter is real and every instrument in Telemetry records. What is still absent is a destination: AddOpenTelemetryConfig exports only when OTEL_EXPORTER_OTLP_ENDPOINT is set, that variable appears nowhere in the charts or in cluster-gitops, and neither host exposes a Prometheus scrape endpoint.

So you can still not plan a dashboard query — but you can read the counters per pod with dotnet-counters, which is the more reliable of the two sources because it does not depend on a log level or a log pipeline. The logs below remain the way to tie an individual observation to a project id, which no metric carries.

While the pilot is dark, only the startup line is emitted

ReviewController.GetFullStats calls IsMaterializedReadRequested(projectId) first, which evaluates the flag and allowlist gate and logs nothing. If any of materializedProjectStatisticsWrites, ...Serving, ...Screening is off, or the project is outside the allowlist, the adapter is never entered and no per-read FEAT-024 line is emitted.

Since #3482 that is no longer indistinguishable from a misconfigured pod. Each host writes exactly one Information line at startup, from SyRF.ProjectManagement.Core.Telemetry.ProjectStatisticsDarkStateStartupLog:

FEAT-024 materialized project statistics startup state: serving requested {MaterializedStatisticsServingRequested}; flags {@MaterializedStatisticsFlags}; allowlisted projects {MaterializedStatisticsAllowlistSize}; meter {MaterializedStatisticsMeterName}.

MaterializedStatisticsFlags carries all twelve materializedProjectStatistics* values as this process resolves them — through the API's runtime-override-aware adapter in the API host, and the shared FeatureFlags singleton in project-management — with the serving entry already conjoined with writes, because serving a projection nothing maintains is never valid. MaterializedStatisticsServingRequested is the same conjoined value, not the raw catalogue flag. MaterializedStatisticsAllowlistSize is a count: the allowlisted project ids are deliberately never logged, because a rollout log is not an access-controlled surface. MaterializedStatisticsMeterName names the meter to attach dotnet-counters to — and resolving it to write the line is also what forces the meter to exist on a freshly rolled pod, rather than at the first statistics read.

Capture both hosts' lines in the evidence file at the start of the soak and again at the end. They are the record that the deployment was in the state the runbook says it was, and they are what distinguishes "the flags were off" from "the pod never restarted after the cluster-gitops sync". Grep them with FEAT-024 materialized project statistics startup state.

Once the gate passes, the adapter logs on every read as well: exactly one of the materialized-serve or fallback lines per call.

The two fallback enums

They are different types and both appear in one log line, separated by a literal /.

ProjectScreeningStatisticsFallbackReason — the adapter's reason, arriving with #3196:

Member Meaning
None the projection served
WritesDisabled materializedProjectStatisticsWrites off
ServingDisabled materializedProjectStatisticsServing off
FamilyDisabled materializedProjectStatisticsScreening off
ProjectNotAllowlisted project outside the pilot allowlist
NoCaller no caller identity to authorize against
NotAvailable bounded not-available: unauthorized, unknown or foreign selection
BundleFallback the reader decided the whole bundle is authoritative
BundleUnavailable the reader refused — visibility token or durable-mode disagreement
CapacityExceeded a bounded capacity ceiling
ScopeTombstoned the selection landed on an explicit deletion tombstone
ScopeRowMissing the materialized bundle carried no row for the scope
CheckpointUnresolved an all-zero (source revision, projection revision, mode epoch) tuple
ReaderFailed the read path threw

ProjectStatisticsFallbackReason — the reader's reason, already on main: None, Missing, Stale, Rebuilding, Incompatible, Disabled, EpochMismatch, Fenced, SnapshotPredicateFailed, InclusionRecalculationInProgress, DefinitionRewriteInProgress, DurableModeDisagreement, ProjectUnavailable.

Both render as their exact PascalCase member names — no attributes change the string form. The fallbackReason field of the parity report carries the reader enum's string.

For the soak, the reasons that should be rare are the interesting ones: a steady stream of BundleFallback/Stale, BundleFallback/Rebuilding, BundleFallback/Incompatible, .../EpochMismatch or .../Fenced is precisely what the soak is for. WritesDisabled, ServingDisabled, FamilyDisabled and ProjectNotAllowlisted mean the pilot is not actually on.

The log statements

All from ProjectScreeningStatisticsQueryAdapter (category SyRF.ProjectManagement.Core.Services.ProjectStatistics.Families.Screening.ProjectScreeningStatisticsQueryAdapter) except the last, which is from ReviewController.

Level Template
Debug FEAT-024 answered the project-screening section authoritatively for project {ProjectId}: {FallbackReason}/{ReaderFallbackReason}.
Debug FEAT-024 served the project-screening section from the projection for project {ProjectId} at checkpoint {CheckpointId}.
Error FEAT-024 project-screening materialized read failed for project {ProjectId}; answering from the authoritative screening facets.
Warning FEAT-024 declined to substitute the materialized screening section for project {ProjectId} at checkpoint {CheckpointId}: the materialized totals and the authoritative screening values differ; retaining the coherent response. Answering authoritatively.

The fallback line is Debug on purpose: while the flags are off every request takes that path, and a dark rollout that fills the log with warnings gets its warnings ignored.

The Warning line is the parity-divergence alarm. It is the only Warning-level FEAT-024 read-path signal and the one worth alerting on during the soak. A single occurrence is a finding: capture the project id and checkpoint, run the parity audit immediately, and record both.

Naming the query

Since syrf #3515 the format is a deployment choice. logging.format in the environment's values is mapped to SYRF__Logging__Console__FormatterName, and the API, project-management and Quartz hosts select their console formatter from it: json emits one JSON object per event, anything else, including unset, renders plain text exactly as before. Identity was always JSON and does not read the variable. Staging already carries logging.format: json, which before #3515 was inert; confirm the pod actually has it before writing a field query:

kubectl -n syrf-staging get deploy/api \
  -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="SYRF__Logging__Console__FormatterName")].value}'

With json, the shape is Serilog's JsonFormatter with renderMessage: true, so every structured property is a field under Properties, the message template is preserved verbatim, and RenderedMessage still carries the sentence the plain-text formatter would have printed:

{"Timestamp":"2026-09-15T09:14:02.1234567+00:00","Level":"Debug","MessageTemplate":"FEAT-024 answered the project-screening section authoritatively for project {ProjectId}: {FallbackReason}/{ReaderFallbackReason}.","RenderedMessage":"FEAT-024 answered the project-screening section authoritatively for project \"…\": \"BundleFallback\"/\"Stale\".","Properties":{"ProjectId":"…","FallbackReason":"BundleFallback","ReaderFallbackReason":"Stale","ThreadId":12}}

RenderedMessage is why a JSON pod is not a downgrade for anyone reading logs by eye: every substring in the plain-text table below still matches, inside that field (jq -r '.RenderedMessage' gives back the pre-#3515 view). Query by Properties.<Name> for anything you want to count.

Fields to query, one JSON object per line, so jq over kubectl logs is enough without a log aggregator:

Purpose Query
fallbacks, with reason .Properties.FallbackReason on the "answered … authoritatively" template
reason histogram for the soak jq -r 'select(.MessageTemplate \| startswith("FEAT-024 answered")) \| "\(.Properties.FallbackReason)/\(.Properties.ReaderFallbackReason)"' \| sort \| uniq -c
materialized serves .MessageTemplate starts with FEAT-024 served the project-screening section
parity divergence — alert .MessageTemplate starts with FEAT-024 declined to substitute (also .Level == "Warning")
host dark-state at startup .Properties.MaterializedStatisticsFlags — the whole flag map as an object, plus .Properties.MaterializedStatisticsAllowlistSize and .Properties.MaterializedStatisticsMeterName
read-path exception .Level == "Error" on the FEAT-024 project-screening materialized read failed template

A worked example — the fallback-reason histogram over the last day:

kubectl -n syrf-staging logs deploy/api --since=24h \
  | grep '^{' \
  | jq -r 'select(.MessageTemplate | startswith("FEAT-024 answered")) |
      "\(.Properties.FallbackReason)/\(.Properties.ReaderFallbackReason)"' \
  | sort | uniq -c | sort -rn

The dark-state flag posture at both ends of the soak, as data rather than as prose:

kubectl -n syrf-staging logs deploy/project-management --since=24h \
  | grep '^{' \
  | jq 'select(.MessageTemplate | startswith("FEAT-024 materialized project statistics startup state")) |
      {at: .Timestamp, allowlist: .Properties.MaterializedStatisticsAllowlistSize,
       flags: .Properties.MaterializedStatisticsFlags}'

The grep '^{' is not decoration, though the reason is narrower than it looks: the bootstrap logger reads the same switch, so its startup lines are JSON. What is not JSON is the one banner each host writes straight to stdout before any logger exists — Console.Out.WriteLineAsync("… is starting…") at SyRF.API.Endpoint/Program.cs:63, SyRF.ProjectManagement.Endpoint/Program.cs:37, SyRF.Quartz/Program.cs:16 and SyRF.Identity.Endpoint/Program.cs:24 — plus anything the .NET runtime itself prints on a crash. jq aborts on the first such line, so filter them out.

Without json — any environment that has not opted in, and every local run — Serilog's MessageTemplateTextFormatter renders the properties into the message string, so FallbackReason="Stale" matches nothing and you must match the message text and parse the trailing <Guid>: <Reason>/<ReaderReason>. Substrings to query in that case:

Purpose Substring
every gated read (count for the soak) FEAT-024 and the project-screening section
fallbacks, with reason FEAT-024 answered the project-screening section authoritatively for project
materialized serves FEAT-024 served the project-screening section from the projection for project
parity divergence — alert FEAT-024 declined to substitute the materialized screening section
host dark-state at startup (capture at both ends of the soak) FEAT-024 materialized project statistics startup state
read-path exception FEAT-024 project-screening materialized read failed for project

The baseline access is pod stdout, e.g. kubectl -n syrf-staging logs deploy/api --since=24h | grep 'FEAT-024'.

Three preconditions before any of the per-read lines yields output:

  • 3196 merged and deployed — these identifiers do not exist on main;

  • the pilot flags on and the project in the allowlist, or no per-read line is emitted at all (the startup line is emitted regardless, and is Information);
  • the effective Serilog minimum level at Debug, for the two Debug lines.

On that last point: SYRF__Serilog__MinimumLevel comes from .Values.logging.level, and staging sets level: debug in lowercase, whereas production carries the comment "Serilog requires capitalized full word". Serilog's enum parse is case-insensitive so it should bind, but verify it live before depending on the Debug lines — confirm a known-Debug line appears in kubectl logs before starting a soak whose read count comes from them.

What is centrally visible

  • Sentry is enabled in staging (monitoring.sentry.enabled: true, environment: staging), so the Error and Warning lines above surface there. The two Debug lines do not, and the Information startup line is below the event threshold too — read it from pod stdout.
  • Elastic APM is disabled in staging (elasticApm: { enabled: false }).
  • No OTLP, and no Loki/Elasticsearch/Fluent/Promtail/Alloy configuration exists in cluster-gitops. docs/architecture/system-overview.md describes "Aggregation: Loki or CloudWatch / Format: Structured JSON"; since syrf #3515 the format half of that is real in any environment that sets logging.format: json, but the aggregation half is still aspirational — there is no collector. Queries in this runbook therefore run over kubectl logs, and you should confirm the actual cluster log pipeline out of band before writing one against a named aggregator.

So, concretely: alert on Sentry for the divergence Warning; once the flags are on, count reads from the syrf.project_statistics.requests counter with dotnet-counters, cross-checked against pod stdout; and take fallback-reason and revision observations from the parity endpoint, whose fallbackReason, readSource, capacityFailure and watermarks fields give the same information without depending on log levels, a log pipeline or a pod's uptime.

Counting mutations

Since syrf #3482 there is a mutation counter: syrf.project_statistics.source_operation_receipts.committed, split by outcome into the point path and the source-only fallback, read per pod with dotnet-counters (see Telemetry). It is recorded only after each transaction's commit is confirmed, so a retried transient abort is not double-counted — but it resets on a restart and is per process, so it is a corroborating series rather than the primary evidence. It also counts only mutations that reached a statistics commit, which while the write flag is off is none of them.

Control revisions establish ordering and freshness, not the number of source mutations. ProjectStatisticsRebuildService advances the projection clock with sourceChanged: false, projectionCommitted: true, so repeated rebuilds can increase committedProjectionRevision without any review action. Source invalidation and client invalidation revisions also describe protocol transitions, not an interchangeable count of user actions.

For source-operation evidence, capture committed pmProjectStatisticsSourceOperationReceipt rows inside the approved observation window, before retention removes them. Deduplicate by (ProjectId, OperationNamespace, OperationId) and require an advancing CommittedSourceBeforeRevision → CommittedSourceAfterRevision. Count one committed transaction, not the revision difference, number of changed Studies, or retries. Use committed/majority reads; ObservedAtUtc must fall inside the run window. Missing receipts leave a gap to investigate, not a licence to replace them with projection-revision differences. The post-commit counter is corroborating evidence only after its per-instance resets and collection gaps are accounted for.

Poll the parity endpoint separately and retain every response, including inconclusive attempts. Parity polls are not consumer-read load and cannot by themselves prove mutation volume or absence of stale serves on ordinary page requests.

Offline evidence accounting

The standard-library validator reads captures and writes a JSON summary; it performs no network, database or runtime operations and needs no feature flag:

python3 scripts/validate-statistics-evidence.py /private/run/evidence.json \
  --reads /private/run/reads.jsonl \
  --receipts /private/run/receipts.jsonl \
  --parity /private/run/parity.jsonl > /private/run/summary.json
python3 scripts/test-statistics-evidence.py

Use the existing evidence template. Fill PilotProject.ProjectId and UTC Soak.StartedAtUtc / Soak.EndedAtUtc; the end cannot be in the future. Captures outside that window are rejected. Submitted ReadCount, MutationCount and ElapsedDays totals are ignored: the summary calculates them from captures. The manifest is limited to 256 KiB; each capture file to 256 MiB, 100,000 records and 256 KiB per JSON line. Keep raw captures private; the summary contains counts and gate results without project/reviewer/operation identities.

Each read JSONL record carries the actual request ID, GET method, request time, pilot project ID, HTTP status and unchanged JSON response body. Give each actual HTTP attempt a unique RequestId; copying a capture does not count again. For example, a consumer capture has this envelope:

{
  "RequestId": "captured-request-id",
  "RequestedAtUtc": "2026-09-21T15:32:11Z",
  "ProjectId": "00000000-0000-0000-0000-000000000102",
  "HttpMethod": "GET",
  "HttpStatus": 200,
  "ResponseBody": {}
}

The example is an envelope with an empty response placeholder, not recorded proof. Preserve the consumer's real response in ResponseBody. Optional Provenance carries readSource, fallbackReason and scopeStates from that same read's server evidence. Many current DTOs expose no provenance: omit it in that case, and the stale-serve gate remains unproven. Never attach a different parity poll as if it described the consumer request. A successful read counts toward the 10,000-read threshold; a rejected HTTP request does not.

Each parity record uses the same envelope with the actual parity response as ResponseBody. Keep parity captures in their separate input, including failed/inconclusive attempts. Acceptance requires Fresh, published, non-tombstone scopes, current configuration, a materialized answer and exact nonempty integer metric comparisons. An inParity: true summary cannot hide a mismatching metric or an inconclusive audit. Parity polls cannot count toward consumer-read volume.

Each receipt JSONL record has CapturedAtUtc and the unchanged exported receipt under Document. The required receipt fields are ProjectId, OperationNamespace, OperationId, ObservedAtUtc, CommittedSourceBeforeRevision and CommittedSourceAfterRevision. Preserve committed receipt identities across capture retries. Mongo extended JSON $date, $numberLong and C# legacy UUID $binary subtype 03 are accepted; subtype 04 is rejected. Non-advancing source receipts and conflicting duplicates fail validation. Bulk receipts count as transactions, not their number of Studies. The validator checks supplied evidence; it does not authenticate captures or prove collection completeness.

For already-authorized fallback drills, mark the associated read with InjectedFallbackScenario: Stale, Rebuilding, Incompatible or KillSwitch. All four need a successful AuthoritativeFallback outcome and a non-None reason. Merely recording that a switch changed is not fallback evidence. This tool neither injects faults nor changes switches.

The summary reports each gate as passed, failed or unproven. Capture span is calculated from actual supplied timestamps, not the claimed run duration. Two observations seven days apart do not prove sustained operation, so continuous coverage remains unproven. Performance, storage/retention, rebuild backlog, rollback and owner approval require their separate evidence and remain unproven here. programmeAccepted is always false: this accounting tool cannot authorize a rollout. Exit 0 means valid captures with no failed evaluated gate, not completed acceptance; exit 1 reports failed evidence/gates, and exit 2 reports unreadable or structurally invalid input. Missing proof remains visible in the gate statuses even on exit 0.

Read performance gate — local harness

The read-performance acceptance gate asks for a controlled, representative HTTP consumer measured before and after on fixed data: the materialized page-load p95 at least 20% lower and at least 80% fewer authoritative aggregations. The repository-adapter benchmark does not close it, because it measures neither HTTP nor the pages' real request sets. e2e/harness/statistics-read-benchmark/ is the harness for the HTTP gate. It runs hermetically on the local E2E stack today and takes the options it needs to point at a deployed environment later.

What it measures

  • Workload. workloads.mjs lists the statistics GETs each page issues on a direct entry, per arm: Project Overview, Stage Overview, Screening Overview, Stage Review progress, Question management and Searches. It covers /screening-stats, the Stage Overview bundle, project and stage reviewer progress, question-answer counts and search counts, plus the legacy full-stats read that the route guards still issue. The lists come from reading the Angular route guards and route-scoped stores; the file names each source. A browser capture has not yet checked them (see follow-ups). Non-statistics requests, such as project details, are the same in both arms and are left out. A page load sends its requests concurrently, as the browser does. Its latency is the time until the last response body arrives.
  • Arms. Both arms run on one API process and one dataset. The deployment's own configuration decides which consumers are on; the --feat024-read-bench deployment does not enable screening history, so neither arm exercises it. The legacy arm uses the operator's own rollback: the statistics-consumers-v2 runtime group is set Off, then every remaining page consumer. The materialized arm clears those overrides (null for the group's member keys and the page consumers) rather than setting the group On, because On would write explicit overrides for every member, including history. The deployment configuration then governs. It runs only after all families are backfilled, the narrow gate is open and the parity endpoint reports Materialized and in parity. Before and after each block, a probe of the Stage Overview bundle must show the expected arm: enabled=false for legacy, or readSource=Materialized for materialized. Every sampled materialized response is also checked where the payload carries its source. A sample that fails or served the wrong source makes the receipt ineligible.
  • Phases. Each block discards --warmup loads of every page. It then runs --iterations sequential loads of every page and the same number again through --concurrency virtual users. The default order LMML (ABBA) spreads host drift across both arms.
  • Authoritative aggregations. The API's syrf.project_statistics.authoritative_aggregations counter only moves on the fallback path inside a materialized consumer. A consumer that is switched off calls full-stats or the authoritative query directly, so the counter records nothing for the legacy arm and cannot measure the reduction. The harness therefore counts at the database. After the timed blocks it runs --count-passes passes of each arm's workload with the MongoDB profiler filtered to aggregate/count/distinct. It then waits through an idle window of the same length and subtracts that window's commands as background Quartz/PM work. An authoritative aggregation is an aggregating command on any collection outside pmProjectStatistics*. The receipt breaks the count down by collection and by page. The API's own FEAT-024 meter is still scraped through its existing Prometheus listener (default off; the harness mode turns it on locally) and recorded per block, including the fallback reason/reader_reason breakdown.
  • Host load. This machine is also the CI host. The receipt records uptime and the load average at the start, after every block and at the end. The check is the 1-minute load divided by the CPU count, against a threshold of 0.5 (--max-load-per-cpu). An acceptance run refuses to start above the threshold unless --allow-loaded-host is given. Crossing the threshold at any point makes the receipt not acceptance evidence.

Fixtures

The data is generated by the PM E2E setup endpoint through the production CSV import pipeline. It is synthetic Bogus content, with no clinical data. The project has three seeded E2E members, a screening-and-annotation stage and two annotation questions. Reviewers save real screening decisions through the API. The fixture is then activated exactly as in the route proof: fleet Enabled, eight family backfills, then the narrow gate.

Fixture Studies Screening decisions Screened / dual-screened Use
smoke 100 90 60 / 30 Proves the harness works; not a performance fixture
standard 5,000 2,200 1,400 / 800 The representative acceptance fixture: same target-study count as PS-DS-02

Neither fixture has annotation answers, so both annotation arms read empty tallies. A heavier annotation fixture is a follow-up. The zero-study route proof's fixture is not used.

The setup retries each step a bounded number of times. Every retry is logged as [FEAT024_READ_BENCH_SETUP_RETRY] and recorded in the receipt's dataset.activation. A FamilyFenced 409 is the route proof's existing case. On the seeded fixture a backfill's publish can collide with a concurrent update to the project's notification-outbox slot (Write conflict during plan execution from ProjectStatisticsRebuildService.CoalesceSlotAsync). The API now re-runs that publication transaction itself and, only if the conflict outlasts its bounded retry, answers a typed 409 with failureReason: "RetryExhausted" (#3826). The harness waits 10 seconds after the decision writes (--settle-seconds), retries a RetryExhausted backfill at most three times, and no longer retries a 500: a 500 is now an unexpected fault and fails the setup. The parity read that follows is what proves the result.

How to run it

The command runs from the PR worktree. It needs the usual E2E prerequisites and a Release build (drop --skip-build to build). Everything after -- goes to the harness.

# Harness smoke (about 10 minutes, mostly stack start-up; small N; never acceptance evidence)
bash e2e/run-local.sh --feat024-read-bench --skip-build -- --mode smoke

# Full measurement. Wait for an idle host: this machine runs CI, and the run refuses above 0.5 load/CPU
bash e2e/run-local.sh --feat024-read-bench --skip-build -- --mode acceptance
#   defaults: --fixture standard --iterations 100 --warmup 10 --concurrency 8 --order LMML --count-passes 3
#   override any of them, for example: -- --mode acceptance --iterations 200 --concurrency 16

Receipts are written to e2e/test-results/statistics-read-benchmark/receipt-<mode>-<UTC>.json, or to --out <path>. Each receipt records the schema version, UTC start and end, the git SHA, branch and dirty flag, the target, the configuration and thresholds, and the fixture description. It also records the seeding and activation steps with their retries, the parity result and the arm probes. Results include page-load p50/p95/p99 per arm and phase, broken down by page and by request, the aggregation counts with their collection and page breakdowns, the app-metric deltas, the host-load snapshots and the gate. gate.verdict has four values. pass or fail appear only on acceptance-eligible runs. not-acceptance-evidence covers a smoke run, a loaded host or a failed sample. single-arm appears when one arm was measured. An acceptance run exits 0 only on pass and exits 2 otherwise. A smoke run exits 0 whenever the harness completed.

The first smoke receipt is evidence/http-read-benchmark/SMOKE-2026-09-29.json. It is a smoke run on a loaded host and is not acceptance evidence: 100 studies, 5 iterations, 2 concurrent users, and a load of 0.50–0.63 per CPU. It was run on commit e51d394 (clean worktree), 11:11:27–11:12:00 UTC. It shows that the harness seeds, activates, switches and verifies both arms, times them and counts aggregations end to end, with zero failed samples.

Its timings say nothing about the gate. At 100 studies the legacy aggregations are cheap, so page-load p95 was roughly equal in the two arms (50.6 versus 54.6 ms sequential, 59.5 versus 55.6 ms concurrent). Several materialized requests were individually slower than their legacy counterparts at this size.

Its aggregation counts show the counting path works. The legacy workload ran 83 authoritative pmStudy aggregations per pass: 12 per page load and 23 on Stage Review. The materialized workload ran 3: 2 on Question management and 1 on Searches, the two pages whose route guards still load full stats in both arms. Per page, the counts suggest one aggregation for full-stats and about eleven for each legacy reviewer-progress read. They also suggest that the materialized question-count read still ran one pmStudy aggregation. The acceptance run should confirm that attribution per request before anyone relies on it.

Pointing it at a deployed environment later

No option below has been exercised against staging. Each is a precondition, not a claim.

  • --setup existing --project-id <id> --stage-id <id> measures an existing enrolled project instead of seeding one. Record that project's study and decision counts beside the receipt.
  • --base-url https://<api-host> with --auth bearer-env (SYRF_BENCH_BEARER) or --auth cookie-env (SYRF_BENCH_COOKIE) authenticates as an operations or service identity, never a personal login. The identity must be able to view the project. Runtime arm control also needs the SyRF administrator groups.
  • --arm-control runtime-overrides changes the environment's runtime flags for everyone reading that project for the whole run. It is a staging change that needs separate approval. Without that approval, use --arm-control none --arms materialized (or legacy) for a single-arm receipt. Comparing two single-arm receipts across time is weaker evidence than one paired run.
  • Aggregations: there is no profiler access to the Atlas preview cluster from this harness. --aggregations app-metrics --metrics-url <scrape URL> reads the API meter. As described above, though, that counter cannot see the legacy arm, so the aggregation half of the gate stays unmeasurable on staging until one of two things exists. One is an Atlas profiler or slow-query export for the run window. The other is a default-off counter on the legacy authoritative paths. Adding either is a separate, flagged change.
  • Host load is then the harness host's load, not the server's. Record the API replica count and the node's load separately.

Operator identity

Operations staff, and Claude through a GitHub workflow, run the FEAT-024 statistics administration operations as a dedicated machine identity instead of anyone's personal login. SyRF Identity can seed a confidential client, syrf-statistics-operator, whose only grantable scope is statistics:operate, requested with the client-credentials grant. Off by default, staging only.

Where to find results. The workflow log carries only the operation, method, path and HTTP status. The sanitised response (the same data as the job summary: known metadata fields only, no investigator ids, no free text) is also uploaded as the artifact statistics-operator-<run_id> (statistics-operator-report.json, kept 7 days), so automation can fetch it with gh run download <run_id> -n statistics-operator-<run_id>; step summaries are not readable through the API.

What it can do. Exactly these routes, and nothing else:

Operation Route
Pending index check / build GET / POST /api/admin/project-statistics/fold/pending-index
Fold status GET /api/admin/project-statistics/{id}/fold/status
Fold enable / disable / reset POST /api/admin/project-statistics/{id}/fold/{enable,disable,reset}
Fold stamp advance (slice 6) POST /api/admin/project-statistics/{id}/fold/advance-stamp
Backfill / forced rebuild POST /api/admin/project-statistics/{id}/backfill, …/{id}/rebuild and …/{id}/{family}/{backfill,rebuild} for the eight further families
Parity read GET /api/admin/project-statistics/{id}/parity
  • Authentication gate. DirectBearerSchemeSelector accepts the scope on those methods and paths only; on every other route (mode, narrow gate, fleet, maintenance, pilot status, fence diagnostics, user administration, every project, study and account route, the notifier and the SignalR hub) the token is not a valid API token and the request is anonymous (401). Pinned by StatisticsOperatorTokenRoutingTests, which sweeps every declared API action.
  • Authorization. The three operator controllers carry ProjectStatisticsOperatorPolicy: the operator client, or a human evaluated exactly as BatchAdminProjectsPolicy evaluates them (human administrator access is unchanged). The parity read's policy also admits it. Every other policy — every project, stage and application policy — refuses it (403), so it grants no general administration or project access even behind the first gate. A token that also carries syrf_api, the read-only evidence scope, a SyRF user id or groups, or that was not produced by OpenIddict introspection, is not the operator. The read-only evidence credential cannot call any operator route except the parity read it already had.
  • Audit. The machine is recorded as service:syrf-statistics-operator: a backfill's lease owner becomes stats.backfill.screening:service:syrf-statistics-operator (administrators stay …:admin:{investigator id}), the parity audit runs under a fixed non-user caller with no groups, and every fold transition and pending-index build request is logged with its actor (Statistics fold {operation} for project {id} by {actor}: {outcome}).
  • The allowlist, flags and fold gates still apply. The operator passes the same typed refusals an administrator does (unlisted project, writes or fold flag off, rollout not confirmed, coverage, pending index).

Threat model. A leaked secret lets its holder run these statistics operations on staging until the secret is removed: rebuild or backfill projections, toggle fold mode for an allowlisted project and start the pending-index build. It cannot read or change study, citation, annotation, user or membership data, open serving gates or change flags. Removing the secret revokes the registration when Identity next starts; already-issued tokens stay valid for up to an hour (as for the evidence credential above).

Enabling it (off by default; staging only). Identity seeds nothing unless a secret is configured; an absent, blank, CHANGEME or shorter-than-32-character secret is "off" and, on startup, deletes every registration holding statistics:operate. In cluster-gitops for staging:

  1. An ExternalSecret syrf-statistics-operator (key clientSecret) generated in-cluster (CreatedOnce, at least 32 characters), next to syrf-statistics-evidence in plugins/local/extra-secrets-staging/values.yaml.
  2. statisticsOperator.enabled: true in the staging Identity values (optionally statisticsOperator.clientSecret.{secretName,key}), rendering SYRF__StatisticsOperator__ClientSecret. The API needs no configuration.
  3. After the sync, copy the secret into the main-only statistics-operator environment without printing it (the environment is created first, see below):
kubectl -n syrf-staging get secret syrf-statistics-operator -o jsonpath='{.data.clientSecret}' \
  | base64 -d | gh secret set SYRF_STATISTICS_OPERATOR_CLIENT_SECRET --env statistics-operator -R camaradesuk/syrf

GitHub environments for the statistics machine credentials

Both machine secrets are environment secrets, never repository secrets. A repository secret is readable by any workflow run of any branch of anyone with push access (a dispatched branch copy of a workflow, or a same-repository pull request), whatever the job's if: says. An environment whose deployment-branch policy admits only main releases its secrets only to jobs running from main (scheduled runs use refs/heads/main and are admitted). The operate job names environment: statistics-operator and the preflight workflow's parity job names environment: statistics-evidence; validate-workflows.sh pins both, and pins that each secret is referenced by its own workflow only. One-time setup (repository settings, by an administrator):

for env in statistics-operator statistics-evidence; do
  gh api -X PUT "repos/camaradesuk/syrf/environments/$env" --input - <<'JSON'
{"deployment_branch_policy":{"protected_branches":false,"custom_branch_policies":true}}
JSON
  gh api -X POST "repos/camaradesuk/syrf/environments/$env/deployment-branch-policies" \
    -f name=main -f type=branch
done
# The evidence secret already exists in-cluster; copy it without printing it.
kubectl -n syrf-staging get secret syrf-statistics-evidence -o jsonpath='{.data.clientSecret}' \
  | base64 -d | gh secret set SYRF_STATISTICS_EVIDENCE_CLIENT_SECRET --env statistics-evidence -R camaradesuk/syrf
# Once the environment-based workflows are on main, remove the repository-level copy.
gh secret delete SYRF_STATISTICS_EVIDENCE_CLIENT_SECRET -R camaradesuk/syrf

The statistics-evidence environment and its secret must exist before the workflow change merges, or the scheduled parity capture has no secret.

Running an operation. Dispatch the Statistics Operator workflow (.github/workflows/statistics-operator.yml) from main: environment (staging only), operation (pending-index-status, pending-index-build, fold-status, fold-enable, fold-disable, fold-reset, backfill, parity), project_id (a GUID; empty for the pending-index operations), family (backfill only), confirm (required for every operation that changes state) and rollout_complete_confirmed (fold-enable only; sent as rolloutCompleteConfirmed: true only when ticked). The workflow and scripts/statistics-operator.py offer no forced rebuild and no advance-stamp, although the API accepts the credential on both; run those with an administrator token. The job runs on the juniper-ci pool with contents: read, masks the token, and writes the HTTP status and a sanitised body to the job summary: only known statistics-metadata fields are kept, free-text messages (detail) and scope keys are dropped, every GUID other than the requested project id is replaced with <id>, and long strings are truncated, so no study content or investigator id can be emitted. The same helper runs locally:

export SYRF_STATISTICS_OPERATOR_CLIENT_SECRET=...   # from the secret store, never on the command line
python3 scripts/statistics-operator.py --operation fold-status --project-id <project GUID>
python3 scripts/test-statistics-operator.py          # offline tests of the helper

Async point-fold: building the pending-statistics index

Slice 0 of the async point-fold design (ADR-019) needs one partial index on pmStudy before fold mode can ever be enabled:

IX_Study_PendingStatistics   { ProjectId: 1, "PendingStatistics.OldestAtUtc": 1 }
partialFilterExpression      { "PendingStatistics.OldestAtUtc": { $exists: true } }

It is a deliberate production operation, not a start-up side effect. Start-up index initialisation never builds it, and no feature flag or deployment setting does either. The only trigger is an administrator calling the endpoint below. Nothing writes PendingStatistics yet, so the index starts empty, but building it still scans the whole collection.

When to run it

  • Before the first fold-mode enable in an environment. Enable (the fold staging pilot) refuses with PendingIndexMissing until this check reports the index ready.
  • Staging first, then production. Staging pmStudy holds about a hundred Studies, so it proves the operation and the endpoint, not the duration.
  • Production only in a separately approved off-peak window. Running it has no effect on users beyond the build load described below.

Run it

Both routes accept an administrator or the statistics operator credential (ProjectStatisticsOperatorPolicy, see Operator identity). On staging, prefer the Statistics Operator workflow from main: operation pending-index-status (read-only) or pending-index-build with confirm ticked, and no project id. The response is in the job summary and the statistics-operator-<run_id> artifact. The equivalent direct calls:

# Read-only check. Never creates, alters or drops anything.
curl -sS -H "Authorization: Bearer $TOKEN" \
  https://api.staging.syrf.org.uk/api/admin/project-statistics/fold/pending-index

# Start the build. Idempotent.
curl -sS -X POST -H "Authorization: Bearer $TOKEN" \
  https://api.staging.syrf.org.uk/api/admin/project-statistics/fold/pending-index

The POST answers immediately; it does not wait for the scan:

Response Meaning
202 with state: 1 (Building) Build started now, or already running (started by this or another replica). Poll the GET
200 with state: 2 (Ready), isReady: true The exact definition already exists and is ready. Nothing was done
409 with reason: PendingIndexConflict An index has this name or this key pattern with a different definition. Nothing was done; detail shows the existing definition

The GET reports state as 0 Absent, 1 Building, 2 Ready or 3 Conflicting, with isReady and a detail line. lastBuildFailure carries the error of a build that this API replica started and that failed; another replica's failure shows only as the state.

The build uses commit quorum votingMembers: the index becomes ready only when every data-bearing voting member has built it, so a failover cannot land on a primary without it. A standalone server (some local setups) refuses that option and the failure appears in lastBuildFailure.

Ready is reported only from listIndexes with includeIndexBuildInfo. A plain listIndexes, and therefore Atlas's index list, shows an index whose build is still scanning exactly as if it were finished. Verified on MongoDB 7.0.40 and 8.0.

Expected duration

Production syrftest.pmStudy on 1 October 2026 ($collStats, metadata only): 3,168,933 documents, 10.9 GB uncompressed (avgObjSize 3,619 bytes), 4.6 GB on disk, 32 indexes totalling 5.6 GB. The smallest existing compound index, ProjectId_1_AllocationBucket_1, is 15 MB.

Local measurement (MongoDB 8.0, single-node replica set, 4 CPUs, 1.5 GB WiredTiger cache, synthetic non-clinical documents): 600,000 documents, 1.42 GB uncompressed, 0.58 GB on disk built in 4.3 s from a cold WiredTiger cache (1.1 s warm). Scaled linearly to production that is 25–35 s of collection scan on local NVMe.

Atlas reads from network storage and has to wait for the secondaries, so expect under five minutes and plan the window for fifteen. That is an estimate, not a measurement: record the real duration from the API log line Built IX_Study_PendingStatistics in N s, and replace this estimate with it. The design asks for a measurement on a restored production-sized copy before production; that has not been done.

Impact

MongoDB 4.4 and later build indexes with the hybrid method: an exclusive lock only for a moment at the start and the end, and reads and writes continue during the scan. The cost is disk read and CPU on every voting member while the scan runs. The finished index is tiny while nothing is pending.

If the API replica that started the build stops, the build carries on: a client disconnect does not abort a createIndexes build (verified on MongoDB 8.0). The status check still sees it. If an Atlas maintenance event interrupts the build itself, check the status; if it reports Absent, run the POST again (it is idempotent).

Verify

  1. GET reports state: 2 and isReady: true.
  2. Independently, in mongosh against the cluster:
db.runCommand({ listIndexes: "pmStudy", includeIndexBuildInfo: true })
  .cursor.firstBatch.filter(i => i.spec.name === "IX_Study_PendingStatistics")

One entry, with spec.key and spec.partialFilterExpression as above and no indexBuildInfo field. 3. Record the environment, the start and end times and the duration in the evidence log.

Stop or drop it

The index is unused until fold mode is enabled for a project, so dropping it is safe while every project is in FoldMode: Disabled. Never drop it while any project is fold-enabled.

// Aborts an in-progress build too.
db.pmStudy.dropIndex("IX_Study_PendingStatistics")

A 409 conflict needs a person to look at the existing index first: it was not created by this operation. Drop it only once you know what created it, then run the POST again.

Async point-fold: the ProjectScreening staging pilot

Slice 2 of the async point-fold design (ADR-019) moves the three screening save shapes onto the fold path for ProjectScreening (fold protocol 1; slice 3 adds the reviewer families and slice 4 the annotation families, below). A screening save in a fold-mode project writes one Study document carrying a pending entry and no statistics document; the project-management fold worker folds the entry into the stored row; the reader serves stored + pending. Slice 3 (fold protocol 2) adds MembershipScreening and ReviewerScreening: the entry carries their transition, so they are kept the same way, for projects with at most 100 members. The annotation families a screening save touches still travel as invalidation intents, so in the pilot they stay Stale and are served authoritatively, exactly as they are today after a save.

Pre-enable gate: the audit overlay (#3902). Slice 2 alone cannot audit a project with fold history: its audit reader has no pending overlay, so the ProjectScreening audit reports Unavailable for the project permanently. #3902 reads the audit through the overlay with a pending-fingerprint bracket. Enable only on an image that contains it (precondition 3); the scheduled staging preflight (scripts/capture-statistics-parity.py) then reports a real pass or fail for the pilot project.

Owner decision (g) allowed this on staging only for slices 2 to 5: a partial-coverage image refuses enable unless the host positively reports IsStaging(), the database is not syrftest and the allow setting below is on, and no runtime flag override reaches any of the three. From slice 6 the image's coverage is complete, so off production (staging, previews, local) the allow setting is no longer needed. The production database (syrftest) is still refused in code until gate (b) passes and the production rollout is separately approved; lifting that refusal is part of that rollout, and nothing in this runbook authorizes it.

Preconditions

  1. The pending index is Ready on staging (above).
  2. The pilot project already passes the Phase 2C proof: fleet and narrow gates open, the ProjectScreening row Fresh, the project in SYRF__ProjectStatistics__ProjectAllowlist on both hosts with identical values.
  3. Staging runs an image that contains slice 2 and the audit overlay (#3902) on both the api and project-management deployments.
  4. The staging application user can create indexes on pmProjectStatisticsFoldQuarantine (the api pods create its indexes on first use; a failure there fails a retried save closed), or the indexes already exist because project-management created them at start-up.

Step 1: turn the fold on in cluster-gitops

One PR to cluster-gitops/syrf/environments/staging, both hosts:

Setting api project-management Why
SYRF__FeatureFlags__MaterializedProjectStatisticsFold "true" "true" Pinned to the deployed value; cannot be flipped at runtime. Starts fleet membership, the worker heartbeat, the sweep schedule and the post-save signal
SYRF__ProjectStatistics__Fold__AllowPartialCoverage "true" not read The staging-only allow setting a partial-coverage image (slices 2 to 5) needed. From slice 6 coverage is complete and the setting is no longer read by any decision; it can be removed in a later cluster-gitops change

ASPNETCORE_ENVIRONMENT must already be Staging on the api host (it is). Merge, then wait for both rollouts to complete: ArgoCD reports Healthy, and read-only kubectl rollout status shows no pod of an older ReplicaSet alive on either deployment. Pre-fold pods write no fleet membership, so the code cannot see them; the enable call's rolloutCompleteConfirmed is your confirmation of this check.

Step 2: read the fold status

curl -sS -H "Authorization: Bearer $TOKEN" \
  https://api.staging.syrf.org.uk/api/admin/project-statistics/$PROJECT/fold/status

Before enabling, expect foldMode: "Disabled", binaryProtocol equal to the deployed image's protocol (1 for slice 2, 2 for slice 3, 3 for slice 4, 4 for slice 5), foldFlagOn: true, allowlisted: true, pendingIndexReady: true, workerHeartbeatLive: true and coverageSufficient: true. A false value names the precondition that is missing. Before slice 6, coverageSufficient also needed isStaging: true and allowPartialCoverage: true; from slice 6 those two are informational off production, and on the production database coverageSufficient is always false. The heartbeat becomes live within 30 seconds of a project-management pod starting with the flag on.

Step 3: enable

curl -sS -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"rolloutCompleteConfirmed": true}' \
  https://api.staging.syrf.org.uk/api/admin/project-statistics/$PROJECT/fold/enable

200 with outcome: Applied (or AlreadyApplied) means one transaction set FoldMode: Enabled, stamped the protocol tripwire, marked every fold family Stale and advanced the clocks. Every refusal is a 409 whose reason names it:

reason Meaning and fix
FoldFlagOff The api pod does not have the fold flag on: Step 1
NotAllowlisted The project is not in this pod's allowlist
RolloutNotConfirmed Send rolloutCompleteConfirmed: true only after checking both rollouts
PendingIndexMissing Build the index first
FoldFleetUnsupported A live api or project-management pod does not declare the protocol being stamped, one declares a lower one, or no fold-worker heartbeat is live; detail says which. From slice 6, enable and reset stamp the lowest protocol any live pod writes that the image accepts (its own, or its previous one from P0 on), so a rolling deploy stamps the older protocol instead of refusing; a pod below that window still refuses. Wait for the rollout
FoldCoverageIncomplete The database is the production one (syrftest): refused until the production rollout, detail says so. Or a partial-coverage image (slices 2 to 5) off the staging pilot: IsStaging() false or the allow setting off
InvalidState The project is still Disabling; wait for Disabled
ConcurrencyLost A concurrent control write; repeat the call

Step 4: republish ProjectScreening

Enable leaves the family Stale, so pages read authoritatively until it is rebuilt. Saves already append from this moment. Run the ordinary backfill; it publishes authoritative - pending from one snapshot, so saves running during it are counted once:

curl -sS -X POST -H "Authorization: Bearer $TOKEN" \
  https://api.staging.syrf.org.uk/api/admin/project-statistics/$PROJECT/backfill

PendingNotDrained (retryable) means something pending cannot be subtracted (an overflow marker, an entry from another configuration, member set or tracking mode, or an unreadable entry). The worker discards and stales those within a minute; repeat the backfill.

The reviewer families need their own family flag. Both MembershipScreening and ReviewerScreening are gated by the single fleet-wide flag materializedProjectStatisticsMembershipScreening, which is off on staging. While it is off:

  • the reviewer backfills (…/membership-screening/backfill, …/reviewer-screening/backfill, or the operator workflow's backfill with family set) answer 409 FamilyNotServable. That is expected; skip them;
  • the reviewer families are never served materialized: reads use the legacy calculation;
  • the fold still receives the slice-3 reviewer transitions, finds no Fresh reviewer guard, and discards them (fold.discarded with family membership_screening / reviewer_screening, reason family_not_servable). That is by design and moves nothing.

Only after the flag is on (see Before turning the reviewer flag on) republish the two reviewer families the same way, with those family routes. Leave the annotation families Stale: every screening save that moves them stales them again in the pilot.

Before turning the reviewer flag on

Turning materializedProjectStatisticsMembershipScreening on is a separate decision from the fold pilot. Do not flip it until all three hold:

  1. Scope. The flag is fleet-wide, not per project: it starts reviewer-family maintenance and serving for every project in SYRF__ProjectStatistics__ProjectAllowlist on both hosts, fold or not. Confirm the staging allowlist holds only the pilot project, or accept the wider scope explicitly.
  2. The exactness check, defined in advance. There is no automated parity audit for the reviewer families. The check is: with GET …/fold/status showing pendingStudies: 0 and no save in progress, read every member's MembershipScreening and ReviewerScreening scope through the statistics query, read the same values from full-stats (the legacy calculation), and diff them key by key; then read fold/status again and repeat if anything was pending or the project's clocks moved between the two reads. Every key of every member must match. Record the diff as the pilot evidence; run it after the backfills and again after a period of real saves.
  3. No stored derived-field drift in any allowlisted project (#3937). Both reviewer families are legacy queries over stored derived fields of each Study: ReviewerScreening reads the inclusion status, MembershipScreening the agreement measure (NumberScreened, AbsoluteAgreementRatio) and inclusion fraction. Every whole-Study save re-persists them recalculated from the screenings, and only the screening saves record a reviewer move for that change: annotation saves on either path, reservation, presence, idle-session and risk-of-bias saves do not. A Study whose stored values differ from the recalculated ones would therefore leave the family that reads them wrong after an ordinary save. Since #3937 that family's rebuild refuses such a project with InclusionStatusDrift (the scope stays Stale and reads fall back), so the flip cannot serve a wrong value, but the family is not materialized until the drift is cleared. For every project in SYRF__ProjectStatistics__ProjectAllowlist (the flag is fleet-wide), before the flip:

  4. Run the read-only check. It writes nothing, needs no statistics flag and reads every Study of the project once, so run it at a quiet time on a large project:

    # Read-only. Administrator token; the statistics operator credential is not accepted here.
    curl -sS -H "Authorization: Bearer $TOKEN" \
      https://api.staging.syrf.org.uk/api/admin/project-statistics/$PROJECT_ID/reviewer-screening/inclusion-status-drift
    

    It answers studiesChecked, statusDriftedStudies (blocks ReviewerScreening), agreementDriftedStudies (blocks MembershipScreening), driftedStudies (either), hasDrift, notResavableStudies and up to 20 sampleDriftedStudyIds. notResavableStudies are documents of the one known legacy shape (a null screening list); only a screening save can save them, so they do not block. Any other Study whose values cannot be recalculated counts as drift for both families. 2. If statusDriftedStudies is above zero, run the project's inclusion recalculation (POST /api/projects/{projectId}/update-study-inclusion with JSON body true), wait for it to complete (the fence diagnostic shows no fence), and run the check again. Its completion advances the project's source clock, so the next backfill records a new bootstrap point. The recalculation's server-side status expression reads each Study's stored agreement fields, whereas the writers recalculate from the screenings, so a Study whose stored agreement fields are themselves stale can still show status drift afterwards. 3. If agreementDriftedStudies (or remaining statusDriftedStudies) is above zero, the recalculation cannot clear it: only a whole-Study save of each listed Study rewrites the agreement measure, and there is no administrative tool for that yet. Record the ids and stop for a decision. Never edit Study documents by hand. 4. With driftedStudies: 0, run the MembershipScreening and ReviewerScreening backfills. A backfill that answers 409 with failureReason: InclusionStatusDrift names the count in each scope's detail; repeat from step 1 (a threshold change in between, for example, runs a new recalculation).

Step 5: verify

  1. Screen a Study in the pilot project. GET …/fold/status shows pendingStudies: 1 for under a second, then 0 once the worker folds it (the post-save signal triggers it; the one-minute sweep is the safety net).
  2. The parity audit for ProjectScreening reports parity. It reads the stored row plus the pending entries, so it is exact while entries are pending; a save landing between the audit's two reads makes that attempt inconclusive (sourceAdvancedDuringAudit: true), and the audit retries once. Repeated inconclusive runs mean the project is too busy to audit, not drift; re-run at a quiet moment. On an image without #3902 the audit reports Unavailable (PendingOverlayUnavailable) permanently, which is not evidence of anything. A bundle that requests ProjectScreening and a reviewer or annotation family falls back for the second or so an entry is pending (the entry stales those families); ProjectScreening-only reads stay exact throughout.
  3. The reviewer families, only once materializedProjectStatisticsMembershipScreening is on: run the exactness check defined in Before turning the reviewer flag on. While the flag is off, a comparison with full-stats proves nothing: the served reviewer values are the legacy calculation, so they always match. A member added to the project has no row until the families are backfilled (reads of that member fall back), and an entry still pending when membership changes is discarded by the fold, which stales both families. Re-run both backfills after a membership change.
  4. Telemetry (instruments): fold.batches and fold.entries grow with saves, pending_entries.appended counts saves, pending_entries.overflowed, fold.quarantined and overlay.invariant_violation stay at 0. fold.discarded counts only families that had moves the fold could not apply; families that travel purely as intents (the annotation families on a screening save) are staled without a discard and never appear there. While the reviewer flag is off it shows membership_screening and reviewer_screening with reason family_not_servable on every batch that moves a reviewer counter, and otherwise stays at 0 for the pilot's families; fold.lag p95 stays under about a second, and fold.worker_heartbeat_age stays under two minutes. There is no alert rule for it in cluster-gitops yet: watch it during the pilot (a value above two minutes means no fold worker is running). Adding the alert is a follow-up. From slice 3, fold.reviewer_fallbacks counts saves that staled the reviewer families instead of moving them (no_threshold, no_evidence, membership_bound); on the pilot project it should stay at 0.
  5. A save never fails because of the fold. A worker outage only delays folding: reads stay exact through the overlay for up to 10 minutes, 128 pending Studies or 1,024 entries, then fall back.

After a fold protocol bump (slice 3 onwards)

Slices 3, 4 and 5 each bump the fold protocol (slice 3 to 2, slice 4 to 3, slice 5 to 4). There is no previous-protocol support before production (owner decision h), so once an image with a new protocol is rolled out on both deployments, the pilot project's control is still stamped with the old protocol: every reader falls back (Incompatible), saves take today's transactional path, and the worker leaves the project alone. Nothing is wrong meanwhile, but nothing is served materialized. Recover it:

  1. Wait until both rollouts have completed (ArgoCD Healthy, no old ReplicaSet pod alive), and GET …/fold/status shows the new binaryProtocol and a live worker heartbeat.
  2. POST …/fold/reset (operator workflow: fold-reset, confirm: true). It re-stamps the tripwire at the new protocol and marks every materialized family Stale. Entries written under the old protocol that are still pending are quarantined by the worker as UnknownProtocol, which also stales their families; that is expected.
  3. Backfill project-screening, and, only if materializedProjectStatisticsMembershipScreening is on, membership-screening and reviewer-screening (Step 4; with the flag off they answer FamilyNotServable, which is expected). From slice 4, and only if the annotation families are on (below), also stage-annotation, membership-stage-annotation and domain-reconciliation. Repeat any that answers PendingNotDrained after a minute.
  4. Verify (Step 5).

From slice 6: the production baseline P0, the N-1 window and the stamp advance

Slice 6 does not bump the fold protocol: protocol 4 (slice 5) becomes the production baseline P0. A pilot already reset at protocol 4 after slice 5 needs no reset after the slice-6 image deploys; check that GET …/fold/status shows stampedProtocol: 4 and binaryProtocol: 4. (A pilot still stamped below 4 follows After a fold protocol bump once more.)

From P0 every later bump ships with previous-protocol (N-1) support (owner decision h):

  • During the rollout of an image at protocol N, the stamp stays N-1. Both images accept N-1, write N-1 entries and fold them, and the N-1 heartbeat stays live (every N pod renews N and N-1). A rollback during this window needs nothing.
  • The stamp advance N-1 to N. After both rollouts report complete (ArgoCD Healthy, read-only kubectl rollout status with no older ReplicaSet pod alive), set SYRF__ProjectStatistics__Fold__StampAdvanceAllowed: "true" on both hosts in cluster-gitops. That setting is the recorded rollout-complete gate: pre-fold pods are invisible to the fleet check. The project-management consumer then advances every fold project stamped N-1 within a minute (it rides the fold sweep and runs only while the image has a previous protocol). The administrator route runs the same code for one project: POST api/admin/project-statistics/{id}/fold/advance-stamp (statistics operator credential or an administrator). It needs no reset and no rebuild. Its typed 409 reasons:
reason Meaning
NothingToAdvance The image is at P0 and has no previous protocol; expected today
NotAllowed StampAdvanceAllowed is not set on this host
FleetNotReady A live pod still writes N-1 only, or a fold worker does not fold N; detail names it
WorkerHeartbeatNotLive No live fold-worker heartbeat for N
NotInFoldMode The project is Disabled and carries no stamp
StampOutsideWindow The control is stamped neither N-1 nor N, or its tripwire disagrees: run reset
PendingBelowPrevious An entry below N-1 is still pending; the worker quarantines it, retry after a minute
ConcurrencyLost A concurrent control write; the next pass (or a repeat) retries

The setting is read once at start-up, so it applies only to pods started after the cluster-gitops change (the change rolls them). After the advance, set StampAdvanceAllowed back to "false" (or leave it until the next bump's rollout starts: it must be "false" before an image with a newer protocol begins rolling out). - After the advance, an N-1 image is outside the project's window and fails closed like a pre-fold image, so rolling back past an advanced stamp follows the rollback runbook. - A breaking bump (a classifier fix or any change to how an existing kind derives) declares the kind breaking in the N image. N pods discard pending N-1 entries of that kind: the overlay falls back, the fold stales the affected families and receipts the entries. After the stamp advance run POST …/fold/reset for each enabled project, then the backfills of the affected families, as in After a fold protocol bump steps 2 to 4.

After a deploy: what to run

Deploy Reset Backfills
An image at the same fold protocol (any deploy from slice 6 that adds no entry meaning) No No, unless the deploy changed a family's calculation or configuration identity
An additive protocol bump N (from P0) No: the N-1 window keeps reading old entries; advance the stamp after both rollouts (above) No
A breaking protocol bump fold-reset for each fold project, after the stamp advance The affected families (PendingNotDrained means repeat after a minute)
A protocol jump outside the window (the pilot between slices 2 and 5; any stamp StampOutsideWindow reports) fold-reset for each fold project Every family that is on: ProjectScreening, then the reviewer and annotation families whose flags are on
After a rollback window (allowlist guard removed) fold-reset twice, around guard removal (design runbook) Every family that is on, after the narrow gate re-opens
The pending index itself Not a deploy: built once per environment with pending-index-build before the first enable; never at start-up —

fold-reset and backfill are Statistics Operator workflow operations (confirm ticked). A reset of a project in fold mode (Enabled or Disabling) passes every enable refusal except the allowlist, including the production-database refusal; a reset of a Disabled project only stales every family. A project that has never been in fold mode answers InvalidState.

Disable and rollback

  • Disable (the normal off switch): POST …/fold/disable. Saves take today's transactional path from then on. The worker finalizes Disabled once the five-minute writer grace has passed and nothing is pending; GET …/fold/status shows foldMode: "Disabled". The served value stays correct throughout.
  • After Disabled: run the backfills of the other fold families (membership-screening, reviewer-screening, stage-annotation, membership-stage-annotation, domain-reconciliation, reviewer-annotation, question-answers). Enable staled them and the pilot kept them Stale, so they are served authoritatively until rebuilt. Then, in cluster-gitops, remove SYRF__ProjectStatistics__Fold__AllowPartialCoverage and set SYRF__FeatureFlags__MaterializedProjectStatisticsFold back to "false" on both hosts, unless fold mode will be re-enabled soon. The project keeps its fold history (FoldEverEnabled), so readers still consult the overlay and find nothing pending.
  • Reset: POST …/fold/reset re-stamps the tripwire at the lowest protocol the live fleet writes that this image accepts, and stales every materialized family. Run it after any slice deploy that bumps the fold protocol (slices 3 to 5 each did; slice 6 does not), and after a breaking bump's stamp advance, then repeat Step 4 and the other families' backfills.
  • Image rollback past slice 2 (or back from slice 4 to slice 2) follows the design's rollback runbook: disable and wait for Disabled, close the narrow gate, apply the allowlist guard alone, then change the images; roll forward with two resets around guard removal.
  • Rollback order, in full. (1) fold-disable for every fold project and wait for foldMode: "Disabled"; (2) close the project narrow gate (and the fleet gate for a full stop, Step 10); (3) the flags off in one cluster-gitops change for both hosts: materializedProjectStatisticsFold first or together with the family, Writes and Serving flags. Flags change only through cluster-gitops (a git revert synced by ArgoCD): never kubectl set env, never helm upgrade, never a runtime flag override, which the project-management host does not see (#3360) and which cannot reach the pinned fold flag. (4) Only for an image rollback past the fold: the allowlist guard alone, then the images, as in the design's rollback runbook.
  • Flag off without disabling first is not a fail-closed stop. With the flag off and fold mode still Enabled, saves take today's transactional point path: Fresh rows keep being served and updated correctly. Only entries still pending at that moment are a problem: no worker folds them, so after 10 minutes the overlay refuses (TooOld) and the project is served authoritatively until fold mode is disabled and drained, or reset and rebuilt. Disable first.

Slice 4: the annotation families

Slice 4 (fold protocol 3, #3925) moves the annotation and reconciliation session save, completion and deletion onto the fold path, and makes StageAnnotation, MembershipStageAnnotation and DomainReconciliation point-maintained: the worker folds them and the overlay serves them exactly. ReviewerAnnotation and QuestionAnswers stay Stale after every save, as today. Reservation claims and releases keep today's path until slice 5.

After the slice-4 image deploys the pilot needs the reset and backfills of After a fold protocol bump, now including the annotation families when they are on. The plan resets the pilot once, after slice 5 deploys, rather than after each of slices 3, 4 and 5: until that reset the pilot is served authoritatively (the cost is speed, never correctness). Then verify as in Step 5. For the annotation families, save and complete a session in the pilot project: pendingStudies returns to 0 within a second, fold.entries grows, and fold.discarded stays at 0 for the three point-maintained families (ReviewerAnnotation and QuestionAnswers are staled by their intents, which are not discards). These are liveness signals, not values: run the exactness check below as well.

Manual exactness check for the annotation families. The parity audit covers ProjectScreening only. Until StageAnnotation, MembershipStageAnnotation and DomainReconciliation have one, check them by hand, after some annotation activity and at a quiet moment:

  1. GET …/fold/status shows pendingStudies: 0, so stored equals stored + pending.
  2. Run the ordinary (non-forced) backfills, …/stage-annotation/backfill, …/membership-stage-annotation/backfill and …/domain-reconciliation/backfill. For each scope the backfill compares the visible row with the authoritative calculation (minus anything pending, which is nothing here) and answers AlreadyCurrent when they are equal.
  3. Pass: every scope of the three families answers AlreadyCurrent. Fail: any scope whose row was Fresh answers Rebuilt. A Fresh row disagreed with the authoritative value, so the fold served a wrong value. The backfill has already republished it; record the scope, disable fold mode (POST …/fold/disable) and report it. A scope that was Stale (for example just after a reset) is expected to answer Rebuilt; run the check again.
  4. Repeat steps 1-3 once more with no save in between; both runs must pass.

Before enabling QuestionAnswers on a fold project. On the fold path every annotation session save, completion and deletion carries a family-wide QuestionAnswers intent, including a re-save that changed no answer. With materializedProjectStatisticsQuestionAnswers on, the family is therefore Stale after almost every annotation save and its reads are served by the live calculation until a rebuild; any bundle that also requests QuestionAnswers falls back while such an entry is pending. That is correct but gives up the materialized read for that family during annotation. On the transactional (non-fold) path the writer stales only the questions whose answer count changed (#3933, a prerequisite for enabling QuestionAnswers anywhere; before it the stale matched no question row). Keep materializedProjectStatisticsAnnotation and materializedProjectStatisticsMembershipAnnotation on as well: with either off the annotation writer does not run and nothing is staled (#3840).

materializedProjectStatisticsMembershipAnnotation also turns on ReviewerAnnotation. The flag map puts MembershipStageAnnotation and ReviewerAnnotation under the same key. On the fold path every annotation save stales ReviewerAnnotation (the affected reviewers' scopes, or the whole family when allocation or inclusion moved). After every save, pilot reviewer-progress reads are therefore served by the authoritative calculation until a ReviewerAnnotation rebuild. Any bundle that also requests ReviewerAnnotation falls back as a whole while such an entry is pending. Reads of the three point-maintained families alone stay exact throughout.

The annotation saves take the fold path only while all three annotation families' writes are on: materializedProjectStatisticsAnnotation and materializedProjectStatisticsMembershipAnnotation on both hosts (today's writer has the same requirement). Staging has both off, so after the deploy alone annotation saves stay plain source saves and the screening save keeps carrying the annotation families as intents. To exercise slice 4, set both to "true" in the same cluster-gitops change for api and project-management, wait for both rollouts, then run the reset and the backfills above. Turning them on also turns on today's transactional reservation writer for the pilot project until slice 5; it moves the Fresh rows directly, which the fold tolerates.

Slice 5: reservations

Slice 5 (fold protocol 4, #3926) moves the reservation writers onto the fold path: the claim (still one atomic update, with the entry written by the same update), releases, timeouts and expiry (the hub's leave and disconnect, the screened-reservation release, the idle and suspended consumers) and the eligibility admission. They keep StageAnnotation, MembershipStageAnnotation and DomainReconciliation exact through the overlay and the fold; every claim stales the claimant's ReviewerAnnotation scope (the whole family on a SufficientlyAllocated crossing).

Like the annotation saves, the reservation writers take the fold path only while both annotation family flags are on, on both hosts (SYRF__FeatureFlags__MaterializedProjectStatisticsAnnotation and SYRF__FeatureFlags__MaterializedProjectStatisticsMembershipAnnotation, set as described in Slice 4). With them off, which is staging's state, a claim or a release appends nothing after the slice-5 deploy and reset: that is expected, not a failed deploy. Order:

  1. Deploy, then the reset and the ProjectScreening backfill of After a fold protocol bump.
  2. One cluster-gitops change setting both flags to "true" for api and project-management; wait for both rollouts. The annotation family backfills refuse while either flag is off on the api pod.
  3. Backfill stage-annotation, membership-stage-annotation and domain-reconciliation (each publishes authoritative - pending; repeat any PendingNotDrained after a minute). Rebuild reviewer-annotation only if a page needs it served from rows: it is Stale-on-save.
  4. Verify. Open an unclaimed Study in an annotation stage of the pilot project (the claim) and leave it (the release). GET …/fold/status shows the Study pending for under a second each time, then nothing; fold.discarded stays at 0 (ReviewerAnnotation travels as an intent, which stales it without a discard). Then run the manual exactness check of Slice 4. A claim never fails because of the fold (an absent, full or malformed pending set still lets it commit), and a release whose Study a fold bumped in the meantime is retried free, never reported as a conflict.
  5. Rollback: set both flags back to "false" on both hosts. Reservation and annotation writers return to today's path and the four families stop being served for the project. Entries already pending are still folded. Before turning the flags on again, rerun the step 3 backfills: while the flags were off, no writer maintained these rows.

A disable request against an already Disabled project reaffirms its intent through the control transaction without advancing its epoch or invalidation revision. A never-enabled project retains its null serving-transition token: its control version orders initial requests, so a later first enable needs no rebuild and preserves backfilled rows. Previously disabled projects rotate their existing transition token. This supersedes an earlier re-enable still rebuilding; that re-enable must refuse if its captured token changed. Rebuild-failure responses report the durable mode and watermarks read after the failed sweep, including any intervening administrative transition. This uses the existing control CAS and administrative gates; no additional feature flag is needed.

A fleet disable also claims or reaffirms the singleton in a transaction. If the singleton was absent, it records Disabled with deployed versions and reviewer mode, leaving write/mode epochs and invalidation at zero. This prevents an earlier pending first enable from committing after the disable reports success. A newly requested first enable remains available; no epoch has been retired.