Check Performance Data - Data egress to LDS (AB#294553 / spec AB#292610)
A back-office admin section where a CYPMD ops user pulls the scrutiny decisions for one checking window's Add and Remove amendment requests, preprocesses the approved ones into LDS-spec CSV files, persists the processed records in the database, and transfers the files to the LDS storage account. Everything lives inside the web app — no separate worker or scheduled job.
1. What it does
Six screens under /admin/egress. The five journey screens are gated by
[RequireAdminSection(AdminNavKeys.Egress)]; the runs history has its own gate (below):
-
Pull (
GET/POST /admin/egress) — choose a checking window and one or more output types (New learners, Remove learners), then pull. Also lists saved (in-progress) and completed runs. -
Results (
GET /admin/egress/runs/{id}/results) — every pulled request, one tab per output type, with its raw Zendesk decision and CYPMD values, before any filtering happens. -
Preprocessing (
GET /admin/egress/runs/{id}/preprocessing, plusGET .../preprocessing/streamfor the SSE progress feed andPOST .../preprocessingfor the no-JS run-to-completion fallback) — runs the eight-step pipeline and shows live progress. -
Failed (
GET /admin/egress/runs/{id}/failed) — reached only when preprocessing found a problem; lists every failure. -
Summary (
GET /admin/egress/runs/{id}/summary,POST .../transfer) — confirms the target container and file list, with a preview (GET .../preview/{outputType}) and download (GET .../download/{outputType}) per file, then transfers. -
Complete (
GET /admin/egress/runs/{id}/complete) — the transfer summary: files, record counts, hashes, who did it and when. -
Runs history (
GET /admin/egress/runs, AB#294590; sidebar tile "Egress runs";EgressRunsController, gated by[RequireAdminSection(AdminNavKeys.EgressRuns)]) — read-only: every run newest first with one of four statuses (EgressRunOutcomes: Success = Transferred; Failed = PreprocessingFailed or TransferFailed; Draft = Pulled, Preprocessing, Preprocessed or Transferring; Abandoned), the output types it covered, the records transferred (the saved row count for a Success, zero for everything else — it answers what LDS received), who started it and when. Filters by checking window and status are cumulative (?windowId=&status=), 20 rows a page (?page=, clamped), and a filter that matches nothing says so instead of rendering an empty table. A Draft's Resume link and a finished run's View link both go to Resume (below). Plain GET form, no script. -
Audit log (
GET /admin/audit-log, AB#294592; root tile "Audit log"; gated by[RequireAdminSection(AdminNavKeys.AuditLog)]) — not an egress screen but where every transfer's audit row is seen: filter by activity "Data egress", checking window and Success/Failed, and export the filtered set as CSV. Seedocs/audit-log.md.
GET /admin/egress/runs/{id} (Resume) sends the browser to whichever of these a run's status
implies, so a saved run reopens without re-pulling. POST /admin/egress/runs/{id}/abandon ends a
run at any non-terminal point.
2. Where the data comes from
Zendesk supplies only the decision: the "Decision status" custom field on the ticket id already
stored on ChangeRequests.CrmId, read via GET /api/v2/tickets/show_many.json (100 ids per call).
Every other value — pupil details, dates, the school's establishment number, the journey's answers
— comes from CYPMD's own data: the ChangeRequests row and the persisted journey blob
(requests/{reference}.json, IRequestStateBlobClient). Production configures no Zendesk custom
fields at all, and an Add (new learner) request carries no establishment number, admission date,
year group or SEN status in Zendesk — so the database has to be the source of truth regardless.
ChangeRequests.OrganisationLaestab is a new column, written at submit time from the DfE Sign-In
organisation_laestab claim (RequestService.OrganisationLaestabOrNull). It exists because a new
learner's pupil record is synthetic and carries no LAESTAB of its own — the school's LAESTAB from
the request row is the only source for that case.
The pull's output, EgressSourceRecord, is the "as pulled" shape shown on Results and persisted as
the run's raw JSON (so a resume never re-pulls): identifiers (ChangeRequestId, ReferenceNumber,
TicketId), the Decision (a Zendesk value, or EgressDecisions.NoTicket / NotFound), the
output type and window type, submission metadata, the school's URN/LAESTAB, every pupil field
CYPMD holds, and the journey's answers flattened by question id (EgressAnswers.Flatten).
3. Run lifecycle
Pulled → Preprocessing → (PreprocessingFailed | Preprocessed) → Transferring → (TransferFailed | Transferred)
Abandoned is reachable from any non-terminal status, including Preprocessing and
Transferring — the lock has no expiry, so a run stuck there after a pod restart or a crashed
tab must always be releasable, and every screen from Preprocessing onwards carries an Abandon
form. Abandoning a Transferring run also sweeps any blob it actually wrote (matched by the
egressRunId metadata stamped on upload) before releasing the pair, so a same-named retry never
collides with an orphaned file; the confirmation banner names what was removed, if anything, or
says nothing was transferred. A run is persisted the moment the pull succeeds
(EgressRunService.StartAsync → IEgressRunRepository.CreateRunAsync), so "Save and exit" on any
screen is just leaving the page — there is no separate save action. Resume
(EgressController.Resume) maps status to screen: Pulled → Results, Preprocessing →
Preprocessing, PreprocessingFailed → Failed, Preprocessed/TransferFailed/Transferring →
Summary, Transferred → Complete; Abandoned (or an unknown run) sends the user home with a
banner. A failed transfer can be re-run from the same screen; a PreprocessingFailed run
cannot — it releases its pair (its outputs go inactive) and the Failed screen's own copy directs
starting a new run instead, so two runs (the retried one and a colleague's fresh one) can never
both hold the same pair.
4. Concurrency
One active run per checking window and output type, enforced by a database constraint, not
just a service check: a partial unique index on egress_run_outputs ("WindowId", "OutputType") WHERE "IsActive". IsActive is true from Pulled all the way through Transferred — a
successful transfer leaves it active forever, so the same window and output type can never be sent
twice — and false for PreprocessingFailed, TransferFailed and Abandoned, so a failed or
abandoned run never blocks a new one.
EgressRunService.StartAsync checks for a blocker first (a friendly refusal naming who holds it
and since when) and EgressRunRepository.CreateRunAsync/TryReactivateAsync also catch the
database's unique-violation as a race guard, in case two admins start at the same instant. The two
refusal sentences (EgressController.Describe, FLAGGED copy):
- "{Output type} for this checking window is already being processed by {name}, started {date} at {time} UTC. Wait for that run to finish or be abandoned."
- "{Output type} for this checking window has already been transferred to LDS by {name} on {date} at {time} UTC. It cannot be sent again."
5. Preprocessing
Eight steps, reported one at a time over the same IAsyncEnumerable<EgressProgress>
(EgressPreprocessor.RunAsync) whether driven by the SSE stream or the no-JS POST:
-
Filter records — keep
approvedandauto_approved; everything else is discarded here, not shown as a failure. -
Derive correction codes — Remove: the bare LDS code from the Zendesk "Correction reason
(31)" option for the journey's
reasonanswer (plus two Post16 reasons mapped by hand); New: the fixed code10. -
Split DfE establishment number — the pupil's LAESTAB first, then the request row's
OrganisationLaestab; 7 digits only, split 3+4. -
Standardise dates — date of birth (and, for New learners, admission date) to
yyyy-MM-dd. -
Build LDS records — the typed
NewLearnerRow/RemoveLearnerRow. - Trim values.
-
Validate against LDS spec — required fields, permitted values, digit widths, date shapes; Surname and Forename (the only free-text cells, typed by the school) must not begin with
=,+,-or@, which a spreadsheet would evaluate as a formula (OWASP CSV injection) — refused here with a named reason rather than rewritten with an apostrophe the LDS spec never asked for (LdsSpecValidator). -
Save to database — all or nothing: if any record failed at any step, nothing is written,
the run becomes
PreprocessingFailed, and every failure (step, ticket id, reference, field, reason) is listed on the Failed screen; otherwise rows go tonew_learners/remove_learners, file names are fixed for the run, and the run becomesPreprocessed. Every independent step (2–4) runs regardless of an earlier one's outcome, so a record with two unrelated problems (say an unmapped correction reason and an invalid LAESTAB) lists both in one pass rather than the second only surfacing on a later re-run.
Leaving the page mid-run (closing the SSE connection, or the no-JS POST being interrupted) puts the
run back to its status before preprocessing started — nothing durable happens until step 8 commits.
If the run was abandoned by someone else while this pipeline was still running, step 8's write is
guarded by the run's expected status and does nothing; the stream reports "this run was abandoned
while preprocessing" rather than claiming Preprocessed or PreprocessingFailed for a run that is
actually Abandoned.
6. Files
Both LDS files use the same writer (EgressCsvWriter): RFC 4180 quoting, \r\n between lines, no
line terminator after the final data row, every value trimmed, UTF-8 without a byte-order mark.
Both files follow LDS_CYPMD_Data specification_v2.4.xlsx (sheets "New Learner" and "Remove
Learner", read top to bottom, keeping the rows marked X for the window's key stage). The column set
is therefore the window's: EgressColumnSets.RemoveLearnersFor(windowType) /
NewLearnersFor(windowType).
Remove learners — every key stage:
Correction_ID, Correction_Type, Correction_Reason, Key_Stage, Establishment_Number, Surname,
Forename, Sex, Date_of_Birth, Cycle_Year, Cycle_Month, Local_Authority, Learner_ID
then, KS4 only: Year_Group (populated for year-group-change removals from the journey's
year-group-higher/lower-moved-to answer, blank otherwise); 16-19 only: Removal_Year_0,
Removal_Year_1, Removal_Year_2 (TRUE/FALSE from the years-to-remove checkbox — Year_0 is
the academic year ending in Cycle_Year; blank when the journey route did not ask).
New learners — every key stage (the spec's Middle_Name is struck through in v2.4, "CYPMD will
not be sending this field from June 2026", so it is not emitted; there is no SEN attribute):
Correction_ID, Correction_Type, Key_Stage, Establishment_Number, Surname, Forename, Sex,
Date_of_Birth, Admission_Date, Post_Code, Cycle_Year, Cycle_Month, Local_Authority, URN, ULN, UPN,
Learner_ID, Year_Group
then, 16-19 only: Attendance_Year_0, Attendance_Year_1, Attendance_Year_2, KS4_Year — all
blank today because no Post16 Add journey exists.
Values: Key_Stage is KS2 / KS4 / 16-19 (EgressOutputTypes.KeyStageValue; the file name
still uses AB#292610's KS5 token). Cycle_Year / Cycle_Month are the checking window's start
year and month (the spec's "month in which the cycle takes place"), not each record's submission
date. Sex is F / M / U on both files. 16-19 Correction_Reason codes are the spec's
(CorrectionCodes): 4 deceased, 325 not at end of study, 326/328/331 not on roll (international /
external / apprentice), 329 other with evidence.
File names are fixed when preprocessing completes: CYPMD_LDS_{stage}_{type}_{yyyy_MM_dd}.csv,
where {stage} is KS2/KS4/KS5 (EgressOutputTypes.StageToken) and the date is the London
calendar date at the moment preprocessing finishes (EgressOutputTypes.FileName).
7. Transfer and audit
EgressTransferService.TransferAsync refuses before touching anything if every output's saved row
count is zero (every record was rejected, undecided, or lost to a misconfigured ticket source) —
Summary shows a plain message instead of Confirm in that case, so LDS is never sent a header-only
file for a pair that then locks forever. Otherwise it builds each file's bytes from the
persisted rows, never from the pulled payload, and uploads with IfNoneMatch: * (create-only —
an existing blob is never overwritten).
Once the run's status has flipped to Transferring, the upload loop, the commit
(MarkTransferredAsync) and all compensation run with CancellationToken.None, not the request's
own token — a browser tab closing mid-upload must not abandon a run with files already in LDS. If
any file fails to upload, every file already written by that attempt is deleted
(IEgressBlobClient.DeleteIfExistsAsync), and the specific file that was mid-upload when the
failure happened is also checked and removed if it turns out to have landed server-side despite
the client seeing a failure (matched by the egressRunId metadata stamped on upload, so a blob
belonging to a different run is never touched); the run becomes TransferFailed with the reason
recorded. A post-upload failure in the commit itself (every file uploaded, but marking the run
Transferred throws, or loses a race because the run was abandoned in between) is caught the same
way and compensated identically — no partial state survives silently. A retry re-activates the
run's outputs first and is refused if another run has since claimed the pair.
A create-only upload that hits an existing blob (EgressBlobAlreadyExistsException) is not
always a real collision. Before giving up, transfer reads the blob's egressRunId stamp
(IEgressBlobClient.GetOwnerRunIdAsync): a file stamped by a run that is now Abandoned (a sweep
the process died in the middle of) or by this very run (an earlier attempt whose compensation
delete failed) is reclaimed — deleted with the ownership-checked delete and the upload retried
exactly once. A file stamped by any live run, or carrying no stamp at all, is left untouched and
the run becomes TransferFailed with a reason that names who wrote it and the way out. Note that
the file name carries the preprocessing date, so retrying the same run on a later day reuses
the same name: the way out of a genuine same-stage/same-day collision is to abandon the run and
start a new one on a later day, or to have LDS remove the file.
Every terminal write (MarkPreprocessingFailedAsync, SavePreprocessedAsync,
MarkTransferredAsync, MarkTransferFailedAsync) is guarded by the status the caller expects the
run to currently hold, and reports rows affected; a caller that gets zero back knows it lost a race
(most often to a concurrent Abandon) and never overwrites what actually happened with a stale
outcome. Success and failure each write an AuditEntry in the same transaction as the guarded
state change: EntityType "EgressRun", Action "Transfer" or "TransferFailed", NewValues a
camelCase JSON object. Both carry the outcome, window id, output types and the person who ran the
transfer (transferredBy); success adds file names, record counts, SHA-256 hashes, target container
and time; failure adds the reason (AB#294592 made the failure payload self-describing so the audit
log needs no join). No audit row is ever written for a write that lost its race, and no
audit row ever claims success for a failed transfer.
8. Configuration
| Setting | Purpose |
|---|---|
ConnectionStrings:EgressStorage |
The LDS storage account. Absent → the Pull page and the Summary show a warning up front (egress-storage-not-configured), the app refuses to transfer with a clear "not configured" message rather than failing at startup, and the dev cleanup skips its blob sweep. Pull, preprocess, preview and download still work without it. appsettings.json deliberately carries no default (pinned by AppSettingsEgressStorageTests): a local-Azurite default there made every deployed environment without an account believe it had one at 127.0.0.1:10000 inside the pod, so transfers and cleanups failed after the SDK's retries instead of refusing. Local runs get it from docker-compose.yaml and the launch profiles. Review apps get the per-PR lds account (terraform/application/application.tf, local.egress_storage_secrets, gated on var.config == "review") so the E2E transfer facts run end to end; the long-lived environments are still unconfigured until the real LDS account is known — see gaps below. |
EgressStorage:Container (default cypmd) / EgressStorage:Prefix (default extracts_input/) |
Where in the account files land. Bindable only so a test can point at a scratch container — never user-editable. |
Zendesk:UseFake (default false) |
Selects the ticket source: the real Zendesk client (ZendeskEgressTicketSource, via the same AddZendeskApiClient the worker uses) unless explicitly set to true, which selects the dev outbox (DevOutboxEgressTicketSource, no Zendesk settings needed). The default matches the worker's own configured default — a fresh environment that sets nothing reads real Zendesk decisions, not the dev outbox. AddCpdEgress refuses to start if UseFake=true is set in Production, regardless of configuration, so the dev outbox can never be reached there. Local/E2E stacks opt in explicitly via Zendesk__UseFake=true (docker-compose.yaml, docker-compose.sandbox.yaml), since neither has real Zendesk credentials. Review apps opt in too (terraform/application/config/review.yml): the E2E egress facts seed their decisions into the outbox via /dev/egress/seed, and against real esfa-preprod every seeded id read back as "Ticket not found", so the suite could not pass there. Web and worker share the ConfigMap, so review-app submissions land in the outbox rather than creating esfa-preprod tickets. |
ZendeskTicketFields:DecisionStatusId |
The real ticket source's required field id; 0 in production today, so it refuses to pull until configured. |
Admin grant egress
|
DefaultAdminAccessSeeder.AllSections — without it a fresh database 404s on /admin/egress even for an admin. |
Admin grant egress-runs
|
DefaultAdminAccessSeeder.AllSections — the runs history's own gate and its sidebar tile's key (AB#294590); the seeder tops the admin role up on every start, so existing databases gain it on deploy. |
Admin grant audit-log
|
DefaultAdminAccessSeeder.AllSections — the Audit log root tile's own gate (AB#294592); topped up on every start. |
9. Local development and E2E
Locally the web container gets a third Azurite account (docker-compose.yaml,
ConnectionStrings__EgressStorage, alongside the app and ingress accounts), and opts in to the
dev outbox fake explicitly with Zendesk__UseFake=true — the code/config default is now the real
Zendesk client, which this stack has no credentials for.
DevEgressController (dev-only, 404 unless Dev:ToolsEnabled and not Production, same rule as
DevPipelineController) stages fixture data with no worker and no real Zendesk:
# Seed 3 approved Remove-learner requests for the KS4 June window
curl -X POST "http://localhost:8080/dev/egress/seed?windowId=F34D285B-8660-4D12-9C30-787328DEAA0A&outputType=RemoveLearners&decision=auto_approved&count=3&laestab=860/4070&urn=142313&reason=pupil-died"
# Remove everything the harness created for that window (runs, requests, blobs)
curl -X POST "http://localhost:8080/dev/egress/cleanup?windowId=F34D285B-8660-4D12-9C30-787328DEAA0A"
tests/DfE.CheckPerformanceData.E2ETests/Admin/DataEgressTests.cs walks the whole journey over
plain HTTP (pull → results → preprocessing → summary → download → transfer → complete → refused),
a failing-record path, and — Linux-only — a real-browser fact for the streamed progress. Run just
this class: docker compose --profile e2e run --rm e2e-tests sh -c 'dotnet test tests/DfE.CheckPerformanceData.E2ETests/ --filter "FullyQualifiedName~DataEgressTests"'.
To inspect what actually landed in the local LDS account, the Storage browser under Danger zone
covers the app account only — the egress account needs a separate client (e.g. Azure Storage
Explorer, or the Azure CLI, pointed at the EgressStorage connection string from
docker-compose.yaml or Properties/launchSettings.json), container cypmd, prefix extracts_input/.
Two rules keep a dev or review environment recoverable after a failed run:
-
POST /dev/egress/cleanupalways deletes its database rows. The blob sweep is best effort — a blob that could not be deleted is counted in the response'sblobErrors, never thrown — because a cleanup that answered 500 and left the runs behind is what broke the next deploy (below). - The dev seeder (
SeedCheckingWindows, run at start-up whereverSeedDevelopmentDatais on) deletesegress_runsbefore it wipesCheckingWindows. The foreign key from a run to its window is RESTRICT on purpose — an egress is an audit record — so a leftover run used to make the wipe throw before the host listened. On a review app that looked like a stalled rollout: the new pod never became Ready, the old one kept serving, terraform reported "old replicas are pending termination" and a re-run reported "No changes". PR #441's review app served a two-day-old image that way;CheckingWindowSeedWithEgressHistoryTestspins the fix.
10. Known gaps and follow-ups
-
Merged learners (and the KS4 June code
20→21correction-code rule, AB#292610) is deliberately out of scope — its own column set and rule, tracked as a follow-up ticket. -
LDS spec v2.4 questions still open with LDS/BA (see the PR notes §10): the struck
Middle_Nameheading is omitted entirely;Key_Stagesays16-19(v2.4 changed the Remove sheet from 16-18, the New Learner sheet was not updated); year-group-change removals go out as Correction_Type 31 / reason 17 (business-confirmed 2026-08-06) although the spec's hidden Addback sheet has them as type 30;Removal_Year_0..2are blank unless the 16-19 journey took the "other" route; 16-19 "other" is always 329 (evidence is always collected, so 330 never occurs); Correction_Type 11 (Include learner) is in the spec but Include is not an egress output yet. -
16-19 gaps: no Add journey exists, so the New learners file for a 16-19 window can only ever
be header-only (and Transfer refuses a run whose outputs are all empty); 16-19 pupil records have
no MATCHREF, so
Learner_IDfails validation for them until the 16-19 pupil file supplies one. - Trailing newline: files end after the last data row with no trailing line terminator, to honour "no additional rows below the final data row" — confirm this reading with LDS.
-
ConnectionStrings__EgressStorageis only in Terraform for review apps (terraform/ application/application.tf,local.egress_storage_secrets, pointing at the per-PRldsaccount). Development, QA, preproduction and production stay unconfigured — add them once the real LDS account details are known. -
Production's
ZendeskTicketFields__DecisionStatusIdis0— the real ticket source refuses to pull until an environment configures it. - The pre-existing gap that
Controllers/WindowAdmin/*carries no[RequireAdminSection]is a separate ticket and was not touched here; the newEgressControlleris gated. - Every copy string introduced by this feature is FLAGGED for content sign-off — see the PR notes.
-
Same-stage same-day file-name collision: two different checking windows of the same stage
(KS4 June and KS4 Autumn both map to the
KS4file token) transferred on the same day produce the same file name, so the second run's transfer fails with "already exists". The file name is fixed at preprocessing, so retrying the same run on a later day collides again; the ways out are to abandon the run and start a new one on a later day, or a manual delete in LDS — and the recorded failure reason now says exactly that (§7). A stable fix means a different naming scheme, which is LDS's to agree. Question for LDS: is one-window-per-stage-per-day a real constraint? Tracked as a follow-up. -
PreprocessingStreamis a state-mutating GET (the acceptedValidateWindowControllerpattern) — the JS closes theEventSourceon a terminal/error event, so the browser's automatic reconnect never restarts the server-side pipeline. Left as-is; noted for the next contributor. -
Abandon crash-window orphan: R1 reordered
AbandonAsyncto writeAbandonedbefore sweeping the run's blobs (closing a data-loss race — see the R1 commit), which moved the crash window rather than removing it: if the process dies after the write commits but before the sweep loop finishes, the run ends upAbandonedwith one or more files still in LDS storage. The recovery is now automatic rather than a manual LDS delete: the next transfer whose file name collides with such a file reclaims it (§7), because it is stamped with anAbandonedrun's id. Two residual caveats: LDS may already have picked the orphan up before it is reclaimed (true of any orphan, reclaimed or not), and an orphan whose name no later transfer ever reuses stays there until LDS or an operator removes it. - Related:
EgressTransferService.FailAsync's compensation delete of its own just-uploaded files (theuploadedlist) callsblobs.DeleteIfExistsAsyncunconditionally, unlike itspossiblyOrphanedcheck, which is ownership-checked viaDeleteIfOwnedByRunAsync. In the create-only-upload case this is safe today (a successful create-only PUT cannot belong to another run), but it is a related blob-lifecycle-under-overlap gap worth tightening for consistency. Follow-up, not an open production incident. - An independent review of the transfer/lock state machine found one Blocker (B1: the web host
defaulted to the dev Zendesk fake with nothing but QA config overriding it, so Production would
have read the dev outbox table instead of real Zendesk decisions) and four Must-fixes (M1:
transfer was not atomic on cancellation or a post-upload database failure, leaving a run stuck
Transferringwith files already in LDS; M2: re-running aPreprocessingFailedrun bypassed its released lock; M3: a run whose approved set was empty could still transfer a header-only file and lock its pair forever; M4: every terminal status write was unconditional, so a concurrent Abandon could be silently overwritten) plus several should-fixes, all addressed in follow-up commits on this branch — see the PR description for the finding-by-finding record.