Scenario · Software release rollback

The Release That Had to Explain Its Rollback

A green overall dashboard hides a failing canary release, seven delivery jobs have unknown outcomes, and the rollback must preserve the truth that code cannot reverse.

9 min readSupply chain and manufacturingIllustrative evidence
What this is
A published decision scenario. Its setting and values are invented for illustration.
What this shows
How declared facts and rules produce a traceable result when conditions change.
What this does not show
A customer deployment, measured outcome, or transfer of authority to software.

Narrated film · 1:10

The Release That Had to Explain Its Rollback

Follow a software release that needs to stop, while preserving an honest record of the work a rollback cannot undo.

Watch narrated film · 1:10 Read the story · 9 min
1:10

Story

Start with the event and the decision it creates.

The story

10:12 — one service is green and one release is red

At 10:12, Ava Chen, Bluejay Works' release duty manager, sees a green whole-service dashboard hiding a failed candidate cohort. Seven document-delivery jobs also have no current terminal state from the external provider. Ava must decide what the evidence supports before authorizing expansion, retry, or rollback; the wrong shortcut could increase exposure or duplicate an effect the provider already accepted.

Thirty minutes earlier, Bluejay's Dispatch service had prepared release dispatch-api@3.8.0+7f2c for a 20-percent canary while retaining known-good coordinate dispatch-api@3.7.4+91ad as its rollback target. The registry pinned both artifacts and an uncomfortable boundary: application rollback cannot undo jobs already submitted to the external provider.

At 10:00, Ava authorized only the canary. The deployment platform acknowledged the request, and a later workload report showed candidate instances serving the declared cohort. The sequence remains intact:

canary authorized → request acknowledged → candidate observed serving

Ava's approval does not move traffic. Platform acceptance does not prove execution. Observing candidate instances does not prove that every eligible job entered the expected cohort. The registry establishes which artifacts Bluejay intended to use, not that the candidate is reliable or that rollback will be safe or complete.

The first closed observation window looks reassuring from far away:

Cohort Eligible jobs Receipt-deadline misses Exact rate Authored gate
Incumbent 3.7.4+91ad 4,800 24 0.5000% Pass
Candidate 3.8.0+7f2c 1,200 23 1.9167% Fail
Blended service 6,000 47 0.7833% Informational pass only

Bluejay's authored policy applies a one-percent timeout ceiling to each release cohort. The candidate therefore fails even though the blended service metric passes. A green whole-service dashboard cannot erase a red candidate gate.

Seven candidate jobs have another problem. Bluejay recorded each submission, but the provider has supplied no current terminal state. The exact job identities remain in the evidence packet; their state is unknown—not failed. Blindly retrying them could duplicate an effect the provider already accepted.

10:13 — the evidence link degrades

The telemetry bridge disappears for seventy-five seconds. Bluejay's authored freshness limit is forty-five. The last accepted observation becomes stale for any current expansion decision, so expansion remains closed. Its recorded candidate failure remains historical evidence of what happened in the closed 10:12 window.

When the bridge returns, it redelivers the same window-close event twice. The first copy is retained against that historical window. The second is marked as a duplicate and changes no count or decision. Neither copy becomes current clearance to expand: a recent delivery time cannot make an old observation current, and silence cannot be filled with invented zeroes. The retained failure can support stopping exposure while fresh evidence remains required for any later expansion.

10:15 — three responses expose the real decision

The authored model evaluates exactly three responses:

Response Result Why
Expand the candidate to 50 percent Rejected The cohort gate fails and seven effects are unknown.
Restore the incumbent and retry all seven jobs Rejected Code can be reversed; unknown provider effects cannot safely be assumed failed.
Stop candidate traffic, restore the incumbent for new work, and quarantine the seven jobs Conditional proposal It limits new exposure while preserving each unknown state for provider reconciliation.

“Complete” means the fixture compared these three authored alternatives. It does not mean the system discovered every possible response, proved a root cause, or found a universal optimum.

10:16 — the model stops at authority

The surviving packet binds the exact release, telemetry, policy, and deployment revisions; the candidate and blended arithmetic; all seven unknown job IDs; the evidence gap and duplicate; the rollback target; validation gates; and the boundary around incident communication.

Ava authorizes alternative C under Bluejay's represented emergency-change policy. Her signed assertion records who approved which exact packet and why. It is not an instruction from Grid, proof of actual delegation, a workload change, a provider reconciliation, or permission to publish a customer message.

10:18 — an accepted rollback is not a rolled-back service

The external platform accepts rollback request RB-419. Only a later platform report shows zero candidate instances and new eligible work returning to 3.7.4+91ad. A separate observation window then passes the incumbent gates.

The full evidence line remains visible:

proposed → authorized → request accepted → deployment observed → service validated

The seven earlier jobs are still outside that rollback. The code has gone back. Its external effects have not.

10:21 — corrected evidence advances the present

Provider revision PROVIDER-889 reports that JOB-1184 and JOB-1190 were delivered. The current unknown count falls from seven to five. The first incident note and support snapshot become stale, but the earlier observation and Ava's decision packet remain intact because seven jobs were genuinely unknown when she acted.

At 10:44, PROVIDER-890 reports six of the original jobs delivered and JOB-1198 rejected before acceptance. Only that exact job becomes eligible for a separately authorized retry. The authored result is not a promise that a real provider, release, or rollback will resolve this way.

A successor release, 3.8.1, must still carry a new artifact digest, repeat the deterministic fixture, pass independent tests, and begin a new authorized canary. A successful rollback cannot promote an untested root-cause hypothesis into truth.

Why this setting is recognizable

CISA's safe software deployment guidance discusses controlled, measured rollouts, monitoring, rollback planning, validation, safeguards, incident communication, and learning. NIST SP 800-128 describes controlled configuration change as a documented sequence rather than a single button press. NIST SP 800-218 supports preserving release files, integrity information, and provenance.

Those official sources ground only the release-management context. They do not establish Bluejay, its one-percent gate, its canary size, the seven jobs, the correct response, a Grid deployment, or any outcome. This Scenario recommends no universal SLO, retry rule, canary percentage, or rollback procedure.

Where the model stops

Grid can represent revisioned evidence, calculate the authored gates, expose stale or contradictory state, compare declared responses, retain a decision packet, and propagate corrections across modeled derivatives. It cannot operate CI/CD, infer authority from authentication, execute a rollback, declare an external effect complete, retry an unknown job, approve communication, or turn a successful scenario recovery into a reliability claim.

What remains to prove

Source revision 2026-08-27.1 has reviewed domain context, product boundaries, and a public-suitability finding, so this authored scenario can be published at R2. The fixture itself remains authored.

The next proof is narrower and harder: implement the exact deterministic counts, gates, seven job identities, stale interval, duplicate event, three alternatives, authority stop, deployment states, and two provider corrections; freeze a separately identified independent oracle; run the negative cases; and retain exact execution and verification evidence. A later intended-environment evaluation would still need Bluejay-owned source, identity, authority, deployment, provider, security, records, incident, and reconciliation controls. None of that would turn this editorial Scenario into proof that Grid manages software releases or that the authored rollback succeeded.

Decision path

Follow the changed fact step by step.

A calculation, proposal, approval, execution report, and outcome are different events. The order keeps those boundaries visible.

  1. 01 · 09:42

    The release arrives with a way back

    The registry pins both release coordinates, artifact evidence, and a rollback candidate without claiming that either release is running.

  2. 02 · 10:00

    A bounded canary begins

    Human authorization, request acknowledgment, and observed candidate service become three distinct evidence states.

  3. 03 · 10:12

    One service is green and one release is red

    The blended metric passes, but the candidate cohort fails its exact gate and seven external submissions remain unknown.

  4. 05 · 10:15

    Three responses expose the real decision

    Expansion and blind retry are rejected; restoring the incumbent for new work while quarantining unknown effects remains a conditional proposal.

  5. 06 · 10:16

    The model stops at authority

    Ava authorizes one exact decision packet, but her assertion neither changes a workload nor resolves a delivery.

  6. 07 · 10:18

    An accepted rollback is not a rolled-back service

    Platform acceptance, observed workload state, validation, and the seven unresolved external effects retain separate evidence.

  7. 08 · 10:21–10:44

    Corrected evidence advances the present

    Provider revisions reduce seven unknown effects to five and then zero without rewriting the original observation or decision packet.

Evidence and limits

What the scenario represents—and what real-world use still requires.

Represented in this scenario

  • Exact release, policy, telemetry, decision, deployment, and provider revisions with their freshness and correction state
  • Candidate and blended gates, bounded alternatives, rejected reasons, and the affected derivative cone
  • Proposal, authorization, request acknowledgment, observed deployment, validation, and external effect as distinct states

Required integration and operating work

  • Author and qualify Bluejay-specific schemas, release policy, source precedence, adapters, identity and delegation, operator experience, reconciliation, and derivative workflow
  • Implement the deterministic fixture, freeze an independent oracle, test stale, duplicate, contradictory, unauthorized, and corrected paths, and retain exact execution and verification evidence

Decisions that remain with people and institutions

  • Canary expansion, rollback, retry, deployment, or release of a successor build
  • Terminal state for an external delivery without provider-owned evidence
  • Customer communication, security finding, root cause, reliability result, or service outcome
Evidence, authority, and publication recordView the scenario contract, capability record, authority stages, verification status, and related work.

Scenario contract

The setting, trigger, decision, and authority boundary.

Setting
Bluejay Works, an invented multi-tenant document-delivery software provider managing one bounded canary release.
Timeframe
09:42 through 10:44 on the scenario's release morning, followed by the requirements for a successor release
Trigger
Candidate release dispatch-api@3.8.0+7f2c fails its cohort-specific receipt-deadline gate while the blended service metric still passes.
Decision
Whether to continue, roll back and retry, or restore the prior release for new work while quarantining seven indeterminate external effects.
Authority
Grid may calculate gates, compare authored responses, and assemble an exact packet; Ava Chen may authorize the represented rollback plan, while the deployment platform, delivery provider, and communications lead retain execution, effect-state, and publication authority.

One rollback, several truths

The code can go back while its effects stay unresolved.

A passing aggregate, failing cohort, seven unknown effects, exact authorization, external deployment, and later corrections remain connected without collapsing into one green or red state.

  1. SourcesRelease and cohort revisions

    Exact coordinates, counts, windows, observation times, and effect identities retain their owners.

  2. ModelCandidate gate fails

    Twenty-three misses among 1,200 candidate jobs fail even while the blended service metric passes.

  3. AlternativesQuarantine beats blind retry

    The authored comparison preserves seven unknown external effects instead of pretending they failed.

  4. Human authorityExact rollback plan

    Ava authorizes one packet under represented policy; authentication alone does not grant that authority.

  5. External evidenceDeployment and provider reports

    Request acceptance, observed workload state, validation, and delivery reconciliation answer different questions.

Every value and result is synthetic; this is an authored R2 Scenario, not an executed Grid fixture or software-release benchmark.

What is established

What is documented, what this scenario combines, and what still needs testing.

This separates documented capabilities from authored combinations in the scenario. Neither proves a complete deployment or outcome.

Documented

Documented building blocks

Capabilities described in maintained Grid documentation or another named source.

  • Reviewed product primitives document reactive values and bindings, predicates, provenance, streams, connectors, governed artifacts, and multiple product surfaces within their stated limits; they do not establish a software-release product or this fixture's execution.
  • Reviewed product primitives document ways to represent revisioned facts, affected dependencies, retained decision evidence, and role-specific views; they do not establish deployment authority, provider state, root cause, or service reliability.
Combined here

Combined in this scenario

Capability combinations represented in this scenario that still require end-to-end evaluation.

  • The authored Scenario composes Bluejay release schemas, exact cohort gates, freshness and duplicate handling, three bounded alternatives, an authority stop, external-effect reconciliation, and correction-capable derivatives.
  • A release console and incident, support, and release-record views would be built around the same exact source and decision revisions rather than inferred from generic Grid primitives.
Needs testing

Not yet proved

Integration, operating, policy, or evidence work that is not complete.

  • No package-owned executable fixture, independent oracle, negative-case suite, retained execution and verification record, or derivative-parity review has yet established the authored behavior.
  • No CI/CD, observability, SLO, incident, identity, deployment, provider, communication, security, or target-environment integration is qualified, and no observed performance or customer outcome exists.

From model result to outcome evidence

A modeled answer does not perform the work.

Calculation, review, authorization, submission acknowledgment, execution reporting, and observed outcome produce different records and must remain independently inspectable.

  1. 01 · ModelEvaluate the declared facts, rules, dependencies, and constraints.

    The result is model output, not an authorized decision.

  2. 02 · ProposalPrepare an exact candidate plan and explanation for review.

    A proposal does not carry institutional authority.

  3. 03 · Human authorizationThe named responsible actor accepts, rejects, or changes the exact reviewed revision.

    An interface action records the scenario step; authority still comes from the responsible institution.

  4. 04 · Submission acknowledgmentThe receiving system records that it accepted the exact instruction for processing.

    Receipt establishes neither execution nor outcome.

  5. 05 · Execution reportThe responsible execution owner separately reports what action was performed.

    Reported execution is not proof of the intended outcome.

  6. 06 · Outcome evidenceAuthoritative observation records what occurred and with what effect.

    An outcome claim requires evidence beyond the model, submission record, and execution report.

Acknowledgment ≠ execution ≠ outcome. Each state requires its own responsible source and evidence record.

Proof and limits

What this scenario supports—and what remains to validate.

These states describe the scenario source and its defined checks. Real-world validation requires separate evidence.

Scenario publication
PublishedReleased August 27, 2026 as an operating scenario.
Source readiness
R2 · Sources reviewedDomain support and product capability boundaries have been reviewed.
Scenario check
Checks not runScenario revision 2026-08-27.1 defines the steps and expected results; the checks have not run yet.
Independent review
PendingThe expected results have not received independent review.
Deployment evidence
NoneNo customer deployment, production performance, or real-world outcome is claimed.
Next proof required
Advance beyond R2Run the defined checks, retain the results, and have an independent reviewer check the expected results.

Evidence and stewardship

What supports this scenario—and when it must be reviewed again.

Illustrative evidence

Authored software-release scenario grounded by reviewed CISA and NIST guidance; it is not a customer deployment, security incident report, service benchmark, runbook, or measured Grid outcome.

Invented elements. Every organization, person, service, release, job, threshold, time, decision, incident, and outcome is invented. Official sources ground only the release-management context, and reviewed Grid documents support only generic primitives—not this composition, fixture behavior, deployment, or result.

Owner
Grid FYI Editorial
Reviewed
August 27, 2026
Review due
February 27, 2027
Source revision
2026-08-27.1
Scenario package
software-release-that-had-to-explain-its-rollback

Related work

Related reading and examples.

These links are chosen as direct companions to this scenario.

Use cases

Insights

White paperReadiness Is a Graph: From Local Green Status to Mission-Qualified AvailabilityA locally valid status does not establish that the right asset, configuration, crew, support, and time window resolve for a particular mission. This paper separates the public readiness record from a bounded Grid proposition and evaluation method.White paperFrom Common Operating Picture to Common Operating ModelSeeing the same facts is not the same as calculating from the same rules. A common operating model connects source identity, dependencies, constraints, alternatives, authority, role-specific views, and replayable evidence.White paperAI Can Propose. Who Authorizes? A Control Model for High-Consequence Government Work.High-consequence AI needs more than a human approval button. This paper separates proposals, facts, models, review, authority, and outcomes, then defines a bounded way to evaluate whether oversight is real.

Films

Continue

Read, watch, or explore the next step.

Thematic companion · 0:36When the Plan BreaksSeveral disruptions invalidate the plan; the affected models recalculate and a person authorizes the response. This released film approaches the same operating theme through a different scenario.Inspect implementation conceptsExplain and validate a changing decisionContinue into Grid Developers for maintained behavior, prerequisites, and implementation limits.Apply the operating patternFrom a Changed Fact to Coordinated ActionWhen one operating fact changes, trace its consequences through the model, compare feasible responses, and update the people and views that depend on the decision.