Use case · Education and research

Trace a Research Claim from Source to Reviewed Release

Keep a consequential finding connected to its exact source revision, eligibility and method rules, arithmetic, affected table and prose, reviewer disposition, and immutable corrected release.

August 26, 202615 minute readoperating guide

Narrated film · 1:10

New evidence changes the lesson students are learning.

See a teacher follow changed evidence into classroom materials while preserving the reasoning and deciding what to teach next.

Watch film · 1:10 Read the use case · 15 min
An authored scenario that explains the idea. Read the story and its limits.
▶ 1:10
Jump to a section in this publication

A published sentence can outlive its denominator

A source steward corrects fifteen records. The analysis reruns. One table changes, but a chart exported yesterday does not. A sentence in the abstract has already been copied into a slide deck. The repository still serves the first release, while an internal folder contains a file called final-corrected-2. Reviewers now have to reconstruct which source, selection rule, run, table, sentence, and approval belonged together.

This is not only a version-control problem. A research claim is assembled across data systems, analysis code, notebooks, spreadsheets, documents, review tools, repositories, and publishing systems. Each tool may preserve its own history while losing the relationship that matters: this exact claim was supported by this exact result, produced from this exact eligible set and method, reviewed in this scope, and released in this exact package.

The objective is not a machine that decides whether research is true. It is a connected research object that makes a claim answerable when a source, assumption, method, interpretation, or review decision changes.

Start with the distinctions research already requires

The National Academies' consensus report on reproducibility and replicability distinguishes obtaining consistent computational results with the same inputs and methods from obtaining consistent findings through a new study. That difference matters here. Replaying an executable package may establish that the declared computation can be reproduced under stated conditions. It does not establish empirical replication, causal validity, generalizability, or truth.

The Institute of Education Sciences' current Standards for Excellence in Education Research emphasize articulated theory, testable hypotheses, objective design, uncertainty, limitations, and accurate reporting. The What Works Clearinghouse handbooks show why sample, outcome, attrition, analysis, and synthesis decisions cannot be collapsed into one generic evidence label. Those requirements apply within the programs and processes that define them; a Grid record cannot award a WWC rating or declare compliance.

For covered federal work, the 2025 OSTP Guidance for Gold Standard Science addresses transparency, assumptions, uncertainty, negative results, conflicts, peer review, and reproducibility within its stated scope. The NIH Data Management and Sharing Policy adds plan, repository, access, and change obligations where applicable. Neither source is a universal software specification, and responsible sharing never means that every datum should be public.

The operating pattern below uses those distinctions as evaluation constraints. The cited institutions do not describe or endorse Grid.

Define a source-to-release claim contract

Before connecting a repository or rebuilding an analysis, select one consequential claim and write down the objects that must remain connected.

Object Minimum identity and state Question it must answer
Research question and plan Claim class, population or system, outcome, inclusion and exclusion rules, method, thresholds, uncertainty, deviation policy, accepted revision What was the team trying to learn, and what was decided before seeing this result?
Source assertion Stable source identity, revision, observation or collection time, units, quality, permissions, correction relation, steward Which evidence was actually eligible for this run?
Eligible set Exact membership or reproducible selection receipt, excluded cases and reasons, plan revision How did the source become the analytic denominator?
Execution Code or model revision, environment, parameters, seeds, order, warnings, failures, external receipts What ran, under what conditions, and did it finish as declared?
Result and claim Exact value, uncertainty, table or figure location, claim text, scope, limitations, supporting and contradicting evidence What does the result support, and what does it not support?
Review and release Objection, response, disposition, reviewer role and scope, accepted revision, immutable package, successor or correction relation Who accepted what for which audience, and what was actually published?

The contract should declare the invalidation rule before the correction arrives. A changed source might affect one cell, an entire eligible set, several tables, a prose claim, an abstract, a teaching slide, and the release that bound them. If the team cannot name that expected reach in advance, it cannot tell the difference between correct propagation and accidental churn.

Synthetic exercise: freeze the Cedar oracle

The Cedar Methods Exercise is wholly synthetic. It contains no people, students, schools, intervention, human-subject record, customer, deployment, or efficacy claim. The round figures are chosen so another evaluator can reproduce every operation without a specialist package.

The question is deliberately descriptive:

Within the generated practice dataset, what proportion of eligible profiles improved by at least 10 points between synthetic pre- and post-scores?

Accepted analysis plan PLAN-4 declares:

  • include only records with account_type = research;
  • require both synthetic pre- and post-scores;
  • calculate change = post_score - pre_score;
  • set met = true when change >= 10;
  • report the pooled descriptive proportion met / eligible;
  • evaluate the two-thirds gate exactly as 3 × met >= 2 × eligible, rather than comparing rounded percentages; and
  • make no causal, effectiveness, generalizability, or real-world educational claim.

The frozen baseline uses roster ROSTER r5, eligible set ELIGIBLE-SET-18, execution RUN-18, Table 2 revision 3, and claim CLAIM-2 r2.

Cohort Eligible synthetic profiles Met the 10-point threshold Descriptive rate
A 30 24 80.0%
B 25 17 68.0%
C 20 12 60.0%
D 25 17 68.0%
Total 100 70 70.0%

The exact gate passes: 3 × 70 = 210, while 2 × 100 = 200. Release CEDAR-REL-1.0.0 therefore carries this bounded sentence:

In the synthetic dataset, 70 of 100 eligible profiles (70.0%) met the 10-point threshold, satisfying the predeclared two-thirds descriptive gate.

The sentence says nothing about why a score changed. It does not describe an intervention, a treatment effect, a validated instrument, a population, or a result that should guide instruction. Its value is that every quantity and boundary is explicit enough to serve as an evaluation oracle.

One corrected source changes the denominator and the conclusion

The source steward receives evidence that the export classified fifteen Cohort D calibration profiles as research profiles. Corrected roster ROSTER r6 supersedes ROSTER r5 and identifies the affected synthetic records as D011 through D025. All fifteen had met = true in the baseline run.

Because PLAN-4 already excludes non-research profiles, applying the corrected account type is a source correction under the accepted plan. The analyst is not inventing a new exclusion after seeing the result. That distinction must be visible in the record.

Cohort Release 1.0 eligible / met Accepted source correction Corrected eligible / met Corrected rate
A 30 / 24 none 30 / 24 80.0%
B 25 / 17 none 25 / 17 68.0%
C 20 / 12 none 20 / 12 60.0%
D 25 / 17 remove 15 / 15 calibration profiles 10 / 2 20.0%
Total 100 / 70 −15 / −15 85 / 55 64.7%

The corrected proportion is 55 / 85 = 11 / 17 = 0.6470588235…, displayed to one decimal as 64.7%. The change from the first release is −5.2941176 percentage points, displayed as −5.3 percentage points.

The exact gate now fails: 3 × 55 = 165, while 2 × 85 = 170. No rounding convention decides the outcome. Candidate claim CLAIM-2 r3 becomes:

After excluding 15 calibration profiles misclassified in roster r5, 55 of 85 eligible synthetic profiles (64.7%) met the threshold; the predeclared two-thirds descriptive gate was not met.

If the team instead changes the threshold, creates a new exclusion, or changes how missing scores are handled, it has changed the method. That should create a new plan revision, such as PLAN-5, and be disclosed as a deviation or exploratory analysis. Calling a post-hoc method change a “data correction” would fail the exercise.

Staleness must travel farther than the calculation

The changed roster should expose the full stale cone rather than silently overwriting it:

ROSTER r5 → ELIGIBLE-SET-18 → RUN-18 → Table 2 cohort D → Table 2 total → CLAIM-2 r2 → abstract sentence → CEDAR-REL-1.0.0

The accepted correction creates a successor path:

ROSTER r6 → ELIGIBLE-SET-19 → RUN-19 → Table 2 r4 → CLAIM-2 r3 → correction notice → CEDAR-REL-1.0.1

Table 2's Cohort D row and total, the prose claim, and any bound abstract, caption, slide, or teaching view become stale. Release 1.0 does not disappear. It remains immutable historical evidence, visibly identified as corrected or superseded, with a navigable relationship to release 1.0.1.

That behavior separates a living research object from a collection of synchronized files. A recalculation is necessary, but it is not sufficient. The system also has to preserve the old context, stop unreviewed language from looking current, identify every affected surface, and let a reader reconstruct what was known and accepted in each release.

Review the correction without collapsing authority

“Human reviewed” is not enough. The record needs to say which person held which role, what evidence that role examined, which disposition was available, and where the person's authority stopped.

Role May decide in the exercise Authority not gained automatically
Source or data steward Correct account type; document affected IDs, basis, time, and revision relationship Method validity, interpretation, authorship, or release
Analyst Recompute both states; expose row and aggregate differences; prepare a candidate correction Approval of the analyst's own method or publication
Methods reviewer Decide whether the exclusion implements PLAN-4; accept, reject, or require a method revision Attestation to source-system truth outside the review scope
Domain reviewer Challenge interpretation, claim class, limitations, and explanatory wording Invisible alteration of data, formula, or provenance
Research lead Accept the corrected scholarly claim and approve an institutional package within actual delegation Journal peer review, publisher acceptance, or repository authority
Repository or publisher Issue a formal version or correction under its process Automatic acceptance of an internal Grid approval
AI assistant Propose mappings, tests, explanations, and wording with attribution Source acceptance, method approval, authorship, review, or release

For the passing path, the methods-review disposition should be reconstructable in substance:

accept_source_correction; PLAN-4 unchanged; exclusion was predeclared; corrected claim required.

A separate release receipt from the actual research lead and repository or publisher process should record that release 1.0 remains retained and that release 1.0.1 was authorized for the named purpose. Methods acceptance is evidence for that later decision; it is not publication authority.

The exact vocabulary can differ. The essential properties are the source of the correction, the unchanged or changed plan, the affected artifacts, objections and responses, the decision, the actor's role and scope, the accepted revision, and the intended release surface. A credential proves identity. It does not prove that the person had the institutional or scholarly authority represented by the action.

Release a portable successor, not a silent replacement

The connected object needs an exit path. The W3C PROV Ontology provides a vocabulary for entities, activities, agents, derivation, attribution, association, primary sources, and revisions. RO-Crate 1.3 describes a JSON-LD research-object package that can identify files, linked resources, people, organizations, software, citations, licenses, and reuse conditions.

Other standards address narrower responsibilities. The DataCite Metadata Schema can represent version and relation metadata for citable objects. CRediT supplies structured contributor roles. CITATION.cff makes software and dataset citation information readable by people and tools. JATS supports journal-content interchange.

Portable element Required content for CEDAR-REL-1.0.1 What its presence does not prove
Source and execution manifest Both roster identities, affected-record receipt under appropriate access, PLAN-4, environment, code or model, runs, warnings, and exact outputs Correct source data, valid method, or a successful independent replay
Provenance and research-object package Qualified relationships among sources, selections, runs, tables, claims, reviews, contributors, rights, and releases PROV-O or RO-Crate conformance without validation and round-trip testing
Citation and contributor metadata Version, predecessor and correction relation, preferred citation, software and data citation, reviewed role assertions DOI registration, authorship determination, peer review, or scholarly merit
Article and correction snapshot Old and corrected tables and claims, correction notice, limitations, accepted article revision Journal acceptance, preservation, retraction status, or publication authority
Independent replay recipe Permitted inputs, dependency and runtime identities, order, parameters, expected outputs, tolerances, and comparison record Empirical replication, construct validity, causal inference, or truth

A public package may expose qualified metadata while keeping row-level or licensed material in an authorized repository. Access state, license, consent, participant protection, data-use agreements, Indigenous or community authority, export controls, student privacy, and retention remain institutional responsibilities. “Portable” cannot become a pretext for indiscriminate disclosure.

Rehearse failure before trusting the happy path

The fixture should fail if any of the following occurs:

  • roster r6 overwrites r5 or deletes the fifteen profiles from history;
  • the exclusion is invented by the analyst instead of derived from the accepted source and PLAN-4;
  • a changed threshold or missing-data rule is disguised as a source correction;
  • Cohort D recalculates while the total, claim, abstract, caption, or export remains current;
  • rounded percentages, rather than the declared integer comparison, determine the gate;
  • one person approves a source, method, interpretation, and release without the declared separation or exception record;
  • an AI-generated correction strengthens the descriptive result into a causal or effectiveness claim;
  • release 1.0 mutates, loses its identity, or cannot be reconstructed byte for byte;
  • a public export leaks fields that the release policy marks restricted;
  • export presence is marketed as standards conformance without validation; or
  • internal approval is described as journal peer review, a WWC rating, publisher acceptance, or independent validation.

Inject each failure deliberately. Preserve the attempted transition, the reason it stopped, and any reviewer disposition. A model that produces the corrected percentage but allows the wrong release or authority transition has not passed.

Evaluate a no-consequence pilot with predeclared evidence

Freeze ROSTER r5, ROSTER r6, PLAN-4, both expected tables, both claim texts, the stale cone, permitted roles, release policy, and export expectations before implementation. Run the existing workflow and the Grid proposition from the same fixture. Do not connect live participant data, grading, publication, policy, funding, or operational consequence.

Evaluation area Minimum passing evidence A pass still does not establish
Arithmetic Independent calculation reproduces 70/100, 55/85, 11/17, −5.2941176 points, and both integer gate comparisons Statistical validity, causal inference, or generalizability
Lineage Every consequential table cell and sentence resolves to the exact source, plan, run, review, and release revision Source truth or completeness of the represented model
Staleness Every predeclared affected artifact is blocked or marked stale; unrelated artifacts do not change Correct handling of every future study design
Change classification Source correction, planned exclusion, method revision, and exploratory branch remain distinguishable That the scientific choice itself was appropriate
Authority Unauthorized review and release transitions stop and remain attributable Institutional delegation, peer review, or publisher acceptance outside the fixture
Replay A second evaluator reconstructs both states within declared tolerances; release 1.0 stays immutable Replication through a new study or environment portability beyond those tested
Portability Independent tools inspect the declared package, provenance, citation, and interchange outputs without undocumented intervention Universal conformance, preservation, or repository acceptance
Responsible access Public views omit restricted fixture fields while authorized reviewers can recover the evidence they need Compliance with privacy, consent, records, security, or data-governance obligations
Claim alignment No surface calls the exercise causal, effective, replicated, peer reviewed, WWC-rated, or generally valid Truth of any real research claim
Correction receipt Source change, affected IDs, arithmetic, objections, disposition, wording, and release relation are recoverable together Research impact, adoption, or improved scholarly outcomes

Measure more than task completion. Record time to identify the affected claim, unexplained manual steps, false-positive and false-negative stale flags, reconciliation effort, reviewer comprehension, unauthorized transitions blocked, export defects, and the work needed for an independent replay. Compare those measures with the existing file-and-copy workflow.

The handoff is a governed model, not a publishing claim

A bounded implementation can use Grid's documented reactive model concepts to represent named dependencies and typed values, and its explain-and-validate workflow to inspect change. External statistical packages, notebooks, repositories, identity systems, data stores, review platforms, DOI services, and publisher workflows remain authoritative for the functions assigned to them. Each connector, import, export, permission, retry, signature, and receipt is configured work that must be tested in the actual environment.

Continue with Make the Structure Visible for the compact model, The Auditable Executable Paper for the wider category and cross-disciplinary architecture, and the circulation-ready white paper for a portable review artifact. Use How to Evaluate Executable Decision Infrastructure and the decision evaluation worksheet to define the evidence boundary before configuring a pilot.

The pass condition is not that the system makes the corrected claim automatically. It is that another qualified person can see what changed, reproduce the represented consequence, inspect disagreement, verify who accepted the correction, recover both releases, and understand exactly what none of that proves.

Related films, scenarios, and next steps

Choose the next move

Test the claim with a different kind of evidence.

For researchersTake the working resource into the conversationUse the printable assessment or brief to make assumptions, authority, and remaining proof concrete.See the modelThe Future of EducationA class investigates why one room loses heat faster than its model predicts, making questions, evidence, assumptions, and student judgment visible.For technical evaluatorsFollow the concept into Grid DevelopersContinue into the linked Grid Developers guide for the exact behavior, prerequisites, and limits used by this explanation.

Continue exploring

Follow the next question.

White paper · 21 min Can every important number, figure, and claim answer where it came from, what produced it, what changed, what contradicts it, who authorized it, and how to reproduce it?

The Auditable Executable Paper

Grids' strongest academic category is the auditable executable paper: a living research object in which every important number, figure, and claim can explain its source, assumptions, revisions, contradictions, review, authorization, and reproduction path.

Understand · EvaluateSource-grounded Explore
Use case · 13 min Can coalition participants calculate the same permitted conclusion, continue locally through a connection loss, and reconcile change without one system assuming national release or action authority?

Coordinate Coalition Decision Logic Without Centralizing Sovereign Data or Authority

Execute an agreed decision contract in each participant’s environment, exchange only authorized claims and results, and reconcile revisions without letting a shared system assume national release or action authority.

Understand · PlanSource-grounded Explore