White paper · Research and learning
Make the Structure Visible
Assumptions, Evidence, and Uncertainty in Research and Learning
Research and learning suffer when a polished result hides its sources, assumptions, exclusions, uncertainty, and revisions. This paper proposes an inspectable model that helps people test the structure of an argument without transferring inquiry, authorship, review, or instructional authority to software.
Narrated film · 1:10
New evidence changes the lesson students are learning.
See a teacher follow changed evidence into classroom materials while preserving the reasoning and deciding what to teach next.
An authored scenario that explains the idea. Read the story and its limits.
▶ 1:10
Inquiry anatomy
A reviewed claim is a six-stage research object.
Each stage can change without becoming a silent edit. The model connects evidence to release while keeping method, interpretation, review, and institutional authority distinct.
- Research and instructional designQuestion and plan
Declare the claim class, population or system, planned method, acceptance tests, owners, and material assumptions.
- Evidence and stewardshipSources and observations
Preserve identity, revision, collection context, time, units, permissions, quality, missingness, and correction relationships.
- Computational modelMethod and execution
Bind selections, exclusions, transformations, code or model, environment, parameters, warnings, failures, and receipts.
- Analysis and argumentResult and claim
Keep the value, uncertainty, sensitivity, scope, limitations, and claim class attached to the exact producing revision.
- Scholarly authorityReview and disagreement
Record objections, counterevidence, responses, deviations, accepted corrections, remaining dissent, and accountable roles.
- Institutional recordRelease and correction
Publish an immutable citable view with contributors, rights, preservation, successor links, and a reproduction package.
Jump to a section in this paper
The answer is often less important than its structure
A chart can look final while its denominator is disputed. A research result can retain a confidence interval but lose the exclusion rule that produced its sample. A course exercise can show the correct number without showing which assumption made it correct. An AI-generated analysis can sound coherent while inventing a source, changing a method, or silently treating an unresolved value as zero.
These failures share a pattern: the visible artifact preserves the answer but hides the structure that makes the answer testable. Sources, definitions, transformations, assumptions, missingness, uncertainty, deviations, and review decisions become prose fragments or private working files. When one of them changes, people reconcile the consequences by hand.
The stronger objective is not automated truth. It is an inspectable chain of inquiry in which a reader can distinguish observation from assumption, description from causal claim, planned analysis from post-hoc exploration, and computational result from an authorized scholarly or instructional judgment.
This paper first describes selected public research and education guidance. It then presents a separate Grid proposition for making that structure executable and reviewable. The cited institutions do not describe or endorse Grid, and none of their guidance establishes that a Grid model is scientifically valid, reproducible, pedagogically effective, legally compliant, or fit for a particular study.
What current public guidance establishes
The White House Office of Science and Technology Policy's June 2025 Guidance for Gold Standard Science applies to federal agencies under its stated executive direction. It addresses reproducibility, transparency, communication of error and uncertainty, skepticism about findings and assumptions, falsifiability, peer review, negative results, conflicts, and other practices. Those are responsibilities for federal scientific activities within that instrument's scope—not a universal rating for every study and not a software specification.
The Institute of Education Sciences makes the importance of visible structure concrete. Its current Standards for Excellence in Education Research say that theory and conceptual frameworks should be articulated, hypotheses should be testable and falsifiable, limitations and uncertainty should be communicated, and applicable work should preserve preregistration and explain deviations. IES also says the standards apply to IES-funded work as specified by the controlling request for applications or proposals. A model cannot declare that a study meets them.
The IES What Works Clearinghouse Procedures and Standards Handbook, Version 5.0 provides defined procedures for screening studies, reviewing findings, and synthesizing evidence. Its existence demonstrates why “evidence” is not one undifferentiated label: design, sample, outcome, attrition, analysis, and synthesis rules affect what a finding can support. A result represented in Grid would not receive a WWC rating unless the actual WWC process assigned one.
Data stewardship has similarly bounded scope. The NIH Data Management and Sharing Policy applies, with stated exceptions, to NIH-funded or conducted research that generates scientific data and requires covered applicants to submit a DMS Plan. NIH's February 2026 updated DMS Plan elements introduced the updated format for competing applications with due dates on or after May 25, 2026. A July 2026 implementation update says that, beginning in fiscal year 2027, other applications and active awards subject to the policy transition through Just-in-Time or the Research Performance Progress Report as specified; effective October 1, 2026, recipients report progress and notify NIH of approved-plan changes through the RPPR rather than the prior-approval process. These requirements do not make every datum public; privacy, consent, repository, access, and program-specific rules remain.
NIST's measurement-uncertainty guidance defines measurement uncertainty as a parameter associated with a measurement result that characterizes the dispersion of values attributable to the measurand, and it presents methods for propagating input uncertainty through a measurement model. That domain is narrower than all research uncertainty, but it supports a useful discipline: uncertainty should remain attached to the quantity and method that produced it, not added as decorative language after the calculation.
For education, the Department of Education's 2023 report, Artificial Intelligence and the Future of Teaching and Learning, recommends keeping educators central as instructional decision makers and treating AI models as incomplete representations to be examined against educational goals. It is advisory, not a finding that any AI system improves learning or that human review alone makes one safe or appropriate.
Together, these sources establish scoped practices and institutional responsibilities. They do not establish one universal research method, compel open disclosure of protected information, replace peer review or an institutional review board, decide authorship, or validate Grid.
Represent the claim, not merely the calculation
Grid proposes a versioned claim package that connects the parts of an analysis while keeping their meanings separate:
- Question and claim class: the research or learning question, population or system of interest, and whether the proposed output is descriptive, associational, predictive, causal, explanatory, or exploratory.
- Sources and observations: origin, collection method, instrument, observation time, units, permissions, transformations, quality state, and supplied uncertainty.
- Definitions and exclusions: constructs, operational definitions, inclusion and exclusion rules, missing-data treatment, subgroup boundaries, and denominators.
- Method and assumptions: equations, code or model revision, parameters, controls, dependency structure, sensitivity cases, and conditions under which the method should not be used.
- Plan and deviations: preregistered or approved analysis, later changes, reason, approving role, timing, and whether a result became exploratory.
- Result and uncertainty: calculated values, intervals or other uncertainty expressions where appropriate, limitations, unresolved conflicts, and sensitivity to material assumptions.
- Review and communication: reviewer comments, accepted corrections, provenance, conflicts, disclosure rules, and the exact artifact approved for a named purpose.
An AI assistant may propose a formula, mapping, explanation, test case, or alternative interpretation. Its proposal remains labeled with the prompt context, model identity where available, source references, and review state. A qualified person must determine whether the question, method, evidence, analysis, and communication are appropriate. Fluent language is not authorship, peer review, instructor approval, or scientific evidence.
Different people need different views of the same inquiry
A researcher may need source-level observations, code, model assumptions, deviations, and diagnostic results. A peer reviewer may need enough method and evidence to challenge the claim without receiving restricted participant data. A student may need a scaffold that reveals dependencies progressively rather than exposing an answer key. An instructor may need misconceptions, submitted reasoning, and an approved rubric while retaining assessment authority. A program leader may need status, limitations, and reproducibility evidence without being shown personally identifiable information.
Those are role-specific projections, not independent copies of the method. Each view should identify the same accepted revision and declare what it omits. Access control, de-identification, consent, research records, intellectual property, student privacy, accessibility, accommodations, and academic-integrity rules remain the institution's responsibility.
The roles also need explicit decision rights. “Human reviewed” is too vague to show what was actually governed.
| Role | May contribute | Authority it does not gain automatically |
|---|---|---|
| Source steward | Accept or correct source identity, metadata, access state, and quality disposition | Authority to choose the method or approve the claim |
| Research or methods lead | Approve represented definitions, exclusions, assumptions, method, and deviation treatment | Journal acceptance, participant consent, or institutional compliance |
| Analyst or computational operator | Execute the accepted method and preserve receipts, warnings, failures, and comparisons | Authority to reinterpret a result or hide an unresolved defect |
| Peer or domain reviewer | Challenge the evidence, method, uncertainty, inference, and communication in the declared review scope | Authorship, institutional release, or authority beyond that review |
| Instructor or assessment owner | Decide instructional use, feedback, accommodations, and grading under the actual policy | Scientific validation or permission to expose restricted research data |
| Release authority | Approve the exact artifact for the named audience, repository, journal, or course surface | Proof that every underlying scholarly judgment is correct |
Fictional exercise: Harbor Basin Methods Studio
Consider a wholly fictional combined research-and-learning fixture. Harbor Basin Methods Studio examines whether an invented shade treatment is associated with lower afternoon surface temperature across four synthetic zones. It uses no real people, schools, environmental claim, or field data.
The approved exercise plan defines a descriptive weighted mean, not a causal effect. Every observation is fictional but expressed in degrees Celsius (°C), and the contrast is untreated minus treated temperature; a positive value therefore describes a lower treated value within the fixture. Zone A has 30 accepted observations with a mean difference of 3.0 °C; B has 25 at 2.0; C has 20 at 1.0; and D has 25 at 2.0. The authored calculation is (30×3 + 25×2 + 20×1 + 25×2) ÷ 100 = 2.10. The result means only that the accepted fictional observations produce a 2.10 °C descriptive difference under fixture revision HB-4.
Students see the declared question, four source groups, units, calculation, and a list of assumptions. They must identify which facts would change the denominator and which would change interpretation. The model can evaluate their arithmetic against the fixture; the instructor determines whether their reasoning meets the course objective and how any work should be graded.
A later calibration review finds that ten Zone A observations used an invalid fictional instrument setting. Those ten excluded records contributed 50 observation-°C to the original weighted numerator; the 20 retained records contribute 40. Source revision A-12 supersedes A-11, and Zone A's accepted mean is therefore 2.0 °C. The revised calculation is (20×2 + 25×2 + 20×1 + 25×2) ÷ 90 = 1.78 after rounding. The model shows the changed source, denominator, contribution, result, and every downstream draft that used HB-4.
It does not conclude that 1.78 is true, causal, important, generalizable, or publishable. The research lead must decide whether the calibration issue requires a protocol deviation, additional collection, different analysis, disclosure, or withdrawal. The instructor decides whether to reveal the correction immediately or use it as a learning exercise. Prior outputs remain reconstructable rather than silently changing.
Unknown, disagreement, and negative results belong in the model
Research structures often fail at their boundaries. Two teams use the same construct name with different definitions. An instrument revision changes comparability. A missingness assumption is undocumented. A planned test yields no supported difference, so attention moves to an unplanned subgroup. An AI assistant supplies a plausible citation that does not exist.
The model should preserve those conditions as explicit states. A missing source remains missing. Conflicting definitions stay unresolved until a named owner adjudicates them. A post-hoc analysis remains distinguishable from the planned analysis. A negative or null result remains connected to the question and method. A retraction, correction, or superseding dataset creates a new state without erasing what earlier readers received.
This visibility does not resolve epistemic disagreement. It gives responsible people a shared object to interrogate.
A bounded evaluation contract
Begin with one completed synthetic or appropriately governed historical analysis and no live grading, publication, participant, or policy consequence. Record the current baseline: time to reconstruct the result, undocumented assumptions found, manual transformations, inconsistent denominators, unreproducible steps, review cycles, and corrections that cannot be traced to affected outputs.
Before configuring Grid, have domain, methods, data-stewardship, education, privacy, accessibility, and security owners define expected cases. Seed a stale source, unit mismatch, duplicate observation, excluded record restored accidentally, changed denominator, missing uncertainty, unapproved method deviation, fabricated citation, restricted field exposed to the wrong view, inaccessible explanation, and AI-authored conclusion presented as approved.
Measure at least six layers:
- Provenance: material inputs retain source, revision, definition, unit, time, transformation, and access state.
- Computational fidelity: independently calculated fixture results and sensitivity cases match; the same accepted revision replays deterministically.
- Claim alignment: descriptive, predictive, causal, and exploratory outputs remain distinguishable; unsupported claim escalation is blocked in the fixture.
- Change control: a corrected source identifies affected calculations, drafts, teaching materials, and reports while preserving prior states.
- Review and authority: AI proposals, researcher acceptance, peer review, instructor decisions, and publication or release approval remain separate and attributable.
- Responsible access: role-specific views omit restricted fields as configured; accessibility, privacy, consent, records, and intellectual-property reviews retain their independent gates.
Retain the source fixtures, data dictionary, plan, assumptions, expected arithmetic, model and environment revisions, deviations, sensitivity cases, prompts and generated proposals, reviewer actions, access tests, explanations, defects, and final evaluation record.
A passing fixture supports at most a product-demonstration claim for the declared materials and conditions. It does not establish scientific validity, replication or reproducibility by independent researchers, research-integrity compliance, privacy or human-subjects compliance, appropriate authorship, instructional effectiveness, accreditation, publication quality, deployment readiness, or improved learning or research outcomes.
Teams can walk the full correction lifecycle in Trace a Research Claim from Source to Reviewed Release, then use The Auditable Executable Paper and its circulation-ready edition for the wider category. How to Evaluate Executable Decision Infrastructure defines acceptance evidence, AI Can Propose. Who Authorizes? separates human and AI roles, and the decision evaluation worksheet records the fixture. Implementation teams can continue with Grid Developers' explain-and-validate guidance.
Making the structure visible does not settle an inquiry. It makes the question, evidence, assumptions, uncertainty, changes, and responsible judgments available for challenge—which is exactly what a trustworthy research or learning process needs.
Related films, scenarios, and next steps
Continue exploring
Follow the next question.
How Grid Works: From Typed Facts to Accountable Action
A technical and operational guide to the six-stage Grid chain: governed sources, a shared and versioned model, reactive constraint evaluation, coordinated surfaces, explanation and evidence, and AI assistance bounded by human authority.
Understand · EvaluateSource-grounded ExploreThe Authority Contract for AI-Assisted Decisions
Where Proposals Stop and Organizational Action BeginsHuman-in-the-loop language is too vague for consequential AI-assisted work. This paper proposes an explicit authority contract that defines what an AI system may observe, propose, validate, call, and never authorize—and how the organization tests that boundary.
Govern · EvaluateSource-grounded ExploreThe Auditable Executable Paper
Grids' strongest academic category is the auditable executable paper: a living research object in which every important number, figure, and claim can explain its source, assumptions, revisions, contradictions, review, authorization, and reproduction path.
Understand · EvaluateSource-grounded Explore