White paper · AI authority

AI Can Propose. Who Authorizes?

A Control Model for High-Consequence Government Work.

High-consequence AI needs more than a human approval button. This paper separates proposals, facts, models, review, authority, and outcomes, then defines a bounded way to evaluate whether oversight is real.

August 25, 20269 minute readdeep-dive

Public funding decisions · 1:10

AI can prepare the award list. It cannot approve it.

Follow the rules, evidence, and exceptions that a responsible person must review before a public award is released.

Watch film · 1:10 Read the publication · 9 min
An authored scenario that explains the idea. Read the story and its limits.
▶ 1:10
Jump to a section in this paper

The authorization gap

An AI system can summarize a case file, draft a course of action, rank applications, identify an apparent threat, or translate policy into executable logic. None of those acts answers the institutional question that follows: who may authorize the consequential action?

Government work assigns authority through law, regulation, delegation, command relationships, professional responsibility, and mission-specific policy. A fluent recommendation cannot inherit that authority. Neither can generated code, a confidence score, a user account, or an approval button.

The central control problem is therefore not simply keeping a person “in the loop.” It is keeping unlike responsibilities separate enough to inspect: what AI proposed, which facts were accepted, which model evaluated them, what a reviewer actually examined, which person or process held authority, and what happened after action. Collapse those boundaries and a human signature can legitimize a decision without revealing how it was produced.

This paper first describes what public institutions say about responsible AI. It then presents a distinct Grid proposition: a control model for making those boundaries executable and testable. The institutional guidance establishes governance needs. It does not endorse Grid or establish that this proposition satisfies them.

What institutional guidance establishes

The DoD Responsible Artificial Intelligence Strategy and Implementation Pathway organizes responsible AI around governance, warfighter trust, the AI product and acquisition lifecycle, requirements validation, an RAI ecosystem, and workforce capability. It operationalizes DoD's five AI Ethical Principles: responsible, equitable, traceable, reliable, and governable. Traceability includes appropriate understanding of technology, development processes, operational methods, data sources, and documentation; reliability depends on explicit uses and lifecycle testing; governability includes detecting and avoiding unintended consequences and the ability to disengage or deactivate systems that behave unintentionally.

The DoD Responsible AI Toolkit provides a voluntary, tailorable process for identifying, tracking, and improving alignment with those principles across the AI lifecycle. Its release describes a living toolkit, not a certification or authorization. “Aligned with” is an evaluation claim that needs evidence; it is not a badge a supplier can award itself.

The NIST Artificial Intelligence Risk Management Framework 1.0 treats governance as continuous and cross-cutting. It calls for documented roles, responsibilities, risk tolerances, system scope, human-oversight processes, testing, evaluation, verification, validation, and monitoring over time. NIST also cautions that transparency alone does not make a system accurate, secure, private, or fair. When consequences are severe, transparency and accountability practices should increase proportionally.

The GAO AI Accountability Framework complements that lifecycle view with four principles: governance, data, performance, and monitoring. GAO asks whether organizations have clear goals and responsibilities, reliable and representative data, performance measures tied to program objectives, and continuing mechanisms to detect drift and reassess results. Its questions are designed for managers, auditors, and third-party assessors—not only developers.

Together, these sources support an institutional conclusion: accountability needs defined roles, attributable evidence, bounded uses, testing, monitoring, and the ability to intervene. They do not select a product architecture, determine the lawful decision maker, or prove that a nominal human review step is meaningful.

One defense source requires especially precise scope. DoD Directive 3000.09, Autonomy in Weapon Systems, applies to autonomous and semi-autonomous weapon systems. It requires those systems to allow commanders and operators to exercise “appropriate levels of human judgment over the use of force” and establishes weapons-specific review, verification, validation, testing, and evaluation requirements. It does not create a universal human-in-the-loop rule for every government AI system, and this paper does not generalize it beyond its weapons context.

Six boundaries for an accountable decision

A high-consequence workflow should preserve six records that can be connected without being confused.

1. Proposal boundary

The proposal is what AI contributed: a summary, extracted claim, candidate rule, priority order, forecast, or course of action. It should retain the instruction, relevant context, model and configuration identity, creation time, output revision, and known limitations. A proposal may be useful and still be incomplete, unsupported, or outside policy.

The boundary answers: What may AI draft or recommend, and what must it never initiate or authorize?

2. Fact boundary

A fact is not a statement made confidently by a model. It is an input accepted for this decision under an organization's source and evidence rules. Its record should identify the source, subject, revision, observation or effective time, classification or handling constraints where applicable, and any dispute, uncertainty, or expiration condition.

The boundary answers: Which claims have authoritative support, as of when, and which remain missing, stale, inferred, or contested?

3. Model boundary

The model contains the approved calculations, rules, thresholds, constraints, permissions, and represented alternatives used to evaluate the proposal. It should distinguish deterministic rules from estimates, record its revision, identify assumptions, and expose why an option passed, failed, or remained unresolved. Generated logic stays a proposal until the relevant owner validates and governs it.

The boundary answers: Which approved interpretation produced this finding, and within what validated use?

4. Review boundary

Review is the work a qualified person performs on a reviewable packet. That packet should show the proposal, source-backed facts, model revision, material assumptions, uncertainty, alternatives, exceptions, and reasons. It should also show what the system cannot conclude.

The boundary answers: What evidence must the reviewer inspect, challenge, correct, or escalate before deciding?

5. Authority boundary

Authority belongs to a named role or external process recognized by the institution. The record should connect the actor to the applicable delegation, scope, conditions, time, and decision. Authentication can establish identity; it does not by itself establish that the person is authorized for this action.

The boundary answers: Who may approve, narrow, defer, reject, or execute the action, and what proves that authority applied?

6. Outcome boundary

An authorized decision is not an observed outcome. The policy may have been weak, conditions may have changed, execution may have failed, or an unforeseen harm may have occurred. Outcome records should preserve what happened, when, to whom, and under which operating conditions without rewriting the original recommendation as a success.

The boundary answers: What occurred after authorization, and what should monitoring, appeal, incident response, or model review learn from it?

Meaningful oversight versus ceremonial review

A reviewer is not meaningful merely because a workflow requires a click. Oversight becomes ceremonial when the reviewer lacks time, competence, relevant evidence, permission to disagree, or a feasible alternative to approval. It is also weak when the recommendation arrives framed as settled, automation bias is not considered, overrides trigger punishment without examination, or the system cannot pause safely.

Meaningful oversight gives the reviewer:

  • A clear decision and scope, not a generic request to “review AI output”
  • Source identity, freshness, uncertainty, and model revision
  • Reasons for acceptance and rejection of represented alternatives
  • Enough time and domain competence for the consequence involved
  • Practical choices to correct, narrow, defer, escalate, or reject
  • A safe way to stop action when facts, policy, or authority are insufficient
  • A durable record of questions, changes, decision, actor, and time

Approval rate is therefore a poor standalone metric. A reviewer who approves 99 percent of recommendations may be seeing an excellent system—or may be functioning as a rubber stamp. Better evidence includes whether reviewers detect seeded defects, request missing evidence, make substantive changes, use deferral appropriately, and can explain the decision basis without relying on the AI's prose.

The Grid proposition

Grid proposes representing this control model as shared, executable decision infrastructure. AI may help author a candidate formula, translate a policy statement, assemble an alternative, or explain an evaluated result. Named sources provide versioned facts. A governed model applies explicit rules and constraints. Explanation exposes dependency paths and reasons. Different interfaces can present the same modeled revision to mission, legal, technical, and leadership roles. An accountable person or approved external process authorizes any consequential effect.

In shorthand:

AI proposes → sources support facts → the model evaluates → a qualified person reviews → recognized authority decides → outcomes are monitored

This is a product and architecture proposition. It does not establish that every relevant policy can be represented, that source data is true, that an explanation is sufficient, that a person exercised sound judgment, or that the resulting action is lawful, fair, safe, or effective. Those claims require evidence in the intended organization and mission.

A bounded evaluation

Begin with one recurring decision class, one actual authority chain, a small source set, and consequences that can be tested without allowing the prototype to take live action. Write the authority contract and evaluation plan before configuring the model.

Build a controlled set of historical, synthetic, and adversarial cases. Include clean cases, ambiguous cases, stale and conflicting facts, missing policy, an out-of-scope request, a proposal that exceeds the reviewer's delegation, a model revision during review, and an apparently plausible recommendation that should be rejected. Compare the current workflow with the proposed separated workflow. Use qualified reviewers who do not know which defects were seeded, and predefine stop conditions for unsafe or unexplained behavior.

Measure at least five groups:

  • Boundary integrity: proposals correctly labeled; material facts linked to source, revision, and time; generated rules prevented from silently becoming approved logic; actions blocked when required authority evidence is absent.
  • Review quality: seeded defects detected; missing evidence requested; substantive corrections, deferrals, escalations, and rejections; reviewer comprehension of the decision basis; time under representative workload.
  • Execution fidelity: repeated cases reproduce the same governed result; every result binds to a model revision; accepted and rejected alternatives carry inspectable reasons; every seeded unauthorized-effect attempt is blocked and logged in the test boundary.
  • Performance and harm: false acceptance and rejection rates; mission-specific error severity; subgroup or stakeholder impacts where relevant; incidents, appeals, overrides, and near misses. Aggregate accuracy should not hide rare catastrophic failure.
  • Monitoring and resilience: drift and stale-source detection; behavior under delayed or disconnected inputs; time to correction; ability to suspend a model, roll back a revision, and reconstruct a sampled decision.

Set acceptance thresholds by consequence and policy, not by what the prototype happens to achieve. Retain fixtures, source snapshots, prompts where permitted, model and package revisions, reviewer actions, test results, and evaluator findings. A successful demonstration can support a product-demonstration claim within those fixtures. It cannot support an operational-effect claim until customer-controlled testing occurs in the intended environment.

Limits that remain outside the model

No computational pattern grants legal or command authority. The responsible organization must determine applicable law, regulation, collective-bargaining obligations, due process, records requirements, procurement rules, civil-rights protections, privacy, classification, releasability, and mission policy. Counsel, privacy officials, security authorities, records officers, commanders, and program owners retain their assigned responsibilities.

An explainable prototype is not an authorization to operate. For DoD systems, DoD Instruction 8510.01 establishes the Risk Management Framework for DoD systems and the system-authorization decision structure. Grid makes no claim here of an ATO, classified-domain accreditation, cross-domain approval, weapons review, or compliance for a particular program.

Mission limits must also be explicit. A model validated for administrative triage is not thereby valid for benefits adjudication, intelligence analysis, targeting, logistics release, safety control, or use of force. Local execution does not prove indefinite operation through denied, disrupted, intermittent, and limited communications. A signed record can show origin and integrity without proving that its content is correct or releasable.

AI can propose at extraordinary speed. The control question is whether the institution can still identify the facts, rules, reviewer, authority, and outcome when the recommendation matters most. The goal is not a ceremonial human at the end of automation. It is a decision whose responsibility remains recognizable from beginning to end.

Related films, scenarios, and next steps

Choose the next move

Test the claim with a different kind of evidence.

For public-sector leadersApply the model to familiar workFollow a bounded use case from a changed fact through evidence, calculation, review, and authorized action.See the modelAI with Human AuthorityFive AI-written versions of the same requirement return different answers; the film moves the rule into one visible model and keeps approval with a person.For technical evaluatorsFollow the concept into Grid DevelopersContinue into the linked Grid Developers guide for the exact behavior, prerequisites, and limits used by this explanation.

Continue exploring

Follow the next question.