White paper · AI governance
The Authority Contract for AI-Assisted Decisions
Where Proposals Stop and Organizational Action Begins
Human-in-the-loop language is too vague for consequential AI-assisted work. This paper proposes an explicit authority contract that defines what an AI system may observe, propose, validate, call, and never authorize—and how the organization tests that boundary.
Narrated film · 0:40
AI with Human Authority
Five AI-written versions of the same requirement return different answers; the film moves the rule into one visible model and keeps approval with a person.
▶ 0:40
Authority boundary
Fluent output cannot promote itself into authority.
The contract distinguishes assistance that software may perform from organizational powers that must be grounded outside the model and bound to an exact reviewed revision.
- AI-assisted componentObserve
Retrieve only permitted sources, fields, time windows, classifications, and quality states.
- AI-assisted componentGenerate
Produce candidate language, logic, or authored alternatives with sources and uncertainty.
- Governed modelEvaluate
Apply declared tests and calculations without expanding the task, evidence, or permission scope.
- AI-assisted componentRecommend
Present an exact proposal revision with reasons, unresolved gates, and prohibited actions.
- Accountable person or institutionAuthorize
Exercise a real legal, command, professional, fiscal, or organizational decision right.
- External effect ownerExecute
Cause an external effect only from the exact instruction, identity, scope, and approval that were authorized.
Jump to a section in this paper
“Human in the loop” does not say who can do what
An AI assistant drafts a recommendation. A person clicks approve. A connected tool creates a purchase order, changes a schedule, sends a notice, or updates a record. Later, the organization discovers that the reviewer saw only a summary, the cited source was stale, the approval applied to an earlier draft, or the tool call carried more authority than either the AI or reviewer possessed.
Everyone can still say a human was in the loop.
That phrase describes a topology, not an accountability model. It does not identify the decision owner, the AI's permitted task, the evidence a reviewer must see, the effect an approval may have, the actions that remain prohibited, or what happens when the model, prompt, source, role, or operating context changes.
An authority contract makes those boundaries explicit and testable. It states what an AI-assisted component may observe, transform, propose, validate, and call; what a responsible person or institution must decide; which external system performs an effect; and which receipts prove only that a declared step occurred.
This paper separates public AI risk-management guidance from the Grid proposition. NIST and GAO provide frameworks and practices for organizations to consider. They do not endorse Grid, confer decision authority, certify an implementation, or establish that any human-AI arrangement is safe, fair, reliable, or compliant.
What the public frameworks establish
The NIST AI Risk Management Framework 1.0 is voluntary, rights-preserving, non-sector-specific, and use-case agnostic. Its core functions—Govern, Map, Measure, and Manage—treat governance as continuous across the AI lifecycle. The framework calls for roles and responsibilities in human-AI configurations to be defined and differentiated, and for context, risk, performance, and monitoring to be considered. NIST's current AI RMF page also states that version 1.0 is being revised; this paper relies on the published 1.0 document and does not assume the content of a future revision.
NIST's 2024 Generative AI Profile, NIST AI 600-1, is a cross-sector companion to AI RMF 1.0. It identifies risks including confabulation, information integrity, privacy, security, value-chain dependencies, and human-AI configuration. Its suggested actions include acceptable-use and refusal policies, content-lineage practices, context-relevant testing, review of sources and citations, sharing pre-deployment test results with release authorities, and monitoring overrides and feedback. The profile also warns against extrapolating capabilities from narrow or anecdotal tests. These are suggested risk-management actions, not a mandatory control catalog or assurance certificate.
The GAO AI Accountability Framework organizes practices around governance, data, performance, and monitoring. It asks organizations to define goals, roles, data quality, metrics, human supervision, monitoring, and conditions for expanded use. GAO says the appropriate degree of supervision depends on purpose and potential consequence. The framework supports auditors and accountable managers; it does not make a generic approval click adequate for any particular use.
For covered federal-agency uses, OMB Memorandum M-25-21 requires minimum risk-management practices for high-impact AI, directs agencies to establish clear expectations for consequential AI-assisted work, and states that officials retain authorities established in other law and policy. The memorandum is not a general authorization for any system or decision. Its scope, exclusions, agency implementation, and any successor policy must be evaluated for the actual use.
Defense and coalition institutions add mission-specific responsible-use frames. The DoD Chief Digital and Artificial Intelligence Office's Generative AI Responsible AI Toolkit provides project teams with questions and tools, including a suitability, feasibility, and advisability assessment; it is not an operational approval or certification. NATO's revised AI strategy describes principles including lawfulness, responsibility and accountability, explainability and traceability, reliability, governability, and bias mitigation, alongside testing and interoperability goals. Those principles still have to be translated into the roles, evidence, systems, and operational conditions of a particular capability.
Together, these sources support a bounded conclusion: organizations should define context, responsibilities, data, testing, performance, human-AI arrangements, monitoring, and change. They do not determine who holds authority under a contract, statute, license, professional duty, corporate delegation, incident plan, or local policy. That must be established outside the AI system.
The contract has eight parts
The Grid proposition represents the authority boundary as a versioned object connected to the model and workflow it governs.
1. Purpose and decision class
Name the organizational objective, intended users, affected parties, consequence level, jurisdiction, operating context, and exact decision class. “Support operations” is too broad. “Draft a capacity-restoration alternative for review by the continuity lead” is testable.
2. Permitted inputs
Declare which sources, fields, time windows, classifications, permissions, and quality states the AI may use. Identify protected or prohibited data, required provenance, freshness rules, conflicts, and the behavior for missing evidence. Retrieval does not make a source accurate, current, licensed, or appropriate.
3. Permitted outputs
Define whether the component may summarize, classify, extract, calculate, explain, draft, rank predefined alternatives, propose a model change, or ask for evidence. Require uncertainty, source references, assumptions, unresolved issues, and the intended recipient where appropriate. A proposal should never inherit the visual status of an approved decision.
4. Tool and effect scope
List every callable tool and the narrowest permitted operation: read, simulate, prepare, stage, submit for review, or execute. Bind identities, credentials, environments, rate and amount limits, destinations, expiry, and denied actions. An AI allowed to draft a transaction should not therefore be allowed to release funds, send a public message, alter a production record, or waive a control.
5. Decision rights
Name the role—not merely a user account—that may accept, reject, modify, defer, or escalate the proposal. Connect that role to the organization's actual delegation. Separate subject-matter review, risk acceptance, legal or policy review, financial approval, release approval, and execution when the process requires them. Grid does not create any of those authorities.
6. Review evidence
Specify what the reviewer must receive: exact proposal revision, material sources, conflicts, calculations, uncertainty, tests, changed facts, affected parties, prohibited alternatives, prior decisions, and model limitations. Record the time spent or sequence followed only if it is a meaningful acceptance criterion; presence on a screen does not prove comprehension or independent judgment.
7. Change, expiry, and revocation
State which changes invalidate approval: source revision, output text, amount, destination, model or prompt version, tool adapter, policy, role, risk tier, or elapsed time. Define expiration, emergency suspension, credential revocation, rollback limits, and treatment of actions already executed. Reverting a model does not reverse a payment, message, personnel action, or physical event.
8. Receipts and monitoring
Retain prompts or structured requests as policy allows, retrieved sources, model and configuration identity, proposal, tests, reviewer actions, approvals, tool calls, responses, exceptions, overrides, incidents, and corrections. A tool receipt shows what that tool reported. It does not prove the person understood, the downstream effect occurred, the decision was lawful, or the outcome was beneficial.
Five verbs that should never collapse
The authority contract distinguishes five verbs:
- Generate: produce candidate language, logic, or alternatives.
- Evaluate: apply declared tests to a candidate or known fixture.
- Recommend: present an option with reasons and limits.
- Authorize: exercise an externally established organizational decision right.
- Execute: cause an effect in another system or the world.
Grid may support the first three. It never supplies authorization, although it can record an externally grounded authorization and the later execution record. When separately configured and permitted, Grid may transmit the exact already-authorized instruction; that places it in the execution path but does not confer a decision right. Host-admitted network, connector, or secret capability is technical permission, not organizational authorization. A configured request-and-response record can show what the integration reported, while external acceptance and any downstream effect remain separate and unproven. The software cannot promote itself from recommendation to authority.
This distinction also limits “agentic” behavior. Chaining more tool calls does not enlarge the contract. A planning agent cannot gain procurement authority because it found a purchasing API. A model cannot use a reviewer credential as evidence that the reviewer approved a changed proposal. A fallback model cannot silently inherit the permissions or accepted performance boundary of the primary model.
| State | Principal actor | Required artifact | What this state cannot establish |
|---|---|---|---|
| Generate | AI-assisted component | Attributed candidate with sources, assumptions, and uncertainty | Correctness, completeness, or permission to proceed |
| Evaluate | Governed model and qualified tester | Declared fixture, expected result, execution receipt, and defects | Performance outside the tested cases |
| Recommend | AI-assisted component or staff author | Exact proposal revision, reasons, unresolved gates, and prohibited actions | Approval, obligation, release, or command decision |
| Authorize | Accountable person or institution | Current delegation, evidence reviewed, scope, conditions, and expiry | That an external system accepted or completed the action |
| Execute | Separately permitted effect owner | Exact authorized instruction plus submission and response receipts | Intended real-world effect or beneficial outcome |
Three different authorities are often collapsed even after those states are separated. System authorization addresses whether a system may operate in an environment under its applicable security process. Information-release authority addresses whether particular data or output may cross a boundary or reach a recipient. Operational or program decision authority addresses whether a person may commit the organization to the actual course, determination, expenditure, or action. Evidence for one does not substitute for the others.
Fictional exercise: Meridian continuity exception
Consider a wholly fictional cross-industry fixture. Meridian Services needs 600 abstract processing units for a declared four-hour window. Its accepted primary source provides 420 units, and an approved backup provides 120, leaving a represented shortfall of 60. These are exercise counters, not real capacity, customers, money, or service commitments.
An AI assistant may retrieve the three current source records, verify 420 + 120 = 540, identify the 60-unit gap, and draft predefined alternatives. It proposes 80 units from a reserve provider, which would create a 20-unit modeled margin. The provider record, however, is candidate, not approved; security review, commercial terms, operational acceptance, and spending authority are missing.
The authority contract permits the AI to prepare a comparison and a draft activation package. It forbids changing the provider status, accepting terms, obligating funds, revising a customer commitment, or calling the activation endpoint. The result remains conditional and names four unresolved gates.
The exercise then supplies security, commercial, and operational reviews, but the fictional institutional delegation record still lacks a finance approver. Under the fixture oracle, release is expected to remain blocked even if a reviewer clicks approve. Next, the organization's designated delegating authority records a current finance delegation, the named finance approver approves the stated spending limit, and all other evidence remains current. The continuity decision owner then accepts proposal revision MC-07. A separate release operator—not the AI—submits the exact authorized revision to a simulated adapter. The adapter acknowledges the request. These are expected fixture dispositions until tested in a configured implementation.
That acknowledgment proves only simulated acceptance. It does not prove reserve capacity existed, service was restored, a contract was valid, funds were available, security was adequate, customers were served, or the decision produced a good outcome.
Next, the source changes from 600 required units to 630. The old approval no longer applies because the contract says a changed demand fact invalidates the proposal. The model recalculates a 90-unit shortfall and shows that the prior 80-unit alternative is insufficient. It does not stretch the earlier approval or silently select another provider.
Evaluate the boundary by trying to break it
Begin with one narrow decision class, synthetic or appropriately governed historical cases, and no live effect. Record the current baseline: untracked AI use, manual rekeys, review time, unsupported citations, approval ambiguity, tool permissions, overrides, incidents, and effort required to reconstruct one decision.
Predefine expected behavior with domain, decision, risk, privacy, security, legal, operations, and technical owners. Seed a confabulated source, stale retrieval, prompt injection, prohibited field, changed amount after approval, wrong role, expired delegation, excessive tool scope, duplicate execution, adapter timeout, fallback model, model update, hidden uncertainty, unsupported recommendation, and reviewer override.
Measure at least:
- Scope fidelity: permitted tasks succeed and prohibited tasks fail under the declared contract.
- Evidence integrity: material sources, revisions, conflicts, calculations, and uncertainty reach the reviewer; invented or stale support is surfaced.
- Approval binding: authorization attaches to the exact proposal, evidence, role, model context, amount, destination, and time required by the fixture.
- Effect isolation: no seeded live or simulated effect occurs without the complete authority set; retries and duplicates behave as declared.
- Human-AI performance: reviewers detect seeded defects under realistic load; overrides, deference, disagreement, and escalation are measured rather than assumed away.
- Change and recovery: material changes invalidate approval as designed; suspension, revocation, fallback, incident reconstruction, and correction are exercised.
Retain the contract, sources, fixtures, expected results, model and prompt configuration, tool scopes, evaluation environment, red-team cases, reviewer evidence, decisions, receipts, overrides, defects, and accepted limitations.
A passing fixture supports at most a product-demonstration claim for its exact configuration and cases. It does not establish model validation, safe or responsible AI, absence of bias, legal or regulatory compliance, security authorization, professional adequacy, deployment readiness, human comprehension, cost or labor savings, decision quality, or organizational outcome. Customer-controlled and independently validated claims require their own disclosed methods and boundaries.
The AI authority control model applies this pattern to high-consequence government work, while AI with Human Authority provides a shorter use-case path. The public-policy paper traces the contract through eligibility, exceptions, notice, and review; the course-of-action paper traces it through recommendation, operational authority, submission, and effect. Teams can define acceptance evidence with How to Evaluate Executable Decision Infrastructure and the AI decision-authority assessment. Implementation teams can continue with Grid Developers' explain-and-validate guidance.
The purpose of an authority contract is not to make AI look cautious. It is to make organizational responsibility executable enough to test: the proposal stops where declared authority begins, and neither fluent output nor connected tools can move that boundary by implication.
Related films, scenarios, and next steps
Continue exploring
Follow the next question.
How Grid Works: From Typed Facts to Accountable Action
A technical and operational guide to the six-stage Grid chain: governed sources, a shared and versioned model, reactive constraint evaluation, coordinated surfaces, explanation and evidence, and AI assistance bounded by human authority.
Understand · EvaluateSource-grounded ExploreMake the Structure Visible
Assumptions, Evidence, and Uncertainty in Research and LearningResearch and learning suffer when a polished result hides its sources, assumptions, exclusions, uncertainty, and revisions. This paper proposes an inspectable model that helps people test the structure of an argument without transferring inquiry, authorship, review, or instructional authority to software.
Understand · EvaluateSource-grounded ExploreHow to Evaluate Executable Decision Infrastructure
A buyer's guide for turning a compelling demonstration into a bounded evaluation of sources, logic, explanations, authority, interoperability, change behavior, and evidence.
Evaluate · PlanSource-grounded Explore