Article reader Listen + reading controls
Article reader
Preparing the reader…
Reading settings
The NIST AI RMF Is an Operating Model, Not a Checklist¶
The Federal Artificial Intelligence Risk Management Act of 2024 proposed requiring federal agencies to use the National Institute of Standards and Technology’s AI Risk Management Framework. The idea was sensible and bipartisan: federal AI should be governed through a common, credible risk-management structure rather than a patchwork of improvised agency practices.
The danger is equally familiar. Institutions can “adopt” a framework by mapping its language into policy, completing templates, and producing inventories while leaving the decisions that determine AI risk largely unchanged.
The NIST AI Risk Management Framework is most valuable when treated as an operating model: a way to connect mission context, evidence, authority, delivery, monitoring, and accountability across the lifecycle of a system.
H.R. 6936 would have directed agencies to incorporate the AI RMF into management of AI use and required OMB guidance, acquisition support, contract language, and oversight. Whether or not a particular bill advances, the implementation problem remains: what would an agency do differently on Monday morning if the framework genuinely governed its AI?
Start With Decisions, Not Documents¶
Risk management exists to improve decisions. An agency should identify the consequential decisions in an AI system’s lifecycle:
- whether the problem warrants AI;
- which mission outcome and affected populations define success;
- whether data may be used for the intended purpose;
- which model or service is acceptable;
- what evidence is required before testing with users;
- who may authorize limited and expanded deployment;
- when the system must be constrained, rolled back, or retired;
- who accepts residual risk;
- and how affected people obtain explanation, review, or remedy.
For each decision, the operating model should specify the decision owner, required evidence, consulted roles, escalation path, and record produced. The AI RMF’s Govern, Map, Measure, and Manage functions then become connected activities rather than four sections of a report.
Without decision rights, “governance” becomes a meeting. Without evidence thresholds, “measurement” becomes a dashboard. Without operational authority, “management” becomes a recommendation.
Map Is a System-Boundary Discipline¶
The Map function is sometimes reduced to a use-case description and risk classification. Its deeper purpose is to define the system that must be governed.
That system includes:
- the model and supporting software;
- data sources, transformations, and knowledge repositories;
- users and affected people;
- policies, incentives, and institutional workflows;
- vendor services and supply-chain dependencies;
- downstream decisions and actions;
- environmental and adversarial conditions;
- and other systems that consume the output.
If an agency maps only the model, it will miss the dominant risks. A benefits model may be statistically sound while the appeals process is inaccessible. A language model may produce acceptable answers in testing while retrieval data is stale or improperly disclosed. A decision-support system may preserve formal human authority while workload and interface design make approval automatic.
Mapping should also identify the baseline. AI risk cannot be judged against perfection. The relevant question is whether the proposed system improves the existing process, including its errors, delays, inequities, and hidden labor.
Measurement Needs Claims and Thresholds¶
Agencies can produce many metrics without establishing whether a system is fit for use. Measurement should be organized around explicit claims:
- The system performs a defined task well enough for a specified population and environment.
- Users understand material limitations and uncertainty.
- The workflow preserves meaningful review and remedy.
- Security controls protect the model, data, and delivery pipeline.
- The system detects conditions outside its evaluated operating envelope.
- Monitoring can reveal material degradation or disparate effects.
Each claim needs evidence, an acceptance threshold, and an owner authorized to determine whether the evidence is sufficient. Some evidence will be quantitative; some will come from legal review, qualitative research, red teaming, accessibility testing, or operational trials.
The Government Accountability Office’s AI accountability framework similarly organizes accountability around governance, data, performance, and monitoring. The practical opportunity is to make these bodies of evidence reusable across oversight, procurement, authorization, and product decisions rather than requiring separate compliance packages.
Risk Tiers Should Change the Delivery Path¶
Classifying an AI system as high risk creates little value unless the classification changes what happens next.
A risk-tier model can determine:
- independence and depth of evaluation;
- required human-factors and rights-impact assessment;
- security and adversarial testing;
- seniority of risk acceptance;
- scope and duration of pilots;
- monitoring frequency;
- incident notification timelines;
- public notice or documentation;
- and conditions for appeal or human review.
OMB’s M-24-10 memorandum later established minimum practices for rights-impacting and safety-impacting AI, including impact assessment, testing, ongoing monitoring, human oversight, and remedies in relevant contexts. That policy illustrates how a general framework can be translated into operational obligations.
The details will evolve. The invariant is that risk classification must alter authority and evidence, not merely labeling.
Governance Must Travel With the Product¶
Federal AI programs often separate policy and delivery. A central governance board reviews plans; a product or contractor team builds the system; an operations organization inherits it. Risk knowledge is lost at every transition.
The product should carry a maintained evidence package:
- intended and prohibited uses;
- decision owners and affected populations;
- data and model provenance;
- evaluation results and limitations;
- human-control design;
- security and privacy evidence;
- deployment constraints;
- monitoring indicators and thresholds;
- incident and rollback procedures;
- and the change history that affects prior claims.
This is not static documentation. It should be versioned with the system and updated through the delivery pipeline. A material model, data, workflow, or population change should trigger targeted re-evaluation.
Responsible AI becomes practical when evidence is produced as part of engineering rather than reconstructed for a periodic review.
Inventory Is Necessary but Not Sufficient¶
Agencies cannot govern AI they do not know they use. Inventories establish visibility, ownership, and public accountability. They also create false confidence if they count use cases without capturing dependencies and risk state.
A useful internal inventory should answer:
- Which operational service or decision depends on the system?
- Who owns the mission outcome and the technical product?
- Which model, vendor, and version are in use?
- Which data sources and downstream systems are connected?
- What is the current scope of authorized use?
- Which evidence is current, expiring, missing, or disputed?
- What incidents, overrides, appeals, and material changes have occurred?
The inventory then becomes a control plane for governance rather than a reporting exercise.
Incident Learning Completes the Framework¶
Predeployment evaluation cannot anticipate every interaction. Agencies need protected, nonpunitive mechanisms for reporting AI incidents and near misses, including unexpected behavior, inappropriate reliance, rights impacts, data leakage, security compromise, and failures of human review.
Incident learning should update:
- system constraints and monitoring;
- test scenarios and acceptance thresholds;
- training and procedures;
- procurement requirements;
- shared agency guidance;
- and, when appropriate, public understanding.
A framework that governs approval but not learning will become stale as systems and uses change.
The Strategic Inference¶
Mandating the NIST AI RMF can create useful consistency. It can also create a new layer of documentation that organizations learn to satisfy while continuing familiar behavior.
The difference is implementation. A real operating model makes decision rights explicit, defines the socio-technical system, ties claims to evidence, changes the delivery path based on risk, maintains evidence with the product, and learns from operation.
The measure of adoption is not whether an agency can show that every RMF category appears in a policy. It is whether leaders and teams make better, more traceable decisions because the framework is present.
This essay was substantially revised in July 2026 to replace the original legislative summary with an evidence-based implementation analysis. It incorporates federal guidance published after the original January 2024 post.
My work on trustworthy AI operations and Responsible AI Development Operations focuses on making this connection between governance evidence and delivery practice.
References¶
- U.S. House of Representatives, H.R. 6936, Federal Artificial Intelligence Risk Management Act of 2024, introduced January 10, 2024.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), January 2023.
- U.S. Government Accountability Office, Artificial Intelligence: An Accountability Framework for Federal Agencies and Other Entities, June 2021.
- Office of Management and Budget, M-24-10: Advancing Governance, Innovation, and Risk Management for Agency Use of Artificial Intelligence, March 28, 2024.
- “House Bill to Advance Federal AI Risk Management Policy Implementation,” ExecutiveGov, January 2024. This report was the historical prompt for the original post.