Article reader Listen + reading controls
Article reader
Preparing the reader…
Reading settings
The best federal AI use case may not need AI¶
A good artificial intelligence portfolio begins with permission to say that artificial intelligence (AI) is not the answer.
That sounds obvious. In practice, organizations often begin in the opposite place. A new model becomes available, leaders announce an adoption goal, and teams are asked to find use cases. The search produces a familiar list—summarization, forecasting, chatbots, anomaly detection, document review—before anyone has defined the mission problem, the current baseline, or the decision the system is supposed to improve.
The result may be technically interesting. It is not yet a strategy.
A federal AI use case should be written as a testable claim: for a defined group of people, performing a specific mission task under known conditions, this capability will improve a measurable outcome enough to justify its cost and risk. If the claim cannot be stated, challenged, and evaluated, the agency does not have a use case. It has a technology theme.
Revised and substantially expanded July 17, 2026, to turn the original short commentary into a practical method for selecting and governing federal AI investments.
Begin where the mission is struggling¶
Mission owners usually know where work is difficult. They see decisions delayed by fragmented information, skilled employees trapped in repetitive review, citizens asked to submit the same facts repeatedly, analysts searching across incompatible systems, and operational teams working around software that no longer fits the job.
Those conditions are better starting points than a catalog of model capabilities.
The first discovery conversation should ask:
- Which outcome are we consistently failing to achieve?
- Where does work wait, repeat, or lose context?
- Which decision depends on more information than a person can reasonably assemble in time?
- Where is scarce expertise being consumed by low-judgment activity?
- Which errors create the greatest cost, delay, safety risk, or public harm?
- What do users already do to compensate for the existing system?
This keeps the problem in operational language long enough for program staff, technologists, acquisition professionals, security teams, and users to build a shared understanding. Once a model enters the conversation, it tends to organize the problem around what the model can do. The order matters.
Establish the baseline before promising improvement¶
“Faster,” “more accurate,” and “more efficient” are not outcomes until they are compared with something.
Before building, measure the current process. How long does the work take from request to completed outcome? Where does queue time accumulate? How often is work returned or corrected? Which groups experience different results? What does the current process cost, including the labor hidden in manual workarounds? Which failures are visible, and which disappear into rework?
The baseline often changes the proposed solution. A team may discover that the largest delay occurs in approval, not analysis. The missing capability may be data access rather than prediction. A rules-based service may handle most cases more transparently and cheaply. Better search, process redesign, or an application programming interface may solve the mission problem without an AI model.
That is not a failed AI initiative. It is successful technical judgment.
Define the decision, not just the output¶
Models generate scores, classifications, recommendations, text, images, code, or ranked options. Missions improve only when those outputs change a decision or action.
A use case therefore needs to specify:
- The decision or task. What will a person or system do differently?
- The user. Who must understand, trust, contest, or act on the output?
- The operating context. Which data, time pressure, environment, workload, and exceptions define real use?
- The authority boundary. What may the system recommend or execute, and what remains a human decision?
- The failure consequence. What happens when the output is wrong, late, unavailable, insecure, or persuasive for the wrong reason?
- The alternative. What other technical or process options could improve the same outcome?
This is where apparently similar applications separate. A language model that drafts an internal meeting summary is not the same use case as one that summarizes a benefits record for an eligibility decision. The interface may look similar; the affected people, evidence requirements, failure modes, and review obligations are not.
The National Institute of Standards and Technology (NIST) makes context central to its AI Risk Management Framework. The “Map” function asks organizations to understand intended purpose, affected parties, context, and potential impacts before selecting and managing risk. That is also good product strategy. A system cannot be evaluated apart from the work it changes.
Choose the smallest credible intervention¶
Once the mission problem and decision are clear, teams can compare technical patterns.
A stable process with explicit rules may need conventional software or automation. A well-defined classification problem with representative historical data may suit traditional machine learning. Work involving large collections of language, images, or code may justify generative models or semantic retrieval. A complex workflow may require several services coordinated through ordinary software rather than an autonomous agent.
The preferred option is not the most advanced technology. It is the least complex intervention capable of producing the required improvement while remaining secure, supportable, and understandable in context.
Complexity creates obligations. A model introduces data provenance, evaluation, monitoring, change control, vendor dependency, and failure modes that ordinary software may not. An agent adds tool permissions, multi-step behavior, state, and a larger space of possible actions. Those costs may be justified, but they belong in the investment decision from the beginning.
Turn the use case into a falsifiable investment thesis¶
A useful one-page case should state:
For these users, performing this task under these conditions, the proposed capability will improve these mission measures from this baseline to this target, while keeping these risks within these limits. We will expand, redesign, or stop the work based on this evidence by this date.
That format forces precision without pretending that every uncertainty is already resolved. It also creates a fair basis for comparison across a portfolio.
The evidence plan should include more than model accuracy. It may need task completion time, downstream rework, user adoption, subgroup performance, override behavior, incident rates, accessibility, security, total operating cost, and effects on the people receiving the service. Measures should be collected in representative conditions, not only in a curated demonstration.
The U.S. Government Accountability Office (GAO) found that federal inventories included incomplete and inaccurate information about AI use cases, including entries that agencies later determined were not AI. Its 2023 government-wide review reinforces the need for common definitions and lifecycle evidence. Better selection begins with a better description of the work.
Design adoption before the pilot¶
Many promising demonstrations fail because the model was treated as the product. Operational capability also requires:
- authoritative and legally usable data;
- integration with systems where work already occurs;
- identity, security, records, privacy, and accessibility controls;
- an acquisition path with adequate data and observability rights;
- trained users and a redesigned workflow;
- monitoring, support, incident response, and change ownership;
- funding beyond the experiment; and
- an accountable mission owner who can accept, constrain, or stop the use.
If those conditions are impossible, the problem may not be ready for AI even when a prototype performs well.
Current Office of Management and Budget (OMB) guidance in M-25-21 encourages federal adoption while requiring additional practices for high-impact uses. Agencies should not interpret faster adoption as permission to defer operating design. Moving quickly means discovering the real constraints early enough to make a decision—not moving a demo into production before ownership and evidence exist.
Manage the portfolio by learning rate¶
No selection method will make every investment succeed. The purpose is to make uncertainty explicit and learning economical.
Fund early exploration in small increments. Require representative tests before major scale. Share evaluation methods and infrastructure across teams. Preserve negative results so the next program does not repeat them. Make stopping a use case an ordinary portfolio outcome.
Leaders should ask which assumptions were tested, which evidence changed the decision, and which reusable capability remains. A stopped pilot can still improve the organization if it produces a better dataset, security pattern, acquisition clause, evaluation harness, or understanding of the mission.
The opposite is also true. A pilot declared successful because the model produced an impressive output may leave no durable capability at all.
Ask the question in the right order¶
“What can we do with AI?” is useful during technical exploration. It is a poor opening question for public investment.
Start with the mission friction. Measure the current state. Define the decision and the people affected. Compare alternatives. State the improvement as a falsifiable claim. Design the operating conditions before celebrating the demonstration.
Sometimes that process will lead to a sophisticated AI system. Sometimes it will lead to better data, simpler software, a clearer policy, or a redesigned handoff. The federal government should want both outcomes. The objective is not to maximize the number of AI systems. It is to improve the mission with evidence, discipline, and technology appropriate to the work.
Sources¶
- National Institute of Standards and Technology, AI Risk Management Framework Core.
- U.S. Government Accountability Office, Artificial Intelligence: Agencies Have Begun Implementation but Need to Complete Key Requirements (December 12, 2023).
- Office of Management and Budget, M-25-21: Accelerating Federal Use of AI through Innovation, Governance, and Public Trust (April 3, 2025).
- U.S. Government Accountability Office, Artificial Intelligence Acquisitions: Agencies Should Collect and Apply Lessons Learned to Improve Future Procurements (2026).