Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

AI assurance has to survive contact with the mission

The first consequential defense artificial intelligence story of 2025 does not arrive as a new model or a weapons demonstration. It arrives as an invitation to test.

The Department of Defense's (DoD) Chief Digital and Artificial Intelligence Office (CDAO) begins January with a crowdsourced assurance pilot in military medicine. The setting matters. A medical system can perform impressively on average and still fail a clinician or patient at exactly the wrong moment. Its quality cannot be separated from the people, workflow, uncertainty, and consequences around it.

Assurance is evidence, not confidence

Artificial intelligence (AI) assurance is sometimes treated as a final checkpoint: run a benchmark, complete a review, sign an approval, and move the system forward. That approach mistakes an artifact for an argument.

A defensible assurance case connects claims about a system to evidence about its behavior under relevant conditions. The National Institute of Standards and Technology (NIST) makes the same point structurally in its AI Risk Management Framework: risk management spans governance, context mapping, measurement, and ongoing management. A score can contribute evidence. It cannot establish that the people using a system will understand it, challenge it, or recover from its failure.

Military medicine makes those missing conditions visible. Data can shift across facilities and populations. Time pressure changes how recommendations are interpreted. Rank, specialty, and workload influence whether someone questions an output. A false positive may consume scarce attention; a false negative may conceal a serious condition. The relevant unit of analysis is therefore not the model alone. It is the human-machine work system.

Why crowdsourcing is promising—and limited

Opening an evaluation to a broader community can diversify the failure modes that emerge. Different clinicians, technologists, security specialists, and operators bring different mental models. That variety is valuable because no single test team can anticipate every way people may use or misunderstand a system.

But the number of testers is not the same as the quality of the test. A challenge needs representative scenarios, protected data, clear reporting rules, and a path from discovery to remediation. It should distinguish a clever model failure from a failure that could plausibly alter a real decision. It also needs to preserve context: what the tester is trying to do, what information is available, why the output seems credible, and what happens next.

Human-factors research has warned for decades that automation can create new supervisory burdens even as it removes manual work. Parasuraman, Sheridan, and Wickens' levels-of-automation framework shows why the allocation of functions between people and machines must be designed deliberately. The question is not simply whether an AI system can perform a task. It is which parts of sensing, analysis, decision, and action it should perform—and how a person can intervene.

Build the learning loop now

Program teams do not need to wait for a department-wide assurance regime. They can create a practical loop around each use case:

  1. State the decision the system is intended to support.
  2. Identify who can be helped, delayed, misled, or excluded.
  3. Test realistic operating conditions, including degraded data and time pressure.
  4. Capture disagreement and near misses, not only pass-or-fail results.
  5. Assign owners and deadlines for remediation.
  6. Retest after model, data, interface, or workflow changes.

The pilot's most important contribution may be cultural. It treats assurance as a participatory practice rather than a specialist's certificate. That is the right direction for national-security AI. Trustworthy systems will not emerge because a model looked good before fielding. They will emerge because organizations become good at finding weakness, learning from it, and changing the complete system before the mission pays the price.

Sources and research trail

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.