Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

Trustworthy decision support must be tested in the moment

The Defense Advanced Research Projects Agency (DARPA) has selected teams for its In the Moment program, which is exploring how machines might support difficult decisions when established rules are incomplete. Initial research focuses include mass-casualty triage and other settings where time, uncertainty, and competing values make judgment unusually demanding.

This is a serious test of human-centered artificial intelligence: not whether a model can produce an answer, but whether a human–machine team can make a better decision under pressure without obscuring who remains responsible.

Hard cases are not missing-rule problems

Many decision-support systems work by applying known criteria to available data. The hard cases DARPA describes are different. Information may be incomplete. Resources may be constrained. Several reasonable choices may conflict. The decision-maker may have seconds to act and little opportunity to explain.

Artificial intelligence (AI) can help organize evidence, compare patterns, or surface relevant experience. It can also create a false sense that the ambiguity has been resolved mathematically.

A recommendation does not make the value judgment disappear. It relocates parts of that judgment into objectives, training data, labels, thresholds, and interface design. The team must be able to say what the system optimized and whose expertise shaped that choice.

Trust should be calibrated to conditions

Lee and See's research on trust in automation argues for appropriate reliance, not maximum trust. Decision-makers need enough information to recognize when the system's competence matches the situation and when it does not.

That is difficult in time-sensitive work. A long explanation may arrive too late. A simple confidence score may conceal the reason for uncertainty. Requiring constant confirmation may overload the user.

The interface should present the few distinctions that change action. Those might include missing critical information, disagreement among models, an unfamiliar case, sensitivity to one assumption, or a recommendation outside the system's validated conditions.

The question is not “Can the system explain everything?” It is “What must this person know now to rely appropriately?”

Test teams, not only algorithms

Laboratory evaluation often separates model performance from human performance. Operational evaluation should examine the joint system.

Useful measures include:

  • decision quality under realistic time and information constraints;
  • cases in which the system changes a correct human judgment to an incorrect one;
  • cases in which the human detects a model failure;
  • time spent resolving disagreement;
  • effects on workload and situation awareness;
  • consistency across users with different experience; and
  • what the team learns after an outcome becomes known.

Endsley's theory of situation awareness is relevant because support that improves a local choice can still degrade the decision-maker's understanding of the larger situation. In critical work, that loss may surface later, after the system encounters something outside its script.

Preserve the reasoning trace without slowing the decision

The decision-maker may not have time to document a full rationale. The system can help preserve a trace automatically: available inputs, recommendation, uncertainty signals, human action, override, and later outcome. That record supports review and learning without turning the operational moment into paperwork.

Post-event review should include the people who performed the work. Their account may reveal cues the system did not capture or reasons an apparently irrational override was correct. Those lessons should update scenarios, training, interfaces, and evaluation—not remain in an after-action report.

DARPA's program is ambitious because it goes directly at decisions where neither rules nor data settle the matter. Success should not be defined as automating judgment. It should be defined as building support that improves perception and deliberation, makes uncertainty actionable, and leaves responsibility clear.

Trustworthy decision support will be proven neither by a benchmark nor by a reassuring explanation after the fact. It has to help in the moment—and help the organization learn once the moment has passed.

Sources and research trail

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.