Article reader Listen + reading controls
Article reader
Preparing the reader…
Reading settings
Human–AI Teaming Requires a Science of the Team¶
The phrase human–machine teaming is used so frequently that it can obscure how little it explains. Placing an AI system in a workflow with a human does not create a team. Nor does assigning the human final authority guarantee meaningful control. A team exists only when participants have interdependent roles, exchange information, adapt to one another, and coordinate their actions toward a shared outcome.
That is why DARPA’s work on quantitative models of human–AI teams is more important than the original 2024 solicitation cycle that prompted this essay. The enduring research problem is not simply whether an AI model performs well or whether a person approves its output. It is whether the combined system behaves competently, legibly, and safely under realistic conditions—including conditions neither participant encountered during development.
DARPA’s Exploratory Models of Human-AI Teams (EMHAT) program identifies a basic gap: methods developed to study human teams have not been adequately extended to human–AI teams, while generative models introduce a wide range of adaptive behaviors that existing evaluation paradigms do not capture. EMHAT proposes to use generative AI, expert feedback, knowledge bases, and simulated “digital twins” of human teammates to evaluate task completion and AI adaptation in proxy operational settings.
The key conceptual move is to treat the team—not the model—as the unit of performance.
Model Performance Does Not Predict Team Performance¶
An AI system can improve on a benchmark while degrading the work of the human–AI team. Several mechanisms make this possible.
Automation bias¶
Users may accept a confident recommendation without sufficiently examining its basis. A model that is correct most of the time can make its rare errors more dangerous if reliability encourages uncritical use.
Algorithm aversion¶
After observing a visible error, users may reject useful assistance even when the system remains better than the available alternative. Trust can collapse faster than statistical performance changes.
Out-of-the-loop degradation¶
As automation assumes more of a task, the person may lose situation awareness and the skill required to intervene. “Human on the loop” becomes nominal if the human cannot reconstruct the state before the decision window closes.
Cognitive displacement¶
The system may reduce one kind of workload while creating another: verifying sources, resolving contradictions, interpreting uncertainty, or monitoring an opaque agent’s actions. A faster interface can conceal a more expensive mental process.
Adaptive interaction¶
People change how they work in response to the system, and the system may change in response to people. Initial performance estimates can become invalid because the joint process is not stationary.
These effects are why a high model-accuracy score is insufficient evidence for a high-stakes deployment. The NIST AI Risk Management Framework emphasizes that AI risks emerge from socio-technical conditions and human behavior, not only from the computational system. Human–AI teaming is the concrete operational expression of that principle.
Role Design Comes Before Interface Design¶
Many teaming efforts begin by deciding how a model’s output should appear on a screen. The prior question is what work should be allocated to each participant and why.
A useful role design distinguishes at least five functions:
- Framing: defining the problem, objective, constraints, and relevant context;
- Generation: producing hypotheses, plans, classifications, or recommended actions;
- Evaluation: testing outputs against evidence, rules, models, and alternatives;
- Decision: selecting or authorizing a course of action;
- Execution and monitoring: acting, observing consequences, and adjusting.
Humans and machines may contribute to every function, but not symmetrically. Machines can search large spaces and maintain consistent procedural checks. Humans can interpret ambiguous intent, recognize when the frame itself is wrong, weigh values that have not been reduced to a metric, and accept institutional or moral responsibility. Those generalizations are not fixed laws; they are hypotheses to be tested for a particular mission.
The most dangerous design is an ambiguous allocation in which the machine effectively frames and recommends the action while the human formally “decides” by approving it. Authority remains with the person on paper, but agency has migrated to the system.
Meaningful Human Control Is a Property of the Workflow¶
Meaningful control cannot be created by adding an approval button. It depends on whether the person has:
- sufficient time to intervene;
- access to relevant evidence and uncertainty;
- a mental model of what the system can and cannot do;
- practical alternatives to accepting its recommendation;
- authority to stop, modify, or escalate the action;
- preserved skills and situation awareness;
- and feedback about the consequences of prior decisions.
This leads to a more rigorous design question: At which points can human judgment materially change the trajectory of the system?
If a model generates hundreds of actions at machine speed, nominal review may be impossible. The design may need to shift human involvement upstream—to constraint-setting, policy, scenario approval, exception definitions, and evaluation—and downstream to monitoring and incident response. Conversely, where context is sparse and consequences are severe, the system may need to slow the interaction and elicit explicit reasoning.
DARPA’s Friction for Accountability in Conversational Transactions (FACT) program makes this insight unusually concrete. Instead of treating a frictionless AI experience as an unconditional good, it explores when dialogue should reveal assumptions, consequences, and alternative paths so that users engage in reflective reasoning. In high-stakes work, well-designed friction can be a safety control.
Trust Should Be Calibrated, Not Maximized¶
Teams perform poorly when people trust AI too much and when they trust it too little. The objective is calibrated reliance: use the system when its competence and the conditions justify doing so; question or reject it when they do not.
Calibration requires more than an aggregate confidence score. Users need information tied to the decision:
- whether the current input resembles evaluated conditions;
- which sources and transformations support the result;
- which assumptions materially affect it;
- how alternative models or analysts disagree;
- which failure modes are known;
- and what signals should trigger escalation.
The interface should not attempt to explain every internal computation. It should expose the evidence and uncertainty needed for the user’s role. Explanations are useful only when they improve a real decision; otherwise they can become persuasive decoration.
Trust is also institutional. A user’s willingness to rely on a system depends on who approved it, how incidents are handled, whether limitations are communicated honestly, and whether operators are punished for declining unsafe automation. Governance shapes team behavior.
Evaluation Must Include Adaptation and Stress¶
A human–AI team should be tested as a dynamic system across time, not as a single user performing a scripted task. An evaluation program should include:
Baseline comparison¶
Compare the team with credible alternatives: the human-only workflow, other tools, and other allocations of work. “The team completed the task” is not evidence of improvement.
Distribution of people¶
Test varied expertise, cognitive styles, language backgrounds, risk preferences, accessibility needs, and levels of familiarity with the system. A design that works for expert evaluators may fail for its actual workforce.
Distribution of conditions¶
Include missing and corrupted data, degraded communications, time pressure, adversarial inputs, ambiguous goals, changing rules, and unfamiliar scenarios.
Longitudinal behavior¶
Measure learning, complacency, skill retention, workarounds, and trust over repeated use. Novelty effects and evaluator attention can make early trials unrepresentative.
Team-level measures¶
Assess mission effectiveness, error recovery, coordination cost, situation awareness, decision latency, appropriate reliance, and resilience—not only model accuracy and user satisfaction.
Consequence-sensitive evaluation¶
Weight errors by operational consequence. A system that improves average performance may still be unacceptable if it concentrates errors in rare, catastrophic situations.
DARPA’s EMHAT use of simulated human teammates is promising because simulation can explore combinations that are expensive, dangerous, or rare in live trials. But simulation introduces its own model risk. A digital twin of a human is not a human; it reflects assumptions about cognition and behavior. The validity of the simulation, the diversity of its agent population, and the transfer of findings to real teams must themselves become objects of evidence.
Knowledge Infrastructure Is Part of the Team¶
Human–AI teaming is often depicted as a dyad: one person and one model. In enterprise and defense settings, a third participant is always present—the knowledge environment.
The system’s outputs depend on doctrine, data, prior decisions, system documentation, operational context, and institutional memory. The human’s interpretation depends on much of the same material. If that knowledge is fragmented, stale, semantically inconsistent, or stripped of provenance, neither teammate can perform reliably.
This makes knowledge engineering a core part of teaming:
- authoritative sources must be identifiable;
- concepts and relationships must be represented consistently;
- provenance and temporal validity must be preserved;
- uncertainty and disagreement must remain visible;
- access controls must reflect mission and policy;
- and feedback from decisions must improve the knowledge base.
Retrieval-augmented generation, knowledge graphs, vector indexes, and semantic layers can help, but only when they are governed as operational infrastructure rather than assembled as a one-time context package. The quality of the team is bounded by the quality of the shared world it can perceive.
From Human-in-the-Loop to Accountable Team Design¶
“Human in the loop” is too coarse a requirement for serious systems. It says where a person appears in a diagram but not whether that person can understand, influence, or take responsibility for the outcome.
A better assurance case documents:
- the mission objective and decision context;
- the allocation of functions between people and machines;
- the conditions under which that allocation changes;
- the evidence each participant receives;
- the time and authority available for intervention;
- the expected failure and recovery behaviors;
- the metrics by which the combined team is judged;
- and the organizational owner accountable for monitoring the system in use.
The Department’s Responsible AI Strategy and Implementation Pathway and the CDAO’s Responsible AI Toolkit provide lifecycle structures within which that evidence can be governed. The practical task is to connect those structures to mission-specific team behavior rather than treat responsible AI as a parallel compliance activity.
The Strategic Inference¶
The frontier of human–AI teaming is not a more conversational interface or a more autonomous agent. It is the ability to engineer and evaluate adaptive joint systems whose competence cannot be inferred from either participant alone.
That requires a science of role allocation, trust calibration, decision rights, knowledge infrastructure, team dynamics, and longitudinal evaluation. It also requires a different standard of leadership: leaders remain accountable not only for what an AI system does, but for the conditions under which people come to rely on it.
DARPA’s research agenda is valuable because it treats these questions as measurable engineering problems without pretending they are only engineering problems. Human values, institutional authority, and mission consequences remain inside the system boundary.
That is where they belong.
This essay was substantially revised in July 2026 to replace the original solicitation summary with an enduring, evidence-based analysis. It incorporates program information published after the original January 2024 post.
My research and practice explore this boundary among trustworthy AI, knowledge infrastructure, and transformation systems. I welcome practitioners working on accountable human–AI collaboration to connect with me on LinkedIn.
References¶
- Defense Advanced Research Projects Agency, “Exploratory Models of Human-AI Teams (EMHAT).”
- Defense Advanced Research Projects Agency, “Friction for Accountability in Conversational Transactions (FACT).”
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), January 2023.
- U.S. Department of Defense, Responsible Artificial Intelligence Strategy and Implementation Pathway, June 2022.
- U.S. Department of Defense, “CDAO Releases Responsible AI Toolkit for Ensuring Alignment With RAI Best Practices,” November 14, 2023.
- “DARPA Issues RFI for Human-Machine Teaming Research Opportunity,” ExecutiveGov, January 2024. This report was the historical prompt for the original post.