Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

Adversarial AI needs a shared language before it needs another tool

Security teams and artificial intelligence (AI) teams can look at the same system and see different attack surfaces. One sees identities, networks, software dependencies, and data flows. The other sees training distributions, model behavior, embeddings, prompts, and evaluation drift.

The National Institute of Standards and Technology (NIST) publishes its adversarial machine-learning taxonomy on March 24 to create a more consistent vocabulary for attacks and mitigations across predictive and generative systems.

The document covers evasion, poisoning, privacy, and misuse attacks, among others. Its immediate value is not a new product. It is the ability to name distinctions that determine who must act.

Ambiguous language creates unmanaged seams

Consider the word “prompt injection.” A developer may treat it as an application-filtering problem. A security engineer may see untrusted input crossing a control boundary. A model evaluator may focus on instruction-following behavior. A mission owner may care only that the system disclosed data or took the wrong action.

Each view is legitimate, but none is complete. If teams do not share a threat model, mitigations can be duplicated in one layer and absent in another.

Taxonomies help by creating categories that can be connected to architecture, tests, controls, and incident reporting. They also make disagreement explicit. A team can ask whether an event represents data poisoning, model evasion, tool misuse, or a conventional access-control failure amplified by AI.

Turn vocabulary into an operating practice

A glossary sitting in a policy library will not change system security. Teams should use the taxonomy in concrete artifacts:

  • threat-model templates;
  • architecture reviews;
  • test-case libraries;
  • security-control mappings;
  • supplier questionnaires;
  • incident classifications;
  • and post-deployment monitoring.

The goal is a traceable line from a named threat to a plausible harm, a system boundary, a mitigation, an owner, and evidence that the mitigation works.

This is also a knowledge-transfer problem. Bechky's research on occupational communities shows that specialists do not automatically share meaning even when they use similar artifacts. Shared language becomes useful through interaction—when teams work through a real case and learn how another profession interprets it.

Preserve uncertainty

A taxonomy can create a false sense that all attacks are neatly classifiable. Real incidents may span categories or reveal a threat the vocabulary does not yet capture. NIST describes the report as guidance that will be updated as the field evolves, which is the right posture.

Programs should keep an “unclassified” path and record where the taxonomy failed to fit. Those cases are signals for learning, not administrative defects to force into the nearest category.

Adversarial machine learning will continue to produce specialized tools. Before buying them, an organization should ask whether its security, data, model, software, and mission teams can describe the same threat in a way that leads to coordinated action.

The shared language is not the defense. It is the infrastructure that lets a defense become coherent.

Sources and research trail

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.