Skip to content

AI Security

Adversarial AI needs a shared language before it needs another tool

Security teams and artificial intelligence (AI) teams can look at the same system and see different attack surfaces. One sees identities, networks, software dependencies, and data flows. The other sees training distributions, model behavior, embeddings, prompts, and evaluation drift.

The National Institute of Standards and Technology (NIST) publishes its adversarial machine-learning taxonomy on March 24 to create a more consistent vocabulary for attacks and mitigations across predictive and generative systems.

AI in Zero Trust: Automate Evidence, Not Accountability

Originally published in 2024; substantially revised in 2026 to deepen the analysis and incorporate additional sources.

Zero trust creates an appealing environment for artificial intelligence. Every access request can generate context: identity, device posture, workload state, data sensitivity, behavior, location, threat intelligence, and prior activity. AI and machine learning can help correlate those signals faster than human analysts can review them individually.

But there is a design trap. If the organization uses AI to make opaque access decisions inside an architecture intended to improve security visibility and control, it can reproduce the very problem zero trust was meant to solve.

AI should automate the collection, interpretation, and routing of security evidence—not dissolve accountability into an inscrutable risk score.

AI Data Poisoning Is a Knowledge-Supply-Chain Problem

Originally published in 2024; substantially revised in 2026 to deepen the analysis and incorporate additional sources.

NIST’s warning about adversarial manipulation of AI systems should change how organizations define the security boundary. Conventional software security focuses heavily on code, dependencies, infrastructure, identity, and configuration. AI systems add another attack surface: the evidence from which system behavior emerges.

Training corpora, feedback data, retrieval indexes, evaluation sets, model artifacts, prompts, and operational context all influence what an AI system learns or produces. If an adversary can shape those inputs, the system may remain technically available while becoming epistemically compromised.

AI expands the software supply chain into a knowledge supply chain. Security must protect not only what the system executes, but what it is permitted to believe.

Scaling policies turn capability into a management trigger

Anthropic has published a Responsible Scaling Policy (RSP) that ties increasingly strong safety and security measures to evidence that a model has reached particular dangerous capabilities. The policy introduces Artificial Intelligence Safety Levels (ASLs), loosely inspired by the graduated containment used for biological hazards.

The specific thresholds will require continued research. The management pattern is already useful: decide in advance which evidence changes the organization's obligations.

Frontier-model security is shared infrastructure

Anthropic has published an initiative focused on the security of advanced artificial intelligence models. The concern is straightforward: as models become more capable and expensive to produce, their weights, training systems, research, and deployment infrastructure become valuable targets for theft or misuse.

The security problem does not sit in one server room. It crosses the organization and its supply chain.

Function calling moves risk beyond the chat window

OpenAI has added function-calling support to its chat models. Developers can describe functions using structured definitions, and the model can return arguments that an application may use to call external tools or retrieve information.

This is an important improvement for building reliable integrations. It also makes a boundary explicit: the model proposes; the application decides what happens next.

Guardrails are architecture, not decoration

NVIDIA has released NeMo Guardrails, an open-source toolkit for adding programmable controls to applications built with large language models. The toolkit is intended to help developers keep conversations on topic, reduce harmful or inaccurate responses, and restrict unsafe connections to external applications.

The word guardrail can sound like a thin layer around an otherwise complete product. In a consequential artificial intelligence system, it should be treated as architecture.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.