Skip to content

AI Security

Agents need a control plane outside their reach

Google DeepMind has published an Artificial Intelligence Control Roadmap for securing internal systems as agents become more capable and operate with greater autonomy. The work considers how organizations can manage systems that may be useful, imperfectly aligned, and able to interact with valuable digital resources.

One principle deserves broad adoption: the mechanisms that observe and constrain an agent should not depend entirely on the agent's own willingness or ability to comply.

There is no final security review for AI

The National Institute of Standards and Technology (NIST) has published a mathematical argument supporting a continuous-monitor-and-update security model for artificial intelligence. The practical conclusion is direct: a fixed set of guardrails cannot remain universally robust against adaptive adversarial prompts.

An organization may approve a system for release. It cannot approve the system out of change.

Independent evaluation needs a secure place to work

The Center for Artificial Intelligence Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST) has entered a cooperative research and development agreement with OpenMined. The collaboration is intended to advance secure methods for evaluating artificial intelligence systems.

The agreement points at a recurring barrier to credible assurance: evaluators need access to meaningful systems and evidence, while model developers, customers, and government organizations need to protect intellectual property, personal information, security-sensitive data, and operational methods.

Capability gating is a cybersecurity control

OpenAI has introduced Trusted Access for Cyber, an approach intended to expand advanced cybersecurity capabilities for legitimate defenders while placing additional controls around uses that could create harm.

The initiative highlights a problem that every organization deploying powerful artificial intelligence now faces: access control cannot stop at whether somebody may use a model. It has to consider which capabilities they may invoke, under what conditions, with what evidence, and through which escalation path.

Agent security starts before the agent acts

The National Institute of Standards and Technology (NIST) has opened a request for information on securing artificial intelligence agent systems. The timing is right. Organizations are moving from systems that generate suggestions toward systems that can call tools, manipulate records, send messages, and coordinate multi-step work.

The security question is no longer only what a model might say. It is what the complete system is allowed to do—and whether the organization can reconstruct what happened afterward.

The protocol is becoming an organizational boundary

Anthropic's June 18 update adds remote Model Context Protocol (MCP) support to Claude Code. Developers can connect the coding agent to hosted tools and knowledge sources without operating each integration locally.

This is convenient infrastructure. It also moves an important boundary: the agent can now cross from a code repository into project systems, observability platforms, knowledge bases, and other services through a common protocol.

AI threat intelligence should change the product backlog

Anthropic's April 23 report documents case studies on the malicious use of Claude. The cases include influence operations, credential-related activity, recruitment fraud, and a novice actor using artificial intelligence to advance malware development.

The details matter, but the report's most important feature is the loop it implies: observe abuse, interpret the pattern, change defenses, and share what others can use.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.