Skip to content

Risk Management

Frontier preparedness needs decision rights before the crisis

OpenAI has announced a Preparedness team and a challenge focused on risks from increasingly capable artificial intelligence models. The effort will examine areas such as cybersecurity, persuasion, autonomy, and other severe harms, with the aim of connecting evaluation to development and deployment decisions.

Preparedness is not only the ability to detect a dangerous capability. It is the ability to decide and act while the evidence is incomplete and the stakes are rising.

Scaling policies turn capability into a management trigger

Anthropic has published a Responsible Scaling Policy (RSP) that ties increasingly strong safety and security measures to evidence that a model has reached particular dangerous capabilities. The policy introduces Artificial Intelligence Safety Levels (ASLs), loosely inspired by the graduated containment used for biological hazards.

The specific thresholds will require continued research. The management pattern is already useful: decide in advance which evidence changes the organization's obligations.

Multimodal AI requires multilayer evaluation

Microsoft Research has published an overview of responsible artificial intelligence work on multimodal systems—models that analyze or generate across text, images, audio, and other forms of data. The research highlights a practical problem: risks can appear in the combination even when each input looks acceptable on its own.

Evaluation must follow the system across modalities, interactions, and real-world effects.

Frontier-model security is shared infrastructure

Anthropic has published an initiative focused on the security of advanced artificial intelligence models. The concern is straightforward: as models become more capable and expensive to produce, their weights, training systems, research, and deployment infrastructure become valuable targets for theft or misuse.

The security problem does not sit in one server room. It crosses the organization and its supply chain.

Model evaluation needs an early-warning function

Researchers from Google DeepMind and several partner organizations have proposed a framework for evaluating general-purpose artificial intelligence models for dangerous capabilities and misalignment. Their central idea is to test for emerging risks early enough that developers can change training, security, or deployment decisions before a capability becomes difficult to contain.

That makes evaluation more than a scorekeeping function. It becomes an early-warning system.

GPT-4 raises the standard for deployment evidence

OpenAI has released Generative Pre-trained Transformer 4 (GPT-4), a multimodal model that accepts image and text inputs and produces text. The accompanying technical report and system card describe strong performance across professional and academic benchmarks, alongside familiar limitations: unreliable facts, reasoning errors, bias, and behavior that can be difficult to characterize completely.

The release offers more than a new capability. It offers a useful distinction between evidence about a model and assurance about a deployed system.

An API turns a model into an organizational dependency

OpenAI has made ChatGPT and Whisper available through application programming interfaces (APIs). Developers can now add conversational language and speech-to-text capabilities to products without training or hosting the underlying models. The lower cost and simpler integration will accelerate experimentation.

It will also make a third-party model part of more organizations' operating machinery.

AI risk management begins with context

The National Institute of Standards and Technology (NIST) has released version 1.0 of its Artificial Intelligence Risk Management Framework (AI RMF). It is voluntary, sector-neutral, and deliberately flexible. That may frustrate anyone looking for a short compliance checklist. It is also the framework's most useful design choice.

Artificial intelligence (AI) risk is not a property of a model in isolation. It emerges from what the system is asked to do, the conditions under which it operates, the people who depend on it, and the organization's capacity to recognize and respond when it fails.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.