Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

AI safety must travel with the product

OpenAI has described its approach to artificial intelligence safety, including pre-release testing, external input, human feedback, monitoring, and phased deployment. The account reflects an important reality: safety is not a property added once to a model and inherited automatically by everything built with it.

Safety has to travel with the product.

The deployment changes the risk

Artificial intelligence (AI) developers can evaluate a foundation model broadly, probe harmful behavior, and apply mitigations. A product team then places that model inside an interface, gives it instructions, connects data, chooses user permissions, and defines a workflow. Each choice changes how capability is expressed and how failure reaches people.

A model may be relatively safe in an isolated chat and less safe when embedded in a medical intake process, software deployment pipeline, or system with external tools. Conversely, a carefully constrained application may reduce risks that appear in open-ended use.

The unit of assurance therefore must be the deployed system—not only the underlying model.

Safety evidence can decay

AI products change quickly. Model versions, prompts, filters, retrieval sources, tools, and user populations evolve. Evidence collected for one configuration may not support the next.

This creates a form of assurance debt. Teams accumulate claims about safety while the system moves away from the configuration that was tested. The interface still works, so the debt can remain invisible until a failure exposes it.

The National Institute of Standards and Technology (NIST) Artificial Intelligence Risk Management Framework addresses this through lifecycle risk management and continuous monitoring. The principle is simple: deployment is not the end of evaluation. It is the beginning of a new source of evidence.

Users encounter contexts, ambiguities, and incentives that a laboratory cannot fully reproduce. Their corrections, complaints, workarounds, and near misses are not merely support data. They are safety signals.

Connect every release to an argument

A practical release process should maintain a trace among four things:

  1. Claim: What do we believe is acceptably safe for this use?
  2. Configuration: Which model, instructions, data, tools, interface, and access controls are included?
  3. Evidence: Which tests and observations support the claim?
  4. Response: Which monitoring signal or incident would require restriction, rollback, or renewed testing?

That trace should be versioned with the product. A change to a prompt may require only focused regression tests. Adding a tool with transaction authority may require a substantially different threat model and approval. The trigger should follow the consequence, not the apparent size of the code change.

Red teams need a route into product management

Adversarial testing is valuable only if findings can alter requirements and release decisions. Red-team results should not sit in a separate security report that product leaders acknowledge and then route around.

For each meaningful finding, capture the exploited condition, expected consequence, mitigation, residual uncertainty, and owner. Then place the mitigation in the same backlog and release gate as performance work. Security, safety, and usability often interact; a refusal mechanism that is too broad may push users toward unsafe workarounds, while a warning that appears constantly will be ignored.

Human-centered design matters here. Amershi and colleagues' guidelines for human–AI interaction emphasize setting expectations, supporting correction, and making clear what the system can do. Those are not only usability concerns. They are part of the safety architecture.

The strongest feature of an iterative deployment approach is not that it promises perfect foresight. It is that it creates repeated opportunities to learn and intervene. Organizations adopting AI should demand that the safety evidence survives those iterations.

The model provider has responsibilities. So does the product team. Safety must be carried from research into configuration, from configuration into deployment, and from operational experience back into the next release. If the evidence cannot make that trip, the claim should not either.

Sources and research trail

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.