Skip to content

AI Engineering

Principles do not field themselves: why I am writing RAIDops

I did not set out to write a book.

RAIDops began during my Master of Science studies in Columbia University's Information & Knowledge Strategy program. An independent study, guided by Blake M. DiCosola III and strengthened by the advice and input of Edward J. Hoffman, started as a literature review of trustworthy artificial intelligence in national defense. Blake and Ed are both IKNS faculty; Ed previously served as NASA's first Chief Knowledge Officer.

The review grew into a long paper. The paper kept returning to an unresolved organizational problem. Eventually, the problem outgrew the paper.

I invented RAIDops and coined the name Responsible AI Development Operations for the framework that emerged from that work. The working monograph is coauthored with Blake and Ed, whose substantive intellectual contributions, guidance, editing, and mentorship have materially shaped it.1 Their collaboration has made the work substantially better.

I have hesitated to write publicly about RAIDops because the research program is active and the manuscript is unfinished. I am not going to reproduce the complete pattern catalog, assessment instruments, or the book's full analytical machinery here. But a framework concerned with reviewability should itself be open to review. This essay offers the public argument: enough to explain and defend RAIDops, while preserving the monograph as the place where the complete derivation, architecture, patterns, evidence controls, and limitations belong.

The thesis is straightforward: trustworthy AI is not merely a property to test in a model. It is an operating achievement that an organization must repeatedly produce, challenge, bound, preserve, and sometimes revoke.

That problem is not abstract to me. It recurs across my work in enterprise AI engineering, knowledge systems, and public-sector technology: building a capable system is only part of the job. The institution must also keep the purpose, evidence, authority, and means of intervention intact as the system crosses teams, contracts, environments, and time.

Multi-agent systems need a shared world, not just shared messages

The proposal window for the Defense Advanced Research Projects Agency's (DARPA) Decentralized Artificial Intelligence through Controlled Emergence program closed yesterday. DICE asks a difficult systems question: can heterogeneous artificial intelligence agents coordinate through peer-to-peer interaction, adapt when individual agents fail or become compromised, and still remain aligned with commander's intent over long missions?

Two days earlier, researchers released a preprint describing an “ontology as a kernel” for language-model agents. The proposed system makes domain concepts, relationships, evidence, and permissible reasoning operations explicit instead of leaving all of them implicit in prompts and unstructured context.

Those developments come from different research communities, and neither proves the other's architecture. Together, they expose the same design boundary: a collection of agents does not become a system merely because the agents can exchange messages. It becomes a system when they can coordinate around a shared, inspectable, and governed model of the world and the work.

That is a semantic-systems problem.

The first deliverable from generative coding should be understanding

The most dangerous sentence in a modernization program may be, “We know what this system does.”

Usually, someone knows what the system is supposed to do. Operators know the screens and workarounds. A few engineers know where the brittle integrations live. Program managers know the contracts and milestones. Cybersecurity teams know some of the exposed surfaces. The source code knows all of it at once—but in a form no single person can hold in mind.

That is why the emerging market for generative coding matters to government for a reason deeper than writing software faster. Its first serious public-sector use may be helping an organization recover a working model of the technology estate it already owns.

Model the builders, not only the models

We usually evaluate high-stakes AI by inspecting the artifact: the model, benchmark, system card, test report, or approval package. But before deployment, accountability is created—or eroded—by a population of builders acting through partial information, uneven authority, deadlines, review queues, and AI-mediated tools.

That makes builder-side AI development a natural candidate for agent-based modeling.

Verification is now part of the product surface

OpenAI has released Generative Pre-trained Transformer 5.5 (GPT-5.5), describing a model that can carry more of a complex task across coding, research, data analysis, document creation, and software tools. This continuing increase in agentic capability changes the user's job.

When a system produces a paragraph, review can happen at the paragraph. When it completes an hour of work across several applications, review has to cover a chain of actions, transformed data, and consequential choices.

Stateful agents make memory a governance problem

OpenAI and Amazon have announced a strategic partnership that includes plans to co-develop a stateful runtime environment for artificial intelligence agents. The phrase “stateful” deserves attention. An agent that can preserve context across steps and sessions may be more useful than one that repeatedly starts from zero.

It may also accumulate assumptions, permissions, and mistakes that no one intended to become durable.

Agent security starts before the agent acts

The National Institute of Standards and Technology (NIST) has opened a request for information on securing artificial intelligence agent systems. The timing is right. Organizations are moving from systems that generate suggestions toward systems that can call tools, manipulate records, send messages, and coordinate multi-step work.

The security question is no longer only what a model might say. It is what the complete system is allowed to do—and whether the organization can reconstruct what happened afterward.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.