Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

Principles do not field themselves: why I am writing RAIDops

I did not set out to write a book.

RAIDops began during my Master of Science studies in Columbia University's Information & Knowledge Strategy program. An independent study, guided by Blake M. DiCosola III and strengthened by the advice and input of Edward J. Hoffman, started as a literature review of trustworthy artificial intelligence in national defense. Blake and Ed are both IKNS faculty; Ed previously served as NASA's first Chief Knowledge Officer.

The review grew into a long paper. The paper kept returning to an unresolved organizational problem. Eventually, the problem outgrew the paper.

I invented RAIDops and coined the name Responsible AI Development Operations for the framework that emerged from that work. The working monograph is coauthored with Blake and Ed, whose substantive intellectual contributions, guidance, editing, and mentorship have materially shaped it.1 Their collaboration has made the work substantially better.

I have hesitated to write publicly about RAIDops because the research program is active and the manuscript is unfinished. I am not going to reproduce the complete pattern catalog, assessment instruments, or the book's full analytical machinery here. But a framework concerned with reviewability should itself be open to review. This essay offers the public argument: enough to explain and defend RAIDops, while preserving the monograph as the place where the complete derivation, architecture, patterns, evidence controls, and limitations belong.

The thesis is straightforward: trustworthy AI is not merely a property to test in a model. It is an operating achievement that an organization must repeatedly produce, challenge, bound, preserve, and sometimes revoke.

That problem is not abstract to me. It recurs across my work in enterprise AI engineering, knowledge systems, and public-sector technology: building a capable system is only part of the job. The institution must also keep the purpose, evidence, authority, and means of intervention intact as the system crosses teams, contracts, environments, and time.

The evidence gap is becoming an operating problem

Two recent federal reports make the need visible.

In April 2026, the U.S. Government Accountability Office reported that federal agencies had more than doubled their reported AI use from 2023 to 2024. In reviewing selected AI acquisitions, GAO identified six recurring challenge areas: access to technical experts; protection of government data and intellectual-property rights; conventional acquisition approaches and timelines; requirements and contract terms; early testing and continuous evaluation; and pricing and total cost. It also found that the Department of Defense, Department of Homeland Security, General Services Administration, and Department of Veterans Affairs were not systematically collecting and sharing acquisition lessons.

That is not one problem for procurement and another for engineering. The requirements determine what the supplier builds. Data and intellectual-property terms determine what the government can inspect, reproduce, improve, and retain. Test provisions determine which claims can be evaluated. Monitoring terms determine whether evidence continues after release. Knowledge-transfer requirements determine whether one program's experience can improve the next. OMB Memorandum M-25-22 already treats cross-functional engagement, performance monitoring, competition, data rights, and knowledge sharing as acquisition concerns; GAO's findings show why those concerns have to become operating practice.

In March 2026, the National Institute of Standards and Technology (NIST) published Challenges to the Monitoring of Deployed AI Systems: Center for AI Standards and Innovation. Its scope extends beyond conventional model telemetry to functionality, operations, human factors, security, compliance, and large-scale impacts. NIST's central warning is appropriately restrained: common terminology, validated methodologies, and best practices for postdeployment AI monitoring remain nascent and scattered.

Together, those reports describe an institutional condition. Organizations can acquire models, build pipelines, convene review boards, and publish principles while still failing to preserve the relationship among a system's purpose, evidence, authority, operating context, and consequences.

That relationship is where RAIDops begins.

The gap is not a lack of principles

Responsible AI has no shortage of normative frameworks. NIST's AI Risk Management Framework organizes continuous work across Govern, Map, Measure, and Manage. ISO/IEC 42001 specifies requirements for establishing, implementing, maintaining, and continually improving an AI management system. The Department of Defense's Responsible AI Strategy and Implementation Pathway connects governance, acquisition, product lifecycle, workforce, and warfighter trust. Other fields contribute model-risk management, safety cases, human-centered design, impact assessment, audit, and professional duties.

Those frameworks matter. RAIDops is not an argument that the field forgot to define responsibility.

The harder problem is that principles do not enact themselves. Brent Mittelstadt's critique that principles alone cannot guarantee ethical AI remains foundational. Morley and colleagues found a persistent gap between the ethical what and the practical how of AI development. Field research has since shown how responsible-AI work collides with organizational incentives, fragmented authority, insufficient resources, and the individualization of risk among people expected to make ethics operational (Rakova et al. 2021; Ali et al. 2023).

I made a related argument in an earlier Field Note: responsible AI needs an operating system. RAIDops is the larger research program behind that intuition. It asks what the operating model must connect when AI work is distributed across product teams, acquisition offices, suppliers, data owners, lawyers, security engineers, independent evaluators, human-factors specialists, executives, operators, and affected institutions.

The missing layer is not another set of values above those people. It is a way to keep their evidence, judgment, authority, and responsibility attached to the same consequential decision while the technology and context change underneath them.

What I mean by RAIDops

On the public RAIDops research page, I define Responsible AI Development Operations as a builder-side, lifecycle-wide socio-technical framework for turning responsible-AI commitments into ordinary development and operational work.

Builder-side does not mean software engineers alone. It includes the people and organizations that conceive, acquire, fund, specify, supply, design, integrate, evaluate, approve, release, monitor, modify, transfer, suspend, and retire an AI-enabled capability.

Lifecycle-wide means responsibility begins before code exists and continues after deployment. Problem formulation determines what will count as success. Contract design distributes inspection rights and duties. Data work determines whose experience becomes visible. Interface design changes what a user can understand and contest. Release binds evidence to a configuration. Monitoring tests whether the conditions supporting the decision still hold. Retirement determines whether the organization can actually stop when they do not.

Socio-technical means the relevant object is not the model in isolation. It is the configured decision system: model, data, dependencies, interface, workflow, users, incentives, policy, authority, intervention, and recourse.

The simplified architecture looks like this:

flowchart TB
    accTitle: RAIDops conceptual operating architecture
    accDescr: Principles and mission intent become bounded claims evaluated across five builder layers. A lifecycle decision leads to release, limitation, redesign, hold, rejection, suspension, or retirement, and operational evidence reopens the claims.
    P["Principles, law, professional duties<br/>and mission intent"] --> C["Purpose- and context-specific claims"]
    C --> B["Five mutually dependent builder layers:<br/>Organizational<br/>Policy & governance<br/>Pipeline & automation<br/>Assurance & monitoring<br/>Socio-technical integration"]
    B --> J["Bounded lifecycle decision"]
    J --> R["Release, limit, redesign, hold,<br/>reject, suspend or retire"]
    R --> F["Operational evidence, incidents<br/>and changed context"]
    F --> C

Figure 1. A conceptual overview of RAIDops. The five builder layers are mutually dependent operating domains, not sequential phases. Operational evidence and material change reopen earlier claims and decisions.

Diagram in words. RAIDops translates principles, law, professional duties, and mission intent into purpose- and context-specific claims. Five mutually dependent builder layers inform a bounded lifecycle decision; operational evidence, incidents, and changed conditions then reopen the claims rather than allowing an approval to become permanent.

RAIDops does not replace MLOps, DevSecOps, systems engineering, safety engineering, law, governance, assurance, or human-centered design. It works through and alongside them. Its contribution is the operating architecture that asks how their outputs meet at an answerable decision.

That is the invention boundary. I did not invent the constituent disciplines, practices, or artifacts. I developed RAIDops as a synthesis and operating architecture for connecting them around a specific problem: keeping responsibility, evidence, credible challenge, and authority attached to consequential AI decisions as systems are defined, built, transferred, changed, used, and retired.

Across those boundaries, RAIDops returns to five recurring questions:

  1. Who owns this decision?
  2. What evidence supports it?
  3. Who can credibly challenge it?
  4. What conditions would invalidate it?
  5. What context, constraints, and accountability must travel when the system changes hands?

These questions are not a certification test. They are a way to discover whether an apparently complete process still contains an unowned decision.

Imagine a contractor-developed model evaluated for one advisory workflow. The receiving organization preserves the weights and model card but connects a new data feed, compresses uncertainty in the interface, and gives the recommendation a different operational role. The build remains reproducible and every encoded test stays green. Yet the evidence no longer describes the complete system in use, and no one has explicitly accepted the changed authority relationship. Every team may have completed its local task while the whole decision has become indefensible.

Product evidence and builder capability are different questions

One of RAIDops's most important distinctions is also one of its simplest.

The field has already developed substantial accounts of what a trustworthy AI product should be able to support. RAIDops groups those claims into four broad families:

  1. Governance and Accountability
  2. Ethical and Legal Risk Management
  3. Technical Design and Assurance
  4. Human Factors and Trust

RAIDops is not a fifth product pillar. It organizes the builder-side capabilities through which an organization produces, judges, preserves, challenges, and reopens the evidence behind those claims. Those capabilities sit in five mutually dependent layers:

  1. Organizational
  2. Policy and Governance
  3. Pipeline and Automation
  4. Assurance and Monitoring
  5. Socio-Technical Integration

Collapsing those dimensions into one maturity or trust score would erase the decision RAIDops is meant to expose.

Dimension Product evidence Builder capability
Central question Does current evidence support this claim about this configured system in this use? Can this builder repeatedly perform, challenge, preserve, and improve the work required for this decision?
Unit of assessment A bounded configuration, use, context, decision, date, and evidence cutoff A named organizational boundary, operating layer, period, and decision scope
Responsible outcomes Support, partial support, non-applicability with rationale, or a finding that the claim is unsupported The demonstrated ability to produce evidence, expose limits, preserve dissent, route authority, intervene, and learn
What it cannot prove That the organization will remain capable or that the claim will survive material change That a particular product is trustworthy, permissible, safe, or ready to deploy

Table 1. Product evidence and builder capability answer different questions. Neither dimension substitutes for the other.

A sophisticated builder can responsibly discover that the product case fails. In fact, the ability to hold, narrow, redesign, reject, suspend, or retire a use may be stronger evidence of institutional capability than an operating model that converts every review into approval.

This is why RAIDops is not designed to make every AI program deployable. It is designed to make consequential decisions more visible, challengeable, bounded, and revisable.

A green pipeline is not an assurance judgment

Modern AI engineering already has powerful operational disciplines. MLOps connects development, deployment, monitoring, data, and model lifecycle work. DevSecOps makes security part of ordinary delivery. NIST's Secure Software Development Framework gives producers and acquirers a common vocabulary for integrating secure practices into the software lifecycle. The Department of Defense's Enterprise DevSecOps Fundamentals provides a rich platform, pipeline, evidence, and continuous-authorization substrate.

RAIDops should reuse those mechanisms. A delivery system can verify that a release candidate has a stable identity, uses an approved data version, passed a named test, meets an encoded threshold, carries a required approval, and is permitted to enter a destination. This is the beginning of an executable evidence chain.

But executable is not synonymous with sufficient.

A pipeline cannot determine, by itself, whether the selected outcome was legitimate, the affected population was properly represented, the metric captured the relevant harm, the threshold was normatively defensible, the reviewer was independent, or residual risk was acceptable. It can prove that encoded checks passed. It cannot prove that the organization encoded the right judgment.

Assurance research makes a similar distinction. Ashmore, Calinescu, and Paterson describe assurance as the lifecycle production of evidence that a machine-learning component is sufficiently safe for its intended use—not a synonym for test accuracy (2021). End-to-end audit research similarly connects organizational processes, documentation, testing, and accountability across the development lifecycle (Raji et al. 2020). Claims, arguments, context, assumptions, and evidence can be made explicit through assurance cases. The notation does not make the evidence true or the judgment wise.

The design target is therefore neither governance as prose nor governance as code. It is a reviewable relationship among claims, evidence, authority, conditions, and technical state, with automation where propositions are determinate and accountable judgment where they are not.

Agentic systems make the operating gap harder to ignore

Agentic and multi-agent systems distribute analysis and action across planners, retrievers, evaluators, and tool-using delegates. A semantic layer or knowledge graph can give those actors a shared representation of the operating domain: purposes, entities, relationships, policies, evidence, decisions, and provenance. That shared world can make coordinated work more coherent. It does not make the representation complete, the retrieved claim true, or the resulting action authorized.

I have written that multi-agent systems need a shared world, not merely shared messages. RAIDops adds the institutional half of the problem. Which human or organization delegated the work? Which version of the shared representation informed it? Which policy bounded the tools? Which evidence was produced outside the agents' ability to revise? Who owns acceptance? Which event revokes authority or reopens the decision?

As agents become more capable of interpreting evidence, proposing plans, invoking tools, and coordinating other agents, the consequential question is no longer merely whether they can complete a task. It is whether the organization can preserve delegated authority, independent evaluation, traceable context, explicit intervention rights, and responsibility for the resulting decision. Those controls are not friction around the “real” work. They are part of the operating design that makes autonomous work answerable.

Evidence must travel with the system

AI systems rarely remain inside the boundary where they were first evaluated. A model is incorporated into a product, delivered by a contractor, connected to new data, placed behind a different interface, moved to another mission, or used by people with a different authority and workload.

The artifact may look technically continuous while its job changes completely.

RAIDops calls the resulting requirement accountability-bearing transfer. The idea is not that one party can hand responsibility to another in a document. Responsibilities may branch among the supplier, integrator, acquirer, operator, product owner, and decision authority. The requirement is that enough context and evidence travel for the receiving organization to make its next decision honestly.

flowchart TB
    accTitle: Accountability-bearing transfer
    accDescr: Intended uses, configuration, evidence, limitations, operating conditions, and authority travel to a receiving organization, which makes a new bounded decision to accept, limit, redesign, hold, reject, suspend, or retire the system.
    O["Originating system, use<br/>and decision context"] --> T["Accountability-bearing transfer package:<br/>Defined and prohibited uses<br/>Configuration, dependencies and provenance<br/>Supported claims and evidence<br/>Uncertainty, limitations and unresolved dissent<br/>Operating conditions, expiry and stop triggers<br/>Decision authority and retained responsibilities"]
    T --> N["Receiving organization makes<br/>a new bounded decision"]
    N --> Y["Accept a bounded use"]
    N --> L["Limit or redesign"]
    N --> H["Hold, reject, suspend or retire"]

Figure 2. Accountability-bearing transfer is a new reviewable decision, not a file copy. Technical continuity does not establish continuity of purpose, evidence, authority, or acceptable consequence.

Diagram in words. A transfer package carries permitted and prohibited uses, technical provenance, supporting evidence, uncertainty, unresolved dissent, operating conditions, stop triggers, and retained responsibilities. The receiving organization must use that context to make a new bounded decision rather than treating delivery as inherited approval.

This is closely related to decision provenance and to research on dislocated accountability in AI supply chains. The difference RAIDops emphasizes is operational: provenance and limitations must remain connected to version, use, authority, and the controls that can alter what happens next.

Documentation is necessary but insufficient. A model card can state an intended use; it does not ensure the recipient sees it, interprets it correctly, or possesses authority to reject a conflicting deployment. A risk register can preserve a concern; it does not ensure the concern receives a disposition. An approval can be authentic and still be applied to the wrong configuration.

Evidence becomes operational when it can change the state of the work.

The operating model must be federated

RAIDops is not a proposal for one central ethics committee to approve every technical decision.

Central functions can establish common language, minimum evidence, independent challenge, reusable methods, infrastructure, and escalation paths. Product and mission teams hold local knowledge about the workflow, affected people, operational constraints, and intended outcome. Technical and domain specialists know different parts of the risk. Executives, authorizing officials, clinicians, commanders, regulators, or other legitimate authorities may own different dispositions.

The design problem is how those forms of knowledge and authority meet without being flattened.

My related theory of networked accountability treats predeployment review as a routing problem across expertise, credibility, reinforcement, and persistence.2 The relevant organizational questions are practical:

  • Can risk-relevant knowledge reach an authorized decision?
  • Can it receive competent and credible challenge?
  • Can the challenge affect the disposition rather than merely be recorded?
  • Can the concern survive organizational and contractual boundaries?
  • Can the reasoning be recovered when the decision must be revisited?

Cross-functional composition alone does not answer those questions. A lawyer consulted after the architecture has hardened, an operator represented only through a manager, a security reviewer without access to the model's dependencies, or a data scientist unable to see the final interface may all be “in the process” while lacking practical influence.

Formal authority can be ceremonial. Credible authority needs relevant evidence, competence, time, organizational standing, an answerable forum, and a usable range of actions.

Monitoring should test the conditions of the claim

An AI system does not need to exhibit conspicuous statistical drift for its assurance case to weaken.

The user population can change. A supplier can update a dependency. An interface can compress uncertainty. A workload can make nominal human review impossible. A policy can change the permitted purpose. A new threat can alter the security assumption. A successful deployment can encourage use outside the conditions that made it successful.

This is why I have argued that monitoring deployed AI is a knowledge practice. Telemetry becomes governance only when observations reach people who can interpret them, connect them to the original claim, and change the product, procedure, authority boundary, or permission to operate.

NIST AI 800-4 makes the breadth of that problem explicit. Monitoring may need to cover system functionality, operational performance, human interaction, security, compliance, and societal effects. RAIDops adds a particular question: which observed change could invalidate the basis on which this use was accepted?

The answer should connect to a predeclared response. Investigate. Narrow the population. Restrict a feature. Revalidate. Roll back. Suspend. Notify. Retire.

Rollback is not moral time travel. It may restore a prior technical version, but it cannot undo a denied benefit, an unsafe recommendation, an information disclosure, or a decision already acted upon. Operational readiness therefore includes recourse, incident learning, and the authority to intervene while intervention still matters.

Defense is the originating stress test, not the boundary

RAIDops grew from research into trustworthy AI in national defense because defense makes the organizational problem difficult to ignore. AI-enabled capabilities cross contractors, program offices, test organizations, classification boundaries, command structures, coalition relationships, acquisition pathways, and operational contexts. Evidence can be restricted. Authority is distributed. Mission pressure is real. Consequences can be irreversible.

Those conditions make defense a demanding stress test. They do not make RAIDops a defense-only framework.

I define high stakes as a relationship among capability, decision, context, authority, and consequence—not as a model type or sector label. A clinical risk score, fraud detector, public-benefits prioritization system, industrial control assistant, or military decision-support system becomes consequential through the work it enters and the power that work exercises.

What can travel across sectors is a common operating requirement: keep purpose, evidence, authority, challenge, change, and intervention connected. What cannot simply travel is the institutional answer. Healthcare, finance, public administration, critical infrastructure, intelligence, and defense each supply different law, professional duties, evidence thresholds, decision authorities, intervention rights, and forms of recourse.

The most credible cross-sector design is therefore one core with domain-specific profiles, not a universal checklist with the job titles renamed. FDA's final guidance on predetermined change-control plans for AI-enabled device software functions, for example, offers a powerful domain precedent for predefining the evidence and controls around material change. It is not a universal legal rule for every AI system.

What leaders and builders should be able to answer

The complete RAIDops work includes a proposed pattern language and assessment instruments. I am intentionally not reproducing them here. The five questions above are a simpler public diagnostic. Leadership teams and builders should extend them:

  • What decision or mission outcome is the AI system actually changing?
  • What is the complete configured system—not only the model?
  • Which bounded claims justify the next lifecycle decision?
  • Which artifacts and observations support each claim, and which uncertainty remains?
  • Which model, data, interface, supplier, workflow, policy, population, or authority changes require reconsideration?
  • Who can hold, narrow, redesign, reject, suspend, or retire the use?
  • Where do exceptions, dissent, near misses, overrides, and workarounds go?
  • How does experience from one program change requirements, contracts, tooling, training, or decisions in the next?

If the answers live only in the memories of well-connected people, the organization has expertise but not yet an operating capability.

Predictable objections are part of the research

Is RAIDops simply MLOps with more governance?

No. MLOps is a vital adjacent discipline, and strong implementations already include governance, monitoring, documentation, and human roles. RAIDops asks a different integrative question: how do technical, organizational, policy, assurance, and human-work evidence meet at a bounded consequential decision? It should strengthen existing pipelines, not build a ceremonial parallel route around them.

Will this create more bureaucracy?

It could. Any framework can become theater. Documents can multiply while authority remains vague and tests remain irrelevant. RAIDops should be evaluated partly by whether it reduces rediscovery, makes handoffs less lossy, automates determinate controls, exposes decision debt earlier, and helps people intervene before commitment. If it merely adds forms, it has failed its own purpose.

Is an assurance case enough?

No. Evidence is not an argument; an argument is not a decision; confidence is not approval; and approval is not permanent. An assurance case can make the reasoning inspectable. Competent people still have to judge the evidence, consider alternatives and affected parties, own the disposition, and revisit it after material change.

Does RAIDops prove that a system is trustworthy?

No. RAIDops is a proposed framework, not a certification scheme, compliance regime, calibrated maturity scale, or validated intervention. Its architecture is grounded in established research and practice, but the integrated framework and its proposed patterns require empirical evaluation. Derivation is not validation.

Why call it an invention if the constituent practices already exist?

Because the original contribution is the synthesis and operating architecture, not a claim to have invented its ingredients. RAIDops connects established bodies of work around an unresolved builder-side problem and proposes an integrated architecture that can be examined, challenged, and evaluated.

That contribution is explicit enough to examine—and bounded enough to prove wrong.

The research is becoming more rigorous as it grows

The initial literature review gave RAIDops its question and much of its intellectual provenance. It should not be used circularly as proof that RAIDops is correct. As the project grew, I separated two workstreams: the framework and monograph continue as an author-directed interdisciplinary synthesis; separately, the surviving evidence record of the original review is being reconstructed, and a prospective de novo systematic scoping-review protocol has been designed—but not yet executed—for reporting with PRISMA-ScR and PRISMA-S.3

That separation matters. A framework should not predetermine the coding or conclusions of the review that may later test its premises. The current RAIDops patterns and assessment instruments remain proposals. Existing responsible-AI pattern work provides a useful precedent for naming reusable interventions (Lu et al. 2024); it does not validate RAIDops's selection, arrangement, or effects. Those remain empirical questions.

This is also why I feel compelled to publish something now. An unfinished monograph should not be insulated from the communities it hopes to serve. Scholars can challenge the conceptual boundaries and research design. Practitioners can identify where the operating model conflicts with actual authority, delivery pressure, acquisition, or field conditions. Leaders can ask whether the framework improves decisions or merely describes them more elegantly.

The public essay is not the whole book. It is an invitation to interrogate the central claim.

Trustworthy AI has to be reproduced through change

The durable problem is not getting an AI system through one review. It is preserving the basis of a consequential decision while the system changes jobs, versions, users, suppliers, environments, and hands.

An organization may have excellent principles and still lose them at a handoff. It may have a green pipeline and still test the wrong claim. It may retain a human “in the loop” while the interface, workload, and institutional incentives have already narrowed that person's practical authority. It may record dissent without allowing dissent to alter the disposition. It may deploy responsibly once and then allow success itself to expand the use beyond the evidence.

RAIDops is my attempt to make those failures harder to hide and the responsible alternatives easier to enact.

The standard should not be whether a framework helps an organization say yes. The test of RAIDops is whether the organization can know—and show—when the responsible answer is no.1

Sources and research trail


  1. West, Bruce B. A., Blake M. DiCosola III, and Edward J. Hoffman. 2026. RAIDops: Responsible AI Development Operations—A Socio-Technical Framework for Building Trustworthy AI in High-Stakes Organizations. Unpublished working monograph, version 0.27, August 19. Cited here for intellectual provenance and the framework's current definition—not as independent validation of RAIDops. 

  2. West, Bruce B. A. 2026. “Networked Accountability in High-Stakes AI Development: A Mechanism-Fit Theory of Predeployment Review.” Unpublished manuscript, version 1, August 12. Cited as a proposed organizational theory, not empirical validation of RAIDops. 

  3. West, Bruce B. A., Blake M. DiCosola III, and Edward J. Hoffman. Trustworthy AI in National Defense: A Comprehensive Literature Review. Unpublished manuscript. The original review is cited here only as project provenance. Its surviving evidence record is being reconstructed; inherited screening counts are quarantined; and fresh search and screening under the prospective de novo protocol have not yet been executed. 

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.