Skip to content

2023

AI safety needs better questions before better rules

The National Institute of Standards and Technology (NIST) has issued a request for information (RFI) on the safe, secure, and trustworthy development and use of artificial intelligence. Responses will inform future guidance on evaluation, red teaming, risk management, and related measurement challenges.

The request arrives after a year of rapid capability releases and equally rapid calls for guardrails. Before guidance becomes more specific, the field needs to ask more precise questions about systems, evidence, and use.

A cyber challenge can build an ecosystem, not just a winner

The Defense Advanced Research Projects Agency (DARPA) has opened registration for the Artificial Intelligence Cyber Challenge (AIxCC), published an exemplar challenge and scoring approach, and added prize funding. Competitors will work toward systems that can find and repair vulnerabilities in widely used software at scale.

Prizes attract teams. The lasting value of a challenge can be the common infrastructure and professional community built around the competition.

Gemini makes evaluation a portfolio capability

Google has introduced Gemini 1.0, a family of multimodal artificial intelligence models in three sizes: Ultra, Pro, and Nano. The models are designed to work across text, images, audio, video, and code, and to run in environments ranging from data centers to mobile devices.

The release adds another capable model family to a fast-changing field. For organizations, the strategic response is not to crown a universal winner. It is to become good at evaluating fit repeatedly.

Retrieval-augmented generation is not a knowledge strategy

Amazon Web Services (AWS) has made Knowledge Bases for Amazon Bedrock generally available. The service can ingest organizational documents, create a searchable vector index, retrieve relevant passages, and use them to ground a foundation model's response—with source attribution included.

Managed retrieval removes a meaningful amount of engineering work. It does not decide which organizational knowledge should be trusted.

AI studios need production discipline

Microsoft has announced the public preview of an artificial intelligence (AI) development environment, Azure AI Studio, at Ignite. It brings model selection, data grounding, evaluation, content safety, and deployment tooling into a common workspace. The platform reflects how quickly generative AI development is moving from isolated notebooks toward managed application delivery.

A studio can make the path to a prototype remarkably short. The path to a dependable product still needs discipline.

GPTs turn prompting into configuration management

At its first developer conference, OpenAI has introduced generative pre-trained transformers (GPTs): custom versions of ChatGPT that can combine instructions, uploaded knowledge, and selected capabilities for a particular purpose. People can build them without conventional programming and share them inside an organization or, eventually, through a public store.

This will make useful experimentation easier. It will also turn a large number of informal prompts into organizational configurations that can affect real work.

Militarizing Innovation: The Path to Global Stability

In a recent dialogue with The Economist, Ukraine’s commander-in-chief, General Valery Zaluzhny, gave a stark assessment of their ongoing conflict with Russia – they lack the technological advantage to make strides and are at a stalemate with Russia. His reflections resonate with a long-standing consensus among military strategists and geopolitical analysts regarding the critical importance of technological advancement in warfare. The Russo-Ukrainian war, despite the employment of modern military technologies, demonstrates a broader imperative for the West: to win wars and mitigate risks, a nation must relentlessly pursue and attain technological dominance over its adversaries.

Frontier preparedness needs decision rights before the crisis

OpenAI has announced a Preparedness team and a challenge focused on risks from increasingly capable artificial intelligence models. The effort will examine areas such as cybersecurity, persuasion, autonomy, and other severe harms, with the aim of connecting evaluation to development and deployment decisions.

Preparedness is not only the ability to detect a dangerous capability. It is the ability to decide and act while the evidence is incomplete and the stakes are rising.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.