Skip to content

AI Strategy

Inference is where AI strategy meets the budget

International Business Machines (IBM) Research has published a timely explanation of artificial intelligence inference—the moment when a trained model receives live input and produces a result. Training attracts attention because it creates the model. Inference is where the model becomes a recurring service, and where much of its lifetime cost and user experience accumulate.

For enterprise leaders, inference is not only an infrastructure concern. It is where an artificial intelligence (AI) portfolio meets a budget.

Open models turn selection into engineering

Meta and Microsoft have announced commercial access to Llama 2, with support across Azure and Windows. The release expands the range of models organizations can host, adapt, and integrate under their own architectural choices rather than consume only through a closed service.

That optionality is valuable. It also moves more of the responsibility from procurement into engineering.

The enterprise AI platform is really a coordination platform

International Business Machines (IBM) has introduced watsonx, an enterprise platform that brings together a studio for foundation models, a data layer, and governance capabilities. The announcement reflects the direction many large organizations are moving: away from isolated model experiments and toward a common environment for building, adapting, and operating artificial intelligence.

The technology matters. The larger challenge is coordination.

Research integration is an organizational-design problem

Google has combined DeepMind and the Google Brain team into a new unit called Google DeepMind. The stated ambition is to bring together talent, computing resources, infrastructure, and research advances to accelerate progress in artificial intelligence.

Mergers of technical groups are often described as exercises in scale. Their success depends just as much on whether distinct communities can combine knowledge without losing the differences that made each one valuable.

Model choice creates a portfolio to govern

Amazon Web Services (AWS) has announced Amazon Bedrock, a managed service intended to give customers access to foundation models from several providers through a common cloud environment. The appeal is obvious: teams can experiment with different models and choose the capability that fits the application without building the underlying infrastructure themselves.

Choice reduces dependence on a single model. It also creates a portfolio that somebody has to govern.

An API turns a model into an organizational dependency

OpenAI has made ChatGPT and Whisper available through application programming interfaces (APIs). Developers can now add conversational language and speech-to-text capabilities to products without training or hosting the underlying models. The lower cost and simpler integration will accelerate experimentation.

It will also make a third-party model part of more organizations' operating machinery.

A research preview has become a service

OpenAI has introduced ChatGPT Plus, a paid pilot that promises general access during busy periods, faster responses, and priority access to improvements. The price and feature list will receive most of the attention. The more consequential change is in the relationship between the user and the system.

A research preview invites exploration. A paid service creates expectations.

Enterprise AI begins at the platform boundary

Microsoft has made Azure OpenAI Service generally available, giving approved customers access to large generative models through Azure's enterprise infrastructure. The announcement will be read primarily as expanded access to capable models. For organizations deciding whether to build with them, the more important development is the boundary being placed around those models.

Enterprise artificial intelligence (AI) begins where a general capability meets identity, data, security, reliability, cost, and accountability.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.