Skip to content

INDEPENDENT RESEARCH / PRACTICE NOTES

Ideas in motion, not ideas behind glass.

This is my working notebook on AI engineering, strategy, knowledge infrastructure, organizational transformation, and the human systems that determine whether innovation becomes real capability.

The notes range from emerging research questions to practical operating models. Some will become papers, tools, talks, or products. Others are here because thinking improves when it is made visible.

Evidence before theater Systems over slogans Useful, accountable AI

Subscribe via RSS Follow Field Notes in any feed reader

Latest writing

The interface should carry more of the context

Google DeepMind is experimenting with an artificial intelligence-enabled mouse pointer that can combine pointing, visual context, and natural language. Instead of describing an object at length or moving material into a separate chat window, a user can indicate “this” or “that” where the work already appears.

The concept addresses a real problem. People should not have to become amateur prompt engineers to communicate context that is already visible on the screen.

Financial agents will be proven in the exception queue

Anthropic has released agent templates for financial services, including work such as preparing pitchbooks, screening Know Your Customer (KYC) files, reviewing valuations, reconciling ledgers, and supporting the monthly close.

These are not toy tasks. They sit inside governed processes with source systems, deadlines, approvals, materiality judgments, and audit expectations. Their automation will be judged less by the clean case than by what happens when the evidence does not line up.

Autonomy is also a materials problem

The Defense Advanced Research Projects Agency (DARPA) is asking researchers to rethink robotics through physical intelligence: materials and structures that integrate sensing, adaptation, computation, and actuation rather than sending every signal through a centralized processor.

The idea is technically ambitious. It is also a useful corrective to the way artificial intelligence conversations often collapse an entire system into its software.

Verification is now part of the product surface

OpenAI has released Generative Pre-trained Transformer 5.5 (GPT-5.5), describing a model that can carry more of a complex task across coding, research, data analysis, document creation, and software tools. This continuing increase in agentic capability changes the user's job.

When a system produces a paragraph, review can happen at the paragraph. When it completes an hour of work across several applications, review has to cover a chain of actions, transformed data, and consequential choices.

Embodied AI turns perception into authority

Google DeepMind has released a new embodied-reasoning model for robotics, intended to improve spatial reasoning and understanding for machines working in real environments. Better embodied reasoning may help robots interpret gauges, locate objects, understand scenes from multiple views, and plan physical tasks.

That progress narrows the distance between perception and action. It also raises the cost of being confidently wrong.

Enterprise AI scales through the operating model

OpenAI's account of the next phase of enterprise artificial intelligence describes rapid growth in organizational use and increasing demand for agents that can operate across real workflows. The direction is clear: companies are moving beyond isolated conversations toward systems that research, update records, create artifacts, and complete multi-step work.

That movement makes the operating model—not access to the model—the limiting factor.

Evaluator disagreement is information

Google Research has published work asking how many human raters an artificial intelligence benchmark needs. The question matters because many model evaluations rely on people to judge qualities that cannot be reduced to exact matching: helpfulness, factuality, relevance, safety, style, or the quality of an explanation.

Adding raters can improve reliability. But disagreement is not always noise that disappears when the sample grows. Sometimes it is the finding.

Independent evaluation needs a secure place to work

The Center for Artificial Intelligence Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST) has entered a cooperative research and development agreement with OpenMined. The collaboration is intended to advance secure methods for evaluating artificial intelligence systems.

The agreement points at a recurring barrier to credible assurance: evaluators need access to meaningful systems and evidence, while model developers, customers, and government organizations need to protect intellectual property, personal information, security-sensitive data, and operational methods.

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.