Skip to content
Article reader Listen + reading controls
LISTEN + READ YOUR WAY

Article reader

Preparing the reader…

0:00 0:00
Reading settings
Text size
100%

The interface should carry more of the context

Google DeepMind is experimenting with an artificial intelligence-enabled mouse pointer that can combine pointing, visual context, and natural language. Instead of describing an object at length or moving material into a separate chat window, a user can indicate “this” or “that” where the work already appears.

The concept addresses a real problem. People should not have to become amateur prompt engineers to communicate context that is already visible on the screen.

Prompting is often a tax on translation

Much of knowledge work is situated. A person sees a paragraph, table, image, record, or map inside an application and understands it in relation to a current goal. A conventional artificial intelligence (AI) interface asks that person to reconstruct the situation in words: identify the object, copy the relevant material, explain the task, and specify where the result belongs.

That translation interrupts flow and discards context. It also favors users who have learned the system's preferred language.

An interface that can combine gesture, location, speech, and application state may reduce that burden. “Compare these” becomes meaningful because the pointing action supplies the referents. The interaction begins to resemble the shorthand people use when they share a physical workspace.

Shared context can be wrong

Human shorthand works because participants continually repair misunderstanding. A colleague can ask which table, point back, or notice hesitation. An AI interface needs equivalent repair mechanisms.

The system may infer the wrong object, include hidden content, miss a selection boundary, or misunderstand whether “move this” means copy, reorganize, or permanently modify. The more seamless the interaction feels, the easier it may be for users to overlook an incorrect inference.

Clark and Brennan's work on grounding in communication explains that shared understanding is established through evidence exchanged by participants. An AI-enabled pointer should show what it believes “this” refers to, preview consequential actions, and make correction inexpensive.

Context is also an access decision

An interface that follows the user across applications may encounter sensitive material. Visual access does not automatically imply permission to ingest, retain, combine, or transmit everything on the screen.

Product teams need clear boundaries:

  • Which application regions are in scope?
  • How does the user know what the system can see?
  • Is context processed temporarily or retained?
  • Can content from one workspace influence another?
  • Which actions require confirmation?
  • How can the user inspect and correct the system's interpretation?

These questions belong in the interaction design, not only in settings or policy text. A visible selection boundary can communicate both context and privacy more effectively than a general disclosure.

Design for intent, evidence, and repair

The most promising aspect of the pointer is not the device itself. It is the principle that AI should meet people inside their work and use natural signals of intent.

Organizations should carry that principle into enterprise tools. Let users point to the source material. Preserve where an answer came from. Show how a transformation will change the artifact. Keep the original available. Make undo reliable. Ask for clarification when ambiguity is consequential, not merely when the model is uncertain in the abstract.

Human-centered AI is sometimes described as adapting technology to people. The deeper standard is mutual intelligibility: the person can express intent without unnatural ceremony, and the system makes its interpretation visible enough to challenge.

The future interface may require fewer elaborate prompts. It will need more careful grounding. “This” and “that” are powerful precisely because context does so much work. Trust depends on letting the user see which context the system thinks it shares.

Sources and research trail

READER-NEUTRAL SUBSCRIPTION

Follow Field Notes via RSS.

Copy this address into the RSS reader you already use. New notes will appear there automatically—no account, email address, or tracking required.