All articlesZH
This version was translated from the original Chinese by AI.

· Elliot Hu

We Built the Ultimate Form of Apple Intelligence

We Built the Ultimate Form of Apple Intelligence cover

Siri launched in 2011, fifteen years ago. Models have gone from recognizing speech to writing code, calling tools, and running agents autonomously. They still don't know what's happening on your computer.

Lately, I've become more convinced that the bottleneck for the next generation of personal AI may be the missing layer between the operating system and the model, rather than the model itself.

Today's AI gets its context through manual ETL

Most AI products still work much the same way underneath.

The user copies in web pages, uploads files, pastes chat histories, then describes the background all over again in a prompt. The model starts reasoning from that manually assembled context. It's manual ETL + LLM.

MCP, Skills, and Tool Calling answer the question of what AI can call. They don't answer what it should look at right now, why it should call something, or how far the current task has progressed.

Those are completely different problems.

image

An agent can have 100 tools. But if it doesn't know that you just switched from Excel to the browser to verify a number, or that the PDF on your desktop and the message you just received in Lark belong to the same task, those 100 tools are mostly just an arsenal.

A full arsenal. No idea what to aim at.

Turning work into state

When we built Lynx, “an even stronger agent” wasn't our starting point.

We started underneath it: how can a machine keep track of the constantly changing work on a computer as context?

Screenshots alone don't give you that context.

You need to know which app owns a window, which tab is open, and which object the user is working on. Are the file and the web page related? What happened before and after? Which pieces belong to the same task? Reconstructing all of that from screenshots, OCR, and a VLM every time is expensive, and you lose continuity between states.

A screen is an image. Work is a continuous stream of state.

That's why so many Computer Use demos look impressive but turn brittle over an eight-hour workday. On a real computer, windows move, pages refresh, files change versions, and users intervene halfway through. A single task crosses five or six applications. A benchmark doesn't capture that mess.

Being able to click a button is an execution problem.

An agent also needs to know why it's clicking the button and what changed afterward.

Why a voice interface isn't enough

We used to think of Siri as a voice interface. So for years, people kept upgrading its recognition, language understanding, and model capabilities.

I think “voice assistant” is too narrow a definition for where Siri should end up.

Siri should be a layer of ambient intelligence inside the operating system, understanding the current state of your work within the permissions you've granted. That doesn't require constant screen recording or monitoring. When you say “that thing from earlier,” “organize these together,” or “continue where we left off,” the system should know what you're referring to.

image

Language expresses intent. Context determines what the sentence actually means.

Without context, “send this to him” is meaningless. With context, it may already specify the current file, a recent contact, the preceding discussion, and the channel to send it through.

Context changes how much information that same sentence carries.

Building that layer in Lynx

Lynx organizes the OS, applications, windows, web pages, files, and interaction events into a continuously updated picture of the user's work and what it means. Models and agents work from that picture.

Planning, Tool Calling, permission confirmation, and execution still matter downstream. We're increasingly convinced, though, that context quality may set the limit on what an agent can do.

Misread what's happening, and even the smartest model will confidently head in the wrong direction.

That's what makes it so painful to build. Beyond hooking up APIs, you're fighting decades of baggage in real operating systems. Accessibility, browsers, native applications, file systems, permissions: every layer has its own quirks.

It's messy work, and a lot of fun.

We're nearly ready to let more people try Lynx. There will be bugs; real people's computers are much messier than any test set we can build ourselves.

image

If you'd like early access, head over to lynxai.work.

We'll keep testing against the real world.

I want Siri to understand why I'm asking a question, beyond finally being able to answer more of them.