All articlesZH
This version was translated from the original Chinese by AI.

· Elliot Hu

I Open-Sourced Cloudflare

I Open-Sourced Cloudflare cover

Agents have been all the rage these past two years.

The grand infrastructure plans the industry's architecture gurus draw up for agents often strike me as absurd. The entire industry seems to be bringing its old habit of infrastructure inflation straight into the AI era.

Today's “standard textbook approach” goes like this:
Need a coding agent to execute code, run a test, or call an API tool?
Spin up a full-featured Linux container sandbox for every agent, sometimes even every subtask.

And then? To babysit those sandbox containers that might live for only 3 seconds, the team has to maintain a massive Kubernetes cluster, complex network namespace isolation, prewarmed image caches, state cleanup jobs, and highly available Redis and persistent storage on top of it all…

The agent's actual task might be a single line of jq to filter a field, or a 30-line JavaScript script. Meanwhile, the infrastructure cleaning up after it consumes several GB of memory and spends tens of seconds pulling images and starting up.

What an outrageous waste of compute and time.

Cloudflare recently published an article called “Your agent needs a computer, not a container.” It argues that in real agent workflows, fewer than 10% of scenarios require heavyweight containers. For the rest—data processing, toolchain calls, code generation, and session control—lighter V8 Isolates are the right answer.

Give the agent a lightweight isolated environment for thinking and executing tools, using tens of MB of memory and a cold start of a few milliseconds. Attach heavier primitives when it needs native Linux programs. That should scale well.

But there's a very real catch: extreme vendor lock-in. Want that experience? You have to weld your company's core data and the agent's entire call chain to Cloudflare's public cloud.

open-compute project overview

To break that monopoly and escape infrastructure inflation, I wrote open-compute, licensed under Apache-2.0.

I built ocd, a single Rust binary, to bring the entire Cloudflare Workers platform onto private hardware. It took a beating from real B2B workloads in my company's production environment before I decided to open-source it.

Imagine a serverless platform with a complete runtime, dynamic scheduling, persistent storage, message queues, state machines for long-running transactions, and even AI search. The whole platform uses only about 60 MB of resident memory—less than a random Chrome tab on your desktop.

It runs on your own server with one command, in under a minute. No K8s, Docker daemon, external database, or other external dependencies.

1. A sweet deal—and a cloud bill you can't escape

Developers who know Cloudflare probably think of it as their patron saint.

CF's offering is about as good as the industry gets today:

  • Need to store an agent's state or session? Lightweight KV.
  • Need structured queries? D1.
  • Large files or Git repositories? R2.
  • Decouple asynchronous tasks? Queues.
  • The real killer features are Durable Objects (DO) and Workflows: a stateful “singleton brain” for agents, and state machines for long-running transactions with persistent checkpoint recovery.

No processes to provision. No networks to configure. It starts on demand with each request, with cold starts measured in milliseconds. For developers and agent architects, that experience is in a different league.

But there's no free lunch. This walled garden has an electric fence.

The public cloud follows a familiar playbook: get you hooked, then make leaving expensive.
Make getting started cheap or free, until developers are used to the convenience and their agents depend on proprietary APIs at every turn: KV, D1, DO. Once the business reaches real scale and thousands of agents are being scheduled concurrently, cross-region network fees, request fees, and storage egress fees all take their share on the month-end bill.

Even harder to get around are data compliance requirements for enterprise deployments.

Over the past few years, my team has worked on-site delivering B2B systems as FDEs. Finance, government and enterprise clients, healthcare, and privacy-conscious international customers: the moment they hear “core business logic and agent data will pass through an overseas public-cloud edge,” the compliance department vetoes the proposal on the spot.

Teams have little choice but to build their own stack on the internal network: K8s + MinIO + Redis + Knative… After all that work, they've traded “serverless heaven” for “full-stack operations hell.”

Cloudflare has open-sourced the underlying workerd, but a runtime (execution engine) isn't a platform. Experienced engineers know the difference.

It's like a manufacturer open-sourcing a spectacular sports-car engine while keeping the chassis, transmission, dashboard, and wheels—multitenant routing, state persistence, lifecycle scheduling, API control plane—in the factory. A V8 box that executes JavaScript alone can't get you on the road.

If nobody was building that chassis, I'd build the whole thing and give it back to the open-source community under Apache-2.0.

2. Architecture: three separate layers on one machine

The worst architectural mistake is overengineering without regard for the actual scale of the business.

In open-compute, I went in the opposite direction: fold distributed paths that often span dozens of microservices into a single process.

Traditional self-hosted cluster             open-compute
(microservice sprawl)                       (folded down to the essentials)
─────────────────────────────              ──────────────────────────────
API Gateway / Ingress                      ┌───────────────────────┐
K8s Scheduler / Controllers                │                       │
Multi-tenant Control Plane  ══════>         │      ocd (1 bin)      │
Redis / Valkey cluster                     │   Rust Async Host     │
PostgreSQL / TiDB                          │                       │
Heavyweight operators                     └───────────────────────┘
                                           Embedded SQLite + local storage

The Rust daemon ocd (Open Compute Daemon) manages the platform in three layers:

  1. Ingress & APIs: ocd is the only public network listener. The Dashboard, CLI, Cloudflare v4 API, and Git Smart HTTP all enter here. Authentication, authorization, multitenant routing, and trusted identity propagation are handled centrally.
  2. Deployment & Resources: maintains every Worker version declaration, environment variable, and Binding topology. Deployment records strictly determine which capabilities a tenant can call.
  3. Coordination & Runtime Products: provides unified scheduling for Queues, Cron, Alarms, Workflows, and AI Search. Handles task claiming, lease renewal, retries, and recovery after crashes.

3. How the platform runs in 60 MB of resident memory

Many people's first reaction is: How could you bring the whole Cloudflare stack down here and use only about 60 MB of resident memory? Surely this is some bare-bones demo toy?

Getting that footprint down for a production system means being ruthless about low-level resource overhead.

open-compute memory usage

1. Reaching 60 MB by cutting network overhead and runtime waste

Where does all the memory go in a traditional serverless architecture?

  • A Node.js gateway process: 150 MB to start.
  • A Java/Go control-plane service: another 300 MB.
  • A local Redis instance and MinIO daemon: several hundred more MB.
  • Add sidecars and communication proxies, and before the system has handled a single request, thousands of OS threads and garbage collectors are already roaring in the background.

In open-compute:

  • One native machine-code binary: ocd is built on Tokio's asynchronous multithreaded runtime and Axum/Hyper. Full LTO (link-time optimization), codegen-units = 1, panic = “abort,” all debug symbols stripped. No virtual machine, no interpreter, no GC threads repeatedly scanning memory.
  • Zero-copy streaming proxy: request bodies received from network sockets are forwarded straight to the execution isolates as raw byte streams. The host never buffers the entire payload in memory. Even uploading several GB of model files barely disturbs host memory usage.
  • Function calls instead of network RPC: authoritative platform metadata and scheduling state live directly in embedded SQLite, in WAL mode. Looking up a route or claiming a task becomes a local in-process function call, eliminating hundreds of MB of heap allocated for network buffers and protocol deserialization.

2. State must outlive the process: persisting Workflows and long-running tasks

The hardest problem in providing an execution environment for AI agents is this: tasks live far longer than any single process.

An agent's planned workflow might start by calling a tool to look up data, then wait for human approval or an external webhook hours later. If the state lives in memory, or even on the function call stack, the moment the process exits or the machine restarts, everything is gone.

open-compute's coordination layer follows one rule: a waiting task shouldn't tie up an execution process.

Workflows uses a state machine:


           [ Step 1: Execute an external tool ]
                          │
                          ▼
    [ Persist step result outside the transaction ] ────► Process sleeps/is reclaimed immediately
                          │
    (Seconds/hours later: a follow-up event or resumption)
                          │
                          ▼
         [ Validate lease / attempt / generation ]
                          │
                          ▼
      [ Reuse historical results, replay subsequent steps ]

Execution follows these rules:

  1. Claim the task and bind a generation identity: record a unique lease, attempt, and generation.
  2. Execute business logic outside database transactions: never let slow or suspended external calls tie up the database connection pool.
  3. Commit with atomic validation: verify that the current generation identity is valid. If it is, persist the result. If an executor stalls, its lease expires, and another node takes over, the old executor's commit will be rejected even if it wakes up and returns later. Split-brain overwrites are prevented outright.

Even after a power failure and restart, scheduler.sqlite can audit the state in milliseconds and resume precisely from the last persisted checkpoint.

3. Separating control data, scheduling state, and files

Putting everything in one big database can sink a monolith. open-compute separates its data:

  • Physically separate control and scheduling data: control.sqlite handles tenant metadata and the resource catalog. scheduler.sqlite handles high-frequency task claiming, retries, and heartbeat leases. They stay out of each other's way, cutting lock contention dramatically.
  • No pointless MinIO proxy: Worker bundles, static assets, R2 storage, persistent snapshots, and AI Search source files all default to Local Direct—direct writes to the local file system. No nonsense like starting a local HTTP object-storage service just to speak the S3 protocol. Removing that detour through the network stack eliminates an entire complex service's memory footprint and delivers raw NVMe disk throughput.

4. Dynamic Cap’n Proto compilation and a “zero-trust” supervisor

ocd dynamically compiles a standard Cap’n Proto binary topology to connect to workerd, rather than using a simple CLI pipe.

  • Strict loopback isolation: ocd launches workerd subprocesses as their supervising parent, exposing communication ports only on the local loopback network.
  • Credentials bound to lifecycle generations (Per-generation Token): privileged credentials used for internal platform control injection never appear in argv startup arguments, environment variables, logs, or system metrics. Even if tenant code escapes its memory boundary and inspects /proc, it cannot intercept the platform's privileged keys.
  • Deterministic recovery throughout the lifecycle: dedicated process-group supervision, bounded log-stream capture, graceful drain, and forced reclamation prevent orphaned processes from getting stuck.

5. Sandboxing document parsing and AI Search

Beyond frequent lightweight code execution, AI workloads have another, much heavier category: PDF/Word parsing, OCR, and local embedding.

Force those computations into the main executor, and memory spikes can take down the whole gateway. open-compute introduces a separate supervised parser execution layer:

  • Document parsing runs in independent, short-lived subprocesses with hard timeouts and bounded memory.
  • Only after parsing completes do the extracted text chunks flow into the derived indexing system.
  • One-way data ownership: business source files in R2 remain the sole originals. AI Search stores only the derived index vectors. Deleting a search index never accidentally deletes the business source documents, and unchanged source files are never parsed again unnecessarily.

4. Testing compatibility across 2,203 API members

We measure compatibility with tests, running the same test cases and fixtures against both open-compute and the real Cloudflare. A claim alone isn't enough.

Dimension / ModuleCompatibilityEngineering Implementation
Total API Coverage2,203 members100% method coverage within the declared core capability scope
Core Storage and Compute100%Native alignment with KV, D1, R2, Durable Objects, Queues, Workflows, Cron, and more
Modern Full-Stack Frameworks100%Next.js 16 builds based on OpenNext run without modification, exactly as on the official edge
Wrangler CLI Toolchain95%Native support for ocd wrangler deploy and real-time logs through ocd wrangler tail

Bring your existing Workers project over and run it without changing a single line of code.

open-compute compatibility matrix

Cloudflare's API surface is enormous and complex. If you find an unsupported API or encounter another problem, please file an issue on GitHub. We're iterating quickly and take every piece of feedback seriously.

5. Giving developers control of their compute

The software industry has always moved between consolidation and fragmentation.

We moved from our own data centers to centralized public clouds and enjoyed the elegant developer experience of cloud serverless. Now we're also deeply entangled in technical lock-in, expensive “cloud rent,” and constantly expanding infrastructure complexity.

Look at today's hardware: a single physical server can have dozens of CPU cores, hundreds of GB of memory, and several GB/s of NVMe throughput. For the vast majority of enterprises, business toolchains, and AI agents, we haven't come close to exhausting what one machine can do.

Blindly chasing bloated distributed clusters often just hides lazy architectural design.

Fold the unnecessary distributed systems and microservices back onto one high-performance machine, using lightweight execution environments built with Rust + V8 Isolates. One machine can support a complete, fast edge-computing platform while keeping full control of the data.

6. Quick start: run it on your own hardware in 1 minute

The full code is open source under Apache-2.0. An ordinary development machine, a dusty mini PC, or a local server is enough to try it:

1. Start the platform core

Install the ocd core engine: curl -fsSL https://open-compute.dev/install.sh | sh

Set up and start an instance: ocd setup --yes

2. Develop and deploy your first Worker

export default {
    fetch(request: Request): Response {
        return Response.json({
            message: "hello from open-compute",
            path: new URL(request.url).pathname,
        });
    },
} satisfies ExportedHandler;
{
  "$schema": "node_modules/wrangler/config-schema.json",
  "name": "hello-worker",
  "main": "src/index.ts",
  "compatibility_date": "2026-09-08",
  "workers_dev": false
}

Develop: wrangler dev

Deploy: ocd wrangler deploy

No complicated compilers or extra dependencies to install on the host. Packaging, validation, and publishing finish instantly.

The repository and documentation have the source code and full technical details.

7. Agents will bring private cloud back

For the past decade-plus, serverless's most successful story has been: you don't need to own infrastructure. Let AWS handle the servers. Let Cloudflare handle the runtime. You handle the business. In the web era, that was a very good trade.

Agents persist, with state, memory, and permissions. An agent changes a file today and comes back tomorrow to continue. It might run a workflow for hours, with ongoing access to enterprise databases, codebases, internal APIs, communication systems, and production environments. It looks less like an HTTP request and more like a digital worker living inside the company.

I think the next few years will bring ownership back into the discussion, after a decade in which it sounded outdated. Who owns the agent's environment, state, runtime, and computer?

Cloudflare recently said: Your agent needs a computer, not a container. They're right. They're working out what that computer should look like. open-compute wants to ask one more question: why can't that computer actually belong to the company itself? That's what we're trying to do.

Bring the serverless model back to your own machines.

A binary. A data directory. Your code. Your data. Your agents. Your machines. That's open-compute.

We didn't open-source it because it isn't valuable. Quite the opposite. We believe the infrastructure that really matters in the agent era should first become a foundation everyone can inspect, own, and build on.

8. Why open-source it?

Buying GitHub stars is practically an industry now. Open source increasingly looks like another business built on chasing attention. I still stubbornly believe there should be room for honest work: tackle the hard problems, solve something real, and share it with everyone.

Thanks for reading this far. If it helps you, I hope you'll give it a star.

Website: https://open-compute.dev
Github: https://github.com/elliothux/open-compute