These days, everything is a subscription, and even the smallest convenience comes with a $20 price tag. Cloudflare Workers still offers remarkably generous free quotas and a $5 starting price. People call it the patron saint of developers. But behind the gratitude, I hear a simple question: you're giving me all this—what's in it for you?
The free Workers plan includes a hundred thousand requests a day. The paid plan, starting at five dollars a month, includes ten million requests and thirty million CPU milliseconds. For a new utility, an API endpoint, or an internal app with few users, it's hard to complain about that starting price.
The stock chart makes for quite a contrast. On September 30, 2026, Cloudflare's market capitalization was approximately $123.66 billion: five dollars on the sign outside a company worth over a hundred billion. (BTW, I'm bullish on CF for the long term. I bought in the $150s and have held it as a long-term investment.)
What looks “cheap” on the pricing page depends on how the company organizes computing resources. What looks “expensive” in the stock market depends on how much future business investors think the platform can support. They're pricing different things.
The engineering between those prices tends to get skipped. How can Cloudflare charge so little to execute code? Do the numbers still work if users exhaust their quotas?
First, understand what's being sold. Then reach for the calculator.
01|Five dollars is just the start
Imagine building an event registration tool over a weekend.
At first, it has one page and one submission endpoint. A Worker receives the form and returns a result. A few days later, you need to save registrations, so you need a database. Participants need to upload photos, so you need object storage. Successful registrations need notifications, preferably without making users wait. Now you need a queue too.
Once people actually use it, you care whether deployments succeeded, why a request failed, and whether you can roll back an update. What looked like a few dozen lines of code gradually becomes a complete application.
Cloudflare's product catalog covers each stage of that growth. Workers Platform has become a collection of services for running applications, beyond executing functions. (These are business groupings, not the company's financial reporting segments.)
| What the Application Needs to Do | Corresponding Products or Capabilities |
|---|---|
| Receive requests, execute code, publish applications | Workers, Static Assets, Pages, builds and deployments |
| Store files and business data | R2, KV, D1 |
| Manage state, messages, and long-running tasks | Durable Objects, Queues, Workflows, Cron |
| Connect models and retrieve knowledge | Workers AI, AI Gateway, Vectorize, AI Search |
| Run more complete programs or browsers | Containers, Browser Run |
| Host other people's code and maintain production environments | Workers for Platforms, logs, observability, control APIs |
The billing models vary. Database operations, object storage, GPU inference, and CPU execution have different cost structures too. Five dollars doesn't buy unlimited use of the entire catalog.
A developer's first payment might be just five dollars. The platform wants the application to keep storing data, executing tasks, and handling traffic on its infrastructure as it grows from a weekend project into a business.
“Number of developers times five dollars” counts entry-level accounts. It misses the revenue from applications that may use the platform for years.
Workers is the entry point. Behind it is an entire platform of storage, state, queues, AI, and operations.
02|What would this platform be worth on its own?
As of September 30, 2026, Cloudflare's market capitalization was approximately $123.66 billion. The midpoint of the company's full-year revenue guidance was $2.867 billion, putting its market valuation at roughly 43.1 times annual revenue. In the second quarter, gross margin was 71.8% and adjusted operating margin was 13.8%, while the company still reported a GAAP net loss. That high valuation clearly isn't just buying current profits.
Using that valuation as a reference, suppose an independent Workers Platform had $500 million in annual revenue, maintained margins close to its parent after infrastructure and other costs, and received the same revenue multiple. That would give it approximately $359 million in annual gross profit, $69 million in adjusted operating profit, and this valuation:
These revenue figures are assumptions, not disclosed results. Cloudflare doesn't report Workers revenue and profit separately, so twenty billion dollars is a scenario, not a confirmed spin-off price. Under those conditions, though, the platform would be worth tens of billions of dollars.
That still leaves the engineering question: how can it sell compute so cheaply? What does the hardware cost if users exhaust their quotas?
03|Thirty million milliseconds is only eight hours and twenty minutes
The easiest mistake in this calculation is seeing “ten million requests” and immediately picturing a server working nonstop for an entire month.
Workers Standard's base compute allowance is ten million requests and thirty million CPU milliseconds per month. Convert the latter into time actually spent executing on a CPU:
If both the request and CPU allowances happen to be fully consumed at the same time, the average CPU time per request is:
Real traffic isn't that uniform, and the platform needs spare capacity for peaks. Still, ten million requests and eight hours and twenty minutes of billable CPU execution are different quantities.
Cloudflare's twelfth-generation server design, published in 2024, uses a single AMD EPYC 9684X with 96 physical cores, 384GB of memory, two 7.68TB NVMe SSDs, and a dual-port 25GbE network interface. Its article reported approximately 600W of total server power draw at an ambient temperature of about 25°C. I'll use that as a reference, not as a claim that the whole fleet uses this configuration in 2026.
We don't know Cloudflare's actual purchase prices, long-term utilization, or internal depreciation policy. Here, we'll use reasonable assumptions:
| Parameter | Value Used Here | Basis |
|---|---|---|
| Server purchase price | 20,000 USD | Calculation assumption |
| Physical cores | 96 | Public hardware specification |
| Depreciation period | 4 years | Calculation assumption; zero residual value |
| Long-term average CPU utilization | 50% | Calculation assumption |
| Total server power draw | 600W | Public figure used as a fixed model input |
| PUE | 1.3 | Calculation assumption |
| Electricity price | 0.10 USD/kWh | Calculation assumption |
At 365 days per year, four years gives us 35,040 hours. With 96 cores and 50% average utilization, the server can deliver:
Now add electricity. Multiply 600W by a PUE of 1.3 and $0.10 per kilowatt-hour, and the server's facility electricity cost is $0.078 per hour. At 50% utilization, 96 cores deliver 48 effective core-hours each hour, so:
To check whether that result depends on one conveniently chosen set of parameters, vary the assumptions:
| Server Price, Depreciation Period, and Utilization | Hardware Depreciation for the CPU Allowance | Including Modeled Electricity Cost |
|---|---|---|
| 15,000 USD, 4 years, 70% | 0.053 USD | 0.063 USD |
| 20,000 USD, 4 years, 50% | 0.099 USD | 0.113 USD |
| 30,000 USD, 3 years, 20% | 0.495 USD | 0.529 USD |
These scenarios aren't upper and lower bounds on actual costs. Across these assumptions, a fully consumed CPU quota might still cost only a few cents to a few tens of cents in allocated hardware and electricity.
Even Cloudflare's own marginal pricing makes this less mysterious. Beyond the base allowance, CPU costs $0.02 per million milliseconds, equivalent to $0.072 per core-hour. At that rate, thirty million CPU milliseconds costs only $0.60. The five-dollar bill was never just selling CPU time.
That doesn't give us “eleven cents in costs and a five-dollar price.” We've counted hardware and electricity for the CPU time on the customer's bill; other overhead along the request path is still missing.
04|What costs are missing from those eleven cents?
The CPU time metered for a user isn't all the CPU time the system consumes on their behalf.
After a request enters the platform, it still goes through protocol handling, routing, scheduling, metering, and other stages. The official definition of Worker CPU time centers on code execution; waiting for networks, databases, and similar operations isn't included. You can't substitute the customer's metered usage directly for the resource consumption of the whole platform.
Suppose each request incurs another millisecond of CPU overhead outside customer metering. Ten million requests adds 2.7778 core-hours. The total rises from 8.3333 to 11.1111 core-hours, and hardware plus electricity in the baseline model rises from roughly $0.113 to $0.150.
That millisecond is an example, not a Cloudflare measurement.
Network differences can be much larger. The same ten million responses, averaging 10KB each, come to about 100GB in decimal units. At 1MB each, that becomes about 10TB. The CPU quota alone doesn't tell you how much data moved, how long connections stayed open, or how much was sitting in memory.
The hardware allocation above also assumes a reasonably balanced use of the machine's resources. If memory is full while the CPU sits idle, you can't keep spreading costs across theoretical CPU capacity. Data centers, network equipment, storage operations, security, support, and R&D don't disappear just because the calculator says eleven cents.
The calculation challenges the intuition that “ten million requests is so much that the platform must lose money if users exhaust the quota.”
It doesn't prove that every five-dollar account is profitable. It certainly doesn't prove that Cloudflare's complete service cost is eleven cents.
Other architectures can also deliver cheap CPU core-hours if utilization is high enough. Isolates don't have a monopoly on low unit compute costs.
The key is sufficiently high utilization.
A platform hosting tens of thousands of independent applications can spend a lot of resources while no code is running. Many apps have only a few visitors a day, yet their environments occupy memory to stay ready for the next request, which can't wait too long.
Isolates reduce the overhead that adds up across those applications: cold starts, resident memory, fixed process costs, and context switches.
So even including the costs of surrounding infrastructure, bandwidth, colocation, storage, the control plane, and so on, a conservative estimate puts this portion at three times the compute cost:
Under this rough worst-case estimate, fully consuming the $5 Workers plan's CPU allowance costs about $0.35 in infrastructure, for a cost ratio of about 6.8% and a gross margin of about 93.2%.
05|How much isolation does three milliseconds of code need?
If someone executes only a few requests a day, each lasting a few milliseconds, keeping a complete resident environment just for them means fixed overhead will dwarf the business logic. Virtual machines and containers both face this problem: the workload is small; the environment sustaining it may not be.
Virtual machines provide full system isolation. Containers share the host kernel and can house their own processes and runtime environments. These designs solve real problems. We shouldn't praise Workers by portraying other platforms as “starting a new virtual machine for every request.” Real systems reuse instances and apply all kinds of optimizations. Cloudflare's main difference is that it shrinks the unit of application isolation further.
A V8 Isolate is an execution instance with independent state. Objects in one Isolate can't be used directly in another. A host can create multiple Isolates and run different ones across multiple threads, but at most one thread can enter a given Isolate to execute at any moment. An Isolate is neither an OS process nor a dedicated thread.
This lets the platform fit many independent applications into one runtime process. Cloudflare's documentation describes a scale of hundreds or even thousands of Isolates per runtime instance. Adding a small application doesn't require adding a complete independent process environment.
The main saving is fixed overhead for every application, alongside a few milliseconds at startup.
We can express this with a simple memory model:
Wasted memory also drags down CPU utilization. If low-traffic applications fill memory first, the machine can't accept more applications even with spare CPU capacity. Reducing their memory footprint lets more people's scattered requests share a machine.
Workers' 128MB isn't a block of memory reserved for each request, either. It's a memory limit per Isolate, and one Isolate can handle multiple concurrent requests. Multiplying concurrent requests directly by 128MB will distort the hardware estimate.
But “shared runtime” doesn't explain enough. What exactly is shared?
V8 has a specific optimization called Embedded Builtins. Earlier, the machine code for some built-in functions contained address references to objects inside a particular Isolate, making it difficult to share across execution environments. V8 later introduced indirection: rather than hard-coding an Isolate's address, the code uses a register pointing to the current Isolate's root table, plus an offset, to find the object it needs. The common machine code can then be embedded in the binary and shared through operating-system mechanisms.
Each Isolate keeps its own state, but multiple Isolates can share the machine code for builtins instead of keeping separate copies.
V8 developed this optimization; it benefits other V8 runtimes too, beyond Cloudflare.
In its workerd architecture introduction, Cloudflare explains another choice: organize basic APIs as shareable native implementations wherever possible, rather than loading a fresh set of JavaScript implementations into every Isolate. At the same time, run many small services in one process to reduce fixed process costs and unnecessary context switches.
Independent processes can also share some read-only code pages. The difference isn't “everyone else shares nothing; Cloudflare shares everything.” It's how much extra memory each new tenant needs, how much initialization it takes, and how many context switches it causes.
These costs may look insignificant for one application. Multiply them by the number of applications on the platform, and they can change the pricing table.
06|The program is waiting; the machine doesn't have to
Consider another, more common kind of waste.
Suppose an endpoint takes 200 milliseconds to return, but executing code consumes only 3 milliseconds. The rest is spent waiting for a database and external APIs. The user did wait 200 milliseconds. The CPU wasn't busy for that long.
Asynchronous execution lets the runtime handle other requests during that wait. The same Workers instance can continue receiving and handling other requests while waiting for asynchronous operations.
Cloudflare didn't invent this trick. Many runtimes support asynchronous I/O.
With asynchronous execution and lightweight environments, the CPU can handle other work while applications wait, without keeping a heavy independent environment for each one. Combining their short bursts of work keeps the machine busy more consistently.
Cold starts need to be understood in that context too.
If rebuilding an environment is slow, the platform has more reason to keep warm instances alive so the next request gets a fast response. If rebuilding is cheaper, reclaiming an idle environment carries a smaller penalty when it needs to be started again.
In 2020, Cloudflare described loading the relevant Worker as soon as the TLS handshake's ClientHello identifies the hostname, where applicable. That moves part of initialization into the wait for the network handshake. The code still has to load, but preparation can start before the complete user request arrives.
Faster startup improves the user experience and reduces the resources the platform needs to keep ready for requests that haven't arrived yet.
Don't overstate this either: an Isolate won't turn a program that really needs a minute of computation into one that takes a millisecond. It mainly reduces the extra burden of residency, waiting, and scheduling. The computation the business actually needs still has to be done.
Putting many people's code in one process raises a reasonable question: is it safe?
Cloudflare doesn't put every tenant in one huge process. It runs multiple runtime processes, groups tenants by trust level, and uses additional process isolation where needed, with OS-level restrictions such as Linux namespaces and seccomp outside that. One concrete example in its documentation: free-account code isn't scheduled into the same process as enterprise-customer code.
The language runtime provides high-density application isolation, with process and operating-system protections adding boundaries. This layered design shares resources where it can and pays for isolation where it must.
Open-source workerd alone isn't the full security system of hosted Workers either. Cloudflare explicitly warns that running potentially malicious code with it still requires appropriate additional sandboxing.
Using V8 alone doesn't give you Cloudflare's costs, security, or operations. Those depend on the arrangements around it.
07|Why does the database have to be in another room?
Cloudflare's other designs ask a question familiar from Isolates: does this step need to cross a boundary?
Take two services, maintained by different teams, deployed and updated independently. It's easy to assume: two services, so deploy them separately and call between them over HTTP.
But team boundaries and machine boundaries aren't necessarily tied together.
Cloudflare's Service Bindings let one Worker call another without going through a public HTTP address. By default, the two can run on the same server and thread; different placement strategies may change their actual locations. Independent releases are preserved while some unnecessary network round trips disappear.
The same applies to databases.
In SQLite-backed Durable Objects, SQLite is called as a library, in the same process and thread as the application code responsible for that state. A query doesn't have to be packaged into a remote database request and wait for another machine to answer.
Having the data next door doesn't mean writes can ignore reliability, though.
There's an interesting detail here called Output Gates. After an application submits a write, it can continue preparing its response. But before “success” actually goes out, the platform waits for confirmation that the relevant writes have been durably stored. Response construction can overlap with durability confirmation. The successful result won't reach the user before that confirmation completes.
KV makes a different trade. It's designed for workloads with many reads and few writes, organizes reads through multiple cache layers, and accepts eventual consistency. A newly written value isn't promised to be immediately visible everywhere. In return, every operation needn't bear the same global coordination cost.
These designs redistribute the work: who owns which data, which operations can stay local, which results must wait for confirmation, and which uses can tolerate delayed propagation. That tells us more than “make the database faster.”
Avoiding an unnecessary remote call is often cheaper than tuning it to perfection.
08|The savings still take engineering
Stopping at V8, Rust, and SQLite makes it sound as though picking the right technologies is enough to cut costs.
Real engineering is much more tedious.
When Cloudflare introduced its Pingora proxy system in 2022, it described a concrete problem. Connection pools in the old architecture were scattered across worker processes. A request landing in one process could often reuse only that process's existing connections. Another process might have an available connection, but the request couldn't necessarily use it. Adding more worker processes fragmented the pools further and could make reuse worse.
Pingora moved to a multithreaded architecture that shared connections and data more effectively, while reducing some copying, allocation, and language-boundary overhead in the old implementation. Cloudflare reported that under the same production traffic at the time, the new system used approximately 70% less CPU and 67% less memory. That's a comparison between specific old and new proxy systems. It doesn't mean the company's total costs fell by seventy percent, and certainly doesn't mean “switch to Rust and automatically save seventy percent.”
Hardware selection isn't simple either. In Cloudflare's twelfth-generation server tests, two 96-core processors had a production-workload performance gap of approximately 22.5% at the same CPU power setting, due to differences including L3 cache. The better-performing Genoa-X had 1,152MB of L3 cache, three times the other processor's. That result is for the mixed production workload at the time, not a Workers-only benchmark.
Core count matters less if the cores spend their time waiting for memory. Adding threads won't necessarily help either if they can't share connections and data sensibly.
Networking belongs in this ledger too. Cloudflare's services are built on shared network and platform capabilities, giving new products the opportunity to reuse existing traffic entry points, network connectivity, and operations rather than building another complete system from scratch for every launch. This is an economic inference from its architecture. It doesn't mean every product runs on every server.
Direct interconnection can also change traffic costs. Cloudflare has explained publicly that peering reduces some traffic's reliance on paid transit. Ports, circuits, equipment, data centers, and staff still cost money. “Direct interconnection” doesn't translate to “free bandwidth.”
Caching reduces repeated work. Tiered Cache has lower-tier caches ask upper-tier caches for content, contacting the origin only when necessary. That avoids many edge locations independently fetching the same thing from the origin. The content still has to reach the user, but the intermediate path doesn't need to start from scratch every time.
Loading fewer copies of a basic implementation, reusing connections, and moving request data fewer times are hard to sell as an exciting business story. Repeated across enough requests, those savings show up on the income statement.
This engineering relies less on a single conspicuous invention than on someone following a request through the system and asking at each step: why are we still spending money here?
09|Once costs fall, why pass the savings to users?
Low costs explain how Cloudflare can make certain services cheap. Choosing to offer free allowances is a separate business decision.
Once costs fall, a company can keep its old prices or pass some savings to users. Passing on the savings can pay off over a longer customer relationship.
In its public explanation of free website security, Cloudflare mentions both lowering network operating costs and gaining richer information about attack patterns by protecting more websites, helping improve overall security. That's its explanation for the website-security free tier. It can't automatically be applied to every free product.
Developer platforms face another constraint: an application that hasn't yet proved it can succeed is particularly reluctant to pay a lot to get started.
Some projects last a weekend or serve a few dozen people; others become companies. The platform doesn't know which will survive. Cheap trials and low starting prices let more projects get going, with revenue coming from those that grow.
This doesn't require every free user to become a paying customer or every five-dollar account to have exactly the same margin. But it does require the platform to manage costs well: failed and dormant projects can't carry heavy resource costs indefinitely, and successful projects need a sufficiently complete set of services to support them.
The business depends on those earlier engineering choices.
Lightweight execution makes long-tail applications easier to serve, and a complete platform gives them room to grow. Investors are betting that many applications will follow that path.
They may overestimate how often it will happen or how much of the resulting business the platform will retain. Even setting those optimistic expectations aside, the engineering behind five dollars is worth serious study.
10|What we want to bring to users' own machines
Our project, open-compute, is a self-hostable platform compatible with Cloudflare Workers Platform. We want this development model to work on users' own machines too.
Saving users five dollars isn't a good reason to build it.
Applications suited to public cloud can keep using Cloudflare. Another option becomes worth building when there are explicit requirements around deployment location, data control, and infrastructure ownership. A team might need to run applications on its internal network, for example, or deliver software together with its runtime rather than leave every dependency in an external service.
In those cases, the question becomes: can we keep the Workers developer experience without switching programming models and reconnecting the storage and task systems?
Cloudflare has already open-sourced workerd. That gives us an excellent runtime foundation, and Cloudflare deserves explicit credit for it. We haven't reinvented the JavaScript engine.
But a runtime that executes code doesn't mean a platform is ready. Someone still has to organize deployment updates, routing to the right version, state ownership, recovery from failed tasks, and the ways users inspect and operate those resources.
That's the work open-compute does.
Written in Rust, ocd manages ingress, control APIs, scheduling, and runtime lifecycles. Worker code executes in a pinned and validated workerd fork. SQLite manages the instance's authoritative local state. Object contents default to local storage, with S3-compatible storage as an option.
We've organized it as one delivery file, one shared daemon, and clearly isolated instance configuration and data directories. Users don't first need Kubernetes, Redis, a separate database, or a service mesh. “One file” describes delivery; managed workerd subprocesses still exist underneath. What users avoid is a pile of services to install, configure, and look after separately.
It follows the same approach as the earlier engineering choices: set a clear goal and leave out the complexity it doesn't require.
A single-machine deployment doesn't necessarily need a distributed control plane first. Local state doesn't necessarily need a trip across the network just to fit a fashionable architecture diagram.
Self-hosting has limits. open-compute doesn't provide Cloudflare's global edge network or a replicated high-availability cluster. The deployer handles host security, outbound network restrictions, backups, and operations. Taking control of deployment means taking responsibility for those things too.
Compatibility must be demonstrated by tests, not by the name. The repository's current public validation matrix tracks 2,256 stable API members and overloads, comparing requests against real Cloudflare across ten product surfaces: Workers, Cache, KV, D1, R2, Durable Objects, Queues, Vectorize, AI Search, and Observability. The documented scope and single-node differences still matter. These records don't mean every behavioral detail of the hosted platform is already identical.
A binary can't reproduce a business worth tens of billions of dollars. Customers, brand, global network, and operations don't come packaged in software.
What we want to preserve is the part that can be written as software, tested, and reused: lightweight execution environments, familiar interfaces, and enough platform capability to keep applications running.
Five dollars turns out to be a useful starting point for understanding the system. Work backward from that price and you reach CPU time, memory, registers, and connection pools, along with decisions about where data belongs and what the system should promise. Those architectural choices end up as a small number on a developer's bill.
Call it the patron saint of developers if you like. I'd rather look at the engineering behind those prices and see how they hold up.
Our implementation is here:
https://github.com/elliothux/open-compute
open-compute is licensed under Apache-2.0. The bundled workerd fork follows the applicable upstream licenses. Stars, issues, and pull requests are welcome.
References
- Cloudflare Workers Pricing
https://developers.cloudflare.com/workers/platform/pricing/ - Cloudflare Workers Limits
https://developers.cloudflare.com/workers/platform/limits/ - How Workers works
https://developers.cloudflare.com/workers/reference/how-workers-works/ - Workers security model
https://developers.cloudflare.com/workers/reference/security-model/ - Service Bindings
https://developers.cloudflare.com/workers/runtime-apis/bindings/service-bindings/ - Workers KV architecture
https://developers.cloudflare.com/kv/concepts/how-kv-works/ - Cloudflare 2026 Q2 financial results
https://www.cloudflare.net/news/news-details/2026/Cloudflare-Announces-Second-Quarter-2026-Financial-Results/default.aspx - Cloudflare Gen 12 servers
https://blog.cloudflare.com/gen-12-servers/ - Cloud Computing without Containers
https://blog.cloudflare.com/cloud-computing-without-containers/ - Eliminating cold starts with Cloudflare Workers
https://blog.cloudflare.com/eliminating-cold-starts-with-cloudflare-workers/ - workerd: the Open Source Workers runtime
https://blog.cloudflare.com/workerd-open-source-workers-runtime/ - SQLite in Durable Objects
https://blog.cloudflare.com/sqlite-in-durable-objects/ - Pingora architecture
https://blog.cloudflare.com/how-we-built-pingora-the-proxy-that-connects-cloudflare-to-the-internet/ - V8 Embedded Builtins
https://v8.dev/blog/embedded-builtins - V8 Isolate API
https://v8.github.io/api/head/classv8_1_1Isolate.html - Vercel Series F
https://vercel.com/blog/series-f - Supabase Series F
https://supabase.com/blog/series-f - open-compute
https://github.com/elliothux/open-compute
