All articlesZH
This version was translated from the original Chinese by AI.

· Elliot Hu

DeepSeek Harness: Good Framework, Bad Product

DeepSeek Harness: Good Framework, Bad Product cover

After DeepSeek Harness launched, the first thing I looked at was BENCHMARK.md in the repository.

The file has three lines: a title, instructions for installing the Python SDK, and a reminder to use different workspace and session IDs when running benchmarks. There are no results, baseline, or test environment.

Meanwhile, the current snapshot has 219 workspace packages. The base composition contains 78 configuration rows; the web patch adds another 51. The standard agent preset runs to 251 lines. Cordis, the underlying framework, has an 88-page paper covering effects, coeffects, fibers, dependency topology, recovery theorems, and “spatiotemporal composability.” Models, tools, sessions, sandboxes, permissions, persistence, and UI are all plugins. The official summary: Everything is a plugin.

Eighty-eight pages of architectural theory alongside three lines of benchmarks with no results. That was my first impression of DeepSeek Harness.

I call it “engineering self-indulgence” because the priorities are wrong. Cordis has technical substance, and many of its designs have clear foundations. The system goes to enormous lengths to explain its elegance without seriously answering why users should pay for it.

What does Cordis solve?

Cordis deserves an explanation before the criticism turns into “complex means bad.”

Think of a shopping mall that has to stay open. Shops can move in, move out, or switch their power supply, payment system, and membership system at any time. Opening a shop means putting up a sign, connecting the power, registering a checkout counter, and subscribing to announcements. If closing means only removing the owner and shelves, all those connections stay behind. Run the system long enough, and you get dead routes, ghost listeners, and timers nobody cleaned up.

Cordis requires a component to register an inverse whenever it creates an effect. In plain English: when you do something, also spell out how to undo it later. When the component unloads, the runtime executes those undo operations in reverse order.

Coeffects handle dependencies. A component first declares which external capabilities it needs. A shop, for example, opens only when both power and payment services are available. Lose either one, and it pauses. When a new provider takes over, it resumes. The paper's “temporal composability” mainly concerns lifecycles and undo operations; “spatial composability” concerns reorganization after dependencies change.

These are real problems. In a long-running plugin host, unloading is usually harder than loading. Many systems eventually rely on a restart to clear everything out. Cordis brings cleanup, dependency switching, and hot swapping into one protocol. I'll give it credit for that.

Harness's append-only session log is good too. System prompts, user messages, reasoning, tool calls, tool results, context injections—anything the model can see must first enter the event stream. In other words, “log it before it enters the model's head.” Resume, fork, trajectory, telemetry, and UI all derive from that ledger. When debugging, you no longer have to guess where a piece of context came from. Recovery and auditing also share one data source.

After reading all this, I rated Cordis more highly than when I first opened the repository. It takes lifecycle tracking, dependency switching, and hot swapping seriously. It doesn't stop at renaming dependency injection. As a runtime for complex plugin hosts, it has research value and solves real engineering problems.

Then DeepSeek made this runtime the center of a coding agent. It hasn't shown that hot-unloading plugins is the problem coding agents most urgently need solved.

When architecture starts making product decisions

We're all familiar with how coding agents fail today. They lose context, choose the wrong tools, cross permission boundaries, run in circles on long tasks, or break code they'd previously fixed. Users judge the results just as directly: did the task get done, how many tokens did it take, how long did I wait, is the diff easy to review, and can it resume after a crash?

DeepSeek Harness spends enormous design effort on replacing plugins without downtime, migrating fibers when providers change, deciding whether a capability belongs in the host plane, agent plane, or browser plane, whether a service should enter an isolate realm, and whether an upper-layer patch should insert a row or replace the entire config.

Those are valid research questions. That doesn't make them the product's highest priority.

In Harness, adding an ordinary capability often starts with translating it into Cordis's language: define a service, implement a provider, hand it to a consumer. Startup configuration passes through profile, bundle, home patch, and command-line overlay layers. Each capability also needs a place in a plane, a realm, and a plugin tree.

The complexity of traditional code hasn't disappeared. It moved into configuration, service graphs, generated directories, and lifecycle protocols. If that trade buys fewer failures, faster recovery, or better task performance, of course it's worth it. The problem is that DeepSeek has shown us the full bill without showing what we're getting for it.

The repository's minimal preset gives us a comparison. It mainly uses two tools, persistent bash and str_replace_editor, with a one-sentence system prompt and none of standard mode's huge composition. The whole file is 62 lines, many of them comments.

Run the same model on the same tasks and measure how much standard actually improves on minimal. Track failures, retries, startup time, memory, tool schemas, and debugging costs too. Those should have been the first numbers published at launch. Right now, none of them are there.

Early community feedback is equally direct. People want a CLI, a TUI, VS Code, SSH, diff views, and token counts. None of that is as pretty as fiber migration, but it determines how people use the product every day. Users are still looking for the door handle. The project is already studying how to replace the building's load-bearing structure while it's open for business.

DeepSeek can build complex systems. My concern is how easily a team gets pulled toward the problems it's best at solving. When complete abstractions are its strength, requirements start bending to fit them. Architecture gradually decides which product problems deserve attention.

You can read the team's priorities in the architecture: which failures it prevents first, which costs it leaves to plugin authors, and which user frustrations it postpones.

DeepSeek Harness pays close attention to internal order while leaving users' immediate frustrations waiting. That explains why it feels overengineered better than the package count does.

The paper's guarantees are smaller than the hype

The Cordis paper tries to prove that components can recover, continue running, and converge after dependency changes. That's serious work, and the paper is clear about its boundaries. Those boundaries often disappear in retellings, leaving something that sounds like “the system can safely keep rewriting itself.”

Page 56 says the runtime doesn't check whether an inverse actually reverses an effect. The plugin author supplies the undo operation; Cordis stores it, orders it, and calls it. The author is still responsible for getting the undo logic right.

Back to the mall. Cordis can guarantee that the shop's closing procedure runs and that the order isn't scrambled. It can't tell whether the owner's checklist forgot the wires inside the wall. It manages the cleanup procedure. It doesn't verify the result.

Page 68 distinguishes acquisition from emission. Opening a file and obtaining a descriptor, or starting a process and obtaining a handle, gives you resources you can close or reclaim. Bytes already written to disk, network requests already sent, and results submitted to external systems have left Cordis's undoable world. After that, you need file deletion, refunds, or compensating transactions.

State inside Context can roll back. The real world has no universal undo button. If an agent sends an email, deletes cloud data, or submits a bad transaction, elegantly unloading the plugin won't help.

The paper's global recovery, progress, and confluence guarantees also depend on a set of assumptions: effects are independent, dependencies are acyclic, there are finitely many fibers, iteration is bounded, no additional failures occur during execution, and components fully provide the services they declared. Cyclic dependencies can be broken out into integration components, but in general, those connecting components may grow to O(n²). The system mainly connects interfaces by key. Matching names alone can't guarantee that both sides understand the interface the same way.

These assumptions are normal in formal proofs, and the Cordis paper states its scope clearly. When publicity strips that scope away, “recovery is possible under these assumptions” can easily become “the system is inherently safe and can evolve arbitrarily.”

Real plugins have already hit the boundaries. An external plugin can declare a new session event at the type level and write it to the log. But standard persistence doesn't recognize event types outside the repository, and the public interface doesn't let an ordinary event be marked as ignorable. The code compiles, yet the session may become unreadable after a restart. The type system says you can extend it. The persistence protocol doesn't follow through.

I'm skeptical of “no privileged core” too. Breaking the core into plugins leaves the trusted parts spread across the loader, Context proxy, event protocol, configuration generator, and multiple providers. The core is still there, distributed and harder to trace when something goes wrong.

Formal methods can answer “what happens if we accept these assumptions?” They don't decide which problems matter most. Mathematics can define a world and prove things inside it. A product can't define its users' problems out of existence.

Self-modification is still a long way from RSI

The separately provided cordis agent preset invites the most speculation. A model can inspect the current plugin tree, define new host or browser code, attach it directly to the running system, then stop or remove it. The community quickly connected this to RSI: Recursive Self-Improvement.

That connection makes some sense. A system needs to observe and modify itself before it can improve itself. Cordis lets an agent attach new tools, prompts, or listeners without restarting the host, then remove them whenever it wants. That makes self-modification experiments much easier.

But between “changed it” and “improved it,” there's an entire feedback loop. Old and new versions must run on the same tasks, be judged by an independent signal, and be compared against other candidates. The winning version must stick around reliably. To call it recursive, you also have to show that the improved system is better at carrying out the next round of improvement, rather than merely scoring a few extra points by chance once.

The core verbs of the current toolset are inspect, define, run, stop, and undefine. Not evaluate, compare, and select. There is no built-in task suite, A/B comparison, independent evaluator, regression detection, version promotion, or curve showing improvement over multiple rounds.

Dynamic packages exist only in the shared process's memory. They vanish after a restart and aren't automatically promoted into persistent plugins. They may also affect other sessions in the same process. The official guidance says to treat the VM sandbox as having the same risk level as bash access; asynchronous host bodies aren't fully constrained by execution timeouts either.

The cordis preset is opt-in, so its high-privilege risks shouldn't be attributed to the default mode. DeepSeek has built a flexible live prototyping tool that lets a model temporarily reshape its runtime. It's useful for research and could become a building block for RSI experiments. Calling it a validated RSI system takes evidence we don't yet have.

The paper is restrained: its conclusion describes self-evolving agent harnesses only as “a direction worth validating in the future.” Secondhand accounts reverse that relationship. Using agent systems to validate Cordis someday becomes Cordis ushering in RSI now.

There's a convenient trick to this narrative. Automatic retries become self-refinement. Tool generation becomes capability growth. Prompt changes become self-evolution. Add hot-loaded plugins, and recursive self-improvement suddenly seems just around the corner. The higher the concept climbs, the less today's product has to face today's comparisons. It needn't first prove that it's more reliable than existing coding agents. It just needs people to believe it's laying the foundations of future intelligence.

But any coding agent with bash and repository access can modify code, run tests, and restart. The hard part has always been deciding what counts as an improvement. That requires explaining who defines the evaluation criteria, whether the system can game the evaluator, whether gains on one local benchmark hurt other tasks, and whether a success survives across versions.

The improvement in RSI needs an external measure. A system can change on its own, but it can't declare those changes progress. Without an evaluator, easier self-modification may just produce changes more efficiently. DeepSeek Harness should meet the same standard.

Show results before talking about the future

DeepSeek Harness needs ordinary controlled experiments now.

Fix the DeepSeek model and reasoning configuration. Assemble a public, reproducible set of short and long tasks. Run minimal and standard separately. Record success rate, retries, input and output tokens, cache hits, time to first token, total duration, peak memory, and human takeovers. Add a mature system as a baseline.

If standard wins consistently, the plugin system's complexity has bought better task quality. If it wins only on certain long tasks, that tells us where it fits. If the difference is small, the team should reconsider which capabilities belong in the default configuration. If dynamic replacement lowers the cost of recovering from failures, recovery time and failure rate can be measured directly too.

Validating RSI is just as concrete: run candidate changes in isolation, use an independent evaluation signal, compare old and new versions fairly, persist the winners, reproduce improvements over multiple consecutive rounds, and show that the system actually gets more efficient at the next round of improvement. Until that evidence exists, runtime self-modification or live prototyping are more accurate descriptions.

I see Cordis as an original dynamic plugin framework with theoretical backing. DeepSeek has used it to build a substantial coding agent. I'm frustrated that the team still hasn't shown that users' problems call for this much complexity.

Engineers sometimes get trapped inside an overly complete theory. The more internally consistent the system becomes, the easier it is to dismiss external counterevidence as “the product isn't finished yet.” The prettier the abstraction, the easier it is to attribute user discomfort to “not understanding the design yet.”

That's how DeepSeek Harness feels to me right now.

I still call it an elaborate exercise in engineering self-indulgence. The engineering deserves “elaborate.” But the evidence so far mainly proves that Cordis can explain Cordis, which is where the self-indulgence comes in.

Put minimal, standard, and an external baseline in the same table. Show the tasks, success rates, costs, and failure modes. That table should arrive before the next architecture paper.