The Cursor Agent SDK as a Subagent in DeepSeek Harness
DeepSeek Harness does not ship subagents as a feature. It ships them as a seam: a provider interface, a registry, and one delegation tool that knows nothing about what sits behind it. That is why the Cursor Agent SDK can be plugged in as a subagent backend instead of a second, Cursor-shaped tool — and why “which model does the coding” becomes a configuration line rather than an architecture decision.
TL;DR
A session that names the cursor-sdk provider gets one extra tool, subagent_cursor. The parent agent — running on a DeepSeek model — delegates a coding task with a workspace; a Cursor Agent on Composer 2.5 executes it in its own process, edits files, runs commands, and reports back through the ordinary subagent result contract. What that buys you: model routing per task instead of per product, delegation onto subscription capacity you already pay for, a child that cannot pollute the parent's context, cancellation that actually kills a process, and a delegated environment that inherits nothing it was not explicitly granted.
LogNroll Team
Developer Experience
This article is written from the seam, not from the marketing page: a working subagent provider that runs the Cursor Agent SDK behind DeepSeek Harness's generic delegation interface, plus the operational details that only show up in a container. Two things it deliberately does not contain: credentials, and anything identifying the specific repositories and hosts it was validated against. The failure modes are real; the names are not.
A provider is not a second tool
The interesting part of this integration is not that a Cursor agent can edit files. It is where it plugs in. DeepSeek Harness exposes subagents as an interface plus a registry: a plugin implements a provider, registers it on ctx.subagents under a name, and from that moment a generic delegation tool can be pointed at that name. The tool layer asks the registry for cursor-sdk and delegates. It contains no Cursor-specific code at all.
The alternative — writing a cursor tool — would push Cursor's vocabulary into the model-facing layer: a second subagent-shaped tool with its own parameters, its own result rendering, and its own opinion about what a failed delegation looks like. Registering as a provider keeps exactly one delegation surface. Selecting Cursor becomes a configuration line, not a different call shape — and a session that switches backends does not have to relearn how to delegate.
DeepSeek Harness session (DeepSeek model)
|
| subagent_cursor(task, cwd)
v
dsh-tool-subagent <- provider-agnostic
|
| ctx.subagents -> 'cursor-sdk'
v
Cursor SDK provider <- this integration
|
| subprocess: cursor-sdk (Python) -> bridge
v
Cursor Agent, Composer 2.5 -> edits the workspaceTwo consequences follow from that shape. The provider can be swapped without touching a prompt, and a second provider can exist beside it — an in-process child for cheap reconnaissance, a Cursor Agent for repository-scale work — with both reached through the same tool call.
What a delegation actually does
The parent calls one tool
A session that mounted the provider has a subagent tool — subagent_cursor — and nothing Cursor-specific in its prompt. The call carries a task and a working directory, and that is the whole vocabulary.
Configuration is validated before anything is spent
Credential present, workspace resolvable, model permitted, provider enabled. A failure here rejects the start: no child is published, no bridge is launched, nothing is billed.
The prompt is composed, not forwarded
The parent’s conversation is not shipped across the boundary. The provider builds a fixed task statement around the caller’s text, the workspace, any explicitly named requirements and context, and a reporting contract.
A driver subprocess launches the bridge
The Cursor SDK is a Python package, so the transport is a process boundary. One agent is created with an explicit model selection: Composer 2.5, standard request path, never a fallback.
Progress arrives as counters
While the child works, the parent gets heartbeats — elapsed time, tool calls, characters of output. Not the child’s transcript: a number, so a long delegation cannot flood the parent’s context.
The result crosses back as an ordinary subagent result
A stop reason and output blocks. The delegating layer does not know what executed the task, which is exactly what makes the seam reusable.
Step three is the one people skip, and the one that decides whether the feature is pleasant or infuriating. The delegated agent has no conversational history, no idea a harness sent it, and no idea what to report. So the prompt is assembled into fixed sections:
You are a delegated coding agent running inside DeepSeek Harness. DeepSeek Harness owns orchestration; you own the execution of the task below. Work only inside the working directory given here. Task: <the caller's task, verbatim> Working directory: /abs/path/to/the/workspace Requirements: - <explicitly named constraint> Context provided by the delegating agent: <explicitly serialized context, if any> When finished, report: - work performed - files changed - tests/checks executed - unresolved issues
The last block is not decoration. A parent that receives free-form prose has to guess what happened; a parent that receives work performed, files changed, checks run, unresolved issues can act on it, quote it, or refuse it. And the context block is deliberately opt-in: a caller that needs a fact to cross the boundary names it at the call site, where it is reviewable, rather than relying on whatever happened to be in scope.
Why the transport is a subprocess at all
The Cursor SDK is a Python package, and a Node plugin installed through a profile symlink cannot resolve bare specifiers outside its own tree — so there is no in-process route to it. The process boundary turned out to be an advantage rather than a compromise: a run becomes something a signal can reach, the child environment becomes exactly the variables it was handed, and the Node half can be tested against a stub that speaks the same JSON, with no credential and no network.
The opportunities, ranked by how much they change your week
Everything below follows from the seam existing. None of it requires the parent to know what a Cursor agent is.
Model routing per task, not per product
The orchestrator stays on a DeepSeek model, which is what a long planning-and-tool-use conversation wants. The repository-scale edit — the 40-file rename, the failing test suite, the dependency bump — goes out to a Cursor Agent. You stop choosing one model for the whole team and start choosing one per delegation: a read-only investigation stays in-process, a multi-file refactor leaves.
Delegation onto capacity you already pay for
If the team already has a Cursor subscription, the marginal cost of a delegated task is drawn from a pool that exists whether or not you use it. The provider records tokens, elapsed time and Cursor’s own dollar figures per run, so “should this go to Composer?” becomes a measurement instead of a preference.
Context isolation, and real concurrency
A Cursor agent greps, reads, edits and runs tests for ten minutes. None of that enters the parent’s transcript: the parent’s window grows by one result, not by a thousand lines of tool output. And because the child is a separate process, the parent can keep reasoning while it runs — or hold several delegations in flight.
A repository agent you did not have to build
Repo-wide search, file edits, a shell, running the test suite — the Cursor Agent arrives with its own tool set and its own view of the workspace. The work on this side of the seam is a transport and a policy, not a toolchain.
A second opinion from another model family
Once delegation exists, asking a different family to review a diff is nearly free to add: same tool, different sibling. Correlated blind spots are the main failure mode of self-review, and cross-family review is the cheapest structural fix — one author, one reviewer, two vendors.
Unattended runs that fail honestly
Configuration failures reject before anything is spent; runtime failures settle as an error and are never papered over by silently trying another model. Cancellation has three layers, up to a kill signal for a bridge that ignores a polite one. A stuck agent cannot pin a session open forever.
The money question is genuinely ambiguous
This one deserves its own heading, because a subscription makes “cost” mean three different things and any tool that picks one silently is lying to you in a helpful tone. A delegated run is recorded with tokens, elapsed time and Cursor's own dollar figures — from the service, not estimated from token rates — under three labels:
undiscounted API rate
What the run would cost at list API prices. This is the figure that draws down a subscription’s included pool, so it is the one that answers “how close am I to the ceiling”.
actually billed
What the invoice says. While a plan covers usage this is genuinely zero — reporting only this number would claim every run was free.
remaining in the pool
The included pool minus the undiscounted figure. Worth stating plainly: on the $200 Ultra plan the included pool is $400. The pool is not the price.
Coverage is stated too, because cost is eventually consistent: a run that ended seconds ago may not be priced yet, and that is reported as unknown, never as zero. The aggregate says out loud how many runs it could price, so a partial figure reads as a floor rather than a total:
Cursor SDK usage - month
plan : Ultra ($200/mo, $400 included API)
runs : 3 (completed: 3)
tokens : 37000 (in 25000, out 6000, cache read 5000, cache write 1000)
time : 1.4m total, 28.3s average
pool drawn down : $0.61 (undiscounted API rate)
billed : $0.00 (0 while the subscription covers usage)
remaining : $399.40 (0.15% used)
warning: cost reported for 2/3 runs - the figures above
are a floor, not a totalThe projection is explicitly linear and says so, because usage is bursty: it answers “where does this land if the rest of the period looks like its start” and nothing more. A period with no priced runs projects nothing rather than $0.00 — zero would assert a fact nobody established. That is the difference between a dashboard and a measurement.
If one model is mandatory, enforce it in four places
This integration fixes the model: Composer 2.5, and nothing else. “Fixed” is easy to write once and easy to erode, so the rule lives in four independent places — one lock is one bug away from being no lock at all.
At plugin load: a profile that configures any model other than composer-2.5 fails to activate.
At request mapping: the selection is constructed as Composer 2.5 with fast=false, always, so a service-side default can never decide.
At the provider boundary: a caller that names another model — or a fast variant, or auto — is refused rather than quietly downgraded.
At the wire: the Python driver re-asserts the same rule before it launches a bridge, so a mismatch is an invalid_model error and no bridge ever starts.
Two subtleties are worth stealing even if you never use Cursor. First, Fast Mode is a parameter of a model, not a mode: omitting it lets the service apply its own default, which is automatic selection wearing a different hat. So the parameter is always sent explicitly. Second, the comparison is an exact string match, not a prefix match — a prefix rule would cheerfully accept composer-2.5-nano the day somebody ships it.
The payoff is a property you can state in one sentence: a failure is a failure, not a downgrade. If the model is unavailable, the run fails with a clear message. It is never quietly resolved by reaching for something cheaper, which is the behaviour that makes AI infrastructure impossible to reason about six months later.
Where the boundary is drawn
Delegation means handing a task — and a shell — to a process you do not control. The interesting engineering is in what that process does not get.
The workspace is preserved or the delegation fails
There is no fallback to the plugin’s own directory. An agent that silently edited the wrong repository would look like success — the worst possible failure — so an unresolvable working directory is an error.
The child inherits almost nothing
The delegated process gets path, home, temp directory and locale, plus the one Cursor variable and anything named explicitly. Not the parent’s environment — which routinely holds ticketing tokens, database URLs and cloud credentials.
The credential is never in a prompt or a log line
A key is not a configuration value either: a profile patch layer is committed and shared, so a key placed there would be published with it. It is read from the environment at delegation time, and only a masked length-and-prefix form is ever logged.
Every message crossing back is redacted
Assigned keys in bare and JSON-quoted form, bearer tokens, PEM blocks and opaque long tokens are stripped from errors before the parent sees them. Request ids are sanitized before they are used as directory names, so a slash-dot-slash in an id cannot escape the state root.
The ledger holds counts, not prose
One append-only record per run: tokens, elapsed time, dollars, status. Not the prompt, not the agent’s output. A ledger is the kind of file that gets copied around, and a delegated prompt can quote a ticket or a customer’s data.
Structured logs, fixed field set
One line per lifecycle point with a closed set of fields and no arbitrary pass-through. “No prompt or credential in a log line” is structural rather than a rule somebody has to remember.
The failure modes that cost the most time
Every one of these was found by running it, not by reading a document. Three of the five produce an error message that points at the wrong problem — which is exactly why they are worth writing down.
A read-only home that reads like a credential problem
EACCES on the agent store
The bridge keeps its per-workspace agent store under the home directory, and the pinned SDK version exposes no override for that root — so home is the only lever. A container whose real home is read-only fails after authenticating, with a permission error that reads like a revoked key and is not one.
The fix is one configuration key pointing the bridge at a writable path on a volume that survives a restart. It is worth setting before the first run rather than after reading the error twice.
A missing key is not a boot failure
unauthenticated on the first delegation
The credential is read at delegation time, not at startup. An absent secret key therefore does not stop the process; it surfaces on the first delegation as an invalid-key error, which reads like a revoked credential rather than a missing one.
Two habits fix this: make the secret reference non-optional so a missing key fails the rollout instead, and probe inside the running container — key present, interpreter can import the SDK, store root writable — before the first delegation spends anything.
The SDK is Python, and the plugin is not
the bridge is 175 MB of somebody else’s Node
A Node plugin cannot import a Python package, and the TypeScript SDK is reachable only through the Python bridge’s bundled modules. The transport is therefore a subprocess, and the wheel has to be importable by the interpreter that driver launches.
It is a large download — most of it the bridge’s own bundled Node runtime — so it is unpacked once onto a persistent volume and put on the interpreter path, rather than committed to a repository or baked into an image that has no package manager to install it with.
Mounting the provider twice aborts the boot
a subagent provider named ... is already registered
A global overlay that mounts the provider with its full configuration, plus a profile patch that also lists it in an agent roster, produces two activation rows. The second one aborts startup and the Web GUI never serves.
The trap is that the same duplication one row above is harmless: a repeated tool plugin id is a plain configuration merge, while a provider that claims a name in the subagent registry cannot be mounted twice. The diagnosis misleads too — the harness reports a configuration failure and names the plugin, so the natural move is to edit the plugin, which is not where the fault is.
Model availability is a runtime answer
the model is unavailable for this account
In the pinned SDK version the model catalog is a cloud call that requires authentication, so “is Composer 2.5 available to this account?” cannot be answered at startup. Configuration is validated at load; availability is discovered on the first real delegation.
It is reported as a failed run with a clear message, and never resolved by trying another model — because a silent substitution is precisely the outcome the whole model policy exists to prevent.
What it deliberately does not do
A capability that is advertised and then ignored is worse than one that is absent, because the caller has already built on it. So the honest list is short and explicit.
One-shot only — no continuable children
A continuable Cursor child would mean the harness’s continuation manager owning an inbox for an agent that lives in another process with its own session model. That is a real design, but not this one — and the seam rejects a continuable start on a provider that does not implement it, rather than accepting one that would not work.
No streaming into the parent
Progress arrives as counters and the final text arrives as the result. Piping a Cursor agent’s text into a parent transcript would need a transport this interface does not have, and inventing a half-working one would be worse than reporting numbers.
No per-child model override
Every capability flag is false, which is the honest answer: Cursor composes its own child, in another process, with its own tool set. Advertising a tool filter, a persona or an output schema would promise a guarantee the provider cannot enforce. A useful consequence: a caller cannot even ask for a different model.
No depth cap owned by the harness
Recursion budget is provider-managed, because Cursor composes its own child and the harness has no depth to enforce. A numeric cap is refused at mount rather than accepted and then ignored — which is the same principle as the capability flags.
Checklist before you wire it up
- Decide the split first: what stays with the orchestrator, what gets delegated
- Set a writable home for the bridge store before the first run, not after the first error
- Make the credential reference non-optional so a missing key fails a rollout
- Probe inside the container: key present, SDK importable, store root writable
- Put the wheel on the interpreter path once, on a volume that persists
- Mount the provider in exactly one place — global overlay or roster, never both
- Give the delegated process a named environment list, not the parent’s environment
- Name the requirements and context explicitly; never rely on implicit forwarding
- Record usage from the first run so the budget question has an answer later
- Read remaining-pool figures, not billed figures, when deciding what to delegate
Common mistakes
Treating the child as a second copy of the parent
A delegated agent has no conversational history and no idea a harness sent it. The prompt has to carry the task, the workspace, the requirements and the shape of the answer you want back — otherwise you get confident prose and no idea what changed.
Assuming a provider is a tool
Providers are mounted once and claim a name in a registry; tools merge. Confusing the two produces a boot that refuses to start, and an error message that points at the wrong repository.
Optimising cost before measuring it
A subscription makes "cost" ambiguous: the bill and the draw on the included pool are different numbers, and reporting only the first one says every run was free. Keep both, label them, and only then decide which delegations are worth it.
What this changes in practice
Put the pieces together and the shape of a working day changes. A long orchestration session — reading tickets, planning, running a delegation, reading the result, deciding the next one — stays on one model that is good at orchestration, with a context window that is not being eaten by other people's grep output. The repository-scale work is a call: rename this across the package and make the tests pass, or add the missing validation and show me the failing case first. What comes back is a short, structured report, a cost line in a ledger, and a diff you can review.
The uncomfortable part is also the useful part: nothing about this tells you whether the change is right in production. An agent that edits forty files in ninety seconds is impressive and completely unverified. The delegation result tells you what was done; it does not tell you how the checkout flow behaved afterwards, whether an error group started growing on the code path that was touched, or whether a slow call became slower.
That gap is why we care about the seam in the first place. Delegating the edit is the cheap half; validating it against real sessions is the half that decides whether the change ships. For the delegation-is-one-call pattern applied to reviewing rather than writing code, see How to Build a PR Review Agent with the Cursor Agent SDK. For the other half of the harness — plugins that give an agent a real browser to inspect — read Build a CDP Plugin for Webpage Debugging: Inside the DSH Plugin System. If your team is still deciding where agents should run at all, Local vs Remote AI Agents covers the trade-off, and if the budget is the constraint, How to Minimize Cursor Token Spending is the practical companion to the ledger described above.
Conclusion
Registering the Cursor Agent SDK as a DeepSeek Harness subagent is a small integration with a disproportionate effect, because it lands on a seam rather than in a feature. One provider, one configuration row, one extra tool — and suddenly model choice is per task, a delegated child cannot flood the parent's context, a runaway agent can be killed, the child environment is a list you wrote rather than the one you happen to have, and the cost of delegating is a number in a ledger instead of a feeling. The limits are real and worth knowing: one-shot children, no streaming, no per-child model override. But they are stated, tested, and refused out loud — which is the property that lets you build on top of a seam instead of working around it.