Research note 001 · Agent control integrity ·

The control layer between intent and action

Agents can click, type, write, and run on schedules. The hard part is keeping a simple human instruction — approve, stop, retry — attached to the exact action that reaches the world. We found four places where that control could drift across CopilotKit’s OpenDots and OpenMuse templates. We built the fixes. All four were reviewed and merged upstream.

CopilotKit / OpenDots CopilotKit / OpenMuse PRs #47 · #48 · #137 · #138 Research by Joseph Zaki / Kronowave
4/4merged upstream

Two repositories. Four separate control boundaries. Each fix started with a source-pinned question, moved through a deterministic local reproduction, and ended as a maintainer-reviewed change in main.

STOP · OpenMuse #138

When the user says stop, does the computer actually stop?

The chat turn could end while an active computer command never received the same cancellation signal.

View merged PR #138 →
RETRY · OpenMuse #137

When a desktop action is retried, is it the same action or a second one?

A repeated desktop call could mint a fresh receipt and dispatch the same click or keystroke again.

View merged PR #137 →
INTERRUPT · OpenDots #47

What if the action happened, but the run never recorded completion?

Lease recovery could replay a scheduled task even though the first side effect had already committed.

View merged PR #47 →
APPROVE · OpenDots #48

Does an approval stay attached to the exact thing the person reviewed?

The review receipt remembered the call, but did not retain the complete original draft it represented.

View merged PR #48 →
The question underneath all four

What does the human control actually bind to?

The UI words are tiny: Stop. Approve. Retry. Underneath them are task IDs, tool calls, leases, receipts, abort signals, destinations, and side effects. We wanted to know whether those pieces still meant the same thing after interruption, recovery, or a changed input.

Stop

Did cancellation reach the effect?

A stopped conversation should not quietly leave its computer command behind.

Retry

Is this still the same operation?

A stable retry identity should recover prior work, not perform it twice.

Interrupt

Did the effect already happen?

An unfinished run and an unfinished real-world action are not the same thing.

Approve

Is this still what was reviewed?

An approval should follow the exact content and destination the person saw.

How we worked

Keep the experiment small enough that the answer is obvious.

The method was intentionally simple. Pin the code. Map the identities involved. Change one thing. Read the actual effect. Then build the smallest fix that restores the intended control boundary.

01 · Freeze

Pin the source

Each research track started from one exact upstream revision so source and runtime behavior stayed aligned.

02 · Map

Name the moving parts

We tracked task, run, operation, tool call, receipt, permission, destination, and effect as separate objects.

03 · Change

Move one variable

Cancel the turn. Expire the lease. Reuse the operation ID. Change the draft. Keep everything else fixed.

04 · Read

Use the effect as truth

Count the page rows, computer inputs, receipts, and command signals instead of trusting a status message.

STOP · OpenMuse PR #138

The chat stopped. The computer command did not get the memo.

This was the easiest one to explain because the user expectation is so clear: pressing Stop should reach the action the agent is performing.

Cancellation propagation

One run, two cancellation paths.

The conversation already used a per-run abort signal for browser and search work. The computer-tool wrapper was created without that signal. Our regression started a computer command, stopped the turn, and inspected the signal that reached command execution.

Merged upstream
1. Computer command startsThe chat run dispatches an effectful computer tool.
→
2. User presses StopThe conversation tears down and cancels its other active work.
→
3. Command signal is missingBefore the patch, the active computer command received no run abort signal.
→
4. Signal reaches executionThe fix passes the existing signal through and blocks later dispatch after cancellation.

What we reproduced

On the pinned baseline, unsubscribing during an active computer command left its command signal undefined.

What changed

computerTools now receives the run’s existing AbortSignal plus a pre-dispatch guard. The established executor handles interruption and preserves the uncertain outcome.

The point

Cancellation is only real when it reaches the layer that can still cause the effect. A stopped chat and a stopped command are two different states until the signal connects them.

Local oracle
The signal supplied to the active computer command before and after unsubscribe.
Research check
21 related chat/computer tests passed on the patched branch with TypeScript and Biome checks.
Maintainer review
21 chat/computer tests plus 41 computer API, Docker-runner, and E2B desktop tests passed; the change was approved and merged.
RETRY · OpenMuse PR #137

A retry needs a memory of the exact action, not just another receipt.

Desktop automation makes retries dangerous because a second click or keystroke is a second real effect. We tested what happened when the same uncertain operation came back.

Durable operation identity

The same desktop action could be dispatched twice.

The desktop tool had no stable operation identity at dispatch. Repeating one use_desktop call could mint a fresh receipt and send the same input again. Our synthetic driver counted the actual inputs, not the returned labels.

Merged upstream
1. Action is requestedA click, key input, or typed action is validated for dispatch.
→
2. New receipt is mintedBefore the fix, a repeated call could get a new random receipt identity.
→
3. Input is sent againThe uncertain retry becomes another desktop effect.
→
4. Durable receipt winsThe same operation ID returns the prior receipt; changed full input conflicts.

What we reproduced

On the baseline, replaying the same logical desktop action pressed Return twice and created two receipts.

What changed

The desktop tool now requires a stable operationId, persists the receipt before dispatch, and hashes the full validated action. An exact retry recovers the receipt without new input; a changed action conflicts.

The point

Idempotency is not “we recognize the run.” It is “we recognize this exact effect.” The retry identity has to bind the complete action that matters at the keyboard or mouse.

Local oracle
Actual synthetic desktop input count, durable receipt identity, restart replay, changed-action conflict, and uncertain-input replay.
Research check
The patched desktop branch passed the full upstream suite: 341 / 341, plus server/mobile typecheck, Biome, and diff checks.
Maintainer review
52 desktop/Docker computer tests passed, including restart replay, changed-action conflicts, and uncertain input; the change was approved and merged.
INTERRUPT · OpenDots PR #47

The action can finish before the run knows it finished.

Schedulers recover work by looking at run state. Real-world effects can live on another timeline. We tested the gap between those two clocks.

Ambiguous completion

Effect committed. Run receipt missing.

Our fixture used the real SQLite task store and authorized local page tool. It created one page, stopped before the run could record completion, advanced beyond the lease, and let the same task recover under a fresh lease.

Merged upstream
1. Task is leasedThe scheduled prompt begins in its original conversation.
→
2. Page is writtenThe side effect commits immediately to the Space store.
→
3. Worker stopsThe effect exists, but the run has no final success receipt.
→
4. Recovery replaysThe same task and prompt return under a new lease.

What we observed

After the first effect there was one page. After simulated replay there were two distinct page IDs in the same Space. The old lease could no longer finish the run.

What changed

Expired leases, shutdowns, and settings interruptions now move the task to a durable Interrupted state. The owner gets Retry after review instead of an automatic replay.

The point

A task lease protects the run record. It does not prove whether the outside effect happened. When completion is ambiguous, recovery needs a stable effect identity or an explicit human decision.

Local oracle
Page count and distinct page IDs before and after replay.
Fixture
Synthetic owner, Dot, Space, thread, scheduled task; real SQLite store and actual page-write path.
Verification
15 focused tests passed; full suite: 167 tests across 35 files; formatter, lint, typecheck, production build, and patch-apply checks passed.
APPROVE · OpenDots PR #48

An approval should follow the exact thing the person saw.

OpenDots already had useful controls around reviewed saves. Our first fixture confirmed them, which made the remaining integrity gap much easier to isolate.

Reviewed-content binding

The receipt remembered the call, but not the whole draft.

The reviewed-save receipt was keyed by threadId + toolCallId and stored the saved page and Space. It did not retain the original reviewed title and body. Reusing that call identity with changed body data recovered the earlier saved page without comparing the new draft to the original reviewed draft.

Merged upstream
1. Draft A is reviewedThe owner approves a specific title, body, and Space.
→
2. Page A is savedA receipt is stored for the thread and tool-call identity.
→
3. Retry carries changed bodyThe same call identity now arrives with different draft data.
→
4. Old save is recoveredThe existing receipt stands in without proving the current draft still matches.

What the fixture told us

Space changes were already blocked with 409. Missing owner authentication returned 401. Cross-origin POST returned 403. Revoked Space access returned 403. The narrow gap was draft binding under the same review identity.

What changed

New receipts store the normalized reviewed draft. Identical retries recover the saved page. Changed title, content, or destination return 409 and require a new review.

The point

A stable tool-call ID is not the same thing as stable human intent. The receipt has to remember the content and destination that actually earned the approval.

Local oracle
HTTP status plus actual page rows across two synthetic Spaces, one Dot, and one owned thread.
Base behavior
Same call ID + changed body → HTTP 201 with the original page and original body. Same call ID + changed Space → HTTP 409.
Verification
Full suite: 166 tests; typecheck, lint, formatter, and production build passed. Coverage included title/content/destination changes, restart/migration behavior, and Space-access checks.
What the four fixes say together

Agent safety is a continuity problem.

The action can be valid at the beginning and still become wrong later. The important question is whether identity, permission, content, cancellation, and effect still line up at the moment something actually happens.

Intent has to survive the trip.

From the user’s point of view, “Approve,” “Stop,” and “Retry” are complete instructions. From the system’s point of view, they are promises that must remain true across several independent state machines.

Our working rule is straightforward: bind the control to the complete effect identity, then carry that binding through dispatch, interruption, recovery, and receipt.

Human controlWhat must stay boundFailure we reproducedMerged repair
StopRun cancellation → effectful commandChat ended while the computer command had no abort signalPropagate the run signal and guard pre-dispatch
RetryOperation identity → full validated desktop actionSame uncertain action could dispatch input againDurable operation ID + full-action hash + receipt recovery
InterruptRun state → already-committed effectLease recovery replayed a task after the first page already existedHold ambiguous runs for explicit review and retry
ApproveReview receipt → exact title, content, destinationChanged draft data could reuse the prior review identityPersist normalized reviewed draft and conflict on change