When the user says stop, does the computer actually stop?
The chat turn could end while an active computer command never received the same cancellation signal.
View merged PR #138 →Agents can click, type, write, and run on schedules. The hard part is keeping a simple human instruction — approve, stop, retry — attached to the exact action that reaches the world. We found four places where that control could drift across CopilotKit’s OpenDots and OpenMuse templates. We built the fixes. All four were reviewed and merged upstream.
Two repositories. Four separate control boundaries. Each fix started with a source-pinned question, moved through a deterministic local reproduction, and ended as a maintainer-reviewed change in main.
The chat turn could end while an active computer command never received the same cancellation signal.
View merged PR #138 →A repeated desktop call could mint a fresh receipt and dispatch the same click or keystroke again.
View merged PR #137 →Lease recovery could replay a scheduled task even though the first side effect had already committed.
View merged PR #47 →The review receipt remembered the call, but did not retain the complete original draft it represented.
View merged PR #48 →The UI words are tiny: Stop. Approve. Retry. Underneath them are task IDs, tool calls, leases, receipts, abort signals, destinations, and side effects. We wanted to know whether those pieces still meant the same thing after interruption, recovery, or a changed input.
A stopped conversation should not quietly leave its computer command behind.
A stable retry identity should recover prior work, not perform it twice.
An unfinished run and an unfinished real-world action are not the same thing.
An approval should follow the exact content and destination the person saw.
The method was intentionally simple. Pin the code. Map the identities involved. Change one thing. Read the actual effect. Then build the smallest fix that restores the intended control boundary.
Each research track started from one exact upstream revision so source and runtime behavior stayed aligned.
We tracked task, run, operation, tool call, receipt, permission, destination, and effect as separate objects.
Cancel the turn. Expire the lease. Reuse the operation ID. Change the draft. Keep everything else fixed.
Count the page rows, computer inputs, receipts, and command signals instead of trusting a status message.
This was the easiest one to explain because the user expectation is so clear: pressing Stop should reach the action the agent is performing.
The conversation already used a per-run abort signal for browser and search work. The computer-tool wrapper was created without that signal. Our regression started a computer command, stopped the turn, and inspected the signal that reached command execution.
On the pinned baseline, unsubscribing during an active computer command left its command signal undefined.
computerTools now receives the run’s existing AbortSignal plus a pre-dispatch guard. The established executor handles interruption and preserves the uncertain outcome.
Cancellation is only real when it reaches the layer that can still cause the effect. A stopped chat and a stopped command are two different states until the signal connects them.
Desktop automation makes retries dangerous because a second click or keystroke is a second real effect. We tested what happened when the same uncertain operation came back.
The desktop tool had no stable operation identity at dispatch. Repeating one use_desktop call could mint a fresh receipt and send the same input again. Our synthetic driver counted the actual inputs, not the returned labels.
On the baseline, replaying the same logical desktop action pressed Return twice and created two receipts.
The desktop tool now requires a stable operationId, persists the receipt before dispatch, and hashes the full validated action. An exact retry recovers the receipt without new input; a changed action conflicts.
Idempotency is not “we recognize the run.” It is “we recognize this exact effect.” The retry identity has to bind the complete action that matters at the keyboard or mouse.
Schedulers recover work by looking at run state. Real-world effects can live on another timeline. We tested the gap between those two clocks.
Our fixture used the real SQLite task store and authorized local page tool. It created one page, stopped before the run could record completion, advanced beyond the lease, and let the same task recover under a fresh lease.
After the first effect there was one page. After simulated replay there were two distinct page IDs in the same Space. The old lease could no longer finish the run.
Expired leases, shutdowns, and settings interruptions now move the task to a durable Interrupted state. The owner gets Retry after review instead of an automatic replay.
A task lease protects the run record. It does not prove whether the outside effect happened. When completion is ambiguous, recovery needs a stable effect identity or an explicit human decision.
OpenDots already had useful controls around reviewed saves. Our first fixture confirmed them, which made the remaining integrity gap much easier to isolate.
The reviewed-save receipt was keyed by threadId + toolCallId and stored the saved page and Space. It did not retain the original reviewed title and body. Reusing that call identity with changed body data recovered the earlier saved page without comparing the new draft to the original reviewed draft.
Space changes were already blocked with 409. Missing owner authentication returned 401. Cross-origin POST returned 403. Revoked Space access returned 403. The narrow gap was draft binding under the same review identity.
New receipts store the normalized reviewed draft. Identical retries recover the saved page. Changed title, content, or destination return 409 and require a new review.
A stable tool-call ID is not the same thing as stable human intent. The receipt has to remember the content and destination that actually earned the approval.
The action can be valid at the beginning and still become wrong later. The important question is whether identity, permission, content, cancellation, and effect still line up at the moment something actually happens.
From the user’s point of view, “Approve,” “Stop,” and “Retry” are complete instructions. From the system’s point of view, they are promises that must remain true across several independent state machines.
Our working rule is straightforward: bind the control to the complete effect identity, then carry that binding through dispatch, interruption, recovery, and receipt.
| Human control | What must stay bound | Failure we reproduced | Merged repair |
|---|---|---|---|
| Stop | Run cancellation → effectful command | Chat ended while the computer command had no abort signal | Propagate the run signal and guard pre-dispatch |
| Retry | Operation identity → full validated desktop action | Same uncertain action could dispatch input again | Durable operation ID + full-action hash + receipt recovery |
| Interrupt | Run state → already-committed effect | Lease recovery replayed a task after the first page already existed | Hold ambiguous runs for explicit review and retry |
| Approve | Review receipt → exact title, content, destination | Changed draft data could reuse the prior review identity | Persist normalized reviewed draft and conflict on change |
All four changes were authored from the jz-krono forks, reviewed by the upstream maintainer, approved, and merged into the CopilotKit repositories on October 5, 2026.