agent-desktop/CLAUDE.md
Lahfir 0c0a5b8dbc
fix: harden agent-desktop reliability (#54)
* feat: add strict ref reliability core

* feat: add session-scoped reliability diagnostics

* fix: skip finder pseudo windows for snapshots

* fix: report wait and trace failure context

* fix: harden ref action reliability

* fix: close reliability review findings

* fix: harden reliability edge cases

* fix: close final reliability edge cases

* fix: address reliability follow-ups

* docs: update reliability docs and skills

* fix: harden wait and ref action reliability

* refactor: centralize reliability helpers

* fix: stabilize macos ref resolution

* refactor: organize binary crate modules

* fix: harden ref action reliability

* fix: harden ref action reliability

* docs: compound reliability patterns

* fix: harden ref reliability edge cases

* fix: harden source-window ref resolution

* fix: preserve safe window title fallback

* fix: make explicit snapshots session-independent

* fix: harden ref fallback resolution

* fix: fail closed on uncertain ref fallback

* chore: strip inline comments and enforce docstrings

Replace inline // comments with /// docstrings where they carry non-obvious
contract, and add a pre-commit guard so inline comments cannot regress.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(review): preserve action results, delegate timeout resolution, tighten core boundaries

Apply verified code-review fixes: ref-action release failures no longer
mask successful action results (prevents double-dispatch on retry);
resolve_element_strict_with_timeout defaults to delegating so strict-only
adapters support wait --element; wait --text reports count only when
--count is requested; latest-refmap refresh logs load failures instead of
silently serving stale refs; InteractionPolicy moved to its own module and
actionability/trace modules scoped pub(crate) per file rules; duplicate
wait test helper extracted to shared support module; timeout error
constructors deduplicated; redaction test covers description; policy
focus-denial path covered; skills document steps array, actionability
details, trace redaction, and batch trace inheritance.

* fix(review): eliminate per-element AX round trips and close remaining review findings

Fold AXPosition/AXSize and the scrollbar probe into the existing
AXUIElementCopyMultipleAttributeValues batch so tree traversal and the
actionability preflight pay one IPC per element instead of up to four;
A/B benchmark shows strictly-faster snapshots with identical ref counts
and scroll capabilities (Finder 4.5s -> 2.4s same-session, Docker
Desktop parity at 440 refs / 933 scroll-capable nodes).

Consolidate the CLI and FFI ref-action pipelines into one core
execute_resolved path (actionability, tracing, and dispatch semantics
live once; FFI passes a default context). Remove the no-context
execute() shims from is/right-click/snapshot/wait and the test-only
helper shims; every command now takes an explicit CommandContext.

Split the macOS resolver into resolve (orchestration), resolve_search
(candidate collection), and resolve_classify (strict classification),
clearing the 400-LOC ceiling with room to grow.

Notification waits now retry transient baseline failures inside the
timeout budget with the same retryable gate as window/text waits and
report last_error in timeout details instead of aborting on the first
flake; a baseline is never fabricated.

Refmap writes clean up their temp file on failure, stale *.tmp orphans
are swept under the store lock, and save_existing_snapshot re-verifies
snapshot ownership inside the owning store's write lock with bounded
re-discovery before deterministically recreating in the caller's store.

Coverage hardening: zero-budget wait timeout shape, wait --text
--count 0 absence detection, ref-action pipeline call-count guard
(1 resolve / 1 live read / 1 dispatch), duplicate snapshot-id collision
on load, pruned-everywhere recreation, tmp sweep and rename-failure
cleanup, FFI AMBIGUOUS_TARGET last-error code assertion.

* docs: document notification retry, error-code contract, and diagnostics sensitivity

Note transient-error retry and last_error timeout detail on wait
--notification; make explicit that agents branch on error.code (message
and suggestion text is informational); warn FFI consumers that
ad_last_error_details may carry on-screen element names, values, and
window titles and should stay out of shared log surfaces.

* fix: validate find roles against the canonical vocabulary

find --role with a role no adapter can emit (textarea, typos) silently
returned ok with zero matches, reading as 'element absent' when the
query could never match. Role queries now resolve through a canonical
vocabulary in core: common text-input aliases (textarea, textbox,
searchfield) normalize to textfield case-insensitively, and unknown
roles fail with INVALID_ARGS carrying details.valid_roles so agents can
self-correct.

The macOS role mapping becomes a sorted single-source table with
binary-search lookup, and a conformance test pins every emitted role
(plus the unknown fallback and the synthesized cell role) to core's
CANONICAL_ROLES — the cross-platform contract Windows/Linux adapters
must map their native vocabularies into, enforced by the same
table + test pattern rather than convention.

* refactor: derive find role hints from the live tree, drop hardcoded vocabulary

Replace the canonical-role allow-list (and its hard INVALID_ARGS
rejection) with a tree-derived approach. A role filter that matches
nothing now returns ok with roles_present — the distinct roles actually
in the searched tree — so the caller distinguishes 'none on screen' from
a wrong role name and self-corrects. This needs no central role list: a
role any adapter newly emits surfaces in roles_present automatically,
with nothing to keep in sync across core and the platform crates.

A tiny role-query normalizer keeps the ergonomic win (textarea, textbox,
searchfield fold to textfield, case-insensitive) but never gates or
rejects — it is a synonym shim, not a vocabulary. The macOS role table
returns to its plain match form; the cross-crate canonical-vocabulary
list and its conformance test are gone.

* refactor: move protected-process knowledge out of core into the adapter

close-app hardcoded macOS/Unix process names (loginwindow, windowserver,
dock, launchd, finder) inside core, baking platform-specific knowledge
into the platform-agnostic crate. Windows would need csrss.exe/
winlogon.exe, Linux gnome-shell/Xorg. Add PlatformAdapter::
is_protected_process (default denies nothing); the macOS adapter owns its
list with substring matching over display and bundle identifiers. core's
close-app just asks the adapter. Also genericize a macOS-flavored test
fixture string so core carries zero native vocabulary even in tests.

Verified: core has no platform-native references in non-test source, no
cfg(target_os) gates; the only remaining cfg(unix)/libc use is securing
core's own refmap/trace/lock files with non-unix fallbacks.

* fix: stop close-app claiming a graceful quit it cannot confirm

close-app returned closed:true the instant a graceful quit was *sent*,
while the app was still running behind an unsaved-changes dialog —
a false completion claim. Empirically (NSWorkspace.runningApplications):
a clean quit completes in ~0.2s, a dialog-blocked quit never completes
on its own, and macOS confirms only that the quit request was sent, not
that the app terminated. Verifying by polling would add seconds of
latency on the exact (blocked) case it is meant to catch, so we do not
poll.

Graceful close now reports { method: graceful, requested: true } —
truthful and instant, no closed claim. --force is a synchronous SIGKILL,
so it reports { method: force, requested: true, closed: true }. Callers
needing graceful confirmation observe via list-apps / wait --window and
can drive a save dialog with snapshot + find, which is the agent-native
path.

* feat: add drag --drop-delay for reliable macOS drop registration

macOS drop targets need the dragged item to dwell over them before they
register as the destination; too short and the gesture lands as a drag
with no drop. The dwell was a hardcoded 500ms dead sleep. Expose it as
--drop-delay <ms> (CLI), drop_delay_ms (DragParams/AdDragParams, 0 =
adapter default sentinel matching duration_ms), and replace the dead
sleep with an event-driven dwell that posts LeftMouseDragged over the
destination every 16ms so the target stays highlighted instead of
dropping the drag mid-pause.

DRY: the C-to-core drag conversion (duration/drop-delay zero-sentinel)
was copied across three FFI sites; collapse them into AdDragParams::
to_core(). FFI ABI: AdDragParams gains drop_delay_ms (header + repr +
header-compile test). Tests: core threads the value into params and
response and omits the field when unset; FFI maps both optionals.

* fix: make action-bearing elements ref-able so scroll/expand can target them

E2E testing against a diverse fixture app surfaced that disclosure
(Expand/Collapse/Click) and scrollarea (Scroll) advertise actions but
never received refs — they are not in INTERACTIVE_ROLES — so the scroll,
expand, and collapse commands required a <REF> their own target roles
could never have. The commands were uninvokable against their primary
targets.

Ref allocation now gates on addressability, not role alone: an element
is ref-able if its role is interactive OR it advertises a primary action
(any action other than a bare SetFocus, which would ref-allocate inert
focusable containers). scrollarea and disclosure become ref-able;
scroll now works against a real app. Ref-count impact is modest
(fixture 61->72, Finder ~262).

Tests assert action-bearing containers get refs, SetFocus-only and inert
elements do not, and interactive roles stay ref-able without actions.
Contract docs (CLAUDE.md, SKILL.md) updated.

* fix: eliminate vacuous AX successes and harden resolution

Dogfooding the binary against a real fixture app surfaced five cases where a
command reported success without producing the effect, or failed with the wrong
error. Each is verified by independent before/after observation in the E2E
harness.

- is_menu_open no longer treats a latent AXMenuBar as an open menu, so select
  and wait --menu-closed stop seeing a permanently-open menu.
- set-value coerces the written AXValue to the element's existing CFNumber/
  CFBoolean/CFString type and verifies numerically, fixing sliders; steppers
  converge via AXIncrement/AXDecrement when AXValue writes are vacuous.
- double-click only claims success when the element advertises AXOpen; otherwise
  it fails closed instead of reporting a non-existent double-click.
- a completed resolution pass that proves a ref absent downgrades a deadline
  TIMEOUT to STALE_REF so removed elements fail with the correct code.
- expand/collapse verify the disclosure state and fall back to a press-toggle
  for press-driven disclosures; press-toggled containers expose EXPAND/COLLAPSE.

roles.rs adds the disclosure expandable role and normalizes textarea/textbox/
searchfield role queries to textfield.

* feat: add Playwright-style headed/headless interaction mode

Ref actions now run in exactly two modes. Headless is the default: semantic
accessibility operations only, no cursor movement, and a fail-closed
POLICY_DENIED when only a physical gesture would work. The global --headed flag
upgrades every ref action to permit focus stealing and cursor movement, so the
chain's physical click/double-click/scroll/keypress fallbacks can complete. The
AX path is always tried first, so --headed never regresses headless-capable
elements; it only adds fallbacks for elements that need a real gesture.

- CommandContext::request(action, base) builds the per-command request: each
  command declares its headless base (pure-AX headless; type uses focus_fallback
  because typing requires focus but never moves the cursor) and --headed upgrades
  any base to the headed policy.
- The internal/FFI "physical" policy is renamed "headed" throughout, including
  the C ABI enum (AD_POLICY_KIND_HEADED keeps discriminant 2) and bindings.
- Raw-input commands (press, hover, drag, mouse-*, key-down/up) are unchanged:
  always physical, mode-independent low-level escape hatch.
- Unit tests assert every ref command is headless by default and headed under
  --headed; docs (CONCEPTS, CLAUDE, skills) describe the two-mode contract.

* test: add E2E fixture app and dual-mode harness

Drives the release binary against a real SwiftUI/AppKit fixture and verifies
every effect by independent before/after observation — never the command's own
ok:true — so a command that reports success without an effect is caught. This is
the layer mock-adapter unit tests cannot cover: it exercises the contract
against the real macOS Accessibility API.

- AgentDeskFixture.swift exposes a fixed, diverse AX surface (native AppKit
  slider/stepper, gesture-only and ambiguous controls, a sheet, a press-toggled
  disclosure, async-appearing elements, a drag canvas). It is never tuned to
  make a command pass; a failure is a finding about the CLI or the harness.
- run.sh drives every ref-action command in BOTH headless and --headed mode with
  mode-specific target values, plus the double-click discriminator (headless
  fails closed with POLICY_DENIED, --headed completes) that proves the two modes
  differ. It also covers strict resolution, wait predicates, skeleton drill-down,
  sessions, trace redaction, surfaces, drag, expand, and force-close.
- The compiled fixture .app is a build artifact (gitignored; built on demand).

Run: cargo build --release && bash tests/e2e/run.sh (needs AX permission).

* fix: close review runtime, correctness, and security gaps

Addresses validated findings from the branch code review:

- chain: thread the chain deadline into increment_to_value so a non-converging
  stepper cannot spin up to 1024 AX round trips and blow past the timeout.
- mouse: a RAII guard posts LeftMouseUp if any fallible step of a drag returns
  early, so an error can never leave the mouse button held down system-wide.
- ffi: cap caller-supplied state/action/path counts before from_raw_parts,
  mirroring the existing MAX_MODIFIERS_PER_COMBO guard, to reject out-of-bounds
  reads from a garbage C count.
- wait: the actionable predicate now forwards the structured ActionabilityReport
  from error.details instead of dropping it to a flat message, so agents can see
  which check is blocking.
- scroll: gate the row-select fallback on policy.allow_focus_steal so a headless
  scroll can no longer silently change the user's table selection.
- actionability: delete the unreachable stable->StaleRef branch (stability_check
  never fails by design) and its now-dead failed_check helper.
- status: delete two pub wrappers with no production callers; the test now drives
  the real execute_with_report_with_context entry point.

* refactor: extract disclosure chain steps under the 400-LOC limit

chain_steps.rs had grown to 402 lines, over the hard per-file limit. The six
press-toggle disclosure helpers form a cohesive group and move cleanly into a
sibling chain_disclosure_steps.rs (following the chain_web_steps/
chain_menu_steps pattern); chain_defs.rs references the new module. No behavior
change.

* feat: focus the target window before ref-addressed physical input

Ref-action physical fallbacks (click/scroll/type) already brought the target
app frontmost before synthesizing CGEvents. The raw-input commands (drag, hover,
mouse-*) resolving a point from a ref did not — the resolver had the pid but
discarded it, so the adapter saw only coordinates and synthetic events could
land on whatever window happened to be frontmost.

Add a best-effort focus_app(pid) to PlatformAdapter (macOS uses
ensure_app_focused; other adapters default to not_supported). The point resolver
now focuses the ref's app before returning, so every physical interaction that
targets a known element raises its window first. Coordinate (--xy) input is
unchanged: the caller owns the target there.

* fix: add AdDragParams size guard and document the ABI breaks

AdDragParams gained a drop_delay_ms field but, unlike AdRefEntry, had no size
guard — an old caller's smaller allocation would let Rust read past it and turn
stack garbage into a real drop delay. Add AD_DRAG_PARAMS_SIZE, ad_drag_params_size(),
a compile-time layout assertion, and a zero-init note, matching the ref-entry
pattern.

This branch makes several consumer-visible contract changes that release
tooling must cut as a major. They are gathered here because the release workflow
ships the C header as an artifact.

BREAKING CHANGE: the C ABI and CLI/JSON contract changed on this branch.
- AdPolicyKind: AD_POLICY_KIND_PHYSICAL is renamed AD_POLICY_KIND_HEADED
  (discriminant 2 unchanged, so compiled binaries are safe; source-level C
  consumers must rename). No back-compat alias is kept — "physical" is gone.
- AdRefEntry grew (caller-allocated input); validate layout with
  AD_REF_ENTRY_SIZE / ad_ref_entry_size().
- AdDragParams grew; validate with AD_DRAG_PARAMS_SIZE / ad_drag_params_size()
  and zero-initialize before use.
- close-app graceful response no longer includes closed:true; it returns
  { method: "graceful", requested: true } because a graceful quit cannot be
  synchronously confirmed.

* test: harden and expand the E2E proof layer

Closes the honesty gaps the review found and covers the interactions that were
missing, all verified by independent before/after observation:

- twins: the fixture twins now record distinct effects (twin-a/twin-b) and the
  assert requires the ADDRESSED twin to fire (or AMBIGUOUS_TARGET), instead of
  passing on any ok:true.
- click: click-status is a counter, so the headed pass must observe a fresh
  increment rather than inheriting the headless pass's value.
- adds triple-click + hover (headed gestures), tab selection (TabView tabs are
  radiobuttons), context-menu open + item selection, and menu-bar enumeration
  via --surface menubar.
- adds a performance section reporting per-command CLI wall-clock (snapshot,
  find, get, click, set-value, type) with a soft <2s snapshot gate.
- documents the SwiftUI CommandMenu and cross-app drop limitations as tracked
  notes, not silent skips.

* docs: trim CLAUDE.md to standards and document gesture headless-capability

- CLAUDE.md: remove ~130 lines of reference material that duplicated code,
  Cargo.toml, or the skills (full PlatformAdapter trait dump, Key Types listing,
  macOS API listings, dependency/build-config tables, the 54-command table).
  Replaced the stale trait dump with a pointer to adapter.rs (which also fixes
  the review's stale-trait-docs finding) and kept only the non-obvious gotchas.
  CLAUDE.md is now standards, invariants, and conventions.
- Document, on macOS (Phase 1), which gestures have a headless path: most ref
  actions do; double-click via AXOpen; triple-click/hover/drag are cursor
  gestures with no AX equivalent (physical only). The command surface is
  platform-agnostic — a future Windows/Linux adapter that exposes a headless
  path lights it up with no command or core change. Added to README and the
  interaction reference.
- FFI skill: AD_POLICY_KIND_PHYSICAL is now AD_POLICY_KIND_HEADED.

* docs: capture gesture headless-capability learning and refresh policy docs

- Add best-practices/macos-gesture-headless-capability: which desktop gestures
  have a headless AX path on macOS (double-click via AXOpen; triple-click/hover/
  drag are physical-only; SwiftUI controls vs native AppKit), and why the command
  never decides — the platform adapter owns headless-vs-physical.
- Refresh two policy learnings for the physical->headed rename: ActionRequest::
  physical -> headed and AD_POLICY_KIND_PHYSICAL -> AD_POLICY_KIND_HEADED, noting
  the new global --headed upgrade path via CommandContext::request.

* fix: address review P2 correctness and reliability findings

- wait: cap each ref-resolution attempt (750ms) so a slow resolve cannot
  consume the whole wait budget on the first poll; the predicate is re-checked
  across the full timeout.
- wait: make LatestRefCache timing fields private (no external readers).
- wait: add a unit test for --menu-closed (asserts it waits for open=false).
- close-app: make the graceful and force responses symmetric — both carry
  `closed` (force confirms true; graceful cannot confirm, so false) instead of
  graceful silently omitting the field.
- close-app: test adapter-error propagation.
- adapter: the default resolve_element_strict_with_timeout now logs that it does
  not enforce the deadline, so an adapter that forgets to override it is visible
  in traces.
- type: a RAII guard restores the user's clipboard on every scope exit (success,
  error, panic) during the paste-based non-ASCII path, shrinking the clobber
  window to an unpreventable SIGKILL.

Note: the reviewer's "fail fast on AMBIGUOUS_TARGET with a pinned snapshot"
suggestion is not applied — existing tests prove transient ambiguity resolves on
retry even with a pinned ref (ambiguity is a property of the live tree, not the
refmap), so failing fast would regress intentional, tested behavior.

* perf: trim actionability preflight and resolve-search allocations

- action_list: gate the AXValue and AXExpanded `is_settable` probes on whether
  the role could plausibly carry that capability (unknown roles always probe),
  skipping up to two AX round trips per preflight on common click-only targets.
  No capability is lost — value/expandable roles still probe.
- resolve_search: reuse one scratch FxHashSet across nodes instead of allocating
  a fresh dedup set per node during path and recursive search.

* refactor: split actionability types into one file each

ActionabilityStatus, ActionabilityCheck, and ActionabilityReport move into their
own files under a new actionability/ module (following the tree/ and actions/
folder pattern); mod.rs keeps the check logic and re-exports the types. Honors
the one-domain-type-per-file rule without changing behavior.

* test: split fixture under the LOC limit and fix swift-ios issues

- Extract the reusable fixture components (status readout, native AppKit
  slider/stepper, drag canvas, card) into FixtureComponents.swift so each file
  is under the 400-LOC limit; build.sh now compiles every .swift in the dir.
- NativeSlider/NativeStepper implement updateNSView so a SwiftUI binding change
  syncs back to the NSView.
- The fixture no longer steals focus unconditionally on launch
  (activate ignoringOtherApps:false), so it cannot mask headless-policy focus
  violations; the harness drives focus explicitly.
- build.sh pins the SDK and a macOS 13 deployment target for reproducible builds.

* docs: document the optional error.details field and roles_present shape

- SKILL.md: note that the error object may carry an optional `details` (the
  actionability report, AMBIGUOUS_TARGET candidates, or a wait TIMEOUT's last
  observed state) and that responses should be parsed leniently — `details` and
  future fields are additive.
- commands-observation: show the no-match `find` response with the
  `roles_present` hint so callers can tell a wrong role name from "none on
  screen".

* fix: gate the ref-gesture focus raise on the interaction policy

Ref-addressed hover/drag raised the target app unconditionally, violating
the headless no-implicit-focus-steal contract. The point resolver now
returns the owning pid instead of focusing, and commands decide: headless
never raises, --headed raises once (drag focuses only the from-app, fixing
the cross-app double-focus). Responses report focused:true so multi-app
agents can detect the frontmost change.

* fix: abort failed drags at the origin and disarm the guard only on success

The mouse-up guard disarmed before the final fallible up-event, so a
failed final post left the button held. Worse, its corrective release
fired at the unreached destination, silently committing an aborted drag
as a completed drop (CGEvents resolve at their embedded coordinates).
The guard now owns the release: it disarms only after the up actually
posts, and an early return cancels by dragging back to the origin and
releasing there.

* fix: enforce the chain deadline inside increment steps

All dispatch sites construct ChainContext with deadline: None, so the
remediation parameter on increment_to_value never received a value and
the 1024-iteration loop ran unbounded by the chain timeout. The chain now
pins its resolved deadline into the context every step observes. Also
extracts the pure write-verification predicates and their tests to
chain_verify.rs, bringing chain.rs back under the 400-LOC limit.

* fix: pin the AdAction ABI layout and bound FFI string and array inputs

AdDragParams is embedded by value in AdAction, so its 8-byte growth grew
the struct C callers pass to ad_execute_action with no size guard —
old-layout callers under-allocate and stack garbage becomes a live
drop_delay_ms. Adds AD_ACTION_SIZE / ad_action_size() with a layout pin,
matching the AdRefEntry pattern.

Also hardens the input boundary: C strings are decoded with a bounded
NUL scan (AD_MAX_STRING_BYTES, sized for CLI argv parity) so a missing
terminator cannot walk arbitrary memory, and the single coarse 1024
array cap becomes published per-field caps (AD_MAX_REF_STATES/ACTIONS/
PATH_DEPTH) with tests just over each limit.

* perf: replace fixed input settle sleeps with state polls

ensure_app_focused slept 50ms per physical input even when the app was
already frontmost; it now polls AXFrontmost (1ms, 50ms deadline).
disclosure_settled slept an unconditional 40ms up to three times per
expand/collapse; it now polls the disclosed state (5ms, 200ms deadline),
converging immediately on fast UIs. The kAXFocusedAttribute settability
probe is gated by role_may_accept_focus, mirroring role_may_bear_value,
and key dispatch reuses ensure_app_focused instead of its inline
duplicate.

* feat: check a specific action in wait --predicate actionable

The actionable predicate hardcoded Click, so wait-then-type flows got a
false ready on fields that cannot accept text (the editability check only
runs for editing actions). --action selects click (default), type,
set-value, or clear, and the preflight mirrors each command's real base
policy (type uses its focus-fallback base).

* fix: redact title, url, help, and placeholder keys in traces

Window titles, URLs, tooltips, and placeholder text carry user content
just like names and values; the redaction list now covers them.

* test: harden the e2e harness and unify fixture AX labels

The harness now fails setup loudly when the fixture build or AX trust is
missing, rebuilds the fixture when sources are newer than the bundle,
asserts hover only against a freshly observed state, and force-collapses
the disclosure so the expand test proves a real flip. The slider/stepper
labels live solely on the NSViews (the AX-actionable elements), removing
the macOS-version-dependent race between two label sources, and the drag
canvas reports a zero frame when detached from a window.

* docs: record the perf commit type, pre-1.0 bump policy, and error details field

Adds perf: to the allowed commit types (release-please already maps it to
a Performance changelog section), records the pre-1.0 versioning policy so
a BREAKING footer is expected to cut a minor rather than a major, and
shows the optional error.details object in the error envelope docs.
The gitignored local AGENTS.md mirror got the same contract sync.

* fix: surface increment deadline expiry as timeout with the observed value

A chain deadline firing mid-increment returned Ok(false), so the step was
recorded as skipped, the chain exhausted into ACTION_FAILED, and the
control sat at a half-applied value the caller could not see — post-state
is only read on success, and ACTION_FAILED recovery guidance points away
from retrying. Expiry is now a TIMEOUT error carrying value_before,
value_at_timeout, target, and a mutated flag in details.

* perf: cap the disclosure settle poll to the chain deadline and widen its interval

The settle poll could spend 3 x 200ms x 5ms-interval reads (~360 IPCs)
per expand/collapse and overshoot the chain's own deadline. A new
CustomWithDeadline chain step threads the chain deadline into the
disclosure steps, the settle budget is min(200ms, remaining chain
budget), and the interval widens to 20ms (~30 IPCs worst case).

* fix: enforce the protected-process guard inside the adapter close path

The guard lived only in the CLI command layer, so ad_close_app could
force-kill session-critical processes (loginwindow, WindowServer, Dock)
that the CLI refuses. close_app_impl now refuses them before any side
effect with the exact CLI error contract, making CLI, FFI, and any future
consumer behave identically; the CLI preflight remains as an earlier
check against the same predicate.

* fix: make focused semantics honest and confirm window focus by polling

ensure_app_focused set AXFrontmost unconditionally and reported success
identically whether or not a raise happened; it now no-ops when the app
is already frontmost, so Ok (and the focused:true response field) means
"frontmost ensured" exactly as documented. focus_window_impl gains the
same confirmation poll after its raise, and the poll interval widens to
5ms (10 reads max in the 50ms window).

* fix: attach abort-state guidance to drag failures and document cancel limits

Drag synthesis errors surfaced as bare INTERNAL with no hint about the
gesture's end state. Failures now carry a suggestion stating the button
was released back at the origin (best-effort), no drop was committed,
and where the cursor ends; the guard doc spells out the two best-effort
limits (corrective posts can fail; a self-drop at the origin is a no-op
for most targets).

* refactor: split oversized files by responsibility under the 400-LOC limit

helpers_tests (426) splits into resolution/window/pipeline tests, a
ref-action+trace test file, and a shared entry-builder support module.
wait_element_tests (414) splits into predicate-behavior tests and
resolution/lifecycle tests over a widened wait_test_support. wait.rs
(395) loses the element-wait loop to wait_element.rs, and refs_store
(397) moves its tmp-cleanup/retention methods to a refs_store_prune
child module (declared via #[path] so the split keeps base_dir and
snapshots_dir private to the store).

* refactor: build the actionable preflight request per action name at parse

The policy mirror lived in a separate helper with a catch-all arm, so a
future action name could silently inherit the headless policy. Parse now
maps every --action name to the exact ActionRequest its real command
runs (type is the only focus-fallback), the catch-all is gone, and a
test pins each name's policy.

* refactor: move point resolution and the focus helper to point_resolve

PointResolveArgs, ResolvedPoint, the ref-or-xy resolver, and
focus_for_physical_input were accumulating in helpers.rs alongside
unrelated ref-action plumbing; they now live in a dedicated
point_resolve module consumed by hover and drag.

* test: keep the drag canvas AX label on the NSView only

The DragCanvas carried two label sources (the NSView and a SwiftUI
modifier on its representable), the same macOS-version-dependent race
the slider/stepper fix removed; the harness-facing label now lives
solely on the AX-actionable NSView.

* docs: sync skills and header with the focused, redaction, and wait contracts

Documents the ensured (best-effort, already-frontmost-aware) semantics of
focused:true and its absence-vs-false meaning, the four redaction keys
added in round 2 plus the substring-match behavior, the INTERNAL error
recovery row, the FFI wait-surface asymmetry, the cross-app drag
occlusion caveat, and the AdDragParams/AdAction layout history with the
adjudicated pre-1.0 breaks so fresh reviews stop re-finding them.

* fix: surface settle-wait deadline truncation as timeout with a schema discriminant

A chain deadline truncating the disclosure settle wait returned a plain
step failure, exhausting into ACTION_FAILED — the same masking class
fixed for increments — even though the triggering action may still land
after the truncated wait. Settle exits are now classified: full-budget
misses stay step failures, deadline-truncated waits raise TIMEOUT with
the wanted/observed state, and the poll sleeps are clamped so a tight
deadline still gets at least one read. All TIMEOUT details now carry a
kind discriminant (wait_timeout vs chain_deadline) so agents can branch
without sniffing field names.

* fix: match protected processes exactly, not by substring

'docker'.contains('dock') permanently blocked close-app for Docker,
FinderSync-class apps, and anything else embedding a protected name.
Matching is now an exact lowercase name or an exact dot-separated
bundle-id component, so Dock and com.apple.dock stay protected while
Docker, Docker Desktop, FinderSync, and PathFinder stay closable —
pinned by false-positive tests.

* fix: raise the element's window before the physical click fallback

CGEvents land on the topmost window at the click point, so an app being
frontmost is not enough when the target element lives in a background
window of that app — the physical fallback clicked whatever overlapped
it, and the skip-raise-when-frontmost optimization widened the window
for that. click_via_bounds now raises the element's own AXWindow (AXRaise,
AXMain fallback, brief confirmation poll) via a shared window_ops helper
that focus_window_impl reuses. Verified live: a headed click on a ref in
an occluded Finder window raises that window and lands the click in it.

* test: guard ref-action policy coverage against silent gaps

A new ref-action command could ship without a base-policy assertion. A
guard test now scans crates/core/src/commands/ for files calling
context.request( and fails unless each stem appears in the
POLICY_TESTED_COMMANDS list backing the policy assertions.

* fix: carry the protected-process suggestion on the CLI preflight

The CLI-layer guard returned bare INVALID_ARGS while the adapter layer
carried recovery guidance, so agents on the primary surface got 'check
command syntax' for a permanently-disallowed operation and looped on
argument fixes. Both layers now state the same suggestion.

* docs: document the TIMEOUT schemas, chain deadline knob, and roles_present scope

Names the two TIMEOUT details schemas by their kind discriminant with
the mutated-flag retry rule, points chain-deadline recovery at
AGENT_DESKTOP_CHAIN_TIMEOUT_MS instead of --timeout, extends the
roles_present hint to all non-count selection-mode misses, and aligns
the STALE_REF recovery row with the richer error.rs suggestion.

* test: extract the scroll card and document drag-canvas data flow

AgentDeskFixture.swift sat at exactly 400 lines; the scroll card is
fully self-contained (its offset state never leaves the card), so it
moves to FixtureCards.swift as a standalone view with zero bindings,
landing the main file at 374. Harness-facing labels are byte-identical.
DragCanvas gains two intent comments distinguishing the deliberately
empty updateNSView from a forgotten sync.

* chore: pin ABI sizes for C consumers and explain the prune module split

C11-gated _Static_asserts mirror the Rust-side layout pins so a C
consumer compiling against a drifted header fails at build time, and
the production #[path] prune module carries its privacy rationale.

* test: give racing wait tests a deterministic budget

Three wait tests used a 1ms timeout that can elapse before the loop's
first resolution attempt under load, flaking on machine pressure (the
ambiguous-resolution test needs at least one attempt to record its
observation). 50ms guarantees the first attempt without slowing the
suite.

* docs: capture three round-4 review learnings and grow the concept map

Documents the abort-state contract for multi-step physical input (guard
disarm ordering, origin release, end-state suggestions), the three-layer
repr(C) size-pinning discipline born from the AdAction silent-growth
incident, and the named-arms-plus-exhaustiveness-guard pattern for
policy/dispatch mirrors. CONCEPTS.md gains Action Chain and Protected
Process and refreshes Coordinate Fallback with the window-topmost rule.

* docs: refresh six learnings against the enhanced-reliability branch

Brings the learning corpus back in line with code that moved this
branch: the gesture-capability and policy docs now describe the
window-level raise in the physical path and link the new abort-state
doc, the reliability contract documents the TIMEOUT kind discriminant
and the wait --action per-name policy variant, the FFI review rule
covers structural repr(C) size drift alongside behavioral parity, the
allocator doc records how the config struct absorbed four more fields
in one place, and the fingerprint doc names the real tri-state decode
error type. Three docs verified accurate with no edits.

* docs: sync roadmap with current reliability contracts

* fix: harden desktop action reliability

* fix: harden reliability review regressions

* fix: close final reliability review gaps

* fix: close reliability review gaps

* test: avoid raw pointer mutation in ffi free tests

* refactor: trim reliability branch dead code

* fix: close reliability review gaps

BREAKING CHANGE: the C ABI AdActionResult layout now includes action steps; C consumers must rebuild against the updated agent_desktop.h header.

* fix: close reliability review gaps

* ci: scope cache hash inputs

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-20 15:43:42 -04:00

19 KiB
Raw Blame History

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Common Commands

cargo build                                    # Debug build
cargo build --release                          # Release build (<15MB target)
cargo test --lib --workspace                   # Run all unit tests
cargo test --lib -p agent-desktop-core         # Test core crate only
cargo test --lib -p agent-desktop-macos        # Test macOS crate only
cargo test test_name                           # Run a single test by name
cargo clippy --all-targets -- -D warnings      # Lint (must pass, zero warnings)
cargo fmt --all -- --check                     # Format check
cargo fmt --all                                # Auto-format
cargo tree -p agent-desktop-core               # Verify no platform crate leaks (CI enforces)
bash tests/e2e/run.sh                          # E2E: real binary vs fixture app, verify by observation (needs --release + AX permission)

Run the binary: ./target/release/agent-desktop snapshot --app Finder -i

The E2E harness drives the release binary against a real SwiftUI/AppKit fixture and asserts every effect by independent observation (never the command's own ok:true), covering every ref action in both headless and --headed mode. See tests/e2e/README.md.

Pre-commit Hook

The repo ships a pre-commit hook at .githooks/pre-commit that runs cargo fmt --check, cargo clippy --all-targets -- -D warnings, and cargo test --lib --workspace against staged Rust changes. Wire it up once after cloning:

git config core.hooksPath .githooks

Bypass for an emergency commit with git commit --no-verify or SKIP_PRECOMMIT=1 git commit ....

Project Overview

Cross-platform Rust CLI + MCP server enabling AI agents to observe and control desktop applications via native OS accessibility trees.

Git & Commits

  • All commits are authored by Lahfir
  • NEVER add Co-Authored-By lines, AI attribution badges, or "Generated with" footers
  • NEVER include co-committers of any kind
  • Conventional Commits required. Every commit message must use a type prefix:
    • feat: — new feature (triggers minor version bump)
    • fix: — bug fix (triggers patch version bump)
    • feat!: or BREAKING CHANGE: footer — breaking change (triggers major version bump)
    • docs: — documentation only
    • style: — formatting, no code change
    • refactor: — code change that neither fixes a bug nor adds a feature
    • perf: — performance improvement with no behavior change
    • chore: — maintenance tasks, dependencies
    • ci: — CI/CD changes
    • test: — adding or fixing tests
  • Format: type: concise imperative description (lowercase type, no capital after colon)
  • Focus on "why" not "what"
  • Examples: feat: add scroll-to command, fix: prevent stale ref on window resize, ci: add binary size check
  • Pre-1.0 versioning policy (release-please bump-minor-pre-major + bump-patch-for-minor-pre-major): while the version is 0.x, a BREAKING CHANGE cuts a minor (0.2 → 0.3) and a feat: cuts a patch. Do not expect a major release before 1.0.

Core Principle

agent-desktop is NOT an AI agent. It is a tool that AI agents invoke. It outputs structured JSON with ref-based element identifiers. The observation-action loop lives in the calling agent.

Architecture

Workspace Layout

agent-desktop/
├── Cargo.toml              # workspace: members, shared deps
├── CONCEPTS.md             # shared domain vocabulary for refs, snapshots, sessions, actionability, and related concepts
├── rust-toolchain.toml     # pinned Rust version
├── clippy.toml             # project-wide lint config
├── crates/
│   ├── core/               # agent-desktop-core (platform-agnostic)
│   │   └── src/
│   │       ├── ref_alloc.rs      # Shared ref helpers (INTERACTIVE_ROLES, is_collapsible)
│   │       ├── snapshot_ref.rs   # Ref-rooted drill-down (run_from_ref)
│   │       └── commands/         # one file per command
│   ├── macos/              # agent-desktop-macos (Phase 1)
│   ├── windows/            # agent-desktop-windows (stub → Phase 2)
│   ├── linux/              # agent-desktop-linux (stub → Phase 2)
│   └── ffi/                # agent-desktop-ffi (cdylib + committed C ABI header)
├── src/                    # agent-desktop binary (entry point)
│   ├── main.rs             # entry point, permission check, JSON envelope
│   ├── batch/              # batch JSON → typed Commands
│   ├── cli/                # clap derive enum, help text, CLI contract tests
│   ├── cli_args/           # command argument structs by domain
│   ├── command_policy/     # permission/ref/side-effect policy
│   ├── dispatch/           # command dispatcher, parse helpers, notifications
│   └── tests/              # binary-level conformance tests
├── docs/
│   └── solutions/          # documented solutions to past problems (bugs, best practices, workflow patterns), organized by category with YAML frontmatter (module, tags, problem_type); relevant when implementing or debugging in documented areas
└── tests/
    ├── fixtures/           # golden JSON snapshots
    └── integration/        # macOS CI integration tests

Dependency Inversion (Non-Negotiable)

  • agent-desktop-core defines the PlatformAdapter trait and all shared types
  • Platform crates (macos, windows, linux) implement the trait
  • Core NEVER imports platform crates. Platform crates NEVER import each other.
  • Two legitimate wiring points bring platform → core together:
    1. The binary crate (src/) — CLI consumers
    2. The FFI crate (crates/ffi/) — cdylib consumers (Python, Swift, Go, Node, C++)
  • CI enforces core isolation: cargo tree -p agent-desktop-core must contain zero platform crate names

Platform Selection

Compile-time via #[cfg(target_os)] in build_adapter(). Agents never specify platform — agent-desktop snapshot -i works identically on macOS, Windows, and Linux.

fn build_adapter() -> impl PlatformAdapter {
    #[cfg(target_os = "macos")]
    { agent_desktop_macos::MacOSAdapter::new() }

    #[cfg(target_os = "windows")]
    { agent_desktop_windows::WindowsAdapter::new() }

    #[cfg(target_os = "linux")]
    { agent_desktop_linux::LinuxAdapter::new() }
}

Target-Gated Dependencies

Binary crate Cargo.toml uses platform-specific deps, NOT unconditional deps with #[cfg] in source:

[target.'cfg(target_os = "macos")'.dependencies]
agent-desktop-macos = { path = "crates/macos" }

[target.'cfg(target_os = "windows")'.dependencies]
agent-desktop-windows = { path = "crates/windows" }

[target.'cfg(target_os = "linux")'.dependencies]
agent-desktop-linux = { path = "crates/linux" }

Command Dispatch

Direct match in the binary crate. No Command trait, no CommandRegistry. Each command is a standalone execute() function under crates/core/src/commands/.

pub fn dispatch(
    cmd: Commands,
    adapter: &dyn PlatformAdapter,
    permission_report: &PermissionReport,
) -> Result<serde_json::Value, AppError> {
    match cmd {
        Commands::Snapshot(args) => commands::snapshot::execute(args, adapter),
        Commands::Click(args) => commands::click::execute(args, adapter),
        // one arm per command
    }
}

Batch is not a second dispatcher. src/batch/mod.rs deserializes JSON entries into the same typed Commands enum, runs the same CommandPolicy preflight, and calls the same dispatch() path as CLI.

Additive Phase Model

  • Phase 1: Foundation + macOS MVP (54 commands, core engine, macOS adapter)
  • Phase 2: Windows + Linux adapters, 10+ new commands — core untouched
  • Phase 3: MCP server mode via --mcp flag — wraps existing commands
  • Phase 4: Daemon, sessions, enterprise quality gates

Phases 24 add adapters, transports, and production readiness work. Nothing in core is rebuilt.

Coding Standards

File Rules

  • 400 LOC hard limit per file. If approaching 400, split by responsibility. No exceptions.
  • No inline comments. Code must be self-documenting through naming. Only Rust doc-comments (///) on public items when the name alone is insufficient.
  • One struct/enum per file for domain types. node.rs defines AccessibilityNode. action.rs defines Action.
  • One command per file. Each CLI command lives in its own file under commands/. Filename matches the command name.
  • No God objects. No struct with more than 7 fields. No function with more than 5 parameters. Use builder patterns or config structs.
  • Explicit pub boundaries. Only lib.rs re-exports public items. Internal modules use pub(crate). No wildcard re-exports.

Error Handling

  • Zero unwrap() in non-test code. All Results propagated with ? or matched explicitly. Panics are test-only.
  • Every error carries: ErrorCode enum (machine-readable), message: String (human-readable), suggestion: Option<String> (recovery hint), platform_detail: Option<String> (OS-specific detail)
  • All platform adapter functions return Result<T, AdapterError>
  • All command handlers return Result<serde_json::Value, AppError>
  • The binary's main() converts AppError to JSON and sets the exit code

Error Codes

PERM_DENIED, ELEMENT_NOT_FOUND, APP_NOT_FOUND, ACTION_FAILED,
ACTION_NOT_SUPPORTED, STALE_REF, AMBIGUOUS_TARGET, WINDOW_NOT_FOUND,
PLATFORM_NOT_SUPPORTED, TIMEOUT, INVALID_ARGS, NOTIFICATION_NOT_FOUND,
SNAPSHOT_NOT_FOUND, POLICY_DENIED, INTERNAL

Exit Codes

  • 0 — success
  • 1 — structured error (JSON with error code)
  • 2 — argument/parse error

Naming Conventions

Element Convention Example
Crate names agent-desktop-{name} agent-desktop-core, agent-desktop-macos
Module files snake_case, singular snapshot.rs, list_windows.rs
Structs PascalCase, descriptive noun SnapshotEngine, RefAllocator
Traits PascalCase, adjective/capability PlatformAdapter, Executable
Enums PascalCase, variants PascalCase Action::Click, ErrorCode::PermDenied
Functions snake_case, verb-first build_tree(), allocate_refs()
Constants SCREAMING_SNAKE_CASE MAX_TREE_DEPTH, DEFAULT_TIMEOUT_MS
CLI flags kebab-case --max-depth, --include-bounds
Ref IDs @e{n} sequential @e1, @e2, @e14

Platform Crate Folder Structure

All platform crates (macos, windows, linux) follow an identical subfolder layout. New files must be placed in the correct subfolder.

crates/{macos,windows,linux}/src/
├── lib.rs              # mod declarations + re-exports only
├── adapter.rs          # PlatformAdapter trait impl (~175 LOC)
├── tree/               # Reading & understanding the UI
│   ├── mod.rs          # re-exports
│   ├── element.rs      # AXElement struct + attribute readers
│   ├── capabilities.rs # AX-supported actions and settable attributes
│   ├── builder.rs      # build_subtree, tree traversal
│   ├── roles.rs        # Role mapping
│   ├── resolve.rs      # Element re-identification
│   └── surfaces.rs     # Surface detection
├── actions/            # Interacting with elements
│   ├── mod.rs          # re-exports
│   ├── dispatch.rs     # perform_action match arms
│   ├── activate.rs     # Smart AX-first activation chain
│   ├── extras.rs       # select_value helpers
│   ├── scroll.rs       # scroll semantics and gated physical fallback
│   └── type_text.rs    # headless text insertion and physical typing
├── input/              # Low-level OS input synthesis
│   ├── mod.rs          # re-exports
│   ├── keyboard.rs     # Key synthesis, text typing
│   ├── mouse.rs        # Mouse events
│   └── clipboard.rs    # Clipboard get/set
└── system/             # App lifecycle, windows, permissions
    ├── mod.rs          # re-exports
    ├── app_ops.rs      # launch, close, focus
    ├── window_ops.rs   # window operations
    ├── key_dispatch.rs # app-targeted key press
    ├── permissions.rs  # permission checks
    ├── screenshot.rs   # screen capture
    └── wait.rs         # wait utilities

Placement rules:

  • Tree reading/traversal/resolution → tree/
  • Element interaction/activation → actions/
  • Raw OS input (keyboard, mouse, clipboard) → input/
  • App lifecycle, windows, permissions, screenshots → system/
  • adapter.rs stays at root — it's the PlatformAdapter impl that wires everything together

Extensibility Pattern

Adding a new command requires exactly these steps:

  1. Create crates/core/src/commands/{name}.rs with an execute() function
  2. Register it in crates/core/src/commands/mod.rs
  3. Add the CLI subcommand variant to src/cli/mod.rs and arguments under src/cli_args/
  4. Add a match arm in dispatch() in the binary crate
  5. If new Action variant needed, add to crates/core/src/action.rs
  6. If new adapter method needed, add to PlatformAdapter trait with a default returning Err(AdapterError::not_supported())

No existing files are modified beyond the registration points. Enforce via code review.

JSON Output Contract

Every command produces a response envelope:

{
  "version": "2.0",
  "ok": true,
  "command": "snapshot",
  "data": {
    "app": "Finder",
    "window": { "id": "w-4521", "title": "Documents" },
    "ref_count": 14,
    "tree": { ... }
  }
}

Error responses:

{
  "version": "2.0",
  "ok": false,
  "command": "click",
  "error": {
    "code": "STALE_REF",
    "message": "Element could not be resolved from the requested snapshot",
    "suggestion": "Run 'snapshot' to refresh, then retry with updated ref"
  }
}

The error object may also carry an optional details object (e.g. the actionability report on an actionability failure, candidate summaries on AMBIGUOUS_TARGET, or the last observed state on a wait TIMEOUT).

Serialization Rules

  • Omit null/None fields (#[serde(skip_serializing_if = "Option::is_none")])
  • Omit empty arrays (#[serde(skip_serializing_if = "Vec::is_empty")])
  • Omit bounds in compact mode
  • ref_count and tree go inside data, not as top-level siblings

Ref System

  • Refs allocated in depth-first document order: @e1, @e2, etc.
  • An element receives a ref when it is addressable for an action: its role is interactive (button, textfield, checkbox, link, menuitem, tab, slider, combobox, treeitem, cell, radiobutton, switch, colorwell, menubutton, incrementor, dockitem), or it advertises an available action regardless of role. Container roles such as scrollarea (Scroll) and disclosure (Expand/Collapse/Click) are not interactive by role but are genuinely actionable, so they are ref-able — scroll / expand / collapse need a ref to target them
  • A bare SetFocus affordance does not qualify on its own (focusability is not a primary action), so inert focusable containers stay ref-less
  • Static text and non-actionable groups/containers do NOT get refs (they remain in tree for context)
  • Refs are deterministic within a snapshot but NOT stable across snapshots if UI changed
  • Snapshot refs are stored by snapshot ID under ~/.agent-desktop/snapshots/{snapshot_id}/refmap.json, with a latest_snapshot_id pointer for commands that omit --snapshot
  • ~/.agent-desktop/last_refmap.json is written only as a latest-snapshot inspection artifact; command code must use RefStore
  • Action commands use optimistic re-identification: (pid, role, name, bounds_hash). Return STALE_REF on mismatch.
  • Progressive traversal: --skeleton clamps depth to 3, annotates truncated containers with children_count. Named/described containers at boundary receive refs as drill-down targets
  • Drill-down: --root @ref starts from a previously-discovered ref with scoped invalidation (only that ref's subtree refs are replaced on re-drill)
  • RefMap size check: write-side guard prevents >1MB refmap files

PlatformAdapter Trait

Core defines PlatformAdapter; platform crates implement it. Methods default to not_supported(), so an adapter only implements what it supports. Read the current signatures in crates/core/src/adapter.rs — notably strict resolution (resolve_element_strict* → STALE_REF on 0, AMBIGUOUS_TARGET on 2+), live reads for the actionability preflight (get_live_*), and is_protected_process (keeps platform-specific process names out of core).

macOS Adapter Gotchas

  • Ancestor-path set, not a global visited set — macOS reuses AXUIElementRef pointers across sibling branches, so a global visited set would prune real subtrees.
  • AXElement memory safety — inner field is pub(crate) (prevents double-free via raw pointer extraction); Clone must CFRetain, Drop must CFRelease.
  • Batch attribute reads — use AXUIElementCopyMultipleAttributeValues (3-5x faster than per-attribute fetches).

Testing

  • Unit tests use an in-memory MockAdapter; golden fixtures in tests/fixtures/ regression-test serialization.
  • macOS CI integration tests drive real apps (Finder, TextEdit, System Settings).
  • tests/e2e/run.sh drives the release binary against the SwiftUI fixture and verifies every effect by independent observation in both headless and --headed mode (see tests/e2e/README.md).

CI Requirements

  • GitHub Actions macOS runner executes full test suite on every PR
  • cargo tree -p agent-desktop-core must not contain platform crate names
  • cargo clippy --all-targets -- -D warnings
  • cargo test --workspace
  • Binary size check: fail if release binary exceeds 15MB

Commands

54 commands spanning App/Window, Observation, Interaction, Scroll, Keyboard, Mouse, Notifications (macOS), Clipboard, Wait, System, and Batch. The full surface and per-command reference live in skills/agent-desktop/. All 54 are implemented on macOS (Phase 1); Windows/Linux (Phase 2/3) target the same surface. Adding a command: see the Extensibility Pattern above.

Non-Goals

  • Does NOT embed or invoke LLMs
  • Does NOT provide a GUI, TUI, or interactive prompt — machine-facing only
  • Does NOT automate web browsers (use agent-browser for that)
  • Does NOT record or replay macros (stateless per invocation until Phase 4 daemon)
  • Does NOT work with custom-rendered or game-engine UIs lacking accessibility exposure

Reference Documents

  • PRD v2.0: docs/agent_desktop_prd_v2.pdf
  • Architecture Brainstorm: docs/brainstorms/2026-02-19-architecture-validation-brainstorm.md
  • Phase 1 Plan: docs/plans/2026-02-19-feat-agent-desktop-phase1-foundation-plan.md