← All releases

PUBLISHED RELEASE

Ankita v2.4.3

2.4.3 ·

Release notes

Browser automation, a live browser stage that follows the app appearance, reliable cancellation and reconnection, filesystem security fixes, and faster command/file/Markdown processing. Detailed changes and verification follow.

Added

  • A packaged desktop verifier exercises the actual executable and shipped modules with disposable settings, a local HTTP/SSE model fixture, appearance switching, browser submission, screenshot pixels, takeover, Stop, command stdin/EOF and the Chromium installer dry run. It downloads no dependencies.
  • browser fill_form fills up to ten fields from one snapshot after checking their refs. A disposable browser verification script covers both backends, screenshots, Stop, reconnect, automatic selection, and preview refresh rate without downloading packages or using a personal browser profile.
  • Browser tools. Ankita can open real pages, read text and snapshots, act on fresh element references, manage tabs, take screenshots, and run batches of up to ten steps. Browser tools load on demand through find_tools.
  • Two browser plugins. Playwright Browser uses a separate Chromium profile; Chrome local connects through the pinned Chrome DevTools MCP server. Desktop setup supports Private Chrome, Debug port, and My session, with Chromium download progress, a background option, and a debug-port connection test. Both plugins start disabled and save their settings across restarts.
  • Live browser beside chat. The stage shows tabs, URL, status, and the current page, with Stop and Take control. Playwright takeover supports clicking, typing, pasting, and scrolling through the preview. Chrome takeover happens in Chrome. Hand back or Esc resumes Playwright control; Ctrl/Cmd+Shift+B toggles the stage.
  • Browser site controls. Each plugin has allowed and blocked domain lists. Browser navigation accepts HTTP/HTTPS; private hosts are blocked by default. The isolated browser also checks page-initiated navigation and resource hosts. Browser actions require approval, and tool-card previews redact entered text.
  • Browser setup and verification documentation, including the browser guide, implementation plan, real Chrome and Playwright benchmark script, desktop UI verification script, and security audit record.

Changed

  • Windows command creation, shell discovery and process-tree work run in a worker so slow native startup can no longer block the owner's event loop. PowerShell selection is preserved. Stdin/output use native acknowledgements and backpressure; input waits for launch readiness before its write deadline.
  • Browser plugin cards now match the existing integration cards in a compact By Ankita team section. Settings open in a details dialog. Plugin cards, dialogs, and the browser stage follow the app's global appearance colors.
  • The browser stage uses the app's restrained tab, address-bar, and footer styling. Opening it collapses the left sidebar and closes workspace review to give chat and the page more room.
  • Visible browser previews target ten updates per second with one capture in flight. Chrome returns JPEG data directly through MCP, avoiding temporary screenshot files. Live frames update independently of chat; tab metadata and chat thumbnails refresh every two seconds. Hidden/lost-connection states poll less frequently. Actual refresh rate depends on screenshot latency.
  • File inspection reuses a resident worker instead of starting one for every call. A local 30-call README sample measured a 3.37 ms warm median and 5.64 ms p95 after 174 ms initial startup.
  • Terminal Markdown redraws coalesce on a fixed 20 ms deadline. Syntax keyword and type Sets are reused, skill catalogs load once per turn, tool schemas are shared within each model round, and tool-call signatures are computed once. History trimming measures message/group costs incrementally and caches image costs instead of repeatedly measuring the entire conversation.
  • Added the Playwright runtime dependency, pinned the Chrome MCP launch version, and raised the desktop IPC contract to 5 for browser events and recovered-message replacement.

Fixed

  • Windows command cleanup keeps the launcher referenced until every pending Stop acknowledgement settles. Native pipe closure can arrive first on fast machines; it no longer leaves cleanup promises unresolved or cancels the remaining command tests on the clean release runner.
  • Recoverable browser errors no longer mark a healthy session disconnected or dump fresh snapshots, page code and raw call logs into the live sidebar. Page/action notices preserve preview refresh and Take control; only connection and startup failures prompt setup. Completed actions keep the last frame.
  • An unexpected Windows launcher worker exit cleans up its identity-checked native process tree before reporting failure; changed or missing process identities are refused instead of killing an unrelated PID.
  • Chrome typing no longer sends an unsupported snapshot argument. Keyboard presses focus the supplied element ref without clicking it, and refuse to send a key when that control cannot receive focus.
  • Browser auto mode uses enabled-backend selection. Tool guidance requires exact opaque refs rather than invented selectors or labels. Invalid refs return a fresh snapshot without replaying the mutation; successful open and action results provide the next refs, avoiding redundant snapshot calls.
  • Playwright snapshots include visible controls beyond hidden input lists, accessible labels and states, custom roles, open shadow roots, and child frames. Label queries reach controls beyond snapshot limits. Password values remain hidden. Re-rendered refs follow only a unique control in the original frame; ambiguous replacements fail safely. Form filling supports selects and checkboxes as well as text fields.
  • Chrome binds refs to their tab, excludes static text from actionable refs, and uses snapshots included in native action and form-fill responses to reduce MCP round trips.
  • Successful browser interactions reset prior browser-read repetition counts, allowing multi-step workflows to continue. Snapshot-only loops and existing tool/round limits still stop. Schema budgeting reserves the actual system prompt and user attachments so tool discovery does not crowd out the task.
  • Browser screenshot receipts attach validated PNG pixels to the next model request through the existing user-image format, retaining text tool messages and original input attachments. Workspace, file, byte, signature, and image count limits apply. MCP shutdown now observes asynchronous process teardown failures rather than leaving an unhandled rejection; the older Chrome verification script handles the adapter's structured screenshot receipt.
  • Disabling Chrome prevents further selection and routes the next page open to enabled Playwright. New sessions prefer Playwright when both plugins are enabled; an existing session keeps its enabled backend. Old tab IDs and element references cannot silently execute against a replacement browser.
  • Chrome connects when a browser request needs it and makes one recovery attempt after losing its connection. Approval is remembered for the exact server command. Read recovery gets fresh references; clicks and form submissions are never automatically replayed.
  • Composer Stop cancels active and queued browser calls, including calls paused for takeover or setup approval. Closing the stage stops the run and disconnects Chrome. Browser/MCP deadlines and transport-close rejection prevent lost connections from waiting indefinitely; cancelled turns avoid spurious errors.
  • Preview captures share pending work, recover their ready state after an error, and avoid retaining stale error frames. Browser sessions close during app shutdown.
  • Interrupted model streams retry the current model step up to twice, preserve completed tool results, discard incomplete tool requests, and replace unfinished text when recovered wording differs. Stop cancels recovery.
  • Desktop replies use consistent Markdown guidance across models. The message view repairs common plain-text headings, bullet glyphs, tabular rows, and whole-answer Markdown fences while preserving code blocks and existing markup.
  • Deferred MCP server activation uses namespaced IDs, preventing collisions with built-in tool names. Browser discovery loads the first-party browser capability and retains registry guidance for other missing capabilities.
  • File tools enforce the workspace boundary for absolute and relative paths, including Git inventories. They reject traversal escapes, interior symlinks/junctions, Windows device aliases, and alternate data streams while supporting a junction at the workspace root. Delete/move refuse both the workspace and current-directory roots.
  • Approved edits, whole-file writes, and patches retain the displayed plan and recheck arguments, workspace identity, file identity, permissions, and bytes. Files changed during approval require a new approval. Atomic writes preserve permissions and use exclusive, unpredictable staging files with no direct-write fallback, closing the temporary-file symlink escape.
  • Empty approval previews no longer bypass required confirmation. mcp_manage declares its approval policy by action. Auto-approval and confirmation results require explicit boolean true, including daemon gates; auto-approval remains an opt-in and still prepares file/process snapshots.
  • Approval fallback and MCP diagnostics redact recognizable credentials plus secrets from environment variables, config files, server settings, and headers. Stderr buffering prevents split-chunk leaks; ordinary command descriptions remain readable.
  • MCP stdout is capped at 16 MiB per message, including unterminated frames. Oversize closes the transport and rejects outstanding calls. Stderr lines are capped at 16 KiB and discarded through their newline when oversized. Split UTF-8 is decoded correctly.
  • A filesystem worker watchdog terminates pathological regex searches or cancelled inspections, then serves queued healthy reads on a replacement. Inspection admission is bounded, and the worker shuts down explicitly with the CLI.
  • Renderer finish flushes pending text, cancels redraw timers, removes resize listeners, and is idempotent. CLI stream failure and replacement paths finish the renderer. History trimming preserves input attachment objects, and optional recall warm-up failures cannot become unhandled rejections.
  • Daemon SIGINT uses the daemon shutdown path, wakes tick sleep, settles queued permits, denies pending approvals, and drains active turns and alert composition. Queued work does not start after Stop.
  • Failed alert enqueue retains one prepared message for delivery retries without repeating the LLM call. New watch changes coalesce from the earliest baseline to the latest reading; the pending backlog caps at 100 distinct changes and logs displacement of its oldest entry.
  • Voice payload reads, writes, and cleanup use asynchronous I/O, including daemon audio staging. Stop/abort during a write cannot start delayed playback or transcription; playback removes its abort listener after exit.
  • The full test command runs serially to avoid Windows process and MCP test timeouts caused by parallel execution.

Verification

  • The clean GitHub Windows release run passed 691 tests, 0 failed, 0 cancelled, 23 skipped (714 total); optional Chromium/Python coverage ran locally instead. All four published release assets were downloaded and their SHA-256 digests checked. The installer's SHA-512, size and version match the update manifest.
  • Release preflight on 2.4.3: 713 passed, 0 failed, 1 skipped (714 total), desktop production build and rendered sidebar round trip passed. The CLI's declared Node minimum now matches Playwright's Node 20 requirement; desktop development requires Node 22.12+ for the Electron build dependencies.
  • Browser continuation: both real backends completed form, advanced controls, screenshot, Stop, reconnect and disabled-Chrome fallback checks. Real stale-ref recovery keeps the live pane usable without displaying snapshot/code text. The packaged executable passed appearance, eight HTTP/SSE model rounds, browser submission, screenshot pixels, takeover, Stop, command stdin/EOF, launcher crash cleanup and Chromium installer dry run. Chrome used cached MCP 1.8.0; configured 1.10.1, real downloads, personal attachment and live provider/site completion remain untested. Traces and coverage are recorded in browser findings.
  • Latest full gate: 713 passed, 0 failed, 1 skipped (714 total). The previous command startup/output failures now pass without weaker deadlines. Desktop TypeScript checking and production build passed; git diff --check and new browser-file whitespace checks passed.
  • Earlier security implementation verification: 660 tests passed, zero failed, one expected Windows skip for POSIX executable permissions. Desktop TypeScript checking and the Vite production build passed, as did git diff --check. Regression tests cover workspace escapes, approval races, stuck workers, credential redaction, stream recovery, browser cancellation, rendering, daemon shutdown, alert retries, and audio staging.
  • The review's reported voice-default, provider-config, and web-fetch failures did not reproduce: their 62 tests passed without changing expected behavior.

Known limits

  • Website challenges and rejected navigation can still prevent a task. The supplied Air India booking URL failed with HTTP/2 in background Chromium and returned 404 in visible Chromium; the sidebar now reports a short page notice.
  • Live Chrome verification used cached MCP 1.8.0 rather than configured 1.10.1. Personal-session permission dialogs, real Chromium downloads, non-Windows packaged builds and live model/vision-provider completion remain uncovered.
  • Closed shadow roots, rare cross-origin frame changes and some Chrome widget operations remain limited. Failed mutations are never automatically replayed, and partial form fills require inspection before continuing.

YOUR NEXT LITTLE COMPANION

Make some room
for Ankita.

Version 2.5.2 · Windows x64

Choose a provider during setup. Model availability and usage limits depend on your provider.

Release notes & other assets