Computer-Use Agent
Agents that read the screen as pixels and drive legacy desktop software that offers no API to call.
Results
What it did
The problem
Why it needed building
Businesses need to automate complex desktop workflows — filling forms, navigating legacy software, extracting data from screens — but traditional RPA is brittle and breaks when UIs change. The challenge was building an AI agent that truly understands what it sees on screen and makes intelligent decisions in real-time, handling unexpected states gracefully.
The approach
How it works
Built a production platform on Claude Computer Use API with a Python/FastAPI backend. The agent captures screenshots, sends them to Claude Vision for analysis, receives structured tool-use commands (click, type, scroll), and executes them via desktop automation. A concurrent multi-session architecture allows parallel desktop control across multiple virtual machines. Session state management ensures the agent can recover from errors and resume interrupted workflows.
- Screen Capture
- Vision Analysis
- Action Planning
- Execution Engine
- State Recovery
- 01Screen CaptureHigh-frequency screenshot pipeline with change detection
- 02Vision AnalysisClaude Vision API interprets UI elements, text, and layout
- 03Action PlanningTool-use system generates precise click/type/scroll commands
- 04Execution EngineDesktop automation layer executes actions with verification
- 05State RecoveryError detection and graceful retry with alternative strategies
The trade-off
What it cost
Vision-driven control is slower and costlier per step than scripted RPA, and it is non-deterministic — every action needs a verification read-back before the next one is safe. What that buys is generality: it works on legacy desktop and native UIs where DOM scraping has nothing to grab.
Live demo
Try it yourself
Tech stack
Built with
Working on something like this?