← back to projects

Computer-Use Agent

Agents that read the screen as pixels and drive legacy desktop software that offers no API to call.

Results

What it did

< 2sscreenshot → action
3desktops driven in parallel

The problem

Why it needed building

Businesses need to automate complex desktop workflows — filling forms, navigating legacy software, extracting data from screens — but traditional RPA is brittle and breaks when UIs change. The challenge was building an AI agent that truly understands what it sees on screen and makes intelligent decisions in real-time, handling unexpected states gracefully.

The approach

How it works

Built a production platform on Claude Computer Use API with a Python/FastAPI backend. The agent captures screenshots, sends them to Claude Vision for analysis, receives structured tool-use commands (click, type, scroll), and executes them via desktop automation. A concurrent multi-session architecture allows parallel desktop control across multiple virtual machines. Session state management ensures the agent can recover from errors and resume interrupted workflows.

  1. Screen Capture
  2. Vision Analysis
  3. Action Planning
  4. Execution Engine
  5. State Recovery
  1. 01
    Screen CaptureHigh-frequency screenshot pipeline with change detection
  2. 02
    Vision AnalysisClaude Vision API interprets UI elements, text, and layout
  3. 03
    Action PlanningTool-use system generates precise click/type/scroll commands
  4. 04
    Execution EngineDesktop automation layer executes actions with verification
  5. 05
    State RecoveryError detection and graceful retry with alternative strategies

The trade-off

What it cost

Vision-driven control is slower and costlier per step than scripted RPA, and it is non-deterministic — every action needs a verification read-back before the next one is safe. What that buys is generality: it works on legacy desktop and native UIs where DOM scraping has nothing to grab.

Live demo

Try it yourself

Desktop — Session #1
Name: Mohamed Hosny
Email: mohamed@example.com
Submit
Capturing Screen
Analyzing UI
Planning Action
Executing Click
Verifying Result

Tech stack

Built with

Claude APIComputer UsePythonFastAPIVision AITool UseWebSocketDocker

Working on something like this?