{
  "id": 12856260,
  "title": "Shadow Automator",
  "url": "https://urgent.news/2026/10/08/shadow-automator",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-08T11:31:49.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/bvrmk/shadow-automator-1c1l"
  },
  "original_language": "en",
  "account": "Shadow Automator is a desktop automation tool that operates fully locally and prioritizes privacy. Powered by a vision-language model, it can locate UI elements using plain-text descriptions and click them directly on the user's hardware, eliminating the need for cloud calls. The team behind this project, Texperts, consists of four members each contributing to different aspects of the software development. Bavan Vijaya Raja handles GPU setup and optimization, Ithyaash focuses on architecture and compliance, Pranav leads UI/UX design, and Bala ensures the OS sandbox and testing. The tool is designed to automate repetitive office tasks that don't have APIs, like logging invoices or filling forms. Traditional RPA tools often struggle with minor UI changes, but Shadow Automator uses recent vision-language models to adapt to these changes. It works by capturing the screen, removing sensitive information, sending the image to a locally hosted model, receiving bounding box coordinates, projecting glowing AR overlays, and clicking the target element. It also includes features like self-healing loops for automatic re-grounding and live ROI dashboards to track cost savings. The implementation relies on AI models like Qwen2.5-VL-3B hosted on a consumer GPU, screen capture via mss, PII redaction through OpenCV, mouse automation with PyAutoGUI, and a custom UI built with CustomTkinter, Tkinter, and PyQt5. All operations occur locally, with no outbound network calls during automation.",
  "summary": "Shadow Automator A fully-local, privacy-first desktop automation tool powered by a vision-language model. It sees your screen, finds UI elements by plain-text description, and clicks them — all on your own hardware, with zero cloud calls. Team Team Name: Texperts | Team Code: HTF 009 Member Role Contribution Bavan Vijaya Raja M.K GPU Lead (6GB VRAM) llama-server setup, Vision-Grounding pipeline,…",
  "key_points": [
    "Shadow Automator operates fully locally, prioritizing privacy.",
    "Powered by vision-language model Qwen2.5-VL-3B for UI element detection.",
    "Designed to automate office tasks without APIs, adapting to UI changes."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}