Urgent.News

What's breaking now, across thousands of outlets.

AI

No API Keys, No Cloud Bills: Running a Document Pipeline Entirely On-Device

No API Keys, No Cloud Bills: Running a Document Pipeline Entirely On-Device We run a document processing pipeline on a Mac mini. No API keys. No cloud inference. No network calls at all during execution. This post is about what that actually looks like in practice, what went wrong along the way, and where we landed on performance. Why we went local Our team builds tooling for due diligence…

A document processing pipeline was set up entirely on a Mac mini, without using any API keys or cloud inference. The team built tooling for due diligence workflows, handling contracts, financial statements, and scanned images that clients would not allow to be uploaded anywhere. They tried using cloud vision APIs with audit logging, but compliance still said no. This led them to develop an on-device GUI agent called Mano-P, using a 4B parameter model with W8A8 quantization.

The agent can see the screen, navigate file managers, open PDFs and images, and does all of this without any data leaving the machine. A typical session involves the agent searching the local filesystem for a target archive, unzipping it with a password stored in a separate local note, reading the contents, and processing batches of contracts with the same extraction routine.

The 4B model runs at around 80 tokens per second decode speed, with GUI interaction latency being the bottleneck rather than model inference. A comparison showed the local 4B model had a 56% pass rate and average 7.9 seconds per step, while the cloud-based Qwen3-VL-Plus had a 39% pass rate and average 10.2 seconds per step. The hardware setup cost under $1500, and the team estimated that the Mac mini paid for itself in just two batches compared to the estimated $800 per batch for cloud pricing.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Friday 9 October →