{
  "id": 13049372,
  "title": "No API Keys, No Cloud Bills: Running a Document Pipeline Entirely On-Device",
  "url": "https://urgent.news/2026/10/09/no-api-keys-no-cloud-bills-running-a-document-pipeline-entirely-on",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-09T06:34:56.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/mininglamp/no-api-keys-no-cloud-bills-running-a-document-pipeline-entirely-on-device-9i8"
  },
  "original_language": "en",
  "account": "A document processing pipeline was set up entirely on a Mac mini, without using any API keys or cloud inference. The team built tooling for due diligence workflows, handling contracts, financial statements, and scanned images that clients would not allow to be uploaded anywhere. They tried using cloud vision APIs with audit logging, but compliance still said no. This led them to develop an on-device GUI agent called Mano-P, using a 4B parameter model with W8A8 quantization. The agent can see the screen, navigate file managers, open PDFs and images, and does all of this without any data leaving the machine. A typical session involves the agent searching the local filesystem for a target archive, unzipping it with a password stored in a separate local note, reading the contents, and processing batches of contracts with the same extraction routine. The 4B model runs at around 80 tokens per second decode speed, with GUI interaction latency being the bottleneck rather than model inference. A comparison showed the local 4B model had a 56% pass rate and average 7.9 seconds per step, while the cloud-based Qwen3-VL-Plus had a 39% pass rate and average 10.2 seconds per step. The hardware setup cost under $1500, and the team estimated that the Mac mini paid for itself in just two batches compared to the estimated $800 per batch for cloud pricing.",
  "summary": "No API Keys, No Cloud Bills: Running a Document Pipeline Entirely On-Device We run a document processing pipeline on a Mac mini. No API keys. No cloud inference. No network calls at all during execution. This post is about what that actually looks like in practice, what went wrong along the way, and where we landed on performance. Why we went local Our team builds tooling for due diligence…",
  "key_points": [
    "Document processing pipeline runs entirely on-device without API keys or cloud inference",
    "Mano-P GUI agent uses 4B parameter model with W8A8 quantization for on-device processing",
    "Local 4B model achieves 56% pass rate vs 39% for cloud-based Qwen3-VL-Plus"
  ],
  "editors_take": "This on-device document processing setup marks a significant shift in how sensitive data is handled, allowing for secure local processing and potentially saving substantial costs compared to cloud-based solutions.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}