{
  "id": 11545187,
  "title": "Running a Small LLM Locally to Parse Bank Notifications: Structured Output with Ollama, Pydantic and a Regex Safety Net",
  "url": "https://urgent.news/2026/10/02/running-a-small-llm-locally-to-parse-bank-notifications-structured",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-02T22:49:08.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/eme_gug_0821b41b948be6516/running-a-small-llm-locally-to-parse-bank-notifications-structured-output-with-ollama-pydantic-4k02"
  },
  "original_language": "en",
  "account": "Recently on Dev.to, a noteworthy project called \"Kharcha\" has been highlighted. This project utilizes a 4B model running locally to parse Indian bank SMS notifications, handling financial data without leaving the user's laptop. On Hacker News, antirez, the creator of Redis, is also exploring running LLMs locally. The author believes this approach aligns well with Vietnamese developers. If you frequently receive notifications of account balance changes from banks like Vietcombank, Techcombank, or MB, setting up a quick app to manage expenses locally is the fastest option. Sending account numbers, balances, and transaction details to a third-party server raises security concerns. This article shares the pipeline the author uses: Ollama + a 3-4B model + JSON Schema structured output + Pydantic validation + regex fallback. This setup works on MacBook M1 8GB or mid-range laptops with GPUs. Running a small model locally for this task may seem unnecessary to some, but it's beneficial for structured extraction tasks. The author's model does not need to write poetry; it only requires reading a short text snippet and returning a few fields: amount, transaction type (credit/debit), balance, content, and time. The pipeline faces several challenges: short inputs (under 300 tokens), a fixed schema with grammatical constraints, variable bank templates, and sensitive data. The author tested models like qwen2.5:3b, llama3.2:3b, and gemma3:4b. Among them, Qwen2.5 3B performed the best for Vietnamese text. Gemma3 4B was also good but slower. The flowchart of the pipeline is: SMS/Push notification -> Pre-filter: is this a bank notification? -> If not, skip; if yes, proceed to Ollama local + JSON Schema -> Pydantic validate -> If valid, cross-check with regex; if not, use regex fallback. The final output is stored in SQLite on the local machine. To set up Ollama and select the model, install Ollama 0.5 or later, and use llama.cpp's grammar-constrained decoding, which prevents the model from generating invalid JSON. This feature is crucial for the pipeline. The author tested Qwen2.5:3b, llama3.2:3b, and gemma3:4b for Vietnamese SMS notifications. Qwen2.5 3B gave the most stable results, while Gemma3 4B was slightly slower.",
  "summary": "Tuần này trên Dev.to có một project khá hay: Kharcha, dùng model 4B chạy local để đọc SMS ngân hàng ở Ấn Độ, dữ liệu tiền bạc không rời khỏi laptop. Trên HN thì antirez (tác giả Redis) cũng đang làm chuyện chạy LLM local. Mình thấy ý tưởng này hợp với dev Việt Nam. Ngày nào mình cũng nhận cả chục tin báo biến động số dư từ Vietcombank, Techcombank, MB... Muốn tự làm app quản lý chi tiêu thì cách…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}