Urgent.News

What's breaking now, across thousands of outlets.

AI

Running LLMs locally on Linux: what actually works on a Raspberry Pi

Running LLMs locally on Linux: what actually works on a Raspberry Pi A 5B-parameter model can run on a Raspberry Pi 5 every day. It is not fast. It is also completely offline, costs nothing per token, and never phones home. That trade is worth making for a specific class of work, and worthless for everything else. Below is what the WIAIA community has found after several months of experimenting.…

Running Large Language Models (LLMs) locally on Linux, specifically on a Raspberry Pi, can be done successfully for certain use cases. However, there are important hardware constraints and trade-offs to consider.

Firstly, a Raspberry Pi 5 has only 8GB of RAM shared with the rest of the system. When a 4-bit quantized 5B model is loaded, the weights alone take up roughly 3.5 to 4GB of that memory. This leaves limited space for context and other processes, resulting in a usable context window of around 2 to 3k tokens before generation starts failing due to an out-of-memory (OOM) error.

Despite the slow raw speed, the biggest advantage of running LLMs locally on a Raspberry Pi is that the request never leaves the machine. This allows jobs to be left running and checked hours later without any data leaving the device. This is particularly valuable for tasks that require privacy, such as summarizing internal documents or classifying files containing sensitive information.

When it comes to actual implementation, several tools have proven effective on a Raspberry Pi. Ollama is the most common default choice, offering a stable HTTP API and systemd support for automatic restarts after reboots. Llama.cpp via llama-server is another option, providing more control over sampling parameters and the ability to pin threads to specific cores for better performance.

The recommended quantization method for 5B models on a Raspberry Pi is 4-bit quantization with the K=1 setting (Q4_K_1). This strikes a balance between model size and performance, resulting in a model size of around 3.5GB on disk. Higher quantization levels, such as Q8, significantly increase the model size without providing noticeable performance gains. Going too low, like Q2, can degrade the output quality and make the model produce nonsensical responses.

One important aspect to note is that quantization incurs a loss of information. Reasoning-heavy tasks may result in confidently incorrect answers, as the model may not explicitly state "I don't know" but rather provide plausible-sounding but inaccurate responses. Prompting a small model requires a different approach compared to larger models. It is best to focus on one task per prompt, keep the output short (typically 3 sentences), and be explicit about the desired format in the system prompt.

One of the key benefits of running LLMs locally is that the data remains on the device. This eliminates concerns about data leakage and the need for expensive hosted API usage. Jobs can be queued up for offline processing without worrying about rate limits or token costs. Additionally, since the model weights are stored locally, it is easy to audit and verify what is running on the device.

However, local inference on a Raspberry Pi has its limitations. The limited memory makes it unsuitable for tasks requiring real-time interaction or large-scale processing. Interactive autocomplete in text editors or code completion for large files are examples of workloads that would struggle on this hardware. The model's output quality is also capped, and it may not perform as well as more powerful models trained on larger datasets.

In conclusion, running LLMs locally on a Raspberry Pi can be a viable and beneficial option for certain use cases, particularly those requiring privacy and offline processing. By choosing the right tools, quantization settings, and understanding the hardware constraints, users can harness the power of LLMs on a personal device. However, it is essential to recognize the limitations of this setup and avoid trying to use it for interactive or highly demanding tasks that are better suited for more powerful hardware.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Buddy: A Private AI Companion Built for a Friend with Local Gemma

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend. What I Built Meet Buddy — a personal AI companion I built for a friend who wanted a more natural and private way to…

  • Buddy is an AI companion for private, natural interactions
  • Uses local Gemma model via Ollama, no cloud data transmission
  • Features voice interaction with ElevenLabs for natural responses

SANKALP: I Built an AI That Turns School Concepts into Interactive 3D Experiences

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built SANKALP for my younger brother, who is currently around Class 6 .

  • SANKALP is AI-powered 3D visualization system for school concepts.
  • System helps students understand complex 3D structures like mitochondrion.
  • Open innovation approach used open technologies for building SANKALP.

More from Sunday 4 October →