Wiring Android's MediaPipe LLM Inference API to a Quantized On-Device Reranker for RAG
--- title : " On-Device RAG on Android: FAISS Retrieval + MediaPipe Cross-Encoder Reranker" published : true description : " Build a full on-device RAG pipeline on Android using FAISS-lite int8 embeddings and a MediaPipe LLM cross-encoder reranker — with real memory budgets and latency on Pixel 9." tags : kotlin, android, architecture, mobile canonical_url :…
This article outlines a method to build an on-device Retrieval-Augmented Generation (RAG) system using Android, MediaPipe, and FAISS. The system uses FAISS-lite for fast retrieval with int8 bi-encoder embeddings and a quantized cross-encoder reranker through MediaPipe's LLM Inference API. The two-stage pipeline is designed to balance speed and relevance, with retrieval being faster but less precise and reranking providing more accurate results.
The memory usage is optimized to fit within the 12GB RAM of a Pixel 9 device, leaving sufficient space for the main LLM context. The author emphasizes the importance of budgeting memory and context window carefully before model size. The system demonstrates that on-device RAG can be practical and efficient, with most of the computation occurring on-device and no reliance on server-based inference.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.