← Back to writing

Building RAG on Android

Why local retrieval-augmented generation is hard on a tablet, and how the pipeline is actually put together.

  • Android
  • AI
  • RAG
  • Edge Computing

The question

When developers discuss building Retrieval-Augmented Generation (RAG) applications, the architecture usually relies on a constellation of cloud services: LangChain or LlamaIndex in Python, Pinecone or Milvus for vector storage, OpenAI or Cohere for embeddings, and an 8-billion parameter model for answering.

When you attempt to take that entire pipeline and condense it onto an Android tablet or an interactive classroom display without an internet connection, a practical engineering question arises:

What breaks first when you try to run an end-to-end RAG pipeline entirely on edge Android hardware?

My initial mental model

When I started experimenting with edge AI, I assumed:

“The language model is the hard part. Retrieval is just calculating dot products across an array of numbers.”

I assumed getting the generative model running smoothly would take 90% of the effort, after which plugging in embeddings and vector lookup would be a trivial weekend task.

That intuition was completely backward.

What actually happens

Running a small language model (SLM) on edge silicon with quantized runtimes like ONNX or llama.cpp is surprisingly deterministic. Once weights are memory-mapped into virtual memory, generation proceeds token by token at a steady pace.

What actually broke first in the real world was retrieval quality and context economics:

[Document Text] ──► [Chunker (200 tokens + Metadata)]
                           │
                           ▼
            [On-Device Embedding Model] (ONNX INT8)
                           │
                           ├──► 384-dim Query Vector ──► [ARM NEON Cosine Match]
                           ▼                                      │
            [SQLite FTS5 BM25 Match] ─────────────────────────────┘
                           │
                           ▼ (Reciprocal Rank Fusion)
                [Top 2 Context Passages] (max 450 tokens)
                           │
                           ▼
            [Edge SLM (Qwen2.5 / SmolLM2)] ──► Grounded Answer

1. Context Scarcity Dictates Everything

On server models like Claude or GPT-4, you have 128K+ token context windows. You can retrieve ten 1,000-word sections, dump them into the prompt, and the model sorts through the noise.

On edge devices, memory limits cap your local SLM context window to 1,024 or 2,048 tokens. If your retriever returns two marginally relevant, bloated chunks:

  • The context window is choked.
  • The model forgets instructions or runs out of tokens before finishing its explanation.
  • The model hallucinates trying to bridge disparate, irrelevant paragraphs.

2. Pure Vector Search Misses Obvious Facts

Educational materials are full of specific identifiers: “Chapter 4, Section 2”, “Theorem 10.1”, or chemical formulas like “CH4 + 2O2”.

Dense embedding models project text into high-dimensional semantic space. In semantic space, “Chapter 4, Section 2” has almost identical cosine distance to “Chapter 4, Section 3”. As a result, dense vector search routinely pulled the wrong section.

We had to implement hybrid search: pairing SQLite’s native FTS5 full-text engine (BM25 keyword matching) with dense vector embeddings.

Let’s look at the implementation

Here is how document chunks are structured and embedded locally using ONNX Runtime Mobile:

data class DocumentChunk(
    val id: String,
    val documentTitle: String,
    val sectionHeader: String,
    val text: String,
    val tokenCount: Int,
    val embedding: FloatArray // 384 dimensions from bge-small
)

class LocalRAGPipeline(
    private val embeddingEngine: OnDeviceEmbeddingModel,
    private val vectorDb: LocalVectorDatabase,
    private val slmEngine: LocalInferenceEngine
) {
    suspend fun answerQuestion(question: String): Flow<String> = flow {
        // 1. Generate query vector locally (< 40ms)
        val queryVector = embeddingEngine.encode(question)

        // 2. Hybrid retrieval combining BM25 keyword search & vector distance (< 25ms)
        val topChunks = vectorDb.searchHybrid(
            query = question,
            queryVector = queryVector,
            limit = 2
        )

        // 3. Assemble tight context budget
        val contextPrompt = buildString {
            append("Answer using ONLY the following verified notes:\n\n")
            topChunks.forEach { chunk ->
                append("### [${chunk.sectionHeader}]\n")
                append(chunk.text).append("\n\n")
            }
            append("Question: $question\nAnswer: ")
        }

        // 4. Stream response from local SLM
        slmEngine.generateStream(contextPrompt).collect { token ->
            emit(token)
        }
    }
}

Notice the critical metadata detail: prepending ### [${chunk.sectionHeader}] before each chunk’s text gives the SLM explicit structural awareness without requiring larger chunk sizes.

The surprising part: Retrieval Trumps Model Size

During internal evaluation, we tested two configurations:

  1. Configuration A: A larger 3B parameter model paired with a naive vector retriever (retrieval precision: ~65%).
  2. Configuration B: A compact 1B parameter model paired with our ARM NEON hybrid retriever (retrieval precision: ~92%).

Configuration B drastically outperformed Configuration A in factual accuracy and hallucination rate.

A smaller model presented with surgically accurate, concise source evidence answers with high fidelity. A larger model fed noisy, semi-relevant context often confabulates plausible-sounding errors.

In edge AI, improving retrieval quality yields exponentially higher ROI than fighting to fit a slightly larger model onto the device.

Practical consequences for Android developers

1. Keep Chunks Small and Overlapped

Chunk size should match your model’s context budget. For a 1,024-token SLM, chunks of 180–220 tokens with a 30-token overlap proved to be the sweet spot.

2. Offload Similarity Calculations to C++ / NEON

Never calculate cosine distance or dot products over thousands of vectors in JVM loops. Delegating to a 15-line C++ function compiled with ARM NEON SIMD cuts retrieval calculation times from ~650ms to ~12ms.

3. Graceful Fallback

If retrieval similarity scores fall below a strict confidence threshold (e.g. cosine distance < 0.65), do not force the SLM to guess. Have the system immediately return: “I couldn’t find information about that in your offline notes.”

What I would remember

  1. Retrieval quality is the real bottleneck: Spend time optimizing chunking, hybrid keyword matching, and metadata indexing before tuning model parameters.
  2. Context economics matter: On edge hardware, every retrieved token costs precious memory and generation headroom. Be surgical.
  3. Hybrid search is mandatory: Semantic vectors capture meaning; BM25 captures exact tokens. You need both.

Further reading

  • Related Project: Android RAG
  • Related Project: On-Device AI
  • ONNX Runtime: Accelerating Mobile & Edge Inference with ONNX Runtime