← Back to projects

exploration · ongoing

Android RAG

Building an entirely on-device Retrieval-Augmented Generation pipeline for educational content on Android.

AI / Edge · Vector Search · SQLiteKotlin · SQLite FTS5 & Room · ONNX Runtime & ARM NEON · Text Embeddings (bge-small / MiniLM) · Small Language Models

Overview

In educational and classroom technology, generic AI chatbots are rarely helpful. Teachers and students need answers grounded in specific curricula: textbook chapters, lesson plans, study guides, and classroom notes.

While a small language model (SLM) can converse fluently, it lacks verified domain facts and readily hallucinates historical dates, scientific formulas, and curriculum standards. Retrieval-Augmented Generation (RAG) bridges this gap by fetching relevant source passages and prompting the model to answer strictly based on provided evidence.

This project engineered a complete, production-viable RAG architecture running 100% locally on Android devices without external vector databases, cloud APIs, or active internet connectivity.

The problem

Cloud RAG architectures assume massive infrastructure: high-memory vector stores (Pinecone, Milvus), multi-gigabyte embedding services, and unconstrained context windows. On Android:

  1. No External Vector Database: Vector storage, indexing, and retrieval must live inside local device storage (SQLite / Room).
  2. Context Window Scarcity: Edge SLMs typically operate within a strict 1,024 to 2,048 token window. Dumping large retrieved documents overflows the prompt budget, leaving no space for generation.
  3. Latency Envelope: The entire pipeline—embedding user queries, searching thousands of vectors, assembling context, and initiating model generation—must complete in milliseconds to feel responsive.

Constraints

  • Storage Footprint: The vector index, document database, and embedding model combined must not exceed 250MB.
  • Retrieval Latency Budget: Vector generation + similarity search must complete under 250ms on ARM silicon.
  • Zero Internet Requirement: Teachers must be able to load textbook PDFs onto a device and query them in air-gapped rural schools.

My approach

We designed an end-to-end edge pipeline combining lightweight ONNX embeddings, SQLite persistence, and ARM NEON-accelerated similarity search:

User Question: "What is Boyle's Law?"
             │
             ▼
[On-Device Embedding Model] (ONNX Runtime, INT8)
             │
             ├──► 384-dimensional Query Vector
             │            │
             │            ▼
             │   [ARM NEON SIMD Dot Product]
             │            │
             │            ├──► Top-K Vector Matches ──┐
             ▼                                        ▼
   [SQLite FTS5 BM25 Engine] ────────► Keyword Matches ──► [Reciprocal Rank Fusion]
                                                                │
                                                                ▼
                                                    Filtered Context Passages (max 384 tokens)
                                                                │
                                                                ▼
                                                    [Local SLM Prompt Assembly]
                                                                │
                                                                ▼
                                                    Grounded Answer Generation
  1. Document Chunking: Input educational materials are chunked into 200–250 token segments with a 30-token rolling overlap, preserving paragraph boundaries.
  2. Quantized Embedding Engine: We deployed a quantized bge-small-en-v1.5 model (INT8 via ONNX Runtime Mobile), converting chunks into compact 384-dimensional vector representations.
  3. Hybrid Search (Dense + Sparse): Pure vector search often fails on exact alphanumeric queries (e.g., “Theorem 3.4” or “NaCl”). We combined SQLite’s native FTS5 full-text keyword search with dense vector cosine similarity, merging candidates via Reciprocal Rank Fusion (RRF).

Important decisions

During testing, students queried specific terms like “Formula for kinetic energy” or “Exercise 12.2”. Dense vector embeddings matched general physics concepts but frequently ranked the exact exercise paragraph lower than conceptual summaries. Combining SQLite FTS5 (BM25) with vector similarity eliminated this blind spot entirely.

2. INT8 Quantized Vector Storage

Storing 10,000 float32 vectors (384 dimensions) requires over 15MB of raw float data in SQLite. By quantizing stored document vectors to signed 8-bit integers (Int8Array) and adjusting cosine distance calculations, index storage shrank by 75% with negligible (<1.2%) degradation in Top-k retrieval recall.

3. Context Budget Allocation

Given a 1,024-token budget on the local SLM:

  • System Instructions & Guardrails: 150 tokens
  • Retrieved Context Passages (Top 2–3 chunks): ~450 tokens
  • User Query & Chat History: 150 tokens
  • Reserved Output Generation Budget: 274 tokens

What went wrong: The JVM Vector Bottleneck

Our first working prototype calculated cosine similarity across 5,000 document vectors in pure Kotlin:

// Naive JVM iteration:
fun cosineSimilarity(v1: FloatArray, v2: FloatArray): Float {
    var dot = 0f; var normA = 0f; var normB = 0f
    for (i in v1.indices) {
        dot += v1[i] * v2[i]
        normA += v1[i] * v1[i]
        normB += v2[i] * v2[i]
    }
    return dot / (sqrt(normA) * sqrt(normB))
}

When querying an index with 5,000 textbook chunks, this loop took 680ms. Under continuous search, it froze the UI thread and consumed unacceptable battery power.

Debugging & Investigation

Profiling identified two root causes:

  1. Garbage collection churn from retrieving and unboxing arrays from SQLite cursors.
  2. Lack of SIMD auto-vectorization in the Android Runtime (ART) JIT compiler.

The Fix: ARM NEON SIMD Acceleration

We shifted vector distance calculations into native C++ using ARM NEON SIMD intrinsics. A single 128-bit NEON vector instruction processes four 32-bit floats (or sixteen 8-bit integers) in a single CPU cycle:

#include <arm_neon.h>

float dot_product_neon(const float* a, const float* b, int n) {
    float32x4_t sum = vdupq_n_f32(0.0f);
    for (int i = 0; i < n; i += 4) {
        float32x4_t va = vld1q_f32(a + i);
        float32x4_t vb = vld1q_f32(b + i);
        sum = vmlaq_f32(sum, va, vb);
    }
    return vaddvq_f32(sum);
}

Result: Vector similarity calculation time across 5,000 vectors dropped from 680ms to 12ms—a 56× speedup!

Result

  • Sub-220ms total retrieval latency: Query embedding + hybrid search completes invisibly before the user even finishes glancing at the screen.
  • High-fidelity grounded generation: Grounding local SLMs in curriculum source chunks virtually eliminated factual hallucination.
  • 100% Offline Autonomy: Teachers can query documents in any classroom without requiring an internet uplink.

What I learned

  1. Context hygiene is paramount: With small models, the quality and brevity of retrieved context dictates answer accuracy far more than model parameter size.
  2. SIMD is a superpower on ARM: Never execute brute-force numerical loops on the Android JVM when 15 lines of C++ with ARM NEON can do the work 50× faster.
  3. Hybrid search is essential for structured content: Dense semantic vectors and sparse BM25 keywords complement each other’s weaknesses perfectly.