exploration · ongoing
On-Device AI
Practical SLM inference and hardware-accelerated model deployment on constrained Android hardware.
Overview
Cloud-hosted LLMs dominate developer discussions, but in real-world educational and enterprise deployments, cloud dependency introduces severe trade-offs: recurring API costs, unpredictable latency jitter, data privacy compliance, and complete failure whenever institutional Wi-Fi drops.
This project explored the feasibility of bringing generative AI completely on-device for Android appliances powered by embedded ARM SoCs (specifically Rockchip RK3576 platforms featuring a dedicated 6 TOPS NPU). The goal was not to build a toy demonstration, but to determine whether small language models (SLMs) can deliver practical, reliable, zero-latency assistance while coexisting with a 60fps 4K user interface.
The problem
Deploying large models on mobile hardware hits a physical wall:
- Memory Ceiling: A typical 7B model at 16-bit precision requires ~14GB of RAM just to hold weights. On dedicated Android appliances with 4GB–8GB of unified system RAM, the AI process would immediately trigger Android’s Low Memory Killer (LMK).
- Thermal & Bus Contention: Continuous inference saturates memory bus bandwidth. If model matrix operations compete with the display compositor (
SurfaceFlinger), the screen stutters visibly. - Usability Threshold: Any generation slower than human reading speed (~6–8 tokens/second) feels unusable.
Constraints
- System RAM Budget: The AI runtime cannot exceed 1.2GB–1.5GB total RSS without endangering foreground applications.
- Compute Substrate: Rockchip RK3576 with an octa-core CPU (4× Cortex-A72 + 4× Cortex-A53) and an integrated 3-core NPU rated at 6 TOPS (INT8).
- Latency Budget: Time-To-First-Token (TTFT) under 1.5 seconds; sustained streaming generation above 12 tokens/sec.
- Complete Air-Gapped Operation: Core text summarization and semantic lookup must function with zero network access.
My approach
We framed on-device AI as an embedded systems problem rather than a pure machine learning problem:
Application Layer (Compose UI)
│ ▲
Streaming Flow │ │ Tokens (15+ tok/s)
▼ │
[Android NDK Inference Bridge (JNI)]
│
┌──────────────────┴──────────────────┐
│ │
Model Weights via mmap() NPU Execution Graph
(Zero JVM Heap Allocation) (RKNN / Int8 Quantized)
│ │
▼ ▼
Flash Storage (eMMC) RK3576 6 TOPS Co-processor
- Model Architecture Selection: Rather than forcing large models to shrink, we started with modern compact architectures engineered for edge efficiency: models in the 0.5B to 1.5B parameter range (such as Qwen2.5-0.5B/1.5B and SmolLM2).
- Aggressive Quantization: Models were quantized to 4-bit (AWQ / GGUF Q4_K_M) and symmetric INT8 targeting the NPU vector instructions. This brought model weight footprints down from ~3.2GB to under 750MB.
- Zero-Copy Memory Mapping (
mmap): We bypassed the Android JVM entirely. Model weights are mapped directly from flash storage into virtual memory via native POSIXmmap()calls in C++. The Linux kernel manages paging on demand without triggering ART garbage collection.
Important decisions
1. NPU vs. GPU Compute Placement
Initial experiments utilized the GPU via Vulkan / OpenCL compute shaders. While throughput reached ~10 tokens/s, running continuous General Matrix Multiplies (GEMM) on the GPU starved RenderThread and SurfaceFlinger. Frame rates dropped from 60 fps to 28 fps during generation.
We shifted matrix operations to the dedicated Rockchip NPU using vendor graph compilation (RKNN-Toolkit). By moving inference off the GPU, the display compositor maintained a locked 60 fps regardless of AI load.
2. Pre-Allocated KV Cache Ring Buffer
Dynamic key-value (KV) cache resizing causes unpredictable heap reallocation spikes during inference. We pre-allocated a static KV cache buffer capped at a 1,024-token context window. This locked memory consumption to a predictable, flat ceiling.
3. Asynchronous Streaming via Kotlin Channels
Native C++ inference callbacks emit individual token IDs into an asynchronous JNI boundary, feeding a Kotlin Channel<String> which safely updates Jetpack Compose text components without dropping frames.
What went wrong: The Thread Throttling Trap
In early NDK builds using CPU fallback, we initialized thread pools to utilize all 8 CPU cores (std::thread::hardware_concurrency()).
Paradoxically, token throughput was slower on 8 cores than on 4 cores. Even worse, the SoC reached 82°C within 45 seconds, triggering thermal frequency throttling across all cores.
Profiling with top and core affinity tools revealed two problems:
- Four of the eight cores are low-power Cortex-A53 efficiency cores. Synchronizing parallel matrix calculations across mismatched core architectures caused the fast A72 cores to spend cycles waiting for slow A53 threads at thread barrier sync points.
- The thermal load rapidly tripped kernel governors.
The Fix
We used Linux sched_setaffinity system calls to restrict inference strictly to the four Cortex-A72 performance cores, letting the efficiency cores handle Android OS background services and keeping thermal equilibrium under 58°C.
Result
- 14–18 tokens/second generation speed: Fast, comfortable streaming text generation on local silicon.
- Sub-800MB memory footprint: Coexists reliably alongside memory-heavy classroom canvas tools.
- Zero UI frame drops: Dedicating the NPU and isolating CPU threads preserved butter-smooth 60 fps 4K UI rendering.
What I learned
- Coexistence over raw benchmark speed: In an embedded consumer product, an AI model that hits 30 tokens/sec but freezes the screen for 200ms is a failure. Balancing memory buses and core affinities is the real challenge.
- Small models are genuinely capable: For domain-specific tasks—summarizing text, rephrasing explanations, extracting key terms—a fine-tuned 1B model running locally often beats a 70B cloud model bogged down by network roundtrip latency.