← Back to projects

work · ongoing

Large-Screen Android Systems

Building high-performance drawing systems and custom Android layers for large-screen interactive touch displays.

Android · Large screens · HardwareKotlin · Jetpack Compose · Android SurfaceView · Coroutines & Flow · Perfetto & Systrace

Overview

Interactive Flat Panels (IFPs) are large-format (65” to 86”) 4K Android-based touch displays engineered for interactive instruction and collaborative workspaces. Unlike standard smartphones or tablets, these panels serve as the central visual appliance in a room: they run continuously throughout the day, host digital whiteboard sessions, split-screen media, and handle simultaneous multi-user touch and stylus inputs.

Working as an Android engineer on dedicated large-screen hardware, my focus was architecting high-responsiveness canvas interactions, orchestrating split-view media controls, and hardening application stability against strict hardware and memory budgets.

The problem

Standard Android application engineering carries implicit assumptions:

  1. Screen scale: Displays hover between 5” and 11”, where pixel fill rate is rarely the primary rendering bottleneck.
  2. Single user input: One finger or stylus at a time triggers touch events.
  3. Lifecycle conventions: The OS aggressively suspends and reclaims background activities to conserve battery.

On an IFP, every one of these assumptions breaks:

  • Driving a 4K resolution (3840×2160) canvas at 60 fps means rendering over 8.2 million pixels per frame within a strict 16.6ms window.
  • Multiple students or teachers write simultaneously, generating hundreds of high-frequency digitizer events per second that threaten to saturate the Android main thread.
  • Kiosk-mode classroom sessions demand persistent stability; dropped frames during live handwriting feel like physical chalk skipping across a board.

Constraints

  • System-on-Chip (SoC) Headroom: Embedded ARM architectures running custom Android board support packages with shared GPU/VPU memory buses.
  • Fill-rate limits: Full-screen redrawing at 4K immediately consumes the GPU fill rate budget, leaving zero headroom for UI animations.
  • Input sampling rates: Hardware digitizers sample touches at 120Hz–200Hz, producing an intense stream of touch coordinates that must be smoothed and rendered with sub-30ms touch-to-pixel latency.
  • Zero toleration for memory leaks: The application must remain responsive across back-to-back lectures without progressive GC degradation.

My approach

To achieve tactile, paper-like handwriting responsiveness on 4K glass, we separated the high-frequency drawing loop from the standard declarative UI tree:

Touch Digitizer (120Hz+)
        │
        ▼
[Native Touch Dispatcher]
        │
        ├── Raw Coordinates ──► Low-latency Canvas Pipeline (SurfaceView / Hardware Layer)
        │                                  │
        │                        Dirty Rect Invalidation & Object Pool
        │                                  │
        ▼                                  ▼
[Path Smoothing Worker] ────────► 4K Direct Pixel Buffer
(Ramer-Douglas-Peucker)
  1. Bypassing Declarative Recomposition: While Jetpack Compose powers high-level chrome, toolbars, and menus, the drawing canvas was isolated into dedicated hardware-accelerated drawing pipelines where dirty-region invalidation was calculated explicitly.
  2. Strict Threading Boundaries: Ingestion of raw touch points and Bezier curve fitting occurred on lightweight coroutine channels, isolating geometric computation from the Main Dispatcher until the exact moment of buffer blitting.
  3. Dirty-Rectangle Clipping: Rather than redrawing the 4K canvas per touch point, the rendering pipeline calculates the bounding box of the active stroke segment plus a padding margin, restricting GPU rasterization to minimal rectangular sub-regions.

Important decisions

1. Object Pooling over Allocation in Touch Loops

Standard Android touch processing allocates point objects and temporary path structures. Under 10-finger multi-touch, this creates severe object churn that triggers ART generational GC pauses. We implemented custom primitive float buffers and object pools for active strokes, achieving a near-zero allocation inner loop.

2. Multi-stage Stroke Simplification

Instead of simplifying complex paths during the active draw gesture, strokes are drawn with quadratic Bezier curves in real time. Once the stylus lifts (ACTION_UP), a background worker runs Ramer-Douglas-Peucker path simplification to reduce vector node count before serializing to disk.

3. Decoupling Tool State from Surface Rendering

Pen color, stroke width, eraser modes, and shape recognizers are modeled as reactive immutable state streams via Kotlin StateFlow. UI controls emit state updates cleanly without re-binding or invalidating the underlying canvas surface.

What went wrong: The Stutter Trap

Early testing revealed that when teachers drew rapid spiral strokes across the screen, visible micro-stutters occurred every few seconds.

The initial assumption was that the embedded GPU was throttling due to thermal limits. However, CPU frequency monitoring showed stable clocks. The problem was subtler: the stroke serialization layer was emitting events to an in-memory undo/redo stack that boxed primitive coordinates into Java ArrayList<Float>.

At 120Hz across multiple fingers, this generated enough heap allocations to provoke concurrent GC sweeps every 4–5 seconds. While ART sweeps are fast on modern phones, on an embedded chipset with high memory bus saturation, GC pauses paused the main thread for 22ms—causing a dropped frame.

Debugging & Investigation

We captured detailed traces using Perfetto and Systrace:

  1. Symptom: Periodic frame drops coinciding with art::gc::Heap::CollectGarbageInternal.
  2. Hypothesis: Heap churn caused by touch event encapsulation.
  3. Verification: Memory profiler confirmed ~40MB of short-lived PointF and boxed collection instances generated within 30 seconds of continuous drawing.
  4. Fix: Replaced collection structures with contiguous primitive float arrays (FloatArray) and reused pre-allocated stroke buffers across gestures.
// Conceptual representation of the zero-allocation touch buffer
internal class TouchPointRingBuffer(private val capacity: Int = 1024) {
    private val xs = FloatArray(capacity)
    private val ys = FloatArray(capacity)
    private var head = 0

    fun push(x: Float, y: Float) {
        xs[head] = x
        ys[head] = y
        head = (head + 1) % capacity
    }
}

Result

  • Sub-30ms touch-to-pixel latency: Natural, chalk-like writing experience across large-scale 4K displays.
  • Stable 60 fps rendering: Maintained consistent frame rates even under continuous multi-touch writing scenarios.
  • Zero-leak longevity: Apps maintain steady memory consumption over 8+ hour continuous classroom sessions.

What I learned

  1. Scale changes physics: Code that performs adequately on a 1080p 6” phone can collapse under the fill-rate and memory-bus demands of a 4K 75” display.
  2. Garbage collection is an architectural constraint: In high-frequency interactive systems, avoiding allocations in the inner loop matters more than micro-optimizing algorithmic constants.
  3. Systrace tells no lies: Intuition about performance bottlenecks is almost always wrong until verified against trace timestamps.