LiteRT

4 updates on LiteRT.

  1. OpenBMB releases MiniCPM5-2B for local deployment

    A 2.5-billion-parameter dense model with a 131,072-token context, released under Apache-2.0 with GGUF, MLX and LiteRT-LM builds.

  2. LiteRT-LM reports 52 decode tokens per second for Gemma 4 E2B on an Android GPU

    Google publishes prefill and decode figures for its on-device runtime, adds Swift and JavaScript APIs, and reports a 2.2x speedup from multi-token prediction.

  3. Gemma 4 lands on the edge with Agent Skills in Google AI Edge Gallery

    Google says Gemma 4 E2B runs in under 1.5 GB on some devices and reaches 3,700 prefill tokens per second on a Qualcomm Dragonwing IQ8 NPU.

  4. Google AI Edge Gallery runs Gemma 3 1B and Qwen2.5 offline on Android

    An experimental Google app downloads LiteRT models from Hugging Face, runs chat, image questions and prompt tests offline, and prints decode speed per reply.