Models
Edge0 releases MoE expert-offloading framework, demos 35B model on iPhone
Edge0 streams mixture-of-experts weights from storage. Its 35B tier reports 2.9 GiB peak memory on a Mac mini M4 Pro; a launch post shows a 35B model on an iPhone.
OpenBMB releases MiniCPM5-2B for local deployment
A 2.5-billion-parameter dense model with a 131,072-token context, released under Apache-2.0 with GGUF, MLX and LiteRT-LM builds.
Google lists first Gemini Nano 4 phones, requires Nano 3 for Gemini Intelligence
ML Kit GenAI documentation now names nano-v4 devices and sets Nano v3 or greater, plus 12 GB of RAM, as the requirement for Gemini Intelligence.
Artificial Analysis benchmarks 33 local models on an iPhone 17 Pro
A benchmark of quantised small models on an iPhone 17 Pro reports intelligence scores, generation times and peak memory between 0.4 and 6.9 GB.
Liquid AI releases LFM2.5-2.6B for on-device agents
Liquid AI reports 30 tokens per second on a phone and under 2.5 GB of memory for its 2.6-billion-parameter model, released with open weights.
Apertus Mini distils an open-data 8B into 0.5B, 1.5B and 4B models
Swiss AI distilled its fully open Apertus 8B into 0.5B, 1.5B and 4B models on 1.7T tokens, with 3-bit to 6-bit MLX builds for Apple devices.
Meta AI builds MobileMoE, 5.3B parameters with 0.9B active per token
Meta AI trained three on-device mixture-of-experts models that store 1.3B to 5.3B parameters and run 272M to 922M of them per token.
LiteRT-LM reports 52 decode tokens per second for Gemma 4 E2B on an Android GPU
Google publishes prefill and decode figures for its on-device runtime, adds Swift and JavaScript APIs, and reports a 2.2x speedup from multi-token prediction.
Tencent open-sources a 440 MB offline translation model for phones
Hy-MT1.5-1.8B-1.25bit compresses a 1.8-billion-parameter translation model from 3.3 GB to 440 MB and runs offline on a phone. Weights and an Android demo are public.
Tempo compresses hour-long video with a 2B vision model so a 4B LLM can answer
Meta AI and KAUST report 52.7 on LVBench for a 6B system in which a 2B vision-language model squeezes long video down to about 3 tokens per frame.
Gemma 4 lands on the edge with Agent Skills in Google AI Edge Gallery
Google says Gemma 4 E2B runs in under 1.5 GB on some devices and reaches 3,700 prefill tokens per second on a Qualcomm Dragonwing IQ8 NPU.
Meta's MobileLLM-Flash goes shallow and wide for 1.8 times faster prefill
Meta designed 350M, 650M and 1.4B models by measuring latency on a Galaxy S25, reversing the deep-and-thin rule of the first MobileLLM.