Android
51 updates on Android.
llm.npu prefills 1,106 tok/s for Qwen1.5-1.8B on a Snapdragon 8 Gen 3
Peking University and BUPT move prompt processing onto the Hexagon NPU and report up to 43.6x faster prefill and up to 59.5x lower energy use.
BUPT measures 22 LLMs on four Android phones at about 200 ms per token
A BUPT team ran 22 models from 0.5B to 7B on four Android phones with llama.cpp and found 7B models take about 4 GB and 200 ms per token.
PowerInfer-2 runs a 47B model on a OnePlus 12 at 11.68 tokens per second
Shanghai Jiao Tong University researchers report a 47B model decoding at 11.68 tokens per second on a OnePlus 12, with weights streamed from flash.
Octopus v2 is a 2B model that calls Android APIs with one token per function
A 2B Gemma fine-tune gives every Android API its own token, and the authors report 99.524% accuracy and 0.38 seconds per call, ahead of GPT-4.
One shared on-device LLM keeps a context per app and switches in 0.27 seconds
Peking University, BUPT and Tsinghua researchers compress and swap each app conversation context in 16-token chunks, switching up to 20 times faster.
Google ships an experimental MediaPipe LLM Inference API for web, Android and iOS
The experimental API runs Gemma 2B, Phi 2, Falcon 1B and Stable LM 3B fully on device, with int8 and int4 weights.
MobiLlama is a fully transparent 0.5B model that runs in 770 MB on a phone
MBZUAI published a 0.5B model that shares one feed-forward block across all layers and reports 7.02 tok/s in 770 MB on a Snapdragon 685 phone.
Qualcomm AI Hub opens 75 pre-optimised models and profiling on hosted Snapdragon phones
Qualcomm opened a library of more than 75 models tuned for Snapdragon, with compilation and profiling on real phones in its cloud and two 7B chat models listed.
Galaxy S24 becomes the second phone line to run Gemini Nano
Google brought Gemini Nano to the Galaxy S24 for on-device Magic Compose in Messages, while the rest of the Galaxy AI features run on Gemini Pro.
Arm runs Llama 2 7B at 9.6 tokens per second on three mobile CPU cores
Arm demonstrated Llama 2 7B generating 9.6 tokens per second on three Cortex-A700 CPU cores, using int4 kernels it says beat llama.cpp by 20 percent.
Gemini Nano ships on the Pixel 8 Pro and Android gets AICore
Google put Gemini Nano on the Pixel 8 Pro for Recorder summaries and Gboard Smart Reply, and introduced AICore as the Android service behind it.
Snapdragon 8 Gen 3 targets 10-billion-parameter models on device
Qualcomm says the new flagship runs generative models with up to 10 billion parameters on device and reaches up to 20 tokens per second for LLMs.