Gemma
10 updates on Gemma.
llama.cpp Hexagon NPU backend tested on a Snapdragon 8 Gen 3 phone
A user report puts Gemma 3 4B at 12.5 tokens per second of generation on a OnePlus 12 using llama.cpp’s Hexagon NPU backend, at about CPU speed but without the heat.
LiteRT-LM reports 52 decode tokens per second for Gemma 4 E2B on an Android GPU
Google publishes prefill and decode figures for its on-device runtime, adds Swift and JavaScript APIs, and reports a 2.2x speedup from multi-token prediction.
Airgap is a React Native kit for support chatbots that answer offline
Xavier Puspus published a React Native kit for support chatbots that answer without a network, running a 2.4 GB Gemma 4 E2B file through llama.rn.
Gemma 4 lands on the edge with Agent Skills in Google AI Edge Gallery
Google says Gemma 4 E2B runs in under 1.5 GB on some devices and reaches 3,700 prefill tokens per second on a Qualcomm Dragonwing IQ8 NPU.
Gemma 3n runs 5B and 8B models in 2 GB and 3 GB of memory
Google previewed a mobile-first model whose Per-Layer Embeddings cut RAM use, and said the same architecture powers the next Gemini Nano.
Gemma 3 adds a 1B size that fits in 0.5 GB as an int4 checkpoint
Google released Gemma 3 at 1B, 4B, 12B and 27B with quantisation-aware int4 checkpoints of 0.5 GB and 2.6 GB for the two smallest sizes.
Gemma 2 2B is distilled from a larger model and scores 1126 on Chatbot Arena
Google released a 2.6B-parameter Gemma 2 trained by distilling a larger model, and reported an Elo of 1126 on the LMSYS Chatbot Arena.
Octopus v2 is a 2B model that calls Android APIs with one token per function
A 2B Gemma fine-tune gives every Android API its own token, and the authors report 99.524% accuracy and 0.38 seconds per call, ahead of GPT-4.
Octopus fine-tunes a 2B model to 93 percent on API function calls
Stanford and Harvard authors fine-tuned open 2B to 7B models on 20,000 RapidAPI functions and report up to 97 percent function call accuracy.
Gemma 2B and 7B open the Gemma line, built on Gemini research
Google released Gemma 2B and 7B with an 8192-token context, weights on Kaggle and Hugging Face under a custom Gemma licence, not an open source one.