Google

16 updates on Google.

  1. Google lists first Gemini Nano 4 phones, requires Nano 3 for Gemini Intelligence

    ML Kit GenAI documentation now names nano-v4 devices and sets Nano v3 or greater, plus 12 GB of RAM, as the requirement for Gemini Intelligence.

  2. Pixel Watch 5 adds offline Gemini commands and faster on-device smart replies

    Offline voice commands on the Pixel Watch 5 run on an on-device model, and a Gemini Nano upgrade makes smart replies 50 percent faster.

  3. Pixel 11 series: Tensor G6 adds 50 percent more TPU compute

    Google says Tensor G6 with the latest Gemini Nano model processes on-device AI tasks up to 3.5 times faster while using up to 3.5 times less energy.

  4. Gemini Nano 4 ships on Samsung foldables with ML Kit Prompt API access

    Google says Samsung’s new foldables carry Gemini Nano 4 with support for more than 140 languages, reachable from apps through ML Kit’s Prompt API.

  5. LiteRT-LM reports 52 decode tokens per second for Gemma 4 E2B on an Android GPU

    Google publishes prefill and decode figures for its on-device runtime, adds Swift and JavaScript APIs, and reports a 2.2x speedup from multi-token prediction.

  6. Gemma 4 lands on the edge with Agent Skills in Google AI Edge Gallery

    Google says Gemma 4 E2B runs in under 1.5 GB on some devices and reaches 3,700 prefill tokens per second on a Qualcomm Dragonwing IQ8 NPU.

  7. Gemma 3n runs 5B and 8B models in 2 GB and 3 GB of memory

    Google previewed a mobile-first model whose Per-Layer Embeddings cut RAM use, and said the same architecture powers the next Gemini Nano.

  8. Google AI Edge Gallery runs Gemma 3 1B and Qwen2.5 offline on Android

    An experimental Google app downloads LiteRT models from Hugging Face, runs chat, image questions and prompt tests offline, and prints decode speed per reply.

  9. Google's ML Drift runs Llama 3.1 8B on a phone GPU at 12.7 tokens per second

    The paper measures 37.1 decode tokens per second for Gemma2 2B and 12.7 for Llama 3.1 8B on the Adreno 750 GPU of a Samsung S24.

  10. Gemma 3 adds a 1B size that fits in 0.5 GB as an int4 checkpoint

    Google released Gemma 3 at 1B, 4B, 12B and 27B with quantisation-aware int4 checkpoints of 0.5 GB and 2.6 GB for the two smallest sizes.

  11. Gemma 2 2B is distilled from a larger model and scores 1126 on Chatbot Arena

    Google released a 2.6B-parameter Gemma 2 trained by distilling a larger model, and reported an Elo of 1126 on the LMSYS Chatbot Arena.

  12. Google ships an experimental MediaPipe LLM Inference API for web, Android and iOS

    The experimental API runs Gemma 2B, Phi 2, Falcon 1B and Stable LM 3B fully on device, with int8 and int4 weights.