On-device

  1. MediaTek launches Dimensity 9600 Pro, a 2nm chip for on-device models up to 30B parameters

    MediaTek says the NPU 1090 in the 2nm Dimensity 9600 Pro raises LLM prefill by 51% and supports on-device models with up to 30B parameters.

  2. Arm recaps Arm Create China and shows Qwen3-TTS 0.6B running on a vivo X300 CPU

    Arm's Create recap shows Qwen3-TTS 0.6B on a vivo X300 CPU with SME2 at a 1.2x real-time factor and 4.7 GB peak memory.

  3. Edge0 releases MoE expert-offloading framework, demos 35B model on iPhone

    Edge0 streams mixture-of-experts weights from storage. Its 35B tier reports 2.9 GiB peak memory on a Mac mini M4 Pro; a launch post shows a 35B model on an iPhone.

  4. llama.cpp Hexagon NPU backend tested on a Snapdragon 8 Gen 3 phone

    A user report puts Gemma 3 4B at 12.5 tokens per second of generation on a OnePlus 12 using llama.cpp’s Hexagon NPU backend, at about CPU speed but without the heat.

  5. iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine

    Apple says the A20 Pro carries 32 Neural Engine cores in total, double the AI processing power of A19 Pro, with 50 percent more memory bandwidth.

  6. Arm unveils Mali G2-Ultra NX GPU with neural accelerators in every shader core

    Arm says the Mali G2-Ultra NX runs neural upscaling and frame generation on accelerators inside its shader cores, with up to 24% higher benchmark performance.

  7. Arm unveils CSS for Mobile 2 with C2 CPU cluster, up to 1.7x faster on AI models

    Arm says its C2-Ultra CPU with two SME2 units delivers up to 1.7x performance across the latest AI models and finishes an agentic workflow 24% faster.

  8. OpenBMB releases MiniCPM5-2B for local deployment

    A 2.5-billion-parameter dense model with a 131,072-token context, released under Apache-2.0 with GGUF, MLX and LiteRT-LM builds.

  9. Google lists first Gemini Nano 4 phones, requires Nano 3 for Gemini Intelligence

    ML Kit GenAI documentation now names nano-v4 devices and sets Nano v3 or greater, plus 12 GB of RAM, as the requirement for Gemini Intelligence.

  10. Artificial Analysis benchmarks 33 local models on an iPhone 17 Pro

    A benchmark of quantised small models on an iPhone 17 Pro reports intelligence scores, generation times and peak memory between 0.4 and 6.9 GB.

  11. Ornith releases Ornith-1.5, a 9B model with a mobile build for iPhone and Android

    Ornith-1.5 comes in 9B, 35B and 397B sizes, and the 9B model scores 70.6 on SWE-bench Verified and has a mobile build for iPhone and Android.

  12. Online-SDFT reports 70.28% routing accuracy and updates a LoRA adapter on the phone

    I-Ju Lin and Zhang-Wei Hong measure 70.28% accuracy against 52.78% for the best baseline, and run the LoRA update itself on an Android phone.