NPU

23 updates on NPU.

  1. llama.cpp Hexagon NPU backend tested on a Snapdragon 8 Gen 3 phone

    A user report puts Gemma 3 4B at 12.5 tokens per second of generation on a OnePlus 12 using llama.cpp’s Hexagon NPU backend, at about CPU speed but without the heat.

  2. iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine

    Apple says the A20 Pro carries 32 Neural Engine cores in total, double the AI processing power of A19 Pro, with 50 percent more memory bandwidth.

  3. Pixel 11 series: Tensor G6 adds 50 percent more TPU compute

    Google says Tensor G6 with the latest Gemini Nano model processes on-device AI tasks up to 3.5 times faster while using up to 3.5 times less energy.

  4. MLPerf Mobile v6.0 adds Llama tests in 1B, 3B and 8B sizes

    MLCommons added Llama tests in 1B, 3B and 8B sizes to its mobile benchmark app, reporting token throughput next to the existing vision and image tests.

  5. llada.cpp runs a diffusion LLM on a Snapdragon NPU up to 42 times faster

    Tsinghua and Beihang researchers report LLaDA-8B generating 128 tokens 17 to 42 times faster on a Hexagon NPU than on the phone CPU.

  6. Quant.npu runs 4-bit LLMs up to 15.1 percent faster on a Qualcomm NPU

    Eight authors report a fully static integer quantisation method that runs 4-bit models up to 15.1 percent faster on a Qualcomm SM8650 NPU.

  7. iPhone 16 Pro settles 41.5 percent below peak over 20 back-to-back prompts

    Four authors ran Qwen 2.5 1.5B 20 times in a row on four edge platforms and report the iPhone 16 Pro settling 41.5 percent below its peak throughput.

  8. FlexServe runs on-device LLM inference inside TrustZone for 4.41 percent more latency

    Shanghai Jiao Tong University researchers kept model weights and prompts inside ARM TrustZone at 4.41 percent more time to first token.

  9. ExecuTorch 1.0 reaches general availability for on-device PyTorch models

    The PyTorch edge runtime promotes Core ML, Qualcomm Hexagon, Arm Ethos-U, Vulkan and XNNPACK backends to production status.

  10. A19 Pro puts Neural Accelerators in every GPU core

    Apple says the iPhone 17 Pro chip pairs Neural Accelerators in each of six GPU cores with a 16-core Neural Engine to run large local language models.

  11. ShadowNPU scores attention on the Snapdragon NPU and reports 4.5 times faster inference

    Peking University and BUPT move attention token scoring onto a Snapdragon NPU and report up to 4.5 times faster inference using a single CPU core.

  12. P/D-Device prefills in the cloud and decodes on the phone, cutting TTFT 60%

    Huawei researchers split prefill and decoding between cloud and phone, reporting time to first token down at least 60% and cloud throughput up to 15x.