<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>MoE · LLMobile.news</title><link>https://llmobile.news/tags/moe/</link><description>A concise news ticker covering AI on mobile devices, local models, apps and hardware.</description><language>en-GB</language><atom:link href="https://llmobile.news/tags/moe/index.xml" rel="self" type="application/rss+xml"/><item><title>Qualcomm teases next-gen Hexagon NPU ahead of Snapdragon Summit 2026</title><link>https://llmobile.news/ticker/qualcomm-hexagon-npu-agentic-ai/</link><guid isPermaLink="true">https://llmobile.news/ticker/qualcomm-hexagon-npu-agentic-ai/</guid><pubDate>Thu, 10 Sep 2026 00:00:00 +0200</pubDate><description>Qualcomm has teased its next-generation Hexagon NPU ahead of Snapdragon Summit 2026 (Sep 22–24), detailing an architecture redesigned around persistent on-device AI agents rather than single-pass inference. The full platform — widely expected to be a Snapdragon 8 Elite Gen 6 — will be formally announced at Summit. Qualcomm has not yet named the chip or provided absolute figures for memory capacity or NPU throughput.
Two architectural changes stand out. The Element Accelerator is a new block inside the Hexagon NPU dedicated to transformer operations, working alongside the existing scalar, vector and matrix units. It targets the compute patterns that modern generative and agentic models use. A 50% larger shared memory subsystem inside the NPU keeps more model state, activations and intermediate tensors on-chip, reducing trips to external LPDDR — a move aimed at the memory wall that has long bottlenecked mobile AI, especially with long context windows.
The NPU also prepares for Mixture-of-Experts (MoE) models. The company cites a 30-billion-parameter MoE that activates only about 3B parameters per token, keeping tens of billions available without the bandwidth cost of dense models. This is paired with flash-to-memory expert management and caching to load the right experts from storage. INT2, INT4, INT8, FP8 and FP16 precision are supported, and the new platform delivers up to 50% higher prefill performance for INT4 models, faster decoding throughput and enhanced speculative decoding.
The design targets always-running AI with long-context reasoning and low-latency action loops, with the CPU handling routing and orchestration of multi-step workloads while Hexagon accelerates the heavy lifting. Adreno handles AI-accelerated graphics rendering and a new Neural Fusion layer extends AI acceleration into gaming. Whether this actually delivers faster, more capable local assistants depends on finished devices — several critical figures (absolute memory size, power consumption, real-world benchmarks) remain undisclosed. The full picture comes at Summit in two weeks.
Source: https://www.qualcomm.com/news/onq/2026/09/hexagon-npu-agentic-ai-architecture
Read the article: https://llmobile.news/ticker/qualcomm-hexagon-npu-agentic-ai/</description><category>NPU</category><category>Qualcomm</category><category>Agentic AI</category><category>MoE</category></item></channel></rss>