<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Arm · LLMobile.news</title><link>https://llmobile.news/companies/arm/</link><description>A concise news ticker covering AI on mobile devices, local models, apps and hardware.</description><language>en-GB</language><atom:link href="https://llmobile.news/companies/arm/index.xml" rel="self" type="application/rss+xml"/><item><title>Arm recaps Arm Create China and shows Qwen3-TTS 0.6B running on a vivo X300 CPU</title><link>https://llmobile.news/ticker/arm-create-china-2026/</link><guid isPermaLink="true">https://llmobile.news/ticker/arm-create-china-2026/</guid><pubDate>Fri, 11 Sep 2026 00:00:00 +0200</pubDate><description>Arm has published five developer takeaways from Arm Create, its developer events in Shanghai and Shenzhen. Two of them concern on-device AI. Arm says model choice starts with the workload and not with model size alone, and that developers should decide which parts of an application stay on the device, which run on nearby edge infrastructure and which need the cloud. The Shenzhen panel included Alibaba Qwen, ModelBest, Tencent Hunyuan and Ultralytics.
In the Shanghai keynote, Shantu Roy, Arm&amp;amp;rsquo;s VP of Developer Relations, discussed the Arm AI Portal. Arm says the portal lists models validated and optimized for Arm-based platforms, together with performance data for specific targets, code and deployment workflows. Coding agents can reach the same information through the Arm MCP Server.
The recap shows the portal&amp;amp;rsquo;s evaluation of Qwen3-TTS 0.6B Custom Voice, a multilingual streaming text-to-speech model from Alibaba, on a mobile CPU. The entry lists a vivo X300 with 8 CPU cores and 16 GB of memory, SME2, the XNNPACK and KleidiAI optimizations, FP16 weights and the LiteRT runtime. It reports a real-time factor of 1.2x against a baseline of 0.28x and a median end-to-end latency of 3,878 ms against 16,877 ms. Peak memory is 4,727 MB against 6,718 MB, and the evaluation uses the English subset of the MiniMaxAI TTS-Multilingual-Test-Set.
Evaluation results in the Arm AI Portal, as shown in Arm&amp;amp;#39;s recap. Source: Arm. Arm also points to Arm CSS for Mobile 2, which combines the Arm C2 CPU Cluster with SME2 and the Mali G2-Ultra NX GPU. Arm says the platform supports new on-device AI experiences on mobile. The next Arm Create event moves to the US, and Arm has not given a date.
Source: https://newsroom.arm.com/blog/takeaways-from-arm-create-china-2026
Read the article: https://llmobile.news/ticker/arm-create-china-2026/</description><category>Chips</category><category>Models</category><category>Android</category><category>TTS</category></item><item><title>Arm unveils Mali G2-Ultra NX GPU with neural accelerators in every shader core</title><link>https://llmobile.news/ticker/arm-mali-g2-ultra-nx/</link><guid isPermaLink="true">https://llmobile.news/ticker/arm-mali-g2-ultra-nx/</guid><pubDate>Tue, 08 Sep 2026 00:00:00 +0200</pubDate><description>Arm has introduced the Mali G2-Ultra NX, a smartphone GPU that places dedicated neural accelerators inside each shader core. The accelerators reuse the GPU&amp;amp;rsquo;s memory system, coherent caches and control structures, support INT8 and INT16 processing and include hardware-accelerated optical flow for motion estimation. Arm calls it the first AI-native Mali GPU and positions the accelerators for neural graphics at 1 W. The GPU is part of the Arm CSS for Mobile 2 platform.
The GPU reaches up to 24% higher benchmark performance than the previous generation and 14% higher performance in non-AI gaming workloads, Arm says. Its neural graphics features are Neural Super Sampling, which reconstructs a higher-resolution image from a lower-resolution render, Neural Frame Rate Upscaling, which generates intermediate frames, and Neural Super Sampling and Denoising, which combines upscaling with denoising for ray-traced scenes. Arm says frame rate upscaling supports up to 120 FPS for longer gaming sessions. In its Neural Dawn demo, Arm reports up to 4x higher performance efficiency and up to 70% lower external memory traffic.
Arm&amp;amp;#39;s overview slide for the Mali G2-Ultra NX. Source: Arm. Arm&amp;amp;#39;s diagram of how Neural Frame Rate Upscaling builds an intermediate frame. Source: Arm. The new execution engine is the largest update to the Mali instruction set architecture in seven generations, Arm says, with up to 2x more registers per warp. The third-generation hardware ray tracing unit adds support for Opacity Micromaps, which handle complex transparent geometry. Arm reports up to 13% lower DRAM traffic on ray tracing benchmarks, a 30% higher frame rate and up to 70% less ray tracing work in a scene from Moku&amp;amp;rsquo;s Central Garden.
For developers, Arm offers the Arm Neural Graphics Development Kit with machine learning extensions for Vulkan, plug-ins for Unreal Engine, an SDK for custom engines and tools for profiling, training and model optimization. Arm names integrations with Tencent Games Central Tech&amp;amp;rsquo;s Magic Dawn engine, Unity China&amp;amp;rsquo;s Tuanjie Engine and Unreal Engine MegaLights. Keli Zhou, engine lead for Where Winds Meet, says the game will be among the first to bring Arm Neural Technology to players. Arm&amp;amp;rsquo;s post names no launch dates for devices with the GPU.
Source: https://newsroom.arm.com/blog/arm-mali-g2-ultra-nx-ai-native-mobile-graphics
Read the article: https://llmobile.news/ticker/arm-mali-g2-ultra-nx/</description><category>Chips</category><category>Android</category></item><item><title>Arm unveils CSS for Mobile 2 with C2 CPU cluster, up to 1.7x faster on AI models</title><link>https://llmobile.news/ticker/arm-css-for-mobile-2/</link><guid isPermaLink="true">https://llmobile.news/ticker/arm-css-for-mobile-2/</guid><pubDate>Tue, 08 Sep 2026 00:00:00 +0200</pubDate><description>Arm has introduced Arm CSS for Mobile 2, a compute platform for smartphone chips that combines the C2 CPU cluster, the Mali G2-Ultra NX GPU and the SI L2 system interconnect. The C2 cluster pairs C2-Ultra and C2-Pro CPUs with two SME2 units, the Scalable Matrix Extension 2 that speeds up matrix math for AI on the CPU. Arm says this doubles the SME2 capability of the previous-generation configuration and reports up to 1.7x performance across the latest AI models.
The cluster delivers up to 15% higher single-thread performance, 15% faster web browsing, 12% faster app launch and 12% higher multi-thread performance, Arm reports. For AI, it cites a peak uplift of up to 70% in selected tasks. Its slide compares speech, personal memory retrieval and prefill, the phase where a model reads the prompt, against the C1-Ultra with SME2. In a representative agentic workflow covering speech processing, memory retrieval, reasoning, app execution and web browsing, the C2-Ultra with two SME2 units finishes 24% faster than the previous generation, according to Arm.
Arm&amp;amp;#39;s own comparison of the C2-Ultra with the C1-Ultra, with the AI tasks measured against the C1-Ultra with SME2. Source: Arm. The example flagship configuration in Arm&amp;amp;rsquo;s slides has two C2-Ultra and six C2-Pro cores. The Mali G2-Ultra NX GPU integrates neural accelerators into its shader cores and adds a new execution engine and a third-generation ray tracing unit for neural graphics. Arm says the SI L2 interconnect provides lower-latency access, higher bandwidth, coherency and quality-of-service controls for CPU, GPU and other resources working at the same time. Partners can use each component on its own or combine them with custom and third-party IP.
Arm&amp;amp;#39;s slide for the C2-Ultra CPU with an example flagship cluster layout. Source: Arm. On the software side, Arm lists KleidiAI, its optimized libraries for Arm CPUs including SME2 paths, plus integrations with common AI frameworks. The Arm AI Portal offers validated models with performance and accuracy data, code examples and deployment resources, and the Arm MCP Server connects them to agentic development tools. vivo says it is bringing Arm Neural Technology to its latest flagship smartphones built on the platform, aimed at mobile gaming. Arm&amp;amp;rsquo;s post names no launch dates for devices with CSS for Mobile 2.
Source: https://newsroom.arm.com/blog/arm-css-for-mobile-2-and-c2-cpu-cluster
Read the article: https://llmobile.news/ticker/arm-css-for-mobile-2/</description><category>Chips</category><category>Android</category></item><item><title>Graphcore runs Llama 3.2 11B Vision on a Pixel 8a with 2.7-bit quantisation</title><link>https://llmobile.news/ticker/graphcore-llama-mobile/</link><guid isPermaLink="true">https://llmobile.news/ticker/graphcore-llama-mobile/</guid><pubDate>Mon, 24 Aug 2026 00:00:00 +0200</pubDate><description>Graphcore Research has fitted the 11-billion-parameter vision-language model Llama-3.2-11B-Vision-Instruct into 3.6 GB, down from 21 GB, using a new 2.7-bit weight format called S3D8. In the team&amp;amp;rsquo;s blog post, the model runs on a Pixel 8a with a custom C++ inference implementation and generates 3.8 tok/s.
Your browser does not support this video. This video could not be loaded. Use the link below to open it directly.
Open video S3D8 stores three weights in 8 bits, which comes to 2.68 bits per parameter. Each group of three weights holds a 5-bit index into a shared table of centroids plus 3 bits of sign information, and the values are decoded to channel-scaled INT8 for compute. On the Pixel 8a using 5 cores, the fused dequantize-and-multiply kernel reaches 33.8 GMAC/s for single-token generation (GMAC/s means giga multiply-accumulate operations per second, a standard measure of AI compute throughput on devices), compared with 26.5 GMAC/s for INT8 and 13.6 GMAC/s for bfloat16, according to Graphcore. Graphcore measured the 3.8 tok/s while generating 100 tokens from a single image tile.
Graphcore measured accuracy on 1,024-example subsets of VQAv2, ChartQA, DocVQA and AI2D. S3D8 averages 0.661 against 0.744 for the bfloat16 original, while a plain integer format of similar size averages 0.347 and a student-t format 0.565. The team recovered quality with quantization-aware training by distillation, where a frozen bfloat16 teacher guides the quantized student over 2,048 steps. Training data consists of ImageNet images paired with text the teacher generated from a pool of 495 prompt seeds, and the team says the wrong data lowers the training loss while the model gets worse on downstream tasks.
The demo needs a Pixel 8a or newer Android phone with 8 GB of RAM and 4 to 5.5 GB of free storage, and a CPU that supports the Arm i8mm and bf16 extensions. Collaborators from Arm contributed to the work. The demo APK can be downloaded from Graphcore, the Android code is on GitHub and the method is described in a paper on arXiv. Graphcore says the model is too large and compute-heavy for the phone, so although it runs, it is not practical.
Update, September 19, 2026. The paper puts the compressed model at 3.73 GB including vocabulary and metadata, of which 3,569 MB is weight storage at 2.68 bits per parameter. The authors also ran S3D8 on a Graviton4 server chip with 96 Neoverse V2 cores, where end-to-end decoding reaches 36.8 tok/s against 26.4 tok/s for INT8. On the Pixel 8a, decoding reads weights at 12.5 GB/s.
Per task, S3D8 scores 0.702 on VQAv2 (bfloat16 0.754), 0.648 on ChartQA (0.747), 0.740 on DocVQA (0.844) and 0.554 on AI2D (0.631). The paper lists its limits as a single evaluated model, CPU-only execution and one fixed format for all weights, since mixed-precision experiments did not work. Training used 2,048 steps at batch size 128 on 1.28 million ImageNet images with teacher-generated responses.
Source: https://graphcore-research.github.io/2026-08-24-llama-mobile/
Read the article: https://llmobile.news/ticker/graphcore-llama-mobile/</description><category>Quantisation</category><category>Research</category><category>Android</category><category>Open source</category></item></channel></rss>