<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>OnePlus · LLMobile.news</title><link>https://llmobile.news/companies/oneplus/</link><description>A concise news ticker covering AI on mobile devices, local models, apps and hardware.</description><language>en-GB</language><atom:link href="https://llmobile.news/companies/oneplus/index.xml" rel="self" type="application/rss+xml"/><item><title>OPPO unveils ColorOS 17 with an on-device linear-attention model and 128K context</title><link>https://llmobile.news/ticker/oppo-coloros-17-on-device-model/</link><guid isPermaLink="true">https://llmobile.news/ticker/oppo-coloros-17-on-device-model/</guid><pubDate>Thu, 17 Sep 2026 00:00:00 +0200</pubDate><description>OPPO presented ColorOS 17 at its developer conference in Zhuhai, with an on-device large language model built on what the company calls its first Linear Attention architecture. Linear attention keeps compute cost proportional to the length of the input instead of growing quadratically. OPPO says the model supports 128K context natively, uses 48% less memory and 55% less energy than the previous generation, and keeps user data on the device while responding faster.
The on-device model is part of AndesGPT — OPPO&amp;amp;rsquo;s proprietary large model trained independently and available in sizes from around 1 billion to over 100 billion parameters, deployed in OPPO&amp;amp;rsquo;s &amp;amp;ldquo;1+N&amp;amp;rdquo; end-cloud architecture (a central large model plus specialized smaller models per task). AndesGPT was the first Chinese on-device LLM to register with the Cyberspace Administration on July 15, 2026, alongside Apple Intelligence, Huawei Xiaoyi, vivo BlueLM, Xiaomi HyperAI, Samsung Galaxy AI and ZTE Nubia Doubao.
The update also brings Persona X, a memory engine that builds on data, environmental and behavioral perception and records, forgets and reflects on what it learns about the user. The Xiao Bu assistant (Breeno outside China) gets longer context and better intent recognition. OPPO reports a 19% gain in internal satisfaction testing with device-cloud collaboration, and the assistant can now open 200 Alipay services, plan routes in Tencent Maps and place WeChat messages and calls by voice command.
ColorOS 17 ships first on the upcoming Find X10 series and OnePlus 16, and OPPO, OnePlus and Realme are upgrading in sync. According to Gizmochina, the rollout to existing devices in China starts on October 8 with the Find X9 series and Find N6. The Find X8 series, Find N5 and tablets follow on October 16, and older flagship and mid-range devices from November 26. No global timeline has been announced.
Source: https://www.jjckb.cn/20260918/447b6648587442e4b8e7b6594720fe3d/c.html
Read the article: https://llmobile.news/ticker/oppo-coloros-17-on-device-model/</description><category>Memory</category><category>Android</category></item><item><title>NPUs lead LLM prefill, CPUs lead decoding on Snapdragon phones</title><link>https://llmobile.news/ticker/mobile-npu-llm-inference-study/</link><guid isPermaLink="true">https://llmobile.news/ticker/mobile-npu-llm-inference-study/</guid><pubDate>Mon, 06 Jul 2026 00:00:00 +0200</pubDate><description>A team of researchers has published a cross-layer measurement study of mobile LLM inference titled &amp;amp;ldquo;Is Your NPU Ready for LLMs?&amp;amp;rdquo;. It covers five frameworks, llama.cpp, MNN, MLC-LLM, MLLM and Qualcomm&amp;amp;rsquo;s GENIE, on CPU, GPU and NPU backends. The study finds that framework-induced performance gaps grow on NPUs and reach up to 10x.
The tests ran on four Snapdragon phones, the Xiaomi 17, OnePlus 15, Xiaomi 15 and Xiaomi 14, with Llama 3.2 1B and 3B, Qwen 2.5 1.5B and 7B and Phi 3.5 3.8B, mostly with 4-bit weights. At a 256-token context, GENIE on the NPU reaches 1,463.7 tok/s in prefill against 115.1 tok/s for llama.cpp. In decoding the order flips, with the CPU at about 50 to 76 tok/s and the NPU at about 10 to 34 tok/s. The authors attribute this to NPUs preferring large, fixed-shape workloads, which conflicts with the small, dynamic kernels of decoding.
For energy, the team built PowerBench, a lightweight C++ library that maps PMIC power zones to CPU, GPU and NPU for backend-specific attribution. They report that suboptimal configurations waste up to 40% of energy and that a tuned setup cuts NPU energy by up to 54.8%. The recommended settings are a 20 μs RPC polling interval, an NPU sleep latency of 65535 μs, the CPU at its lowest frequency and full-graph offloading with QNN.
Source: https://arxiv.org/abs/2607.05475
Read the article: https://llmobile.news/ticker/mobile-npu-llm-inference-study/</description><category>Research</category><category>Benchmarks</category><category>NPU</category><category>Android</category><category>Runtimes</category></item></channel></rss>