<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Smartphone · LLMobile.news</title><link>https://llmobile.news/categories/smartphone/</link><description>A concise news ticker covering AI on mobile devices, local models, apps and hardware.</description><language>en-GB</language><atom:link href="https://llmobile.news/categories/smartphone/index.xml" rel="self" type="application/rss+xml"/><item><title>iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine</title><link>https://llmobile.news/ticker/iphone-18-pro-a20-pro-neural-engine/</link><guid isPermaLink="true">https://llmobile.news/ticker/iphone-18-pro-a20-pro-neural-engine/</guid><pubDate>Mon, 21 Sep 2026 00:00:00 +0200</pubDate><description>Apple has introduced the iPhone 18 Pro and iPhone 18 Pro Max with the A20 Pro chip. According to Apple, the chip has a new dual 16-core Neural Engine, 32 cores in total, which the company describes as double the AI processing power of A19 Pro. Apple says the Neural Engine accelerates on-device AI models and computational photography.
Image: Apple. The A20 Pro also offers 50 percent more memory bandwidth than A19 Pro, and its 6-core CPU includes integrated Neural Accelerators. Apple positions the chip for &amp;amp;ldquo;more advanced on-device AI workloads&amp;amp;rdquo; alongside games.
The phones ship with iOS 27, Apple Intelligence and the new Siri AI. Apple states that Apple Intelligence uses on-device processing together with Private Cloud Compute. Pre-orders start on 12 September, with availability from 18 September; prices start at $1,199 for the iPhone 18 Pro and $1,299 for the iPhone 18 Pro Max, both with 256 GB.
Update, September 20, 2026. The retailer ZEERA reports results from a Max Tech side-by-side comparison of the iPhone 18 Pro Max and iPhone 17 Pro Max. Local LLM generation rose from 33.5 to 55.4 tok/s, a gain of 65 percent, and measured memory bandwidth from 66.5 to 111.8 GB/s, a gain of 68 percent. The comparison also lists 25 percent higher single-core and 29 percent higher multi-core CPU scores in Geekbench 7, and a 48 percent higher graphics score.
Apple&amp;amp;rsquo;s developer page for Apple Intelligence says the Foundation Models framework gives Swift apps direct access to the same on-device model that powers Apple Intelligence. New multimodal prompts pass images alongside text, and the model can call Vision framework tools such as OCR and barcode readers, all on-device according to Apple. Apps can also work with any model that conforms to the Language Model protocol, including cloud models like Claude and Gemini. Members of the App Store Small Business Program with fewer than 2 million first-time downloads get the next generation of Apple Foundation Models on Private Cloud Compute at no cloud API cost.
The new Siri AI is a three-tier hybrid system, not fully on-device. A new System Orchestrator in iOS 27 routes every request to the right tier:
Tier 1 — On-Device: Simple tasks (timers, app launches, local context search) run entirely on Apple&amp;amp;rsquo;s Foundation Models (~3B parameters, 4-bit quantized) on the Neural Engine. Nothing leaves the iPhone. Tier 2 — Private Cloud Compute: Moderately complex requests route to Apple&amp;amp;rsquo;s own servers on Apple Silicon (M2 Ultra, now M5) in hardware-isolated Secure Enclaves. Personal data is stripped and tokenized. Apple retains control; no data retention. Tier 3 — Google Gemini: The heaviest reasoning, real-time knowledge, and multimodal tasks are sent through PCC&amp;amp;rsquo;s privacy proxy to a custom 1.2-trillion-parameter Gemini model on Nvidia Blackwell B200 GPUs in Google Cloud. Before any query reaches Google, the user&amp;amp;rsquo;s identity and personal data are removed. Google is contractually barred from training on these queries. Apple reportedly pays Google ~$1 billion per year for this partnership. The System Orchestrator decides the routing threshold in real time based on task complexity and context size. Apple&amp;amp;rsquo;s public materials only name On-Device and PCC as the two processing locations — Gemini is never mentioned on Apple&amp;amp;rsquo;s consumer pages.
Foundation Models can now work with any language model, including Apple&amp;amp;rsquo;s own, Claude, and Gemini, through a single Swift interface. Apple did not publish details of the routing logic or which types of queries are sent off-device.
Update, September 21, 2026. Adrien Grondin, who works on the Locally AI app, says on X that the iPhone 18 Pro runs a 27B model at double the speed of the iPhone 17 Pro. His 11-second video shows the demo, but the post names neither the model nor its quantization and gives no tok/s figures.
Wccftech reports on the same demo, attributing it to Grondin and saying the model is Bonsai 2 27B from Prism ML — a ternary-weight version of Qwen3.8 27B at 1.76 bits per weight, 5.9 GB total. PrismML describes Bonsai 2 27B as released under Apache 2.0, running on Apple devices through MLX and on NVIDIA GPUs through CUDA.
The iPhone 18 Pro is much better than what I expected for on-device AI
It can run a 27B model at double the speed of the 17 Pro
The new A20 Pro chip is a beast pic.twitter.com/gb11TvBfn0
&amp;amp;mdash; Adrien Grondin (@adrgrondin) September 19, 2026
Source: https://www.apple.com/newsroom/2026/09/apple-debuts-iphone-18-pro-and-iphone-18-pro-max/
Read the article: https://llmobile.news/ticker/iphone-18-pro-a20-pro-neural-engine/</description><category>iPhone</category><category>Chips</category><category>NPU</category></item><item><title>vivo previews BlueCode phone coding agent and researches a 30B MoE on-device model</title><link>https://llmobile.news/ticker/vivo-bluecode/</link><guid isPermaLink="true">https://llmobile.news/ticker/vivo-bluecode/</guid><pubDate>Sat, 19 Sep 2026 12:19:04 +0200</pubDate><description>vivo has released a preview version of BlueCode, a native coding agent that lets users complete software development tasks on the phone itself. The company presented it at its developer conference in Shenzhen and describes the goal as turning phones from content consumption terminals into productivity tools. The reports name no supported devices, release date or technical details for the preview.
vivo is also researching a 30B-parameter BlueLM MoE model for on-device use, according to Tencent News. MoE means the model is split into expert sub-networks and only part of it is active per token. Vice President Zhou Wei said the research aims to explore trillion-scale model applications on phones, and that memory and compute are the main barriers to running large models on them. If model capability improves substantially, each load would need only about 3 to 12 GB of memory, the report quotes him as saying.
The 30B model sits next to the BlueLM lineup that vivo announced at the same event, which has four models for speech (BlueLM-RealTime), on-device use (BlueLM-Nano) and the cloud (BlueLM-Flash and BlueLM-Pro). According to NetEase, BlueLM-Nano handles local perception and personal memory on the device. vivo also launched OriginOS 7, and BlueOS 4 arrives on the vivo WATCH 6.
Source: https://developers.vivo.com/product/ai/bluecode
Read the article: https://llmobile.news/ticker/vivo-bluecode/</description><category>Developer tools</category><category>Business</category></item><item><title>OPPO unveils ColorOS 17 with an on-device linear-attention model and 128K context</title><link>https://llmobile.news/ticker/oppo-coloros-17-on-device-model/</link><guid isPermaLink="true">https://llmobile.news/ticker/oppo-coloros-17-on-device-model/</guid><pubDate>Sat, 19 Sep 2026 12:19:04 +0200</pubDate><description>OPPO presented ColorOS 17 at its developer conference in Zhuhai, with an on-device large language model built on what the company calls its first Linear Attention architecture. Linear attention keeps compute cost proportional to the length of the input instead of growing quadratically. OPPO says the model supports 128K context natively, uses 48% less memory and 55% less energy than the previous generation, and keeps user data on the device while responding faster.
The on-device model is part of AndesGPT — OPPO&amp;amp;rsquo;s proprietary large model trained independently and available in sizes from around 1 billion to over 100 billion parameters, deployed in OPPO&amp;amp;rsquo;s &amp;amp;ldquo;1+N&amp;amp;rdquo; end-cloud architecture (a central large model plus specialized smaller models per task). AndesGPT was the first Chinese on-device LLM to register with the Cyberspace Administration on July 15, 2026, alongside Apple Intelligence, Huawei Xiaoyi, vivo BlueLM, Xiaomi HyperAI, Samsung Galaxy AI and ZTE Nubia Doubao.
The update also brings Persona X, a memory engine that builds on data, environmental and behavioral perception and records, forgets and reflects on what it learns about the user. The Xiao Bu assistant (Breeno outside China) gets longer context and better intent recognition. OPPO reports a 19% gain in internal satisfaction testing with device-cloud collaboration, and the assistant can now open 200 Alipay services, plan routes in Tencent Maps and place WeChat messages and calls by voice command.
ColorOS 17 ships first on the upcoming Find X10 series and OnePlus 16, and OPPO, OnePlus and Realme are upgrading in sync. According to Gizmochina, the rollout to existing devices in China starts on October 8 with the Find X9 series and Find N6. The Find X8 series, Find N5 and tablets follow on October 16, and older flagship and mid-range devices from November 26. No global timeline has been announced.
Source: https://www.jjckb.cn/20260918/447b6648587442e4b8e7b6594720fe3d/c.html
Read the article: https://llmobile.news/ticker/oppo-coloros-17-on-device-model/</description><category>Memory</category><category>Android</category></item><item><title>Arm unveils CSS for Mobile 2 with C2 CPU cluster, up to 1.7x faster on AI models</title><link>https://llmobile.news/ticker/arm-css-for-mobile-2/</link><guid isPermaLink="true">https://llmobile.news/ticker/arm-css-for-mobile-2/</guid><pubDate>Fri, 18 Sep 2026 20:26:52 +0200</pubDate><description>Arm has introduced Arm CSS for Mobile 2, a compute platform for smartphone chips that combines the C2 CPU cluster, the Mali G2-Ultra NX GPU and the SI L2 system interconnect. The C2 cluster pairs C2-Ultra and C2-Pro CPUs with two SME2 units, the Scalable Matrix Extension 2 that speeds up matrix math for AI on the CPU. Arm says this doubles the SME2 capability of the previous-generation configuration and reports up to 1.7x performance across the latest AI models.
The cluster delivers up to 15% higher single-thread performance, 15% faster web browsing, 12% faster app launch and 12% higher multi-thread performance, Arm reports. For AI, it cites a peak uplift of up to 70% in selected tasks. Its slide compares speech, personal memory retrieval and prefill, the phase where a model reads the prompt, against the C1-Ultra with SME2. In a representative agentic workflow covering speech processing, memory retrieval, reasoning, app execution and web browsing, the C2-Ultra with two SME2 units finishes 24% faster than the previous generation, according to Arm.
Arm&amp;amp;#39;s own comparison of the C2-Ultra with the C1-Ultra, with the AI tasks measured against the C1-Ultra with SME2Credit: Arm The example flagship configuration in Arm&amp;amp;rsquo;s slides has two C2-Ultra and six C2-Pro cores. The Mali G2-Ultra NX GPU integrates neural accelerators into its shader cores and adds a new execution engine and a third-generation ray tracing unit for neural graphics. Arm says the SI L2 interconnect provides lower-latency access, higher bandwidth, coherency and quality-of-service controls for CPU, GPU and other resources working at the same time. Partners can use each component on its own or combine them with custom and third-party IP.
Arm&amp;amp;#39;s slide for the C2-Ultra CPU with an example flagship cluster layoutCredit: Arm On the software side, Arm lists KleidiAI, its optimized libraries for Arm CPUs including SME2 paths, plus integrations with common AI frameworks. The Arm AI Portal offers validated models with performance and accuracy data, code examples and deployment resources, and the Arm MCP Server connects them to agentic development tools. vivo says it is bringing Arm Neural Technology to its latest flagship smartphones built on the platform, aimed at mobile gaming. Arm&amp;amp;rsquo;s post names no launch dates for devices with CSS for Mobile 2.
Source: https://newsroom.arm.com/blog/arm-css-for-mobile-2-and-c2-cpu-cluster
Read the article: https://llmobile.news/ticker/arm-css-for-mobile-2/</description><category>Chips</category><category>Android</category></item><item><title>MediaTek launches Dimensity 9600 Pro, a 2nm chip for on-device models up to 30B parameters</title><link>https://llmobile.news/ticker/mediatek-dimensity-9600-pro/</link><guid isPermaLink="true">https://llmobile.news/ticker/mediatek-dimensity-9600-pro/</guid><pubDate>Fri, 18 Sep 2026 17:55:18 +0200</pubDate><description>MediaTek announced the Dimensity 9600 Pro, a flagship smartphone chip built on a 2nm process. According to the company, its NPU 1090 supports on-device applications with models of up to 30B parameters. MediaTek reports 51% higher LLM prefill performance, the phase where the model reads the prompt, and 55% higher token generation per watt, measured on demo devices in its own labs.
The chip pairs the NPU 1090 with a second-generation Super Efficient NPU, which MediaTek says cuts power consumption for always-on AI by 40%. The platform supports LPDDR6 memory and UFS 5.0 storage.
In September 2025, MediaTek announced that it had completed the tape-out, the final design handoff to the fab, of a flagship chip on TSMC&amp;amp;rsquo;s N2P 2nm process, with volume production expected in late 2026. The CPU uses a 2+3+3 layout of eight big cores, with two C2-Ultra cores at up to 4.55 GHz. MediaTek states up to 17% higher single-core and up to 15% higher multi-core performance over the previous generation, and up to 61% lower multi-core power consumption.
The first smartphones with the Dimensity 9600 Pro and the related Dimensity 9600M are expected to launch this quarter, according to MediaTek.
Source: https://www.mediatek.com/press-room/mediatek-dimensity-9600-pro-sets-new-standard-for-flagship-smartphone-chips
Read the article: https://llmobile.news/ticker/mediatek-dimensity-9600-pro/</description><category>Chips</category><category>Android</category></item><item><title>Google lists first Gemini Nano 4 phones, requires Nano 3 for Gemini Intelligence</title><link>https://llmobile.news/ticker/gemini-nano-4-first-devices/</link><guid isPermaLink="true">https://llmobile.news/ticker/gemini-nano-4-first-devices/</guid><pubDate>Thu, 27 Aug 2026 16:00:00 +0200</pubDate><description>Google&amp;amp;rsquo;s ML Kit GenAI documentation now lists the first devices running nano-v4, 9to5Google reports. The list covers the Pixel 11, Pixel 11 Pro, Pixel 11 Pro XL and Pixel 11 Pro Fold, plus Samsung&amp;amp;rsquo;s Galaxy Z Flip8, Galaxy Z Fold8 and Galaxy Z Fold8 Ultra.
The same documentation sets Nano v3 or greater as the requirement for Gemini Intelligence, Google&amp;amp;rsquo;s on-device feature set. According to the report, that requirement first appeared in May 2026, was removed, and has now been reinstated. Listed hardware requirements include 12 GB or more of RAM, a qualified flagship system-on-chip, five or more OS upgrades and six years of security support.
Gemini Intelligence features named in the report include Rambler and Proactive Assistance on Pixel 11, and task automation across more than 40 apps on Samsung&amp;amp;rsquo;s foldables.
Source: https://9to5google.com/2026/08/27/gemini-intelligence-nano-4/
Read the article: https://llmobile.news/ticker/gemini-nano-4-first-devices/</description><category>Android</category></item><item><title>Pixel 11 series: Tensor G6 adds 50 percent more TPU compute</title><link>https://llmobile.news/ticker/pixel-11-tensor-g6/</link><guid isPermaLink="true">https://llmobile.news/ticker/pixel-11-tensor-g6/</guid><pubDate>Wed, 12 Aug 2026 19:00:00 +0200</pubDate><description>Google has announced the Pixel 11, Pixel 11 Pro and Pixel 11 Pro XL, built around the Tensor G6 chip. Google states that Tensor G6 packs 50 percent more TPU compute and, paired with the latest Gemini Nano model, processes on-device AI tasks up to 3.5 times faster while using up to 3.5 times less energy.
The company also cites an upgraded CPU with 25 percent faster web browsing and 15 percent quicker app launches, and says the chip powers the 30x Super Zoom on the 5x telephoto lens. Google does not publish RAM figures, model sizes or per-task latency in the announcement.
Pre-orders opened on 12 August, with retail availability from 20 August.
▶Meet Google Pixel 11 ProLoading connects your browser to www.youtube-nocookie.com, which may process your IP address and use cookies.
Load contentOpen externallyVideo: Made by Google.
Source: https://blog.google/products-and-platforms/devices/pixel/google-pixel-11-pro-xl/
Read the article: https://llmobile.news/ticker/pixel-11-tensor-g6/</description><category>Chips</category><category>NPU</category></item><item><title>Gemini Nano 4 ships on Samsung foldables with ML Kit Prompt API access</title><link>https://llmobile.news/ticker/gemini-nano-4-mlkit-prompt-api/</link><guid isPermaLink="true">https://llmobile.news/ticker/gemini-nano-4-mlkit-prompt-api/</guid><pubDate>Wed, 22 Jul 2026 18:00:00 +0200</pubDate><description>Google&amp;amp;rsquo;s Android developer blog states that Samsung&amp;amp;rsquo;s new foldable devices come with Gemini Nano 4, which it calls its latest on-device model. The post credits Nano 4 with support for more than 140 languages and better multimodal understanding.
Apps reach the model through ML Kit&amp;amp;rsquo;s Prompt API, which sends natural language requests on-device to Gemini Nano. It takes text, or a combination of image and text, and returns text or structured output. Google names structured output and thinking mode as the features to use for on-device intelligence.
The ML Kit release notes dated 14 July 2026 record the structured output API, system instructions and thinking mode arriving in the Prompt API, along with multi-image support and an output token limit raised to 4,096 tokens. A note dated 21 July records a fix for Gemini Nano v4 compatibility in the Prompt API on non-Pixel devices. The Prompt API moved from alpha to beta in January 2026 and carries no service level agreement or deprecation policy.
The post also points developers to app functions, which share an app&amp;amp;rsquo;s capabilities with the Gemini Intelligence system. Its remaining sections cover adaptive layouts, fold-aware design, CameraX and Wear OS widgets.
Source: https://android-developers.googleblog.com/2026/07/optimize-galaxy-screen-sizes.html
Read the article: https://llmobile.news/ticker/gemini-nano-4-mlkit-prompt-api/</description><category>Android</category><category>Developer tools</category></item><item><title>Galaxy S24 becomes the second phone line to run Gemini Nano</title><link>https://llmobile.news/ticker/galaxy-s24-gemini-nano/</link><guid isPermaLink="true">https://llmobile.news/ticker/galaxy-s24-gemini-nano/</guid><pubDate>Wed, 17 Jan 2024 20:00:00 +0100</pubDate><description>Google announced on January 17, 2024 that the Galaxy S24 series runs Gemini Nano on device, which made it the first phone line outside the Pixel 8 Pro to do so. Google names Magic Compose in Google Messages as the feature that runs locally, and states that the data does not leave the phone.
Image: Google. The rest of the announced features run in the cloud. Google lists Gemini Pro behind summarisation in Samsung Notes and Voice Recorder as well as keyboard features, with Generative Edit in the Gallery app built on Imagen 2. Gemini Ultra was still in testing at the time.
The launch also introduced Circle to Search, a gesture that searches whatever is circled or highlighted on screen without switching apps. Samsung published its own account of the launch on the Samsung Newsroom.
Source: https://blog.google/products/android/google-ai-samsung-galaxy-s24/
Read the article: https://llmobile.news/ticker/galaxy-s24-gemini-nano/</description><category>Android</category></item><item><title>Gemini Nano ships on the Pixel 8 Pro and Android gets AICore</title><link>https://llmobile.news/ticker/gemini-nano-pixel-8-pro/</link><guid isPermaLink="true">https://llmobile.news/ticker/gemini-nano-pixel-8-pro/</guid><pubDate>Wed, 06 Dec 2023 18:00:00 +0100</pubDate><description>Google brought Gemini Nano to the Pixel 8 Pro in its December 2023 feature drop, where it powers Summarize in Recorder and Smart Reply in Gboard. Google calls the Pixel 8 Pro the first smartphone engineered for Gemini Nano and runs the model on the Tensor G3.
Your browser does not support this video. This video could not be loaded. Use the link below to open it directly.
Video: Google. Gemini Nano summarising a recording in the Recorder app. Open the Android Developers post According to Google, running the model locally helps prevent sensitive data from leaving the phone and lets the features work without a network connection. Summarize in Recorder launched in English. Smart Reply in Gboard launched globally on the United States English keyboard layout, starting with WhatsApp, Line and KakaoTalk.
On the same day Google introduced AICore, a system service in Android 14 that handles model management, runtimes and safety features for Gemini Nano. It supports Low Rank Adaptation, so developers can build small adapters trained on their own data, and it targets the Google Tensor TPU as well as NPUs from Qualcomm, Samsung and MediaTek. Google describes the service as isolated from the network by design and opened access through an early access programme.
Diagram: Google.
Source: https://blog.google/products/pixel/pixel-feature-drop-december-2023/
Read the article: https://llmobile.news/ticker/gemini-nano-pixel-8-pro/</description><category>Android</category><category>Developer tools</category></item></channel></rss>