<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Apple · LLMobile.news</title><link>https://llmobile.news/companies/apple/</link><description>A concise news ticker covering AI on mobile devices, local models, apps and hardware.</description><language>en-GB</language><atom:link href="https://llmobile.news/companies/apple/index.xml" rel="self" type="application/rss+xml"/><item><title>Hugging Face adds native GGUF inference to Transformers, starting with Qwen3.5 on Macs</title><link>https://llmobile.news/ticker/transformers-gguf-metal/</link><guid isPermaLink="true">https://llmobile.news/ticker/transformers-gguf-metal/</guid><pubDate>Wed, 23 Sep 2026 21:56:00 +0200</pubDate><description>Hugging Face added native support in Transformers for GGUF, the quantized file format used by llama.cpp, so developers can load those checkpoints through the library&amp;amp;rsquo;s standard from_pretrained() API instead of switching to a separate runtime. Engineers Marc Sun, Arthur Zucker and Lysandre wrote in the announcement that on Apple Silicon, weights stay packed in their GGUF blocks and Metal GPU kernels run the matmuls directly on the compressed data, skipping a full dequantization step. The feature currently covers Qwen3.5, including its dense and MoE variants, plus compatible Qwen3.8 checkpoints.
The packed inference path draws its Metal kernels from a kernels library that fetches them from the Hub, including ggml-attn, the same flash attention implementation llama.cpp runs for both prefill and decode. Separate kernels handle quantization, normalization and the gated delta net used in the hybrid attention layers of both Qwen3.5 and Qwen3.8. Architectures outside this initial set, or devices other than Apple Silicon, fall back to a legacy loader that dequantizes GGUF weights into a dense model at load time, a path that already covers Llama, Mistral, Qwen2, Qwen2Moe, Phi3, Bloom, Falcon, StableLM, GPT2 and Starcoder2, among others.
On a MacBook Pro M2 Max with 32 GB of memory, Hugging Face reports about 70 tok/s generating with Qwen3.5-4B at Q4_K_M quantization, close to the 72 tok/s the company measured for llama.cpp on the same hardware. The company says its new layer kernels add 51 to 109 percent more throughput over the packed-quantization-only path, depending on the model, with Qwen3.5-4B gaining 59 percent, Qwen3.8-27B 51 percent and Qwen3.5-35B-A3B 109 percent. Further gains come from generation-loop changes such as dropping unneeded attention masks early and checking stop conditions asynchronously. Quantization options for the 4B model range from a 2.74 GB Q4_K_M file up to an 8.42 GB BF16 file.
The packed path requires Transformers&amp;amp;rsquo; main development branch together with a matching version of the kernels library, and Hugging Face says it is limited to MPS devices for now, with padding and batched requests still needing optimization. Models can also be served through an OpenAI-compatible API with transformers serve, which identifies a GGUF checkpoint as &amp;amp;lt;repo&amp;amp;gt;:&amp;amp;lt;file&amp;amp;gt;.gguf since a single Hub repository can hold several quantizations of the same model.
Source: https://huggingface.co/blog/transformers-llama-cpp-quants
Read the article: https://llmobile.news/ticker/transformers-gguf-metal/</description><category>Quantisation</category><category>GPU</category><category>Benchmarks</category></item><item><title>Apple launches M6 Mac mini and M5 Ultra Mac Studio for local AI</title><link>https://llmobile.news/ticker/apple-mac-mini-mac-studio-m5/</link><guid isPermaLink="true">https://llmobile.news/ticker/apple-mac-mini-mac-studio-m5/</guid><pubDate>Wed, 23 Sep 2026 09:35:00 +0200</pubDate><description>Apple has started shipping an updated Mac mini and a new Mac Studio topping out at the M5 Ultra chip. Apple says the configuration lets &amp;amp;ldquo;AI researchers run enormous LLMs entirely on device,&amp;amp;rdquo; citing up to 512GB of unified memory and 1.2TB/s of memory bandwidth on the top Mac Studio.
The Mac mini starts with a 12-core M6 chip, a 12-core GPU with Neural Accelerators in every core, and a dual 16-core Neural Engine, running 16GB of unified memory at up to 170GB/s of bandwidth. A Mac mini with the M5 Pro chip steps up to an 18-core CPU, a 20-core GPU and up to 64GB of memory at 307GB/s. Apple prices the M6 Mac mini from $899 and the M5 Pro model from $1,699.
The Mac Studio comes in M5 Max and M5 Ultra configurations, with the Max offering an 18-core CPU and up to a 40-core GPU with up to 128GB of unified memory, starting at $2,499. The M5 Ultra model reaches a 36-core CPU, an 80-core GPU and up to 512GB of memory, starting at $5,499. Apple says Thunderbolt 5 with Remote Direct Memory Access lets multiple Mac Studio systems cluster together, giving up to 3x faster performance for distributed AI inference compared with a single system.
The new Mac mini and Mac Studio models became available on September 22, running macOS 27 with Apple Intelligence and Siri AI features. Apple also offers leasing plans starting at $48.99 a month for the M5 Max Mac Studio and $110.10 a month for the M5 Ultra model over 36 months.
Source: https://www.apple.com/newsroom/2026/09/the-new-mac-mini-and-mac-studio-are-available-today/
Read the article: https://llmobile.news/ticker/apple-mac-mini-mac-studio-m5/</description><category>Chips</category><category>Memory</category></item><item><title>iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine</title><link>https://llmobile.news/ticker/iphone-18-pro-a20-pro-neural-engine/</link><guid isPermaLink="true">https://llmobile.news/ticker/iphone-18-pro-a20-pro-neural-engine/</guid><pubDate>Mon, 21 Sep 2026 00:00:00 +0200</pubDate><description>Apple has introduced the iPhone 18 Pro and iPhone 18 Pro Max with the A20 Pro chip. According to Apple, the chip has a new dual 16-core Neural Engine, 32 cores in total, which the company describes as double the AI processing power of A19 Pro. Apple says the Neural Engine accelerates on-device AI models and computational photography.
Image: Apple. The A20 Pro also offers 50 percent more memory bandwidth than A19 Pro, and its 6-core CPU includes integrated Neural Accelerators. Apple positions the chip for &amp;amp;ldquo;more advanced on-device AI workloads&amp;amp;rdquo; alongside games.
The phones ship with iOS 27, Apple Intelligence and the new Siri AI. Apple states that Apple Intelligence uses on-device processing together with Private Cloud Compute. Pre-orders start on 12 September, with availability from 18 September; prices start at $1,199 for the iPhone 18 Pro and $1,299 for the iPhone 18 Pro Max, both with 256 GB.
Update, September 20, 2026. The retailer ZEERA reports results from a Max Tech side-by-side comparison of the iPhone 18 Pro Max and iPhone 17 Pro Max. Local LLM generation rose from 33.5 to 55.4 tok/s, a gain of 65 percent, and measured memory bandwidth from 66.5 to 111.8 GB/s, a gain of 68 percent. The comparison also lists 25 percent higher single-core and 29 percent higher multi-core CPU scores in Geekbench 7, and a 48 percent higher graphics score.
Apple&amp;amp;rsquo;s developer page for Apple Intelligence says the Foundation Models framework gives Swift apps direct access to the same on-device model that powers Apple Intelligence. New multimodal prompts pass images alongside text, and the model can call Vision framework tools such as OCR and barcode readers, all on-device according to Apple. Apps can also work with any model that conforms to the Language Model protocol, including cloud models like Claude and Gemini. Members of the App Store Small Business Program with fewer than 2 million first-time downloads get the next generation of Apple Foundation Models on Private Cloud Compute at no cloud API cost.
The new Siri AI is a three-tier hybrid system, not fully on-device. A new System Orchestrator in iOS 27 routes every request to the right tier:
Tier 1 — On-Device: Simple tasks (timers, app launches, local context search) run entirely on Apple&amp;amp;rsquo;s Foundation Models (~3B parameters, 4-bit quantized) on the Neural Engine. Nothing leaves the iPhone. Tier 2 — Private Cloud Compute: Moderately complex requests route to Apple&amp;amp;rsquo;s own servers on Apple Silicon (M2 Ultra, now M5) in hardware-isolated Secure Enclaves. Personal data is stripped and tokenized. Apple retains control; no data retention. Tier 3 — Google Gemini: The heaviest reasoning, real-time knowledge, and multimodal tasks are sent through PCC&amp;amp;rsquo;s privacy proxy to a custom 1.2-trillion-parameter Gemini model on Nvidia Blackwell B200 GPUs in Google Cloud. Before any query reaches Google, the user&amp;amp;rsquo;s identity and personal data are removed. Google is contractually barred from training on these queries. Apple reportedly pays Google ~$1 billion per year for this partnership. The System Orchestrator decides the routing threshold in real time based on task complexity and context size. Apple&amp;amp;rsquo;s public materials only name On-Device and PCC as the two processing locations — Gemini is never mentioned on Apple&amp;amp;rsquo;s consumer pages.
Foundation Models can now work with any language model, including Apple&amp;amp;rsquo;s own, Claude, and Gemini, through a single Swift interface. Apple did not publish details of the routing logic or which types of queries are sent off-device.
Update, September 21, 2026. Adrien Grondin, who works on the Locally AI app, says on X that the iPhone 18 Pro runs a 27B model at double the speed of the iPhone 17 Pro. His 11-second video shows the demo, but the post names neither the model nor its quantization and gives no tok/s figures.
Wccftech reports on the same demo, attributing it to Grondin and saying the model is Bonsai 2 27B from Prism ML — a ternary-weight version of Qwen3.8 27B at 1.76 bits per weight, 5.9 GB total. PrismML describes Bonsai 2 27B as released under Apache 2.0, running on Apple devices through MLX and on NVIDIA GPUs through CUDA.
The iPhone 18 Pro is much better than what I expected for on-device AI
It can run a 27B model at double the speed of the 17 Pro
The new A20 Pro chip is a beast pic.twitter.com/gb11TvBfn0
&amp;amp;mdash; Adrien Grondin (@adrgrondin) September 19, 2026
Source: https://www.apple.com/newsroom/2026/09/apple-debuts-iphone-18-pro-and-iphone-18-pro-max/
Read the article: https://llmobile.news/ticker/iphone-18-pro-a20-pro-neural-engine/</description><category>iPhone</category><category>Chips</category><category>NPU</category></item><item><title>Prism ML releases Bonsai 2 27B for iPhone, iPad and Mac, with a 3.9 GB 1-bit version</title><link>https://llmobile.news/ticker/bonsai-2-27b/</link><guid isPermaLink="true">https://llmobile.news/ticker/bonsai-2-27b/</guid><pubDate>Thu, 17 Sep 2026 00:00:00 +0200</pubDate><description>Prism ML has released Bonsai 2 27B, a compressed version of Qwen3.8 27B that runs on Apple devices including iPhone, iPad and Mac via MLX, using custom low-bit kernels. On NVIDIA GPUs it runs via CUDA. For the iPhone 17 Pro, Prism lists a 1-bit version at 3.9 GB with a 262K-token context window.
The main Ternary Bonsai 2 27B model stores each weight as -1, 0 or +1 with FP16 scaling per group of weights. This works out to 1.76 effective bits per weight and a total footprint of 5.9 GB, which Prism says is more than 9x smaller than the full-precision model. The model accepts text and image input.
Prism reports an aggregate benchmark score of 83.9 for the ternary model, compared with 85.4 for Qwen3.8 27B, which it describes as 98.2% retention. The model scores 82.66 in instruction following against 81.25 for the original, and trails it in knowledge and reasoning (83.95 against 86.66), vision (78.59 against 81.64) and agentic tool calling (77.57 against 79.74). Prism also gives 96.57 for math and 81.58 for coding, against 97.06 and 82.17 for Qwen3.8 27B.
According to Prism, the ternary model generates up to 143 tok/s on an NVIDIA GeForce RTX 5090 and 46.8 tok/s on an Apple M5 Max. On an RTX 4090 it uses 0.714 mWh per token, which the company says is 40% less than a full-precision 8B model. The project repository lists peak memory for text-only use at about 7.8 GB at 4K context and 13.7 GB at 100K context, with images adding about 0.9 GB.
The model is released under the Apache 2.0 license. The Hugging Face collection holds GGUF and MLX 2-bit versions, and the repository documents running it with llama.cpp on Macs with Apple silicon and on CUDA, Vulkan, ROCm and CPU. A WebGPU demo runs in the browser on Hugging Face Spaces.
Source: https://prismml.com/news/bonsai-2-27b
Read the article: https://llmobile.news/ticker/bonsai-2-27b/</description><category>iPhone</category></item><item><title>Desert Ant Labs releases 18 on-device models for speech, audio and privacy tasks</title><link>https://llmobile.news/ticker/desert-ant-labs/</link><guid isPermaLink="true">https://llmobile.news/ticker/desert-ant-labs/</guid><pubDate>Tue, 08 Sep 2026 00:00:00 +0200</pubDate><description>Desert Ant Labs, a newly launched company, has introduced 18 on-device models as its first release, each built for a single task such as transcription, audio enhancement, personal data masking, language detection and video summarization, according to its launch post. Twelve of the models are stable and six are in beta.
The first collection of Desert Ant Labs models, as shown in the launch post.Desert Ant Labs
Voz, the transcription model, converts 10 minutes of audio in 2 seconds on an iPhone, according to the company. Desert Ant Labs reports 298x realtime speed on an iPhone 17 Pro and 319x on an M3 Ultra, and says Voz is 4.7x faster than Whisper and returns word-level timestamps. The audio enhancement model Clear is 9 MB and processes a 5-minute recording in 1 second, at 302x realtime on an iPhone 16 Pro and 345x on a MacBook Pro with M5.
Redact masks personal data in 27 languages, takes up 12 MB and detects 88.8% of personal data, the company says. Tongue identifies 84 languages in 2 MB and reaches an accuracy of 0.933 from three words. Clips turns a 10-minute video into clips in 5 seconds with a 284 MB model, which Desert Ant Labs says is 10x faster than Claude Sonnet and uses 470x less energy at the same quality.
All models are free up to 100k monthly active devices, with no tokens and no logins, according to the Desert Ant Labs website. Developers integrate them through SDKs for Swift, Kotlin and JavaScript, which are published on GitHub. The models are also on Hugging Face, where browser demos let users try them, and the company also offers a command-line tool for testing. Documentation is at desertant.com/docs.
Source: https://desertant.com/blog/introducing-desert-ant-labs/
Read the article: https://llmobile.news/ticker/desert-ant-labs/</description><category>Developer tools</category><category>Benchmarks</category></item><item><title>Ornith releases Ornith-1.5, a 9B model with a mobile build for iPhone and Android</title><link>https://llmobile.news/ticker/ornith-1-5/</link><guid isPermaLink="true">https://llmobile.news/ticker/ornith-1-5/</guid><pubDate>Wed, 19 Aug 2026 00:00:00 +0200</pubDate><description>Ornith has released Ornith-1.5, a model family with a 9B dense model, a 35B mixture-of-experts model that activates about 3B parameters per token and a 397B mixture-of-experts model. The 9B model also comes as Ornith-1.5-9B-Mobile, a quantized variant (about 1,5 GB on disk) designed for iPhone and Android.
The 9B model scores 47.0 on Terminal-Bench 2.1 with the Claude Code harness and 70.6 on SWE-bench Verified in Ornith&amp;amp;rsquo;s tests. The company says the model matches or exceeds much larger models such as Gemma 4-31B and Qwen 3.6-35B. In the company&amp;amp;rsquo;s chart, Qwen3.6-35B-A3B leads on SWE-bench Verified with 73.4 and on Terminal-Bench 2.1 with 52.5, while the 9B model scores 86.4 on GPQA Diamond and 54.2 on MCP-Atlas. The previous Ornith-1.0-9B reaches 43.1 on Terminal-Bench 2.1 in the same chart.
Ornith&amp;amp;#39;s own benchmark figures for the 9B model. According to Ornith, its training loop lets the model propose new tasks, generate task-specific scaffolds and produce solution rollouts, with the reward from the rollouts propagated across all three stages. The company reports that the 35B model scores 67.8 on Terminal-Bench 2.1 with the Terminus-2 harness, against 52.5 for Qwen 3.6-35B. The models are on Hugging Face, with GGUF builds of all three sizes and MLX builds of the 9B and 35B models.
Source: https://ornith.ai/ornith_1_5.html
Read the article: https://llmobile.news/ticker/ornith-1-5/</description><category>Android</category></item><item><title>Prism ML releases Bonsai 27B, a 3.9 GB 1-bit model that fits an iPhone 17 Pro</title><link>https://llmobile.news/ticker/bonsai-27b/</link><guid isPermaLink="true">https://llmobile.news/ticker/bonsai-27b/</guid><pubDate>Tue, 14 Jul 2026 00:00:00 +0200</pubDate><description>Prism ML has released Bonsai 27B, a compressed version of Qwen3.6 27B for Mac, iPhone, iPad and NVIDIA GPUs. The 1-bit variant needs 3.9 GB, which Prism says fits within the memory budget of an iPhone 17 Pro. The company notes that a phone never exposes its full memory to an app and that a 12 GB iPhone offers about 6 GB to a model.
The ternary variant stores each weight as -1, 0 or +1 and takes up 5.9 GB at 1.71 effective bits per weight. The 1-bit variant stores weights as -1 or +1 at 1.125 effective bits. Prism applies the low-bit format to embeddings, attention layers, MLPs and the output head. Both models accept images through a 4-bit vision tower and offer a 262K-token context window.
Across a 15-benchmark suite, Prism reports an overall score of 85.0 for Qwen3.6 27B, 80.5 for the ternary model and 76.1 for the 1-bit model, which it describes as retaining 95% and 90% of the original. The largest gaps show up in agentic tool calling, at 80.0 for the original, 74.0 for ternary and 66.0 for 1-bit, and in vision, at 72.6, 65.2 and 59.6. Prism also publishes an &amp;amp;ldquo;intelligence density&amp;amp;rdquo; figure, defined as the negative log of the error rate divided by model size in GB, of 0.530 for the 1-bit model and 0.400 for the ternary one.
On an NVIDIA GeForce RTX 5090, Prism measures 163 tok/s for the 1-bit model and 134 tok/s for the ternary model. On an Apple M5 Max the figures are 87 and 58 tok/s. The models also support speculative decoding, where a small draft model proposes tokens that the main model then verifies.
Bonsai 27B is released under the Apache 2.0 license, with builds for MLX on Apple devices and for CUDA on NVIDIA GPUs. Weights are in the Hugging Face collection, and the Bonsai-demo repository covers running them. Prism also lists the Locally AI iOS app and a free, limited-time developer API on Together.ai as ways to try the model.
Source: https://prismml.com/news/bonsai-27b
Read the article: https://llmobile.news/ticker/bonsai-27b/</description><category>Quantisation</category><category>Benchmarks</category><category>Open source</category></item><item><title>Prism ML releases Bonsai Image 4B, a 1-bit image model that runs on iPhone</title><link>https://llmobile.news/ticker/bonsai-image-4b/</link><guid isPermaLink="true">https://llmobile.news/ticker/bonsai-image-4b/</guid><pubDate>Tue, 26 May 2026 00:00:00 +0200</pubDate><description>Prism ML has released Bonsai Image 4B, an image generation model with 4 billion parameters in a 1-bit and a ternary version. The company says it is the first image model in its parameter class to run directly on an iPhone. It generates a 512x512 image in 9.4 seconds on an iPhone 17 Pro Max and in about 6 seconds on an M4 Pro Mac, where Prism reports up to 5.6 times the speed of the stock full-precision MFLUX pipeline.
Sample images from the Bonsai Image 4B announcement.Prism ML
The diffusion transformer takes 0.93 GB in the 1-bit version, which uses 1.125 effective bits per weight, and 1.21 GB in the ternary version at 1.71 bits. Prism lists 7.75 GB for the comparison model FLUX.2 Klein 4B. The full deployment payload on Apple Silicon is 3.42 GB and 3.88 GB, against 15.97 GB for FLUX.2 Klein 4B. At runtime, 512x512 generation uses 1.5 GB and 1.96 GB against 11.74 GB, and 1024x1024 generation uses 1.95 GB and 2.38 GB against 14.39 GB.
Prism scores the ternary version at 0.723 on GenEval, 12.22 on HPSv3 and 0.851 on DPG-Bench, which it summarizes as 95% of FLUX.2 Klein 4B overall. The 1-bit version scores 0.671, 11.15 and 0.822, or 88%, while FLUX.2 Klein 4B scores 0.819, 12.84 and 0.853. GenEval measures object composition and attribute binding, HPSv3 human preference and DPG-Bench how well dense prompts are followed.
The models run on iPhone, iPad and Mac via MLX low-bit paths and on CUDA GPUs via Gemlite low-bit kernels. They are released as open weights and code under the Apache 2.0 license, and Prism ships the iOS app Bonsai Studio for trying the model on iPhone. Weights are in the Hugging Face collection, code is in the GitHub repository, and a WebGPU demo runs in the browser.
Source: https://prismml.com/news/bonsai-image-4b
Read the article: https://llmobile.news/ticker/bonsai-image-4b/</description><category>Quantisation</category><category>Demos</category><category>Open source</category></item><item><title>Locally AI runs Llama, Gemma and Qwen offline on iPhone and iPad through MLX</title><link>https://llmobile.news/ticker/locally-ai/</link><guid isPermaLink="true">https://llmobile.news/ticker/locally-ai/</guid><pubDate>Wed, 08 Apr 2026 18:00:00 +0200</pubDate><description>Adrien Grondin released Locally AI on April 13, 2025, an iPhone and iPad app that downloads open-weight language models and runs them entirely on the device. The App Store listing gives that release date and lists the app as free with no in-app purchases. Once a model is on disk the app needs no account and no internet connection, and Grondin states that no data is collected and nothing is sent to a cloud service.
The site at launch named Meta Llama 3.2 and Llama 3.1, Google Gemma 2, Qwen 2 and DeepSeek R1 as the models on offer, and described them as language and vision models, so a picture can go into the prompt alongside text. A user picks a model, downloads it once and can edit the system prompt the model runs under.
Grondin builds the app on MLX, Apple&amp;amp;rsquo;s machine learning framework for Apple silicon, and attributes the performance to unified memory, where the CPU and the GPU reach the same memory pool so arrays need not be copied between them. The app fetches open-weight model files itself rather than calling any model that ships with the operating system.
Neither the site nor the App Store entry names a minimum chip. The site says only that Locally AI covers recent iPhone and iPad models and is optimised for Apple silicon, and the listing asks for iOS 18.1 or iPadOS 18.1 or later. No generation speed, memory requirement or model file size is given anywhere in the developer&amp;amp;rsquo;s own material, and the launch site was a bought template whose team, testimonial and FAQ sections still carried placeholder Latin text.
Update, April 8, 2026. Grondin added support for Apple&amp;amp;rsquo;s on-device Foundation Models framework in September 2025, so the app can open a conversation with the model built into iOS 26 without downloading anything, which Simon Willison recorded on September 21, 2025. A Mac version followed and requires macOS 26 or later. LM Studio announced on April 8, 2026 that Locally AI had joined the company and that Grondin had moved with it, and the App Store entry now reads Locally AI by LM Studio with Element Labs Inc as the seller.
Source: https://locallyai.app/
Read the article: https://llmobile.news/ticker/locally-ai/</description><category>Apple Silicon</category><category>MLX</category><category>iOS</category><category>iPhone</category></item><item><title>Prism ML releases 1-bit Bonsai 8B, an 8.2B-parameter model that needs 1.15 GB</title><link>https://llmobile.news/ticker/1-bit-bonsai-8b/</link><guid isPermaLink="true">https://llmobile.news/ticker/1-bit-bonsai-8b/</guid><pubDate>Tue, 31 Mar 2026 00:00:00 +0200</pubDate><description>Prism ML has released 1-bit Bonsai 8B, a language model with 8.2 billion parameters that stores every weight in a single bit. Embeddings, attention layers, MLP layers and the LM head are all 1-bit, which brings the model down to 1.15 GB. The company also released 4B and 1.7B versions of the family.
On Apple hardware Prism reports about 40 tok/s on an iPhone 17 Pro and about 44 tok/s on an iPhone 17 Pro Max, against 131 tok/s on an M4 Pro Mac. An NVIDIA RTX 4090 reaches 368 tok/s, compared with 59 tok/s for a 16-bit 8B model in Prism&amp;amp;rsquo;s chart. The 16-bit model reaches 16 tok/s on the M4 Pro and does not fit on the iPhone 17 Pro Max.
Prism measured 0.074 mWh per token on the M4 Pro and 0.068 mWh per token on the iPhone 17 Pro Max, which it puts at 4 to 5 times less than 16-bit counterparts. On the RTX 4090 the figures are 0.276 mWh for 1-bit Bonsai 8B and 1.134 mWh for the 16-bit model.
Energy per token for 1-bit Bonsai 8B and a 16-bit 8B model, as measured by Prism ML.Prism ML
The company gives the 8B model an average benchmark score of 70.5, made up of 65.7 on MMLU Redux, 50.0 on MuSR, 88.0 on GSM8K, 73.8 on HumanEval+, 79.8 on IFEval and 65.7 on BFCLv3. It compares the model with 8B-class models such as Qwen3 8B, Llama 3.1 8B, Ministral3 8B and LFM2 8B. On its own &amp;amp;ldquo;intelligence density&amp;amp;rdquo; metric, a score per GB of model size, Prism lists 1.06 for Bonsai 8B and 0.10 for Qwen3 8B.
The models are released under the Apache 2.0 license. They run on Apple devices via MLX and on NVIDIA GPUs via llama.cpp CUDA, and the iOS app Locally AI can run them on iPhone. Weights are in the Hugging Face collection, and the demo repository holds a whitepaper and quick start instructions, along with a Colab notebook.
Source: https://prismml.com/news/bonsai-8b
Read the article: https://llmobile.news/ticker/1-bit-bonsai-8b/</description><category>Quantisation</category><category>Benchmarks</category><category>Open source</category></item><item><title>Apple ships Python bindings for the on-device Foundation Models framework</title><link>https://llmobile.news/ticker/apple-foundation-models-sdk/</link><guid isPermaLink="true">https://llmobile.news/ticker/apple-foundation-models-sdk/</guid><pubDate>Wed, 25 Feb 2026 20:47:00 +0100</pubDate><description>Apple published python-apple-fm-sdk, Python bindings for its Foundation Models framework, with a first beta release on February 25, 2026 according to the repository&amp;amp;rsquo;s release history. The package reaches the on-device model at the core of Apple Intelligence from Python on macOS, and installs with pip install apple-fm-sdk.
The Python API follows the Swift one. SystemLanguageModel reports whether the model is available and, if not, why, while LanguageModelSession holds a conversation and its async respond call returns text or streams it. A @fm.generable decorator marks a Python class for the model to fill in, and fm.guide adds per-field constraints such as a numeric range, as the README shows with a Cat class whose age is bounded to 0 through 20. Tool calling and session transcripts are exposed as well.
The bindings do not reimplement inference. The repository carries a Swift package named foundation-models-c that wraps the framework behind a C interface and builds it as a dynamic library, which the Python layer then calls. Apple writes in the documentation that the bindings run the Swift framework underneath, so evaluations reflect real on-device performance and behaviour, and its evaluation guide notes that inference calls are processed one at a time rather than in parallel at the macOS hardware level.
Apple lists macOS 26.0 or later, Xcode 26.0 or later with its agreement accepted in the Xcode app, Python 3.10 or later, and Apple Intelligence switched on. PyPI carries only source distributions for the package, so the Swift component is compiled locally during installation, which is what the Xcode requirement covers. Apple Intelligence on the Mac runs on Apple silicon machines only.
The code is under the Apache 2.0 license and the package metadata on PyPI classifies it as alpha. Apple pushed 0.1.0 to PyPI on March 8, 2026, added the Attachment API for sending images alongside text in 0.2.0 on June 8, 2026 for WWDC 2026, and exposed the model&amp;amp;rsquo;s context size and the token count of a given input in 0.2.1 on June 29, 2026. The README states the project is not yet taking contributions.
Source: https://github.com/apple/python-apple-fm-sdk
Read the article: https://llmobile.news/ticker/apple-foundation-models-sdk/</description><category>Developer tools</category><category>Apple Silicon</category><category>Open source</category></item><item><title>A19 Pro puts Neural Accelerators in every GPU core</title><link>https://llmobile.news/ticker/a19-pro-neural-accelerators/</link><guid isPermaLink="true">https://llmobile.news/ticker/a19-pro-neural-accelerators/</guid><pubDate>Tue, 09 Sep 2025 20:00:00 +0200</pubDate><description>Apple announced the A19 Pro alongside the iPhone 17 Pro on September 9, 2025. Its 6-core GPU carries Neural Accelerators built into each core, which Apple says work together with the 16-core Neural Engine to power AI models, graphics and games. The company states that the chip enables running large local language models on the phone.
The A19 Pro also has a 6-core CPU that Apple calls the fastest in any smartphone, along with a larger cache and more memory than the A18 Pro. Combined with a vapor chamber cooling system, Apple puts sustained performance at up to 40 percent above the previous generation.
Image: Apple. The vapor chamber Apple credits for the sustained performance gain.
Source: https://www.apple.com/newsroom/2025/09/apple-unveils-iphone-17-pro-and-iphone-17-pro-max/
Read the article: https://llmobile.news/ticker/a19-pro-neural-accelerators/</description><category>iPhone</category><category>Chips</category><category>Apple Silicon</category><category>NPU</category></item><item><title>Apple puts the cost of 2-bit compression at 3.4 MMLU points</title><link>https://llmobile.news/ticker/apple-foundation-models-report-2025/</link><guid isPermaLink="true">https://llmobile.news/ticker/apple-foundation-models-report-2025/</guid><pubDate>Fri, 18 Jul 2025 01:37:00 +0200</pubDate><description>Apple posted its 2025 foundation models tech report to arXiv on July 17, 2025, giving the numbers behind the roughly 3-billion-parameter on-device model it shipped at WWDC25. The report&amp;amp;rsquo;s Table 3 puts the price of compressing that model to 2 bits per weight at 3.4 points of MMLU, from 67.8 in 16 bits down to 64.4, with IFEval falling from 85.1 to 82.3.
Apple reaches 2 bits through quantisation-aware training with a learnable scaling factor per weight tensor, and reports that a balanced set of four levels at -1.5, -0.5, 0.5 and 1.5 trains more smoothly than the usual -2, -1, 0 and 1. The embedding table is quantised to 4 bits and the key-value cache to 8 bits, and low-rank adapters are fine-tuned afterwards to claw back the quality the compression costs. The server model takes a different route, compressed after training to 3.56 bits per weight with Adaptive Scalable Texture Compression, a GPU texture format whose fixed-function decode hardware on Apple GPUs unpacks the weights at no cost to the compute cores.
The report also spells out the KV cache sharing the framework announcement only named. Apple splits the transformer so that Block 1 holds 62.5 percent of the layers and Block 2 the remaining 37.5 percent with its key and value projections removed, reusing Block 1&amp;amp;rsquo;s cache instead. That cuts KV cache memory by 37.5 percent, and because Block 2 produces no keys or values, prefill skips its computation entirely, which Apple says cuts time to first token by about the same 37.5 percent.
Apple rebuilt how the on-device model is trained. The company first trained a dense model on about 14T tokens, sparse-upcycled it into a 64-expert mixture-of-experts version, where only part of the model runs per token, using 1T high-quality tokens, then retrained the dense model on its last 10 percent of tokens against a distillation loss from that MoE teacher. Apple states this cut the cost of training the teacher by 90 percent and removed the need for structural pruning. The server model was trained on 8192 v5p Cloud TPU chips for 13.4T tokens.
For the language expansion the report states the models are designed to support 16 languages, built on a text tokenizer grown from a 100k to a 150k vocabulary and a continued pre-training mix whose multilingual share rose from 8 percent to 30 percent. In Apple&amp;amp;rsquo;s own human evaluation on US English text prompts, graders preferred the on-device model over Qwen-2.5-3B on 35.3 percent of prompts against 12.9 percent the other way, while against Gemma-3-4B the split was 21.0 percent to 21.9 percent. The report is on arXiv under a Creative Commons Attribution 4.0 license.
Chart: Apple&amp;amp;#39;s own human evaluation, from the tech report. Rows marked with an asterisk were tested against the compressed on-device model.
Source: https://arxiv.org/abs/2507.13575
Read the article: https://llmobile.news/ticker/apple-foundation-models-report-2025/</description><category>Apple Silicon</category><category>Quantisation</category><category>Distillation</category><category>Benchmarks</category></item><item><title>Apple opens its on-device model to all apps with the Foundation Models framework</title><link>https://llmobile.news/ticker/apple-foundation-models-framework/</link><guid isPermaLink="true">https://llmobile.news/ticker/apple-foundation-models-framework/</guid><pubDate>Mon, 09 Jun 2025 21:00:00 +0200</pubDate><description>Apple opened its on-device foundation model to third-party apps at WWDC25 on June 9, 2025. The Foundation Models framework gives any app direct access to the roughly 3-billion-parameter model, which runs offline and costs developers nothing per call.
Developers annotate Swift data structures with the @Generable macro, and the framework applies constrained decoding, which Apple calls guided generation, so the model returns those structures rather than free text. Tool calling lets the model invoke functions the app supplies. Developers who need more control can train rank 32 adapters with a Python toolkit.
Apple compressed the model to 2 bits per weight using quantisation-aware training. It splits the transformer into two blocks at a 5:3 depth ratio and shares the key-value cache between them, which Apple says cuts KV cache memory use by 37.5 percent. Images are handled by a 300M-parameter ViTDet-L vision encoder.
The model is designed to support 15 languages. The companion server model on Private Cloud Compute uses what Apple calls a Parallel-Track Mixture-of-Experts design, where independent transformer tracks cut how often the tracks have to synchronise with each other.
Diagram: Apple. The server model, not the on-device one.
Source: https://machinelearning.apple.com/research/apple-foundation-models-2025-updates
Read the article: https://llmobile.news/ticker/apple-foundation-models-framework/</description><category>iOS</category><category>Developer tools</category><category>Quantisation</category><category>Agents</category></item><item><title>Apple team finds H100 last on tokens per dollar for models up to 2B</title><link>https://llmobile.news/ticker/small-model-training-bottlenecks/</link><guid isPermaLink="true">https://llmobile.news/ticker/small-model-training-bottlenecks/</guid><pubDate>Fri, 25 Oct 2024 12:30:00 +0200</pubDate><description>Seven Apple researchers measured what it costs to train language models up to 2B parameters on rented Nvidia GPUs and found the most expensive card was the worst buy. The H100-80GB returned the fewest tokens per dollar of the three cards tested at every one of their four model sizes, 100M, 500M, 1B and 2B, while costing 2.53 times as much per hour as the A100-40GB.
The price scale comes from normalising published hourly rates at Lambda Labs and Google Cloud, with the A100-40GB set at 1.00 and the A100-80GB landing at 1.46. The same comparison shows cost efficiency falling as GPUs are added, with a single card the best or near-best point at all four model sizes. For the 1B model, tokens per dollar drop from about 17,000 on one GPU to about 10,300 on 64.
FlashAttention mattered more at small scale than at large. Reading the authors&amp;amp;rsquo; first figure, the 100M model gained roughly three times the tokens per dollar over vanilla attention, while at 2B the gain was under two times. The authors put this down to the cost of attention being quadratic in context length, so it dominates once the hidden dimension shrinks and the small models become bound by data movement between CPU and GPU and between GPUs. Vanilla attention also ran out of memory at a global batch size of 1024 on the 1B and 2B models, where FlashAttention did not.
Two habits carried over from large-model training did not pay off. Plain Distributed Data Parallel, which keeps a full copy of the model on every GPU, beat both sharded FSDP variants for the smaller models because it moves less data, and FSDP only pulled ahead at 2B, where it kept training at a global batch size that made DDP run out of memory. Filling the cards was not cost-optimal either, since the batch size figure peaks around a global batch size of 128 for the 100M model, 32 for 500M, 16 for 1B and 8 for 2B, in every case well short of what the memory allowed.
These are training runs on datacentre GPUs, not anything running on a phone, and the authors frame the work as guidance for labs with small budgets rather than for deployment. They name their own limits, running each setup at least three times for only 10 training steps, reporting normalised price ratios rather than absolute dollars, using no gradient accumulation, and leaving GPU utilisation and memory bandwidth measurements to future work. They also assume that settings such as the learning rate can be retuned to converge at whichever hardware configuration turns out cheapest. The 8-page paper is free to read on arXiv under the site&amp;amp;rsquo;s non-exclusive distribution licence.
Source: https://arxiv.org/abs/2410.19456
Read the article: https://llmobile.news/ticker/small-model-training-bottlenecks/</description><category>Research</category><category>Benchmarks</category><category>PyTorch</category></item><item><title>Apple's MobileCLIP-S0 encodes an image in 1.5 ms on an iPhone 12 Pro Max</title><link>https://llmobile.news/ticker/mobileclip/</link><guid isPermaLink="true">https://llmobile.news/ticker/mobileclip/</guid><pubDate>Fri, 21 Jun 2024 02:15:00 +0200</pubDate><description>Five Apple researchers presented MobileCLIP at CVPR 2024 on June 20, a family of image-text models built for phones. They exported each model with Core ML Tools 7.0 and timed it at batch size 1 on an iPhone 12 Pro Max running iOS 17.0.3, and their Table 7 gives the smallest variant, MobileCLIP-S0, 1.5 ms for the image encoder and 1.6 ms for the text encoder at 67.8% zero-shot accuracy on ImageNet. The same authors measure OpenAI&amp;amp;rsquo;s ViT-B/16 CLIP on that phone at 11.5 ms and 3.3 ms for 68.3%, and both models average 58.1 across the 38 datasets of the DataComp evaluation suite.
Multi-modal reinforced training, the method the paper introduces, moves the expensive part of distillation out of the training loop and into the dataset. Apple reinforces the DataComp dataset once, writing five synthetic captions per image with the CoCa model in OpenCLIP, storing the parameters of strong random image augmentations, 30 per image for the 12M subset and 10 for the billion-sample set, and storing the embeddings that two ViT-L/14 teacher models produce for those augmented images and for the real and synthetic captions. Training afterwards reads the captions and the teacher embeddings from disk, so the captioning model and the teachers never run again however many models are trained on the set.
The paper puts numbers on what that saves. One epoch over the 12.8M samples of the reinforced set on a single node of 8 A100-80GB GPUs takes 1.3 hours, the same as an epoch on the unreinforced data, against 21.1 hours when the captioning model and the teacher ensemble run during training, according to Table 4d. The cost moves to storage, and Table 4c puts DataComp-1B at 90 TB against 140 TB for the reinforced DataCompDR-1B, with the 12M subset going from 0.9 TB to 1.9 TB.
MobileCLIP comes in four sizes, and Table 7 lists image and text encoder parameters separately. S0 pairs an 11.4M-parameter image encoder with a 42.4M-parameter text encoder, S1 uses 21.5M and 63.4M at 2.5 ms and 3.3 ms, S2 uses 35.7M and 63.4M at 3.6 ms and 3.3 ms, and MobileCLIP-B keeps a standard ViT-B/16 image encoder at 86.3M and 10.4 ms. Zero-shot ImageNet accuracy across the four runs 67.8%, 72.6%, 74.4% and 76.8%, and a version of B trained on 39B seen samples instead of 13B reaches 77.2%. All three small variants replace the vision transformer with a convolution-transformer hybrid derived from Apple&amp;amp;rsquo;s FastViT, and S0 also swaps the text encoder for one that mixes 1-D convolutions with four self-attention layers, which the authors report as 5% smaller and 15.8% faster than the 12-layer transformer text encoder that S1 and S2 keep.
Apple released the code, the weights and the reinforced data at apple/ml-mobileclip, with checkpoints for all four variants and the longer-trained B, and with DataCompDR-1B and DataCompDR-12M on Hugging Face as synthetic captions, augmentation parameters and teacher embeddings keyed to DataComp images rather than as the images themselves. The repository lists the code under MIT, the models under Apple&amp;amp;rsquo;s own machine learning research model terms, and the data under CC BY-NC-ND. Apple showed zero-shot scene classification running in real time on an iPhone in the CVPR exhibit hall, and added an iOS app for the same task to the repository in November 2024.
Source: https://openaccess.thecvf.com/content/CVPR2024/html/Vasu_MobileCLIP_Fast_Image-Text_Models_through_Multi-Modal_Reinforced_Training_CVPR_2024_paper.html
Read the article: https://llmobile.news/ticker/mobileclip/</description><category>iPhone</category><category>Research</category><category>Benchmarks</category><category>Open weights</category></item><item><title>Apple Intelligence pairs a 3-billion-parameter on-device model with a server model</title><link>https://llmobile.news/ticker/apple-intelligence-foundation-models/</link><guid isPermaLink="true">https://llmobile.news/ticker/apple-intelligence-foundation-models/</guid><pubDate>Mon, 10 Jun 2024 21:00:00 +0200</pubDate><description>Apple introduced Apple Intelligence on June 10, 2024, built on a roughly 3-billion-parameter language model that runs on the device and a larger model that runs on Apple silicon servers through Private Cloud Compute. The on-device model uses a 49K token vocabulary, the server model 100K.
Apple compresses the on-device model with a mix of 2-bit and 4-bit weights, averaging 3.7 bits per weight, and says an option at 3.5 bits per weight costs no significant quality. On an iPhone 15 Pro the company reports a time to first token of about 0.6 milliseconds per prompt token and 30 tokens per second, both measured before token speculation.
Rather than ship a separate model per feature, Apple swaps LoRA adapters on top of the shared base model. The adapters are stored in 16 bits and, at rank 16, take tens of megabytes each. Apple says human graders preferred its models over comparable ones including GPT-3.5, GPT-4, Llama, Mistral and Gemma across its evaluations.
Diagram: Apple.
Source: https://machinelearning.apple.com/research/introducing-apple-foundation-models
Read the article: https://llmobile.news/ticker/apple-intelligence-foundation-models/</description><category>iPhone</category><category>iOS</category><category>Quantisation</category><category>Apple Silicon</category></item><item><title>ExecuTorch alpha runs Llama 2 7B on iPhone 15 Pro and Galaxy phones</title><link>https://llmobile.news/ticker/executorch-alpha/</link><guid isPermaLink="true">https://llmobile.news/ticker/executorch-alpha/</guid><pubDate>Tue, 30 Apr 2024 18:00:00 +0200</pubDate><description>PyTorch released ExecuTorch alpha on April 30, 2024, the first version of its edge runtime aimed at large language models rather than small vision and speech networks. PyTorch says the work enables Llama 2 7B to run on iPhone 15 Pro, iPhone 15 Pro Max and Samsung Galaxy S22, S23 and S24 phones, with early support for Llama 3 8B. The example project shipped with the release reports 8.15 to 11.6 tok/s for the quantised 7B model, measured by PyTorch over adb on a Galaxy S22, a Galaxy S24 and a OnePlus 12. The 0.1 preview from October 2023 had no such path.
Fitting a 7B model on a phone came down to quantisation. Every linear layer carries 4-bit weights in groups of 128 or 256 values, with activations quantised to 8-bit at runtime from their observed range. PyTorch measured WikiText perplexity at 9.16 for the full-precision model against 10.2 and 10.7 for the two group sizes, and notes that groups smaller than 128 were not enabled because the resulting files were still too large, since the embedding table and the weight scales remained in 32-bit floats. Support for 16-bit floats was under way.
Arm, Apple and Qualcomm Technologies worked on the release, and PyTorch describes the mobile accelerator paths through Core ML, MPS, TOSA and the Qualcomm AI Stack as ongoing work rather than finished. The release notes add Snapdragon 8 Gen 3 support and 4-bit and 16-bit quantisation on the Qualcomm side, over 100 operators on Apple&amp;amp;rsquo;s MPS backend, and a Vulkan delegate for mobile GPUs. PyTorch also credits Google for joint work on XNNPACK, the CPU library that carried the language model numbers, and says MediaTek was enabling the Llama models on its own chips.
PyTorch was explicit about what alpha did not yet do. The release notes mark XNNPACK as the only finished performance path for generative models and everything else as in progress, and the announcement calls on-device LLM support early while asking the community for help building GPU and NPU delegates and expanding the quantisation schemes. Meta was already shipping the runtime elsewhere, using it for hand tracking on Quest 3, for several models on the Ray-Ban Meta smart glasses and in a beginning rollout inside Instagram. The Python package became installable from PyPI with this release.
Source: https://pytorch.org/blog/executorch-alpha/
Read the article: https://llmobile.news/ticker/executorch-alpha/</description><category>PyTorch</category><category>Quantisation</category></item><item><title>Apple releases OpenELM at 270M to 3B, with parameters spread unevenly across layers</title><link>https://llmobile.news/ticker/openelm/</link><guid isPermaLink="true">https://llmobile.news/ticker/openelm/</guid><pubDate>Tue, 23 Apr 2024 01:12:00 +0200</pubDate><description>Apple posted OpenELM to arXiv on April 22, 2024, a family of four language models at 270M, 450M, 1.1B and 3B parameters, each also released in an instruction-tuned version. The name stands for Open-source Efficient Language Models, and the efficiency comes from layer-wise scaling, which spreads the parameter budget unevenly across the transformer&amp;amp;rsquo;s layers instead of giving every layer the same shape. Table 1 of the paper gives the 1.1B model an average of 45.93% across the OpenLLM leaderboard tasks against 43.57% for OLMo 1.2B, a gap the authors state as 2.36%, reached on 1.5T pre-training tokens against OLMo&amp;amp;rsquo;s 3.0T.
Layer-wise scaling varies the number of attention heads and the width of the feed-forward network from layer to layer, keeping layers near the input narrower and widening them towards the output, so a layer late in the stack holds more parameters than an early one. Apple trained the four variants for 350k iterations with its CoreNet library on public data only, drawing on RefinedWeb, a deduplicated PILE, a subset of RedPajama and a subset of Dolma v1.6, about 1.8T tokens in total. Text is filtered and tokenised on the fly rather than pre-tokenised, using the same tokeniser as Llama.
The paper benchmarks the models on an Apple MacBook Pro with an M2 Max chip and 64 GiB of RAM running macOS 14.4.1, through a port of OpenELM to MLX v0.10.0, and separately on a Linux workstation with an Intel i9-13900KF and an NVIDIA RTX 4090. On the MacBook in BFloat16, Apple measures the 270M model generating 212.40 tok/s and the 3B model 33.96 tok/s, with the 4-bit versions at 256.35 and 60.33 tok/s. On the CUDA machine OpenELM 1.08B is slower than OLMo 1.18B, at 92.15 tok/s against 203.40, which the authors attribute to a naive RMSNorm implementation and to OpenELM having 113 normalisation layers where OLMo has 33.
Apple presents the release as the complete framework rather than weights and inference code, covering the training and evaluation code for public datasets, the training logs, multiple checkpoints, the pre-training configurations and the code that converts a model to MLX for inference and fine-tuning on Apple devices, with pre-converted bf16 and 4-bit checkpoints for all four sizes. The eight models are on Hugging Face and the code is in CoreNet. Both the model cards and CoreNet name the terms as the Apple sample code license, which is Apple&amp;amp;rsquo;s own license and not one of the standard open source licenses.
Source: https://arxiv.org/abs/2404.14619
Read the article: https://llmobile.news/ticker/openelm/</description><category>Apple Silicon</category><category>Research</category><category>Open weights</category><category>Benchmarks</category></item><item><title>Apple researchers run models twice the size of available DRAM from flash</title><link>https://llmobile.news/ticker/llm-in-a-flash/</link><guid isPermaLink="true">https://llmobile.news/ticker/llm-in-a-flash/</guid><pubDate>Tue, 12 Dec 2023 12:00:00 +0100</pubDate><description>Apple researchers published LLM in a flash on December 12, 2023, a method for running language models that do not fit in a device&amp;amp;rsquo;s memory. By keeping parameters in flash storage and loading them into DRAM only when needed, the paper reports running models up to twice the size of the available DRAM.
Two techniques carry the result. Windowing reuses neurons that were already loaded for recent tokens, which reduces how much data has to move. Row-column bundling arranges the weights so that reads follow the sequential access patterns flash memory handles best.
Measured against naive loading from flash, the authors report a 4 to 5 times increase in inference speed on CPU and a 20 to 25 times increase on GPU. The paper was last revised in July 2024.
Chart: Alizadeh et al. Latency per token when only half the model fits in memory.
Source: https://arxiv.org/abs/2312.11514
Read the article: https://llmobile.news/ticker/llm-in-a-flash/</description><category>Memory</category><category>Research</category></item><item><title>Apple publishes MLX, where CPU and GPU share arrays without copies</title><link>https://llmobile.news/ticker/apple-mlx/</link><guid isPermaLink="true">https://llmobile.news/ticker/apple-mlx/</guid><pubDate>Tue, 05 Dec 2023 01:04:00 +0100</pubDate><description>Apple machine learning research shipped the first public release of MLX on December 5, 2023, an array framework for machine learning on Apple silicon. The repository carries a v0.0.2 tag dated that day and the same version went up on PyPI minutes later, after an initial commit on November 28. Arrays in MLX live in shared memory, so an operation can run on the CPU or on the GPU without the data being copied from one place to the other.
That behaviour comes from the hardware. Apple silicon uses a unified memory architecture in which the CPU and the GPU have direct access to the same memory pool, and the MLX documentation states that an array therefore has no device of its own. Code names the device when it runs an operation instead of moving an array to a device first, and when two operations on different devices depend on each other, the documentation says the MLX scheduler inserts the dependency between the streams automatically.
Computation is lazy. The documentation states that operations only record a compute graph and that nothing is computed until an eval call asks for a result, which is what lets MLX apply function transformations such as automatic differentiation and automatic vectorisation to the graph. It gives memory as the second reason, since a model whose weights are created as float32 and then replaced with float16 before any evaluation peaks at half the memory eager computation would need.
The MLX Swift bindings followed a week later, with an initial commit on December 12, 2023, and the package manifest lists macOS 14, iOS 17, tvOS 17 and visionOS 1 as supported platforms. The example apps build for iOS as well as macOS and include a chat client for language and vision-language models, a text generation demo that downloads weights from Hugging Face, and Stable Diffusion image generation. The LLMEval example uses the Increased Memory Limit entitlement on iOS, which its README attributes to the size of language model weights, and caps the MLX buffer cache at 20 MB.
MLX LM is the Python package that runs and fine-tunes language models on top of MLX, first published to PyPI on January 12, 2024. It installs the mlx_lm.generate and mlx_lm.chat command line tools, whose default model is a 4-bit quantised Llama 3.2 3B Instruct from the mlx-community organisation on Hugging Face, and it can quantise and upload converted models back to the Hub. MLX is published under the MIT license, and the install documentation lists Apple silicon, macOS 14.0 or newer and Python 3.10 or newer for the macOS package, alongside CUDA and CPU-only builds for Linux added later.
Source: https://github.com/ml-explore/mlx
Read the article: https://llmobile.news/ticker/apple-mlx/</description><category>Apple Silicon</category><category>Developer tools</category><category>Open source</category><category>iOS</category></item></channel></rss>