Apple
21 updates on Apple.
Hugging Face adds native GGUF inference to Transformers, starting with Qwen3.5 on Macs
Transformers now loads GGUF models in packed form and runs them on Apple Silicon GPUs via Metal kernels, starting with Qwen3.5.
Apple launches M6 Mac mini and M5 Ultra Mac Studio for local AI
Apple’s new Mac Studio scales to 512GB of unified memory, which Apple says lets researchers run large language models entirely on device.
Prism ML releases Bonsai 2 27B for iPhone, iPad and Mac, with a 3.9 GB 1-bit version
Bonsai 2 27B runs on iPhone, iPad and Mac via MLX, and a 1-bit version needs 3.9 GB for the iPhone 17 Pro, according to Prism ML.
iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine
Apple says the A20 Pro carries 32 Neural Engine cores in total, double the AI processing power of A19 Pro, with 50 percent more memory bandwidth.
Desert Ant Labs releases 18 on-device models for speech, audio and privacy tasks
Desert Ant Labs released 18 single-task models that run locally on phones and laptops and are free up to 100k monthly active devices.
Ornith releases Ornith-1.5, a 9B model with a mobile build for iPhone and Android
Ornith-1.5 comes in 9B, 35B and 397B sizes, and the 9B model scores 70.6 on SWE-bench Verified and has a mobile build for iPhone and Android.
Prism ML releases Bonsai 27B, a 3.9 GB 1-bit model that fits an iPhone 17 Pro
Prism ML’s 1-bit Bonsai 27B compresses Qwen3.6 27B to 3.9 GB and runs on iPhone, iPad, Mac and NVIDIA GPUs under Apache 2.0.
Prism ML releases Bonsai Image 4B, a 1-bit image model that runs on iPhone
Bonsai Image 4B generates a 512x512 image in 9.4 seconds on an iPhone 17 Pro Max, and its 1-bit diffusion transformer needs 0.93 GB.
Prism ML releases 1-bit Bonsai 8B, an 8.2B-parameter model that needs 1.15 GB
1-bit Bonsai 8B stores all weights in one bit, needs 1.15 GB and runs at about 40 tok/s on an iPhone 17 Pro, according to Prism ML.
Apple ships Python bindings for the on-device Foundation Models framework
The apple-fm-sdk package calls the on-device Apple Intelligence model from Python on macOS 26, for scripting and batch evaluation outside Swift.
A19 Pro puts Neural Accelerators in every GPU core
Apple says the iPhone 17 Pro chip pairs Neural Accelerators in each of six GPU cores with a 16-core Neural Engine to run large local language models.
Apple puts the cost of 2-bit compression at 3.4 MMLU points
Apple measures its on-device model at 67.8 MMLU in 16 bits and 64.4 after compression to 2 bits per weight, and details the distillation pipeline behind it.