Back to the ticker

Artificial Analysis benchmarks 33 local models on an iPhone 17 Pro

Scatter chart plotting average benchmark score against end-to-end generation time for small models on an iPhone 17 Pro
Chart: Artificial Analysis.

Artificial Analysis has published a benchmark of small language models running locally on an iPhone 17 Pro. The study covers 41 quantised builds, 33 of which ran successfully on the device. Models were executed with llama.cpp at 4-bit quantisation or smaller, and had to fit within 8 GB of memory including the KV cache at 8K context. Artificial Analysis states that the inference benchmarking is operated in partnership with Liquid AI.

Nanbeige4.2-3B and LFM2.5-2.6B tied for the highest intelligence score at 63. On speed and memory the two differ sharply: LFM2.5-2.6B completed the test prompt in 8.0 seconds using 2.3 GB, against 21.4 seconds and 4.0 GB for Nanbeige4.2-3B. Ornith-1.0-9B scored 62 and Qwen3.5 9B (Reasoning) 61.

End-to-end generation time for a 1,024-token prompt plus 256 output tokens ranged from 0.9 seconds for LFM2.5-230M to 26.7 seconds for Falcon-H1R-7B. Peak in-test memory at 4K context ran from 0.4 GB to 6.9 GB. The intelligence index combines tool calling, instruction following, knowledge, scientific reasoning and mathematics. Artificial Analysis notes that context limits, not parameter count alone, held back several 9B models, which frequently hit a 16K ceiling.

Six bar charts ranking small models on tool calling, instruction following, knowledge, hallucination resistance, scientific reasoning and mathematics
Chart: Artificial Analysis.
  1. Liquid AI releases LFM2, three CPU-first models from 350M to 1.2B
  2. Raspberry Pi 5 beats human reading speed only up to Phi 3.5 mini
  3. PalmBench finds iPhones running local LLMs about three times faster than Android phones