Hardware
EXO Labs benchmarks put Llama 3.1 8B at 14 tok/s on an iPhone 15 Pro
EXO Labs published automated inference benchmarks from real devices, covering iPhone 15 Pro, Galaxy S24 Ultra, Mac mini M4 Pro clusters and desktop GPUs.
Apple team finds H100 last on tokens per dollar for models up to 2B
Seven Apple authors measured tokens per dollar for LLaMA-style models from 100M to 2B and found H100s last at every size, behind cheaper A100s.
PalmBench finds iPhones running local LLMs about three times faster than Android phones
A benchmark of quantised LLMs on eight phones and boards reports throughput, memory, power and heat, with iPhones ahead and 4-bit drawing more power than 3-bit.
Arm runs Llama 2 7B at 9.6 tokens per second on three mobile CPU cores
Arm demonstrated Llama 2 7B generating 9.6 tokens per second on three Cortex-A700 CPU cores, using int4 kernels it says beat llama.cpp by 20 percent.
Snapdragon 8 Gen 3 targets 10-billion-parameter models on device
Qualcomm says the new flagship runs generative models with up to 10 billion parameters on device and reaches up to 20 tokens per second for LLMs.
Qualcomm runs Stable Diffusion on an Android phone for the first time
Qualcomm AI Research generated 512x512 images in under 15 seconds on a Snapdragon 8 Gen 2 phone after quantising the model from FP32 to INT8.