Nvidia
8 updates on Nvidia.
Nemotron-Flash-1B decodes 1.9 times faster than Qwen3-0.6B on an H100
Nvidia designed a hybrid 1B and 3B model family around measured decoding latency on an H100 rather than parameter count, and released three checkpoints.
Intelligence per watt puts local model coverage at 88.7% of real queries
Stanford and Together AI measured 20+ local models on 1M real queries and propose accuracy per watt as the metric for local inference.
D2MoE picks a bit-width per token, 1.39 times the throughput at up to 53 percent less memory
A MobiCom 2025 paper routes every token to an expert and a bit-width, reporting up to 1.39 times the throughput of EdgeMoE at up to 53 percent less memory.
Raspberry Pi 5 beats human reading speed only up to Phi 3.5 mini
Helsinki and EURECOM researchers measured 11 models on a Raspberry Pi 5 and a Jetson Orin Nano, where memory and battery gave out before speed did.
MobiLLM fine-tunes OPT-1.3B in 4.5 GB on device by moving backpropagation to a server
Researchers fine-tuned OPT-1.3B on a Jetson Xavier NX in 4.5 GB by keeping the frozen model on the device and the trainable side network on a server.
Apple team finds H100 last on tokens per dollar for models up to 2B
Seven Apple authors measured tokens per dollar for LLaMA-style models from 100M to 2B and found H100s last at every size, behind cheaper A100s.
MeRino designs sub-100M language models that run 4.9 times faster than OPT-350M
Researchers design 52M to 64M parameter transformers by maximising entropy under a compute budget, matching OPT-350M accuracy on an NVIDIA Jetson Nano.
MobileVLM V2 runs 1.7B, 3B and 7B vision models on 144 image tokens
Meituan and Zhejiang University report a 1.7B vision language model at 64.2 on six benchmarks and 51.63 tok/s on an NVIDIA Jetson Orin.