Qwen
9 updates on Qwen.
RikkaHub Agent test: Android phone agent compiles whisper.cpp on its own
XDA runs RikkaHub Agent on an Oppo Find N5 against a self-hosted Qwen 3.6 27B. The agent installed dependencies and built whisper.cpp in Termux in seven minutes.
Tempo compresses hour-long video with a 2B vision model so a 4B LLM can answer
Meta AI and KAUST report 52.7 on LVBench for a 6B system in which a 2B vision-language model squeezes long video down to about 3 tokens per frame.
Alibaba adds 0.8B and 2B sizes to Qwen3.5, with 262K context and a vision encoder
Alibaba released Qwen3.5-0.8B and Qwen3.5-2B, dense vision-language models with a 262,144-token context, Apache 2.0 weights and 4-bit MNN builds.
Alibaba releases GUI-Owl-1.5 agent models from 2B to 32B under MIT
Tongyi Lab open-sourced six GUI agent checkpoints from 2B to 32B and reports 71.6 on AndroidWorld, with every benchmark run server-side, not on a phone.
Intelligence per watt puts local model coverage at 88.7% of real queries
Stanford and Together AI measured 20+ local models on 1M real queries and propose accuracy per watt as the metric for local inference.
ShadowNPU scores attention on the Snapdragon NPU and reports 4.5 times faster inference
Peking University and BUPT move attention token scoring onto a Snapdragon NPU and report up to 4.5 times faster inference using a single CPU core.
Alibaba gives Qwen3's 0.6B and 1.7B models a reasoning switch
Qwen3-0.6B and Qwen3-1.7B carry the family's switch between a reasoning mode and a fast mode, with 32K context and Apache 2.0 weights.
Alibaba MNN runs 4-bit LLMs on phone CPUs and GPUs, with a multimodal Android app
MNN-LLM converts PyTorch checkpoints into a 4-bit MNN format for phones, and Alibaba reports prefill 8.6 times faster than llama.cpp on an Android CPU.
Alibaba builds Qwen2's 0.5B and 1.5B sizes for phones, earphones and glasses
Alibaba built Qwen2-0.5B and Qwen2-1.5B for smartphones, earphones and smart glasses, with 32K context and Apache 2.0 weights.