Open weights

42 updates on Open weights.

  1. Edge0 releases MoE expert-offloading framework, demos 35B model on iPhone

    Edge0 streams mixture-of-experts weights from storage. Its 35B tier reports 2.9 GiB peak memory on a Mac mini M4 Pro; a launch post shows a 35B model on an iPhone.

  2. OpenBMB releases MiniCPM5-2B for local deployment

    A 2.5-billion-parameter dense model with a 131,072-token context, released under Apache-2.0 with GGUF, MLX and LiteRT-LM builds.

  3. Liquid AI releases LFM2.5-2.6B for on-device agents

    Liquid AI reports 30 tokens per second on a phone and under 2.5 GB of memory for its 2.6-billion-parameter model, released with open weights.

  4. Apertus Mini distils an open-data 8B into 0.5B, 1.5B and 4B models

    Swiss AI distilled its fully open Apertus 8B into 0.5B, 1.5B and 4B models on 1.7T tokens, with 3-bit to 6-bit MLX builds for Apple devices.

  5. Tencent open-sources a 440 MB offline translation model for phones

    Hy-MT1.5-1.8B-1.25bit compresses a 1.8-billion-parameter translation model from 3.3 GB to 440 MB and runs offline on a phone. Weights and an Android demo are public.

  6. Gemma 4 lands on the edge with Agent Skills in Google AI Edge Gallery

    Google says Gemma 4 E2B runs in under 1.5 GB on some devices and reaches 3,700 prefill tokens per second on a Qualcomm Dragonwing IQ8 NPU.

  7. Alibaba adds 0.8B and 2B sizes to Qwen3.5, with 262K context and a vision encoder

    Alibaba released Qwen3.5-0.8B and Qwen3.5-2B, dense vision-language models with a 262,144-token context, Apache 2.0 weights and 4-bit MNN builds.

  8. Alibaba releases GUI-Owl-1.5 agent models from 2B to 32B under MIT

    Tongyi Lab open-sourced six GUI agent checkpoints from 2B to 32B and reports 71.6 on AndroidWorld, with every benchmark run server-side, not on a phone.

  9. Nemotron-Flash-1B decodes 1.9 times faster than Qwen3-0.6B on an H100

    Nvidia designed a hybrid 1B and 3B model family around measured decoding latency on an H100 rather than parameter count, and released three checkpoints.

  10. Liquid AI releases LFM2, three CPU-first models from 350M to 1.2B

    Liquid AI released open-weight models of 350M, 700M and 1.2B parameters and reports 2x faster decode and prefill than Qwen3 on CPU.

  11. Hugging Face's SmolLM3 is a 3B model with 128k context and two reasoning modes

    Hugging Face released a 3B model with a 128k context window, six languages and a switchable reasoning mode, along with the full training recipe.

  12. Gemma 3n runs 5B and 8B models in 2 GB and 3 GB of memory

    Google previewed a mobile-first model whose Per-Layer Embeddings cut RAM use, and said the same architecture powers the next Gemini Nano.