SLM Daily #8: a 2.6B model with a 128K context window, verified at 1.68 GB

August 6, 2026 · 5 min read

Liquid AI shipped LFM2.5-2.6B on August 4. bartowski's Q4_K_M quant landed the same week at a verified 1.68 GB — comfortably inside our 2.5 GB cap, five days old, and in the catalog as of today. That's the whole catalog news this edition; everything else this week either doesn't fit on a phone or doesn't exist as downloadable weights yet.

What shipped: LFM2.5-2.6B

Liquid AI's announcement and Hugging Face blog post describe LFM2.5-2.6B as a dense hybrid model built for agentic and tool-calling workloads, 128K context, pretrained on roughly 34 trillion tokens. Liquid's own numbers claim around 220 tokens/second on an M5 Max and roughly 30 tokens/second on a phone — we haven't independently reproduced either figure, so treat them as vendor-reported, not verified here.

What we did verify ourselves: bartowski's LiquidAI_LFM2.5-2.6B-GGUF Q4_K_M file returns a real HTTP HEAD Content-Length of 1,684,023,584 bytes — 1.68 GB, well under our cap. Architecture tag is lfm2, the same family as the LFM2/LFM2MoE support that landed in mainline llama.cpp back in July (see SLM Daily #5) — already-proven engine compatibility, not a fresh gate to clear. License is the LFM Open License v1.0, the same terms we reviewed and accepted for LFM2.5-230M back on July 18: free commercial use under $10M in annual revenue, a paid license above that. It's in the catalog now as lfm2.5-2.6b-q4.

Sources: liquid.ai/blog/lfm2-5-2-6b, huggingface.co/LiquidAI/LFM2.5-2.6B, bartowski/LiquidAI_LFM2.5-2.6B-GGUF, all checked 2026-08-06.

What we skipped

  • Qwen3.5-0.8B (bartowski/unsloth/lmstudio-community/mradermacher GGUFs) — shows up in every "new model" search this week because of upload-date metadata, but the underlying release is roughly five months old. Newly noticed by us, not newly released.
  • MiniCPM5-2B — teased publicly as a sub-4B leaderboard contender, but as of today OpenBMB's own model collection only lists MiniCPM5-1B (May 2026). The weights for the 2B version haven't actually been published. Worth checking again in a few days; nothing to evaluate yet.
  • Thinking Machines' "Inkling Small" — genuinely interesting (scores within a point of its own 3x-larger sibling on Artificial Analysis's Intelligence Index), but it's a 276B-total/12B-active MoE reasoning model. Not a phone candidate at any quant.

Engine: high cadence, nothing phone-relevant

llama.cpp's latest tag is b10290 (August 6, ~01:33 UTC). The merged work since our last check was almost entirely multi-token-prediction support for large models — DeepSeek V3.2, DeepSeek V4 Flash/MTP+DSpark, Qwen3-Next, GLM-4.7-Flash — plus MiniMax M3 stability fixes and a breaking change to llama-tts for Qwen3-TTS multimodal support. None of it touches a model small enough for our catalog. Nanbeige4.2-3b-q4's compatibility, confirmed in our last edition and re-verified again during this week's consolidation work, remains the most recent phone-relevant engine change.

Also this edition

Today's Edge in the Wild spotlight is the most extreme "no OS between the model and the hardware" project we've covered: a UEFI application that boots straight into a chat interface on a real Raspberry Pi 5, with nothing else running on the board at all. Read it here.

The catalog now

22 models, one more than last edition. Same honest caveat as every edition since SLM Daily #7: we're nominally holding ourselves to a cap of 8, and 14 of these are pre-cap specialist entries (medical, legal, finance and similar) we haven't unilaterally pulled from a live catalog. Flagging it again rather than letting it go quiet.

Discuss this on the forum →