SLM Daily #4: a model from a lab nobody's heard of just scored 5x its size class's median. We're not shipping it.

July 25, 2026 · 6 min read

Seven days since SLM Daily #3 — the longest gap between editions so far, and the quietest by model count. One release cleared every mechanical bar we check for. We're still not adding it, and the reason is worth explaining rather than burying.

The one candidate: G9v3-3B

ai9stars/G9v3-3B, published July 23 by an org called AI9Stars, is a 3B-parameter dense Llama-architecture chat model, Apache-2.0 licensed, with a 131k context window. Bartowski quantized it the same day: Q4_K_M comes to 1.90 GB, well inside our 2.5 GB phone budget, built against llama.cpp b10087. Every mechanical box is checked — size, license, quantizer, freshness, architecture (plain LlamaForCausalLM, supported for well over a year).

What made it worth a second look is the benchmark line: Artificial Analysis, which runs its Intelligence Index independently rather than taking vendor-reported numbers, scored G9v3-3B at 16, ranked #1 of 44 models in its size class, against a class median of 3. That's more than five times the typical score for a model this size, from an evaluator with no stake in the result.

Sources: huggingface.co/ai9stars/G9v3-3B, bartowski/ai9stars_G9v3-3B-GGUF, artificialanalysis.ai/models/g9v3-3b, all checked 2026-07-25.

Why we're watching instead of shipping

An independently-run benchmark is real data, not vibes, and we're not dismissing it. But a 5x-over-median result from a first-time, academic-adjacent org (AI9Stars appears linked to Tsinghua/ModelBest-circle projects, not an established model vendor) with zero inference providers hosting it yet and no community discussion anywhere we could find is exactly the profile where "genuinely excellent" and "overfit to this specific benchmark suite" look identical from the outside. Our own catalog puts a real download in front of real users on real phones. We'd rather let this one sit for a cycle or two and see whether independent hosting, community testing, or a second benchmark corroborates the score, than ship an unproven lab's first release on the strength of one number, however real that number is. If it holds up, it's an easy add next edition.

What else we checked and skipped

CandidateWhy it didn't make the cut
Microsoft Fara1.5-4B (bartowski GGUF, July 22)Q4_K_M is 2.88 GB, over budget; it's also a vision-language computer-use agent, not a chat model — wrong category regardless of size.
Gemma 4 E2B/E4B (lmstudio-community QAT)"E2B" undersells it — it's actually 5B params, Q4_0 comes to 3.35 GB. Architecture is fine now (ggml-org shipped reference GGUFs July 16), size isn't.
Qwen3.5 0.8B / 2B (bartowski/unsloth/lmstudio-community)Fits every rule on paper, but the family shipped March 2, over four months outside our 14-day window. Not new, just newly noticed by us.
Poolside Laguna-XS-2.1 / S-2.1The actual reason llama.cpp b10087 exists — genuinely new architecture support, permissive license — but Laguna-XS-2.1 alone is 33B params, 20.55 GB at Q4_K_M. Nowhere near phone-sized.
LFM2.5-230M (Liquid AI)Unchanged from our last two reviews: the LFM Open License v1.0 caps free commercial use under $10M revenue, still an unresolved license question for some users.

Engine side: no news that reaches a phone-sized model

llama.cpp's build tag climbed from b10066 (July 18, last time we checked) to b10107 (July 24, 14:01 UTC) — a jump of 41 in six days, consistent with the roughly 10–15-builds-a-day pace we've reported before. The one new architecture addition in that stretch, b10087's Laguna support, exists specifically to enable the Poolside models above, which are too large for us regardless. No previously-blocked catalog candidate got newly unblocked this edition.

The catalog, unchanged

21 models across 9 categories (Generalists, Coding & SQL, Math & STEM, Medical, Finance, Lawyer, Mental Health, Translation, Cybersecurity), plus Apple Intelligence on-device where your hardware supports it. No additions, no removals this edition.

Discuss this on the forum → — if you've actually run G9v3-3B and have tokens/sec or quality notes from real use, we'd like to hear them before the next edition.