Every model below downloads once and runs fully offline on your iPhone, iPad or Mac — checked daily, with the RAM it needs stated up front. Comparing apps? See privateSLM vs Private LLM and privateSLM vs PocketPal AI.
Smallest & fastest — replaces Llama 3.2 1B, which it beats on every published benchmark (GPQA, MMLU-Pro, IFEval, tool-use) at a smaller file size. 32K context, …
Strong for its size. Balanced everyday choice.
Very fluid conversation for a compact model.
Best all-round chat. Huge 128K context. 6 GB+ RAM.
Microsoft model that punches above its size. Recent phones.
Long-context (131K) general assistant with a Think/No-Think toggle. Tops Artificial Analysis's Intelligence Index for open-weight models ≤4B params (score 16 vs…
New looped-transformer architecture, chat + reasoning + tool-use template. IQ4_XS (not Q4_K_M — the default quant misses our size cap by 184 MB, this one doesn'…
Liquid AI's agentic/tool-calling generalist, 128K context. LFM Open License v1.0 (free commercial use under $10M revenue — same terms as our other LFM2 entries)…
Fast, accurate code & SQL. Great on-device autocomplete.
Benchmark leader for local coding. Refactors & finds bugs. 8 GB+.
Strong dedicated code model — completion, refactoring, many languages. 8 GB+.
Finance-domain model for reports, statements & market terms. 8 GB+.
Conversational finance assistant — earnings, filings, market questions. Not financial advice. 8 GB+.
Legal-domain specialist — contracts, clauses, statutes. Not legal advice. 8 GB+.
Conversational legal assistant for contracts & regulations. Not legal advice. 8 GB+.
Dedicated math engine — algebra, calculus, LaTeX output. 8 GB+.
Math-olympiad winner. Multi-step problem solving with tool-integrated reasoning. 8 GB+.
Fine-tuned on PubMed. General health & biomedical Q&A. 8 GB+.
Clinical model (EPFL) for medical Q&A and terminology. Not medical advice. 8 GB+.