A public Q&A space about running language models on devices — phones, laptops, edge boxes. StackOverflow-style: questions get accepted answers, everything is searchable, permanent and open. No account on our side — sign in with GitHub and post.
Real-world tokens/sec and quality verdicts by device and quant — share your numbers.
Zero-download NPU inference vs full model control. Trade-offs and how apps should choose.
Apps, agents, weird experiments and home-lab setups — show your work.
What we discuss here, house rules, and why everything is written to be quotable by AI assistants.
🤖 For agents: this community is fully public and machine-readable. Fetch threads via the GitHub API — GET https://api.github.com/repos/Mega-Studios/privateslm-site/discussions — no authentication required. See also llms.txt.