search
local AI models
Trends
- 1Qwen 125B model claims fast inference on a single RTX 4090●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A project called Strata claims to run Qwen's 125B-parameter Flash Next model on a single consumer RTX 4090 GPU at roughly 100 tokens per second. The claim drew heavy attention on Hacker News, where it ranked near the top of the site. If verified, it would represent a significant step for running large language models on everyday hardware rather than data-centre equipment.
- 2Turning off Apple Intelligence frees up disk space on macOS 27●Turn off Apple Intelligence on macOS 27 and get its disk space back
Mac users are discussing how to disable Apple Intelligence in macOS 27 and reclaim the disk space its models and caches occupy. A community script shared online automates the removal, and commenters are weighing whether the storage savings are worth losing the AI features, with some also raising privacy concerns about on-device models.
- 3Jevstiller distills Jev into a local model with disagreement bound●Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
A developer has launched Jevstiller, a tool that distills Jev into a local AI model, claiming a built-in disagreement bound that limits how far the distilled model can diverge from its source. The project was shared on Hacker News under the Show HN format, drawing moderate attention from the community. Details on the guarantee behind the disagreement bound are published on the project's site.
- 4Hacker turns iPhone into a second GPU for MacBook AI speedup●I made my iPhone a second GPU for my MacBook-Qwen 3.8 27B prefills 29–44% faster
A developer has shown how to use an iPhone as a second GPU alongside a MacBook, reporting that prefill times for running the Qwen 3.8 27B model are 29 to 44 percent faster. The experiment exploits Apple's unified connectivity to share compute between devices for local AI inference. Commenters are debating how the setup works, its practicality, and what it means for running large models on consumer hardware.
- 5Developer builds offline AI planner that picks your best hour outdoors●Say what you want to do outside. Gemma, on your own computer, reads it, code picks the hour from an open forecast, and y
A developer has built a small open-source tool that lets users type what they want to do outside; Google's Gemma model runs locally on the user's own computer, code picks the best hour from an open weather forecast, and the result comes back as a reminder and a pocket card that works without any signal. The project is being shared as part of an open-source development challenge.
Repos
- debpalash/VoiceStudio VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictati
- Ebony-Vinyl/dsh-our-free-model 在 dsh 里装上这个插件即可,无需登录、注册或填 API Key,就能使用包括 DeepSeek V4.1 Flash、Kimi K3 在内的前沿模型——完全免费,不限量。 All you do is install this plugi
- yetone/magpie Every agent's model. One place. Codex on DeepSeek, Claude Code on Kimi, from the menu bar.
- zouyuxuan122/dsh-our-free-model 在 dsh 里装上这个插件即可,无需登录、注册或填 API Key,就能使用包括 DeepSeek V4.1 Flash、Kimi K3 在内的前沿模型——完全免费,不限量。 All you do is install this plugi
- openclaw/openclaw The AI that really does things. Any OS. Any Platform. The lobster way. 🦞
- allenv0/SCM Deep AI search for every photo and every frame of video in any folder on macOS
- antirez/ds4 DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
- glanderness/BeefTV Local-first, lightweight, AI-native video workspace.
- Vibra-Ingenn/Janus Janus is a API router for AI models written in Go and has a Vulkan Model runner
- Rizzo-AI-Academy/rizzo-flow The open, local take on Jev: typed decisions from an LLM, without generating a single token
- tursomari/machtiani Empowering users.