search
local AI models
Trends
- 1Running Qwen 3.8 Flash Next on a single RTX 4090●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A newly shared open-source project claims to run Alibaba's Qwen 3.8 Flash Next, a 125-billion-parameter model, on a single consumer RTX 4090 GPU at around 100 tokens per second. If the benchmarks hold up, it would make very large language models practical for hobbyists and local inference without datacenter hardware. Developers in the discussion are examining the approach and questioning the real-world performance figures.
- 2Janus tool runs GGUF models on any GPU via Vulkan●Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia
A new open-source project called Janus has been released on GitHub, offering a single Go binary that runs GGUF-format language models through Vulkan graphics drivers. It works across AMD, Intel and Nvidia hardware without needing CUDA or vendor-specific toolchains, and it drew early attention and discussion among developers on Hacker News.
- 3
Salvatore Sanfilippo, the creator of Redis known as antirez, has released ds4, a local inference engine for DeepSeek 4 Flash and PRO models. The project, written in C, supports Metal, CUDA and ROCm, meaning it runs on Apple, Nvidia and AMD hardware. It is drawing attention as a lightweight option for running the Chinese models entirely on local machines.
- 4Developer turns iPhone into second GPU for MacBook AI speedups●I made my iPhone a second GPU for my MacBook-Qwen 3.8 27B prefills 29–44% faster
A developer has shared a method of using an iPhone as a secondary GPU for a MacBook, reporting that Qwen 3 27B model prefill speeds improve by 29–44%. The trick routes the phone's hardware alongside the laptop's chip when running local language models. The claim has drawn attention from people experimenting with local AI setups and squeezing more performance out of consumer hardware.
- 5Redis creator launches ds4 for running LLMs locally●From the creator of Redis; run LLM locally with ds4
Salvatore Sanfilippo, the creator of Redis, has released ds4, a tool for running large language models locally on a computer. The project, hosted at dwarfstar.sh, is drawing attention on Hacker News, where it has attracted several hundred upvotes. Commenters are discussing what the database pioneer's move into local AI tooling could mean for the space.
- 6
An open-source tool has been highlighted for allowing users to run artificial intelligence models directly on their own machines, without relying on cloud services. Coverage in the open-source software community points to growing interest in local AI for privacy, cost savings, and independence from major providers. Enthusiasts say the approach puts control of data and computing back in users' hands.
- 7Free tool helps estimate GPU memory needed to run AI locally●Quanta memoria serve per far girare un'IA in locale? Uno strumento gratuito per scegliere il server GPU https:// diggita
A new free tool aims to help users work out how much GPU memory is required to run artificial intelligence models on local hardware, particularly when choosing a GPU server. It is drawing interest among hobbyists and professionals who want to self-host AI models instead of relying on cloud services, a topic of growing debate as local AI deployments become more practical.
- 8India seeks deeper access to Anthropic's advanced AI models▼India wants deeper access to Anthropic’s advanced models
India is pressing for expanded access to Anthropic's advanced AI models, according to The Banker. The report highlights growing tensions between national ambitions to build AI capacity and the controlled release policies of leading US artificial intelligence firms. New Delhi's interest reflects its broader push to secure cutting-edge AI tools for domestic development, as governments worldwide negotiate with a small group of frontier model providers over access, pricing and local deployment.
- 9
A developer has released a Google Maps Scraper MCP server, a tool that lets AI models and applications pull business data such as names, addresses, reviews and contact details directly from Google Maps. The launch drew attention on Hacker News, where users are weighing its usefulness for lead generation and local data projects against questions about scraping terms of service and Google's restrictions.
Repos
- Ebony-Vinyl/dsh-our-free-model 在 dsh 里装上这个插件即可,无需登录、注册或填 API Key,就能使用包括 DeepSeek V4.1 Flash、Kimi K3 在内的前沿模型——完全免费,不限量。 All you do is install this plugi
- zouyuxuan122/dsh-our-free-model 在 dsh 里装上这个插件即可,无需登录、注册或填 API Key,就能使用包括 DeepSeek V4.1 Flash、Kimi K3 在内的前沿模型——完全免费,不限量。 All you do is install this plugi
- debpalash/VoiceStudio VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictati
- yetone/magpie Every agent's model. One place. Codex on DeepSeek, Claude Code on Kimi, from the menu bar.
- allenv0/SCM Deep AI search for every photo and every frame of video in any folder on macOS
- openclaw/openclaw The AI that really does things. Any OS. Any Platform. The lobster way. 🦞
- antirez/ds4 DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
- glanderness/BeefTV Local-first, lightweight, AI-native video workspace.
- Rizzo-AI-Academy/rizzo-flow The open, local take on Jev: typed decisions from an LLM, without generating a single token
- Vibra-Ingenn/Janus Janus is a API router for AI models written in Go and has a Vulkan Model runner
- LockedinLabs-AI/agent-console Local-first observability for AI coding agents. Every Claude Code and Codex session's tokens, cache, models and cos
- tursomari/machtiani Empowering users.