⬢github C · 23.4K ★ +115 since we first saw it · pushed 14 d ago · MIT
antirez/ds4
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
DwarfStar (antirez/ds4) is a small, self-contained C inference engine for running a handful of specific open-weight LLMs—DeepSeek V4 Flash/PRO, GLM 5.x, Qwen3.8—on consumer hardware like 96GB+ Macs, DGX Spark, and Strix Halo. It supports Metal, CUDA (including multi-GPU setups like 8x L40S), and ROCm, with SSD streaming for machines short on RAM, tensor/pipeline parallelism, an HTTP server, and its own GGUF files.
Why now: It's gaining attention as antirez's specialized alternative to llama.cpp for a few curated models, with an open 'AI full disclosure' about heavy AI-agent-assisted development sparking discussion.
Who it is for: Hobbyists and developers with high-RAM consumer machines who want fast local inference of specific frontier open-weight models.
Stars over our 29 snapshots: 23.3K to 23.4K, since 7 h ago.
Where people talked about it
- ⬢github antirez/ds4 just now
API: https://socialmediatrends-api.osmike.com/v1/repos/antirez/ds4