search
AI inference hardware
Trends
- 1Running a 125B Qwen model fast on an RTX 4090●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A new open-source project, Strata, claims to run the Qwen 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 graphics card at roughly 100 tokens per second. The claim is drawing attention among AI enthusiasts because a model of that size is normally considered far too large for consumer hardware, suggesting new efficiency techniques for local inference.
Repos
- antirez/ds4 DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm