search
Niko1221
Trends
- 1Project claims 125B model runs fast on RTX 4090●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s Article URL: https:// github.com/Niko1221/Strat
A GitHub project called Strata claims it can run the Qwen 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 GPU at roughly 100 tokens per second. The claim is drawing attention from developers on Hacker News, where it has collected 15 points and a handful of comments, with discussion focused on whether such speed and memory efficiency on consumer hardware is realistic.
Repos
- Niko1221/Strata Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anth