Yhn SportFootball first seen 19 h ago, last 12 min ago, peak #1
Running Qwen 3.8 Flash Next on a single RTX 4090
Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A newly shared open-source project claims to run Alibaba's Qwen 3.8 Flash Next, a 125-billion-parameter model, on a single consumer RTX 4090 GPU at around 100 tokens per second. If the benchmarks hold up, it would make very large language models practical for hobbyists and local inference without datacenter hardware. Developers in the discussion are examining the approach and questioning the real-world performance figures.
Why now: Running a 125B model at high speed on consumer hardware would be a major breakthrough for local AI inference
QwenAlibabaNvidia RTX 4090StrataGitHub
Rank over time, top of the chart is #1. 11 snapshots from 1 h ago to 12 min ago.
Evidence
API: https://socialmediatrends-api.osmike.com/v1/trends/988765