Yhn SportFootball first seen 15 h ago, last 2 min ago, peak #1
Qwen 3.8 Flash Next 125B runs at 100 tokens/s on RTX 4090
Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A new open-source project called Strata claims it can run Alibaba's Qwen 3.8 Flash Next, a 125-billion-parameter model, on a single consumer RTX 4090 graphics card at roughly 100 tokens per second. If the benchmarks hold up, it would make a frontier-class large language model practical on gaming-grade hardware without a data centre, which is why developers are scrutinising the code and debating how the performance is achieved.
Why now: Running a 125B model on a single consumer GPU at high speed would be a major shift for local AI deployment.
Qwen 3.8 Flash NextStrataNVIDIA RTX 4090Alibaba
Rank over time, top of the chart is #1. 38 snapshots from 5 h ago to 2 min ago.
Evidence
API: https://socialmediatrends-api.osmike.com/v1/trends/988765