MikeTrendsTrends right now

Yhn SportFootball first seen 15 h ago, last 2 min ago, peak #1

Qwen 3.8 Flash Next 125B runs at 100 tokens/s on RTX 4090

Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

A new open-source project called Strata claims it can run Alibaba's Qwen 3.8 Flash Next, a 125-billion-parameter model, on a single consumer RTX 4090 graphics card at roughly 100 tokens per second. If the benchmarks hold up, it would make a frontier-class large language model practical on gaming-grade hardware without a data centre, which is why developers are scrutinising the code and debating how the performance is achieved.

Why now: Running a 125B model on a single consumer GPU at high speed would be a major shift for local AI deployment.

Qwen 3.8 Flash NextStrataNVIDIA RTX 4090Alibaba

Open on hn →

Rank over time, top of the chart is #1. 38 snapshots from 5 h ago to 2 min ago.

Evidence

API: https://socialmediatrends-api.osmike.com/v1/trends/988765