MikeTrendsTrends right now

Mmastodon TechnologyAI first seen 9 h ago, last 9 h ago, peak #5

Qwen 3.8 Flash Next Runs at 100 T/s on One RTX 4090

Original: Qwen 3.8 Flash Next on a Single RTX 4090: How Consumer‑Grade GPUs Reach 100 T/s By Senior Editor – October 2026 “A singl

Reports circulating in tech circles claim that Alibaba's Qwen 3.8 Flash Next, a 125-billion-parameter model, can run at roughly 100 trillion tokens per second on a single consumer RTX 4090 GPU — a throughput previously associated with multi-node H100 clusters. Enthusiasts are discussing what this means for local AI inference and the collapsing cost barrier between consumer and data-center hardware.

Why now: Claimed near-data-center inference speeds on a single consumer GPU would be a major shift for local AI deployment.

QwenRTX 4090NVIDIAAlibaba

Open on mastodon →

Evidence

API: https://socialmediatrends-api.osmike.com/v1/trends/1079390