MikeTrendsTrends right now

Mmastodon TechnologyTechnology first seen 6 h ago, last 6 h ago, peak #1

Project claims 125B model runs fast on RTX 4090

Original: Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s Article URL: https:// github.com/Niko1221/Strat

A GitHub project called Strata claims it can run the Qwen 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 GPU at roughly 100 tokens per second. The claim is drawing attention from developers on Hacker News, where it has collected 15 points and a handful of comments, with discussion focused on whether such speed and memory efficiency on consumer hardware is realistic.

Why now: Running a 125B model at 100 tokens per second on a single consumer GPU would be a notable efficiency breakthrough, prompting skepticism and interest.

Qwen 3.8 Flash NextStrataNiko1221Nvidia RTX 4090Hacker News

Open on mastodon →

Evidence

API: https://socialmediatrends-api.osmike.com/v1/trends/1022056