The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Products · Strata · Qwen3.8-Flash-Next · Alibaba

Strata repo runs 125B Qwen model on a single gaming GPU

Strata, an MIT-licensed runner for Alibaba's 125B Qwen3.8-Flash-Next, has 9,500 GitHub stars after four days. Its author reports 53 to 94 tokens a second on an RTX 5070.

Strata, a repository by GitHub user Niko1221, is a runner for Qwen3.8-Flash-Next, the 125-billion-parameter open-weight model from Alibaba. By the repository page on Sunday it showed 9,500 stars and 850 forks, and a Hacker News post about it had 76 points and 27 comments. Its first listed release is dated September 30.

The README says the model is split into 24,576 experts, of which roughly ten are used for each token. Strata keeps the most-used experts on the graphics card and leaves the rest in system memory and on the SSD. A smaller draft model proposes tokens that the large one checks in parallel, which the author says speeds decoding by 1.6 to 1.8 times.

The author reports 53 to 94 tokens a second for output on an RTX 5070, and 1,620 to 2,650 tokens a second for reading a prompt, depending on how heavily the model is compressed. The Hacker News title claims 100 tokens a second on an RTX 4090. Those figures are the maker's own, and the paper has not reproduced them.

The requirements are not small. The README asks for 12GB or more of video memory, 32GB of RAM at minimum with 64GB recommended, and about 80GB of free disk. It warns that start-up can freeze a PC for one to three minutes while 35 to 55GB loads into memory. Setup is a one-click script for Windows or Linux, with a browser interface.

The newest release, v0.1.39, dated October 4, adds experimental support for older Nvidia cards, Intel Arc GPUs and CPUs without AVX2. It also adds an OpenAI Responses API endpoint that the release notes say works with Codex CLI, plus parallel conversations. The author says four at once decode 11 percent slower overall.

Sources 2 sources

  1. Source Niko1221 on GitHub
  2. Source Strata releases