Strata runs Qwen 3.8 Flash Next 125B on a single RTX 4090 at over 100 tokens/sec via 2-bit quant
- Strata, a GitHub repo by Niko1221 with 10.3k stars, runs Qwen 3.8 Flash Next (125B) on one RTX 4090 at roughly 100 tokens/sec; a commenter reports 124 tok/s on a 4090 with 128GB DDR5 and a Ryzen 7950X3D, while a Ryzen 3600X with 48GB RAM and an RTX 3080 reaches 30 tok/s on the Coder build.
- The speed comes from aggressive quantization: the smallest Coder variant drops half the experts and runs at Q2, fits in 32GB of RAM, and the authors measure it at 91% of the full model's SWE-bench Verified score.
- Community benchmarks in bench/results report Q2_0 at 33 tok/s decode and about 600 tok/s prompt processing at 128K context on an RTX 2060 8GB, plus a 2x AMD Instinct MI50 (gfx906) run; one user gets 60 tok/s on IQ3_XXS in Strata against 21 tok/s in llama.cpp.
- Setup can be handed to an AI agent: docs/AI_SETUP.md and an MCP server let a coding assistant install Strata, which drew comparisons to piping curl into bash, and AMD support is still partial with the RX 9070 XT baseline left as a TODO placeholder.
- Critics argue the low-bit quants make the numbers hollow: a cited paper finds 4-bit quantization usually preserves performance while 2-bit causes broad degradation, and several commenters want speed figures published alongside accuracy benchmarks.
Hacker News opinions
Gave it a shot on my 4090 with 128GB DDR5 and a 7950X3D and I'm getting 124 tokens per second. Figured that was worth sharing here.
How does it compare to Qwen 3.8 27B? I really want to see the distilled ones with a harness up against the full MoE versions.
Why is that surprising? It's 2.5x faster than Anthropic's models, you get data sovereignty and privacy, and it's a strong model. Sounds like a best case scenario to me.
Coder build gives me 30 tok/s on a Ryzen 3600X with 48GB of RAM and a 3080. That's not a fast desktop, memory is around 2000MHz, and I still have Chromium, video streams and an agent running. Only change I made was setting thinking to low.
Which quantization are you using to hit those numbers?
Has anyone actually measured the effective intelligence of these quantized models? Publishing benchmarks with the quantized weights should be standard practice.
The README says the Coder version cuts half the experts, hits 91% of the full model's SWE-bench Verified, and fits in 32GB of RAM.
There's a recent paper on quantization degradation that found 4-bit usually preserves performance while 2-bit causes broad degradation. This repo uses 2-bit and strips experts for its fastest model, so make of that what you will.
Generation speed is the easy half for MoE offload. What does your prompt processing look like at 16K context?
They publish community benchmarks. Q2_0 does 33 tok/s decode and about 600 tok/s prompt processing at 128K on an RTX 2060 with 8GB of VRAM.
The setup instructions literally tell you to paste "set up Strata on this PC" into an agent and point it at docs/AI_SETUP.md. And I thought piping to bash was bad.
I've never understood the security complaint about curl foo | bash. You're already installing software from that same domain. If they wanted to do something nasty they'd do it in the software, not the setup script.
Been playing with this on a 3090 and it flies. Does a decent job on the PHP codebase security audits I've thrown at it.
Every one of these low-spec 100 tok/s projects is the same 2-bit quant with nothing else behind it, and conveniently none of them publish accuracy numbers. 4 bit is the floor.
Sure, quant it to Q2 and rip out half the experts and it goes fast. But you can't rely on that for long-horizon coding, which is where the reasoning lives. Get enough VRAM for Q4 or use a smaller model.
You can run IQ3_XXS, IQ3_S and IQ4_XS too. I switched to IQ3_XXS and get 60 tok/s on Strata versus 21 in llama.cpp, with better output.
Dwarfstar already supports this and I use the Q4 quant daily. Works really well for me.
I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford. Show me what runs on a Chromebook or an 8GB phone.
That card launched at $1600 MSRP. We went from needing supercomputers to needing high-end PCs to needing a $1600 GPU, the same path image rendering took.