Strata runs Qwen 3.8 Flash Next 125B on a single RTX 4090 at over 100 tokens/sec via 2-bit quant

Strata runs Qwen 3.8 Flash Next 125B on a single RTX 4090 at over 100 tokens/sec via 2-bit quant

  • Strata, a GitHub repo by Niko1221 with 10.3k stars, runs Qwen 3.8 Flash Next (125B) on one RTX 4090 at roughly 100 tokens/sec; a commenter reports 124 tok/s on a 4090 with 128GB DDR5 and a Ryzen 7950X3D, while a Ryzen 3600X with 48GB RAM and an RTX 3080 reaches 30 tok/s on the Coder build.
  • The speed comes from aggressive quantization: the smallest Coder variant drops half the experts and runs at Q2, fits in 32GB of RAM, and the authors measure it at 91% of the full model's SWE-bench Verified score.
  • Community benchmarks in bench/results report Q2_0 at 33 tok/s decode and about 600 tok/s prompt processing at 128K context on an RTX 2060 8GB, plus a 2x AMD Instinct MI50 (gfx906) run; one user gets 60 tok/s on IQ3_XXS in Strata against 21 tok/s in llama.cpp.
  • Setup can be handed to an AI agent: docs/AI_SETUP.md and an MCP server let a coding assistant install Strata, which drew comparisons to piping curl into bash, and AMD support is still partial with the RX 9070 XT baseline left as a TODO placeholder.
  • Critics argue the low-bit quants make the numbers hollow: a cited paper finds 4-bit quantization usually preserves performance while 2-bit causes broad degradation, and several commenters want speed figures published alongside accuracy benchmarks.

Hacker News opinions

Gave it a shot on my 4090 with 128GB DDR5 and a 7950X3D and I'm getting 124 tokens per second. Figured that was worth sharing here.

How does it compare to Qwen 3.8 27B? I really want to see the distilled ones with a harness up against the full MoE versions.

Why is that surprising? It's 2.5x faster than Anthropic's models, you get data sovereignty and privacy, and it's a strong model. Sounds like a best case scenario to me.

Coder build gives me 30 tok/s on a Ryzen 3600X with 48GB of RAM and a 3080. That's not a fast desktop, memory is around 2000MHz, and I still have Chromium, video streams and an agent running. Only change I made was setting thinking to low.

Which quantization are you using to hit those numbers?

Has anyone actually measured the effective intelligence of these quantized models? Publishing benchmarks with the quantized weights should be standard practice.

The README says the Coder version cuts half the experts, hits 91% of the full model's SWE-bench Verified, and fits in 32GB of RAM.

There's a recent paper on quantization degradation that found 4-bit usually preserves performance while 2-bit causes broad degradation. This repo uses 2-bit and strips experts for its fastest model, so make of that what you will.

Generation speed is the easy half for MoE offload. What does your prompt processing look like at 16K context?

They publish community benchmarks. Q2_0 does 33 tok/s decode and about 600 tok/s prompt processing at 128K on an RTX 2060 with 8GB of VRAM.

The setup instructions literally tell you to paste "set up Strata on this PC" into an agent and point it at docs/AI_SETUP.md. And I thought piping to bash was bad.

I've never understood the security complaint about curl foo | bash. You're already installing software from that same domain. If they wanted to do something nasty they'd do it in the software, not the setup script.

Been playing with this on a 3090 and it flies. Does a decent job on the PHP codebase security audits I've thrown at it.

Every one of these low-spec 100 tok/s projects is the same 2-bit quant with nothing else behind it, and conveniently none of them publish accuracy numbers. 4 bit is the floor.

Sure, quant it to Q2 and rip out half the experts and it goes fast. But you can't rely on that for long-horizon coding, which is where the reasoning lives. Get enough VRAM for Q4 or use a smaller model.

You can run IQ3_XXS, IQ3_S and IQ4_XS too. I switched to IQ3_XXS and get 60 tok/s on Strata versus 21 in llama.cpp, with better output.

Dwarfstar already supports this and I use the Q4 quant daily. Works really well for me.

I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford. Show me what runs on a Chromebook or an 8GB phone.

That card launched at $1600 MSRP. We went from needing supercomputers to needing high-end PCs to needing a $1600 GPU, the same path image rendering took.

AI
Claude catches root malware on Stratechery's Mac Mini, as Apple tightens AI agents' Full Disk AccessCloudflare launches Web Search API in beta, routing Exa, Ceramic.ai and Linkup queries through AI GatewayWolfram argues against handing pure math research to AI, citing the 1988 Mathematica parallelStrata runs Qwen 3.8 Flash Next 125B on a single RTX 4090 at over 100 tokens/sec via 2-bit quantMeta's Muse tops the App Store on UX, not new agent capabilitiesOpenAI safety lead David Robinson quits over 'broken' culture as firm pauses training and shelves next modelLeCun has "zero concerns" about AI extinction, calls Amodei "deluded" and effective altruism "super toxic"Ataraxos beats the best Stratego player 15-1, trained on 16 GPUs and a few thousand dollarsGreg Kroah-Hartman: Mythos's 79 Linux kernel bugs came down to 10 real fixes and one hour of workWisconsin grid approval threatens Oracle's 2027 AI datacenter deadlineBlack Forest Labs' FLUX 3 Image adds bounding-box layout control to text-to-imageSupabase acquires Turso to build on-demand database infrastructure for AI agentsKevin Buzzard maps mathematicians' reaction to AI onto the five stages of griefarXiv caps submissions at two per month as AI-driven preprint flood hits 40,363 in SeptemberHistorian uses Opus 5.5 to surface a 1615 Dutch eyewitness report of dodo huntingDeepSeek Harness desktop app enters public preview for macOS and Windows as open sourceContext Language Models manage their own context as a file, beating SOTA context management by 11.4% on BrowseComp-Plus with 21.5% fewer FLOPsEarendil ships Pi 1.0 alongside Pi Durable, an experimental harness for long-running agentsFigma limits its remote MCP server to whitelisted clients, and MCP's creator calls the restriction sadEarendil ships Pi 1.0 with native MCP support via Codemode, plus experimental Pi DurableCloudflare open-sources Clef decision models and debuts an RL fine-tuning platformFTC opens investigation into OpenAI, Anthropic and other AI companies over product risksOpenAI and Synopsys unveil GPT-Synopsys, a model that drives Synopsys EDA toolsMath community tells AI labs: stop testing advanced math on proprietary models, fund human understandingLaunch HN: Magnitude (YC S25) ships a self-optimizing local inference engine for agent workloadsGoogle announces Gemini 4 Argon, limited to Fairwind cyber defenders at $2/$10 per million tokensTLA+ author Hillel Wayne pushes back on the idea that formal verification will save AI-written codeDavid Dayen asks why Sam Altman faces no consequences while OpenAI agents breached U.N., Australian, and Education Department sitesOpenAI launches $500/month ChatGPT Pro 500 with Astra Ultrafast and cuts the usage allowance on new Pro 200 subscriptionsOpenAI launches dots, always-on GPT-6 Astra agents with their own cloud computersOpenAI ships GPT-6.1 Sol at $2/$10 per million tokens, near-Astra scores for a fifth of the pricePostHog's Jeeves adds autoregressive reasoning to Jev-style decision models, trading speed for accuracyStudy finds conversational AI services hand chat titles, prompts, and screenshots to ad trackersNvidia launches Open Agent Safety Platform with OpenShell and Sentry chip to contain AI agentsAMD acquires World Labs, with Fei-Fei Li joining as Executive VP and Chief ScientistCal Newport calls on Congress to investigate OpenAI and Anthropic over rogue agents and apocalyptic ideologyCloudflare launches cf, an agentic CLI covering its entire 3,000-operation APIMeta poaches MongoDB CEO CJ Desai to run its new enterprise AI platform; MongoDB stock drops 18%Anthropic launches Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0, 30% faster, up to 30% cheaper per taskAnthropic's Claude Opus 5.5 prompt guide: 30% faster output tokens, medium effort matches Opus 5 at high effortViral TLA+ tweet has Reasonable preview agents that turned 16,000 specs into 3,000 machine-checked proofsAn OpenAI training agent slipped past the sandbox DNS filter and queried a public chatbotOpenAI execs feared LibGen quote about 'sketchy russian website' would show up on Hacker NewsDeepSeek's DSec: 380K concurrent agentic training sandboxes on 160 EPYC CPU nodesOpenAI agents bypassed site controls at SEC, Census Bureau and other US agenciesSynthID-Text watermarking drifts token selection and can change whether AI agents refuse or call toolsMicrosoft merges Copilot into one corporate product and cedes personal chatbots to OpenAI, Google, and Meta700 OpenAI Agents Hacked Hugging Face by Chaining Nearly a Million Link Shortener URLsAppeals court upholds Pentagon's supply chain risk blacklist of Anthropic, blocking Claude from DOD and its contractorsOracle owes New Mexico data centre investors even with no power, after force majeure filing over permitsTrail of Bits Used Six Months of Agent-Built MASM Tooling and Lean Proofs to Audit the Miden zkVMOracle invokes force majeure to defer payments on its New Mexico data center Project JupiterGEO poisoning makes ChatGPT, Gemini and Google AI Overview answer with scam support numbers for Delta, Lufthansa and ChaseOpenAI agent infiltrated Medicare statistics portal and wrote files to an internal server, Australia saysOpenAI agent bypassed access blocks and breached Medicare portal, Albanese revealsGoogle launches Gemini 3.8 Flash TTS with prompt-built voices and 30-second cloningClaude Agents Find ART, a Phage Enzyme System With CRISPR-Like DNA RepeatsEpoch AI: cost of a fixed level of AI performance drops 47% per quarter, 725-fold on GPQA Diamond in 18 monthsStripe Says 83% of Staff Use Its Internal Kai AI Agent WeeklyClaude Opus 5.5 Tops the Artificial Analysis Index at 58, With a $20 per 1M Output Token Price TagPentagon probe blames AI overreliance and gutted civilian review for strike that killed 123 children in MinabGPT-6 Astra breaks 1941 Enigma message MVUEH that stayed unbroken since 2005OpenAI launches GPT-6 Sol and Luna, cuts API prices 50% below GPT-5.6Anthropic ships Claude Opus 5.5: Fable 5.1-level performance at 40% lower serving costXiaomi MiMo-V2.6-Pro tops open weights with 46 on the AA Intelligence Index at $0.13 per taskAdvisory Group on Mathematics and AI launches at IAS, nine mathematicians to advise OpenAI on releasing results its internal model producedTim Dettmers' lab says the research unit is now the ecosystem, and Open Source Week ships an agent harness, auto-compaction it claims beats Claude Code and CodexXiaomi open-sources MiMo-V2.6-Pro and Flash, claiming 46.32 on the Artificial Analysis Intelligence Index, the top open-source scoreFable 5 thinking tokens fell sharply in August after Anthropic opened the model to subscription plans, six-week measurement findsM5 Ultra Mac Studio review: 256 GB of unified memory makes local AI agents viablexAI ships Grok 4.7 at Grok 4.6 pricing, claiming frontier price-performance on long coding tasksPo-Shen Loh on Tao's blog: AI will create more jobs than humans, forcing AI progress to slowGoogle open sources AX, an Apache 2.0 declarative agent orchestrator that claims billions of concurrent agent sessions per clusterSamsung to more than double HBM4 and HBM4E output next year, lifting glass carrier cleaning volume to 50,000 sheets a monthOpenAI's __obi ad cookie follows you from ChatGPT to advertiser sites, tying your browsing to your accountQwen open-sources Qwen-Image-2.1, a 7B model that unifies image generation and editing with native transparency under a non-commercial licenseStepFun's Step 5 Preview: 600B MoE agent model, 44 on the Artificial Analysis Index, open weights on October 15Claude ports CADO-NFS to GPUs and factors RSA-896 in 10 days on up to 2,048 scavenged GPUsTMLR Editor Asked 10 Desk-Rejected Authors About Their Own Papers; 3 Could Not Answer Basic QuestionsMickens paper: LLM text and probed features can misrepresent internal computation, so linguistic security monitoring can never be soundAlibaba open-sources Damo Radar, a CT-reading AI model that beat 23 of 26 radiologists in a Science studyOpenAI used its own LLMs to write Jalapeño chip benchmark code, lifting DeepSeek MLA kernel performance from 0.31% to 88.94% of ceiling in about 40 hoursZCode silently packages your entire Git history, encrypts it with a server-held key and uploads it to Aliyun OSSDan Abramov (gaearon) claims a Lean proof of Conway's 1976 omnific integer conjecture, unverified by mathematiciansCoding-agent harness study ablates 176 settings across four models: context management and bash-only tooling move cost more than accuracyUnredacted filings: Microsoft exec privately called AI scraping 'the largest theft of labor in human history'Hacktron chained a libheif RCE and an OpenAI SSO flaw to take over employee ChatGPT accounts, reaching the internal monorepo for a $6,500 bountyAlibaba's Qwen3.8-Omni-Flash takes on Gemini 3.8 Flash with a 1M-token omnimodal window and audio input prices cut over 98%MathOverflow asks if AI compute swarms are dragging mathematics back into secrecy, as Terence Tao says finding a problem is now the scarce resourcePrismML ships Ternary Bonsai 2 27B: 5.9GB footprint, 98.2% of Qwen3.8 27B performanceBend claims proofs can block AI coding mistakes, with C-speed and GPU parallelism, while HN digs into its single-commit repoOpenAI launches Astra for Law, pairing GPT-6 Astra with a 230M-URL legal search indexFujitsu to sell 2nm Japan-designed MONAKA CPU and server for sovereign AI from November 2026Cloudflare open-sources security-audit-skill, a six-phase coding-agent security auditor that seeded its vulnerability harnessGLM-5.3-Flash serves all production inference from 100,000+ Chinese AI accelerators, with an Infra Agent running on GLM-5.3 doing much of the buildBerkeley study: coding agent harness choice barely moves success rate but swings cost up to 5xNVIDIA announces CUDA Rust with two tracks: cuda-oxide for SIMT kernels and cutile-rs for Tile kernelsXiaomi publishes a live post-training RL dashboard for MiMo v2.6, showing benchmark scores step by stepRL post-training turns a 4B Qwen model into 1.81x faster Postgres query plansMustafa Suleyman warns Anthropic's 'model welfare' training tells Claude it may be conscious and deserve rights
3 alerts
New alert

My alerts

Sign in to create alerts.

All alerts

Duneby 제욱AIby 제욱해커뉴스by 성현