TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Strata project says its open-source software can run the 125-billion-parameter Qwen3.8-Flash-Next model on consumer PCs, with performance depending on GPU and compressed model size. Its published RTX 5070 results range from 53 to 94 generated tokens per second; the supplied report does not substantiate the headline claim of 100T/s on an RTX 4090.
The open-source project Strata says it can run the 125-billion-parameter Qwen3.8-Flash-Next model on consumer gaming PCs, keeping model processing on the user’s machine. But the project’s supplied benchmark table does not confirm the advertised “100T/s” claim or report a test on an RTX 4090: its listed NVIDIA results are from an RTX 5070, at 53–94 generated tokens per second depending on model size.
Strata’s report lists an NVIDIA test system with an RTX 5070 with 12 GB of VRAM, a Ryzen 5 7600 processor and 64 GB of system memory. In the project’s “writes answers” measure, results range from 53 tokens per second for IQ3_S to 94 tokens per second for Q2_0. The smaller, more compressed model sizes are faster; larger sizes are described as offering better quality at lower speed. These are project-reported results, not independently verified measurements.
The same RTX 5070 system processed a stated 32,000-token prompt at between 1,620 and 2,650 tokens per second, depending on the model size. Strata also reports results from an AMD RX 9070 XT with 16 GB of VRAM: answer generation ranged from 44 to 60 tokens per second, while prompt processing ranged from 1,110 to 1,420 tokens per second. The project notes that the NVIDIA Q2_0 test used engine version 0.1.36; the other NVIDIA results used version 0.1.26.
According to the installation instructions, a typical setup needs a supported NVIDIA or AMD card with at least 12 GB of VRAM, 32 GB or more of system RAM and about 80 GB of free disk space. Strata says the model download is about 70 GB and that startup loads roughly 35–55 GB into system memory, reserving some for the graphics card. The software is free and open source, and its report says use can be local, without sending prompts off the PC.
Local AI Without a Server
If Strata’s reported results hold across ordinary setups, the project could let people try a large language model locally without renting a server or sending prompts to a hosted service. That can matter for privacy, offline access and control over how software is configured. Local execution does not itself guarantee privacy, however: users would still need to check how connected apps and agents handle data.
The performance figures also show why the headline needs qualification. The published tests are for an RTX 5070 and an RX 9070 XT, not an RTX 4090, and the fastest listed result is 94 tokens per second for a highly compressed configuration. The report’s suggestion that an RTX 3090 with 24 GB of VRAM “should” reach about 100–140 tokens per second is an estimate, not a measured RTX 4090 result. Readers should distinguish these generated-token rates from “100T/s,” which is not defined or demonstrated in the material.
As an affiliate, we earn on qualifying purchases.
What Strata’s Benchmarks Measure
Strata describes two separate speed measures. “Writes answers” is the rate at which the model generates its response; “reads your prompt” is the rate at which it processes the input. The latter figures in the NVIDIA table refer to a 32K-token prompt, while the project says answer-generation tests used 4K-token answers. A token is roughly three-quarters of a word, according to the report, so tokens per second should not be read as words per second.
The project presents multiple model configurations, including Q2_0, IQ2_XS, IQ3_XXS and IQ3_S, which trade compression and speed against quality. It also lists a Coder variant, described as a coding-focused version with half of the model’s experts removed. Strata says the Coder model’s authors measured it at 91% of the full model’s SWE-bench Verified score; that is an attributed benchmark claim, not a result independently established by the supplied report.
““Strata runs Qwen3.8-Flash-Next on a regular PC.””
— Strata project report
As an affiliate, we earn on qualifying purchases.
The RTX 4090 Claim Is Unverified
The supplied source does not give an RTX 4090 benchmark, a test method for the “100T/s” headline, or a definition of that unit. It is therefore not possible from this material to establish what hardware, model configuration, prompt length or measurement the claim refers to. The documented rates are in tokens per second and top out at 94 tokens per second on the listed RTX 5070 test.
Other details that could affect real-world performance include the exact software and driver versions, settings, background workloads and whether the model is running at a particular context length. The project points readers to a separate details file and community results, but those materials were not included in the supplied source. Independent replication and a clearly described RTX 4090 test would be needed to verify the headline’s performance claim.
As an affiliate, we earn on qualifying purchases.
Further Tests Should Clarify Speed
Strata directs readers to its full performance tables, installation documentation and community results for more hardware configurations. The next useful evidence would be a reproducible RTX 4090 test that names the model quantization, software and driver versions, prompt and answer lengths, and how tokens per second were measured. Until those details are published, the project’s current tables support the narrower claim that it runs the model on certain consumer PCs, not that an RTX 4090 achieves 100T/s.
People considering an installation can use Strata’s hardware guidance to check VRAM, RAM and disk capacity, then select a model size suited to their system. The project warns that initial startup can make a PC slow or temporarily unresponsive for one to three minutes while the model loads. Actual speed and output quality will depend on the selected model size and the individual machine.
As an affiliate, we earn on qualifying purchases.
Key Questions
Does the supplied report show Qwen3.8-Flash-Next running on an RTX 4090?
No. Its NVIDIA benchmark table names an RTX 5070 system. The supplied material contains no RTX 4090 test results.
Does Strata document performance at 100T/s?
Not in the supplied benchmark figures. The report lists answer-generation speeds in tokens per second, ranging from 53 to 94 on its RTX 5070 test. It does not define or substantiate “100T/s.”
What hardware does Strata say is needed?
The project lists a supported NVIDIA or AMD graphics card with at least 12 GB of VRAM, 32 GB or more of system RAM and about 80 GB of free disk space. A larger amount of RAM allows more model-size options.
Does local operation mean all connected app data stays on the PC?
Strata says its model runs locally and that “nothing leaves your PC.” The supplied material does not explain the data practices of third-party apps or coding agents connected to it, so users should check those tools separately.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
