AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Strata project says its open-source software can run the 125-billion-parameter Qwen3.8-Flash-Next model on consumer PCs, with performance depending on GPU and compressed model size. Its published RTX 5070 results range from 53 to 94 generated tokens per second; the supplied report does not substantiate the headline claim of 100T/s on an RTX 4090.

The open-source project Strata says it can run the 125-billion-parameter Qwen3.8-Flash-Next model on consumer gaming PCs, keeping model processing on the user’s machine. But the project’s supplied benchmark table does not confirm the advertised “100T/s” claim or report a test on an RTX 4090: its listed NVIDIA results are from an RTX 5070, at 53–94 generated tokens per second depending on model size.

Strata’s report lists an NVIDIA test system with an RTX 5070 with 12 GB of VRAM, a Ryzen 5 7600 processor and 64 GB of system memory. In the project’s “writes answers” measure, results range from 53 tokens per second for IQ3_S to 94 tokens per second for Q2_0. The smaller, more compressed model sizes are faster; larger sizes are described as offering better quality at lower speed. These are project-reported results, not independently verified measurements.

The same RTX 5070 system processed a stated 32,000-token prompt at between 1,620 and 2,650 tokens per second, depending on the model size. Strata also reports results from an AMD RX 9070 XT with 16 GB of VRAM: answer generation ranged from 44 to 60 tokens per second, while prompt processing ranged from 1,110 to 1,420 tokens per second. The project notes that the NVIDIA Q2_0 test used engine version 0.1.36; the other NVIDIA results used version 0.1.26.

According to the installation instructions, a typical setup needs a supported NVIDIA or AMD card with at least 12 GB of VRAM, 32 GB or more of system RAM and about 80 GB of free disk space. Strata says the model download is about 70 GB and that startup loads roughly 35–55 GB into system memory, reserving some for the graphics card. The software is free and open source, and its report says use can be local, without sending prompts off the PC.

At a glance
reportWhen: Current project report; publication dat…
The developmentStrata has published instructions and performance figures for running Qwen3.8-Flash-Next locally, but its supplied benchmarks do not show a 100T/s result or an RTX 4090 test.

Local AI Without a Server

If Strata’s reported results hold across ordinary setups, the project could let people try a large language model locally without renting a server or sending prompts to a hosted service. That can matter for privacy, offline access and control over how software is configured. Local execution does not itself guarantee privacy, however: users would still need to check how connected apps and agents handle data.

The performance figures also show why the headline needs qualification. The published tests are for an RTX 5070 and an RX 9070 XT, not an RTX 4090, and the fastest listed result is 94 tokens per second for a highly compressed configuration. The report’s suggestion that an RTX 3090 with 24 GB of VRAM “should” reach about 100–140 tokens per second is an estimate, not a measured RTX 4090 result. Readers should distinguish these generated-token rates from “100T/s,” which is not defined or demonstrated in the material.

Amazon

NVIDIA RTX 4090 graphics card

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Strata’s Benchmarks Measure

Strata describes two separate speed measures. “Writes answers” is the rate at which the model generates its response; “reads your prompt” is the rate at which it processes the input. The latter figures in the NVIDIA table refer to a 32K-token prompt, while the project says answer-generation tests used 4K-token answers. A token is roughly three-quarters of a word, according to the report, so tokens per second should not be read as words per second.

The project presents multiple model configurations, including Q2_0, IQ2_XS, IQ3_XXS and IQ3_S, which trade compression and speed against quality. It also lists a Coder variant, described as a coding-focused version with half of the model’s experts removed. Strata says the Coder model’s authors measured it at 91% of the full model’s SWE-bench Verified score; that is an attributed benchmark claim, not a result independently established by the supplied report.

““Strata runs Qwen3.8-Flash-Next on a regular PC.””

— Strata project report

Amazon

high VRAM gaming GPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The RTX 4090 Claim Is Unverified

The supplied source does not give an RTX 4090 benchmark, a test method for the “100T/s” headline, or a definition of that unit. It is therefore not possible from this material to establish what hardware, model configuration, prompt length or measurement the claim refers to. The documented rates are in tokens per second and top out at 94 tokens per second on the listed RTX 5070 test.

Other details that could affect real-world performance include the exact software and driver versions, settings, background workloads and whether the model is running at a particular context length. The project points readers to a separate details file and community results, but those materials were not included in the supplied source. Independent replication and a clearly described RTX 4090 test would be needed to verify the headline’s performance claim.

Amazon

large language model local setup

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Tests Should Clarify Speed

Strata directs readers to its full performance tables, installation documentation and community results for more hardware configurations. The next useful evidence would be a reproducible RTX 4090 test that names the model quantization, software and driver versions, prompt and answer lengths, and how tokens per second were measured. Until those details are published, the project’s current tables support the narrower claim that it runs the model on certain consumer PCs, not that an RTX 4090 achieves 100T/s.

People considering an installation can use Strata’s hardware guidance to check VRAM, RAM and disk capacity, then select a model size suited to their system. The project warns that initial startup can make a PC slow or temporarily unresponsive for one to three minutes while the model loads. Actual speed and output quality will depend on the selected model size and the individual machine.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Does the supplied report show Qwen3.8-Flash-Next running on an RTX 4090?

No. Its NVIDIA benchmark table names an RTX 5070 system. The supplied material contains no RTX 4090 test results.

Does Strata document performance at 100T/s?

Not in the supplied benchmark figures. The report lists answer-generation speeds in tokens per second, ranging from 53 to 94 on its RTX 5070 test. It does not define or substantiate “100T/s.”

What hardware does Strata say is needed?

The project lists a supported NVIDIA or AMD graphics card with at least 12 GB of VRAM, 32 GB or more of system RAM and about 80 GB of free disk space. A larger amount of RAM allows more model-size options.

Does local operation mean all connected app data stays on the PC?

Strata says its model runs locally and that “nothing leaves your PC.” The supplied material does not explain the data practices of third-party apps or coding agents connected to it, so users should check those tools separately.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Kimi K3: Open Frontier Intelligence

Kimi K3 introduces Open Frontier Intelligence, a new platform for real-time data analysis and performance metrics, marking a significant step in AI transparency.

What Does Thinking Machines’ Inkling Tell Us About AI’s Next Step?

Thinking Machines released Inkling, a 975B parameter open-weight model, highlighting transparency and new benchmarks in AI development.

Vertigo relief app

A new vertigo relief app is being tested for adults with BPPV, offering guided repositioning maneuvers and symptom tracking, with potential for clinic integration.

AI‑Powered Image Generation for Blogs

Suppose AI-powered image generation can revolutionize your blog—discover how it can elevate your visuals and why you should consider integrating it today.