Strata Brings 125B-Parameter AI to Gaming PCs

AI-generated image · US National Wire
New open-source software allows users to run the Qwen3.8-Flash-Next model locally on NVIDIA and AMD hardware.
Strata, a free and open-source project hosted on GitHub, now enables users to run the Qwen3.8-Flash-Next AI model on standard gaming PCs, as Hacker News first reported. According to documentation from Hacker News, the 125-billion-parameter model—which typically requires server-grade hardware—can operate locally on Windows or Linux systems using NVIDIA or AMD graphics cards with at least 12 GB of VRAM.
Performance varies by hardware and model compression. Hacker News reporting indicates that an NVIDIA RTX 5070 (12 GB) paired with a Ryzen 5 7600 and 64 GB of RAM can achieve write speeds between 53 and 94 tokens per second, depending on the model size used. For users with higher-end hardware, the source notes that an RTX 3090 (24 GB) should reach speeds of approximately 100 to 140 tokens per second. AMD users with an RX 9070 XT (16 GB) and a Ryzen 9 3900X saw write speeds ranging from 44 to 60 tokens per second across different model configurations, such as the Coder and Q2_0 versions.
To function, Strata requires a minimum of 32 GB of RAM, though 64 GB is recommended to run all model sizes. The software supports various configurations, including a "Coder" version optimized for programming and an experimental "Unsloth UD-Q4_K_XL" version; the latter is the closest to the full model but operates slower—writing only 7-8.5 tokens/s on a 64 GB PC—because it reads most of its data from the SSD. The system is designed to keep all data on the local machine, supporting tasks such as coding, chatting, and image reading.

