Engine used: llama.cpp branch pr27742
Model used: Qwen3.8-Flash-Next-UD-IQ4_XS
System used: Strix Halo (Beelink GTR9 PRO 128GB)
Operating system used: Arch Linux, Vulkan VRAM 96GB allocated
Reference: Link reference
After testing it a few times with the settings,
it seems that setting the ctx length to around 128k is probably optimal.
When set to around 64k, the response speed really comes close to that of a commercial service.
Setting it to around 256k and running the server will cause the system to be fully mobilized by consuming all CPU processing resources and system memory. Of course, the speed also decreases as the context length increases.
However, even in this case, it seems that it will run well enough if you have a little patience.
Looking at the benchmarking results, it definitely surpasses Deepseek 4 flash, but the speed is much faster, so the development speed is truly remarkable.
It's very exciting that we can get results with a local model that has a usable speed and public service quality.
Actually, I've been hoping for models like this to come out from places like Dokpaimo besides Chinese models. Dokpaimo focuses on creating cutting-edge models, and the companies involved don't seem to have the resources or motivation to put much effort into open weights for commercial reasons. It's a shame that they don't push local models alongside them for commercial reasons.
▶ Original source: https://www.reddit.com/r/StrixHalo/comments/1vz5yb3/qwen38flashnext_125ba6b_running_on_strix_halo/?solution=957fe3a6a547c072957fe3a6a547c072&js_challenge=1&token=7afd7253fec22262ff1c52b1703fe9ec7dbcc5482b6045fe1f50de322cb8d1eb&jsc_orig_r=