Testing the Qwen3.8-Flash-Next model on a Strix Halo 128GB.

58.183.***.***
8
  • Engine used: llama.cpp branch pr27742

  • Model used: Qwen3.8-Flash-Next-UD-IQ4_XS

  • System used: Strix Halo (Beelink GTR9 PRO 128GB)

  • Operating system used: Arch Linux, Vulkan VRAM 96GB allocated

  • Reference: Link reference

After testing it a few times with the settings,

it seems that setting the ctx length to around 128k is probably optimal.

When set to around 64k, the response speed really comes close to that of a commercial service.

Setting it to around 256k and running the server will cause the system to be fully mobilized by consuming all CPU processing resources and system memory. Of course, the speed also decreases as the context length increases.

However, even in this case, it seems that it will run well enough if you have a little patience.

Looking at the benchmarking results, it definitely surpasses Deepseek 4 flash, but the speed is much faster, so the development speed is truly remarkable.

It's very exciting that we can get results with a local model that has a usable speed and public service quality.

Actually, I've been hoping for models like this to come out from places like Dokpaimo besides Chinese models. Dokpaimo focuses on creating cutting-edge models, and the companies involved don't seem to have the resources or motivation to put much effort into open weights for commercial reasons. It's a shame that they don't push local models alongside them for commercial reasons.

▶ Original source: https://www.reddit.com/r/StrixHalo/comments/1vz5yb3/qwen38flashnext_125ba6b_running_on_strix_halo/?solution=957fe3a6a547c072957fe3a6a547c072&js_challenge=1&token=7afd7253fec22262ff1c52b1703fe9ec7dbcc5482b6045fe1f50de322cb8d1eb&jsc_orig_r=

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.78 -0.01

2026.08.28 KEB 하나은행 고시회차 121회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!