Simple review of QWEN3.8 27B

170

I'll briefly talk about my current system and loading conditions to start.

  • T7910 used + dual NVIDIA GV100 (32GB VRAM per card) - NVLink is flying around on eBay, but I'm crying because I found results saying that the speed drops critically when combined with NVLink + MTP.

  • Model Q8, KV cache FP16

  • Parameters are set to the recommended configuration (excluding context)

  • 2 slots, 200K context per slot (configured this way, GV100 consumes 31.5GB)

I'm using an HP Strix Halo Mini Workstation as a server and previously used and sold an Asus GX10.

Please note that these are simple test results I conducted yesterday after loading, not controlled environment benchmarks.

  • Speed varies depending on the MTP hit rate, ranging from around 30 tps to 50 tps (MTP=4 set)

  • Inference tokens are still spewing out tremendously. However, I believe this is the driving force behind achieving these results with a 27B parameter model, so I don't disable or limit inference separately.

  • I originally used it for OpenCloze, not for single calls, but this time I wanted to see how aware it was of its surroundings and how it reacted when the context got longer. So I tried using it in OpenCloze.

  • Subjectively, I feel that the tool calling ability has significantly improved. I observed QWEN3.6 often failing or unable to find tools, while QWEN3.8 found and used them well.

  • When asked "What model are you?" QWEN3.6 only said its injected model name in the session and gave the same answer when questioned again. On the other hand, QWEN3.8 showed the ability to check its model name, context size, quantization, etc. by sending a Curl request. When asked if it was really true, it even found the actual model file name and reported what model it was.

  • I still observe that when long tasks continue, it sometimes responds mixing English and Korean or Chinese. This is an intermittent phenomenon that also appeared in GPT 5.6 Luna, so I'm not particularly concerned.

  • I instructed it to install the CODEX Extension on a Raspberry Pi 5 inside the Tailscale network and log in. The installation was successful, but it's struggling with CODEX login. GPT 5.5 also failed (it's unclear if OAuth login is even possible), so I'm interested to see if it will succeed (after tinkering with it for a while, I'm now searching the internet...).

  • However, it doesn't stop even when it seems impossible and falls into an infinite loop.... This is still observed.

  • If you only expose essential tools and assign tasks, it seems to produce quite good results as a local model (exposing too many tools increases inference time...).

While it may vary depending on the development area, I'm considering whether the Frontier + QWEN3.8 combination will run smoothly.

When I was using GX10, I had a GPT Plus subscription and tried various things with GPT 5.5 + QWEN3 Coder Next through Roo Code. Now that QWEN3.8 is out, I'm looking forward to hearing from others.

I put QWEN3.8 back to work diligently on LLM Wiki management and email collection/classification AI. ㅎㅎ

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.52 ▼ -0.01

2026.10.09 KEB 하나은행 고시회차 3469회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!