StrixHalo Halogen Flash Server User Guide

64

This is an implementation focused solely on StrixHalo's Qwen3.8 Flash Next.

It's still in a closed source state, shared only with a few developers, so I don't know exactly how it works, but it runs.

Filling 256K Full context maintains almost 1000t/s and TG is over 30 in that state ;;;

Vision was initially removed, but now it works, and support has started for Qwen 3.8 FN gguf besides the model file created by the developer.

Of course, their own model is the fastest. (Perplexity has dropped slightly)

It seems to be structured so that KV cache consumes 20GB of memory per 256K.

Currently, it's in a separate container format, so it's a bit awkward to use with other llamacpp. For now, I've attached it as an external service to lemonade and set it to terminate the container after 10 minutes of idle time, but its performance is truly monstrous.

There were some initial concerns because it wasn't open source, but everyone is quiet after seeing the performance haha

Those using AIMAX 395+ 128GB should definitely check out

https://github.com/peonist-ai/halogen-flash-server

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.52 ▼ -0.01

2026.10.09 KEB 하나은행 고시회차 1858회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!