DFlash is really fast.

61.188.***.***
30

DFlash has been integrated into llamacpp,

so I tried running it once with Qwen 3.6, which is said to work well.

In Strix Halo, which was slow, based on Q6KL

27B : 15-16tps -> 15-23tps fluctuating

35BA3B: 60tps -> 60-90tps fluctuating

Since there is still memory left, I thought it would be enough to just add DFlash, but when I tried it, the diffusion series seemed

somewhat crude and dull ;;

For MoE models, it doesn't seem to be able to understand speech properly either

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
8.28 0.01

2026.07.21 KEB 하나은행 고시회차 655회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!