DFlash has been integrated into llamacpp,
so I tried running it once with Qwen 3.6, which is said to work well.
In Strix Halo, which was slow, based on Q6KL
27B : 15-16tps -> 15-23tps fluctuating
35BA3B: 60tps -> 60-90tps fluctuating
Since there is still memory left, I thought it would be enough to just add DFlash, but when I tried it, the diffusion series seemed
somewhat crude and dull ;;
For MoE models, it doesn't seem to be able to understand speech properly either