https://arca.live/b/alpaca/184157616
I saw that you introduced it on this board, and it seems like the Strata engine can optimize and run the Qwen 3.8 NF model. It's almost like the ultimate optimization tool.
Suitable installation conditions are...
VRAM 12GB or more, probably an NVIDIA graphics card
System RAM is 64GB or more
If you install it under these conditions, the output speed will be around 40~60t/s.
Of course, they are Q2 and Q3 quantization models, but
Is it really possible to do this with a VRAM 12GB graphics card????
I was suspicious, so I installed it right away and it was true.
Of course, if you increase the context size, the speed will decrease accordingly.
Anyway, I've tried about four other optimized things here and there,
but in terms of speed, it's completely superior.
The fact that the Qwen 3.8 NF model can be installed with 12GB VRAM is unbelievable.
Anyway, I deleted the existing Bonsai 2 that was set up on my company computer and switched to this.
My company computer has 16GB VRAM and 128GB system RAM, so it seemed optimal for this.
I set the ctx to 128k and ran it, and it outputs 40~50t/s.
Anyway, it seems like developers are doing all sorts of things with the Qwen3.8NF model,
and the results are really, really good.