Strata - Running the Qwen3.8NF model at high speed with 12GB VRAM

31

https://arca.live/b/alpaca/184157616

I saw that you introduced it on this board, and it seems like the Strata engine can optimize and run the Qwen 3.8 NF model. It's almost like the ultimate optimization tool.

Suitable installation conditions are...

VRAM 12GB or more, probably an NVIDIA graphics card

System RAM is 64GB or more

If you install it under these conditions, the output speed will be around 40~60t/s.

Of course, they are Q2 and Q3 quantization models, but

Is it really possible to do this with a VRAM 12GB graphics card????

I was suspicious, so I installed it right away and it was true.

Of course, if you increase the context size, the speed will decrease accordingly.

Anyway, I've tried about four other optimized things here and there,

but in terms of speed, it's completely superior.

The fact that the Qwen 3.8 NF model can be installed with 12GB VRAM is unbelievable.

Anyway, I deleted the existing Bonsai 2 that was set up on my company computer and switched to this.

My company computer has 16GB VRAM and 128GB system RAM, so it seemed optimal for this.

I set the ctx to 128k and ran it, and it outputs 40~50t/s.

Anyway, it seems like developers are doing all sorts of things with the Qwen3.8NF model,

and the results are really, really good.

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.52 ▼ -0.01

2026.10.09 KEB 하나은행 고시회차 2027회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!