Strix Halo also tried deep frying.

125.250.***.***
13

DeepSpeech V4 0731 was rolled back to Unslos Q2KL

It goes up to IQ3XXS, but it's too tight and the speed is a bit slow,

Anyway, when I run it with ROCm, the prefill seems to stay around 100 seconds, and the tokens stay around 12-13 per second.

If you run it on ds4, the prefill is 150-130 and the tokens are 15-16 more

It's a hassle to go back and forth with ds4, so I switched it to llama for now.

The picture is the result of running the prompt on the link.

There seems to be something wrong with vulkan, so I think some modifications will be needed. If you use the strix halo kv cache modified version, it will generally increase by about 30%, but llama hasn't merged it yet.ㅜㅡㅠ

▶ Original source: https://sunsetcompare.web.app/

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
8.05 =0.00

2026.08.02 KEB 하나은행 고시회차 2385회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!