Certainly, frontier-level models are prone to some tinkering locally.

222.217.***.***
63

Right now, I'm running DeepSpeech with 0731 Q2KL,

and it consumes about 100W when running and outputs around 10 tokens per second.

It takes 27 hours to output 1 million tokens -_-ㅋ

But if I use it on OpenRouter, I can process it quickly for almost free.

I think my electricity bill will go up ㅋㅋㅋ

Even if I push it hard, Qwen 3.6 27B Q6, 35B A3B Q6, or Step 3.7 Q4 seem to be the baseline;

Of course, DeepSpeech also does well with Q2KL.

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.83 0.01

2026.08.24 KEB 하나은행 고시회차 482회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!