Right now, I'm running DeepSpeech with 0731 Q2KL,
and it consumes about 100W when running and outputs around 10 tokens per second.
It takes 27 hours to output 1 million tokens -_-ㅋ
But if I use it on OpenRouter, I can process it quickly for almost free.
I think my electricity bill will go up ㅋㅋㅋ
Even if I push it hard, Qwen 3.6 27B Q6, 35B A3B Q6, or Step 3.7 Q4 seem to be the baseline;
Of course, DeepSpeech also does well with Q2KL.