The existing qwen3.5-122b I was using seemed less intelligent than expected and consumed a lot of RAM, leaving insufficient free memory, so I decided to look for an alternative and searched around
A version modified by NVIDIA was uploaded to Hugging Face
qwen3.6-27b dense, qwen3.6-35b-a3b like this
At first, I heard the dense model was good, so I downloaded and tried using it... but it's very slow
So I was thinking about going back to 122b, but since I heard 35b had similar performance, I downloaded it first, and it has Spark options
So I applied everything and uploaded it to vLLM and measured the token generation speed
| Item | 27B (Previous) | 35B-A3B (Current) |
| ------- | -------------- | ----------------- |
| Context | 131K | 262K |
| VRAM | 98GB | 22GB |
| RAM | 109GB | 79GB (including cache) |
| Speed | 9 tok/s | 105 tok/s |
Roughly this is what I got. The 105 tokens were measured without OpenCode applied, so when OpenCode is actually applied, I get lower token generation, but it still produces more than 80
At this level, there's no problem with actual use
I heard that with 2 Sparks you can operate DeepSeek-4 Flash.... I heard MiniMax3 works too.... I'm jealous