I previously posted an article about Qwen 3.8 Splash.
https://damoang.net/ai/7949
There were comments saying that the quality was disappointing due to 4-bit quantization, so I asked Gemini a few more questions about it and got the following new information.
※ This is a tip for those who are disappointed with the performance (inference cliff) of Qwen 3.8 Splash due to 4-bit quantization.
The original Splash engine released initially was 4-bit based, so there was a lot of feedback that its intelligence was lacking in complex coding or mathematical reasoning. Open source users on Reddit (r/LocalLLaMA) and elsewhere have successfully modified the engine code to run an uncompressed native 8-bit version (Native 8-bit, also known as Splash-HQ).
For Mac users with ample RAM capacity, I highly recommend this.
1. Inference Ability (Intelligence) Fully Restored
2. Paradoxically Improved Speed
Contrary to expectations, the uncompressed native 8-bit weights actually increase the hit rate of speculative decoding compared to the heavily compressed quantization model.
As a result, despite significantly higher intelligence, the actual speed is maintained at 37~55 tokens per second, which is much more comfortable and faster than the original Ollama.
3. Recommended for These Users
"It's good that it's fast, but I don't like it when it sometimes talks nonsense or the code breaks."
Users with integrated memory of 48GB ~ 64GB or more, who can comfortably load a model of about 27GB.
For those who need to analyze source code exceeding tens of thousands of lines, or create error-free precise game development/work automation scripts, apply the Native 8-bit patch and weight file released on Reddit instead of using 4-bit. The balance is perfect.