I couldn't find any related information even after searching, so I'm posting this.
I found a video on YouTube saying that running 'Qwen 3.8 Splash' on an M6 Mac mini 32G improved performance by twice. I was curious about what it was, so I did some research. While I know that performance has been significantly improved on Macs, I don't have the confidence to explain it in detail, so I asked Gemini for help ㅋㅋㅋ
To get straight to the point, the newly released Qwen 3.8 Splash is not a new AI model but a 'tuning inference engine that has been insanely optimized for Apple Silicon (M series)'. It was open-sourced by Inco AI and recently officially integrated into LM Studio.
Here's a summary of the key points:
1. Insane execution speed (2~3 times faster than existing engines)
They said that only two models, Qwen 3.8 27B / 35B MoE, were targeted and custom optimized with C++ and Metal kernels.
In a typical M series Pro/Max chipset environment, the speed is 70~140 tokens per second for short prompts. It's the fastest local LLM in terms of perceived speed.
Even with long contexts (over 150,000 characters), it maintains a stable speed of 20~30 tokens per second without any slowdown.
2. Intelligence level (perfect replacement for commercial mini models)
While it's slightly less capable than paid flagship models like GPT-4o or 3.5 Sonnet, it easily outperforms cost-effective commercial API levels such as GPT-4o-mini or Claude 3.5 Haiku.
It's more than enough for everyday coding, data extraction, paper summarization, and work automation agents.
3. Reasons why it is strongly recommended for Mac users
You can perform unlimited calculations at a commercial level speed within your local Mac without paying for paid API fees and worrying about your personal code or important data being leaked to external servers.
Especially for those who use Macs with ample integrated memory (RAM) capacity, you can comfortably run even 8-bit (HQ) quantization models with minimal performance degradation.