BottleCap AI said that ThinkingCap optimized the thinking process of Qwen3.6-27B without losing its intelligence.
By reducing unnecessary tokens in the thinking process, it reduced the thinking process by an average of 45%.
In some tests, the output time until the result came out was about 6 times faster.
It also supports MTP.

huggingface : bottlecapai/ThinkingCap-Qwen3.6-27B
Test
oMLX v0.5.1 M5 Pro (20c) 64GB
Prompt : Create a standalone html file simulating the solar system
Model | PREFILL(T/S) | THINKING(S) | TOKEN GEN (T/S) |
|---|
Qwen3.6-27B-oQ4e-mtp | 5.3t/s (23 tok) | 19.3s (545 tok) | 30.2t/s (7630 tok) |
ThinkingCap-Qwen3.6-27B-oQ4e-mtp | 5.2t/s (23 tok) | 6.4s (108 tok) | 29.5t/s (7053 tok) |
We were able to confirm that the THINKING time was significantly reduced.
Results
General model
ThinkingCap