The ads made users tempted, claiming GPT 5.6 Terra has the same ability as 5.5 while using only half the tokens. But when I actually used it, on a medium-sized project it does exactly twice as much floundering in code analysis and redesign. Just in case, it even proceeds without verification when I say something wrong about the code. It has a different character from 5.5.
In the end, for raw brains the order is: 5.5 early model > 5.5 late model > 5.6 early model.
The more it claims to save tokens, the dumber it gets -_-
Grok 4.5
I installed Grok Build too, and it's way better than 5.6 Terra. It's about as dumb as 5.6 Terra, but if you give very detailed and precise instructions, the implementation quality is fantastic. And it's twice as fast. It feels like an upgraded version of Qwen 27B.
Conclusion:
GPT 5.6 Terra: a 5.5 that can't catch what you mean
Grok 4.5: a 5.6 Sol that does the job well but is forgetful