It is said to be an agent-specialized model, but it argues that the difference is that it has applied a self-improvement loop and improved compared to version 1.0, which just received such evaluations.
397B (Dense)
35B-A3B (MoE)
9B
These are three models.
The 397B model has achieved a benchmark score comparable to opus 4.8...

It is said that the 9B model surpasses Gemma4 31B in benchmarks...

Anyway, when I tried it through ollama and openclaw, the 9B model also answered quite intelligently. However, the loop running when it gets longer remains unchanged, and unexpectedly, Chinese still doesn't disappoint, but it does feel like "it's talking compared to version 1.0".
Please refer to it~
▶ Original source: https://ornith.ai/ornith_1_5.html