Actually, I was just using installation-type LLMs for fun, but I came across Ornith while looking for a lightweight model to run on my notebook.
One-line summary: Can't code, but good at following instructions
Two-line summary: If you tell it not to follow instructions and do it yourself, it can't do that either.
The following is a Codex Python click benchmark (aka laptop running standard) that I created and use.
Conditions are raw Ollama model, num_ctx=4096, temperature 0, Python 3.9 compatible
Total Score
Model | Task Pass | Case Score | Accuracy | Avg tok/s |
|---|
Qwen3-Coder 30B-A3B | 5/6 | 35/37 | 94.59% | 42.27 |
Ornith 35B Q4 4K | 4/6 | 35/37 | 94.59% | 11.06 |
Devstral 24B Dense | 3/6 | 31/37 | 83.78% | 15.30 |
Ornith 9B Q8 | 4/6 | 29/37 | 78.38% | 23.00 |
Item Score
Model | Code Writing | Error Correction/Bugfix | Data Structure Implementation | Error Detection/Review |
|---|
Qwen3-Coder 30B | 18/18 | 12/12 | 2/2 | 3/5 |
Ornith 35B | 17/18 | 11/12 | 2/2 | 5/5 |
Devstral 24B | 14/18 | 12/12 | 2/2 | 3/5 |
Ornith 9B Q8 | 12/18 | 12/12 | 0/2 | 5/5 |
Error Detection
Model | Exact | Candidate Acc | Precision | Recall | F1 |
|---|
Ornith 9B | 3/5 | 90.32% | 88.89% | 94.12% | 91.43% |
Qwen3-Coder 30B | 3/5 | 87.10% | 84.21% | 94.12% | 88.89% |
Code Structure Understanding
Model | Exact | Candidate Acc | Precision | Recall | F1 |
|---|
Ornith 9B | 3/4 | 95.83% | 92.86% | 100.00% | 96.30% |
Qwen3-Coder 30B | 3/4 | 95.83% | 100.00% | 92.31% | 96.00% |
Interpretation:
9B is truly usable in review/detection. It tends to have higher recall than missing things.
Qwen 30B is more conservative. There are fewer false positives in code structure understanding, but it missed one reachable bug.
Both had weaknesses in false positive control. 9B flagged two false bugs, and Qwen flagged three.
The format is riskier for 9B. The content is correct, but it doesn't adhere to the JSON array format, outputting something like {"A","B"}. A parser correction would be needed if using it as a reviewer.
The recommended structure has become clearer:
Code writing/modification: Qwen3-Coder 30B
Error detection/initial code review: Ornith 9B
Final application decision: Cross-verification of Qwen results and 9B review results
Output such as strict block_patch.json should not be left to 9B alone.