Rejected no longer in llama-swap.yaml — entry and/or weights removed after this verdict.
| Run | Score | tok/step | Log |
|---|---|---|---|
| No results-*.log found under this name. | |||
Why rejected: Dense 72B model (Qwen2.5-72B base) at extreme quant, same problems as [[deepseek-r1-distill-llama-70b]].
1. Dense 72B at IQ2_XXS = extreme quant degradation (same class as rejected DeepSeek)
2. ~10 t/s (6-8x slower than MoE models at 60-85 t/s)
3. Max 64K context (vs 128-256K for MoE models)
4. Produces ◁think▷ tags that break grill code extraction
5. Extremely verbose: hits 16K token budget per task (25+ min/task on this hardware)
6. 412 MiB VRAM margin at 64K (no room for larger context or better KV)
Comparison: existing MoE models (Ornith-1.0-35b, Agents-A1-35b, Qwen3.6-35b) all provide better quality, speed, and context.
Removed from: not added to llama-swap.yaml or models.json