← all models

deepseek-r1-distill-llama-70b

Rejected  no longer in llama-swap.yaml — entry and/or weights removed after this verdict.

Grill run history

RunScoretok/stepLog
No results-*.log found under this name.

Memory note

DeepSeek-R1-Distill-Llama-70B — REJECTED

Why rejected: Dense 70B model needs extreme quant to fit 32GB VRAM, destroying quality.

  • Q2_K (25GB): OOM at any useful context (>32K)
  • IQ2_S (21GB): OOM at >48K context
  • IQ2_XXS (18GB): Fits 128K but grill 0/3 — model produces garbage at this quant.

Even basic tasks like 232 returned 2.0 instead of 512.0.

Also ~13.6 t/s (5-6x slower than MoE models) and produces 8-11K thinking tokens per task.

The existing MoE models (Ornith-1.0-35b, Agents-A1-35b, Qwen3.6-35b) all provide better quality

at 5-6x the speed with 256K context. No reason to keep this model.

Removed from: llama-swap.yaml (port 9105), models.json