Models on server

50 model entries ever added to llama-swap.yaml (50 currently active) — grill scores, generic-capability sanity checks, and the verdict behind each one. Click a model for full detail.

Model Status Grill (/23) R5 agentic lm_eval HumanEval Verdict
coder-agentic
active
Kept / trial 22/23 WHY this one: superseded as driver by gemma-awq on 2026-08-10 (text-only,
fable-fusion
active
Active 22/23 fable-fusion (llama.cpp GGUF) regrilled 2026-08-09: the ONLY model on this box that closes the KLayout loop WITHOUT an API reference (6/7, everything else is exactly 0/4). Best KLayout from memory (5/16), best with a reference (15/16), best office (15/16), zero cap-hits in 24 suite-runs. Costs 33 tok/s.
pocket-35b
active
Kept / trial 22/23 FINAL-Bench/POCKET-35B-GGUF:Q4_K_M (19.71GiB) — TRIAL 2026-07-28.
north-mini
active
Kept / trial 22/23 unsloth/North-Mini-Code-1.0-GGUF:UD-Q4_K_XL (17.93GiB) — TRIAL 2026-07-28.
qwen38-gsq
active
Kept / trial 21/23 TRIAL 2026-09-09 — ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF, file
laguna-s
active
Kept / trial 20/23 50.3 GiB. RE-TRIED 2026-09-01 after the 2026-08-03 trial was retired
bonsai
active
Active 20/23 PQ2_0 and the retired dspark drafter were left behind on ssk500). Switched
kat-coder
active
Active 19/23 (REJECTED), aquila 4/4 ("a deterministic nameable defect"), bigbang too.
coder-agentic-q3
active
Kept / trial All three entries are the SAME base weights, kept side by side on purpose:
fable-turbo
active
Kept / trial TRIAL 2026-09-09 — DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-
ornith-15-35b
active
Kept / trial (0.84 GiB) — TRIAL 2026-08-21. Ornith-1.5-35B-A3B, arch qwen3_5_moe
tommy
active
Kept / trial thomsonreuters/Thomson-1.0-Small — TRIAL 2026-08-26. Thomson Reuters'
tommy2
active
Active tommy2 (Thomson-1.0-Small IQ2_M, single-card 4060 Ti, 131k) grilled 2026-08-29 with R5 x6: 0 LOOPS in all six runs, 35/36 complete, coding 22/23 — the box's BEST agentic-coding candidate. But static 2-bit DESTROYED API RECALL: office 0/9 unaided vs the parent's 9/9, with 32k-token runaways. Excellent WITH API references, unusable without.
aquila
active
Kept / trial XYZAILab/XYZ-Aquila-mini — TRIAL 2026-08-30. Candidate REPLACEMENT for
nex-n25-mini
active
Kept / trial nex-agi/Nex-N2.5-mini — TRIAL 2026-09-10. Agentic/computer-use model,
apodex
active
Rejected gsm8k 30% invents APIs unaided ([[apodex-11-mini-trial]]). Loadable again for
glm-flash-ck
active
Kept / trial CONTROLLED QUANTIZER A/B against `glm-flash-awq` — TRIAL 2026-08-31.
gpt-oss-20b
active
Promoted RETIRED from this box for exactly that. Promoted on coding strength with the
gpt-oss-120b
active
Active gpt-oss-120b MXFP4 (63.39 GB) SERVED on the 32GB box via --n-cpu-moe spill, added 2026-08-31 as `gpt-oss-120b` (ncmoe=22, --tensor-split 75,25, 131072, q8_0 KV, ~14-15 tok/s). Three reusable findings: --tensor-split is NOT a layer split under -ncmoe, q8_0 KV buys COMPUTE not context, and short-prompt prefill numbers are meaningless.
nemotron-lightning
active
Kept / trial TRIAL 2026-08-12, PARTIAL. ARCH: nemotron_h_moe (hybrid Mamba-2 / MoE /
muse-glimmer
active
Kept / trial meta-models/Muse-Glimmer-30B-GGUF — TRIAL 2026-08-11. Meta Superintelligence
embed
active
Kept / trial Kept resident by the embed-persistent group below (never swaps out).
glm-flash-awq
active
Active 2026-07-19: raggo.net backend retired (only the SearXNG web search at
qwen3-vl-thinking
active
Active QuantTrio/Qwen3-VL-30B-A3B-Thinking-AWQ trialled 2026-08-07 as a controlled A/B against its Instruct sibling (qwen3-vision). The closed-loop hypothesis FAILED — its apparent 2/3 was a grader false pass on geometry 1000x oversize. Real gain is office (3/5, first model to finish the formula->PDF->read-back chain). Coding 3/8 at 16k is a BUDGET ARTIFACT: 6/8 at 32k, with 2 true runaways.
qwen3-vision
active
Active
qwen3vl
active
Promoted grilled n=3 on 2026-08-24, is the OPPOSITE on the metric that retired Mellum:
qwen3-coder
active
Kept / trial ~141 tok/s. Kept because kimi-code/pi/omp reference it by this exact id —
qwen3-thinking
active
Kept / trial TRIAL 2026-08-20: the box's one open gap is a FAST TEXT-ONLY THINKING MoE --
qwen3-instruct
active
Kept / trial whether the Thinking tune's CoT buys the score. TRIAL 2026-08-20.
thinkingcap
active
Kept / trial 18/18, and ZERO cap-hits in 24 suite runs. See auto-memory thinkingcap-awq-trial.
omni
active
Active
gemma-awq
active
Active Its GGUF sibling was REJECTED 2026-07-31 for deterministic 16k runaways
qwen3next-thinking
active
Promoted across its 3 reps (the suite that most drove the PROMOTED verdict,
fable-711-gptq
active
Active
qwen36-35b
active
Kept / trial Qwen3_5MoeForConditionalGeneration. See auto-memory qwen36-35b-awq-trial.
bonsai-awq
active
Active Ternary-Bonsai-27B AWQ 4-bit trialled 2026-08-10, served 2026-08-11 as `bonsai-awq`: 41/46 coding and 10/12 vision, but 0/16 KLayout from memory and a GENUINE 0/8 on BOTH closed-loop arms, at 168 s/task — the slowest model on the box. It does NOT settle the ternary-vs-4-bit question, because the ternary sibling was never scored on this instrument.
qwen38-mtp
active
Kept / trial TRIAL 2026-08-24 — shawnw3i/Qwen3.8-27B-AWQ-MTP: the SAME Qwen3.8-27B AWQ
qwen38-awq
active
Kept / trial auto-memory/qwen38-awq-vllm-trial.md.
qwen38
active
Active Qwen3.8-Flash-Next (qwen4exp) SERVES via next-llama.cpp but only at 6.7 tok/s — this box is RAM-poor, not VRAM-poor; revisit after the 64 GB upgrade.
gemma12-solo
active
Active
gemma12
active
Active gemma-4-12B QAT AWQ single-card — n=3 grill: office 0/9 unaided vs 9/9 WITH the API ref (3x each), zero R5 loops, realcase 2/3, but KLayout+ref a hard 2/8; its weakness is API RECALL, not capability.
glm-ocr
active
Active GLM-OCR 0.9B is the OCR pick (llama.cpp GGUF, #1 OmniDocBench); confine with CUDA_VISIBLE_DEVICES not --device, and parallelism only comes from separate processes.
mellum
active
Active
qwen3vl-8b
active
Active
qwen38-exl3
active
Active
kat-coder-fast
active
Active KAT-Coder MTP GGUF Q4_K_XL — 20.33/23 (identical to the EXL3 build) at 127 tok/s (+67%), works with Claude Code; interrupt_replan loops 1/14, and FOUR candidate sampler fixes all failed.
gemma-exl3
active
Kept / trial ([[kat-coder-exl3-trial]]). Never trust a vision claim without grepping
qwen38-flash-next
active
Active Qwen3.8-Flash-Next (qwen4exp) SERVES via next-llama.cpp but only at 6.7 tok/s — this box is RAM-poor, not VRAM-poor; revisit after the 64 GB upgrade.
k2-horizon
active
Active IFM K2-Horizon MoVA-36B-A4B — SOLVED 2026-09-06 by building the vendor fork side-by-side as ifm-llama.cpp/ (the prism_llama_bin pattern). Serves at 131072/q4_0 KV, ~63-66 tok/s, wired as `k2-horizon` on :9193. The earlier verdict was wrong three ways: transformers trust_remote_code always worked, the fork does NOT shadow mainline, and MoVA is not a novel attention KERNEL.
glm-flash-sglang
active
Active GLM-4.7-Flash added to llama-swap.yaml as \"glm-flash\" — arch/quant/context tuning gotchas and grill results

Also tested — no longer in llama-swap.yaml

Entry and/or weights deleted after the verdict; the record survives only in auto-memory.

ModelStatusGrill (/23)Verdict
tess-4-27b Retired 22/23 Tess-4-27B — RETIRED 2026-07-27 (entry + 27GB weights deleted); WAS 22/23, ties vision-coder, one below fable-fusion; a genuine R5 miss and no longer justified a slot
kimi-vl Retired 10/23 kimi-vl (Kimi-VL-A3B-Thinking-2506) — RETIRED 2026-07-27, entry + 18GB weights deleted; was vision 17/21 but coding 10/23, NO tool-call format, ◁think▷ unparsed
bigbang-v1-trial Rejected endless-frontier/BigBang-v1 (Qwen3.6-35B-A3B frontier-task post-train) grilled 2026-08-30 — REJECTED. Loses to both tommy and aquila: worst coding (20/23), office 3/9 in BOTH arms (capability ceiling,
davidau-catalog-verdicts Rejected DavidAU HF catalog scan 2026-07-26: 3 models evaluated (gpt-oss-20b abliterated, GLM-Grande-42B, Deckard-40B) all REJECTED; the generalizable pattern is that his Brainstorm expansions + abliterations
deepseek-r1-distill-llama-70b Rejected DeepSeek-R1-Distill-Llama-70B REJECTED — IQ2_XXS too degraded, Q2_K/IQ2_S can't fit 128K context on 32GB
devstral-2-123b-iq1s Rejected Devstral-2-123B UD-IQ1_S REJECTED 2026-07-28 — first 1-bit quant tried on this box; 2/23 with both points from abstention tasks, zero working tool calls; also the load-vs-decode VRAM gotcha
devstral-small-2-24b Retired devstral — RETIRED 2026-07-27 (entry + 19GB weights deleted); WAS 21/23 grill, the box's only Mistral lineage, but a genuine R5 miss and 6x slower than kimi-distill's tied 21/23
ernie-45-vl-28b-thinking Rejected cyankiwi/ERNIE-4.5-VL-28B-A3B-Thinking-AWQ-4bit REJECTED 2026-08-07. The non-Qwen lineage probe. Perceives fine (vision 4/6) and tool-calls 5/5, but scores 0/8 on KLayout WITH the API reference — the
fable711-awq-self-quantize-todo Rejected REJECTED after 4 attempts 2026-08-06/07: EVERY llmcompressor-AWQ build of Fable-Fusion-711 is functionally broken (empty output), including one quantizing ONLY standard full_attention+MLP via llmcompr
froggeric-template-trial Rejected froggeric/Qwen-Fixed-Chat-Templates v22.4 — REJECTED on analysis (no model loaded); box already fixed the one bug minimally, froggeric is a maximal rewrite that adds real risk (hermes xml/json tool-fo
gemma-4-12b-agentic-yuxinlu1 Rejected yuxinlu1 gemma-4-12B-agentic-fable5-composer2.5-v2 — TRIED AND REJECTED 2026-07-06 at Q8_0: 19/23 and a 135-turn agentic runaway; the author's tau2-telecom claim did not transfer
gemma-4-26b-a4b-moe Rejected google/gemma-4-26B-A4B-it TRIED AND REJECTED 2026-07-31 — 2 reproducible 16000-token runaways in 6 tasks, 21/23 ceiling, crashed under load; also the per-token-speed trap (4x faster t/s but ~6x the to
gemma-4-26b-awq-vllm Rejected gemma-4-26B-A4B-it as AWQ on vLLM, 2026-08-10: the deterministic 16k runaways that got the GGUF REJECTED are GONE, and it scores 44/46 — tied BEST coding on the box — at 93 tok/s in 16 GiB. Only non-Q
genesis-hermes-v5 Rejected REJECTED — LuffyTheFox Qwen3.6-35B Genesis-Hermes-V5-APEX: controlled A/B vs the uncensored entry it shares a base with; 21/23 at 2x the tokens, weights deleted 2026-07-26
glm-4-7-flash-reap-23b Rejected GLM-4.7-Flash-REAP-23B TRIED AND REJECTED 2026-07-31 across THREE quants — the clean 2x2 that proved expert-pruning is ~free but higher weight-quant is NOT better (Q6 scored BELOW Q4 at identical expe
kimi-dev-72b Rejected Kimi-Dev-72B REJECTED — dense 72B at IQ2_XXS: 10 t/s, 64K max ctx, grill FAIL
kimi-distill-x2-trial Rejected kimi-distill-x2 — 2-slot parallel trial of kimi-distill (fastest 256k model, ~100-106 t/s); FULLY GRILLED clean, zero loops/redundant calls in every round4/round5 run (solo + 2 concurrent); two failur
kimi-linear-48b Retired kimi-linear — RETIRED 2026-07-27 (entry + 22GB weights deleted); WAS 18/23 grill (weakest algorithmic score on the box) + a genuine R5 loop; redundant once fixed
laguna-xs-2-1 Retired poolside Laguna family — laguna-xs RETIRED 2026-07-27 (entry + 24GB weights deleted) despite 23/23; a genuine R5 loop and the score wasn't ladder-comparable (32k vs the standard 16k budget); S-2.1 118
llama-3-3-70b-iq3xxs Rejected Llama-3.3-70B-Instruct UD-IQ3_XXS — REJECTED TWICE and REMOVED 2026-07-29; the complete grill (16/23, R5 1/6 with 5 loops) plus the parallel-tool-call template limit and the no-mix-guard confound
llama4-scout-17b-16e-tq1 Rejected Llama-4-Scout-17B-16E-Instruct UD-TQ1_0 (near-1-bit ternary, unsloth GGUF, 27.25 GiB) TRIALLED AND REJECTED 2026-08-16. No role: weaker than coder-agentic on every fair axis, weaker than Laguna-S-2.1
mellum2-thinking-vllm-entry Retired RETIRED 2026-08-26 (entry removed, 23 GB safetensors deleted) — Mellum2-12B-A2.5B-Thinking BF16 on vLLM TP=2 was dominated by coder-agentic; the Mellum2 build that stays is the Q8_0 llama.cpp `mellum`
nemotron-3-nano-30b-a3b Rejected NVIDIA-Nemotron-3-Nano-30B-A3B (hybrid Mamba MoE) TRIED AND REJECTED 2026-07-29 — 20/23 + R5 5/6 with 0 loops at ~90 t/s, but dominated by pocket-35b and its advertised 1M context does not retrieve; a
nemotron-35-lightning-30b Rejected NVIDIA-Nemotron-3.5-Lightning-30B-A3B. FINAL 2026-08-15: KEEP bartowski Q5_K_M, deploy client max_tokens=32000 (llama-swap.yaml has no budget flag). realcase 2/3->3/3 PASS and office 2.7->4.7/6 reacha
ornith-awq-trial Rejected ulkaa/Ornith-1.5-35B-A3B-AWQ-INT4 (compressed-tensors W4A16, vLLM TP=2) — GO probe (loads + coherent, REFUTES the fable711 'AWQ on qwen3_5_moe is always broken' concern), but 1× grill is WORSE than th
parallel-agent-slots Retired coder-x4 — RETIRED 2026-07-27 along with its base gpt-oss-20b (coder) entry; WAS 1.86x aggregate throughput at 4 concurrent agents; no model on the box currently serves --parallel > 1 for chat