server50 model entries ever added to llama-swap.yaml (50 currently active) —
grill scores, generic-capability sanity checks, and the verdict behind each one. Click a model for full detail.
| Model | Status | Grill (/23) | R5 agentic | lm_eval | HumanEval | Verdict |
|---|---|---|---|---|---|---|
| coder-agentic active |
Kept / trial | 22/23 | — | — | — | WHY this one: superseded as driver by gemma-awq on 2026-08-10 (text-only, |
| fable-fusion active |
Active | 22/23 | — | — | — | fable-fusion (llama.cpp GGUF) regrilled 2026-08-09: the ONLY model on this box that closes the KLayout loop WITHOUT an API reference (6/7, everything else is exactly 0/4). Best KLayout from memory (5/16), best with a reference (15/16), best office (15/16), zero cap-hits in 24 suite-runs. Costs 33 tok/s. |
| pocket-35b active |
Kept / trial | 22/23 | — | — | — | FINAL-Bench/POCKET-35B-GGUF:Q4_K_M (19.71GiB) — TRIAL 2026-07-28. |
| north-mini active |
Kept / trial | 22/23 | — | — | — | unsloth/North-Mini-Code-1.0-GGUF:UD-Q4_K_XL (17.93GiB) — TRIAL 2026-07-28. |
| qwen38-gsq active |
Kept / trial | 21/23 | — | — | — | TRIAL 2026-09-09 — ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF, file |
| laguna-s active |
Kept / trial | 20/23 | — | — | — | 50.3 GiB. RE-TRIED 2026-09-01 after the 2026-08-03 trial was retired |
| bonsai active |
Active | 20/23 | — | — | — | PQ2_0 and the retired dspark drafter were left behind on ssk500). Switched |
| kat-coder active |
Active | 19/23 | — | — | — | (REJECTED), aquila 4/4 ("a deterministic nameable defect"), bigbang too. |
| coder-agentic-q3 active |
Kept / trial | — | — | — | — | All three entries are the SAME base weights, kept side by side on purpose: |
| fable-turbo active |
Kept / trial | — | — | — | — | TRIAL 2026-09-09 — DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic- |
| ornith-15-35b active |
Kept / trial | — | — | — | — | (0.84 GiB) — TRIAL 2026-08-21. Ornith-1.5-35B-A3B, arch qwen3_5_moe |
| tommy active |
Kept / trial | — | — | — | — | thomsonreuters/Thomson-1.0-Small — TRIAL 2026-08-26. Thomson Reuters' |
| tommy2 active |
Active | — | — | — | — | tommy2 (Thomson-1.0-Small IQ2_M, single-card 4060 Ti, 131k) grilled 2026-08-29 with R5 x6: 0 LOOPS in all six runs, 35/36 complete, coding 22/23 — the box's BEST agentic-coding candidate. But static 2-bit DESTROYED API RECALL: office 0/9 unaided vs the parent's 9/9, with 32k-token runaways. Excellent WITH API references, unusable without. |
| aquila active |
Kept / trial | — | — | — | — | XYZAILab/XYZ-Aquila-mini — TRIAL 2026-08-30. Candidate REPLACEMENT for |
| nex-n25-mini active |
Kept / trial | — | — | — | — | nex-agi/Nex-N2.5-mini — TRIAL 2026-09-10. Agentic/computer-use model, |
| apodex active |
Rejected | — | — | gsm8k 30% | — | invents APIs unaided ([[apodex-11-mini-trial]]). Loadable again for |
| glm-flash-ck active |
Kept / trial | — | — | — | — | CONTROLLED QUANTIZER A/B against `glm-flash-awq` — TRIAL 2026-08-31. |
| gpt-oss-20b active |
Promoted | — | — | — | — | RETIRED from this box for exactly that. Promoted on coding strength with the |
| gpt-oss-120b active |
Active | — | — | — | — | gpt-oss-120b MXFP4 (63.39 GB) SERVED on the 32GB box via --n-cpu-moe spill, added 2026-08-31 as `gpt-oss-120b` (ncmoe=22, --tensor-split 75,25, 131072, q8_0 KV, ~14-15 tok/s). Three reusable findings: --tensor-split is NOT a layer split under -ncmoe, q8_0 KV buys COMPUTE not context, and short-prompt prefill numbers are meaningless. |
| nemotron-lightning active |
Kept / trial | — | — | — | — | TRIAL 2026-08-12, PARTIAL. ARCH: nemotron_h_moe (hybrid Mamba-2 / MoE / |
| muse-glimmer active |
Kept / trial | — | — | — | — | meta-models/Muse-Glimmer-30B-GGUF — TRIAL 2026-08-11. Meta Superintelligence |
| embed active |
Kept / trial | — | — | — | — | Kept resident by the embed-persistent group below (never swaps out). |
| glm-flash-awq active |
Active | — | — | — | — | 2026-07-19: raggo.net backend retired (only the SearXNG web search at |
| qwen3-vl-thinking active |
Active | — | — | — | — | QuantTrio/Qwen3-VL-30B-A3B-Thinking-AWQ trialled 2026-08-07 as a controlled A/B against its Instruct sibling (qwen3-vision). The closed-loop hypothesis FAILED — its apparent 2/3 was a grader false pass on geometry 1000x oversize. Real gain is office (3/5, first model to finish the formula->PDF->read-back chain). Coding 3/8 at 16k is a BUDGET ARTIFACT: 6/8 at 32k, with 2 true runaways. |
| qwen3-vision active |
Active | — | — | — | — | |
| qwen3vl active |
Promoted | — | — | — | — | grilled n=3 on 2026-08-24, is the OPPOSITE on the metric that retired Mellum: |
| qwen3-coder active |
Kept / trial | — | — | — | — | ~141 tok/s. Kept because kimi-code/pi/omp reference it by this exact id — |
| qwen3-thinking active |
Kept / trial | — | — | — | — | TRIAL 2026-08-20: the box's one open gap is a FAST TEXT-ONLY THINKING MoE -- |
| qwen3-instruct active |
Kept / trial | — | — | — | — | whether the Thinking tune's CoT buys the score. TRIAL 2026-08-20. |
| thinkingcap active |
Kept / trial | — | — | — | — | 18/18, and ZERO cap-hits in 24 suite runs. See auto-memory thinkingcap-awq-trial. |
| omni active |
Active | — | — | — | — | |
| gemma-awq active |
Active | — | — | — | — | Its GGUF sibling was REJECTED 2026-07-31 for deterministic 16k runaways |
| qwen3next-thinking active |
Promoted | — | — | — | — | across its 3 reps (the suite that most drove the PROMOTED verdict, |
| fable-711-gptq active |
Active | — | — | — | — | |
| qwen36-35b active |
Kept / trial | — | — | — | — | Qwen3_5MoeForConditionalGeneration. See auto-memory qwen36-35b-awq-trial. |
| bonsai-awq active |
Active | — | — | — | — | Ternary-Bonsai-27B AWQ 4-bit trialled 2026-08-10, served 2026-08-11 as `bonsai-awq`: 41/46 coding and 10/12 vision, but 0/16 KLayout from memory and a GENUINE 0/8 on BOTH closed-loop arms, at 168 s/task — the slowest model on the box. It does NOT settle the ternary-vs-4-bit question, because the ternary sibling was never scored on this instrument. |
| qwen38-mtp active |
Kept / trial | — | — | — | — | TRIAL 2026-08-24 — shawnw3i/Qwen3.8-27B-AWQ-MTP: the SAME Qwen3.8-27B AWQ |
| qwen38-awq active |
Kept / trial | — | — | — | — | auto-memory/qwen38-awq-vllm-trial.md. |
| qwen38 active |
Active | — | — | — | — | Qwen3.8-Flash-Next (qwen4exp) SERVES via next-llama.cpp but only at 6.7 tok/s — this box is RAM-poor, not VRAM-poor; revisit after the 64 GB upgrade. |
| gemma12-solo active |
Active | — | — | — | — | |
| gemma12 active |
Active | — | — | — | — | gemma-4-12B QAT AWQ single-card — n=3 grill: office 0/9 unaided vs 9/9 WITH the API ref (3x each), zero R5 loops, realcase 2/3, but KLayout+ref a hard 2/8; its weakness is API RECALL, not capability. |
| glm-ocr active |
Active | — | — | — | — | GLM-OCR 0.9B is the OCR pick (llama.cpp GGUF, #1 OmniDocBench); confine with CUDA_VISIBLE_DEVICES not --device, and parallelism only comes from separate processes. |
| mellum active |
Active | — | — | — | — | |
| qwen3vl-8b active |
Active | — | — | — | — | |
| qwen38-exl3 active |
Active | — | — | — | — | |
| kat-coder-fast active |
Active | — | — | — | — | KAT-Coder MTP GGUF Q4_K_XL — 20.33/23 (identical to the EXL3 build) at 127 tok/s (+67%), works with Claude Code; interrupt_replan loops 1/14, and FOUR candidate sampler fixes all failed. |
| gemma-exl3 active |
Kept / trial | — | — | — | — | ([[kat-coder-exl3-trial]]). Never trust a vision claim without grepping |
| qwen38-flash-next active |
Active | — | — | — | — | Qwen3.8-Flash-Next (qwen4exp) SERVES via next-llama.cpp but only at 6.7 tok/s — this box is RAM-poor, not VRAM-poor; revisit after the 64 GB upgrade. |
| k2-horizon active |
Active | — | — | — | — | IFM K2-Horizon MoVA-36B-A4B — SOLVED 2026-09-06 by building the vendor fork side-by-side as ifm-llama.cpp/ (the prism_llama_bin pattern). Serves at 131072/q4_0 KV, ~63-66 tok/s, wired as `k2-horizon` on :9193. The earlier verdict was wrong three ways: transformers trust_remote_code always worked, the fork does NOT shadow mainline, and MoVA is not a novel attention KERNEL. |
| glm-flash-sglang active |
Active | — | — | — | — | GLM-4.7-Flash added to llama-swap.yaml as \"glm-flash\" — arch/quant/context tuning gotchas and grill results |
Entry and/or weights deleted after the verdict; the record survives only in auto-memory.
| Model | Status | Grill (/23) | Verdict |
|---|---|---|---|
| tess-4-27b | Retired | 22/23 | Tess-4-27B — RETIRED 2026-07-27 (entry + 27GB weights deleted); WAS 22/23, ties vision-coder, one below fable-fusion; a genuine R5 miss and no longer justified a slot |
| kimi-vl | Retired | 10/23 | kimi-vl (Kimi-VL-A3B-Thinking-2506) — RETIRED 2026-07-27, entry + 18GB weights deleted; was vision 17/21 but coding 10/23, NO tool-call format, ◁think▷ unparsed |
| bigbang-v1-trial | Rejected | — | endless-frontier/BigBang-v1 (Qwen3.6-35B-A3B frontier-task post-train) grilled 2026-08-30 — REJECTED. Loses to both tommy and aquila: worst coding (20/23), office 3/9 in BOTH arms (capability ceiling, |
| davidau-catalog-verdicts | Rejected | — | DavidAU HF catalog scan 2026-07-26: 3 models evaluated (gpt-oss-20b abliterated, GLM-Grande-42B, Deckard-40B) all REJECTED; the generalizable pattern is that his Brainstorm expansions + abliterations |
| deepseek-r1-distill-llama-70b | Rejected | — | DeepSeek-R1-Distill-Llama-70B REJECTED — IQ2_XXS too degraded, Q2_K/IQ2_S can't fit 128K context on 32GB |
| devstral-2-123b-iq1s | Rejected | — | Devstral-2-123B UD-IQ1_S REJECTED 2026-07-28 — first 1-bit quant tried on this box; 2/23 with both points from abstention tasks, zero working tool calls; also the load-vs-decode VRAM gotcha |
| devstral-small-2-24b | Retired | — | devstral — RETIRED 2026-07-27 (entry + 19GB weights deleted); WAS 21/23 grill, the box's only Mistral lineage, but a genuine R5 miss and 6x slower than kimi-distill's tied 21/23 |
| ernie-45-vl-28b-thinking | Rejected | — | cyankiwi/ERNIE-4.5-VL-28B-A3B-Thinking-AWQ-4bit REJECTED 2026-08-07. The non-Qwen lineage probe. Perceives fine (vision 4/6) and tool-calls 5/5, but scores 0/8 on KLayout WITH the API reference — the |
| fable711-awq-self-quantize-todo | Rejected | — | REJECTED after 4 attempts 2026-08-06/07: EVERY llmcompressor-AWQ build of Fable-Fusion-711 is functionally broken (empty output), including one quantizing ONLY standard full_attention+MLP via llmcompr |
| froggeric-template-trial | Rejected | — | froggeric/Qwen-Fixed-Chat-Templates v22.4 — REJECTED on analysis (no model loaded); box already fixed the one bug minimally, froggeric is a maximal rewrite that adds real risk (hermes xml/json tool-fo |
| gemma-4-12b-agentic-yuxinlu1 | Rejected | — | yuxinlu1 gemma-4-12B-agentic-fable5-composer2.5-v2 — TRIED AND REJECTED 2026-07-06 at Q8_0: 19/23 and a 135-turn agentic runaway; the author's tau2-telecom claim did not transfer |
| gemma-4-26b-a4b-moe | Rejected | — | google/gemma-4-26B-A4B-it TRIED AND REJECTED 2026-07-31 — 2 reproducible 16000-token runaways in 6 tasks, 21/23 ceiling, crashed under load; also the per-token-speed trap (4x faster t/s but ~6x the to |
| gemma-4-26b-awq-vllm | Rejected | — | gemma-4-26B-A4B-it as AWQ on vLLM, 2026-08-10: the deterministic 16k runaways that got the GGUF REJECTED are GONE, and it scores 44/46 — tied BEST coding on the box — at 93 tok/s in 16 GiB. Only non-Q |
| genesis-hermes-v5 | Rejected | — | REJECTED — LuffyTheFox Qwen3.6-35B Genesis-Hermes-V5-APEX: controlled A/B vs the uncensored entry it shares a base with; 21/23 at 2x the tokens, weights deleted 2026-07-26 |
| glm-4-7-flash-reap-23b | Rejected | — | GLM-4.7-Flash-REAP-23B TRIED AND REJECTED 2026-07-31 across THREE quants — the clean 2x2 that proved expert-pruning is ~free but higher weight-quant is NOT better (Q6 scored BELOW Q4 at identical expe |
| kimi-dev-72b | Rejected | — | Kimi-Dev-72B REJECTED — dense 72B at IQ2_XXS: 10 t/s, 64K max ctx, grill FAIL |
| kimi-distill-x2-trial | Rejected | — | kimi-distill-x2 — 2-slot parallel trial of kimi-distill (fastest 256k model, ~100-106 t/s); FULLY GRILLED clean, zero loops/redundant calls in every round4/round5 run (solo + 2 concurrent); two failur |
| kimi-linear-48b | Retired | — | kimi-linear — RETIRED 2026-07-27 (entry + 22GB weights deleted); WAS 18/23 grill (weakest algorithmic score on the box) + a genuine R5 loop; redundant once fixed |
| laguna-xs-2-1 | Retired | — | poolside Laguna family — laguna-xs RETIRED 2026-07-27 (entry + 24GB weights deleted) despite 23/23; a genuine R5 loop and the score wasn't ladder-comparable (32k vs the standard 16k budget); S-2.1 118 |
| llama-3-3-70b-iq3xxs | Rejected | — | Llama-3.3-70B-Instruct UD-IQ3_XXS — REJECTED TWICE and REMOVED 2026-07-29; the complete grill (16/23, R5 1/6 with 5 loops) plus the parallel-tool-call template limit and the no-mix-guard confound |
| llama4-scout-17b-16e-tq1 | Rejected | — | Llama-4-Scout-17B-16E-Instruct UD-TQ1_0 (near-1-bit ternary, unsloth GGUF, 27.25 GiB) TRIALLED AND REJECTED 2026-08-16. No role: weaker than coder-agentic on every fair axis, weaker than Laguna-S-2.1 |
| mellum2-thinking-vllm-entry | Retired | — | RETIRED 2026-08-26 (entry removed, 23 GB safetensors deleted) — Mellum2-12B-A2.5B-Thinking BF16 on vLLM TP=2 was dominated by coder-agentic; the Mellum2 build that stays is the Q8_0 llama.cpp `mellum` |
| nemotron-3-nano-30b-a3b | Rejected | — | NVIDIA-Nemotron-3-Nano-30B-A3B (hybrid Mamba MoE) TRIED AND REJECTED 2026-07-29 — 20/23 + R5 5/6 with 0 loops at ~90 t/s, but dominated by pocket-35b and its advertised 1M context does not retrieve; a |
| nemotron-35-lightning-30b | Rejected | — | NVIDIA-Nemotron-3.5-Lightning-30B-A3B. FINAL 2026-08-15: KEEP bartowski Q5_K_M, deploy client max_tokens=32000 (llama-swap.yaml has no budget flag). realcase 2/3->3/3 PASS and office 2.7->4.7/6 reacha |
| ornith-awq-trial | Rejected | — | ulkaa/Ornith-1.5-35B-A3B-AWQ-INT4 (compressed-tensors W4A16, vLLM TP=2) — GO probe (loads + coherent, REFUTES the fable711 'AWQ on qwen3_5_moe is always broken' concern), but 1× grill is WORSE than th |
| parallel-agent-slots | Retired | — | coder-x4 — RETIRED 2026-07-27 along with its base gpt-oss-20b (coder) entry; WAS 1.86x aggregate throughput at 4 concurrent agents; no model on the box currently serves --parallel > 1 for chat |