Rejected no longer in llama-swap.yaml — entry and/or weights removed after this verdict.
| Run | Score | tok/step | Log |
|---|---|---|---|
| No results-*.log found under this name. | |||
yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF — REJECTEDPulled at Q8_0 (12.67 GB) and grilled 2026-07-06 (session 0b02a136,
log bench/results-gemma-4-12b-agentic-v2-20260706-225728.log). Entry removed
and HF cache deleted — nothing on disk now.
Score: 19/23 at ~21-22 t/s. r1 7/8 (expr_eval), r2 3/5 (regex_dot_star,
weighted_interval_scheduling), r3 5/5, r4 4/5 (repo_fix). The decisive
failure was live, not in the harness: on the buggy-repo agentic fix it **ran
away for 135 turns**, where Qwen3-VL-30B-A3B — evaluated in the same session —
solved it in 6 steps and terminated with a summary.
Why it looked good on paper: card claims tau2-bench telecom ~15% → ~55%
vs the base google/gemma-4-12B-it (3.5x), on the author's own local harness,
20 tasks, self-simulated user. It did not transfer. This is the third
independent case of the same pattern, after [[davidau-catalog-verdicts]]
(glm-grande, coder-uncensored): **an author's own agentic benchmark does not
predict [[grill-round5-agentic-loops]] on this box.** Demand our own round-5
numbers before believing an "agentic" tune.
Secondary reasons not to revisit: dense 12B so decode is weight-read-bound
(~21 t/s, vs 90-100 for the box's 27-35B MoEs); the GGUF ships no mmproj
even though the gemma4_unified base is text+image+audio, so it is text-only;
and its peg-gemma4 tool parser is PEG-native, the family that produced the
[[llama-cpp-peg-native-tool-parser-500]] mixed-output 500.
Cheap for the box, though: ~8.7 KB/token KV at q8_0 (48 layers but only 8
full-attention layers with 1 KV head; the other 40 are sliding-window capped at
1024), so the full 262k window costs only ~2.2 GiB. If a *v3* ever ships, the
sizing work is done — Q8_0 at full context fits in ~15 GiB. Only re-try on a
new version, and grill round 5 first.
Re-derived from bench logs 2026-07-29 when the same repo was proposed again and
neither memory nor the repo held the verdict — hence this file.