← all models

gemma-4-12b-agentic-yuxinlu1

Rejected  no longer in llama-swap.yaml — entry and/or weights removed after this verdict.

Grill run history

RunScoretok/stepLog
No results-*.log found under this name.

Memory note

yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF — REJECTED

Pulled at Q8_0 (12.67 GB) and grilled 2026-07-06 (session 0b02a136,

log bench/results-gemma-4-12b-agentic-v2-20260706-225728.log). Entry removed

and HF cache deleted — nothing on disk now.

Score: 19/23 at ~21-22 t/s. r1 7/8 (expr_eval), r2 3/5 (regex_dot_star,

weighted_interval_scheduling), r3 5/5, r4 4/5 (repo_fix). The decisive

failure was live, not in the harness: on the buggy-repo agentic fix it **ran

away for 135 turns**, where Qwen3-VL-30B-A3B — evaluated in the same session —

solved it in 6 steps and terminated with a summary.

Why it looked good on paper: card claims tau2-bench telecom ~15% → ~55%

vs the base google/gemma-4-12B-it (3.5x), on the author's own local harness,

20 tasks, self-simulated user. It did not transfer. This is the third

independent case of the same pattern, after [[davidau-catalog-verdicts]]

(glm-grande, coder-uncensored): **an author's own agentic benchmark does not

predict [[grill-round5-agentic-loops]] on this box.** Demand our own round-5

numbers before believing an "agentic" tune.

Secondary reasons not to revisit: dense 12B so decode is weight-read-bound

(~21 t/s, vs 90-100 for the box's 27-35B MoEs); the GGUF ships no mmproj

even though the gemma4_unified base is text+image+audio, so it is text-only;

and its peg-gemma4 tool parser is PEG-native, the family that produced the

[[llama-cpp-peg-native-tool-parser-500]] mixed-output 500.

Cheap for the box, though: ~8.7 KB/token KV at q8_0 (48 layers but only 8

full-attention layers with 1 KV head; the other 40 are sliding-window capped at

1024), so the full 262k window costs only ~2.2 GiB. If a *v3* ever ships, the

sizing work is done — Q8_0 at full context fits in ~15 GiB. Only re-try on a

new version, and grill round 5 first.

Re-derived from bench logs 2026-07-29 when the same repo was proposed again and

neither memory nor the repo held the verdict — hence this file.