Rejected active in llama-swap.yaml · aliases: apodex-1.1-mini-qwen35-256k
| Run | Score | tok/step | R5 | Log |
|---|---|---|---|---|
| 20260901-222422 | 6/9 | — | — | results-apodex-office-apihelp-20260901-222422.log |
| 20260901-221223 | 0/9 | — | — | results-apodex-office-20260901-221223.log |
| 20260901-214339 | 4/8 | — | — | results-apodex-klayout-apihelp-20260901-214339.log |
gsm8k / ifeval / truthfulqa_gen via EleutherAI lm-evaluation-harness, generate_until only (this backend can't reach multiple-choice tasks). Separate axis, not part of the 23-task grill score.
| gsm8k | sample_len=20.000 exact_match:strict-match=0.300 exact_match:flexible-extract=0.300 |
| ifeval | sample_len=20.000 prompt_level_strict_acc:none=0.100 inst_level_strict_acc:none=0.067 prompt_level_loose_acc:none=0.100 inst_level_loose_acc:none=0.067 |
| truthfulqa_gen | sample_len=20.000 bleu_max:none=17.930 bleu_acc:none=0.600 bleu_diff:none=3.373 rouge1_max:none=37.818 rouge1_acc:none=0.500 rouge1_diff:none=7.352 rouge2_max:none=27.706 rouge2_acc:none=0.500 rouge2_diff:none=5.567 rougeL_max:none=36.134 rougeL_acc:none=0.550 rougeL_diff:none=6.282 |
apodex/Apodex-1.1-mini, bartowski Q4_K_M (21.86 GiB) + mmproj-bf16 (903 MB),
served as apodex on port 9181 (entry now commented out). Fine-tune of Qwen3.5-35B-A3B ("reasoning-first
for long-horizon research"), apache-2.0, 262144 ctx, vision depth 27 — arch
verified IDENTICAL to aquila/tommy before download.
| suite | unaided | + API ref |
|---|---|---|
| textturn tag leak | 0/15 CLEAN | — |
| textturn fabrication | 3/5 no-tools, 0/10 tools-bound | — |
| KLayout | — | 4/8 |
| office | 0/9 | 6/9 (vision 2/2) |
| R5 (n=3) | complete 4/5/4 of 6, looped 2/1/2 | — |
Speed ~94 tok/s. Split 45,55 (see [[mmproj-caps-cuda0-tensor-split]]).
interrupt_replan loops 3/3run1 complete 4/6 looped 2/6 ['multifile_refactor','interrupt_replan']
run2 complete 5/6 looped 1/6 ['interrupt_replan']
run3 complete 4/6 looped 2/6 ['multifile_refactor','interrupt_replan']
Completion wobbled 4-5/6 = noise, exactly as [[round5-is-a-sample-not-a-measurement]]
predicts. The looping-task IDENTITY was perfectly stable, which is the signal
per [[grill-round5-agentic-loops]]: interrupt_replan failed every run. A model
that cannot re-plan after an interrupt must not drive a tool loop.
Two independent libraries, same failure class:
* PolygonWithProperties.is_polygon — verified absent from klayout 0.30.9 AND
appearing 0 times in the API reference we supplied. Cost 2 of 4 KLayout
failures plus a 16k-token runaway on min_spacing.
* chart.title.text = '...' on a fresh openpyxl chart (which is None) — wrong
in 3/3 office stages. Correct idiom chart.title = "str" is stated
verbatim in OFFICE_API_HELP.
But this is RECALL, not a ceiling — and I nearly got that wrong. On the
unaided 0/9 I was ready to reject it as a BigBang-style capability ceiling
([[bigbang-v1-trial]]). The +ref arm took office to 6/9, i.e. the gemma12
pattern ([[gemma12-grill]]), not BigBang's. **Always run the ref arm before
calling something a ceiling** — 0/N unaided discriminates nothing on its own.
It clears the think-tag leak gate that disqualified aquila — 0/15 leaks
where aquila leaked into message.content on its first real Claude Code turn
([[aquila-think-tag-leak-real-use]]). --reasoning-format deepseek works here
because Apodex emits <think>. Vision is real (office vision 2/2 off rendered
PDFs). Fabricated execution appears only in the no-tools arm (3/5); with tools
bound it is 0/10, so the agent-shaped path is clean.
RETIRED 2026-09-03 — entry commented out, weights PARKED not deleted
([[ask-before-deleting-weights]]). The -sm tensor arm was abandoned UNGRADED
after the box crash it coincided with ([[xid79-gpu-fell-off-bus-sm-tensor]]):
no payoff in re-running a box-crashing arm for a model that will never take
traffic. Weights moved to /mnt/wd1tb/models/apodex/ (USB spinning HDD,
the parked tier — [[storage-tiers]]), freeing 22 GB on /mnt/models; the
commented entry's -m/--mmproj paths already point there, but **do not serve
from that tier** — move them back to /mnt/models or /mnt/ssk500 first.
REJECTED as a driver, not deleted. It does
not beat the incumbents it would displace — qwen36-35b (KLayout+ref 15/16),
glm-flash (13/16), fable-711-gptq (7/8) — and per
[[model-triage-checklist]] check 8 a derivative must justify itself. The patched
template chat-templates/apodex-1.1-256k.jinja is kept alongside the commented
entry so the measurements stay reproducible.
/mnt/wd1tb (where the weights were parked) died and was replaced by
/mnt/ssd1700, a genuinely fast USB SSD. Weights re-downloaded byte-exact
from bartowski/apodex_Apodex-1.1-mini-GGUF. Giovanni asked to uncomment
apodex along with the two genuinely storage-blocked entries (gpt-oss-120b,
laguna-s) since the drive is fast now — but **the storage-speed framing
does not actually apply to apodex**: it loads fully into VRAM (-ngl 99, no
--n-cpu-moe spill), so wd1tb's slowness was never why it was pulled. The
REAL reason (interrupt_replan loops 3/3) is independent of storage and
still stands. Re-enabled anyway per the explicit request, verified
standalone (loads in ~15s, free VRAM 2964/2169 MiB, non-first-system-message
guard confirmed still working, generates correctly), and HUP-reloaded into
production llama-swap — but flagged this distinction clearly rather than
silently re-enabling a rejected driver under a rationale that didn't apply
to it. Still do NOT wire to Claude Code / pi / kimi-code tool loops.
# ===== RE-ENABLED 2026-09-11 — weights moved off the dead /mnt/wd1tb HDD
# onto /mnt/ssd1700 (fast USB SSD, confirmed non-rotational, ~1.8 GB/s
# read), which resolves the storage-tier concern this entry carried while
# parked ([[storage-tiers]]). THE QUALITY REJECTION FROM THE 2026-09-01
# GRILL STILL STANDS AND IS UNRELATED TO STORAGE — this model loads fully
# into VRAM (-ngl 99, no --n-cpu-moe spill), so disk speed was never the
# issue for apodex specifically: `interrupt_replan` looped 3/3 and it
# invents APIs unaided ([[apodex-11-mini-trial]]). Loadable again for
# testing/comparison — do NOT wire it to Claude Code / pi / kimi-code tool
# loops as a driver. A related `-sm tensor` variant coincided with the
# box's only Xid 79 (5060 Ti fell off the PCIe bus,
# [[xid79-gpu-fell-off-bus-sm-tensor]]); that arm stays abandoned — this
# entry does not use -sm tensor.
# chat-templates/apodex-1.1-256k.jinja is kept, and
# --chat-template-file is load-bearing (the embedded template 400s on any
# non-first system message).
"apodex":
aliases: [apodex-1.1-mini-qwen35-256k]
# apodex/Apodex-1.1-mini — TRIAL 2026-09-01. Another Qwen3.5-35B-A3B
# derivative in the crowded 35B-A3B slot, this one "reasoning-first" for
# agentic / long-horizon research (papers, datasets, images, code) with
# native function calling. Apache-2.0. Competes with `aquila` (Deep Search)
# and `tommy` (RAG fidelity); see those entries for what it has to beat.
#
# ARCH: verified IDENTICAL to aquila/tommy before download —
# Qwen3_5MoeForConditionalGeneration, qwen3_5_moe, 256 experts 8/tok,
# 40 layers, full_attention_interval 4 (= 10 full-attn layers),
# 262144 ctx, vision_config depth 27. So the box's qwen3_5_moe
# infrastructure (KV sizing, bandwidth ceiling, the non-first-system
# jinja guard) all applies.
#
# CHAT TEMPLATE IS NOT BYTE-IDENTICAL — unlike aquila (which shares
# Thomson's template, md5 52b6d51ae5b2), Apodex ships a CUSTOM 8831-char
# template with a baked-in identity preamble ("Always respond as Apodex
# and never pretend to be any other AI model"). It INHERITED the
# non-first-system raise_exception guard (line 104: 'System message must
# be at the beginning.') that 400s on Claude Code / tool-use clients
# ([[jinja-system-guard-tool-parser]]). Patched into a dedicated file —
# chat-templates/apodex-1.1-256k.jinja — using the SAME one-line fix as
# qwen35moe-nonfirst-system-256k.jinja (guard -> render the block).
# --chat-template-file is LOAD-BEARING; without it the embedded template
# 400s on any non-first system message.
#
# QUANT = Q4_K_M (21.86 GiB), NOT Q5_K_M. Apodex's Q5_K_M is 25.5 GiB —
# +2.2 GiB over aquila's Q5 (23.30 GiB), which left only ~1,112 MiB on
# the binding card at split 45,55. Q5 at 45,55 is a ~100 MiB DEFICIT ->
# OOM, and retuning the split blind is unreliable: non-weight allocs
# (mmproj + KV + compute + ViT) land preferentially on CUDA0 so the
# layer-share model does not hold ([[gpu-device-ordering]],
# [[gpu-card-assignment-policy]] — splits are measurement-tuned, not
# inherited). Q4_K_M is SMALLER than aquila's Q5, so aquila's measured
# 45,55 envelope carries over with extra headroom — reliable for
# grilling without an OOM confounding the vision stages. Q5_K_M + a
# measurement-tuned split is the PROMOTION-time task, not a trial one.
#
# REASONING: Qwen3-native `mid`/`` tags (same family as tommy); card
# specifies the "qwen3 reasoning parser". --reasoning-format deepseek
# extracts thoughts into reasoning_content (only two modes exist: none /
# deepseek). VALIDATE with bench/grill_textturn.py before driving Claude
# Code — aquila's grill scored CLEAN then leaked think-tags in real use
# ([[aquila-think-tag-leak-real-use]], [[grill-does-not-validate-real-use]]).
#
# STORAGE: /mnt/ssd1700 (fast USB SSD, replaced the dead /mnt/wd1tb HDD
# 2026-09-11; weights re-downloaded byte-exact from
# bartowski/apodex_Apodex-1.1-mini-GGUF). See [[storage-tiers]].
#
# ===== GRILLED 2026-09-01 — REJECTED AS DRIVER. See
# [[apodex-11-mini-trial]]. Loadable for comparison; do NOT wire it
# to Claude Code / pi / kimi-code tool loops.
#
# textturn tag leak 0/15 CLEAN (BEATS aquila, which leaked in real
# use); fabrication 3/5 no-tools but 0/10 tools-bound
# KLayout +ref 4/8 <- bottom of field (fable-711 7/8, glm-flash 13/16,
# qwen36-35b 15/16), and 2 failures were an INVENTED method
# `PolygonWithProperties.is_polygon` (absent from klayout
# 0.30.9 AND from the ref we supplied) + a 16k RUNAWAY
# office 0/9 unaided -> 6/9 WITH ref (vision 2/2 — mmproj works).
# RECALL defect, NOT a BigBang-style ceiling.
# R5 n=3 complete 4/5/4 of 6 (NOISE) but `interrupt_replan`
# LOOPED 3/3 — THE discriminating task
# ([[grill-round5-agentic-loops]]). This is the reject.
# speed ~94 tok/s
#
# --tensor-split 45,55 inherited from aquila and AUDITED, not assumed: a
# 5-point sweep showed 50,50 is +1.3% (95.8 vs 94.6 tok/s) and 55,45 /
# 58,42 / 61,39 ALL OOM on the same 860.98 MiB alloc — the mmproj loads
# wholly onto CUDA0 and is not placed by --tensor-split. 45,55 KEPT for
# the bigger CUDA0 vision margin. See [[mmproj-caps-cuda0-tensor-split]];
# the lopsided 29%/63% GPU util it produces is COSMETIC, not a bug.
cmd: |
/usr/bin/env
CUDA_DEVICE_ORDER=PCI_BUS_ID
${llama_bin}
-m /mnt/ssd1700/models/apodex/apodex_Apodex-1.1-mini-Q4_K_M.gguf
--mmproj /mnt/ssd1700/models/apodex/mmproj-apodex_Apodex-1.1-mini-bf16.gguf
--alias apodex
--jinja --chat-template-file chat-templates/apodex-1.1-256k.jinja
--reasoning-format deepseek
-ngl 99 -c 262144 -fa on
--tensor-split 45,55
-b 2048 -ub 512
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0
--cache-type-k q8_0 --cache-type-v q8_0
--host 127.0.0.1 --port 9181 --parallel 1
proxy: http://127.0.0.1:9181