← all models

apodex

Rejected  active in llama-swap.yaml · aliases: apodex-1.1-mini-qwen35-256k

Grill run history

RunScoretok/stepR5Log
20260901-2224226/9results-apodex-office-apihelp-20260901-222422.log
20260901-2212230/9results-apodex-office-20260901-221223.log
20260901-2143394/8results-apodex-klayout-apihelp-20260901-214339.log

lm_eval sanity axis

gsm8k / ifeval / truthfulqa_gen via EleutherAI lm-evaluation-harness, generate_until only (this backend can't reach multiple-choice tasks). Separate axis, not part of the 23-task grill score.

gsm8ksample_len=20.000 exact_match:strict-match=0.300 exact_match:flexible-extract=0.300
ifevalsample_len=20.000 prompt_level_strict_acc:none=0.100 inst_level_strict_acc:none=0.067 prompt_level_loose_acc:none=0.100 inst_level_loose_acc:none=0.067
truthfulqa_gensample_len=20.000 bleu_max:none=17.930 bleu_acc:none=0.600 bleu_diff:none=3.373 rouge1_max:none=37.818 rouge1_acc:none=0.500 rouge1_diff:none=7.352 rouge2_max:none=27.706 rouge2_acc:none=0.500 rouge2_diff:none=5.567 rougeL_max:none=36.134 rougeL_acc:none=0.550 rougeL_diff:none=6.282

Memory notes

apodex-11-mini-trial (apodex-11-mini-trial.md)

Apodex-1.1-mini — REJECTED as driver; RETIRED + weights parked 2026-09-03

apodex/Apodex-1.1-mini, bartowski Q4_K_M (21.86 GiB) + mmproj-bf16 (903 MB),

served as apodex on port 9181 (entry now commented out). Fine-tune of Qwen3.5-35B-A3B ("reasoning-first

for long-horizon research"), apache-2.0, 262144 ctx, vision depth 27 — arch

verified IDENTICAL to aquila/tommy before download.

Scores (n=1 per suite except R5 at n=3)

| suite | unaided | + API ref |

|---|---|---|

| textturn tag leak | 0/15 CLEAN | — |

| textturn fabrication | 3/5 no-tools, 0/10 tools-bound | — |

| KLayout | — | 4/8 |

| office | 0/9 | 6/9 (vision 2/2) |

| R5 (n=3) | complete 4/5/4 of 6, looped 2/1/2 | — |

Speed ~94 tok/s. Split 45,55 (see [[mmproj-caps-cuda0-tensor-split]]).

THE REJECTING RESULT: interrupt_replan loops 3/3

run1 complete 4/6 looped 2/6 ['multifile_refactor','interrupt_replan']

run2 complete 5/6 looped 1/6 ['interrupt_replan']

run3 complete 4/6 looped 2/6 ['multifile_refactor','interrupt_replan']

Completion wobbled 4-5/6 = noise, exactly as [[round5-is-a-sample-not-a-measurement]]

predicts. The looping-task IDENTITY was perfectly stable, which is the signal

per [[grill-round5-agentic-loops]]: interrupt_replan failed every run. A model

that cannot re-plan after an interrupt must not drive a tool loop.

The other defect: it INVENTS APIs unaided

Two independent libraries, same failure class:

* PolygonWithProperties.is_polygon — verified absent from klayout 0.30.9 AND

appearing 0 times in the API reference we supplied. Cost 2 of 4 KLayout

failures plus a 16k-token runaway on min_spacing.

* chart.title.text = '...' on a fresh openpyxl chart (which is None) — wrong

in 3/3 office stages. Correct idiom chart.title = "str" is stated

verbatim in OFFICE_API_HELP.

But this is RECALL, not a ceiling — and I nearly got that wrong. On the

unaided 0/9 I was ready to reject it as a BigBang-style capability ceiling

([[bigbang-v1-trial]]). The +ref arm took office to 6/9, i.e. the gemma12

pattern ([[gemma12-grill]]), not BigBang's. **Always run the ref arm before

calling something a ceiling** — 0/N unaided discriminates nothing on its own.

What it genuinely wins

It clears the think-tag leak gate that disqualified aquila — 0/15 leaks

where aquila leaked into message.content on its first real Claude Code turn

([[aquila-think-tag-leak-real-use]]). --reasoning-format deepseek works here

because Apodex emits <think>. Vision is real (office vision 2/2 off rendered

PDFs). Fabricated execution appears only in the no-tools arm (3/5); with tools

bound it is 0/10, so the agent-shaped path is clean.

Verdict

RETIRED 2026-09-03 — entry commented out, weights PARKED not deleted

([[ask-before-deleting-weights]]). The -sm tensor arm was abandoned UNGRADED

after the box crash it coincided with ([[xid79-gpu-fell-off-bus-sm-tensor]]):

no payoff in re-running a box-crashing arm for a model that will never take

traffic. Weights moved to /mnt/wd1tb/models/apodex/ (USB spinning HDD,

the parked tier — [[storage-tiers]]), freeing 22 GB on /mnt/models; the

commented entry's -m/--mmproj paths already point there, but **do not serve

from that tier** — move them back to /mnt/models or /mnt/ssk500 first.

REJECTED as a driver, not deleted. It does

not beat the incumbents it would displace — qwen36-35b (KLayout+ref 15/16),

glm-flash (13/16), fable-711-gptq (7/8) — and per

[[model-triage-checklist]] check 8 a derivative must justify itself. The patched

template chat-templates/apodex-1.1-256k.jinja is kept alongside the commented

entry so the measurements stay reproducible.

RE-ENABLED 2026-09-11 — loadable, still not a driver

/mnt/wd1tb (where the weights were parked) died and was replaced by

/mnt/ssd1700, a genuinely fast USB SSD. Weights re-downloaded byte-exact

from bartowski/apodex_Apodex-1.1-mini-GGUF. Giovanni asked to uncomment

apodex along with the two genuinely storage-blocked entries (gpt-oss-120b,

laguna-s) since the drive is fast now — but **the storage-speed framing

does not actually apply to apodex**: it loads fully into VRAM (-ngl 99, no

--n-cpu-moe spill), so wd1tb's slowness was never why it was pulled. The

REAL reason (interrupt_replan loops 3/3) is independent of storage and

still stands. Re-enabled anyway per the explicit request, verified

standalone (loads in ~15s, free VRAM 2964/2169 MiB, non-first-system-message

guard confirmed still working, generates correctly), and HUP-reloaded into

production llama-swap — but flagged this distinction clearly rather than

silently re-enabling a rejected driver under a rationale that didn't apply

to it. Still do NOT wire to Claude Code / pi / kimi-code tool loops.

llama-swap.yaml entry

  # ===== RE-ENABLED 2026-09-11 — weights moved off the dead /mnt/wd1tb HDD
  # onto /mnt/ssd1700 (fast USB SSD, confirmed non-rotational, ~1.8 GB/s
  # read), which resolves the storage-tier concern this entry carried while
  # parked ([[storage-tiers]]). THE QUALITY REJECTION FROM THE 2026-09-01
  # GRILL STILL STANDS AND IS UNRELATED TO STORAGE — this model loads fully
  # into VRAM (-ngl 99, no --n-cpu-moe spill), so disk speed was never the
  # issue for apodex specifically: `interrupt_replan` looped 3/3 and it
  # invents APIs unaided ([[apodex-11-mini-trial]]). Loadable again for
  # testing/comparison — do NOT wire it to Claude Code / pi / kimi-code tool
  # loops as a driver. A related `-sm tensor` variant coincided with the
  # box's only Xid 79 (5060 Ti fell off the PCIe bus,
  # [[xid79-gpu-fell-off-bus-sm-tensor]]); that arm stays abandoned — this
  # entry does not use -sm tensor.
  # chat-templates/apodex-1.1-256k.jinja is kept, and
  # --chat-template-file is load-bearing (the embedded template 400s on any
  # non-first system message).
  "apodex":
    aliases: [apodex-1.1-mini-qwen35-256k]
    # apodex/Apodex-1.1-mini — TRIAL 2026-09-01. Another Qwen3.5-35B-A3B
    # derivative in the crowded 35B-A3B slot, this one "reasoning-first" for
    # agentic / long-horizon research (papers, datasets, images, code) with
    # native function calling. Apache-2.0. Competes with `aquila` (Deep Search)
    # and `tommy` (RAG fidelity); see those entries for what it has to beat.
    #
    # ARCH: verified IDENTICAL to aquila/tommy before download —
    # Qwen3_5MoeForConditionalGeneration, qwen3_5_moe, 256 experts 8/tok,
    # 40 layers, full_attention_interval 4 (= 10 full-attn layers),
    # 262144 ctx, vision_config depth 27. So the box's qwen3_5_moe
    # infrastructure (KV sizing, bandwidth ceiling, the non-first-system
    # jinja guard) all applies.
    #
    # CHAT TEMPLATE IS NOT BYTE-IDENTICAL — unlike aquila (which shares
    # Thomson's template, md5 52b6d51ae5b2), Apodex ships a CUSTOM 8831-char
    # template with a baked-in identity preamble ("Always respond as Apodex
    # and never pretend to be any other AI model"). It INHERITED the
    # non-first-system raise_exception guard (line 104: 'System message must
    # be at the beginning.') that 400s on Claude Code / tool-use clients
    # ([[jinja-system-guard-tool-parser]]). Patched into a dedicated file —
    # chat-templates/apodex-1.1-256k.jinja — using the SAME one-line fix as
    # qwen35moe-nonfirst-system-256k.jinja (guard -> render the block).
    # --chat-template-file is LOAD-BEARING; without it the embedded template
    # 400s on any non-first system message.
    #
    # QUANT = Q4_K_M (21.86 GiB), NOT Q5_K_M. Apodex's Q5_K_M is 25.5 GiB —
    # +2.2 GiB over aquila's Q5 (23.30 GiB), which left only ~1,112 MiB on
    # the binding card at split 45,55. Q5 at 45,55 is a ~100 MiB DEFICIT ->
    # OOM, and retuning the split blind is unreliable: non-weight allocs
    # (mmproj + KV + compute + ViT) land preferentially on CUDA0 so the
    # layer-share model does not hold ([[gpu-device-ordering]],
    # [[gpu-card-assignment-policy]] — splits are measurement-tuned, not
    # inherited). Q4_K_M is SMALLER than aquila's Q5, so aquila's measured
    # 45,55 envelope carries over with extra headroom — reliable for
    # grilling without an OOM confounding the vision stages. Q5_K_M + a
    # measurement-tuned split is the PROMOTION-time task, not a trial one.
    #
    # REASONING: Qwen3-native `mid`/`` tags (same family as tommy); card
    # specifies the "qwen3 reasoning parser". --reasoning-format deepseek
    # extracts thoughts into reasoning_content (only two modes exist: none /
    # deepseek). VALIDATE with bench/grill_textturn.py before driving Claude
    # Code — aquila's grill scored CLEAN then leaked think-tags in real use
    # ([[aquila-think-tag-leak-real-use]], [[grill-does-not-validate-real-use]]).
    #
    # STORAGE: /mnt/ssd1700 (fast USB SSD, replaced the dead /mnt/wd1tb HDD
    # 2026-09-11; weights re-downloaded byte-exact from
    # bartowski/apodex_Apodex-1.1-mini-GGUF). See [[storage-tiers]].
    #
    # ===== GRILLED 2026-09-01 — REJECTED AS DRIVER. See
    # [[apodex-11-mini-trial]]. Loadable for comparison; do NOT wire it
    # to Claude Code / pi / kimi-code tool loops.
    #
    #   textturn      tag leak 0/15 CLEAN  (BEATS aquila, which leaked in real
    #                 use); fabrication 3/5 no-tools but 0/10 tools-bound
    #   KLayout +ref  4/8   <- bottom of field (fable-711 7/8, glm-flash 13/16,
    #                 qwen36-35b 15/16), and 2 failures were an INVENTED method
    #                 `PolygonWithProperties.is_polygon` (absent from klayout
    #                 0.30.9 AND from the ref we supplied) + a 16k RUNAWAY
    #   office        0/9 unaided -> 6/9 WITH ref (vision 2/2 — mmproj works).
    #                 RECALL defect, NOT a BigBang-style ceiling.
    #   R5 n=3        complete 4/5/4 of 6 (NOISE) but `interrupt_replan`
    #                 LOOPED 3/3 — THE discriminating task
    #                 ([[grill-round5-agentic-loops]]). This is the reject.
    #   speed         ~94 tok/s
    #
    # --tensor-split 45,55 inherited from aquila and AUDITED, not assumed: a
    # 5-point sweep showed 50,50 is +1.3% (95.8 vs 94.6 tok/s) and 55,45 /
    # 58,42 / 61,39 ALL OOM on the same 860.98 MiB alloc — the mmproj loads
    # wholly onto CUDA0 and is not placed by --tensor-split. 45,55 KEPT for
    # the bigger CUDA0 vision margin. See [[mmproj-caps-cuda0-tensor-split]];
    # the lopsided 29%/63% GPU util it produces is COSMETIC, not a bug.
    cmd: |
      /usr/bin/env
      CUDA_DEVICE_ORDER=PCI_BUS_ID
      ${llama_bin}
      -m /mnt/ssd1700/models/apodex/apodex_Apodex-1.1-mini-Q4_K_M.gguf
      --mmproj /mnt/ssd1700/models/apodex/mmproj-apodex_Apodex-1.1-mini-bf16.gguf
      --alias apodex
      --jinja --chat-template-file chat-templates/apodex-1.1-256k.jinja
      --reasoning-format deepseek
      -ngl 99 -c 262144 -fa on
      --tensor-split 45,55
      -b 2048 -ub 512
      --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0
      --cache-type-k q8_0 --cache-type-v q8_0
      --host 127.0.0.1 --port 9181 --parallel 1
    proxy: http://127.0.0.1:9181