This run (2×RTX A6000, GDDR6 ~768 GB/s each) — the same hybrid
qwen3_5_moe (Gated-DeltaNet × 256-expert MoE VLM) that limps on edge silicon
simply works on discrete CUDA:
-ngl 99) decodes
correct tokens at temp 0 (the Xe3 lm_head corruption is a Vulkan kernel bug, not a model property), and the
CLIP/mmproj vision encoder runs clean (no DeviceLost) — UI localization lands within ~2 px.Why this comparison matters. NUC16 measures what a fanless edge box can do today; this run measures the same agent with the serving bottleneck removed — isolating model capability from serving physics. Score deltas between the two runs bound the quality cost of edge serving (quant + latency-induced behavior differences), while the step-latency/energy gap quantifies its price.