Commit Graph

15 Commits

Author SHA1 Message Date
Entrpi 72badaf39a serve: enable OpenAI auto tool-calling (qwen3_xml) + reasoning split by default
A forum user hit HTTP 400 "'auto' tool choice requires --enable-auto-tool-choice
and --tool-call-parser to be set" — the server shipped without tool-calling
enabled, which is exactly what agent/tool workloads (the whole point) need.

- serve.sh + mtp_serve.sh now pass --enable-auto-tool-choice --tool-call-parser
  qwen3_xml --reasoning-parser qwen3 by default (overridable/disable via
  TOOL_PARSER= / REASONING_PARSER=).
- Parser matters: Qwen3.5 emits the XML tool format
  (<function=name><parameter=k>v</parameter></function>), so hermes returns 200
  but with EMPTY tool_calls (the call lands in content). qwen3_xml parses it.
  Verified on the box: tool_choice=auto -> tool_calls=[get_weather {city:Paris}];
  plain chat unaffected; <think> -> reasoning_content.
- README: document tool-calling defaults + that the server binds 0.0.0.0
  (LAN-reachable at http://<spark-ip>:8000; no auth).
2026-06-24 23:08:31 +10:00
Entrpi 2cb5033fb8 kv: recover dense pool 376k->427k (+13%) at the same 0.82 headroom
Reclaims over-reserved memory rather than spending headroom, so it honors
the conservative 0.82 / 14 GB-OS-reserve choice (still ~15 GiB free under
peak load) while closing most of the gap to the dflash pool (456k):

- int8 lm-head: build int8 on first warmup forward, then free the dead bf16
  copy (~1.4 GiB) + empty_cache() so the block lands in the KV pool. Safe:
  the DFlash drafter shares the one int8 lm_head module (verified one int8
  build; drafter ckpt has no lm_head/embed) and tie_word_embeddings=False,
  so bf16 has no reader (bias fallback = int8-GEMV + add-bias).
  SPARK_KEEP_BF16_LMHEAD=1 restores keep-bf16.
- VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 in serve.sh/mtp_serve.sh:
  returns the CUDA-graph over-estimate (~0.6 GiB; actual capture ~0.14 GiB
  from the wide 0.82 headroom). Overridable with =1.
- scripts/stress_mem.sh: drive the concurrency bank while sampling host RAM
  to prove a gpu-mem setting survives peak load without swapping.

Validated: pool 426,610 tokens, coherent output, drafter acceptance 4-12,
no new swap, throughput unchanged. max-batched-tokens 8192->4096 tested as
a KV lever and rejected (only +2.5k tokens, costs prefill speed).
2026-06-24 17:41:07 +10:00
ent 1d0f2e1af3 docs: fix Reproducing — --start defaults to dense now; dflash is the explicit opt-in 2026-06-24 16:46:22 +10:00
ent 4e5c9552fe docs: refactor README so dense is unambiguously the default and the hosted hybrid (bleysg/...int4-fp8-hybrid) is the primary/chosen checkpoint (title link, model bullet, quick start, related-work) 2026-06-24 16:45:57 +10:00
ent d3688cc833 load speed: persist compile cache at $HF_HOME/.vllm_cache (was wiped each container boot, -42s/boot); propagate --load-format to mtp_serve; document startup time 2026-06-24 16:32:02 +10:00
ent a0b661b016 make dense the default profile (validated download-only at 0.82/seqs3); README: per-workload concurrency table + dense/dflash pool tradeoff (376k vs 456k) 2026-06-24 15:40:44 +10:00
ent efd80e8fe7 docs: dense profile downloads the prebuilt hybrid (no build step) 2026-06-24 15:08:13 +10:00
ent 314ae39000 docs: refactor README to a consistent neutral technical-documentation voice (drop second-person/session-narrative); add conc_workloads.py 2026-06-24 14:54:40 +10:00
ent df8981d54c VALIDATED defaults: gpu-mem 0.82 + seqs 3 (0.88+ over-subscribes/swaps); README corrected with measured KV pool (457k tok, 1.74x) + concurrency curve 2026-06-24 14:30:56 +10:00
ent dabc866507 tune for 1 orchestrator + up to 3 subagents: gpu-mem 0.88 (~15GB reserve), seqs=4, ctx 262144; document use case + tradeoffs 2026-06-24 13:31:36 +10:00
ent 6797a9f6dd default max-num-seqs=4 (>=2 with headroom; ~5 fit full native ctx each) 2026-06-24 13:21:52 +10:00
ent 7dd6a74d38 tune defaults for max context + KV depth: gpu-mem 0.89 (reserve ~14GB), ctx 262144, max-num-seqs 1, decouple max-batched-tokens (8192) 2026-06-24 13:18:32 +10:00
ent 70c91f94a7 docs: GB10 memory is 128 GB (119 GiB), not 128 GiB 2026-06-24 13:10:36 +10:00
ent dc9d40a1c4 docs: credit albond's recipe as the foundation + head-to-head comparison (59.0 vs 51.58 e2e, same method) 2026-06-24 13:08:52 +10:00
ent 60bf1b7b02 qwen3.5-122B-A10B on DGX Spark: vLLM + DFlash + dense-bandwidth stack, one-shot installer 2026-06-24 13:02:35 +10:00