Commit Graph

19 Commits

Author SHA1 Message Date
Entrpi 2cb5033fb8 kv: recover dense pool 376k->427k (+13%) at the same 0.82 headroom
Reclaims over-reserved memory rather than spending headroom, so it honors
the conservative 0.82 / 14 GB-OS-reserve choice (still ~15 GiB free under
peak load) while closing most of the gap to the dflash pool (456k):

- int8 lm-head: build int8 on first warmup forward, then free the dead bf16
  copy (~1.4 GiB) + empty_cache() so the block lands in the KV pool. Safe:
  the DFlash drafter shares the one int8 lm_head module (verified one int8
  build; drafter ckpt has no lm_head/embed) and tie_word_embeddings=False,
  so bf16 has no reader (bias fallback = int8-GEMV + add-bias).
  SPARK_KEEP_BF16_LMHEAD=1 restores keep-bf16.
- VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 in serve.sh/mtp_serve.sh:
  returns the CUDA-graph over-estimate (~0.6 GiB; actual capture ~0.14 GiB
  from the wide 0.82 headroom). Overridable with =1.
- scripts/stress_mem.sh: drive the concurrency bank while sampling host RAM
  to prove a gpu-mem setting survives peak load without swapping.

Validated: pool 426,610 tokens, coherent output, drafter acceptance 4-12,
no new swap, throughput unchanged. max-batched-tokens 8192->4096 tested as
a KV lever and rejected (only +2.5k tokens, costs prefill speed).
2026-06-24 17:41:07 +10:00
ent 1d0f2e1af3 docs: fix Reproducing — --start defaults to dense now; dflash is the explicit opt-in 2026-06-24 16:46:22 +10:00
ent 4e5c9552fe docs: refactor README so dense is unambiguously the default and the hosted hybrid (bleysg/...int4-fp8-hybrid) is the primary/chosen checkpoint (title link, model bullet, quick start, related-work) 2026-06-24 16:45:57 +10:00
ent d3688cc833 load speed: persist compile cache at $HF_HOME/.vllm_cache (was wiped each container boot, -42s/boot); propagate --load-format to mtp_serve; document startup time 2026-06-24 16:32:02 +10:00
ent bdb57d776d load speed: --load-format fastsafetensors (read straight to device, no mmap) — weight load 463s -> 32s (14.5x), time-to-READY ~12min -> ~3min on GB10. Falls back to nogds automatically. 2026-06-24 16:30:33 +10:00
ent 9b7be16bee int8 lm-head: output bf16 (match stock lm-head dtype) instead of fp32 — halves the logits buffer that was inflating vLLM's profiled activation reserve and shrinking the dense KV pool 2026-06-24 15:53:58 +10:00
ent a0b661b016 make dense the default profile (validated download-only at 0.82/seqs3); README: per-workload concurrency table + dense/dflash pool tradeoff (376k vs 456k) 2026-06-24 15:40:44 +10:00
ent efd80e8fe7 docs: dense profile downloads the prebuilt hybrid (no build step) 2026-06-24 15:08:13 +10:00
ent 17bad6c39b dense profile: download prebuilt hybrid checkpoint (HYBRID_REPO) instead of requiring --build-hybrid; serve as HF repo id (no /model mount); --build-hybrid kept as opt-in 2026-06-24 15:07:57 +10:00
ent 314ae39000 docs: refactor README to a consistent neutral technical-documentation voice (drop second-person/session-narrative); add conc_workloads.py 2026-06-24 14:54:40 +10:00
ent cadfa61184 install.sh: fix false-negative GPU-access probe (--entrypoint true) 2026-06-24 14:44:29 +10:00
ent df8981d54c VALIDATED defaults: gpu-mem 0.82 + seqs 3 (0.88+ over-subscribes/swaps); README corrected with measured KV pool (457k tok, 1.74x) + concurrency curve 2026-06-24 14:30:56 +10:00
ent dabc866507 tune for 1 orchestrator + up to 3 subagents: gpu-mem 0.88 (~15GB reserve), seqs=4, ctx 262144; document use case + tradeoffs 2026-06-24 13:31:36 +10:00
ent 6797a9f6dd default max-num-seqs=4 (>=2 with headroom; ~5 fit full native ctx each) 2026-06-24 13:21:52 +10:00
ent 7dd6a74d38 tune defaults for max context + KV depth: gpu-mem 0.89 (reserve ~14GB), ctx 262144, max-num-seqs 1, decouple max-batched-tokens (8192) 2026-06-24 13:18:32 +10:00
ent 70c91f94a7 docs: GB10 memory is 128 GB (119 GiB), not 128 GiB 2026-06-24 13:10:36 +10:00
ent dc9d40a1c4 docs: credit albond's recipe as the foundation + head-to-head comparison (59.0 vs 51.58 e2e, same method) 2026-06-24 13:08:52 +10:00
ent 884945ada4 install.sh: fix set-e exit from bare 'return' in build_hybrid guard 2026-06-24 13:06:03 +10:00
ent 60bf1b7b02 qwen3.5-122B-A10B on DGX Spark: vLLM + DFlash + dense-bandwidth stack, one-shot installer 2026-06-24 13:02:35 +10:00