Logo
Explore Help
Register Sign In
TBNilles/qwen3.5-122B-A10B-on-spark
Watch 1
Star 0
Fork 0
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
18 Commits 1 Branch 0 Tags
1d0f2e1af3c4953783af554a07110738bc891d6b
Commit Graph

6 Commits

Author SHA1 Message Date
ent d3688cc833 load speed: persist compile cache at $HF_HOME/.vllm_cache (was wiped each container boot, -42s/boot); propagate --load-format to mtp_serve; document startup time 2026-06-24 16:32:02 +10:00
ent df8981d54c VALIDATED defaults: gpu-mem 0.82 + seqs 3 (0.88+ over-subscribes/swaps); README corrected with measured KV pool (457k tok, 1.74x) + concurrency curve 2026-06-24 14:30:56 +10:00
ent dabc866507 tune for 1 orchestrator + up to 3 subagents: gpu-mem 0.88 (~15GB reserve), seqs=4, ctx 262144; document use case + tradeoffs 2026-06-24 13:31:36 +10:00
ent 6797a9f6dd default max-num-seqs=4 (>=2 with headroom; ~5 fit full native ctx each) 2026-06-24 13:21:52 +10:00
ent 7dd6a74d38 tune defaults for max context + KV depth: gpu-mem 0.89 (reserve ~14GB), ctx 262144, max-num-seqs 1, decouple max-batched-tokens (8192) 2026-06-24 13:18:32 +10:00
ent 60bf1b7b02 qwen3.5-122B-A10B on DGX Spark: vLLM + DFlash + dense-bandwidth stack, one-shot installer 2026-06-24 13:02:35 +10:00
Powered by Gitea Version: 1.27.0+dev-308-g6f4027a6be Page: 11ms Template: 1ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API