Logo
Explore Help
Register Sign In
TBNilles/qwen3.5-122B-A10B-on-spark
Watch 1
Star 0
Fork 0
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
12 Commits 1 Branch 0 Tags
efd80e8fe7d66c9919e0346c5edf3e8e514fe113
Commit Graph

9 Commits

Author SHA1 Message Date
ent efd80e8fe7 docs: dense profile downloads the prebuilt hybrid (no build step) 2026-06-24 15:08:13 +10:00
ent 314ae39000 docs: refactor README to a consistent neutral technical-documentation voice (drop second-person/session-narrative); add conc_workloads.py 2026-06-24 14:54:40 +10:00
ent df8981d54c VALIDATED defaults: gpu-mem 0.82 + seqs 3 (0.88+ over-subscribes/swaps); README corrected with measured KV pool (457k tok, 1.74x) + concurrency curve 2026-06-24 14:30:56 +10:00
ent dabc866507 tune for 1 orchestrator + up to 3 subagents: gpu-mem 0.88 (~15GB reserve), seqs=4, ctx 262144; document use case + tradeoffs 2026-06-24 13:31:36 +10:00
ent 6797a9f6dd default max-num-seqs=4 (>=2 with headroom; ~5 fit full native ctx each) 2026-06-24 13:21:52 +10:00
ent 7dd6a74d38 tune defaults for max context + KV depth: gpu-mem 0.89 (reserve ~14GB), ctx 262144, max-num-seqs 1, decouple max-batched-tokens (8192) 2026-06-24 13:18:32 +10:00
ent 70c91f94a7 docs: GB10 memory is 128 GB (119 GiB), not 128 GiB 2026-06-24 13:10:36 +10:00
ent dc9d40a1c4 docs: credit albond's recipe as the foundation + head-to-head comparison (59.0 vs 51.58 e2e, same method) 2026-06-24 13:08:52 +10:00
ent 60bf1b7b02 qwen3.5-122B-A10B on DGX Spark: vLLM + DFlash + dense-bandwidth stack, one-shot installer 2026-06-24 13:02:35 +10:00
Powered by Gitea Version: 1.27.0+dev-308-g6f4027a6be Page: 18ms Template: 2ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API