docs: refactor README to a consistent neutral technical-documentation voice (drop second-person/session-narrative); add conc_workloads.py

This commit is contained in:
ent
2026-06-24 14:54:40 +10:00
parent cadfa61184
commit 314ae39000
2 changed files with 283 additions and 138 deletions
+138 -134
View File
@@ -1,58 +1,58 @@
# qwen3.5-122B-A10B-on-spark
[`Qwen3.5-122B-A10B`](https://huggingface.co/Intel/Qwen3.5-122B-A10B-int4-AutoRound)
(hybrid GDN + mamba + 128-expert MoE, ~10B active) running on a single
(hybrid GDN + mamba + 128-expert MoE, ~10B active) on a single
**NVIDIA DGX Spark** (GB10 / SM121, 128 GB / 119 GiB unified) under **vLLM**, with
**[DFlash](https://modal.com/blog/spec-is-all-u-need) block-diffusion speculative
decode** and an optional **dense-bandwidth patch stack** — measured end-to-end,
with a per-token bandwidth model that explains every number.
with a per-token bandwidth model behind the numbers.
**Status:** Working end-to-end, one-shot install. On real Hermes-agent
tool-call turns, **DFlash decode reaches a median ~81 tok/s on GB10**
**~2× the native MTP-2 head (~40 tok/s)** on the same workload, and above
**Status:** working end-to-end, one-shot install. On real agent tool-call turns,
**DFlash decode reaches a median ~81 tok/s on GB10** about **2× the native
MTP-2 head (~40 tok/s)** on the same workload, and above
[**albond's**](https://github.com/albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4)
fully-patched MTP recipe (51.58 tok/s end-to-end) — **the recipe this repo
gratefully builds on** (see [Building on the albond recipe](#building-on-the-albond-recipe)).
DFlash's acceptance is task-dependent (it block-drafts 12 tokens in one parallel
forward), so the win is largest on structured/tool-call/code traffic and
collapses to parity on open-ended prose.
fully-patched MTP recipe (51.58 tok/s end-to-end), the recipe this project builds
on (see [Building on the albond recipe](#building-on-the-albond-recipe)). DFlash
acceptance is task-dependent (it block-drafts 12 tokens in one parallel forward),
so the gain is largest on structured / tool-call / code traffic and falls to
parity on open-ended prose.
A separate **dense-bandwidth stack** (hybrid INT4+FP8 shared experts + int8
lm-head) adds **+28 % to no-spec / base decode** (28.2 → 36.0 tok/s) but, by the
amortization law below, washes out to ~null on high-acceptance agent traffic —
so it's a lever for *base / low-acceptance* serving, not for the agent path.
lm-head) adds **+28 % to no-spec / base decode** (28.2 → 36.0 tok/s) but, per the
[amortization law](#the-amortization-law) below, falls to ~nil on high-acceptance
agent traffic — it is a lever for base / low-acceptance serving, not the agent path.
- **Engine:** [`vLLM`](https://github.com/vllm-project/vllm) 0.23, sm121 build with the DFlash PRs, via the prebuilt image `ghcr.io/aeon-7/aeon-vllm-ultimate:2026-06-18-v0.23.0-dflashfix`. No host build — the four runtime patches in [`runtime/`](runtime/) are applied at serve time.
- **Target:** [`Intel/Qwen3.5-122B-A10B-int4-AutoRound`](https://huggingface.co/Intel/Qwen3.5-122B-A10B-int4-AutoRound) — INT4 (AutoRound/GPTQ) routed experts + attention, BF16 shared experts/embeddings/head, ~62 GiB. (Safetensors, *not* GGUF — vLLM serves HF checkpoints directly.)
- **Drafter:** [`z-lab/Qwen3.5-122B-A10B-DFlash`](https://huggingface.co/z-lab/Qwen3.5-122B-A10B-DFlash) — 0.8B / 6-layer non-causal block-diffusion drafter (block 16), shares the target's `embed_tokens` + `lm_head`, ~1.6 GiB.
- **Target:** [`Intel/Qwen3.5-122B-A10B-int4-AutoRound`](https://huggingface.co/Intel/Qwen3.5-122B-A10B-int4-AutoRound) — INT4 (AutoRound/GPTQ) routed experts + attention, BF16 shared experts / embeddings / head, ~62 GiB. Safetensors, *not* GGUF — vLLM serves the HF checkpoint directly.
- **Drafter:** [`z-lab/Qwen3.5-122B-A10B-DFlash`](https://huggingface.co/z-lab/Qwen3.5-122B-A10B-DFlash) — 0.8B / 6-layer non-causal block-diffusion drafter (block 16), sharing the target's `embed_tokens` + `lm_head`, ~1.6 GiB.
- **Hardware:** NVIDIA DGX Spark, GB10, SM121, 128 GB LPDDR5X unified (~119 GiB usable), ~273 GB/s.
## Quick start
On a DGX Spark with Docker + the NVIDIA container runtime:
On a DGX Spark with Docker and the NVIDIA container runtime:
```bash
curl -sSL https://raw.githubusercontent.com/Entrpi/qwen3.5-122B-A10B-on-spark/main/install.sh | bash -s -- --start
```
That one command:
This command:
1. Verifies the host (aarch64, GB10 / SM121, Docker GPU access, free disk).
2. Pulls the sm121 vLLM image (~40 GiB, one-time).
3. Downloads the INT4 target (~62 GiB) + DFlash drafter (~1.6 GiB) into the HF cache.
3. Downloads the INT4 target (~62 GiB) and the DFlash drafter (~1.6 GiB) into the HF cache.
4. Starts the `dflash` profile on `:8000`, waits until READY, and runs the
"capital of France" smoke test (asserts "Paris").
**Already have the model?** Skip the 62 GiB download:
To reuse a checkpoint that is already present and skip the ~62 GiB download:
```bash
# point at a checkpoint dir you already have (mounted read-only at /model):
# point at an existing checkpoint directory (mounted read-only at /model):
./install.sh --start --model-dir /path/to/Qwen3.5-122B-A10B-int4-AutoRound
# or reuse an existing HF cache (download becomes a no-op if already present):
# or reuse an existing HF cache (the download becomes a no-op if already present):
./install.sh --start --hf-home /mnt/big/hf
```
Preview without running: `... | bash -s -- --help`.
Preview without running: append `--help`.
## Hardware requirements
@@ -61,67 +61,68 @@ Preview without running: `... | bash -s -- --help`.
| Validated on | NVIDIA DGX Spark (GB10, SM121, 128 GB / 119 GiB unified) |
| Likely to work | other Blackwell with `--force` (untested) |
| Runtime | Docker + NVIDIA container runtime (`docker run --gpus all`) |
| Disk | ≥ 75 GiB free (image + weights); ≥ 150 GiB if `--build-hybrid` |
| Disk | ≥ 75 GiB free (image + weights); ≥ 150 GiB with `--build-hybrid` |
| OS | aarch64 Linux (Grace) |
| Memory | 128 GB / 119 GiB unified is enough for the model + DFlash drafter + KV @ 16k |
| Memory | 128 GB / 119 GiB unified holds the model + DFlash drafter + KV |
GB10 is detected via `nvidia-smi --query-gpu=compute_cap` returning `12.1`;
anything else needs `--force`.
GB10 is detected via `nvidia-smi --query-gpu=compute_cap` returning `12.1`; other
hardware requires `--force`.
### Memory & context (defaults tuned for: 1 orchestrator + up to 3 subagents)
### Memory and context (defaults tuned for up to 3 concurrent streams)
Target use case — **one user**, one main orchestrator thread (~100 k context
average, safe up to the model's 262 144 max) that dispatches **up to 3 subagents**
(<100 k each). With prefix-caching off (DFlash requires it) the orchestrator is
*not* a persistent stream — when it dispatches it has ended its turn and is idle
until the subagents return — so the concurrent peak is the **3 subagents**, not 4.
The intended deployment is a single user with up to **3 concurrent decode
streams** (for example, a main agent thread plus subagents). A single stream can
use the full **262 144** context; the defaults assume concurrent streams stay
under ~100 k tokens each. The default operating point is one stream with no
contention; two to three streams are additive (see the per-workload table below).
These numbers are **measured on the box**, not estimated (an earlier naive
"24 KiB/token → ~1.3 M pool" projection was wrong — the per-sequence GDN/mamba
state and the activation reserve cost far more than the attention KV alone). At
the shipped defaults (`gpu-mem 0.82`, `ctx 262144`, `seqs 3`):
Attention KV is small for this model (~24 KiB/token — 12 of 48 layers are full
attention, with GQA `num_key_value_heads=2`), but the per-sequence GDN/mamba state
and the activation reserve dominate the footprint, so the usable pool is far
smaller than a KV-only estimate suggests. The values below are measured on the
target hardware at the shipped defaults (`gpu-mem 0.82`, `ctx 262144`, `seqs 3`):
| | measured |
| Measurement | Value |
|---|---|
| free memory at READY | **~14 GiB** (responsive, no swap) |
| Free memory at READY | **~14 GiB** (responsive, no swap) |
| GPU KV cache pool | **456,664 tokens** |
| max concurrency @ full 262 144 | **1.74×** |
| concurrency 1 → 2 → 3 (prose) | per-request **26.5 → 18.8 → 17.4** tok/s · aggregate **26.5 → 36.8 → 51.6** |
| Max concurrency at full 262 144 | **1.74×** |
| Concurrency 1 → 2 → 3 (prose) | per-stream **26.5 → 18.8 → 17.4** tok/s · aggregate **26.5 → 36.8 → 51.6** |
Your realistic peak — 3 subagents <100 k (≈ <300 k tokens) — fits the 456 k pool
with margin, and the orchestrator can still run to the full 262 k when it's the
active stream. **Why not higher:** `gpu-mem 0.880.89` over-subscribes this box —
the static footprint left only ~5 GiB free, it swapped, and requests hung. `0.82`
is the validated sweet spot (~14 GiB headroom). Defaults (override via flags/env):
A typical load — three streams under ~100 k each (≈ <300 k tokens) — fits the
456 k pool with margin, and a single stream can still reach the full 262 144
context. At `gpu-mem` 0.880.89 the static footprint leaves only ~5 GiB free; the
host then swaps and requests stall. `0.82` is the validated value (~14 GiB free).
Defaults (override via flags or environment variables):
| Flag / env | Default | Note |
|---|---|---|
| `--gpu-mem` / `GPU_MEM` | **0.82** | ~14 GiB free (validated). 0.88+ over-subscribes swap → hangs. |
| `--ctx` / `CTX` (`MAX_MODEL_LEN`) | **262144** | model native max — orchestrator safe at any length (costs only KV-pool sizing; graph range is tied to `max-batched-tokens`) |
| `--max-num-seqs` / `MAX_NUM_SEQS` | **3** | the concurrent-subagent peak; pool ÷ ctx ≈ 1.74× full-context, ample for <100 k subagents |
| `--max-batched-tokens` / `MAX_BATCHED_TOKENS` | **8192** | chunked-prefill chunk kept **below** ctx so a long prefill doesn't batch all at once |
| `--gpu-mem` / `GPU_MEM` | **0.82** | ~14 GiB free (validated); 0.88+ over-subscribes and swaps |
| `--ctx` / `CTX` (`MAX_MODEL_LEN`) | **262144** | model native max; a single stream can reach any length up to this. Costs only KV-pool sizing — the CUDA-graph compile range tracks `max-batched-tokens`, not `ctx` |
| `--max-num-seqs` / `MAX_NUM_SEQS` | **3** | concurrent-stream cap; the pool holds 1.74× a full-262 k context, ample for <100 k streams |
| `--max-batched-tokens` / `MAX_BATCHED_TOKENS` | **8192** | chunked-prefill chunk, kept **below** `ctx` so a long prefill does not batch all at once |
**Single-stream is the default operating point** (DFlash speculative decode is a
single-stream lever; at 1 active stream there's no contention — ~81 tok/s on agent
turns). 23 concurrent is additive: per-request tok/s eases down as the decode
batch grows (more routed-expert traffic/step) while aggregate rises — the cost is
throughput-vs-latency, not safety. For a 4th simultaneous stream raise
`--max-num-seqs 4` (it'll queue, not crash) or trim `--ctx`.
The default operating point is a single stream (no contention; ~81 tok/s on agent
turns). Two to three concurrent streams are additive: per-stream throughput
decreases as the decode batch grows (more routed-expert traffic per step) while
aggregate throughput rises — a throughput-versus-latency trade, not a safety
limit. A fourth simultaneous stream requires `--max-num-seqs 4` (additional
streams queue rather than fail) or a smaller `--ctx`.
> Unified-memory OOM **hard-freezes** the box, and vLLM's profiler can undershoot
> peak by a couple GB — always bring the server up under
> Unified-memory OOM hard-freezes the host, and the vLLM profiler can undershoot
> peak by a couple of GB. Bring the server up under
> [`scripts/monitor.sh`](scripts/monitor.sh) (OOM auto-kill guard) the first time
> at a new `gpu-mem`/`ctx`.
> at any new `gpu-mem` or `ctx`.
## What you get — profiles
## Profiles
Pick with `--profile`:
Selected with `--profile`:
| Profile | Stack | Best for | Measured |
|---|---|---|---|
| **`dflash`** *(default)* | INT4 + DFlash n=12 | agents / tool-calls / code | **~81 tok/s** Hermes · 53.7 albond-bench |
| `dense` | hybrid INT4+FP8 + int8 lm-head + DFlash n=12 | base / low-accept serving | 36.0 base (+28%) · 59.0 albond-bench |
| `base` | plain INT4, no spec | airtight baseline | 28.2 tok/s c=1 |
| `dense` | hybrid INT4+FP8 + int8 lm-head + DFlash n=12 | base / low-acceptance serving | 36.0 base (+28%) · 59.0 albond-bench |
| `base` | plain INT4, no speculative decode | airtight baseline | 28.2 tok/s c=1 |
| `mtp` | INT4 + native MTP-2 head | comparison | ~40 tok/s Hermes |
The server is OpenAI-compatible (`/v1/chat/completions` with tool calls + SSE,
@@ -129,11 +130,11 @@ The server is OpenAI-compatible (`/v1/chat/completions` with tool calls + SSE,
## Benchmarks
All single-stream (c=1), temperature 0, GB10. "Hermes" = regenerating the next
All single-stream (c=1), temperature 0, GB10. "Hermes" regenerates the next
assistant turn over 10 real conversations from a live agent's `state.db` (73 %
tool-calls); "albond-bench" = albond's own end-to-end harness (completion_tokens
/ total wallclock incl. prefill, 5 prompts, run-1 discarded — directly
comparable to his published 51.58).
tool-calls); "albond-bench" is albond's end-to-end harness (completion tokens /
total wallclock incl. prefill, 5 prompts, run 1 discarded — directly comparable to
the published 51.58).
### DFlash vs MTP, same harness, unpatched
@@ -145,16 +146,17 @@ comparable to his published 51.58).
| **Hermes, real turns (8.3)** | — | **39.9** | **~81** |
| albond-bench e2e (6.5) | — | — | **53.7** |
MTP-2 drafts 2 tokens *sequentially* (acceptance caps at ~3); DFlash block-drafts
12 in **one parallel forward**, so on predictable/agent traffic it accepts 511
and pulls ~2× ahead. They tie only on low-acceptance prose. **53.7 unpatched
already clears albond's fully-patched MTP (51.58)** under his own method.
MTP-2 drafts 2 tokens *sequentially*, so acceptance caps at ~3; DFlash block-drafts
12 in **one parallel forward**, accepting 511 on predictable / agent traffic and
running ~2× ahead. The two tie only on low-acceptance prose. Unpatched DFlash
(53.7) already clears albond's fully-patched MTP (51.58) under the same end-to-end
method.
### The dense-bandwidth stack (`dense` profile)
### Dense-bandwidth stack (`dense` profile)
Two independent always-on levers, ported to vLLM 0.23 as runtime patches:
hybrid INT4+FP8 (BF16 shared experts → calibrated FP8) and int8 lm-head (the
248 320-row vocab projection → int8 w8a16 GEMV, ~2× the bf16 read).
Two independent always-on levers, ported to vLLM 0.23 as runtime patches: hybrid
INT4+FP8 (BF16 shared experts → calibrated FP8) and int8 lm-head (the 248 320-row
vocab projection → int8 w8a16 GEMV, ~2× the bf16 read).
| Config | base (acc 1) | DFlash spec, albond-bench (acc 6.4) | Hermes (acc 8.3) |
|---|---|---|---|
@@ -165,60 +167,59 @@ hybrid INT4+FP8 (BF16 shared experts → calibrated FP8) and int8 lm-head (the
## Building on the albond recipe
This repo stands on **[albond's DGX-Spark Qwen3.5-122B recipe](https://github.com/albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4)** —
This project builds on
**[albond's DGX-Spark Qwen3.5-122B recipe](https://github.com/albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4)** —
the first working high-throughput recipe for this model on Spark, and the
reference we measured everything against. On eugr's vLLM 0.19.1 fork, albond
reference used for the comparisons here. On eugr's vLLM 0.19.1 fork, albond
established:
- the **rebuilt hybrid INT4+FP8 checkpoint** (BF16 dense → calibrated FP8),
- the **INT8 lm-head** patch (the single biggest dense-bandwidth lever),
- the **INT8 lm-head** patch (the single largest dense-bandwidth lever),
- **MTP-2** native speculative decode, and
- the **end-to-end benchmark methodology** (completion_tokens / total wallclock,
incl. prefill) we report against — reproduced verbatim in
[`scripts/bench_albond.py`](scripts/bench_albond.py).
- the **end-to-end benchmark methodology** (completion tokens / total wallclock,
incl. prefill), reproduced verbatim in [`scripts/bench_albond.py`](scripts/bench_albond.py).
We gratefully build on all of it. What this repo adds is **carrying that recipe
forward to the latest vLLM and a stronger drafter**, and getting the community
patches to stack together there:
This project carries that recipe forward to the latest vLLM and a stronger
drafter, and composes the community patches there:
1. **Forward-ported the dense levers to vLLM 0.23.** albond's patches target the
0.19 fork and don't drop in cleanly. The hybrid-FP8 dispatch had to be
re-expressed against 0.23's `maybe_update_config(model_name, hf_config=…)`
quant-config hook; the INT8 lm-head needed a from-scratch integration (the
prior port zeroed the lm-head weight — which corrupts the **DFlash-shared**
head → garbage — and looped per-row for batch>4, which is *slower* under spec;
both fixed in [`patch_int8_lmhead_v3.py`](runtime/patch_int8_lmhead_v3.py)).
2. **Swapped MTP-2 for the DFlash block-diffusion drafter** and got it running on
the hybrid 122B in vLLM (the [KV-unify fix](runtime/patch_unify2.py)). DFlash
block-drafts 12 tokens in one parallel forward vs MTP's sequential head
(acceptance-capped at ~3), so it pulls ~2× ahead on agent/code traffic.
3. **Composed the dense levers *with* DFlash** instead of MTP.
1. **Dense levers forward-ported to vLLM 0.23.** albond's patches target the 0.19
fork and do not apply cleanly. The hybrid-FP8 dispatch is re-expressed against
0.23's `maybe_update_config(model_name, hf_config=…)` quant-config hook; the
INT8 lm-head is reintegrated from scratch (the prior port zeroed the lm-head
weight — which corrupts the **DFlash-shared** head — and looped per-row for
batch > 4, which is slower under speculative decode; both are fixed in
[`patch_int8_lmhead_v3.py`](runtime/patch_int8_lmhead_v3.py)).
2. **MTP-2 replaced by the DFlash block-diffusion drafter,** running on the hybrid
122B in vLLM via the [KV-unify fix](runtime/patch_unify2.py). DFlash block-drafts
12 tokens in one parallel forward versus MTP's sequential head (acceptance-capped
at ~3), running ~2× ahead on agent / code traffic.
3. **Dense levers composed with DFlash** rather than MTP.
### Where we land — albond's own end-to-end method, same hardware class
### Comparison — albond's end-to-end method, same hardware class
| Stack | Spec | Dense patches | e2e tok/s |
|---|---|---|---|
| **albond** (published) | MTP-2 | hybrid-FP8 + INT8 lm-head + PR#38325 | 51.58 |
| this repo`dflash` | DFlash n=12 | *none* | **53.7** (+4%) |
| this repo`dense` | DFlash n=12 | hybrid-FP8 + INT8 lm-head | **59.0** (+14%) |
| this project`dflash` | DFlash n=12 | *none* | **53.7** (+4%) |
| this project`dense` | DFlash n=12 | hybrid-FP8 + INT8 lm-head | **59.0** (+14%) |
On the **real agent workload** (decode-only, regenerating live tool-call turns),
DFlash's parallel block-drafting pulls further ahead of MTP's sequential head:
On the real agent workload (decode-only, regenerating live tool-call turns),
DFlash's parallel block-drafting runs further ahead of MTP's sequential head:
**~81 vs ~40 tok/s**.
> **Stated plainly:** albond's 51.58 is his published figure on his stack
> (vLLM 0.19 + MTP); our figures are on this stack (vLLM 0.23 + DFlash). Both use
> the **same e2e harness** on the **same hardware class** (DGX Spark / GB10) — a
> fair best-on-each-stack comparison, not a single-variable controlled run. Note
> that **unpatched DFlash (53.7) already clears albond's fully-patched MTP**, so
> the dense stack is upside on top of the drafter swap, not the source of the win.
> **Note on the comparison.** The 51.58 is the published figure for albond's stack
> (vLLM 0.19 + MTP); the figures here are for this stack (vLLM 0.23 + DFlash). Both
> use the same end-to-end harness on the same hardware class (DGX Spark / GB10) — a
> best-on-each-stack comparison, not a single-variable controlled run. Unpatched
> DFlash (53.7) already exceeds the fully-patched MTP result, so the dense stack is
> additional headroom rather than the source of the difference.
## The amortization law
The dense levers cut **always-on** weight reads (shared experts + lm-head, read
The dense levers cut **always-on** weight reads (shared experts and lm-head, read
every token). Under speculative decode the verify forward reads those weights
**once and amortizes them across the accepted block**, so the gain shrinks as
acceptance rises — monotonically, across the whole curve:
acceptance rises — monotonically, across the curve:
```
dense stack uplift: +28% (base, accept 1)
@@ -226,25 +227,26 @@ dense stack uplift: +28% (base, accept 1)
→ ~0% (Hermes, accept ~8.3)
```
**Consequence:** for the **agent path (`dflash`)**, DFlash's own high acceptance
already saturates the dense levers — its remaining bottleneck is *routed-expert*
verify-batch reads, which no dense-weight quant touches. For **base / low-accept
serving (`dense`)**, the stack is a real +28 %. See [`docs/FINDINGS.md`](docs/FINDINGS.md).
For the agent path (`dflash`), DFlash's high acceptance already saturates the
dense levers; the remaining bottleneck is *routed-expert* verify-batch reads,
which no dense-weight quantization touches. For base / low-acceptance serving
(`dense`), the stack is a real +28 %. Full derivation in
[`docs/FINDINGS.md`](docs/FINDINGS.md).
## Under the hood: the four runtime patches
vLLM is unmodified on disk; [`runtime/serve.sh`](runtime/serve.sh) edits the
installed package in-place before `vllm serve` (idempotent, sentinel-guarded):
| Patch | What it does | Needed by |
| Patch | Effect | Required by |
|---|---|---|
| [`patch_unify2.py`](runtime/patch_unify2.py) | scale-block KV-cache **unify** so the hybrid GDN+mamba target absorbs the drafter's attention spec (the original assert can't); + `--no-enable-prefix-caching` routes to the no-hash-assert coordinator | **DFlash** (any spec profile) |
| [`patch_unify2.py`](runtime/patch_unify2.py) | scale-block KV-cache **unify** so the hybrid GDN+mamba target absorbs the drafter's attention spec (the upstream assert cannot); paired with `--no-enable-prefix-caching` to route to the no-hash-assert coordinator | DFlash (any spec profile) |
| [`patch_inc_hybrid.py`](runtime/patch_inc_hybrid.py) | adds an `INCConfig.maybe_update_config` override that detects FP8 dense layers in the hybrid checkpoint and dispatches `Fp8LinearMethod` for `shared_expert` | `dense` |
| [`patch_int8_lmhead_v3.py`](runtime/patch_int8_lmhead_v3.py) | replaces the lm-head matmul in `_get_logits` with a batched int8 w8a16 Triton GEMV (keeps bf16 weight for the shared drafter) | `dense` |
| [`patch_fla_shmem.py`](runtime/patch_fla_shmem.py) | lets the FLA GDN chunk kernels use big tiles on sm121's 99 KiB shmem (prefill/TTFT only; harmless) | always (free) |
| [`patch_int8_lmhead_v3.py`](runtime/patch_int8_lmhead_v3.py) | replaces the lm-head matmul in `_get_logits` with a batched int8 w8a16 Triton GEMV (keeps the bf16 weight for the shared drafter) | `dense` |
| [`patch_fla_shmem.py`](runtime/patch_fla_shmem.py) | allows the FLA GDN chunk kernels to use large tiles on sm121's 99 KiB shmem (prefill / TTFT only; harmless) | always (free) |
Why DFlash needs the unify patch at all, why the drafter must run `FLASH_ATTN`
(non-causal), and the full vLLM-vs-SGLANG dead-end history are in
The rationale for the unify patch, the requirement that the drafter run
`FLASH_ATTN` (non-causal), and the full vLLM-vs-SGLang history are in
[`docs/FINDINGS.md`](docs/FINDINGS.md).
## Repo layout
@@ -261,15 +263,16 @@ runtime/ Mounted read-only at /host inside the container:
scripts/ Host-side helpers:
monitor.sh Container-startup monitor with OOM auto-kill guard
bench_decode.py Decode-only tok/s (excludes TTFT)
bench_albond.py albond's e2e method (comparable to his 51.58)
bench_albond.py albond's e2e method (comparable to 51.58)
hermes_bench.py Real agent turns from ~/.hermes/state.db
run_bank.sh prose/code/counting/hermes bank on any server
conc_workloads.py Per-workload concurrency sweep (decode-only)
run_bank.sh prose / code / counting / hermes bank
tools/
build-hybrid-checkpoint.py Build the hybrid INT4+FP8 ckpt (for --build-hybrid)
inspect_ckpt.py Which layers are INT4 vs BF16 vs FP8
build-hybrid-checkpoint.py Builds the hybrid INT4+FP8 checkpoint (--build-hybrid)
inspect_ckpt.py Reports which layers are INT4 / BF16 / FP8
validate_*.py Standalone correctness checks for the patches
docs/
FINDINGS.md The full investigation, methodology, and the
FINDINGS.md Full investigation, methodology, and the
amortization-law derivation
```
@@ -279,39 +282,40 @@ docs/
# default agent path (DFlash) + smoke test:
./install.sh --start
# the dense stack (build the hybrid ckpt once, ~20 min, then serve):
# the dense stack (build the hybrid checkpoint once, ~20 min, then serve):
./install.sh --build-hybrid
./install.sh --start --profile dense
# benches (run on the host against the server; need: pip install requests):
# benchmarks (run on the host against the server; requires: pip install requests):
python3 scripts/bench_decode.py --base-url http://127.0.0.1:8000 --model qwen \
--prompt "Write a detailed essay about the history of tea."
python3 scripts/bench_albond.py http://127.0.0.1:8000 "dflash" # e2e, vs 51.58
python3 scripts/hermes_bench.py --base-url http://127.0.0.1:8000 # real agent turns
python3 scripts/conc_workloads.py --base-url http://127.0.0.1:8000 # concurrency sweep
# MTP comparison:
./install.sh --start --profile mtp
```
## How this fits with related work
## Related work
| Piece | Role |
| Project | Role |
|---|---|
| [`vLLM`](https://github.com/vllm-project/vllm) | the inference engine; this repo serves Qwen3.5 + DFlash on it, unmodified-on-disk |
| [`Intel/...int4-AutoRound`](https://huggingface.co/Intel/Qwen3.5-122B-A10B-int4-AutoRound) · [`z-lab/...DFlash`](https://huggingface.co/z-lab/Qwen3.5-122B-A10B-DFlash) | the target + drafter weights |
| [`albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4`](https://github.com/albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4) | the MTP + hybrid-FP8 + int8-lmhead recipe we benchmarked against and ported the dense levers from |
| [`vLLM`](https://github.com/vllm-project/vllm) | the inference engine; served unmodified-on-disk |
| [`Intel/...int4-AutoRound`](https://huggingface.co/Intel/Qwen3.5-122B-A10B-int4-AutoRound) · [`z-lab/...DFlash`](https://huggingface.co/z-lab/Qwen3.5-122B-A10B-DFlash) | the target and drafter weights |
| [`albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4`](https://github.com/albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4) | the MTP + hybrid-FP8 + int8-lm-head recipe; benchmark reference and source of the ported dense levers |
| [`Entrpi/ds4-on-spark`](https://github.com/Entrpi/ds4-on-spark) | sibling repo, same hardware, different model (DeepSeek-V4-Flash via ds4) |
| [Modal: *Speculative decoding is all you need*](https://modal.com/blog/spec-is-all-u-need) | the DFlash block-diffusion drafter and the task-dependent-acceptance framing |
## Acknowledgements
- **[`albond`](https://github.com/albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4) — the foundation this builds on.** The hybrid INT4+FP8 checkpoint, the INT8 lm-head, MTP-2, and the end-to-end benchmark methodology. This repo is a grateful forward-port and recomposition of that recipe onto vLLM 0.23 + DFlash.
- **[`albond`](https://github.com/albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4)** — the foundation this project builds on: the hybrid INT4+FP8 checkpoint, the INT8 lm-head, MTP-2, and the end-to-end benchmark methodology, forward-ported and recomposed here onto vLLM 0.23 + DFlash.
- [`z-lab`](https://huggingface.co/z-lab) / [Modal](https://modal.com/blog/spec-is-all-u-need) — the DFlash drafter and the block-diffusion speculative-decode work.
- [`Intel/AutoRound`](https://huggingface.co/Intel) — the INT4 target quantization.
- [`vLLM`](https://github.com/vllm-project/vllm) and the AEON sm121 image maintainers — the engine and the DFlash-enabled GB10 build.
- [`eugr/spark-vllm-docker`](https://github.com/eugr/spark-vllm-docker) — the vLLM-on-Spark base that albond's recipe (and much of this ecosystem) started from.
- [`eugr/spark-vllm-docker`](https://github.com/eugr/spark-vllm-docker) — the vLLM-on-Spark base much of this ecosystem started from.
## License
MIT — see [LICENSE](LICENSE). The patches are original; vendored third-party
files (`tools/build-hybrid-checkpoint.py`) retain their upstream attribution.
MIT — see [LICENSE](LICENSE). The patches are original; vendored third-party files
(`tools/build-hybrid-checkpoint.py`) retain their upstream attribution.
+141
View File
@@ -0,0 +1,141 @@
#!/usr/bin/env python3
"""Per-workload concurrency sweep. Reports DECODE-only tok/s (streaming, excludes
prefill/TTFT) for prose / code / agentic at K concurrent streams. "agentic" uses
real tool-call contexts from ~/.hermes/state.db (the workload DFlash wins biggest
on); prose/code use fixed prompts. Per-stream = mean decode rate of the K streams;
aggregate = sum (concurrent total throughput).
python3 conc_workloads.py --base-url http://127.0.0.1:8000 --levels 1,2,3
"""
import argparse
import concurrent.futures as cf
import json
import sqlite3
import statistics
import sys
import time
import urllib.request
DB = "/home/ent/.hermes/state.db"
PROSE = [{"role": "user", "content":
"Write a detailed, flowing essay about the history and cultural "
"significance of tea across civilizations."}]
CODE = [{"role": "user", "content":
"Implement a complete red-black tree in Python: insert, delete, search, "
"rotations, with type hints and docstrings. Output only the code."}]
def stream_chat(base, model, messages, max_tokens, timeout=600):
body = json.dumps({"model": model, "messages": messages, "max_tokens": max_tokens,
"temperature": 0.0, "stream": True,
"stream_options": {"include_usage": True}}).encode()
req = urllib.request.Request(base.rstrip("/") + "/v1/chat/completions", data=body,
headers={"Content-Type": "application/json"})
t_first = t_last = None
toks = 0
with urllib.request.urlopen(req, timeout=timeout) as r:
for raw in r:
ln = raw.decode("utf-8", "replace").strip()
if not ln.startswith("data:"):
continue
d = ln[5:].strip()
if d == "[DONE]":
break
try:
o = json.loads(d)
except json.JSONDecodeError:
continue
u = o.get("usage")
if u and u.get("completion_tokens"):
toks = u["completion_tokens"]
ch = o.get("choices") or []
if ch and (ch[0].get("delta") or {}).get("content"):
now = time.perf_counter()
if t_first is None:
t_first = now
t_last = now
return toks, t_first, t_last
def decode_rate(res):
toks, tf, tl = res
return (toks - 1) / (tl - tf) if (toks and tf and tl and tl > tf) else 0.0
def load_agentic(min_msgs=8, char_budget=60000):
"""Pick a representative real tool-call context (truncate to ~budget)."""
db = sqlite3.connect(DB)
cur = db.cursor()
for sid, sysp in cur.execute(
"SELECT id, system_prompt FROM sessions WHERE message_count >= ? "
"ORDER BY started_at DESC LIMIT 40", (min_msgs,)):
rows = list(cur.execute(
"SELECT role, content, tool_calls, tool_call_id FROM messages "
"WHERE session_id=? ORDER BY id", (sid,)))
last = max((i for i, r in enumerate(rows) if r[0] == "assistant"), default=None)
if not last:
continue
msgs = []
for role, content, tc, tcid in rows[:last]:
content = content or ""
if role == "user":
msgs.append({"role": "user", "content": content})
elif role == "assistant":
m = {"role": "assistant", "content": content}
if tc:
try:
j = json.loads(tc)
if isinstance(j, list) and j:
m["tool_calls"] = j
except Exception:
pass
msgs.append(m)
elif role == "tool":
msgs.append({"role": "tool", "content": content, "tool_call_id": tcid or "call_0"})
if not msgs:
continue
sysm = [{"role": "system", "content": sysp}] if sysp else []
while sum(len(json.dumps(m)) for m in msgs) > char_budget and len(msgs) > 1:
msgs.pop(0)
approx = sum(len(json.dumps(m)) for m in (sysm + msgs)) // 4
return sysm + msgs, approx
return None, 0
def sweep(base, model, label, messages, max_tokens, levels):
stream_chat(base, model, messages, 16) # warm
row = [label]
for K in levels:
with cf.ThreadPoolExecutor(max_workers=K) as ex:
res = list(ex.map(lambda _: stream_chat(base, model, messages, max_tokens), range(K)))
rates = [decode_rate(r) for r in res]
row.append((statistics.mean(rates), sum(rates)))
print(f" {label:8s} K={K}: per-stream {statistics.mean(rates):5.1f} tok/s "
f"(min {min(rates):.1f}/max {max(rates):.1f}) aggregate {sum(rates):6.1f}")
return row
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--base-url", default="http://127.0.0.1:8000")
ap.add_argument("--model", default="qwen")
ap.add_argument("--max-tokens", type=int, default=200)
ap.add_argument("--levels", default="1,2,3")
args = ap.parse_args()
levels = [int(x) for x in args.levels.split(",")]
agentic, approx = load_agentic()
print(f"=== concurrency sweep (decode-only tok/s, streaming) levels={levels} ===")
if agentic:
print(f" (agentic context ~{approx} tokens, real tool-call turn)")
sweep(args.base_url, args.model, "prose", PROSE, args.max_tokens, levels)
sweep(args.base_url, args.model, "code", CODE, args.max_tokens, levels)
if agentic:
sweep(args.base_url, args.model, "agentic", agentic, args.max_tokens, levels)
else:
print(" (no Hermes DB found — skipped agentic)", file=sys.stderr)
if __name__ == "__main__":
main()