serve: enable OpenAI auto tool-calling (qwen3_xml) + reasoning split by default

A forum user hit HTTP 400 "'auto' tool choice requires --enable-auto-tool-choice
and --tool-call-parser to be set" — the server shipped without tool-calling
enabled, which is exactly what agent/tool workloads (the whole point) need.

- serve.sh + mtp_serve.sh now pass --enable-auto-tool-choice --tool-call-parser
  qwen3_xml --reasoning-parser qwen3 by default (overridable/disable via
  TOOL_PARSER= / REASONING_PARSER=).
- Parser matters: Qwen3.5 emits the XML tool format
  (<function=name><parameter=k>v</parameter></function>), so hermes returns 200
  but with EMPTY tool_calls (the call lands in content). qwen3_xml parses it.
  Verified on the box: tool_choice=auto -> tool_calls=[get_weather {city:Paris}];
  plain chat unaffected; <think> -> reasoning_content.
- README: document tool-calling defaults + that the server binds 0.0.0.0
  (LAN-reachable at http://<spark-ip>:8000; no auth).
This commit is contained in:
Entrpi
2026-06-24 23:08:31 +10:00
parent 2cb5033fb8
commit 72badaf39a
3 changed files with 35 additions and 0 deletions
+7
View File
@@ -15,6 +15,12 @@ LOAD_FORMAT="${LOAD_FORMAT:-fastsafetensors}"
PORT="${PORT:-8000}"
# Reclaim the CUDA-graph memory over-estimate to KV (see serve.sh). Set =1 to restore.
export VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS="${VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS:-0}"
# Auto tool-calling + reasoning split (see serve.sh). Off via TOOL_PARSER="" etc.
TOOL_PARSER="${TOOL_PARSER:-qwen3_xml}"
REASONING_PARSER="${REASONING_PARSER:-qwen3}"
TOOL_ARG=()
[ -n "$TOOL_PARSER" ] && TOOL_ARG+=(--enable-auto-tool-choice --tool-call-parser "$TOOL_PARSER")
[ -n "$REASONING_PARSER" ] && TOOL_ARG+=(--reasoning-parser "$REASONING_PARSER")
echo "[mtp] qwen3_5_mtp — backend=$BACKEND, num_speculative_tokens=$NSPEC, model=$MODEL"
exec vllm serve "$MODEL" \
--served-model-name qwen \
@@ -28,4 +34,5 @@ exec vllm serve "$MODEL" \
--trust-remote-code \
--load-format "$LOAD_FORMAT" \
--attention-backend "$BACKEND" \
"${TOOL_ARG[@]}" \
--speculative-config "{\"method\":\"qwen3_5_mtp\",\"num_speculative_tokens\":$NSPEC,\"model\":\"$MODEL\"}"