命令:

cd chapter2/local_llm_serving

../../.venv/bin/python run_experiment.py \
  --output runs/exp2-1-qwen3-0.6b-20260919T134503Z

运行结果:

official_complete: true
model: qwen3:0.6b
model_digest: 7df6b6e09427a769808717c0a93cadc4ae99ed4eb8bf5ca557c90846becea435
tool_case_passed: true
mean_decode_throughput: 268.60 tok/s
cache_hit_mean_ttft: 88.03 ms
cache_miss_mean_ttft: 497.75 ms
cache_miss_over_hit: 5.65x
credential_scan_passed: true
cost: 0 USD

模型首轮同时调用 get_current_time 和 get_current_temperature,两个工具通过线程池并行执行;第二轮消费工具结果后正常终止。六项工具调用门禁全部通过,5 组缓存配对中命中请求均快于未命中请求。

回归测试:

uv run --locked --extra ch2 --with pytest \
  python -m pytest \
  chapter2/local_llm_serving/test_run_experiment.py \
  chapter2/local_llm_serving/test_benchmark.py -q
5 passed, 1 warning in 2.13s

限制:缓存请求没有生成可见正文,因此 runner 报告的 TTFT 使用总请求时间作为回退值;它能反映前缀重算的延迟差异,但不是严格意义上的首个可见 token 时间。最终回答还把温哥华经度写成了 -123.12°E,规范表达应为约 123.12°W,该语义错误不在正式工具调用门禁范围内。