Swarm complete
Mission outcome
The swarm identified a composite disaggregated hosting topology that can beat a well-tuned monolithic vLLM* baseline on cost per token (ESTIMATE ~1.25–2.5× depending on SLO and prefix reuse), without claiming new GPU measurements in this environment.
Recommended method (primary)
Hypothesis sh_lphBfIh_oTE2 (promoted): heterogeneous prefill/decode fleets (e.g., H100-class prefill + A100-class FP8 KV decode pool), NIXL/LMCache MultiConnector P/D handoff, and break-even–gated external KV admission (load LMCache/Mooncake tier only when prefix length ≥ L*). Mechanism: phase decoupling + SKU right-sizing + optional cheap storage for reused prefixes (Splitwise, DistServe, LMCache, py-kvcache).
Fair baseline
vLLM* = vLLM continuous batching + PagedAttention + Sarathi-style chunked/stall-free scheduling + FP8 KV + APC (Sarathi-Serve, FP8 KV blog). Config-only wins (FP8 alone) are hygiene, not mission winners.
What does NOT beat vLLM* on default high-goodput chat
- FP8-only (sh_Q5ZdrV1rxF5u): part of vLLM*.
- FrugalGPT cascade (sh_ddcnRBuw2Qzp): multi-model routing vs API $; weak same-engine novelty.
- CPU/edge (alt-scout survey): no beat path at 512/128 high λ.
- Scale-to-zero burst (sh_ARBATA9U5VwO): wins only vs over-provisioned always-on GPU, not rubric anchor λ≈400 rps.
Deliverables
- report.md (sar_cxqve8Z-QQjJ) — verified conditional pass (verify/verify-result.json)
- design/prototype-composite-pd.md (sar_IXxGu1CfZNVV) — build plan using vLLM production-stack P/D
- Surveys: kv-memory, disaggregated, batching-speculative, disagg-composition, alternative-stacks; ranking/cost-per-token-rubric.md
- 8 hypotheses, all critiqued
Prototype / next phase
Human GPU repo required for measured A/B vs vLLM*; see report Phase 1–3. Optional follow-up: disaggregated async speculative decoding (SwiftSpec). No request_phase_change filed (no repository in swarm).