Milestones
List view
0.2.0: speculative decoding grows from a shipped feature into a subsystem (rejection sampling #512, shared primitives #443), GLM5.2 decode serving converges on the vLLM DP8/EP8 reference (#542 + perf backlog, #590 DSpark×prefix-cache), batch-invariant Qwen3 inference reaches its e2e bar (#435), and Prometheus observability stabilizes on the Qwen3 line (#602: real counters + Grafana dashboard). Plus frontend error-surface polish (#294, #584).
No due date•8/19 issues closed