RebaseKV: Optimistic Cache Alignment for Multi-Agent LLM Serving
A training-free architecture that decodes from inherited, semantically relevant KV states while the exact prompt-conditioned cache is materialized asynchronously, then rebases onto the fresh cache.
- Up to 3.9× TTFT reduction across GSM8K, MMLU, and HumanEval.
- Preserves answer-level accuracy and restores role-specific programming behavior from 5% to 97%.
Manuscript under review