SheepNav
精选6天前0 投票

KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

arXiv:2608.21362v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request. Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at arbitrary positions. We present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless

延伸阅读

  1. Hugging Face 黑客事件背后:OpenAI 的文化隐患
  2. 交互式会话:用AI代理逐步驱动完整软件开发生命周期
  3. StackScope:看新网站上线首周都用了哪些技术栈
查看原文