First open-source implementation of a Google DeepMind paper — 4-7x KV cache compression for LLMs at inference time with near-zero quality loss. Benchmarked on 6 models up to 70B params.
← All work
AI · ML infrastructure
TurboQuant
First open-source implementation of a Google DeepMind paper — 4-7x KV cache compression for LLMs at inference time with near-zero quality loss. Benchmarked on 6 models up to 70B params.