KV (key-value) cache is the intermediate data a large language model stores as it generates text, so it doesn't recompute earlier tokens for every new one. As context windows and user sessions grow, KV cache balloons, and holding it all in costly DRAM strains data-center budgets. That is pushing chipmakers toward offloading it to faster NAND, which is why the term matters for memory-stock demand.