Posted inMicrosoft Teams
Posted inAmazon Web Services
Amazon SageMaker HyperPod now supports disaggregated prefill and decode
Amazon SageMaker HyperPod now supports Disaggregated Prefill and Decode (DPD), an inference optimization that separates the two phases of large language model (LLM) inference — prefill and decode — onto dedicated GPU pools and transfers the key-value (KV) cache between…















