TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference

Publication
ASPLOS ‘26: Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, pp. 2048–2062