People
Publications
Resources
Contact
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference
X. Tang
,
F. Meng
,
P. Tang
,
Y. Wang
,
D. Yin
,
X. Sun
,
M. Zhang
March 2026
Code
Type
Conference paper
Publication
ASPLOS ‘26: Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, pp. 2048–2062
Cite
×