GQLSA: Grouped-Query Latent Sparse Attention — A hardware-native attention mechanism combining latent compression, grouped-query sharing, and block-sparse attention for linear O(T) complexity.
machine-learning deep-learning language-modeling pytorch artificial-intelligence transformer attention-mechanism research-paper efficient-transformers linear-complexity sparse-attention llm grouped-query-attention latent-compression hardware-native
-
Updated
Sep 8, 2026 - Python