Research Scientist
CurrentMy research mainly focuses on developing training and inference algorithms for large language models. I am particularly interested in efficient techniques leveraging different forms of sparsity and low precision. Projects:• Training compute-efficient models using sparsely-activated mixture-of-experts layers.• Sparse attention techniques for fast autoregressive transformer inference.• Quantisation and low-precision numerical formats for model compression.Publications:• SparQ Attention: Bandwidth-Efficient LLM Inference (ICML 2024)