Repository Issues

microsoft/MInference

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

View on GitHub
Stars
 (1,226 stars)
Forks
 (81 forks)
Indexed issues
 (0 indexed issues)
open beginner issues
 (0 open beginner issues)
Latest indexed
Aug 9, 2026
Last GitHub push
Mar 9, 2026
Contributing guide
No contributing guide
Code of conduct
Code of conduct
Dominant language
Python
PR merge metrics
 (No merged PRs in 30d)
Beginner labels
No beginner labels indexed

Issues

0 indexed issues

No indexed issues found for this repository.