vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

89.8k
Stars
+27.1k
Gained
43.2%
Growth
Python
Language

🎭 Best For

⚖️ Compare With

🏷️ Topics & Ecosystem

amd blackwell cuda deepseek deepseek-v3 gpt gpt-oss inference kimi llama llm llm-serving model-serving moe openai pytorch qwen qwen3 tpu transformer

📊 Activity

Latest commit: 2026-08-24. Over the past 285 days, this repository gained 27.1k stars (+43.2% growth). Activity data is based on daily RepoPi snapshots of the GitHub repository.