vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

93.4k
Stars
+30.7k
Gained
48.9%
Growth
Python
Language

🎭 Best For

⚖️ Compare With

🏷️ Topics & Ecosystem

amd blackwell cuda deepseek deepseek-v3 gpt gpt-oss inference kimi llama llm llm-serving model-serving moe openai pytorch qwen qwen3 tpu transformer

📊 Activity

Latest commit: 2026-10-08. Over the past 330 days, this repository gained 30.7k stars (+48.9% growth). Activity data is based on daily RepoPi snapshots of the GitHub repository.