Back to Repository Index

vllm-project/vllm

PythonApache License 2.090,641 stars

A high-throughput and memory-efficient inference and serving engine for LLMs

#amd#blackwell#cuda#deepseek#deepseek-v3#gpt#gpt-oss#inference#kimi#llama#llm#llm-serving#model-serving#moe#openai#pytorch#qwen#qwen3#tpu#transformerllm

Repository Info

Stars90,641
Forks21,511
Watchers90,641
Open Issues7,291
LicenseApache License 2.0
Last Pushed2h ago