Skip to content
All comparisons
Comparisons8 min read

vLLM vs llama.cpp for serving models locally

Across eleven sources read on 30 September 2026, vLLM v0.30.0 and llama.cpp v0.5.0 serve the same open-weight models from opposite ends: vLLM turns concurrency into throughput and needs the model in accelerator memory, while llama.cpp runs one stream at a time on whatever hardware is already there.

  • a-vs-b
  • ai-libraries
  • autopilot
  • hermes

One of these engines turns an accelerator into throughput for many callers. The other runs a model on the machine that is already there. Both serve open-weight models behind an OpenAI-compatible API, and both are current: vLLM's latest tagged release is v0.30.0 (22 September 2026) and llama.cpp's is v0.5.0 (23 September 2026, with nightly builds running to b11265 on 29 September), all from the projects' own release pages.

Each side rests on eight sources, eleven in all, across five tiers: both projects' documentation and release notes, their repository figures as the GitHub REST API reported them on 30 September 2026 (92,960 stars for vllm-project/vllm, 129,897 for ggml-org/llama.cpp), two editorial comparisons from Red Hat Developer, a dated energy test in a public repository, a practitioner write-up with its own numbers and method, and one Hacker News thread per engine. The sources date from 31 March 2023 to 29 September 2026 and were all read on 30 September 2026. A software engine has no retailer page, so the buyer tier is empty here. Nothing was installed, run or benchmarked for this piece.

Consensus

Four sources set the two engines against each other: the Red Hat Developer articles of 30 September 2025 and 15 June 2026, a practitioner write-up on selfhostindex.com, and a dated energy test in the vreys-ai repository on GitHub. Three of them make the difference a function of concurrency, with vLLM's throughput growing as simultaneous requests rise while llama.cpp's stays flat, and the selfhostindex.com write-up puts the crossover at roughly four to eight sustained streams. The fourth, the vreys-ai test, ran both engines at a single stream on a Colab L4 and found them within one percent of each other in duration and GPU energy.

Confidence is strong on the shape, thin on the magnitude: every magnitude here came off somebody else's hardware, the editorial benchmark used one NVIDIA H200 and versions a year old (vLLM v0.10.0, llama.cpp b6100), the vreys-ai test one L4 at batch size one, and the selfhostindex.com figures are that site's own estimates on an RTX 4090. Neither project publishes a figure for a reader's own request pattern.

Recurring strengths

Throughput. vLLM's documentation states its own case: PagedAttention for key and value memory, continuous batching, chunked prefill, prefix caching, more than 200 supported model architectures on Hugging Face, and an OpenAI-compatible API server. The editorial benchmark measured what that adds up to on one NVIDIA H200: at peak load across one to 64 concurrent users, vLLM delivered more than 35 times the request throughput and more than 44 times the output tokens per second of llama.cpp (Red Hat Developer, 30 September 2025). At a concurrency of one, the same test had the two engines comparable.

Hardware. The llama.cpp README lists 16 backends, among them CUDA, HIP, Metal for Apple silicon, Vulkan, WebGPU, SYCL, OpenCL and Hexagon, and describes CPU plus GPU hybrid inference for models larger than the available VRAM, with no dependencies to install. vLLM's documentation lists NVIDIA and AMD GPUs and x86, ARM and PowerPC CPUs, with hardware plugins for Google TPUs, Intel Gaudi and Apple silicon among others: real breadth, arriving as plugins rather than as the core path.

Formats and licence. llama.cpp is where GGUF lives, and the Red Hat Developer article of 15 June 2026 calls the single-file format the de facto standard for local model distribution, with Ollama and LM Studio built on the engine. vLLM's documentation leads with FP8, MXFP8, MXFP4, NVFP4, INT8 and INT4, GPTQ and AWQ, and lists GGUF among its quantisation workflows. Both ship permissive licences, the GitHub REST API reporting Apache-2.0 for vllm-project/vllm and MIT for ggml-org/llama.cpp. Whether either licence suits a reader's own use is a legal question this piece does not answer.

Recurring complaints

The first cost of vLLM is where the model has to live. The selfhostindex.com write-up states it: the model must fit in GPU memory or be split across GPUs with tensor parallelism, and CUDA or ROCm is effectively required. The Red Hat Developer article of 15 June 2026 routes it the same way, sending vLLM to teams with data-centre GPUs and llama.cpp to the rest.

The second is operational weight, named by the selfhostindex.com write-up and the Red Hat Developer article of 15 June 2026: two of the four. vLLM is a Python service with a dependency chain of its own and a container as the recommended deployment, where llama.cpp ships as one binary with no dependencies, and the selfhostindex.com figures put the two within 10 to 15 percent of each other below four to eight sustained streams. The same write-up notes that vLLM preallocates about 90 percent of VRAM by default through its memory-utilisation setting, which reads as a leak in monitoring and is not one.

llama.cpp's cost mirrors its portability: throughput does not grow with the queue. The H200 benchmark found its throughput almost perfectly flat as load rose, with P99 time to first token rising exponentially, which that article attributes to the queuing model, and the vreys-ai limits section warns that llama.cpp's one-stream advantage narrows under high-concurrency serving. That complaint appears in the H200 benchmark, the vreys-ai limits section and the selfhostindex.com table, three of the four sources that look at load.

The second cost is the choice it hands back. The vreys-ai README states that with llama.cpp a reader must pick a GGUF precision and that the choice changes the question being answered, and the selfhostindex.com table lists Q2 to Q8, the K-quants and the IQ variants as the engine's own axis. vLLM hands that decision over less often, and what a llama.cpp user trades for it is a model that has to fit.

Where reviewers split

Where the engines part company depends on how many callers there are. At one stream they converge: the vreys-ai energy test measured identical-precision runs within one percent of each other in duration and GPU energy, and concluded the pair should be chosen on operational grounds rather than on energy. At 64 concurrent users the two are not the same measurement at all, the H200 benchmark having vLLM more than 35 times ahead on requests per second. Both are true at once, and the selfhostindex.com read puts the crossover around four to eight sustained streams.

The second split is about where each engine runs, and there the sources disagree more than the benchmarks do. vLLM's documentation lists Apple silicon and CPU support among its hardware plugins, while the selfhostindex.com write-up and the Red Hat Developer article of 15 June 2026 both put vLLM in the data-centre bracket and hand whatever hardware is already owned to llama.cpp. Our read: both are true, since vLLM's support list is broader than its reputation while its design centre remains the accelerator, which is why those paths sit in the plugin list rather than the quickstart.

The third split reverses inside a single source. In the H200 benchmark vLLM's P99 time to first token stayed nearly flat out to 64 users while llama.cpp's rose exponentially, and the same article reports the opposite for inter-token latency: llama.cpp's P99 inter-token latency is extremely low and above four users the roles reverse, because vLLM's larger batches slightly slow each token inside them.

Who each one is for

Take vLLM when more than a handful of people or programs are hitting the same model and the model fits in the accelerator. The H200 benchmark is the evidence on throughput, the documentation on features, and the selfhostindex.com write-up on the floor: below four to eight sustained streams a team pays vLLM's dependency tree for throughput nobody is asking for.

Take llama.cpp when the hardware is the constraint: one or two streams on a laptop, a Mac or a CPU box, a model larger than the card with layers left on the processor, GGUF files swapped freely, and offline running where nothing is rented. Its README's 16 backends and its single binary are the evidence, and the selfhostindex.com disqualification list is the reminder that the throughput figures above do not describe a single caller.

One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it. Read the weekly verdict.

Sources

Method: eleven sources across five tiers, eight carrying each engine, published between 31 March 2023 and 29 September 2026 and read on 30 September 2026, read and synthesised with AI; nothing was installed, prompted or rerun, and the Hacker News figures are that API's totals for two single threads.

One verdict a week.

Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.

By subscribing you agree to receive one weekly email from Review Machine. You can unsubscribe at any time.

Share by email