
vLLM
situationalHigh-throughput OSS inference engine — PagedAttention, OpenAI-compatible server, you bring GPUs.
Visit github.com →> **vLLM** — situational > Verified 2026-08-19 > Limitation: You operate CUDA/ROCm, drivers, and capacity; this is not serverless inference. > > https://noemium.com/tools/vllm/Add to kitOpen in NoemiumFact-check this catalog entry for vLLM (dev-infra). Verdict: situational Verdict text: The default way to serve open weights if you already have NVIDIA (or a growing list of other accelerators). OpenAI-compatible HTTP, continuous batching, quantization zoo. The bill is hardware and the on-call, not a SaaS SKU. Not a hosted model API. Not field-run in this pass. Limitations: - You operate CUDA/ROCm, drivers, and capacity; this is not serverless inference. - Day-0 support for a brand-new architecture can lag the lab's own stack. - Multi-node parallelism is real ops, not a checkbox. - No public vLLM Cloud price — third-party hosts are other cards. Receipts: https://github.com/vllm-project/vllm https://docs.vllm.ai/en/latest/ Last verified: 2026-08-19 Find outdated prices, wrong claims, or missing limitations. Cite primary sources (docs, pricing pages, benchmarks, model cards). Do not use marketing pages.[](https://noemium.com/tools/vllm/)
ledger entry · observed by @whysanesanders
The default way to serve open weights if you already have NVIDIA (or a growing list of other accelerators). OpenAI-compatible HTTP, continuous batching, quantization zoo. The bill is hardware and the on-call, not a SaaS SKU. Not a hosted model API. Not field-run in this pass.
Strengths
- Default open-source engine for serving open weights at high throughput.
- OpenAI-compatible HTTP server with continuous batching and quantization.
- No SaaS seat fee; the bill is your hardware and on-call time.
Known limitations
- You operate CUDA/ROCm, drivers, and capacity; this is not serverless inference.
- Day-0 support for a brand-new architecture can lag the lab's own stack.
- Multi-node parallelism is real ops, not a checkbox.
- No public vLLM Cloud price — third-party hosts are other cards.
Use it when
- Self-hosted open-weight inference on NVIDIA or supported accelerators.
- Teams that need quantization and batching control under their own stack.
Skip it when
- You want serverless inference without operating GPUs and drivers.
- Day-0 model support or multi-node parallelism would block you.
- You cannot staff the on-call needed for a self-hosted inference layer.
Buy-check
- ·Price — free, Apache-licensed engine free; you pay GPUs/cloud; no vLLM seat fee as of 2026-08-19. Open the pricing page, not a screenshot.
- ·Free tier — Free tier exists — burn it first.
- ·Cancel fallout — What happens to your data/projects on cancel — check the ToS, not the FAQ.
- ·Seat math — Per-seat pricing multiplies quietly. Count seats before the annual plan.
- ·Who skips it — Also skipped when: You want serverless inference without operating GPUs and drivers.; Day-0 model support or multi-node parallelism would block you..
Price trailchangelog →
No price move in this repo yet.
Facts
- pricing
- free1
- price note
- Apache-licensed engine free; you pay GPUs/cloud; no vLLM seat fee
- free tier
- yes
- open source
- yes
- api
- yes
- self-host
- yes
- category
- dev-infra
- evidence
- source-verified
- Models
- BYOK
Data: trains on inputs unknown · local processing yes
Sources
No source, no number.
Similar joball alternatives →
Compare:vLLM vs Ollama
| tool | verdict | price | oss | self-host |
|---|---|---|---|---|
| vLLM | situational | free | yes | yes |
| Modal | ship | freemium | no | no |
| OpenRouter | ship | freemium | no | no |
| Replicate | ship | paid | no | no |
Facts from the catalog files — not scores you can buy. Quality and speed stay in the briefing, not in a fake 1–5 grid.
Modal
Serverless GPU and compute platform for running ML and AI workloads.
APIFREE TIER
freemiumfree tier ($30/mo compute); pay-as-you-go beyond
OpenRouteranchor
One API for every major model — unified billing, fallbacks and routing.
APIFREE TIER
freemiumpay-per-token pass-through plus a small fee
Replicate
Run any community model via API — pay per second, no GPU wrangling.
API
paidpay per run; image models from ~$0.003/image