Skip to content

Catalogdev-infra

vLLM

situationalBlueshift — gaining momentum

High-throughput OSS inference engine — PagedAttention, OpenAI-compatible server, you bring GPUs.

Visit github.com →Add to kitOpen in Noemium

ledger entry · observed by @whysanesanders

The default way to serve open weights if you already have NVIDIA (or a growing list of other accelerators). OpenAI-compatible HTTP, continuous batching, quantization zoo. The bill is hardware and the on-call, not a SaaS SKU. Not a hosted model API. Not field-run in this pass.

Strengths

  • Default open-source engine for serving open weights at high throughput.
  • OpenAI-compatible HTTP server with continuous batching and quantization.
  • No SaaS seat fee; the bill is your hardware and on-call time.

Known limitations

  • You operate CUDA/ROCm, drivers, and capacity; this is not serverless inference.
  • Day-0 support for a brand-new architecture can lag the lab's own stack.
  • Multi-node parallelism is real ops, not a checkbox.
  • No public vLLM Cloud price — third-party hosts are other cards.

Use it when

  • Self-hosted open-weight inference on NVIDIA or supported accelerators.
  • Teams that need quantization and batching control under their own stack.

Skip it when

  • You want serverless inference without operating GPUs and drivers.
  • Day-0 model support or multi-node parallelism would block you.
  • You cannot staff the on-call needed for a self-hosted inference layer.

Buy-check

  • ·Price — free, Apache-licensed engine free; you pay GPUs/cloud; no vLLM seat fee as of 2026-08-19. Open the pricing page, not a screenshot.
  • ·Free tier — Free tier exists — burn it first.
  • ·Cancel fallout — What happens to your data/projects on cancel — check the ToS, not the FAQ.
  • ·Seat math — Per-seat pricing multiplies quietly. Count seats before the annual plan.
  • ·Who skips it — Also skipped when: You want serverless inference without operating GPUs and drivers.; Day-0 model support or multi-node parallelism would block you..

Price trailchangelog →

No price move in this repo yet.

Facts

pricing
free1
price note
Apache-licensed engine free; you pay GPUs/cloud; no vLLM seat fee
free tier
yes
open source
yes
api
yes
self-host
yes
category
dev-infra
evidence
source-verified
Models
BYOK

Data: trains on inputs unknown · local processing yes

Sources

  1. 1github.com/vllm-project/vllm[snapshot]
  2. 2docs.vllm.ai/en/latest/[snapshot]

No source, no number.

Edit this page on GitHub

Similar joball alternatives →

Compare:vLLM vs Ollama

toolverdictpriceossself-host
vLLMsituationalfreeyesyes
Modalshipfreemiumnono
OpenRoutershipfreemiumnono
Replicateshippaidnono

Facts from the catalog files — not scores you can buy. Quality and speed stay in the briefing, not in a fake 1–5 grid.

Modal
dev-infraship

Modal

Serverless GPU and compute platform for running ML and AI workloads.

APIFREE TIER

freemiumfree tier ($30/mo compute); pay-as-you-go beyond

Blueshift — gaining momentum
OpenRouter
dev-infraship

OpenRouteranchor

One API for every major model — unified billing, fallbacks and routing.

APIFREE TIER

freemiumpay-per-token pass-through plus a small fee

Blueshift — gaining momentum
Replicate
dev-infraship

Replicate

Run any community model via API — pay per second, no GPU wrangling.

API

paidpay per run; image models from ~$0.003/image

Steady momentum