Catalog/ № 016
Fireworks AI
situationalFast inference platform for open models — serverless and on-demand GPUs.
fireworks.aiVerified 2026-08-15
ledger entry · observed by @whysanesanders
Genuinely fast, and the compound-AI tooling (function calling, multi-modal) is well done. But it competes with Together and Groq on the same models — pick whoever serves your model cheapest this quarter.
Known limitations
- Differentiation vs Together/Groq unclear for standard inference.
- Enterprise pricing opaque.
- Smaller free tier than rivals.
- GPU on-demand prices rise 2026-09-01 (H100 $7 to $8/hr).
Facts
- pricing
- freemium1
- price note
- $1 starter credits; serverless from $0.10/1M tokens; on-demand GPU from $7/hr
- free tier
- yes
- open source
- no
- api
- yes
- self-host
- no
- category
- dev-infra
Models used
Sources
If we can't show the source, we don't print the number.
Same shelf — dev-infra
Groq
dev-infra
Ultra-fast inference on custom LPU hardware — open models at 500+ tok/s.
APIFREE TIER
freemiumfree tier; from $0.05/1M tokens (small models)
Hugging Face
dev-infra
The GitHub of AI — models, datasets, Spaces and inference endpoints.
APIFREE TIER
freemiumfree; PRO $9/mo; inference pay-as-you-go
OpenRouterfeatured
dev-infra
One API for every major model — unified billing, fallbacks and routing.
APIFREE TIER
freemiumpay-per-token pass-through plus a small fee