
Ollama
situationalRun open models locally (CLI, app, API) — optional Ollama Cloud Pro/Max for hosted large models.
Visit ollama.com →> **Ollama** — situational > Verified 2026-08-19 > Limitation: Local quality and speed are your GPU; Ollama is the runner, not the weights. > > https://noemium.com/tools/ollama/Add to kitOpen in NoemiumFact-check this catalog entry for Ollama (dev-infra). Verdict: situational Verdict text: The local-model default: one CLI, an OpenAI-shaped API, desktop apps, and you keep weights on disk. Cloud Pro is a separate $20 meter with 5-hour and weekly caps; Max is capacity-paused for new buyers. Local is the product; cloud is overflow. Not a quality claim about any one model. Limitations: - Local quality and speed are your GPU; Ollama is the runner, not the weights. - Cloud usage is token-and-model-weighted, not a flat token bucket. - New Max subscriptions paused while they add capacity (existing Max kept). - Cloud hosts primarily in the US (EU/Singapore overflow); data-residency still a question. Receipts: https://ollama.com/pricing https://github.com/ollama/ollama Last verified: 2026-08-19 Find outdated prices, wrong claims, or missing limitations. Cite primary sources (docs, pricing pages, benchmarks, model cards). Do not use marketing pages.[](https://noemium.com/tools/ollama/)
ledger entry · observed by @whysanesanders
The local-model default: one CLI, an OpenAI-shaped API, desktop apps, and you keep weights on disk. Cloud Pro is a separate $20 meter with 5-hour and weekly caps; Max is capacity-paused for new buyers. Local is the product; cloud is overflow. Not a quality claim about any one model.
Strengths
- Local-model default with one CLI, an OpenAI-shaped API, and desktop apps.
- You keep model weights on disk and run inference on your own hardware.
- Free local runtime with no usage caps beyond what your machine can handle.
Known limitations
- Local quality and speed are your GPU; Ollama is the runner, not the weights.
- Cloud usage is token-and-model-weighted, not a flat token bucket.
- New Max subscriptions paused while they add capacity (existing Max kept).
- Cloud hosts primarily in the US (EU/Singapore overflow); data-residency still a question.
Use it when
- Offline or private inference where data never leaves the machine.
- Prototyping with open weights before deciding on cloud spend.
Skip it when
- Your GPU cannot deliver the quality or speed the use case demands.
- You need guaranteed data residency from a cloud region.
- You were planning to sign up for Max while new subscriptions are paused.
Buy-check
- ·Price — freemium, local runtime $0 unlimited on your hardware; Cloud Free light usage; Pro $20/mo ($200/yr); Max $100/mo (new Max sign-ups paused) as of 2026-08-19. Open the pricing page, not a screenshot.
- ·Free tier — Free tier exists — burn it first.
- ·Cancel fallout — What happens to your data/projects on cancel — check the ToS, not the FAQ.
- ·Seat math — Per-seat pricing multiplies quietly. Count seats before the annual plan.
- ·Who skips it — Also skipped when: Your GPU cannot deliver the quality or speed the use case demands.; You need guaranteed data residency from a cloud region..
Price trailchangelog →
No price move in this repo yet.
Facts
- pricing
- freemium1
- price note
- local runtime $0 unlimited on your hardware; Cloud Free light usage; Pro $20/mo ($200/yr); Max $100/mo (new Max sign-ups paused)
- free tier
- yes
- open source
- yes
- api
- yes
- self-host
- yes
- category
- dev-infra
- evidence
- source-verified
- Models
- BYOK
Data: trains on inputs unknown · local processing yes
Sources
No source, no number.
Similar joball alternatives →
Compare:Ollama vs vLLMOllama vs Open WebUI
| tool | verdict | price | oss | self-host |
|---|---|---|---|---|
| Ollama | situational | freemium | yes | yes |
| Modal | ship | freemium | no | no |
| OpenRouter | ship | freemium | no | no |
| Replicate | ship | paid | no | no |
Facts from the catalog files — not scores you can buy. Quality and speed stay in the briefing, not in a fake 1–5 grid.
Modal
Serverless GPU and compute platform for running ML and AI workloads.
APIFREE TIER
freemiumfree tier ($30/mo compute); pay-as-you-go beyond
OpenRouteranchor
One API for every major model — unified billing, fallbacks and routing.
APIFREE TIER
freemiumpay-per-token pass-through plus a small fee
Replicate
Run any community model via API — pay per second, no GPU wrangling.
API
paidpay per run; image models from ~$0.003/image