03 /MODELS
Models
What a task actually costs, not what a landing page implies. Sort by price per task; unit-priced media models stay in their own group at the bottom.
Prices from Helicone (Apache-2.0) + LiteLLM (MIT).Verified 2026-08-15. Vendors change prices — check before billing decisions.
task
price per task≈75k in / 2k out
| open | best for | ||||||
|---|---|---|---|---|---|---|---|
| Qwen3 235B A22B | Alibaba (Qwen) | 262k | $0.071 | $0.100 | $0.0055 | yes | self-hosted multilingual chat (Apache-2.0); commercial self-host without license friction |
| Llama 4 Maverick | Meta | 131k | $0.200 | $0.600 | $0.016 | yes | self-hosted multimodal MoE; fine-tuning base for custom models |
| DeepSeek R1 | DeepSeek | 131k | $0.280 | $0.420 | $0.022 | yes | reasoning on a budget; self-hosted chain-of-thought workloads |
| DeepSeek V3 | DeepSeek | 131k | $0.280 | $0.420 | $0.022 | yes | cheap general-purpose chat at scale; self-hosted MoE deployments |
| GPT-5 mini | OpenAI | 272k | $0.250 | $2.00 | $0.023 | no | cheap batch classification; high-throughput support chat |
| Gemini 2.5 Flash | 1M | $0.300 | $2.50 | $0.028 | no | cheap batch classification at scale; low-latency production chat | |
| Mistral Large | Mistral AI | 262k | $0.500 | $1.50 | $0.041 | no | european data-residency deployments; strong multilingual chat (FR/DE/ES) |
| Claude Haiku 3.5 | Anthropic | 200k | $0.800 | $4.00 | $0.068 | no | cheap high-volume classification; low-latency autocomplete-style tasks |
| o4-mini | OpenAI | 200k | $1.10 | $4.40 | $0.091 | no | fast reasoning on a budget; math and coding at scale |
| Gemini 2.5 Pro | 1M | $1.25 | $10.00 | $0.114 | no | long-document analysis (1M token context); multimodal video/audio understanding | |
| GPT-5 | OpenAI | 272k | $1.25 | $10.00 | $0.114 | no | agentic coding and multi-step tool use; complex reasoning over mixed text/image input |
| o3 | OpenAI | 200k | $2.00 | $8.00 | $0.166 | no | deep math/science reasoning; complex debugging and code analysis |
| Claude Sonnet 4.5 | Anthropic | 200k | $3.00 | $15.00 | $0.255 | no | agentic coding and multi-step tool use; long-document analysis |
| Grok 4 | xAI | 256k | $3.00 | $15.00 | $0.255 | no | reasoning with real-time X/Twitter context; long-context analysis (256k) |
| Claude Opus 4.1 | Anthropic | 200k | $15.00 | $75.00 | $1.27 | no | hardest agentic coding tasks; long-horizon autonomous agents |
| Eleven v3 | ElevenLabs | — | $0.00018/char | — | no | expressive text-to-speech with emotion tags; multilingual voiceover and audiobooks | |
| FLUX 1.1 Pro | Black Forest Labs | — | $0.040/image | — | no | production text-to-image via API; fast high-quality product/marketing shots | |
| Sora 2 | OpenAI | — | $0.100/video-sec | — | no | text-to-video generation; creative storyboarding and concept clips | |
| Veo 3 | — | $0.400/video-sec | — | no | video generation with native audio; high-fidelity short clips via Vertex AI | ||
| Whisper Large v3 | OpenAI | — | $0.000031/audio-sec | — | yes | self-hosted speech-to-text (MIT weights); cheap batch transcription via Groq | |
20 / 20 models · unit-priced media models stay at the bottom