{"name":"Together AI","tagline":"Cloud inference for open models — serverless endpoints and fine-tuning.","url":"https://www.together.ai","category":"dev-infra","pricing":"freemium","price_note":"starter credits; from $0.10/1M tokens (small open models)","free_tier":true,"open_source":false,"api":true,"self_host":false,"models_used":["llama-4-maverick","qwen3-8-max","deepseek-v4-flash"],"model_routing":"locked","verdict":"ship","verdict_text":"Solid uptime, broad catalog, serverless or dedicated endpoints. A safe place to move an open-weight prototype into production. Costs at scale push you toward dedicated GPUs, and the newest models sometimes land later than on specialist hosts.","limitations":["Costs at scale push you toward dedicated GPU contracts.","Newest models sometimes arrive later than at specialists.","Fine-tuning UX less polished than the inference side."],"strengths":["Solid uptime and a broad open-model catalog, serverless or dedicated.","Safe place to move an open-weight prototype into production."],"use_for":["Hosted open models without picking a GPU vendor yet.","Fine-tunes that can start on their inference side."],"skip_when":["Scale already points at dedicated GPU contracts.","You need the newest model the day a specialist hosts it.","Fine-tuning UX is the product, not an afterthought."],"receipts":["https://www.together.ai","https://www.together.ai/pricing"],"affiliate":"none","evidence_tier":"source-verified","momentum":"steady","featured":false,"last_verified":"2026-08-19","observed_by":"whysanesanders","slug":"together-ai","canonical_url":"https://noemium.com/tools/together-ai/","attribution":"Noemium catalog data is licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). See https://noemium.com/method/ for how entries are verified."}