{"name":"Groq","tagline":"Ultra-fast inference on custom LPU hardware — open models at 500+ tok/s.","url":"https://groq.com","category":"dev-infra","pricing":"freemium","price_note":"free tier; from $0.075/1M tokens (small models)","free_tier":true,"open_source":false,"api":true,"self_host":false,"model_routing":"locked","verdict":"ship","verdict_text":"Sub-second responses really do change what you can build — voice agents, autocomplete, anything real-time. The free tier is big enough to prototype on. Just remember the model catalog is limited to what fits their own hardware.","limitations":["Model catalog limited to what fits their hardware.","Context windows capped below the biggest hosted rivals.","Free-tier throughput limits are tight for production."],"strengths":["Sub-second tokens on their LPUs — voice agents and autocomplete actually feel live.","Free tier is big enough to prototype the latency, not just read a blog post."],"use_for":["Real-time UIs that die at 2-second model latency.","Open-model inference when the model already fits their hardware."],"skip_when":["The model you need is not on their catalog.","You need the largest hosted context windows.","Free-tier throughput would be the production path."],"receipts":["https://groq.com","https://console.groq.com/docs/models","https://console.groq.com/docs/deprecations"],"affiliate":"none","evidence_tier":"source-verified","momentum":"blueshift","featured":false,"last_verified":"2026-08-24","observed_by":"whysanesanders","slug":"groq","canonical_url":"https://noemium.com/tools/groq/","attribution":"Noemium catalog data is licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). See https://noemium.com/method/ for how entries are verified."}