Groq LPU
SRAM-based language processing units sold as a token-priced inference cloud.
Full Groq LPU record →Head to head
Groq LPU (Groq) and Google TPU Ironwood (Google) both sit in Chips & Compute. Groq LPU is the cheaper entry point at Pricing not published. Contact sales or check account console for usage-based rates.. Google TPU Ironwood shipped more tracked changes in the last 30 days (12 vs 4). Every value below comes from the latest crawl of the vendors' own pages.
| Attribute | Groq LPU Groq | Google TPU Ironwood |
|---|---|---|
| Segment | Chips & Compute | Chips & Compute |
| Version | LPU | Ironwood (7th generation) |
| Entry cost | Pricing not published. Contact sales or check account console for usage-based rates. | from $5.4/unit |
| Pricing tiers | GroqCloud Console Free Tier / Pay-As-You-Go free | On-Demand (us-central1 Iowa) $12 · DWS Flex-start price (us-central1 Iowa) $6 · DWS Calendar Mode price (us-central1 Iowa) $8.4 · 1-year Commitment (us-central1 Iowa) $8.4 · 3-year Commitment (us-central1 Iowa) $5.4 · On-Demand (europe-west2 London) $13.2 · DWS Flex-start price (europe-west2 London) $6 · DWS Calendar Mode price (europe-west2 London) $8.4 |
| Model stack | canopylabs/orpheus-arabic-saudi · canopylabs/orpheus-v1-english · kimi-k2-instruct-0905 · Llama / Kimi / GPT-OSS · Llama 3.1 8B Instant · Llama 3.3 70B Versatile · LPU v2 · meta-llama/llama-4-maverick-17b-128e-instruct · meta-llama/llama-4-scout-17b-16e-instruct · minimax-m2.5 · minimax/minimax-m2.5 · minimaxai/minimax-m2.5 · moonshotai/kimi-k2-instruct-0905 · openai/gpt-oss-120b · openai/gpt-oss-20b · openai/gpt-oss-safeguard-20b · Orpheus English · orpheus-arabic-saudi · orpheus-v1-english · Qwen 3.6 27B · qwen/qwen3-vl-32b-instruct · qwen3-vl-32b-instruct · Whisper V3 Large · whisper-large-v3 | Gemini · Ironwood · Ironwood (TPU v7) · Ironwood TPU · Ironwood TPU (7th Generation) · TPU v7 · TPU7x · TPU7x (Ironwood) · XLA compiler |
| Context window | 131k | — |
| Public API | ||
| Routes models | ||
| Changes / 30d | 4 | 12 |
| Origin | Global | Global |
| Capabilities | deterministic latency, OpenAI-compatible API, open-model catalogue, batch API, Fast LLM inference, OpenAI Compatibility, Prompt Caching, Speech to Text | 9,216-chip pods, inference optimised, liquid cooled, JAX + PyTorch/XLA, Large-scale training, Reasoning and inference, 9,216 liquid-cooled chips per pod, 42.5 ExaFlops performance |
SRAM-based language processing units sold as a token-priced inference cloud.
Full Groq LPU record →v7 TPU pods scaling to 9,216 chips, available through Google Cloud and used for Gemini.
Full Google TPU Ironwood record →