Cerebras
Wafer-scale inference and training
- Version
- WSE-3 / Inference
- Cost
- Free tier, then from $0.06 / M tokens
- Model
- Llama / Qwen / GPT-OSS
AWS · #6 most active of 18 in Chips & Compute
Trn2 UltraServers and Project Rainier capacity, programmed through the Neuron SDK.
Current version
Trainium2
Entry cost
from ~$1.3 per accelerator-hour
Changes / 30d
0
from ~$1.3 per accelerator-hour
Entry price over time
Not enough pricing history yet — we start charting from the second observation.
Trainium2
AWS
Neuron compiler
AWS
Context window
—
Public API
yes
Multi-model routing
no
Other tracked products running on the same foundation models — a quick read on how much of the catalog moves when one of these models changes.
green disclosed · amber inferred · grey unknown
Reference data. Each figure is whatever the source actually said — weekly users, downloads, revenue run-rate — with its own definition and date. These are not comparable between tools and are never used to rank anything.
No public usage figure on record for this product.
Cerebras
Wafer-scale inference and training
Inference-first TPU pods
Groq
Deterministic low-latency inference
High-memory GPU alternative
On-device inference across Mac, iPhone and iPad.
Broadcom
Co-designed XPUs behind most hyperscaler silicon.