Chips & Compute

Cerebras WSE-3

Cerebras · #1 most active of 18 in Chips & Compute

Compare →

Single-wafer accelerators delivering the fastest published open-model token rates.

Current version

WSE-3 / Inference

Entry cost

Free tier, then from $0.06 / M tokens

Changes / 30d

1

Capabilities

  • wafer-scale engine
  • record token/s on open models
  • CS-3 systems
  • training clusters

Pricing

  • Free

    Free

  • Pay-as-you-go

    per M input tokens

    $0.06

Entry price over time

Not enough pricing history yet — we start charting from the second observation.

Geek mode

Models under the hood

  • Llama / Qwen / GPT-OSS

    open weights

    disclosed
  • WSE-3

    Cerebras

    disclosed

Context window

Public API

yes

Multi-model routing

yes

Stack signals

  • wafer-scale integration
  • weight streaming
  • SwarmX fabric

Shared model stack

Other tracked products running on the same foundation models — a quick read on how much of the catalog moves when one of these models changes.

green disclosed · amber inferred · grey unknown

Reported scale

Reference data. Each figure is whatever the source actually said — weekly users, downloads, revenue run-rate — with its own definition and date. These are not comparable between tools and are never used to rank anything.

No public usage figure on record for this product.

Also in Chips & Compute

All Chips & Compute

Inference-first TPU pods

Version
v7 Ironwood
Cost
from ~$1.2 per chip-hour
Model
TPU v7

Deterministic low-latency inference

Version
GroqCloud
Cost
Free tier, then from $0.05 / M tokens
Model
Llama / Kimi / GPT-OSS

High-memory GPU alternative

Version
MI355X
Cost
~$2-3 per GPU-hour on clouds
Model
CDNA 4

On-device inference across Mac, iPhone and iPad.

Version
M5
Cost
Bundled with hardware
Model
Apple Foundation Models

Cloud-native training silicon

Version
Trainium2
Cost
from ~$1.3 per accelerator-hour
Model
Trainium2

Co-designed XPUs behind most hyperscaler silicon.

Version
2025
Cost
Enterprise supply agreements
Model
powers Google TPU

Change history

  • capability

    New throughput record published on a large open model

    record token/s

    source