Chips & Compute

Cerebras WSE-3

Cerebras · #3 most active of 18 in Chips & Compute

Compare →

Single-wafer accelerators delivering the fastest published open-model token rates.

67

Trust score · Mixed

3 data points checked · 2026-09-12

Model stack: VerifiedEntry price: DisputedVersion: DisputedHow this works →

Current version

WSE-3

Entry cost

Free tier / developer tier pricing available (contact/partner integrations)

Changes / 30d

9

The short answers · verified September 15, 2026

What is the latest version of Cerebras WSE-3?
The current shipped version of Cerebras WSE-3 is WSE-3, as of September 15, 2026.
How much does Cerebras WSE-3 cost?
Cerebras WSE-3 has no published flat entry price. Free tier / developer tier pricing available (contact/partner integrations)
What AI model does Cerebras WSE-3 use?
Cerebras WSE-3 runs primarily on Codex-Spark.

Source: Cerebraswww.cerebras.ai · Reusable under CC BY 4.0 — cite Tomorrow

Capabilities

  • wafer-scale engine
  • record token/s on open models
  • CS-3 systems
  • training clusters
  • Wafer-scale AI inference
  • Cerebras PyTorch
  • ModelZoo
  • Low-level SDK
  • OpenAI API compatibility
  • Disaggregated Inference
  • Wafer-Scale Engine (58x larger than GPUs)
  • 15x faster than GPUs
  • Cerebras PyTorch & ModelZoo support
  • Ultra-fast token generation
  • Low-latency inference
  • Model fine-tuning and training services
  • Wafer-Scale Engine
  • Ultra-fast AI Inference
  • Model fine-tuning and training
  • Custom model weights support
  • PyTorch support
  • 58x larger than GPUs
  • World's largest AI-optimized processor
  • Model training and fine-tuning
  • Ultra-fast AI inference
  • Wafer-Scale Engine AI hardware
  • AI model training and fine-tuning
  • Drop-in OpenAI API compatibility
  • Inference API
  • PyTorch Training
  • Model fine-tuning
  • AI training
  • AI inference
  • Low-latency programming
  • AI Inference
  • Model Fine-tuning
  • Model Training
  • OpenAI API Compatibility
  • Training
  • Fine-tuning
  • Training and fine-tuning services
  • Dedicated queue priority
  • Custom model weights
  • AI Hardware / Chip
  • AI Training & Fine-tuning
  • Wafer-scale hardware acceleration
  • LLM Inference API
  • Model training
  • SDK
  • Developer support
  • AI Training
  • PyTorch Support
  • AI Hardware Accelerator
  • Inference Cloud API
  • World-record speeds
  • AWS Marketplace Integration
  • OpenRouter Integration
  • HuggingFace Integration
  • Vercel Integration
  • Training services
  • Fastest AI inference
  • Developer API
  • AWS Marketplace integration
  • OpenRouter integration
  • HuggingFace integration
  • Vercel integration
  • AI Accelerator
  • AWS Marketplace
  • AI Inference API
  • Model training services
  • Low latency
  • Inference
  • HPC applications
  • Up to 30x faster than GPUs
  • Dedicated on-premise deployment
  • Pytorch Support
  • inference
  • training
  • fine-tuning
  • api-access
  • on-premise
  • Developer SDK
  • Low-latency queue priority
  • GPU-free Wafer-Scale Engine
  • OpenAI-compatible API
  • Fine-tuning and training services
  • Developer Tier Pricing
  • Enterprise SLA
  • Low latency inference
  • Dedicated support
  • Up to 30x faster inference than GPU
  • Pre-training
  • PyTorch SDK
  • Low-latency token generation
  • OpenAI API compatible
  • Custom kernels
  • On-Premise Deployment
  • Cloud API

Pricing

  • Developer Tier

    Free

Entry price over time

11 of 49 points reconstructed from Internet Archive captures (hollow dots) — dated by capture, not by when we first saw the page. View a capture

9/2/2025 · $15009/15/2026 · $0

Geek mode

Models under the hood

  • Codex-Spark

    Cerebras

    disclosed
  • Gemma

    Google

    inferred
  • Gemma-4-31B

    Google

    inferred
  • GLM-4.7

    Zhipu

    disclosed
  • GPT 5.6 Sol

    OpenAI

    disclosed
  • GPT-5.6-Sol-Ultrafast

    OpenAI

    disclosed
  • Kimi K2.6

    Moonshot

    disclosed
  • Llama

    Meta

    inferred
  • Llama / Qwen / GPT-OSS

    open weights

    disclosed
  • Qwen

    Alibaba

    inferred
  • WSE-3

    Cerebras

    disclosed

Context window

Public API

yes

Multi-model routing

yes

Stack signals

  • PyTorch

Shared model stack

Other tracked products running on the same foundation models — a quick read on how much of the catalog moves when one of these models changes.

green disclosed · amber inferred · grey unattributed. Aliases are folded into one model; provider concentration counts only models with an established vendor.

Reported scale

Reference data. Each figure is whatever the source actually said — weekly users, downloads, revenue run-rate — with its own definition and date. These are not comparable between tools and are never used to rank anything.

No public usage figure on record for this product.

Also in Chips & Compute

All Chips & Compute

Open RISC-V AI hardware

Version
Blackhole
Cost
from $999/unit
Model
Tensix cores

Inference-first TPU pods

Version
Ironwood (7th generation)
Cost
from $5.4/unit
Model
Gemini

GB200/B200 rack-scale AI systems

Version
Blackwell
Cost
Pricing is not publicly listed (enterprise hardware architecture/system)
Model
Blackwell

Cloud-native training silicon

Version
Trainium2
Cost
Pricing not published; available via AWS EC2 instance pricing.
Model
Neuron compiler

Cost-focused training and inference accelerator.

Version
1.24.1
Cost
Pricing not published. Available via Intel Tiber AI Cloud or OEM platforms.
Model
Accelerator
5

Dataflow chips for model serving

Version
SN40L
Cost
Free tier available. Paid plans include Developer (Pay-as-you-go) and Enterprise (Subscription-based).
Model
DeepSeek-V3.1

Elsewhere on the same models

Products in other segments that name one of Cerebras WSE-3's models in their stack.

Fireworks AI

Fireworks AI

Low-latency serving for open models and custom fine-tunes.

Version
2025
Cost
Per-token usage
Model
Llama
Kimi

Moonshot AIChina

10

Long-context assistant and agentic K-series models.

Version
Kimi K3
Cost
API access with pay-as-you-go pricing based on token consumption. File APIs are currently free.
Model
Kimi K2.6
Ollama

Ollama

Local model runner that made desktop inference trivial.

Version
0.x
Cost
Free
Model
Gemma
Qwen-Image

AlibabaChina

1

Open-weight image generation and editing from the Qwen line.

Version
Qwen-Image
Cost
Pricing not specified in the provided documentation.
Model
Qwen
Together AI

Together AI

Open-model inference and fine-tuning at frontier speed.

Version
2025
Cost
Per-token, from $0.06/M
Model
Llama

Change history

  • capability

    New capabilities: On-Premise Deployment, Cloud API

    96 tracked98 tracked · +On-Premise Deployment, Cloud API

    source
  • pricing

    Entry price decreased to $0/mo

    $10$0

    source
  • capability

    New capabilities: OpenAI API compatible, Custom kernels

    94 tracked96 tracked · +OpenAI API compatible, Custom kernels

    source
  • capability

    New capabilities: PyTorch SDK, Low-latency token generation

    92 tracked94 tracked · +PyTorch SDK, Low-latency token generation

    source
  • capability

    New capabilities: Up to 30x faster inference than GPU, Pre-training

    90 tracked92 tracked · +Up to 30x faster inference than GPU, Pre-training

    source
  • capability

    New capabilities: Low latency inference, Dedicated support

    88 tracked90 tracked · +Low latency inference, Dedicated support

    source
  • capability

    New capabilities: Developer Tier Pricing, Enterprise SLA

    86 tracked88 tracked · +Developer Tier Pricing, Enterprise SLA

    source
  • capability

    New capabilities: GPU-free Wafer-Scale Engine, OpenAI-compatible API, Fine-tuning and training services

    83 tracked86 tracked · +GPU-free Wafer-Scale Engine, OpenAI-compatible API, Fine-tuning and training services

    source
  • capability

    New capabilities: Developer SDK, Low-latency queue priority

    81 tracked83 tracked · +Developer SDK, Low-latency queue priority

    source
  • capability

    New capabilities: inference, training, fine-tuning

    76 tracked81 tracked · +inference, training, fine-tuning, api-access, on-premise

    source
  • capability

    New capabilities: Pytorch Support

    75 tracked76 tracked · +Pytorch Support

    source
  • capability

    New capabilities: Up to 30x faster than GPUs, Dedicated on-premise deployment

    73 tracked75 tracked · +Up to 30x faster than GPUs, Dedicated on-premise deployment

    source
  • capability

    New capabilities: HPC applications

    72 tracked73 tracked · +HPC applications

    source
  • capability

    New capabilities: Inference

    71 tracked72 tracked · +Inference

    source
  • capability

    New capabilities: Low latency

    70 tracked71 tracked · +Low latency

    source
  • version

    Cerebras WSE-3 moved to WSE-3

    CS-3WSE-3

    source
  • version

    Cerebras WSE-3 moved to CS-3

    WSE-3CS-3

    source
  • capability

    New capabilities: AI Inference API, Model training services

    68 tracked70 tracked · +AI Inference API, Model training services

    source
  • version

    Cerebras WSE-3 moved to WSE-3

    CS-4WSE-3

    source
  • version

    Cerebras WSE-3 moved to CS-4

    WSE-3CS-4

    source
  • capability

    New capabilities: AI Accelerator, AWS Marketplace

    66 tracked68 tracked · +AI Accelerator, AWS Marketplace

    source
  • capability

    New capabilities: Fastest AI inference, Developer API, AWS Marketplace integration

    60 tracked66 tracked · +Fastest AI inference, Developer API, AWS Marketplace integration, OpenRouter integration, HuggingFace integration, Vercel integration

    source
  • capability

    New capabilities: Training services

    59 tracked60 tracked · +Training services

    source
  • version

    Cerebras WSE-3 moved to WSE-3

    CS-4WSE-3

    source
  • version

    Cerebras WSE-3 moved to CS-4

    CS-3CS-4

    source
  • model

    Cerebras WSE-3 added Codex-Spark to its model stack

    GLM-4.7, GPT 5.6 Sol, GPT-5.6-Sol-Ultrafast, Gemma, Gemma-4-31B, Kimi K2.6, Llama, Llama / Qwen / GPT-OSS, Qwen, WSE-3Codex-Spark, GLM-4.7, GPT 5.6 Sol, GPT-5.6-Sol-Ultrafast, Gemma, Gemma-4-31B, Kimi K2.6, Llama, Llama / Qwen / GPT-OSS, Qwen, WSE-3

    source
  • capability

    New capabilities: AWS Marketplace Integration, OpenRouter Integration, HuggingFace Integration

    55 tracked59 tracked · +AWS Marketplace Integration, OpenRouter Integration, HuggingFace Integration, Vercel Integration

    source
  • version

    Cerebras WSE-3 moved to CS-3

    WSE-3CS-3

    source
  • capability

    New capabilities: AI Hardware Accelerator, Inference Cloud API, World-record speeds

    52 tracked55 tracked · +AI Hardware Accelerator, Inference Cloud API, World-record speeds

    source
  • capability

    New capabilities: AI Training, PyTorch Support

    50 tracked52 tracked · +AI Training, PyTorch Support

    source
  • model

    Cerebras WSE-3 added GPT 5.6 Sol, GPT-5.6-Sol-Ultrafast to its model stack

    GLM-4.7, Gemma, Gemma-4-31B, Kimi K2.6, Llama, Llama / Qwen / GPT-OSS, Qwen, WSE-3GLM-4.7, GPT 5.6 Sol, GPT-5.6-Sol-Ultrafast, Gemma, Gemma-4-31B, Kimi K2.6, Llama, Llama / Qwen / GPT-OSS, Qwen, WSE-3

    source
  • capability

    New capabilities: Model training, SDK, Developer support

    47 tracked50 tracked · +Model training, SDK, Developer support

    source
  • capability

    New capabilities: Wafer-scale hardware acceleration, LLM Inference API

    45 tracked47 tracked · +Wafer-scale hardware acceleration, LLM Inference API

    source
  • capability

    New capabilities: AI Hardware / Chip, AI Training & Fine-tuning

    43 tracked45 tracked · +AI Hardware / Chip, AI Training & Fine-tuning

    source
  • capability

    New capabilities: Training and fine-tuning services, Dedicated queue priority, Custom model weights

    40 tracked43 tracked · +Training and fine-tuning services, Dedicated queue priority, Custom model weights

    source
  • capability

    New capabilities: Training, Fine-tuning

    38 tracked40 tracked · +Training, Fine-tuning

    source
  • capability

    New capabilities: AI Inference, Model Fine-tuning, Model Training

    34 tracked38 tracked · +AI Inference, Model Fine-tuning, Model Training, OpenAI API Compatibility

    source
  • version

    Cerebras WSE-3 moved to WSE-3

    CS-3WSE-3

    source
  • capability

    New capabilities: AI training, AI inference, Low-latency programming

    31 tracked34 tracked · +AI training, AI inference, Low-latency programming

    source
  • version

    Cerebras WSE-3 moved to CS-3

    WSE-3CS-3

    source
Subscribe to Cerebras WSE-3 changes

Cerebras WSE-3 compared

Straight head-to-head pages against the busiest products in Chips & Compute.

How to cite this page

Free to cite and reuse under CC BY 4.0. Permalink: https://tomorrow.aliensquad.ai/tools/cerebras

APA
Tomorrow. (2026). Cerebras WSE-3 — version, pricing and model stack [Data set entry]. AlienSquad. Retrieved 2026-09-17, from https://tomorrow.aliensquad.ai/tools/cerebras
BibTeX
@misc{tomorrow-tools-cerebras,
  author       = {{Tomorrow}},
  title        = {Cerebras WSE-3 — version, pricing and model stack},
  year         = {2026},
  publisher    = {AlienSquad},
  howpublished = {\url{https://tomorrow.aliensquad.ai/tools/cerebras}},
  note         = {Accessed: 2026-09-17}
}