← Segments

Audio

Voice, music, transcription and dubbing

15 tools · 1 changes / 30d

AssemblyAI

AssemblyAI

Speech intelligence API

Version
Universal-2
Cost
Pay as you go from $0.12/hr
Model
Universal-2
Cartesia

Cartesia

Ultra-low-latency voice models built on state space models.

Version
Sonic 3
Cost
Free tier, $5/mo
Model
Sonic 3
Deepgram

Deepgram

Real-time speech to text

Version
Nova-3
Cost
Pay as you go from $0.0043/min
Model
Nova-3
ElevenLabs

ElevenLabs

Lifelike text to speech

Version
v3
Cost
Free tier, Starter $5/mo
Model
Eleven v3
Fireflies.ai

Fireflies.ai

Notetaker with CRM-aware conversation intelligence.

Version
2025
Cost
Free tier, $18/mo Pro
Model
in-house ASR
Fish Audio

Fish AudioChina

Open-weight multilingual TTS and voice cloning.

Version
S1
Cost
Open weights; API from $15/M chars
Model
Fish Speech S1
Granola

Granola

Local-first meeting notes that augment your own typing.

Version
2025
Cost
Free tier, $18/mo
Model
Claude
MiniMax Speech

MiniMaxChina

1

Multilingual TTS and voice cloning at low per-character rates.

Version
2.6
Cost
Free tier, API per character
Model
MiniMax Speech 2.6
Mureka

Kunlun TechChina

Music generation with stem export and a public API.

Version
V8
Cost
from $8/mo
Model
Mureka V8
Otter.ai

Otter.ai

Meeting transcription and AI meeting agents.

Version
2025
Cost
Free tier, $16.99/mo Pro
Model
in-house ASR
PlayAI

PlayAI

Voice agents and TTS for conversational products.

Version
PlayDialog
Cost
Free tier, $31.20/mo
Model
PlayDialog
Speechmatics

Speechmatics

Accuracy-focused speech recognition for broadcast and enterprise.

Version
Ursa 2
Cost
Usage-based, self-host option
Model
Ursa 2
Suno

Suno

Full songs from a text prompt

Version
v5
Cost
Free tier, Pro $10/mo
Model
Suno v5
Udio

Udio

Music generation and remixing

Version
v2
Cost
Free tier, Standard $10/mo
Model
Udio v2
Whisper

OpenAI

Open-weight speech recognition

Version
large-v3
Cost
Free weights, API $0.006/min
Model
Whisper large-v3

Recent changes in Audio

  • capability

    MiniMax SpeechEmotion control exposed in the speech API

    Voice cloning, 30+ languagesVoice cloning, 30+ languages, Emotion control

    source
  • version

    ElevenLabsEleven v3 released with emotional tags and 70+ languages

    v2.5v3

    source
  • version

    SunoSuno v5 improves vocal clarity and instrument separation

    v4.5v5

    source
  • capability

    ElevenLabsMusic generation added to the platform

    Voice cloning, Dubbing studioVoice cloning, Dubbing studio, Music

    source
  • pricing

    DeepgramNova-3 streaming price reduced to $0.0043 per minute

    $0.0059 per minute$0.0043 per minute

    source
  • version

    AssemblyAIUniversal-2 model rolled out as the default

    Universal-1Universal-2

    source
  • model

    Whisperlarge-v3-turbo variant published for faster inference

    large-v3large-v3, large-v3-turbo

    source