← Segments

Audio

Voice, music, transcription and dubbing

15 tools · 29 changes / 30d

AssemblyAI

AssemblyAI

5

Speech intelligence API

Version
Universal-3.5 Pro
Cost
Free tier with $50 free credits, then pay-as-you-go starting at $0.15/hr for pre-recorded STT (Universal-2)
Model
Claude
Cartesia

Cartesia

Ultra-low-latency voice models built on state space models.

Version
Sonic 3
Cost
Free tier · from $5/mo
Model
Sonic 3
Deepgram

Deepgram

6

Real-time speech to text

Version
Nova-3
Cost
Free tier · from $4,000/min
Model
Aura-1
ElevenLabs

ElevenLabs

3

Lifelike text to speech

Version
Eleven v3
Cost
Free tier · from $6/mo
Model
Dubbing v2
Fireflies.ai

Fireflies.ai

Notetaker with CRM-aware conversation intelligence.

Version
2025
Cost
Free tier · from $18/mo
Model
in-house ASR
Fish Audio

Fish AudioChina

Open-weight multilingual TTS and voice cloning.

Version
S1
Cost
Open weights; API from $15/M chars
Model
Fish Speech S1
Granola

Granola

Local-first meeting notes that augment your own typing.

Version
2025
Cost
Free tier · from $18/mo
Model
Claude
MiniMax Speech

MiniMaxChina

Multilingual TTS and voice cloning at low per-character rates.

Version
2.8
Cost
from $30/mo
Model
MiniMax Music 3.0
Mureka

Kunlun TechChina

7

Music generation with stem export and a public API.

Version
Mureka AI MV Director
Cost
Free tier · from $8/mo
Model
Mureka V8
Otter.ai

Otter.ai

Meeting transcription and AI meeting agents.

Version
2025
Cost
Free tier · from $16.99/mo
Model
in-house ASR
PlayAI

PlayAI

Voice agents and TTS for conversational products.

Version
PlayDialog
Cost
Free tier · from $31.2/mo
Model
PlayDialog
Speechmatics

Speechmatics

Accuracy-focused speech recognition for broadcast and enterprise.

Version
Ursa 2
Cost
Usage-based, self-host option
Model
Ursa 2
Suno

Suno

6

Full songs from a text prompt

Version
v6
Cost
Free tier · from $10/mo
Model
Bark
Udio

Udio

Music generation and remixing

Version
v2
Cost
Free tier · from $10/mo
Model
Udio v2
Whisper

OpenAI

2

Open-weight speech recognition

Version
gpt-transcribe
Cost
Pricing not specified in the document
Model
gpt-transcribe

Recent changes in Audio

  • capability

    MurekaNew capabilities: video creator

    81 tracked82 tracked · +video creator

    source
  • pricing

    SunoEntry price increased to $10/mo

    $8$10

    source
  • pricing

    AssemblyAIEntry price decreased to $0/mo

    $0.15$0

    source
  • pricing

    DeepgramEntry price increased to $4000/mo

    $333.33$4000

    source
  • version

    MurekaMureka moved to Mureka AI MV Director

    MurekaMureka AI MV Director

    source
  • capability

    MurekaNew capabilities: video generator

    80 tracked81 tracked · +video generator

    source
  • pricing

    DeepgramEntry price decreased to $333.33/mo

    $4000$333.33

    source
  • capability

    MurekaNew capabilities: Audio Downloads (MP3/WAV), Video Downloads (MP4)

    78 tracked80 tracked · +Audio Downloads (MP3/WAV), Video Downloads (MP4)

    source
  • capability

    WhisperNew capabilities: Audio to text

    42 tracked43 tracked · +Audio to text

    source
  • capability

    DeepgramNew capabilities: Custom Models

    26 tracked27 tracked · +Custom Models

    source
  • capability

    MurekaNew capabilities: custom soundtrack library, ai video generator

    76 tracked78 tracked · +custom soundtrack library, ai video generator

    source
  • pricing

    AssemblyAIEntry price increased to $0.15/mo

    $0$0.15

    source
  • pricing

    SunoEntry price decreased to $8/mo

    $10$8

    source
  • capability

    AssemblyAINew capabilities: Streaming STT, Synchronous STT

    33 tracked35 tracked · +Streaming STT, Synchronous STT

    source
  • pricing

    DeepgramEntry price increased to $4000/mo

    $333.33$4000

    source
  • capability

    SunoNew capabilities: Vocal generation, Instrumental generation

    80 tracked82 tracked · +Vocal generation, Instrumental generation

    source
  • pricing

    AssemblyAIEntry price decreased to $0/mo

    $0.15$0

    source
  • capability

    MurekaNew capabilities: Royalty-Free Downloads, Custom Soundtracks

    74 tracked76 tracked · +Royalty-Free Downloads, Custom Soundtracks

    source
  • version

    SunoSuno moved to v6

    v5.5v6

    source
  • model

    SunoSuno added v6, v6-mini, v6-wild to its model stack

    Bark, Chirp, Suno v5, v4, v4.5+, v4.5-all, v5, v5.0, v5.5Bark, Chirp, Suno v5, v4, v4.5+, v4.5-all, v5, v5.0, v5.5, v6, v6-mini, v6-wild

    source
  • capability

    AssemblyAINew capabilities: Pre-recorded Speech-to-Text, Streaming Speech-to-Text, Synchronous Speech-to-Text

    30 tracked33 tracked · +Pre-recorded Speech-to-Text, Streaming Speech-to-Text, Synchronous Speech-to-Text

    source
  • pricing

    DeepgramEntry price decreased to $333.33/mo

    $4000$333.33

    source
  • capability

    MurekaNew capabilities: royalty free music, speech generation

    72 tracked74 tracked · +royalty free music, speech generation

    source
  • capability

    ElevenLabsNew capabilities: API Access

    34 tracked35 tracked · +API Access

    source
  • capability

    WhisperNew capabilities: Transcription

    41 tracked42 tracked · +Transcription

    source