Speechmatics Pricing
From startups to enterprise - scalable pricing with the accuracy, support and control you need. Start with $100 in credit, no card required.
Pro
No commitment required55+ languages
$100 in credit to get started, add a card when you're ready
50 concurrent real-time sessions
10 file jobs per second
Multi-region cloud options
Sign up now to lock in lowest price
Low-latency (ideal for voice agents)
English - more languages coming soon
Enterprise
Get in touch55+ languages
No rate limits
Privacy-first deployment options
Multi-region cloud options
Custom models
SaaS or On-premises deployment
Lowest-latency, highest privacy with STT & TTS in your environment
Highest concurrency
Custom voice development
Custom language development
Built-in value, whatever your needs.
Your data is encrypted in transit and at rest, with compliance-ready infrastructure built for peace of mind.
Transcribe in 55+ languages and dialects, with the market leading accuracy across the board - reaching over 4 billion people.
Fine-tune transcription with custom vocabularies, formatting rules, and flexible deployment options to fit your workflow.
Deliver the best outcomes with the expert guidance from your dedicated Customer Success Manager and Solutions Engineer.
Your data is encrypted in transit and at rest, with compliance-ready infrastructure built for peace of mind.
Transcribe in 55+ languages and dialects, with the market leading accuracy across the board - reaching over 4 billion people.
Fine-tune transcription with custom vocabularies, formatting rules, and flexible deployment options to fit your workflow.
Deliver the best outcomes with the expert guidance from your dedicated Customer Success Manager and Solutions Engineer.
Compare plans
Pricing
| Speech-to-Text models | |||
|---|---|---|---|
Batch Melia 1 | - | $0.129/hr | Custom |
Batch Standard | - | $0.24/hr | Custom |
Batch Enhanced | - | $0.40/hr | Custom |
Real-time Standard | - | $0.24/hr | Custom |
Real-time Enhanced | - | $0.43/hr | Custom |
| Discounts | - | 20% discount over 500 hr/month | Volume discounts available |
Batch Melia 1 | ||
|---|---|---|
| - | $0.129/hr | Custom |
Batch Standard | ||
|---|---|---|
| - | $0.24/hr | Custom |
Batch Enhanced | ||
|---|---|---|
| - | $0.40/hr | Custom |
Real-time Standard | ||
|---|---|---|
| - | $0.24/hr | Custom |
Real-time Enhanced | ||
|---|---|---|
| - | $0.43/hr | Custom |
| Discounts | ||
|---|---|---|
| - | 20% discount over 500 hr/month | Volume discounts available |
| Speech-to-Text bolt-ons | |||
|---|---|---|---|
Translation | - | $0.65/hr | Custom |
Summaries | - | $0.12/hr | Custom |
Chapters | - | $0.40/hr | Custom |
Sentiment | - | $0.12/hr | Custom |
Topics | - | $0.20/hr | Custom |
Translation | ||
|---|---|---|
| - | $0.65/hr | Custom |
Summaries | ||
|---|---|---|
| - | $0.12/hr | Custom |
Chapters | ||
|---|---|---|
| - | $0.40/hr | Custom |
Sentiment | ||
|---|---|---|
| - | $0.12/hr | Custom |
Topics | ||
|---|---|---|
| - | $0.20/hr | Custom |
| Text-to-Speech | |||
|---|---|---|---|
Text-to-Speech | - | $0.011/1k characters | Custom |
Text-to-Speech | ||
|---|---|---|
| - | $0.011/1k characters | Custom |
Speech-to-Text features
Global language coverage | 55+ languages | 55+ languages | 55+ languages |
Standard and Enhanced accuracy | |||
Industry-leading accent coverage | |||
Real-time latency <1s | |||
Language identification | |||
Speaker diarization | |||
Custom dictionary | |||
Precise timestamps | |||
Advanced punctuation and casing | |||
Numeral formatting | |||
Profanity and disfluency detection | |||
Multi-channel support | |||
Subtitle formatting options | |||
Audio events | |||
Audio alignment | - | - | |
Early access to new features | - | - |
Global language coverage | ||
|---|---|---|
| 55+ languages | 55+ languages | 55+ languages |
Standard and Enhanced accuracy | ||
|---|---|---|
Industry-leading accent coverage | ||
|---|---|---|
Real-time latency <1s | ||
|---|---|---|
Language identification | ||
|---|---|---|
Speaker diarization | ||
|---|---|---|
Custom dictionary | ||
|---|---|---|
Precise timestamps | ||
|---|---|---|
Advanced punctuation and casing | ||
|---|---|---|
Numeral formatting | ||
|---|---|---|
Profanity and disfluency detection | ||
|---|---|---|
Multi-channel support | ||
|---|---|---|
Subtitle formatting options | ||
|---|---|---|
Audio events | ||
|---|---|---|
Audio alignment | ||
|---|---|---|
| - | - | |
Early access to new features | ||
|---|---|---|
| - | - | |
Speech-to-Text bolt ons
Translation | |||
Summaries | |||
Chapters | |||
Sentiment | |||
Topics | |||
Early access to new capabilities | - | - |
Translation | ||
|---|---|---|
Summaries | ||
|---|---|---|
Chapters | ||
|---|---|---|
Sentiment | ||
|---|---|---|
Topics | ||
|---|---|---|
Early access to new capabilities | ||
|---|---|---|
| - | - | |
Speech-to-Text deployment
SaaS | |||
Private Cloud | - | - | |
Container | - | - | |
Virtual Appliance | - | - | |
On-Device | - | - | |
GPU & CPU based models | - | - | |
Multi-region cloud | US, EU or Australia | US, EU or Australia | US, EU or Australia |
SaaS | ||
|---|---|---|
Private Cloud | ||
|---|---|---|
| - | - | |
Container | ||
|---|---|---|
| - | - | |
Virtual Appliance | ||
|---|---|---|
| - | - | |
On-Device | ||
|---|---|---|
| - | - | |
GPU & CPU based models | ||
|---|---|---|
| - | - | |
Multi-region cloud | ||
|---|---|---|
| US, EU or Australia | US, EU or Australia | US, EU or Australia |
Speech-to-Text Rate Limits
Real-time session concurrency | 2 sessions | 50 sessions | Unlimited |
Batch job creation | 1 job per second | 10 jobs per second | Unlimited |
Voice Agent conversation concurrency | 3 conversations | 6 conversations | Unlimited |
Real-time session concurrency | ||
|---|---|---|
| 2 sessions | 50 sessions | Unlimited |
Batch job creation | ||
|---|---|---|
| 1 job per second | 10 jobs per second | Unlimited |
Voice Agent conversation concurrency | ||
|---|---|---|
| 3 conversations | 6 conversations | Unlimited |
Text-to-Speech features
Early access to new features | - | - | |
On-premises deployment | - | - |
Early access to new features | ||
|---|---|---|
| - | - | |
On-premises deployment | ||
|---|---|---|
| - | - | |
Text-to-Speech deployment
SaaS | |||
Private Cloud | - | - | |
Container | - | - | |
Virtual Appliance | - | - | |
On-Device | - | - | |
GPU & CPU based models | - | - | |
Multi-region cloud | - | - | US, EU or Australia |
SaaS | ||
|---|---|---|
Private Cloud | ||
|---|---|---|
| - | - | |
Container | ||
|---|---|---|
| - | - | |
Virtual Appliance | ||
|---|---|---|
| - | - | |
On-Device | ||
|---|---|---|
| - | - | |
GPU & CPU based models | ||
|---|---|---|
| - | - | |
Multi-region cloud | ||
|---|---|---|
| - | - | US, EU or Australia |
Service
Online email support | - | ||
Prioritized email support | - | - | |
Dedicated Customer Success Manager | - | - | |
Dedicated Solutions Engineer | - | - | |
Custom models | - | - | |
Customer community | - | - |
Online email support | ||
|---|---|---|
| - | ||
Prioritized email support | ||
|---|---|---|
| - | - | |
Dedicated Customer Success Manager | ||
|---|---|---|
| - | - | |
Dedicated Solutions Engineer | ||
|---|---|---|
| - | - | |
Custom models | ||
|---|---|---|
| - | - | |
Customer community | ||
|---|---|---|
| - | - | |
FAQs
It's an opt-in programme that takes 33% off Speech to Text rates. By opting in, you grant permission for your audio and transcripts to be used to help improve our models; however, opting in does not guarantee that your data will be used for training. We may, at our discretion, store and use some, all, or none of the data you've made available. The discount applies regardless of whether your data is ultimately stored and used.
The programme is off by default, can be reversed at any time, and applies only to future usage from the point of opt-in or subsequent opt-out. It doesn't apply to Text to Speech.
Volume discounts are automatically applied on any billable usage above 500 hours for each type of Speech-To-Text in a given month. For example, if you use 800 hours of Real-time enhanced accuracy and 400 hours of Real-time standard accuracy, you will be billed as follows:
500 hours Real-time enhanced accuracy at base rate
300 hours Real-time enhanced accuracy at 20% discount
400 hours Real-time standard accuracy at base rate
Additional discounts are available starting from 24,000 hours usage per year. Speak to us to find out more!
We bill Pro tier customers on the 1st of each month for the previous months usage. Costs are billed to the second, based on the cost per hour.
Billing for Enterprise customers is on a custom basis.
Yes, absolutely! Sign up and you'll get $100 in credit to try our award-winning technology — no card required.
When you've used your credit, simply add your credit card details in your account settings to keep going.
Credits are how usage is priced under our new model. 1 credit equals $1, so your credit balance maps directly to the numbers on your invoice.
New accounts start with $100 in credit, no card required. Once it's used, add a card in your account settings to switch to pay-as-you-go: usage from then on is billed at the standard rate shown in the pricing table above. Without a card on file, service access pauses once your credit balance reaches zero — adding one keeps you running without interruption.
Turn on Model Training in your account settings for a discount on usage while it's active.
We offer three proprietary transcription models, available to all customers:
Enhanced: when accuracy matters most, our Enhanced model delivers our highest accuracy across all our languages.
Standard: when you need strong accuracy but file turnaround time or cost control are the priorities. Note that Standard does not provide turnaround time benefits when using Realtime speech-to-text.
Melia 1: when your audio contains more than one language, our Melia 1 model transcribes multilingual speech in a single transcript, including speakers who switch language mid-conversation, with no need to select a language. It matches Standard on accuracy and is available for Batch transcription.
Different models suit different jobs, so you can always choose the right one for the task at hand.
Our AI model supports 55+ languages for transcription, with 69 pairs supported for AI translation.
Transcription
Arabic - Bashkir - Basque - Belarusian - Bengali - Bulgarian - Cantonese - Catalan - Croatian - Czech - Danish - Dutch - English - Esperanto - Estonian - Finnish - French - Galician - German - Greek - Hebrew - Hindi - Hungarian - Indonesian - Interlingua - Irish - Italian - Japanese - Korean - Latvian - Lithuanian - Malay - Maltese - Mandarin (Traditional - & - Simplified) - Marathi - Mongolian - Norwegian - Persian - Polish - Portuguese - Romanian - Russian - Slovak - Slovenian - Spanish - Swahili - Swedish - Tamil - Thai - Turkish - Ukrainian - Urdu - Uyghur - Vietnamese - Welsh
Translation
Bulgarian - Catalan - Croatian - Czech - Danish - Dutch - English - Estonian - Finnish - French - Galician - German - Greek - Hindi - Hungarian - Indonesian - Italian - Japanese - Korean - Latvian - Lithuanian - Malay - Mandarin (Traditional - & - Simplified) - Polish - Portuguese - Romanian - Russian - Slovak - Slovenian - Spanish - Swedish - Turkish - Ukrainian - Vietnamese - Bokmål > Nynorsk
Multilingual
Arabic - Spanish - Mandarin - Malay - Tamil
We're certified to SOC 2 Type II, ISO/IEC 27001:2022, GDPR and HIPAA. Full details and additional documents (such as our DPA and Sub-BAA) are available in the Trust Centre.
If you have not opted in to the model training discount programme, Realtime audio is never stored, and Batch data auto-deletes after 7 days or sooner via the API, and your data is never used to improve our models.
Please feel free to email us at hello@speechmatics.com - we're here to help!