Melia 1
Our next-generation multilingual model. Transcribes 56 languages without choosing one up front, switches language mid-sentence.
World class Speech to Text
Our next-generation multilingual model. Transcribes 56 languages without choosing one up front, switches language mid-sentence.
Our highest accuracy. You pick the language for each transcription. Supports custom dictionary and advanced punctuation.
Our latest medical model. For transcription of clinical language: drug names, dosages, abbreviations and procedure terms, with full multilingual support.
Only available as a
Our voice agent model, with built-in turn detection for conversational use cases.
Cost-efficient and feature-rich: the same features as Enhanced at a lower rate, with lower accuracy.
Included for free.
| General purpose | Medical | More models | |||
|---|---|---|---|---|---|
| Model | Melia 1 | Enhanced | Oak 1 | Standard | |
| Accuracy | High | Highest | Highest | Good | |
| Use case | Multilingual and mixed-language audio | Single-language audio, maximum accuracy | Clinical conversations and dictation | High volume, cost-sensitive workloads | |
| 55+ languages | |||||
| Mixed-language transcription | – | – | |||
| Code-switching | – | – | |||
| Healthcare vocabulary | – | – | – | – | |
| Speaker diarization | |||||
| Language hints | – | – | |||
| Language labeling | – | – | |||
| Custom dictionary | – | – | |||
| Smart formatting | |||||
| Timings, word-level | |||||
Extra processing you can bolt onto transcription.
| Add-on | Rate | Melia 1 | Enhanced | Oak 1 | Standard | |
|---|---|---|---|---|---|---|
| PII redaction | On request | – | – | – | – | |
| Entity detection | On request | – | – | |||
| Translation | $0.49 / hr | – | – | |||
| Chapters | $0.30 / hr | – | – | |||
| Topics | $0.15 / hr | – | – | |||
| Summaries | $0.09 / hr | – | – | |||
| Sentiment | $0.09 / hr | – | – | |||
| Audio alignment | On request | – | – |
| Self-serve vs Enterprise | Self-serveStart free · $100 | EnterpriseContact sales |
|---|---|---|
| Concurrent sessions (Real-time) | 50 (2 on Free) | Dynamically adjusts to demand |
| Files per second (Pre-recorded) | 10 | Unlimited |
| Support | Community, portal chat & email | Dedicated CSM + Solutions Engineer |
| Early feature access | – | |
| SOC 2 Type II, ISO/IEC 27001:2022, GDPR, HIPAA | ||
| Trust Center access | ||
| No data logging for model training by default | ||
| Data processing location choice (EU, US, AUS) | ||
| UI and API access | ||
| Workspace account | ||
| Unlimited seats | ||
| Projects | ||
| Domain verification | ||
| API keys | ||
| Management tokens | ||
| SSO | – | Paid add-on, $280/mo |
| SaaS on cloud | ||
| Multi-region cloud | ✓ (STT) | ✓ (STT and TTS) |
| Containers (on-prem) | – | |
| Virtual appliance (on-prem) | – | |
| On-device | – | |
| Hybrid (cloud plus on-prem) | – | |
| Payment methods | Card: pay as you go, top up, subscribe | Invoicing, ACH / Direct Debit |
| Contract | Standard terms of service | Master Services Agreement |
Feature availability can differ by model, product package and deployment. Payment methods and discounts are covered under .
Natural, low-latency speech synthesis for Voice AI use cases.
$0.011per 1,000 characters
First 1 million characters free (about 20 hours of speech).
Built for
Voice AI
Voice assistants, chatbots and IVR systems.
Translation
Real-time translation of live events or media.
Accessibility
Screen readers and assistive technologies.
Content creation
Podcasts, dubbing, audiobooks and voice-overs.
Media production
News broadcasts and automated announcements.