Medical voice infrastructure

Voice AI built for healthcare.

Add medical speech to the tools you are building—from a weekend prototype to a production healthcare AI product. Use the hosted API or deploy our open model yourself.

For clinicians who code · Healthtech startups · Healthcare AI companies

25 free audio-hours monthlySelf-serve BAANever used for model training
Consultation audio
“Start amoxicillin 500 mg and repeat HbA1c at review.”
✓ Drug and dosage preserved
#1 / 30medical-term accuracy
0.00%drug-name error
97.7%dosage F1 · 86/89 events
Built for the workflow

Turn clinical speech into product data.

Use one medical voice layer for the moments your users already speak through.

Consultations

Ambient documentation

Capture speakers, timing and the medical language needed for notes and downstream agents.

Dictation

Clinical text entry

Give clinicians a faster input method inside EHR, dental, therapy and specialty workflows.

Patient calls

Voice workflows

Feed confirmed medical text into triage, follow-up and care-navigation logic.

Local AI

Private on-device speech

Run the open model in your own application when audio cannot leave the environment.

Hosted API plans

Build for free. Add a card for production.

Every plan gets the complete medical API. Payment protects continuity and raises capacity.

Builder
Free
25 audio-hours includedPooled across batch and live. Resets every month.

For clinicians who code, independent builders and small healthcare AI projects.

  • Complete medical batch and live API
  • Medical vocabulary, speaker labels and timestamps included
  • No card required · no rollover
  • Self-serve BAA and DPA
Create a Builder key
Enterprise
Custom
Contract pricing and capacitySized around your production workload.

For healthcare AI companies that need contracted capacity, guarantees or private deployment.

  • Capacity sized for your workload
  • Private cloud and supported SDK
  • Service-level agreement
  • Security and integration support
Talk to us
No transcription add-onsNever used for model training
Open-source medical speech

Trust it because you can run it.

omi-medical-edge-1 has open weights under CC-BY-4.0 and runs on Mac, NVIDIA CUDA or CPU. Inspect the model, test it on your own clinical audio, and keep every byte inside your environment.

Open weightsBenchmark, inspect and adapt the model yourself.
Runs locallyMac, CUDA and CPU runtimes.
0.6B parametersSmall enough to embed in real products.
Hosted when readyMove to the flagship API without rebuilding your product story.
Benchmark results

Every layer of the medical record, measured.

Overall transcription, medical terminology and dosage accuracy are scored independently on the same clinical audio.

Overall word error · lower is better

Azure5.97
Omi5.99
AWS6.12
ElevenLabs6.17

Medical-term error · lower is better

Omi0.94
ElevenLabs0.97
Google1.11
AssemblyAI1.43

Dosage F1 · higher is better

Omi97.7
Deepgram86.8
ElevenLabs85.4
Azure83.3
Fresh from Omi

Build from evidence, not marketing claims.

Practical guides, reproducible comparisons and open model work for people putting medical speech into products. Follow the RSS feed →

New benchmark update

Three new transcription models on medical audio

Azure, OpenAI and Google joined the same sealed 1,513-clip clinical board.

See what changed →
Guide · updated Aug 2026

How to choose a medical speech-to-text API

Evaluate clinical accuracy, dosage safety, BAA access, speakers, languages and real price.

Read the guide →
Builder guide

HIPAA-compliant speech-to-text: a practical checklist

What a healthcare builder should verify before sending protected health information.

Use the checklist →
Trust center

Built for clinical data.

The hosted API keeps the operational choices visible, while the open model gives you a path to keep audio entirely inside your environment.

01

EU processing

Hosted API audio and transcripts are processed in the European Union.

02

Self-serve agreements

Sign the BAA and DPA in the console on Builder, before sending PHI.

03

1–72 hour retention

Choose how long stored job results remain available for retrieval.

04

Never used for training

Customer audio and transcripts are not used to train Omi models.

Start with your own audio.

Use the hosted API or download the open model.