Rankings · Medical speech-to-text
Best medical speech-to-text APIs, ranked
Speech recognition engines scored on the same rubric as scribe APIs, so the gap between a transcript and a chartable note is visible rather than assumed.
Short answer
Which medical speech-to-text API should an integrator choose?
Deepgram leads this group at 7.6 of 10 on developer experience, streaming latency and commercial transparency. The more consequential decision comes first: every engine here returns words, not documentation. Choosing ASR means owning note generation, template adherence, clinician review and clinical-safety review yourself — typically two to four engineering quarters plus permanent ongoing load.
Cite as: Medical speech-to-text API ranking, Compare Healthcare API, last reviewed 2026-09-01.
Key findings
- Published per-minute pricing is the norm in this group and the exception among scribe vendors — a structural transparency difference, not a coincidence.
- Every major cloud ASR provider will sign a BAA, but default retention and opt-out behaviour for model improvement differs materially and must be configured, not assumed.
- Streaming latency claims are published far more often than accuracy methodology, because latency is measurable and reproducible while WER as published is not.
- Specialty vocabulary is usually a customisation surface you populate, not a capability you receive.
The ranking
Rubric v1.0, applied identically to ASR engines and scribe APIs.
| # | Vendor | Weighted score | API22% | Accuracy20% | Interop16% | Compliance16% | Latency10% | Coverage8% | Commercial8% |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Developer speech-to-text API Best pure speech-to-text developer experience | 7.6/10 | 9.6 | 7.8 | 3.2 | 8.3 | 9.4 | 5.4 | 9.2 |
| 2 | Hyperscaler medical speech service Best if you are already committed to AWS | 7.5/10 | 7.8 | 7.6 | 5.2 | 9.5 | 7.9 | 6.2 | 8.4 |
| 3 | Hyperscaler speech service | 7.3/10 | 7.6 | 7.2 | 5.0 | 9.4 | 8.0 | 5.8 | 8.2 |
| 4 | Developer speech-to-text API | 7.3/10 | 9.2 | 7.4 | 3.0 | 8.0 | 8.6 | 5.2 | 8.9 |
| 5 | Speech-to-text with deployment flexibility Best for on-premise and data-residency constraints | 7.2/10 | 8.2 | 7.7 | 3.0 | 8.8 | 8.1 | 6.6 | 7.8 |
What you still have to build
The honest cost of buying words instead of notes.
- Encounter segmentation and speaker attribution tuned to real exam-room audio.
- Note generation with template adherence per specialty, and a hallucination-control strategy.
- A clinician review and attestation UI, plus an audit trail of what was edited.
- Structured extraction for problems, medications and orders if anything beyond prose is required.
- FHIR write-back, per-EHR app review, and per-customer enablement.
- Ongoing clinical-safety review and incident handling, permanently, as your responsibility.
Evidence & sources
Every factual claim on this page traces to one of the primary references below. Each entry records what it supports and its evidence tier, so documentation can be told apart from judgement.
NIST speech recognition evaluation literature · Methodology · Tier B — published methodology
Supports: Why a headline WER figure without a stated corpus, audio condition and reference-transcript protocol is not comparable across vendors.
U.S. Department of Health & Human Services · Regulation · Tier A — primary documentation
Supports: The contractual clauses a documentation vendor's BAA must contain.
U.S. Department of Health & Human Services · Regulation · Tier A — primary documentation
Supports: Administrative, physical and technical safeguards a vendor handling recorded encounter audio must implement.
AICPA · Standard · Tier A — primary documentation
Supports: What a SOC 2 Type II report does and does not attest to during a vendor security review.
NIST · Standard · Tier A — primary documentation
Supports: A defensible structure for governing an AI documentation feature you ship to clinicians.
HL7 International · Standard · Tier A — primary documentation
Supports: Resource definitions (DocumentReference, Composition, Encounter, Condition, MedicationRequest) that clinical documentation output must map onto.
Source tiers are defined on the methodology page. Outbound links are unaffiliated and carry no commercial relationship.
Frequently asked questions
- What is the best medical speech-to-text API?
- Among raw ASR engines, Deepgram ranks first at 7.6 of 10 on the integrator rubric. But every engine in this group returns words, not a clinical note — the summarisation, template adherence, review workflow and clinical-safety ownership stay with you. Buyer guide.
- Should I buy ASR or a scribe API?
- Buy ASR if you need transcripts, dictation or your own note-generation pipeline and you have clinical NLP capability in-house. Buy a scribe API if what you actually need to ship is a reviewable draft note, because reproducing that from raw transcripts is a multi-quarter project with ongoing clinical-safety obligations. Build vs buy.
- Can I compare vendors on published word error rate?
- Not reliably. Published WER figures in this category omit the corpus, audio conditions or reference-transcript protocol needed to reproduce them. Run a clinically weighted error count on a few hundred of your own de-identified encounters instead. How to test.
- Do general-purpose speech APIs work for healthcare?
- Sometimes — for clear single-speaker dictation with light medical vocabulary. They degrade fastest on drug names, dosages, multi-party rooms and accented speech, which is exactly where a documentation error becomes a clinical risk. Healthcare ASR guide.