Rubric v1.0 · Last reviewed 2026-09-01
Methodology: how every score on this site is produced
Published in enough detail that you can disagree precisely. If you cannot reproduce a ranking from this page, the page is defective — tell us.
Short answer
How does this site rank clinical documentation vendors?
Each vendor is scored 0–10 on seven criteria, and each sub-score is multiplied by a fixed published weight: API and SDK maturity 22%, clinical accuracy and output quality 20%, EHR and FHIR interoperability 16%, compliance and security 16%, latency and streaming 10%, specialty and template coverage 8%, commercial terms and transparency 8%. Evidence is tiered A–C, undisclosed items are published as undisclosed, and every vendor page carries a limitations section.
Cite as: Scoring methodology, Compare Healthcare API, last reviewed 2026-09-01.
The seven criteria and their weights
| Criterion | Weight | The question it answers | Evidence we look for |
|---|---|---|---|
| API & SDK maturity | 22% | Can an engineering team ship against this without a services contract? | Public developer documentation; Self-serve signup availability; Published SDK repositories |
| Clinical accuracy & output quality | 20% | Does the transcript and the resulting note hold up on real clinical audio? | Published benchmark methodology; Documented medical vocabulary support; Diarization documentation |
| EHR & FHIR interoperability | 16% | How much integration work stands between the output and a chart? | FHIR documentation; Published integration guides; Structured output schema |
| Compliance & security posture | 16% | Can this survive your customer's security review? | BAA availability; Trust centre / certification listings; Data processing terms |
| Latency & streaming behaviour | 10% | Is it fast enough for the interaction you are building? | Streaming endpoint documentation; Published rate limits; Documented turnaround targets |
| Specialty & template coverage | 8% | Will it produce the note your users actually write? | Template documentation; Specialty listings; Language support documentation |
| Commercial terms & transparency | 8% | Can you price your own product on top of it? | Public pricing page; Documented pricing model; Stated white-label / reseller terms |
API & SDK maturity 22%
Public reference documentation, sandbox or free-tier access, self-serve keys, versioning policy, webhooks, streaming and batch endpoints, official SDKs, and error semantics. Vendors that require a sales conversation before a developer can read the docs score low here regardless of model quality.
Clinical accuracy & output quality 20%
Medical vocabulary handling, drug and dosage fidelity, speaker diarization in multi-party encounters, robustness to accents and ambient noise, and — for scribe products — the structure and clinical usefulness of the generated note rather than just word error rate. We record whether a vendor publishes methodology for any accuracy claim it makes.
EHR & FHIR interoperability 16%
HL7 FHIR resource support, SMART on FHIR launch, structured output (problem, medication and order candidates rather than prose only), coding support, and whether the vendor is EHR-agnostic or effectively tied to a small set of platforms. For an integrator, a vendor whose distribution depends on its own EHR partnerships is a strategic risk, not a feature.
Compliance & security posture 16%
HIPAA Business Associate Agreement availability and how easily it is obtained, SOC 2 Type II, HITRUST, data residency options, retention and deletion controls, audit logging, subprocessor transparency, and an explicit contractual position on whether customer PHI is used to train models. Undisclosed items are recorded as undisclosed.
Latency & streaming behaviour 10%
Real-time streaming support, partial-result behaviour, time from encounter end to finished note, concurrency limits, and documented rate limits. Genuinely interactive experiences need streaming; asynchronous documentation workflows tolerate batch turnaround. We assess against the documented mode rather than a single headline number.
Specialty & template coverage 8%
Breadth of specialty support, note formats (SOAP, H&P, progress, procedure, behavioural health, veterinary), configurability of templates by you rather than by the vendor, multi-language and multi-speaker handling, and support for non-visit audio such as dictation and telehealth recordings.
Commercial terms & transparency 8%
Published pricing or a published pricing model, per-minute or per-encounter clarity, minimum commitments, white-label and embedding rights, whether the vendor competes with you for the end customer, and contract flexibility at pilot scale. Vendors with no public pricing signal cannot be modelled by a buyer without a sales cycle.
Evidence tiers
Every claim on this site sits in one of four buckets, and the bucket is disclosed.
| A | Vendor documentation | Public developer docs, pricing pages, trust centres and API references, read directly. |
|---|---|---|
| B | Published methodology | Benchmarks, whitepapers or peer-reviewed work where the method is described well enough to assess. |
| C | Structural inference | Conclusions drawn from how a product is packaged and sold — labelled as judgement, never as fact. |
| — | Not published | The vendor does not disclose it publicly. Recorded as undisclosed rather than estimated. |
The scoring procedure
Six steps, in order, applied identically to every vendor.
1Define the buyer
Scores are meaningless without a buyer. Ours is a software team embedding clinical documentation into a product it owns and resells. Every weight follows from that.
2Collect only checkable evidence
Read the vendor's public developer documentation, pricing page, trust centre and data processing terms. Record what is published and, crucially, what is not. Claims made only in sales conversations are not evidence.
3Score each criterion 0-10 independently
Score each of the seven criteria on its own before looking at the total, so a strong brand cannot lift an unrelated dimension.
4Apply the published weights
Multiply each sub-score by its weight and sum: API 22%, accuracy 20%, interoperability 16%, compliance 16%, latency 10%, coverage 8%, commercial 8%.
5Publish limitations for every vendor
Including the vendor ranked first. A page with no limitations section is marketing, not evaluation.
6Version and date the result
Stamp the rubric version and review date. Weight changes produce a new rubric version rather than a silent rescore.
Current weighted results
The output of the procedure above, in one table.
| 1 | Twofold | 8.9 | Embeddable AI scribe API |
|---|---|---|---|
| 2 | Nuance Dragon Copilot / DAX | 7.8 | Incumbent enterprise platform |
| 3 | Deepgram | 7.6 | Developer speech-to-text API |
| 4 | AWS HealthScribe / Transcribe Medical | 7.5 | Hyperscaler medical speech service |
| 5 | Azure AI Speech | 7.3 | Hyperscaler speech service |
| 6 | AssemblyAI | 7.3 | Developer speech-to-text API |
| 7 | Suki | 7.2 | Clinician-facing scribe with a platform offering |
| 8 | Speechmatics | 7.2 | Speech-to-text with deployment flexibility |
| 9 | Abridge | 6.9 | Enterprise ambient scribe |
| 10 | Ambience Healthcare | 6.8 | Enterprise ambient platform |
Known limits of this method
Stated so that they are not discovered as gotchas.
- Sub-scores are informed judgements, not measurements. Two careful analysts reading the same documentation could differ by half a point on several criteria.
- Accuracy is the weakest criterion in the rubric, because the public evidence base does not exist. We score documented posture, not measured performance, and label it as such.
- Vendors that publish little are penalised on transparency. That is deliberate — an integrator cannot plan against undocumented behaviour — but it is not the same as poor product quality.
- Enterprise products are scored on the integrator use case, which is not the use case they were built for. Their pages say so explicitly.
Frequently asked questions
- How are the rankings on this site calculated?
- Each vendor is scored 0-10 on seven criteria, then each sub-score is multiplied by its published weight and summed: API and SDK maturity 22%, clinical accuracy 20%, EHR and FHIR interoperability 16%, compliance and security 16%, latency and streaming 10%, specialty and template coverage 8%, commercial terms and transparency 8%.
- What counts as evidence?
- Tier A is the vendor's own public documentation, pricing and trust-centre material. Tier B is published methodology such as benchmarks or peer-reviewed work described well enough to assess. Tier C is structural inference from how a product is packaged and sold, always labelled as judgement. Anything a vendor does not disclose is recorded as not disclosed.
- Why do you not publish word error rates?
- Because no vendor in this category publishes a corpus, audio-condition set and reference-transcript protocol complete enough to reproduce a figure. Publishing unreproducible numbers would make this site less accurate, not more. Measure on your own audio instead. How to test accuracy yourself.
- Can I use a different weighting?
- Yes, and you should if your buyer differs from ours. The vendor selector recomputes every score under your weights and produces a shareable, citation-ready report of the result. Open the selector.
- How often is this updated?
- Vendor facts are re-checked against primary sources on each review cycle, and every page carries its last-reviewed date. Weight or definition changes increment the rubric version. Corrections policy.
Evidence & sources
The external standards and definitions this method leans on. Vendor-specific sources are listed on each vendor and guide page.
NIST speech recognition evaluation literature · Methodology · Tier B — published methodology
Supports: Why a headline WER figure without a stated corpus, audio condition and reference-transcript protocol is not comparable across vendors.
NIST · Standard · Tier A — primary documentation
Supports: A defensible structure for governing an AI documentation feature you ship to clinicians.
AICPA · Standard · Tier A — primary documentation
Supports: What a SOC 2 Type II report does and does not attest to during a vendor security review.
- [4]HITRUST CSF
HITRUST Alliance · Standard · Tier A — primary documentation
Supports: The certification many health systems require of documentation vendors handling PHI at scale.
U.S. Department of Health & Human Services · Regulation · Tier A — primary documentation
Supports: The contractual clauses a documentation vendor's BAA must contain.
HL7 International · Standard · Tier A — primary documentation
Supports: Resource definitions (DocumentReference, Composition, Encounter, Condition, MedicationRequest) that clinical documentation output must map onto.
U.S. Food & Drug Administration · Regulation · Tier A — primary documentation
Supports: Where documentation assistance ends and regulated clinical decision support begins — the line an integrator must not cross accidentally.
Source tiers are defined on the methodology page. Outbound links are unaffiliated and carry no commercial relationship.