Skip to content

Rubric v1.0 · Last reviewed 2026-09-01

Methodology: how every score on this site is produced

Published in enough detail that you can disagree precisely. If you cannot reproduce a ranking from this page, the page is defective — tell us.

Written by Compare Healthcare API Editorial DeskReviewed by Technical ReviewLast reviewed Rubric v1.0

Short answer

How does this site rank clinical documentation vendors?

Each vendor is scored 0–10 on seven criteria, and each sub-score is multiplied by a fixed published weight: API and SDK maturity 22%, clinical accuracy and output quality 20%, EHR and FHIR interoperability 16%, compliance and security 16%, latency and streaming 10%, specialty and template coverage 8%, commercial terms and transparency 8%. Evidence is tiered A–C, undisclosed items are published as undisclosed, and every vendor page carries a limitations section.

Cite as: Scoring methodology, Compare Healthcare API, last reviewed 2026-09-01.

The seven criteria and their weights

CriterionWeightThe question it answersEvidence we look for
API & SDK maturity22%Can an engineering team ship against this without a services contract?Public developer documentation; Self-serve signup availability; Published SDK repositories
Clinical accuracy & output quality20%Does the transcript and the resulting note hold up on real clinical audio?Published benchmark methodology; Documented medical vocabulary support; Diarization documentation
EHR & FHIR interoperability16%How much integration work stands between the output and a chart?FHIR documentation; Published integration guides; Structured output schema
Compliance & security posture16%Can this survive your customer's security review?BAA availability; Trust centre / certification listings; Data processing terms
Latency & streaming behaviour10%Is it fast enough for the interaction you are building?Streaming endpoint documentation; Published rate limits; Documented turnaround targets
Specialty & template coverage8%Will it produce the note your users actually write?Template documentation; Specialty listings; Language support documentation
Commercial terms & transparency8%Can you price your own product on top of it?Public pricing page; Documented pricing model; Stated white-label / reseller terms
Rubric v1.0. Weights sum to 100 and are identical for every vendor.

API & SDK maturity 22%

Public reference documentation, sandbox or free-tier access, self-serve keys, versioning policy, webhooks, streaming and batch endpoints, official SDKs, and error semantics. Vendors that require a sales conversation before a developer can read the docs score low here regardless of model quality.

Clinical accuracy & output quality 20%

Medical vocabulary handling, drug and dosage fidelity, speaker diarization in multi-party encounters, robustness to accents and ambient noise, and — for scribe products — the structure and clinical usefulness of the generated note rather than just word error rate. We record whether a vendor publishes methodology for any accuracy claim it makes.

EHR & FHIR interoperability 16%

HL7 FHIR resource support, SMART on FHIR launch, structured output (problem, medication and order candidates rather than prose only), coding support, and whether the vendor is EHR-agnostic or effectively tied to a small set of platforms. For an integrator, a vendor whose distribution depends on its own EHR partnerships is a strategic risk, not a feature.

Compliance & security posture 16%

HIPAA Business Associate Agreement availability and how easily it is obtained, SOC 2 Type II, HITRUST, data residency options, retention and deletion controls, audit logging, subprocessor transparency, and an explicit contractual position on whether customer PHI is used to train models. Undisclosed items are recorded as undisclosed.

Latency & streaming behaviour 10%

Real-time streaming support, partial-result behaviour, time from encounter end to finished note, concurrency limits, and documented rate limits. Genuinely interactive experiences need streaming; asynchronous documentation workflows tolerate batch turnaround. We assess against the documented mode rather than a single headline number.

Specialty & template coverage 8%

Breadth of specialty support, note formats (SOAP, H&P, progress, procedure, behavioural health, veterinary), configurability of templates by you rather than by the vendor, multi-language and multi-speaker handling, and support for non-visit audio such as dictation and telehealth recordings.

Commercial terms & transparency 8%

Published pricing or a published pricing model, per-minute or per-encounter clarity, minimum commitments, white-label and embedding rights, whether the vendor competes with you for the end customer, and contract flexibility at pilot scale. Vendors with no public pricing signal cannot be modelled by a buyer without a sales cycle.

Evidence tiers

Every claim on this site sits in one of four buckets, and the bucket is disclosed.

AVendor documentationPublic developer docs, pricing pages, trust centres and API references, read directly.
BPublished methodologyBenchmarks, whitepapers or peer-reviewed work where the method is described well enough to assess.
CStructural inferenceConclusions drawn from how a product is packaged and sold — labelled as judgement, never as fact.
Not publishedThe vendor does not disclose it publicly. Recorded as undisclosed rather than estimated.
Tiering exists so a reader can tell documentation from judgement without reading our minds.

The scoring procedure

Six steps, in order, applied identically to every vendor.

  1. 1Define the buyer

    Scores are meaningless without a buyer. Ours is a software team embedding clinical documentation into a product it owns and resells. Every weight follows from that.

  2. 2Collect only checkable evidence

    Read the vendor's public developer documentation, pricing page, trust centre and data processing terms. Record what is published and, crucially, what is not. Claims made only in sales conversations are not evidence.

  3. 3Score each criterion 0-10 independently

    Score each of the seven criteria on its own before looking at the total, so a strong brand cannot lift an unrelated dimension.

  4. 4Apply the published weights

    Multiply each sub-score by its weight and sum: API 22%, accuracy 20%, interoperability 16%, compliance 16%, latency 10%, coverage 8%, commercial 8%.

  5. 5Publish limitations for every vendor

    Including the vendor ranked first. A page with no limitations section is marketing, not evaluation.

  6. 6Version and date the result

    Stamp the rubric version and review date. Weight changes produce a new rubric version rather than a silent rescore.

Current weighted results

The output of the procedure above, in one table.

1Twofold8.9Embeddable AI scribe API
2Nuance Dragon Copilot / DAX7.8Incumbent enterprise platform
3Deepgram7.6Developer speech-to-text API
4AWS HealthScribe / Transcribe Medical7.5Hyperscaler medical speech service
5Azure AI Speech7.3Hyperscaler speech service
6AssemblyAI7.3Developer speech-to-text API
7Suki7.2Clinician-facing scribe with a platform offering
8Speechmatics7.2Speech-to-text with deployment flexibility
9Abridge6.9Enterprise ambient scribe
10Ambience Healthcare6.8Enterprise ambient platform
Weighted totals under rubric v1.0.

Known limits of this method

Stated so that they are not discovered as gotchas.

  • Sub-scores are informed judgements, not measurements. Two careful analysts reading the same documentation could differ by half a point on several criteria.
  • Accuracy is the weakest criterion in the rubric, because the public evidence base does not exist. We score documented posture, not measured performance, and label it as such.
  • Vendors that publish little are penalised on transparency. That is deliberate — an integrator cannot plan against undocumented behaviour — but it is not the same as poor product quality.
  • Enterprise products are scored on the integrator use case, which is not the use case they were built for. Their pages say so explicitly.

Frequently asked questions

How are the rankings on this site calculated?
Each vendor is scored 0-10 on seven criteria, then each sub-score is multiplied by its published weight and summed: API and SDK maturity 22%, clinical accuracy 20%, EHR and FHIR interoperability 16%, compliance and security 16%, latency and streaming 10%, specialty and template coverage 8%, commercial terms and transparency 8%.
What counts as evidence?
Tier A is the vendor's own public documentation, pricing and trust-centre material. Tier B is published methodology such as benchmarks or peer-reviewed work described well enough to assess. Tier C is structural inference from how a product is packaged and sold, always labelled as judgement. Anything a vendor does not disclose is recorded as not disclosed.
Why do you not publish word error rates?
Because no vendor in this category publishes a corpus, audio-condition set and reference-transcript protocol complete enough to reproduce a figure. Publishing unreproducible numbers would make this site less accurate, not more. Measure on your own audio instead. How to test accuracy yourself.
Can I use a different weighting?
Yes, and you should if your buyer differs from ours. The vendor selector recomputes every score under your weights and produces a shareable, citation-ready report of the result. Open the selector.
How often is this updated?
Vendor facts are re-checked against primary sources on each review cycle, and every page carries its last-reviewed date. Weight or definition changes increment the rubric version. Corrections policy.

Evidence & sources

The external standards and definitions this method leans on. Vendor-specific sources are listed on each vendor and guide page.

  1. NIST speech recognition evaluation literature · Methodology · Tier B — published methodology

    Supports: Why a headline WER figure without a stated corpus, audio condition and reference-transcript protocol is not comparable across vendors.

  2. NIST · Standard · Tier A — primary documentation

    Supports: A defensible structure for governing an AI documentation feature you ship to clinicians.

  3. AICPA · Standard · Tier A — primary documentation

    Supports: What a SOC 2 Type II report does and does not attest to during a vendor security review.

  4. HITRUST Alliance · Standard · Tier A — primary documentation

    Supports: The certification many health systems require of documentation vendors handling PHI at scale.

  5. U.S. Department of Health & Human Services · Regulation · Tier A — primary documentation

    Supports: The contractual clauses a documentation vendor's BAA must contain.

  6. HL7 International · Standard · Tier A — primary documentation

    Supports: Resource definitions (DocumentReference, Composition, Encounter, Condition, MedicationRequest) that clinical documentation output must map onto.

  7. U.S. Food & Drug Administration · Regulation · Tier A — primary documentation

    Supports: Where documentation assistance ends and regulated clinical decision support begins — the line an integrator must not cross accidentally.

Source tiers are defined on the methodology page. Outbound links are unaffiliated and carry no commercial relationship.

Continue