CLINICAL TECHNOLOGY, PUT TO THE TEST
AI scribes, under the microscope.
We test AI medical scribes on the same recorded consultations. Explore how they compare on transcription, clinical notes and speed—then look closer at the cases behind the numbers.
Checklists were drafted by AI from each source's human transcript or published answer key, cross-checked against the consulting doctors' own notes where the dataset includes them, and frozen before any notes were captured. Every fact was then audited against its source: none were unsupported or contradicted.
Benchmark · 22 September 2026
7 AI scribes, 11 consultations, 5+ runs each
Every note graded blind. Each figure is the average across runs; select any cell to see every run.
Swipe or scroll the table to see every scribe →
| Measure | Hanah | Heidi1 | Freed2 | CliniScripts3 | Lyrebird4 | Preve5 | PatientNotes |
|---|---|---|---|---|---|---|---|
| Time to a signable noteWait from pressing stop to the graded note, plus review time: how long until the note can be signed if it is reviewed straight away. | |||||||
| Notes ready to sign as writtenShare of notes that captured every checklist fact in full with no hallucination flags, so nothing needed editing. | |||||||
| Editing effort0% means no editing required; 100% means writing the full note by hand. Fixes are weighted by the work they take: partial fact 1, missing fact 2, hallucination flag 3, and +2 when a fact is captured wrongly. | |||||||
| Review time per noteEstimated clinician time to proofread and fix one note (190 wpm reading, 5 s to recall and 40 wpm to type each missing or partial fact, 12 s to verify and delete each hallucination flag, including a wrong fact before it is retyped). | |||||||
| Worst note, any runThe single slowest note to review and fix across every case and run. | |||||||
| Time saved per noteEstimated time to write the note yourself (5 s to recall each checklist fact, typing it at 40 wpm) minus review time. | |||||||
| Significant note errors per runMissed facts, partial facts and hallucinations that a blind AI rater judged could change diagnosis, treatment, follow-up, safety or the legal record, across 11 cases. Significance is a judgement call; the flat count is in Hallucination flags. | |||||||
| Hallucination flags per runEvery claim in the note with no support in the consultation, counted flat: each one counts, however minor. No severity judgement. Across 11 cases. | |||||||
| Edit actions per noteMissing items, partial items and hallucination flags a clinician would need to fix, per note. | |||||||
| Clinical coverageShare of each case checklist captured by the note (partial counts as half). | |||||||
| Significant transcript errors per runMis-heard words that change clinical meaning, rated blind, across 11 cases. | |||||||
| Word error ratePooled word error rate per run against the human reference. Every word counts equally, including fillers. | |||||||
| Stop to prompted noteMedian time from pressing stop to the note written to our instruction, per run. | — | ||||||
| Data per consultAverage network data per consultation, per run. |
- Heidi has 6 runs; every other product has 5.
- Two notes are counted as product failures: Preve day1_01 and PatientNotes day5 in run 2.
- 1 Heidi: Wide variation between runs, additional run included for fine tuning.
- 2 Freed: Notes graded on default template.
- 3 CliniScripts: Notes graded on default template.
- 4 Lyrebird: 10 cases (no French).
- 5 Preve: 9 cases (no French or Hindi); no separate prompted-note step.
Download everything: one JSON file with every run’s scores, transcripts, notes and jury reasons, or every run’s scores as CSV.
Face-off
Pick two scribes
Each rope is pulled toward the scribe that did better on that measure. The further the ribbon moves from the centre line, the bigger the gap. The shaded band is where it would land between each scribe’s best and worst run: a wide band means inconsistent results. Small figures are the run range.
GO ONE LEVEL DEEPER
Results by clinical case
Each consultation compares the scribes on the same audio across every graded run. Open one to read each run's transcript and note, with every error marked.
English consultations
Six English consultation tests.
Language-specific consultations
Separate comparisons with capability exclusions and translation-aware interpretation.
Physician dictation
Scored against verbatim references built from the publisher’s answer keys. Spoken dictation commands are not scored.
Extended-session stress test
A 50-minute synthetic fixture with background distractors. Included in the overall results.