CLINICAL TECHNOLOGY REVIEWEDITION 01 / SEPTEMBER 2026
ScribeStandard.Explore results

CLINICAL TECHNOLOGY, PUT TO THE TEST

AI scribes, under the microscope.

We test AI medical scribes on the same recorded consultations. Explore how they compare on transcription, clinical notes and speed—then look closer at the cases behind the numbers.

Checklists were drafted by AI from each source's human transcript or published answer key, cross-checked against the consulting doctors' own notes where the dataset includes them, and frozen before any notes were captured. Every fact was then audited against its source: none were unsupported or contradicted.

7 scribes tested5+ full runs each, graded blindHow we test

Benchmark · 22 September 2026

7 AI scribes, 11 consultations, 5+ runs each

Every note graded blind. Each figure is the average across runs; select any cell to see every run.

Swipe or scroll the table to see every scribe →

Average result per scribe across all runs. Lower is better except clinical coverage. Select a cell to open the run-by-run spread.
MeasureHanahHeidi1Freed2CliniScripts3Lyrebird4Preve5PatientNotes
Time to a signable noteWait from pressing stop to the graded note, plus review time: how long until the note can be signed if it is reviewed straight away.
Notes ready to sign as writtenShare of notes that captured every checklist fact in full with no hallucination flags, so nothing needed editing.
Editing effort0% means no editing required; 100% means writing the full note by hand. Fixes are weighted by the work they take: partial fact 1, missing fact 2, hallucination flag 3, and +2 when a fact is captured wrongly.
Review time per noteEstimated clinician time to proofread and fix one note (190 wpm reading, 5 s to recall and 40 wpm to type each missing or partial fact, 12 s to verify and delete each hallucination flag, including a wrong fact before it is retyped).
Worst note, any runThe single slowest note to review and fix across every case and run.
Time saved per noteEstimated time to write the note yourself (5 s to recall each checklist fact, typing it at 40 wpm) minus review time.
Significant note errors per runMissed facts, partial facts and hallucinations that a blind AI rater judged could change diagnosis, treatment, follow-up, safety or the legal record, across 11 cases. Significance is a judgement call; the flat count is in Hallucination flags.
Hallucination flags per runEvery claim in the note with no support in the consultation, counted flat: each one counts, however minor. No severity judgement. Across 11 cases.
Edit actions per noteMissing items, partial items and hallucination flags a clinician would need to fix, per note.
Clinical coverageShare of each case checklist captured by the note (partial counts as half).
Significant transcript errors per runMis-heard words that change clinical meaning, rated blind, across 11 cases.
Word error ratePooled word error rate per run against the human reference. Every word counts equally, including fillers.
Stop to prompted noteMedian time from pressing stop to the note written to our instruction, per run.
Data per consultAverage network data per consultation, per run.

Download everything: one JSON file with every run’s scores, transcripts, notes and jury reasons, or every run’s scores as CSV.

GO ONE LEVEL DEEPER

Results by clinical case

View the aggregate →

Each consultation compares the scribes on the same audio across every graded run. Open one to read each run's transcript and note, with every error marked.

Language-specific consultations

Separate comparisons with capability exclusions and translation-aware interpretation.

Physician dictation

Scored against verbatim references built from the publisher’s answer keys. Spoken dictation commands are not scored.

Extended-session stress test

A 50-minute synthetic fixture with background distractors. Included in the overall results.