Evaluation

How we know
it works.

Every change to dictation is graded on one suite, run through the shipped app rather than a copy of it. This is the whole record, the runs we are proud of and the ones we are not. Nothing here is hand-picked.

Latest run · 6 July 2026, 13:56 · eval-engine · 0169abd

96%of cases cleaned correctlylight mode · typed text · 89 cases
231 msfor a typical cleanup389 ms at the 90th percentile
3.4 to 23.5%of words misheard, best condition to worstnever quoted as one number, see below
0bytes of audio that leave the Macevery number here was measured on device

Where the hearing gets hard

The average share of words the microphone gets wrong, by recording condition. Clean human speech is the number to trust. The noisy room and the Bluetooth headset are where we currently do worst, and they are on this page for exactly that reason.

Clean human speech20 clips · real recordings3.4%
Bluetooth headset15 clips · HFP, the worst microphone people use16.3%
Synthetic voice89 clips · generated speech19.2%
Noisy room16 clips · speech mixed with room noise23.5%
105failures where the microphone misheard the words
19failures where the cleanup itself slipped

How each version scored

Every stable and beta build grades itself as it is published, against this same suite. So the question is not how some commit did, it is how the copy you are running did.

No released build has been graded yet. From the next release on, every stable and beta build runs this suite as it is published and its score appears here.

What happens to one clip

Every case rides the same chain the app uses, then is graded at the end.

  1. 01

    Speak

    Real recorded human speech, plus a synthetic voice, a noisy-room mix and a Bluetooth-headset mix of every written case.

  2. 02

    Transcribe

    The exact speech model the app ships, running on a Mac. No stand-in, no cloud, no second implementation.

  3. 03

    Clean up

    The shipped cleanup pass runs inside the real app, behind the same faithfulness guard your copy has.

  4. 04

    Grade

    Keyword rules, word error against the reference, and the guard's verdict. Every failure is attributed to hearing or to cleanup.

Is it getting better?

Share of cases that pass, run by run, oldest first. Typed and spoken cases are drawn apart on purpose: a spoken case can fail because the microphone misheard it, so one line across both would hide which half moved.

0255075100
Typed text, the cleanup on its ownSpoken, transcription and cleanup together

Every run

Each row is one full pass of the suite against one build. Open a row to read every case in it, marked up.

Check it yourself

The cases, the prompts and the scorer are in the repository, and this page is generated from the stored runs rather than written by hand.