← All runs

Evaluation run

self-correction chains

6 July 2026, 13:53 · eval-engine · 232478e · 89 cases

Did it hear the words right?

This run graded typed text only, so there is nothing here to have misheard.

Did it clean up correctly and safely?

96%
Light, ships on by default
85/89 cases passed · 231 ms typical
92%
Polish, experimental
82/89 cases passed · 265 ms typical
When a case fails, whose fault is it?
0 the microphone misheard the words  ·  11 the cleanup itself slipped

Every case

Each case shows the instant on-device cleanup, then what each mode did to it. Added words are marked, removed words are struck. For a spoken case, Heard is what the microphone transcribed and Was said is the reference it is scored against.

89 cases match

real-longTypedclean
Instantso i talked to the team today and we decided that we're going to push the launch to next week because the the testing isn't done and we want to make sure everything is solid
LightSo I talked to the team today and we decided that we're going to push the launch to next week because the the testing isn't done and we want to make sure everything is solid. 704 ms
PolishSo I talked to the team today, and we decided that we're going to push the launch to next week because the the testing isn't done, and we want to make sure everything is solid. 695 ms
real-messageTypedclean
Instantrunning like 10 minutes late traffic is crazy be there as soon as i can
LightRunning like 10 minutes late, traffic is crazy, be there as soon as I can. 397 ms
PolishRunning like 10 minutes late. Traffic is crazy. Will be there as soon as I can. 415 ms
real-noteTypedclean
Instanttodo for tomorrow finish the report call the bank and pick up groceries
LightTodo for tomorrow: finish the report, call the bank, and pick up groceries. 353 ms
PolishTodo for tomorrow: finish the report, call the bank, and pick up groceries. 365 ms
sens-passwordTypedclean
Instantmy wifi password is hunter two three four
LightMy wifi password is hunter two three four. 284 ms
Polishmy Your Wi-Fi password is "hunter two three four". 274 ms
vocab-acronymTypedneeds review
Instantour a p i is getting rate limited
Lightour a p i is getting rate limited 195 msthe cleanup was rejected, so the original was kept
Polishour a p i is getting rate limited 196 msthe cleanup was rejected, so the original was kept
Why?the result is missing “API”
vocab-namesTypedclean
Instantloop in raj and priya on the nvidia integration
LightLoop in Raj and Priya on the Nvidia integration. 259 ms
PolishLoop in Raj and Priya on the NVIDIA integration. 275 ms
vocab-preserveTypedclean
Instantwe deployed the parakeet model to production
LightWe deployed the parakeet model to production. 294 ms
PolishWe deployed the Parakeet model to production. 315 ms
vocab-ragTypedclean
Instantthe rag pipeline retrieves documents before the model answers
LightThe rag pipeline retrieves documents before the model answers. 242 ms
PolishThe Rag pipeline retrieves documents before the model answers. 264 ms
vocab-techTypedclean
Instantwe deploy with kubernetes and terraform on aws
LightWe deploy with Kubernetes and Terraform on AWS. 336 ms
PolishWe deploy with Kubernetes and Terraform on AWS. 268 ms