/
VonFlow
1

Hold a key, speak, let go

The text lands where the cursor already was, in whatever was already open. From releasing the key to the words appearing takes about a second and a half at the median, spoken at a median of 123 words a minute; the transcription inside that runs at about 13 milliseconds per second of audio, roughly 76 times faster than real time. 506 dictations have gone through it so far: 70,218 words across 10.3 hours of speech on 40 days of work.

2

It hears two languages without being told

The language is read per dictation rather than set once in a preference: 439 English and 67 Dutch so far, with nothing to flip between them. Switching halfway through the morning costs nothing, because there was never a switch.

3

Record a call, read it afterwards

A call is captured as two separate tracks and written up with the speakers labelled, so the transcript reads as a conversation rather than a wall. 8 calls, 1.6 hours and 905 labelled segments since 10 July 2026. The one accuracy soak on record, on 31 July 2026, put word error at roughly 2 to 3 percent, clustered in proper nouns the model had never met.

4

Drop in a voice note

A voice note, or any recording the Mac can open, goes down the same path and comes back as text with no export step in the middle. What somebody left as audio becomes something you can read at your own pace.

5

It tidies up what was actually said

Fillers come out, 2,008 of them so far and about four per dictation. Punctuation goes in. A correction spoken mid-sentence is applied rather than transcribed. Names it has never heard are fixed against a dictionary of 39 entries, every one of them fed in by hand.

6

Nothing leaves the Mac

No server in the path, no account, no upload. The longest single dictation so far ran 1,517 words in one breath of 10 minutes and 43 seconds, and not a syllable of it touched the network.