Talking to your own Mac shouldn't cost a subscription. Hold one key, yap, release: clean text at your cursor, in any app, and your voice never leaves the machine.
$0 · no account · no cloud · also for windows & linux ↓the 2.0 app, live in your browser · hover the rail, click around · the real one types at your cursor
illustrative of what dictation tends to cost, rounded on purpose. we are not comparing yapping to any particular product; prices vary by vendor and change often. the zero does not. much of that transcription happens in a data center. yours happens in the metal you already paid for.
bars float above your dock and dance with your voice. release, they fade. no window, no chrome.
cleanup prompts you can read and edit. styles per app. your names spelled right. zero setup by default.
history shows raw and cleaned side by side. if the AI ever changes your meaning, you'll see it. raw words win on doubt.
Your words appear above the Dock as you speak, in a pill that stays through cleanup and flashes green when the text lands. You never talk blind.
Transcribe whatever your Mac is playing, then one click turns an hour of meeting into a TL;DR, key points, and action items. Fully local, kept in your transcripts library.
Words, pace, streaks, time saved versus typing, and where you yap most. Computed on your Mac, never phoned home, never your words.
Double-tap the globe key to talk without holding. Say "send it" and the message fires itself.
Select text, hold the key, say "make this more formal". Rewritten in place.
Audio streams into the recognizer while you talk. Text lands the moment you release.
Apple's on-device model cleans transcripts out of the box. Ollama and custom endpoints for tinkerers.
Dictating to Claude or ChatGPT stays terse and keeps paths, flags, and identifiers verbatim.
Drop audio or video on the menu bar icon; transcript arrives faster than real time.
Correct yourself mid-sentence and only the final version survives.
"New paragraph", "comma", "scratch that". Say the punctuation, get the punctuation.
A field that says it holds a password never starts the mic, and nothing is stored.
Speak one language, paste another. Locally, on the same model that cleans your text.
A live meter, your device and format, and a self test that runs the whole path. Bug reports write themselves.
An icon rail that expands on hover, seven sections, and your local model always in view. Try it above.
Build from source (zero Gatekeeper friction) or grab the zip from releases.
The setup assistant walks you through mic, input monitoring, and accessibility with live checkmarks.
That's it. That's the manual. There is no step four.
A dictation app hears everything. That's why this one can't phone home.
complete list of network connections: localhost (your Ollama, if you use it) + GitHub, to check for updates (once a day and on demand) and to download the update when you click Update Now. the windows and linux betas also download their speech models from Hugging Face, once. read the source and verify.
macOS 26 on Apple Silicon, 64-bit Windows 10/11, or a 64-bit Linux desktop (X11 or Wayland) on a 2024-era distro or newer (Ubuntu 24.04+, Fedora 40+, Debian 13+) for the betas. On the Mac, cleanup works out of the box with Apple's on-device model; Ollama and OpenAI-compatible endpoints are there for tinkerers. Without any of them, yapping pastes accurate raw transcripts.
It shouldn't anymore: releases from 2.0 onward are signed with a Developer ID and notarized by Apple, so macOS opens them without complaint. If you're on an older build, either update from inside the app (which skips Gatekeeper entirely) or allow it once under System Settings, Privacy & Security. Building from source always works too; the code is public, read it first if you like.
It's a beta, rebuilt in Rust around the same idea. Hold Ctrl+Win, speak, release; the transcript pastes at your cursor. Transcription runs fully on-device (the speech model downloads once on first run, about 640 MB), and the app now carries most of the Mac feature set: cleanup via local Ollama or a custom endpoint, history, settings, Listen mode for system audio, and file transcription. The recording overlay and per-app styles are still on the way. The installer isn't code-signed yet, so SmartScreen will warn once: choose More info, then Run anyway, or read the source first.
Also a beta, rebuilt in Rust and GTK4 for X11 and Wayland. Hold Right Alt (rebindable), speak, release; hands-free double-tap, per-app styles, the personal dictionary, and history all made the trip. Speech runs on-device via faster-whisper (the model downloads once on first dictation), cleanup via your local Ollama. Grab the AppImage, or the .deb / .rpm from releases; two one-time setup steps (input group, speech engine) are walked through by the built-in Setup Assistant, or read the source first.
Audio is never written to disk; it streams from the mic straight into the on-device recognizer. Your last 200 dictations are kept in a local, searchable history file so you can inspect what cleanup changed, and you can clear it any time.
Anywhere you can type: browsers, editors, chat apps, terminals. Terminals and code editors get verbatim mode by default so cleanup never touches your commands.
It sits in the corner doing almost nothing, your thumb already knows where it is, and holding it feels like a walkie-talkie. Set "Press globe key to: Do Nothing" in System Settings and yapping takes it from there. Or double-tap it and go hands-free. On Windows the chord is Ctrl+Win, on Linux it's Right Alt (rebindable): same walkie-talkie, different corner.
One download, one permission screen, zero dollars. The $720 stays yours.
macOS 26 on Apple Silicon · 64-bit Windows 10/11, beta · 64-bit Linux, X11 & Wayland, beta