Local speech-to-text for macOS: a CLI that transcribes audio with hands the same capability to an agent whose model cannot process audio. No API key. No audio leaves the machine. An agent driven by a model that cannot hear still needs to read recordings. Sending them to a cloud API costs money per minute and puts the audio on someone