I pay $0 to transcribe video. It runs on my laptop, nothing gets uploaded, and there is no per-minute meter.
I put my whole setup in a repo: local-transcribe.
To install it, open Claude Code and say:
install this repo: https://github.com/mattycartwright/local-transcribe
It sets up the two pieces for you. ffmpeg pulls the audio track out of a video file. mlx-whisper runs OpenAI's Whisper speech model on Apple silicon, which is the part that makes it fast.
Then hand it a file. A 23 minute podcast comes back in 40 seconds. A 43 minute interview takes 3 minutes. You get a markdown transcript with real paragraphs, an srt for captions, and timestamps.
What an hour of audio costs everywhere else:
| Rev | $15.00 |
|---|---|
| Sonix | $10.00 |
| Otter Pro | $8.33/mo, capped at 20 hours |
| OpenAI's API | $0.36 |
| this | nothing, run it a thousand times |
One catch: Apple silicon only, M1 or newer. That chip is the free GPU doing the work.
Cheers,
Matty C