We don't talk enough about speech-to-text
The new speech-to-text actually works. And what that says about the balance of power between software makers and users.

We don’t talk enough about speech-to-text.
Not the 2015 thing that understood one word in three. The new one. The one that actually works.
Switching from French to English mid-sentence? OK. Dropping “Kubernetes”, “index.ts”, “Loki”? OK. The AI transcribes the exact word.
The advances of the last 12 months have made this genuinely usable. And Mistral just dropped Voxtral, which beats OpenAI on transcription. French pride. 🐓
Today, dictating your text is a niche thing. A few devs, a few early adopters.
But think about it for 2 seconds.
Preparing a PowerPoint presentation? You dictate. Describing an Excel formula? You dictate. Explaining a feature to code? You dictate.
It’s just faster to explain than to type. Especially when more context = better.
In 3 years, I bet most people will dictate everything. Devs just adopt tools faster than everyone else. As usual.
The other interesting thing is what it says about the software industry.
There are apps like SuperWhisper. $9/month. Solid product, well thought out, real distribution strategy. A real business.
I used it for a few weeks. And then I thought: I use 20% of the features. What if I just rebuilt that?
I launched Claude Code in near-autonomy. One day. I did the QA and steered the tech choices. I’ve never written a line of Swift in my life.
Result: Whisper Voice, a native Mac + Windows app. Functional.
Today it’s simple products you can rebuild like this. Tomorrow it’ll be more complex ones. The balance of power between software makers and users is shifting.
Anyway. I put the app open source if you’re interested.
Whisper Voice. Free. Mac + Windows. You enter your API key (OpenAI or Mistral), you speak, the text appears.