01 The problem
Everyone who uses AI writing tools hits the same wall eventually: the output is competent but sounds nothing like them. It's flat, generic, over-polished — recognisably "AI." Fixing that by hand, every time, defeats the point of the tool.
The usual fix is to paste a few writing samples into a prompt. But most people express themselves more naturally by speaking than by curating written samples — and the way you talk carries your rhythm, your phrasing, your fillers, the texture that makes a voice yours.
Capture voice through speech, not written samples. VoiceSkill records how you actually talk, transcribes it, and distils it into a reusable profile — a portable style guide you can hand to any AI tool so it writes in your voice from the start.
02 How it works
Record naturally
You talk the way you normally would — the most honest signal of your actual voice.
Groq Whisper
Speech turned to text accurately, preserving your real phrasing rather than cleaning it up.
Style profile
The transcript analysed into a reusable voice profile — tone, rhythm, characteristic patterns.
Export & library
Export the profile as a portable spec; multiple users each keep their own in a profile library.
03 The design decisions
- Voice-in, not text-in. The whole premise rests on speaking being a truer signal than curated writing samples. It's the differentiator, and it's why Whisper is central rather than incidental.
- A portable profile, not a walled garden. The output is an exportable spec you can use in other tools — VoiceSkill is a capture layer, not a lock-in.
- Multi-user from the start. A profile library means teams or individuals can each maintain a distinct voice — the design assumes voice is personal and plural.
04 What I'd watch next
VoiceSkill is live but hasn't been validated with real users. The open questions:
- Does the profile actually capture voice? The test is whether output using the profile reads as authentically "you" — if it still sounds generic, the distillation step needs work.
- Is speaking really easier than pasting samples? The core bet. If people find recording awkward, the premise needs rethinking.
- Where does a portable profile get used? Understanding which downstream tools people plug it into would shape what the export format should contain.
05 What building it taught me
- Input modality is a product decision. Choosing voice over text wasn't a technical afterthought — it's the entire reason the product exists, and it flows from a real observation about how people express themselves.
- Capture layers should be portable. The value grows when the profile works everywhere, not just inside my app — resisting lock-in was a deliberate call.
- The best signal is the most natural one. People reveal their real voice when they talk, not when they curate — designing for that honesty is the whole game.