Speak once.It lands. It stays.It acts.
For people who talk faster than they type. What you say lands in the window you already had open, stays as a file in a folder you named, and can be handed to one agent that acts on it.
- AGPL-3.0
- macOS, Windows and Linux
- cloud, local, your server, enterprise
It lands in whatever window has focus.
- Neovim
- Sublime Text
- IntelliJ IDEA
- Chrome
- Firefox
- Gmail
- Thunderbird
- Discord
- Telegram
- Notion
- Obsidian
- Google Docs
- Linear
- Jira
- GitHub
Not integrations. There is nothing to install into any of these: WordScript types where you type.
Where your words go.
Putting text at the cursor is the easy third. Everyone can do it, and soon everyone will do it for free.
The cursor. Hold the hotkey and speak. The mode decides which text lands, and the mode is the only thing the transform stage reads.
lands at the cursor, then the surface asks what else
Context. A dictation, a meeting, an uploaded file and a pasted link are the same kind of record, and they accumulate in a directory you own.
Plain files, in a directory you named. Your editor, your grep and your agent already know how to open them.
The agent. One orchestrator drives your coding agents and WordScript talks to nobody else. When it cannot decide, it asks you out loud and waits.
Most dictation tools end at the cursor. The ones that keep anything keep it inside their own product. WordScript keeps it in a directory your own tools already open.
Every destination is a connection somebody has to build, authenticate and keep working. Which places your words can go is then their roadmap, not yours.
Connections the product has to build and keep working: 5
It counts what it did for you.
That speaking beats typing is not our finding, and for most of the time computers could listen at all it was not even true. The three measurements below are other people's.
- 1999Speaking lost.spoken13.6wpmtyped32.5wpm
Transcription by voice against keyboard and mouse, 24 users, three commercial recognisers. Karat, Halverson, Horn and Karat, Patterns of Entry and Correction in Large Vocabulary Continuous Speech Recognition Systems, CHI 1999. Figures as reported in Ruan et al. 2016.
- 2016It turned over.spoken161.2wpmtyped53.5wpm
English short messages, 32 participants, with 20.4 percent fewer errors than the keyboard. Measured on a phone. Ruan, Wobbrock, Liou, Ng and Landay, Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones, 2016.
- 2018The desk, checked.typed, physical keyboard51.6wpm
Mean 51.56, SD 20.2, across 168,000 participants. Dhakal, Feit, Kristensson and Oulasvirta, Observations on Typing from 136 Million Keystrokes, CHI 2018.
The 2016 study ran on a phone, and that is the objection worth answering rather than hiding: its touchscreen keyboard came out at 53.5 wpm, and a physical keyboard measured across 168,000 people comes out at 51.6. The typing side barely moves. What moved is the other side, and what held it back in 1999 was correction, not speed. That is the half the modes do.
These five are yours, and every one of them is read off your own disk.
One column is a week, one row is a weekday, Monday at the top. Hover a day for what it holds: the dictations, the words, the two clocks, and the meetings and uploads that landed on it. Read from the same history file the four figures below are computed from.
- 142Words per minutemedian, over 991 dictations
- 5h36mTime savedan estimate, over 4 weeks
against 40 wpm typing - 0.9sTurnaroundmedian wait, over every run
- 3Languagesmostly English, 71 percent
measured on 1,083 of 1,152
A constructed example, not an average of anybody. What is real is that the app measures these five things, out of a file you can open.
The dictation half, done properly.
- Profiles
A profile that changes the output.
Vocabulary, style, delivery and model belong to whatever you are working on. A commit message and a mail to a client are not the same job.
- Models
Four lanes, picked per job.
The local runtime as a real option, not a weekend of configuration. Picked per profile, not once for the whole app.
- The record
It shows what it changed.
Raw and delivered are stored apart. Every rule that fired is named, and audio it could not hear is marked as a gap instead of guessed at.
- Credentials
Your keys, your machine.
API keys live in the operating system secret store, and the config file is scrubbed on save. Nothing leaves for anywhere you did not pick.
- Sources
Bring audio in.
An uploaded file, or a pasted video, podcast or direct audio link, becomes the same kind of record as something you said out loud.
- Licence
AGPL-3.0, the whole thing.
No open core with the good half behind a login. Read it before you decide anything.
Nothing runs on one engine.
A profile decides what handles each job, and the jobs do not have to agree. Cleanup on a small local model, the agent on a frontier one, dictation wherever it is fastest that week. Four lanes, and you pick by what it costs you to run rather than by whichever model is fashionable this month.
A vendor's own hosted API, on an account you hold.
needs a key, in your operating system's secret store
- Dictationwhat you said, into textwhisper-large-v3-turbo
- Meetingsa room, while it runswhisper-large-v3
- Uploads and linksa file or a pasted URLwhisper-1
- Cleanuprepairs what you saidopenai/gpt-oss-20b
- Rewritechanges how it readsopenai/gpt-oss-120b
- Translatesays it in another languageclaude-sonnet-5
- Prompt Enhanceturns a thought into a promptopenai/gpt-oss-120b
- The agentanswers, and can actclaude-sonnet-5
GroqOpenAIAnthropicOpenRouterCartesia
On this machine, on its own disk and memory.
needs none, by construction
- Dictationwhat you said, into textggml-base148 MB
- Meetingsa room, while it runsggml-small488 MB
- Uploads and linksa file or a pasted URLggml-small488 MB
- Cleanuprepairs what you saidllama-3.2-3b-instruct2.0 GB
- Rewritechanges how it readsqwen2.5-7b-instruct4.7 GB
- Translatesays it in another languageqwen2.5-7b-instruct4.7 GB
- Prompt Enhanceturns a thought into a promptqwen2.5-7b-instruct4.7 GB
- The agentanswers, and can actqwen2.5-7b-instruct4.7 GB
WhisperLlamaQwenGemmaOllama
An OpenAI-compatible server you run, on another machine.
needs a base URL, a typed model id, a token if you set one
https://your-host:8000/v1typed on the endpoint
Every one of the eight, on whatever that server serves. The model list is yours, so this page does not have one.
all eightDictationMeetingsUploads and linksCleanupRewriteTranslatePrompt EnhanceThe agent
A cloud account with a region, a tenant and an audit trail. Only Azure transcribes among the three, so the listening jobs go elsewhere.
needs three shapes, one per vendor
- Dictationwhat you said, into textnot on this lane
- Meetingsa room, while it runsnot on this lane
- Uploads and linksa file or a pasted URLnot on this lane
- Cleanuprepairs what you saidanthropic.claude-haiku-4-5
- Rewritechanges how it readsanthropic.claude-sonnet-5
- Translatesays it in another languageanthropic.claude-sonnet-5
- Prompt Enhanceturns a thought into a promptanthropic.claude-haiku-4-5
- The agentanswers, and can actanthropic.claude-sonnet-5
AWS BedrockAzure OpenAIGCP Vertex AI
- Dictationwhat you said, into textCloudwhisper-large-v3-turbothe fastest lane, and you are waiting on this one
- Cleanuprepairs what you saidLocalllama-3.2-3b-instructa repair job a small model does well, and it never leaves the machine
- Translatesays it in another languageYour servertyped on the endpointthe box in the next room, already running
- The agentanswers, and can actCloudclaude-sonnet-5the one job worth a frontier model
Switch language mid-paragraph.
Nothing has to be set first and nothing has to be set back. The language is read off the text that came back, not off a dropdown and not from the provider — on the two lanes most dictations go through, no provider reports one.
- 99+
- languages the recogniser transcribes
- 70
- it names in your own record, so the app can count what you actually dictate in
- English
- Deutsch
- español
- français
- português
- italiano
- Nederlands
- polski
- svenska
- українська
- русский
- Türkçe
- العربية
- हिन्दी
- 中文
- 日本語
- 한국어
- +53 more
With the network off, it is the same app.
The Local lane is the whole product, not a degraded one: dictation, cleanup, rewrite, translate, the agent. Weights on your disk, a runner on 127.0.0.1, no account and no key.
- No keynothing to sign up for, nothing to revoke
- No requestno bytes leave the machine, so nothing can be logged
- 148 MB
- if you only dictate — one speech model, and that is a working install
- 7.3 GB
- if you run all eight jobs locally — 4 files, fetched once
4 lanes and 35 models across 7 vendors, in the catalogue both runtimes read. Every row above is that file's own default for that job. Last checked 2026-08-17.
The honest answers.
Can I download it yet?
Not yet. No release, no date, and no waitlist to put your address on. It runs from source on all three desktops today, and Discord is where you will hear about it first.
Is this another wrapper around Whisper?
The recogniser is swappable and Whisper is one of the options. Speech to text is turning into a commodity, so it is not the product. What is: what accumulates from what you said, and what can act on it.
Does my audio leave the machine?
Only if you point that profile at a cloud model. The local lane runs on your own hardware and reaches nothing. It is a per-profile choice, not one switch for the whole app.
How is this different from the other open dictation apps?
They are good, and the dictation half genuinely overlaps. The difference is where what you said ends up. Theirs keeps it inside the product. WordScript writes it into a directory your own tools already open, so there is no second integration surface to maintain just to get it back out.
When will it be released?
There is no date. The gates still open are recorded in the repository, and that list is the only honest answer to this.
Can I help build it?
Yes, and that is the whole point of this page. The repository is open and the arguing happens on Discord.
Nothing is released yet.
That is the point.
I never wanted a better ear. They all have an ear. What I kept missing is what happens after it: my words stay mine, in folders I own, and exactly one agent gets to touch them. It should be the best dictation app on the desk too. Give me a year.
Or write to
