Open Source Alternatives LogoOpen Source Alternatives
AlternativesBlogAdvertise
Open Source Alternatives LogoOpen Source Alternatives

Stay Updated

Subscribe to our newsletter for the latest news and updates about Alternatives

Open Source Alternatives LogoOpen Source Alternatives

Handpicked Open Source Alternatives to Paid Softwares

Product
  • Categories
  • Tag
  • Advertise
Resources
  • Blog
  • Collection
  • Submit
  • Advertise your tool
Company
  • Privacy Policy
  • Terms of Service
  • Refund Policy
  • Sitemap
Alternatives
  • Claude Code
  • Jira
  • Notion
  • Slack
  • Linear
  • Wispr Flow
  • All alternatives
Copyright © 2026 All Rights Reserved.
Home/Categories/AI & Machine Learning/Diction
icon of Diction

Diction

Open source alternative to Wispr Flow, Superwhisper, Aqua Voice and VoiceDash

Run voice dictation in any iPhone app with an open-source, self-hosted speech-to-text gateway. MIT licensed, no word limits, no audio tracking.

216 starsGoMITActive this week
Visit websiteGitHub repoDeployDeploy on Hostinger
image of Diction
Contents
  1. 01Who Diction is for
  2. 02The problem it solves
  3. 03How it solves it
  4. 04Strengths and trade-offs
  5. 05Diction vs alternatives
  6. 06Quick start
  7. 07Tech stack
  8. 08FAQ
  9. 09Similar open-source tools
TL;DR

Diction is an iOS keyboard that transcribes speech to text in any app without switching contexts. The open-source gateway, written in Go and MIT licensed, replaces paid services like Wispr Flow by running transcription on your own server with Docker. It supports Parakeet and Whisper models, adds AES-256-GCM encrypted audio streaming, and imposes no word limits or daily caps. Best for developers and privacy-focused users who dictate often and want full control over where audio is processed.MIT · Go · 216 stars · Active this week

who it's for

Who Diction is for#

iPhone-first professionals who dictate daily

Users who draft emails, messages, and documents frequently on iPhone benefit directly from unlimited transcription volume and in-field text insertion without app switching. Diction saves the round-trip to a dedicated dictation app by placing transcribed text wherever the cursor is active.

Skip if:

You primarily write on a Mac or PC. Diction is an iOS system keyboard; there is no desktop keyboard equivalent in this repository.

Privacy-focused users who want audio on-premise

On-device mode transcribes audio locally on the iPhone with no network request at all. The self-hosted gateway mode sends audio only to your own server over an encrypted connection. Neither mode routes audio through a third-party API, which matters for anyone working with sensitive personal or professional content.

Skip if:

You are comfortable with managed cloud transcription and do not want to maintain server infrastructure. In that case, Wispr Flow or Superwhisper offer maintained managed services with no setup.

Developers testing or benchmarking speech models

The gateway's OpenAI-compatible API makes it straightforward to compare Parakeet-v3 against Whisper small, medium, and large-v3-turbo using the same test audio. The X-Diction-Whisper-Ms, X-Diction-Route-Model, and X-Diction-Route-Lang headers give per-request timing and routing data.

Skip if:

You need transcription across platforms beyond iPhone. The gateway server runs on any Docker host, but the keyboard client is iOS-only.

Teams with data-residency or compliance requirements

Organizations where audio data cannot leave internal infrastructure can run the gateway on a self-managed server inside their network. The gateway adds no analytics layer; audio is forwarded only to the configured speech model backend, which also runs on infrastructure the team controls.

Skip if:

Your compliance requirements are already met by a vetted managed transcription vendor. The gateway adds self-hosting control but not regulatory certification.

the problem

The problem it solves#

The leading iOS voice-to-text tools are closed, cloud-dependent services that process audio on proprietary servers. Wispr Flow, Superwhisper, and similar paid tools impose word limits or daily dictation caps, and none offer a self-hosting path. Every dictation session passes through infrastructure you do not control, with no way to audit where the audio goes or how long it is retained.

For users who dictate frequently, subscription pricing scales with volume in ways that quickly outpace what a self-hosted server costs. For teams handling sensitive data, sending audio to a third-party vendor can be a blocker. The built-in iOS dictation and competing keyboards lack the accuracy of dedicated Whisper-based models and offer no flexibility to swap backends or bring your own speech model.

how Diction solves it

How it solves it#

Real-time streaming transcription

The gateway adds a WebSocket layer between the iOS keyboard and the speech model. Audio is transcribed as you speak, so results arrive before you tap stop. The X-Diction-Whisper-Ms response header reports the speech model's inference latency in milliseconds on each request, so you can measure backend performance directly.

AES-256-GCM end-to-end encryption

Audio travels from the iPhone keyboard to the gateway over AES-256-GCM with X25519 key exchange, the same cryptographic primitives used by Signal and WireGuard. Encryption runs before audio leaves the device; the gateway decrypts it to forward to the speech model, which runs on your own server.

OpenAI-compatible transcription API

The gateway implements POST /v1/audio/transcriptions from the OpenAI API spec, so any Whisper-compatible server works as the backend without modification. Swap models by changing DEFAULT_MODEL in your compose file. The X-Diction-Route-Model header confirms which backend actually served each request.

On-device transcription mode

The iOS app can transcribe speech locally on the iPhone with no gateway required. On-device mode processes audio in the app itself; nothing is transmitted to a server. This is the simplest path for users who do not want to maintain a server and whose primary languages are covered by the on-device model.

Multi-backend speech model support

Four model profiles ship with the project: Parakeet-v3 (NVIDIA GPU, 25 European languages, weights baked into the image), Whisper small (~850 MB RAM, best for CPU), Whisper medium (~2.1 GB), and Whisper large-v3-turbo (~2.3 GB, highest accuracy). Switching is one DEFAULT_MODEL change in the gateway compose file.

Zero telemetry

The self-hosted gateway collects no analytics, no crash data, and no usage telemetry. The source code is auditable on GitHub. In self-hosted and on-device modes, audio routes only to your configured speech model backend; nothing passes through Diction Labs' infrastructure.

strengths · trade-offs

Strengths and trade-offs#

Strengths

  • MIT license for the gateway serverThe gateway code is MIT licensed, so you can run it on your own infrastructure, modify the source, and use it commercially without licensing fees. The iOS app is distributed separately via the App Store; the MIT license governs the server-side code in this repository.
  • No transcription caps or rate limitsSelf-hosted and on-device modes have no word limits, no daily quotas, and no rate limits. Paid services like Wispr Flow gate higher volumes behind subscription tiers. Running your own gateway removes that ceiling; the only constraint is the compute on your server.
  • Works in any iOS app via system keyboardDiction installs as a system-level iOS keyboard, so it inserts transcribed text into any text field across any app on the device without needing the target app to support dictation. The keyboard switch is a long-press on the globe icon, then a tap on Diction.
  • Model-agnostic gatewayThe gateway forwards requests to any server implementing the OpenAI transcription API spec. Parakeet gives low-latency transcription for 25 European languages on NVIDIA GPU; Whisper small, medium, and large-v3-turbo cover 99 languages on CPU or GPU. Switching is one line in the compose file.

Trade-offs

  • -NVIDIA GPU recommended for best performanceThe default Parakeet stack targets NVIDIA GPUs via the Container Toolkit. The CPU-only Whisper path runs on any hardware but is slower, particularly on longer dictations. The Parakeet INT8 image needs approximately 2 GB of VRAM; hosts without a compatible NVIDIA GPU must use the Whisper CPU path instead.
  • -Server setup required for self-hosted modeDiction is not a zero-infrastructure option. Self-hosting requires a machine running Docker, network routing from the iPhone to the server (either a local IP reservation or a VPN like Tailscale), and ongoing maintenance of the gateway container. On-device mode sidesteps the server requirement but offers a single fixed model with no swappable backends.
  • -iOS only for the keyboard clientThe iOS app requires iOS 17.0 or later on iPhone. No Android app is included in this repository, and the gateway does not ship with a web or desktop keyboard client. The gateway server runs on any Docker host, but the keyboard frontend is iPhone-only.
versus alternatives

Diction vs alternatives#

Diction vs Wispr Flow

Both tools deliver voice-to-text directly into any active text field on iPhone without switching apps. The key difference is where transcription happens.

FeatureDictionWispr Flow
LicenseMIT (gateway)Proprietary
Self-hostingYes (Docker)No
On-device modeYesLimited
Word limitsNone (self-hosted)Tier-dependent
Audio routed to third partyNo (self-hosted)Yes

Diction is the stronger choice when you want audio processed on your own server, need unlimited dictation volume without a subscription cap, or work in environments where third-party cloud audio processing is not acceptable. Setup requires Docker and a host machine reachable from the iPhone.

Wispr Flow is the better choice when you want a managed service with no server infrastructure, need AI-powered writing assistance beyond basic transcription cleanup, or prefer a polished consumer experience without infrastructure to maintain.

Diction vs Superwhisper

Superwhisper is macOS-native; Diction is iPhone-native. Both use Whisper-based transcription but target different primary devices. Superwhisper processes audio on-device using Apple Silicon by default; Diction offers on-device transcription on iPhone and a self-hosted server path for streaming use cases where latency matters.

Choose Diction when your primary dictation device is an iPhone and you want no word caps or subscription cost. Choose Superwhisper when your primary device is a Mac and you want AI text transformations built in without managing a server.

install · quick start

Quick start#

bash
Self-hosting deploys the Diction gateway and a speech model backend using Docker Compose on any Linux server or VPS.
```bash
docker compose up -d
```
tech stack · detected from GitHub

What it's built on#

Languages
Go
frequently asked

FAQ#

Is the Diction gateway free to self-host?

Yes. The gateway server is MIT licensed and free to run on your own infrastructure with no usage fees. Self-hosted and on-device modes have no word limits and no rate limits, per the project README. The iOS keyboard app is available separately on the App Store (ID 6759807364); check the App Store listing for its current pricing.

Do I need an NVIDIA GPU to run the Diction gateway?

No. Diction ships two backend paths: Parakeet (requires NVIDIA GPU via the Container Toolkit, approximately 2 GB VRAM) and Whisper (CPU-only, runs on any Docker host). The CPU Whisper path is slower on long audio but runs on a NUC, VPS, or any machine without a GPU. The README's No GPU section covers the CPU compose setup in detail.

What is the difference between the iOS app and the GitHub repository?

The iOS app is a keyboard extension distributed on the App Store. The GitHub repository hosts the gateway: a Go server that sits between the keyboard and your speech-to-text backend. The gateway handles WebSocket streaming, AES-256-GCM encryption, and model routing. You self-host the gateway; the iOS app connects to it.

Which languages does Diction support?

The product website states support for 99 languages in cloud and on-device modes. The Parakeet backend covers 25 European languages. Whisper small, medium, and large-v3-turbo cover a broader language range. The X-Diction-Route-Lang response header returns the detected language on each request when language detection runs.

How does self-hosting Diction differ from using a managed service like Wispr Flow?

With Diction's self-hosted gateway, audio is processed on your own server or on-device; it never routes through Diction Labs' infrastructure. You choose the speech model, pay only for the server, and face no word caps. Managed services like Wispr Flow handle the server for you but process audio on their own infrastructure, apply usage limits on lower tiers, and charge a monthly subscription.

also worth a look

Similar open-source tools#

FU Flow

FU Flow

Free offline voice typing for Windows 10/11 and Ubuntu Linux

0RustMIT License with third-party component notice
Voquill

Voquill

Open source voice dictation with local AI and custom glossary

1KTypeScriptAGPL-3.0
VoiceInk

VoiceInk

Private voice dictation for Mac, no subscription required.

6.7KSwiftGPL-3.0-only
FreeFlow

FreeFlow

Free, open source Mac dictation with AI cleanup and voice macros

2.8KSwiftMIT
SpeakoFlow

SpeakoFlow

Voice dictation and AI assistant for your desktop, fully offline

260RustMIT
OpenWhispr

OpenWhispr

Local-first voice dictation with Whisper, offline mode, and AI cleanup

9KJavaScriptMIT

Repository

Stars
216
Forks
12
License
MIT
Latest
v14.0
Last commit
today
Last verified
Oct 7, 2026
Repo
DictionLabs/Diction ↗

Additional details

Language
Go
Open issues
0
Contributors
6
First release
2026

Categories

AI & Machine LearningCommunication & CollaborationProduct & Project Management

Tags

AI Coding AssistantSelf HostedDeveloper Tools