Open Source Alternatives LogoOpen Source Alternatives
AlternativesBlogAdvertise
Open Source Alternatives LogoOpen Source Alternatives

Stay Updated

Subscribe to our newsletter for the latest news and updates about Alternatives

Open Source Alternatives LogoOpen Source Alternatives

Handpicked Open Source Alternatives to Paid Softwares

Product
  • Categories
  • Tag
  • Advertise
Resources
  • Blog
  • Collection
  • Submit
  • Advertise your tool
Company
  • Privacy Policy
  • Terms of Service
  • Refund Policy
  • Sitemap
Alternatives
  • Claude Code
  • Jira
  • Notion
  • Slack
  • Linear
  • Wispr Flow
  • All alternatives
Copyright © 2026 All Rights Reserved.
Home/Categories/AI & Machine Learning/Openlive
icon of Openlive

Openlive

Open source alternative to OpenAI Realtime API, Gemini Live and ElevenLabs Agents

Run a local voice and vision layer for AI agents and coding assistants, with on-device speech processing, barge-in, and support for any AI provider.

329 starsTypeScriptApache-2.0Active this week
Visit websiteGitHub repo
image of Openlive
Contents
  1. 01Who Openlive is for
  2. 02The problem it solves
  3. 03How it solves it
  4. 04Strengths and trade-offs
  5. 05Openlive vs alternatives
  6. 06Tech stack
  7. 07FAQ
  8. 08Similar open-source tools
TL;DR

Openlive is an on-device voice and vision layer for AI agents and coding assistants, released under the Apache-2.0 license. It replaces per-minute hosted pipelines like OpenAI Realtime API, Gemini Live, and ElevenLabs Agents by running the full speech loop (VAD, STT, TTS, barge-in) locally on your machine. Available as a desktop app for macOS, Windows, and Linux. Best for developers who voice-drive coding agents like Claude Code, Cursor, or Codex, and for technical users who want a system-wide voice assistant without uploading audio to a cloud service.Apache-2.0 · TypeScript · 329 stars · Active this week

who it's for

Who Openlive is for#

Developers who voice-drive coding agents

Talk to Claude Code, Cursor, Codex, or other supported agents during a coding session: describe changes, answer permission prompts by voice, and hear the agent's plan steps narrated as it works. Conversations land in the agent's own session history, and the agent's existing CLI sessions show up in OpenLive's History.

Skip if:

Skip if your preferred agent is not yet supported by the Agent Client Protocol, or if you primarily use browser-based AI tools with no terminal component.

Technical users who want a system-wide voice assistant

Flow activates from any app with a double-tap of Control, routes to your own API key or Ollama, and can type at the cursor, open apps, click UI elements, and run shell commands. No subscription required; no audio leaves the machine.

Skip if:

Skip if you need a voice assistant on mobile or on a device you do not control. OpenLive is a desktop app for macOS, Windows, and Linux only.

Users who need fully local voice processing

Every stage of the voice loop runs on-device. The STT, TTS, and VAD never send audio to any server. The only network traffic is the transcript sent to whichever model answers, and Ollama keeps even that local.

Skip if:

Skip if you are on a very low-spec machine. Running Whisper or native STT models locally adds CPU or GPU load that hosted alternatives offload entirely.

Multilingual voice typing into any application

Dictate types your speech into any text box in ten languages: English, Spanish, French, German, Italian, Portuguese, Hindi, Chinese, Japanese, and Korean. Local rules handle punctuation and filler words. Optional AI polish rewrites the output in a tone you choose, using your API key.

Skip if:

Skip if you need real-time collaborative transcription across multiple speakers. Dictate is single-user and single-device.

the problem

The problem it solves#

Building a real conversational loop around an AI takes more than an API call. Voice activity detection, reliable end-of-turn detection, streaming speech-to-text, model inference, streaming text-to-speech, and barge-in support all need to compose correctly. Commercial platforms like OpenAI Realtime API, Gemini Live, and ElevenLabs Agents package this as a hosted service, but they charge per minute for audio processing and route your voice data through their infrastructure.

The result is audio fees that compound quickly on long sessions, no path to fully local inference, and a hard dependency on a specific provider's pipeline. Developers who want to use Ollama, route through a different model provider, or keep voice data off cloud servers have no drop-in option. Coding agent users who want to talk to Claude Code or Cursor in real time face the same gap: the agents have no voice interface, and building one from scratch means solving every piece of the speech loop independently.

how Openlive solves it

How it solves it#

On-device speech loop with VAD, STT, TTS, and barge-in

The full voice pipeline runs on-device: Silero VAD detects speech, Smart-Turn handles end-of-turn, STT options include Whisper (WebGPU), Nemotron, Parakeet, Moonshine, and Canary. TTS options include Kokoro (28 voices), Supertonic (44.1 kHz, GPU), and Piper. No audio is uploaded; only the final transcript leaves the machine.

14 model providers and 9 coding agents

Bring any model: Anthropic, OpenAI, Google Gemini, xAI, DeepSeek, Groq, Ollama (fully local), MiniMax, OpenRouter, Mistral, Together, Fireworks, Cerebras, and Perplexity. Or connect a coding agent (Claude Code, Codex, Cursor, OpenCode, Hermes, Gemini CLI, GitHub Copilot, Kiro, Pi) over the Agent Client Protocol (JSON-RPC over stdio).

Three modes: Chat, Flow, and Dictate

Chat is a voice call in a window with any AI or agent. Flow is a system-wide assistant triggered with a double-tap of Control from any app: it can answer out loud, type at the cursor, and drive the machine (open apps, click, scroll, run shell commands). Dictate types what you say into any text box, with optional AI polish.

Zero-shot local voice cloning

Settings > Voice records 5 to 30 seconds of audio, and the assistant speaks in your voice from then on. Zero-shot cloning runs locally via ZipVoice (approximately 208 MB download, optional). Profiles can be exported and imported between machines.

Camera and screen vision per turn

Camera or screen frames ride each conversation turn. A text-only model can borrow a separate vision model's description of each frame. The `look` tool grabs a hi-res frame on demand mid-conversation without interrupting the voice loop.

MCP connectors, tools, and ten spoken languages

Built-in tools, skills, and MCP connectors are managed in Settings > Capabilities. Connectors import directly from Claude Desktop, Claude Code, Codex, Cursor, Gemini CLI, and VS Code. Ten spoken languages are supported: English, Spanish, French, German, Italian, Portuguese, Hindi, Chinese, Japanese, and Korean.

strengths · trade-offs

Strengths and trade-offs#

Strengths

  • No per-minute audio feesHosted voice pipelines charge for every second of audio processed. OpenLive processes all audio locally, so there are no audio API costs. You pay only for the model inference you would pay for anyway with a direct API key or Ollama.
  • Works with any model, including fully local OllamaUnlike OpenAI Realtime API or ElevenLabs Agents, which are tied to specific models, OpenLive is model-agnostic. Ollama support means you can run the entire stack without any external API calls: local audio, local model, local TTS.
  • Apache-2.0 license with encrypted API key storageApache-2.0 allows commercial use, modification, and redistribution. Voice audio never uploads: only the transcript (and optionally screen frames) reaches the model. API keys are encrypted at rest with AES-256-GCM; only the last four digits are shown in the UI.
  • Voice-drives coding agents over ACP with permission relayClaude Code, Codex, Cursor, and other agents connect via the Agent Client Protocol. When the agent wants to run a command or edit files, OpenLive speaks the question; you answer by voice. The agent's working plan renders as a checklist; optional narrated progress reports steps out loud while the agent works.

Trade-offs

  • -Cascaded pipeline, not full-duplex speech-to-speechOpenLive is a cascade (speech to text, then model, then TTS), not a full-duplex speech-to-speech model like GPT-Live. The README explicitly notes this trade: a speech-native model can overlap listening and speaking in ways a cascade cannot, but the cascade is what makes provider-agnosticism and local audio processing possible.
  • -Local STT and TTS require local computeSTT and TTS models download on demand and run on your CPU or GPU. Low-spec machines will have slower transcription and higher latency than hosted alternatives that offload this compute. A capable GPU accelerates inference but is not required.
  • -Flow mode needs platform-specific permission setupThe Flow mode requires Microphone, Accessibility, and Screen Recording permissions on macOS. Windows and Linux have their own setup steps; on Linux, the `input` group is required for global keybindings (both X11 and Wayland). Until access is granted, global hotkeys do not work.
versus alternatives

Openlive vs alternatives#

OpenLive vs OpenAI Realtime API

OpenAI Realtime API provides a hosted, full-duplex speech-to-speech pipeline tied to GPT-4o-Realtime. OpenLive uses a local cascaded pipeline that works with any of 14 providers, including OpenAI as one option.

FeatureOpenLiveOpenAI Realtime API
LicenseApache-2.0Proprietary
Audio processingOn-deviceCloud (OpenAI servers)
Model support14 providers + 9 agentsGPT-4o-Realtime only
Audio billingNone (local)Per-second
Barge-inYesYes
Full-duplex overlapNo (cascaded)Yes

OpenLive is the better choice when you need to keep voice data local, route to Ollama or a non-OpenAI provider, or avoid per-second audio billing. OpenAI Realtime API remains preferable when full-duplex, speech-native behavior is a hard requirement, where overlapping speech and listening matters more than local processing or model flexibility.

OpenLive vs ElevenLabs Agents

ElevenLabs Agents is a hosted conversational AI platform with high-quality managed TTS voices. OpenLive runs TTS models locally (Kokoro, Supertonic, Piper) and adds zero-shot voice cloning via ZipVoice.

FeatureOpenLiveElevenLabs Agents
LicenseApache-2.0Proprietary
TTS processingOn-deviceCloud
Voice cloningZero-shot, local (ZipVoice)Cloud-based
PricingFree (local compute)Per-character billing
Coding agent support9 agents via ACPNot supported

ElevenLabs Agents has a larger library of high-quality managed voice profiles and a simpler hosted setup. OpenLive is the better fit when you need local voice cloning, zero audio billing, or integration with coding agents like Claude Code or Cursor via the Agent Client Protocol.

OpenLive vs Gemini Live

Gemini Live is Google's multimodal voice conversation interface tied to the Gemini model family. OpenLive supports Google Gemini as one of 14 providers while processing audio on-device.

FeatureOpenLiveGemini Live
LicenseApache-2.0Proprietary
Model lock-inNone (14 providers)Gemini only
Audio processingOn-deviceCloud
Coding agent support9 agents via ACPNot supported

Gemini Live is faster to set up and benefits from tight Gemini model integration and Google's voice infrastructure. OpenLive is preferable when model-agnosticism, local audio, or coding agent voice control matters more than a zero-setup hosted experience.

tech stack · detected from GitHub

What it's built on#

Languages
JavaScriptRustTypeScript
Frameworks
Next.jsReact
Tooling
esbuild
frequently asked

FAQ#

Does OpenLive send my voice audio to any cloud service?

No. All audio processing, including VAD, speech-to-text, and text-to-speech, runs on your local machine. The only data that leaves the device is the final text transcript (and screen frames if you enable vision), sent to whichever AI model or agent you have configured. If you use Ollama as your model provider, no data leaves the machine at all.

Is OpenLive the same as OpenAI Realtime API or a wrapper around it?

No. OpenLive is a separate, Apache-2.0 licensed cascaded pipeline that handles STT, TTS, VAD, and barge-in locally. It is not a wrapper around OpenAI Realtime, GPT-Live, or any hosted voice API. You can connect it to OpenAI as one of 14 supported model providers, but the audio pipeline itself always runs on-device.

Which coding agents does OpenLive support?

Claude Code, Codex, Cursor, OpenCode, Hermes, Gemini CLI, GitHub Copilot, Kiro, and Pi. Agents connect over the Agent Client Protocol (ACP), which uses JSON-RPC over stdio. OpenLive can install, sign in, and update each agent's CLI from within the app, and the agent's existing CLI sessions appear in OpenLive's History.

What STT and TTS engines does OpenLive use?

For STT: Whisper on WebGPU, Nemotron (transcribes while you talk), Parakeet, Moonshine, and Canary. For TTS: Kokoro (28 voices), Supertonic (10 voices, 44.1 kHz), Piper, and others. Engines download on demand; OpenLive picks the one fastest on your hardware and falls back to Whisper or the browser voice when a selected engine is missing.

Can I clone my own voice in OpenLive?

Yes. Settings > Voice records 5 to 30 seconds of audio, then the assistant speaks in your voice using ZipVoice, a zero-shot cloning model that runs entirely on your device. The model is approximately 208 MB and is optional. It requires enabling models with a restricted license in Settings because ZipVoice's weights were trained on non-commercial data. The README notes this is intended for your own voice or one you have clear permission to use.

also worth a look

Similar open-source tools#

RealtimeSTT

RealtimeSTT

Real-time speech-to-text library with VAD and wake words

10.2KPythonMIT
hindsight

hindsight

Self-hosted memory that makes AI agents smarter over time

45.2KPythonMIT
cli

cli

Official Lark/Feishu CLI with 200+ commands and AI Agent Skills

17.5KGoMIT
LibreChat

LibreChat

One self-hosted interface for every AI model you use

45.2KTypeScriptMIT
FckSignups

FckSignups

Open-source tools that work instantly, no signup required

4.4KTypeScriptGPL-3.0
FreeFlow

FreeFlow

Free, open source Mac dictation with AI cleanup and voice macros

2.8KSwiftMIT

Repository

Stars
329
Forks
77
License
Apache-2.0
Latest
v0.3.1
Last commit
1 day ago
Last verified
Oct 6, 2026
Repo
katipally/openlive ↗

Additional details

Language
TypeScript
Open issues
7
Contributors
1
First release
2026

Categories

AI & Machine LearningDeveloper ToolsCommunication & Collaboration

Tags

AI AgentsLocal-first