Audio Engine · v0.6.0

A local audio AI engine, on your machine.

Speech-to-text, text-to-speech and more — powered by audio.cpp, a pure-C++/ggml engine. Install it once; it keeps a local server running and puts an OpenAI-compatible API on 127.0.0.1. No Python, no account, nothing uploaded.

WhisperKey's heavyweight sibling: where WhisperKey is a tiny, single-purpose dictation tool, this is the full engine for when you need one.

What it does

Speech to text

Transcribe audio locally — SenseVoice, fast and accurate, with punctuation.

Text to speech

Natural neural voices with Pocket-TTS; more voices and cloning available.

Voice cloning

Clone a voice from a short reference clip (via the underlying engine).

More

audio.cpp also brings diarization, VAD, source separation and music — room to grow.

Use it right in the app

No code needed — the app has a built-in panel for each job, all running on your machine.

Speak

Type or paste text, pick a voice and speed, and get a voiceover you can play and save.

Transcribe

Record from your mic or drop in an audio file and get the text back — recording is native, not the browser.

Podcast

Paste notes and split them into a back-and-forth, assign two host voices, and render one stitched show.

Self-contained

The engine ships inside the app. The model weights (~390 MB) download themselves the first time you run it, after which it works entirely offline. Nothing to install alongside it, no toolchain, no venv.

Use it from anything

# LLMs: point an MCP client at
node mcp/server.mjs        # AUDIO_ENGINE_URL=http://127.0.0.1:8080

# Any app: the OpenAI audio API
curl :8080/v1/audio/speech -d '{"model":"tts","input":"hello"}' -o out.wav

Download

Linux

v0.6.0

x86_64 · GNOME, KDE or any X11/Wayland desktop

  • .deb package — Debian, Ubuntu, Mint, Pop!_OS
  • AppImage — Any distribution, nothing to install

Windows

v0.6.0

Windows 10 1803 or newer · 64-bit

  • Installer — Per-user, no administrator prompt

macOS

Not yet built

A macOS build is on the way — not yet packaged for download.

SHA-256 checksums

Verify with sha256sum on Linux or Get-FileHash in PowerShell. The builds are unsigned, so this is the only way to confirm you have the file that was actually published.

AudioEngine_0.6.0_amd64.deb 083131cf64ad5a66081186049d0ac6c327c88980be83f64f37759a77bba735f0
AudioEngine_0.6.0_amd64.AppImage 93f5fc1799cb84c43c187265fe3cbd029b66f0bddd643fc8a8406485203e7c69
AudioEngine_0.6.0_x64-setup.exe 23aaecb54cd5ad550fdb9e78f420d2c6ac11c4f42fa63870086a075cd9ee02c9

The builds are unsigned; on first launch your OS may warn, and every download has a published SHA-256 above. The Windows installer is cross-compiled from Linux — SmartScreen will warn (More info → Run anyway).

Prefer to build Windows yourself?

The Windows download above is prebuilt. To build it natively instead (~15 min), against an audio.cpp checkout:

# PowerShell
irm https://dl.vibedout.app/audioengine/build-windows.ps1 -OutFile build-windows.ps1
$env:AUDIO_CPP_SRC = "C:\path\to\audio.cpp"
powershell -ExecutionPolicy Bypass -File build-windows.ps1

Direct link: build-windows.ps1. It installs any missing tools, builds the engine with MSVC, and produces the same installer.