100% Local — Your voice never leaves your machine

Voice to text that never leaves your computer

You're talking. The words are landing — in your editor, your terminal, a chat box, a form, wherever the cursor sits. And every one of them stays on the machine you said them on.

Brethof Voice Pro is voice-to-text and translation that runs entirely on your own computer. That isn't a privacy feature. It's the whole product.

100% Local — No cloud ever
30 ASR + 38 Translation languages
GPU Accelerated — NVIDIA, AMD, Intel
Linux & Windows
Works Offline

Everything you need, nothing you don't

One program, the work of four. It transcribes your voice, translates it across 38 languages, subtitles your videos, and types for you wherever a keyboard would. It runs offline — on a plane, in a hotel with bad wifi, through a tunnel — and on the laptop you already own, with or without a graphics card. No subscription. No per-minute meter. No uploads. Give your AI agent ears, too: it can transcribe, translate, and record a meeting of any length through MCP, on your machine, with no API key. Buy it once, and it keeps working.

Three things touch the network, and nothing else does

Your voice never leaves your computer. That sentence is precise, and it's the pitch. Three things reach the network at launch, and only three: a licence check, an update check, and the model downloads you start yourself. The licence check sends a licence key, the update check a version string — no audio, no text, ever. Turn the update check off, and once the models are on disk, nothing goes out at all.

30 languages, 22 dialects, and your accent stops being a problem

Transcription in 30+ languages, and 22 Chinese regional dialects understood on their own — no setup, no training. And your accent stops coming out wrong. Brethof Voice Pro learns your voice, your jargon, and your proper nouns, on your own machine. Speak your grandmother's Sichuan and it comes out right.

Speak once, read it in 38 languages

Speak Polish, read English. Speak English, read Japanese. Or say it once and have it typed in English, Japanese, French and Spanish at the same time — one per line, or side by side in the same line. Pick as many target languages as you want; 38 are there, and 23 work in both directions. The translation happens on your machine too, in Tencent's Hunyuan MT2. At the quality size it reaches 97.9% of Google's Gemini 3.1 Pro on the FLORES-200 benchmark, and beats it on real-world and minority-language translation. Two sizes, both optional downloads: a fast one for everyday, a quality one for the words that matter. Subtitles, too. Plain text or SRT, timed by sentence, or word by word when you need to edit precisely. Translate a subtitle file and every timing survives. Subtitle your own video in 38 languages without paying by the minute.

Fully Offline

That's why it works where nothing else will. A plane. A hotel with bad wifi. A train through a tunnel. There's nowhere for your recording to go, so it doesn't need to go anywhere. Drop in a two-hour meeting recording and go make coffee; the transcript is there when you're back. Forget you left it recording and nothing is lost — it finishes the job on its own and keeps the transcript.

Two Model Sizes

Two transcription models, and it's a choice, not a tier — you can switch any time.

It learns from the fixes you already make

Most voice software makes you read out sample sentences at setup. This one doesn't. It learns from the corrections you already make: you fix a misheard word, and that fix becomes the training example. Pin your brand names, your jargon, your colleagues' names once, and they stop coming out wrong — in the transcript and in the translation, because the same list steers both. A name is never localised by accident. Training is one click, it runs on your hardware, and the new model is ready to switch to when it's done.

Pay once, own forever

Buy it once and it's yours. A perpetual licence with a year of updates included.

Noise reduction, only when you need it

A noisy room stops being a problem. Noise reduction is built in, for recordings made in a café, a trade show floor, a car. It's off by default, because we measured it: on short, clean microphone clips it makes things worse. For the two-hour meeting with the espresso machine behind you, switch it on.

The keyboard where you're already typing

The voice keyboard types wherever your cursor is — a browser, an IDE, a chat app, any field that takes a keyboard. And it can type the translation instead of the transcription. Say the sentence in your language, and the right words land in the form, in the language the form needs. One line, no cutting and pasting.

Two sizes, your call

The 0.6B is the everyday one. It runs on any 4 GB graphics card, on integrated graphics, and on laptops with no graphics card at all. Fast, light, there when you need it.

1.7B

Large · Vulkan / CPU
~2–3 GB

No GPU at all? It still runs. CPU-only needs 8 GB of RAM, a four-core processor, and about 5 GB of free disk; everything else is 7–9 GB. Any Vulkan 1.2+ card works — NVIDIA, AMD, Intel Arc. And transcription and translation each pick their own processor, so one can run on the graphics card while the other runs on the CPU.

Optional add-ons download on demand from Settings → Models:

Forced Aligner (~540 MB) for word-level timestamps · Hunyuan MT2 Fast (~1 GB) or Quality (~4.3 GB) for translation.

How we compare

Feature Brethof Voice Pro Dragon Google STT Otter.ai Whisper (OSS)
100% local processing ✓ ✓ ✗ ✗ ✓
Perpetual license ✓ ~ ✗ ✗ ✓
Native Linux support ✓ ✗ ~ ✗ ✓
Native Windows support ✓ ✓ ~ ✗ ~
30 ASR languages + auto-detect ✓ ✗ ✓ ~ ✓
Offline translation (38 languages) ✓ ✗ ✗ ✗ ✗
GPU acceleration (NVIDIA + AMD + Intel) ✓ ✗ N/A N/A ~
Personal model fine-tuning (LoRA) ✓ ✓ ✗ ✗ ✗
MCP server for AI agents ✓ ✗ ✗ ✗ ✗
Built-in noise reduction ✓ ✓ ✓ ✓ ✗
Direct text injection ✓ ✓ ✗ ✗ ✗
Polished desktop GUI ✓ ✓ ✗ ✓ ✗
Typical cost $49 once $350+/yr $17/mo $17/mo Free

Pay once. Own forever.

No monthly fees. No usage limits. Perpetual license with 1 year of updates included.

Free Trial
$ 0

No credit card required. Just an email to verify your trial.

  • ✓ Both model sizes (0.6B + 1.7B)
  • ✓ GPU acceleration
  • ✓ 30 ASR + 38 translation languages
  • ✓ Noise reduction
  • × No personal training (paid plans only)
  • × No MCP server (paid plans only)
  • × 14-day limit
Start Free Trial
Business
$ 249 /seat

Per-seat perpetual license. Team & organization use. 1 year of updates.

  • ✓ Perpetual license (per seat)
  • ✓ Team & organization use
  • ✓ Both ASR sizes (0.6B + 1.7B)
  • ✓ Offline translation (38 languages)
  • ✓ Personal voice training (LoRA)
  • ✓ MCP server for AI agents
  • ✓ Priority support
  • ✓ Volume discounts (10+ seats)
Buy Business License

Prices excl. tax. Then $20/seat/year for updates (optional)

Frequently asked questions

No. Brethof Voice Pro processes everything locally on your device. No audio or text data ever leaves your computer. There is no cloud component, no telemetry, and no analytics.

Any modern GPU works. NVIDIA, AMD, and Intel Arc all use Vulkan acceleration. You can also run on CPU only, though transcription will be slower. The 0.6B model runs comfortably on integrated graphics or any 4 GB+ Vulkan card.

Start with the 0.6B model — it is the recommended default and runs great on most GPUs (and even on CPU on most modern machines). If you need higher accuracy on accented or noisy audio, switch to the 1.7B model (needs 6 GB+ VRAM). You can switch sizes at any time from Settings → Models without re-downloading.

Yes. Brethof Voice Pro supports both Linux and Windows natively. On Linux it works with X11 and Wayland. On Windows it runs as a standard desktop application.

Your license is perpetual — the app keeps working forever with whatever version you have. The optional $20/year Update Pass gives you access to new features and model improvements. Without it, you simply stay on your current version.

Yes — personal voice training is included in v2.0.0 and runs end-to-end on your machine. Every time you correct a misrecognised word, the {clip, correction} pair is auto-saved to your local training dataset. The main window's training card shows total samples and minutes captured at a glance — click "Start training" in the Training tab to fine-tune a LoRA on your accent. The result auto-exports to GGUF and you switch to it in one click. Free for every paid licence, your voice data never leaves your machine.

Ready to try it?

14-day free trial. No credit card. No cloud. No compromises.

Hear it when it ships

New releases, real benchmarks and the occasional deep-dive. No spam, unsubscribe in one click.

Everything we build

External:   YouTube · GitHub