ExLlamaV3New
Fast quantised-LLM inference on consumer NVIDIA GPUs; successor to ExLlamaV2 (archived by its author). Serve it with TabbyAPI.
★ 1.6kMITv1.5.4 · 2026-10-03verified
Curated list of AI tools that run 100% on your machine — no cloud, no telemetry, no "local-ish" setups that secretly phone home.
Showing 0 of 84
NameCategoryTagsStarsLatestListedNothing matches. Clear a filter or search for something shorter.
Engines that load LLMs, vision models, and other neural networks for inference on your hardware.
Fast quantised-LLM inference on consumer NVIDIA GPUs; successor to ExLlamaV2 (archived by its author). Serve it with TabbyAPI.
★ 1.6kMITv1.5.4 · 2026-10-03verified
Open-source ChatGPT alternative. Bundles llama.cpp + a clean UI.
verified
Single-binary llama.cpp wrapper with KoboldAI UI for chat, story-writing, RP.
★ 11.9kAGPL-3.0v1.122.1 · 2026-09-26verified
One Go binary and one config file that hot-swaps models on demand in front of llama.cpp, vLLM, stable-diffusion.cpp, ComfyUI and other local servers, behind a single OpenAI/Anthropic-compatible endpoint. MIT.
★ 5.8kMITv262 · 2026-10-03verified
Reference C++ implementation for running LLaMA-family and other transformer models with GGUF quantization. Powers most of the others in this section.
★ 130.4kMITv0.5.0 · 2026-09-23verified
Mozilla's one-file LLM: model and llama.cpp runtime in a single executable that runs on Linux, Windows and macOS without an install.
★ 26.2k0.10.6 · 2026-09-15verified
Polished desktop app for discovering, downloading and running local LLMs (llama.cpp and MLX engines), with an OpenAI-compatible server, the lms CLI and the headless llmster daemon. Free for personal and commercial use. Its sibling app Bionic is an agent with optional paid cloud models (account needed only for those).
verified
Self-hosted, OpenAI-compatible inference server. Text, image, audio, embeddings — all on your machine.
verified
Rust LLM inference platform with quantization, vision, MoE, and speculative decoding.
★ 7.7kMITv0.9.4 · 2026-09-24verified
Compile-once, deploy-anywhere LLM runtime. Targets WebGPU, Vulkan, CUDA, Metal, iOS, and Android from a single source.
verified
Single-binary server with a built-in model library. Pull, run, and swap models with one command.
verified
Runs and serves models from Hugging Face, Ollama or OCI registries inside GPU-matched containers (Podman/Docker) running llama.cpp or vLLM, so the host needs no setup. Containers run with network access off. MIT.
★ 3.1kMITv0.25.0 · 2026-09-25verified
Fast serving framework for LLMs, VLMs and diffusion models (built-in SGLang Diffusion), with RadixAttention prefix caching and structured output. NVIDIA, AMD, Intel, TPU and Apple Silicon.
★ 36.8kApache-2.0v0.5.21 · 2026-10-02verified
oobabooga's desktop app for local LLMs, renamed from text-generation-webui. llama.cpp, ExLlamaV3, Transformers and TensorRT-LLM backends; vision and tool-calling.
★ 47.7kAGPL-3.0v4.9 · 2026-05-20verified
High-throughput inference engine with PagedAttention. Designed for serving, not desktop chat — pair with Open WebUI or LiteLLM.
verified
GUI applications wrapping a local runtime in a chat interface.
Workspace-style chat with built-in RAG. Works fully offline with a local LLM provider.
verified
Desktop chat with branching conversations and parallel-model comparison; local engines (Ollama, llama.cpp, MLX) are in the free tier.
verified
Self-hosted ChatGPT-style web UI. Pair with Ollama or any OpenAI-compatible local server. Source-available under the Open WebUI License (BSD-style plus a branding clause).
verified
Voice-to-text, translation and subtitles, all on your own computer: transcription in 30 languages plus 22 Chinese dialects (Qwen3-ASR 0.6B / 1.7B), offline translation across 38 languages (Hunyuan MT2), text / SRT / VTT subtitles whose timings survive translation, and a voice keyboard that types into any app — the transcript or its translation. Listens to the microphone, a file, or system audio; runs on CPU or any Vulkan 1.2+ GPU (NVIDIA, AMD, Intel). Voice training from your own corrections and an MCP server for agents come with a paid licence; 14-day trial. Network use is a licence check, an update check and the model downloads you start — no audio, no text. Disclosure: maintained by us.
Desktop app that transcribes and translates audio offline with Whisper (whisper.cpp with Vulkan, CUDA, Apple Silicon). Handles live microphone input, files and folders, speaker identification, and exports TXT/SRT/VTT. MIT.
★ 21.8kMITv1.4.5 · 2026-08-23verified
Hold-to-talk dictation into any text field, with local Whisper, Parakeet v3 and Canary models that work offline after one download, no account or API key. Cloud transcription and AI clean-up are optional extras, off the offline path.
★ 8MITv1.0.0 · 2026-09-08verified
Whisper reimplemented on CTranslate2, up to 4x faster than openai/whisper at the same accuracy with less memory, plus 8-bit quantisation on CPU and GPU. The engine under WhisperX and RealtimeSTT. MIT.
★ 25.7kMITv1.2.1 · 2025-10-31verified
Free, open-source push-to-talk dictation into any app, working completely offline with Whisper or Parakeet models.
★ 32.9kMITv0.9.8 · 2026-10-03verified
Very low-latency streaming speech-to-text for on-device use; no account or API key. Code and default models MIT (legacy non-English models are non-commercial).
★ 11.2kv0.1.5 · 2026-08-24verified
Fast multilingual ASR model (CC-BY-4.0). Run it locally with onnx-asr, parakeet-mlx or NVIDIA NeMo.
verified
Reference Python implementation. Accurate but slower than the C++ ports; useful when you need the exact research behaviour.
★ 110kMITv20250625 · 2025-06-26verified
Low-latency streaming wrapper around faster-whisper for live dictation pipelines.
★ 10.2kMITv1.1.2 · 2026-08-30verified
Offline speech toolkit from the Next-gen Kaldi team, on onnxruntime with no internet: streaming and file STT, TTS, VAD, diarization, keyword spotting and speech enhancement on Linux, Windows, macOS, Android, iOS and Raspberry Pi. Apache-2.0.
★ 15.1kApache-2.0v1.13.8 · 2026-09-10verified
Microsoft's MIT-licensed voice models: VibeVoice-ASR (long-form transcription with speaker labels, 50+ languages, streaming variant) and a realtime 0.5B TTS.
★ 54.6kMITlast push verified
Lightweight offline speech recognizer with 20+ language models. Real-time on CPU.
verified
C++ port of OpenAI Whisper with GGUF quantization. Runs on CPU, Metal, CUDA, Vulkan.
★ 54.1kMITv1.9.4 · 2026-09-11verified
faster-whisper plus forced alignment, voice-activity detection, and speaker diarization.
★ 24.4kBSD-2-Clausev3.8.6 · 2026-05-25verified
Resemble AI's MIT-licensed voice-cloning TTS family: Multilingual (0.5B), Turbo (350M, supports tags like [laugh]) and Nano (110M, CPU).
★ 26.7kMITv0.1.2 · 2025-06-13verified
The maintained fork of the Coqui TTS toolkit (the original repo is unmaintained). Multiple architectures (VITS, XTTS) and voice cloning. pip install coqui-tts.
★ 2.3kMPL-2.0v0.27.5 · 2026-01-26verified
Tiny 82M-param TTS model (Apache-2.0), surprisingly natural for the size and fine on low-end hardware. Run it with Kokoro-FastAPI (OpenAI-compatible server) or kokoro-onnx.
verified
Lightweight TTS from Kyutai built to run on a CPU, no GPU PyTorch needed. Code MIT, weights CC-BY-4.0.
★ 9.8kMITv3.3.0 · 2026-09-24verified
Fast neural TTS, dozens of voices and languages, built for Raspberry Pi-class hardware. Development moved from rhasspy/piper (archived) to the Open Home Foundation; now GPL-3.0.
★ 5.8kGPL-3.0v1.8.0 · 2026-09-04verified
Multilingual TTS with voice design and cloning from OpenBMB; Apache-2.0 code and weights. pip install voxcpm.
★ 38.3kApache-2.02.0.3 · 2026-05-11verified
Node-graph workflow editor for diffusion models. Powers most modern local image and video pipelines.
★ 136.1kGPL-3.0v0.38.0 · 2026-09-29verified
The maintained continuation of lllyasviel's Forge (the original has not moved since mid-2025). Low-VRAM A1111-style UI; the default Neo branch adds newer model support. AGPL-3.0.
★ 1.8kAGPL-3.02.29.2 · 2026-10-01verified
Pro-grade image-generation app with a layer-based canvas, inpainting and node workflows (FLUX, SDXL and more). Apache-2.0 and community-maintained since the founding team joined Adobe and the hosted service shut down (Oct 2025).
verified
CLI-native local image, video and 3D generation in Rust (Candle, CUDA/Metal), with an MCP server, for people, scripts and agents.
★ 51MITv0.32.0 · 2026-09-26verified
Native Mac image generator and editor running open models (FLUX.2 Klein, Krea 2, Qwen Image Edit, Z-Image and more) entirely on Apple silicon via MLX. No account, no telemetry. Closed source; free tier includes every model, PRO adds formats and advanced nodes. macOS 26.2+.
verified
All-in-one WebUI for image and video generation on Diffusers, with dozens of models, SDNQ quantisation and balanced offload for low VRAM. Runs on NVIDIA CUDA, AMD ROCm/ZLUDA, Intel Arc, OpenVINO, DirectML and Apple MPS.
★ 7.4kApache-2.0last push verified
Modular UI built on top of ComfyUI. User-friendly mode out of the box, full node-graph available when you need it.
★ 4.6kMIT0.9.8-Beta · 2026-02-06verified
ComfyUI nodes drive Lightricks LTX video models for text-to-video and image-to-video generation. The chunked-loop pattern (released in our comfyui-workflows) produces longer outputs than vanilla LTX allows.
★ 136.1kGPL-3.0v0.38.0 · 2026-09-29verified
Audio-driven talking avatars generated on an Android phone; Experience mode runs fully offline, no account or API key. Code MIT; the lip-sync weights and default avatar are CC BY-NC 4.0 (non-commercial).
★ 20v1.0.0-20260912 · 2026-09-12verified
Low-VRAM video, image and audio generator for consumer GPUs ("for the GPU Poor"), covering Wan 2.1/2.2, LTX-2, Hunyuan Video, Qwen Image, Z-Image, Flux and Qwen3 TTS in one web UI.
★ 10klast push verified
Terminal pair-programming. Bring-your-own-LLM via LiteLLM — run with Ollama or any OpenAI-compatible local endpoint.
verified
Open-source coding agent for VS Code, JetBrains, the terminal and a desktop app, with human-in-the-loop approval and MCP support. Runs fully locally against Ollama, LM Studio or any OpenAI-compatible server. Apache-2.0.
★ 69.9kApache-2.0desktop-v0.0.43 · 2026-10-02verified
Vim/Neovim plugin for fill-in-the-middle completions and instruction-based editing from a local llama.cpp server. No cloud.
★ 2.2kMITv0.1.0 · 2026-08-24verified
Self-hosted GitHub Copilot alternative: local model serving with completion, chat and IDE plugins, no DBMS or cloud needed. Development has slowed (last release v0.32.0, Jan 2026).
★ 33.9kv0.32.0 · 2026-01-25verified
VS Code assistant that stays on your network: autocomplete, chat, inline edit, code review and an experimental agent mode against Ollama, LM Studio or llama.cpp. MIT, no telemetry, no sign-in; an optional team gateway is paid above 5 seats.
★ 3.7kMITv4.0.20 · 2026-09-22verified
Aider's planning mode separates "decide" and "edit" steps; works well with strong local reasoning models.
verified
General-purpose agent (desktop app, CLI and API) that runs on your machine, with 70+ MCP extensions. Use Ollama for a fully local setup. A Linux Foundation (Agentic AI Foundation) project, formerly block/goose. Apache-2.0.
★ 55kApache-2.0v1.53.0 · 2026-10-02verified
Encrypted, append-only knowledge store for local agents (Rust CLI, MIT) with scoped, expiring grants over MCP/HTTP; no hosted service required, device sync optional. Developer alpha with no independent audit yet; installer builds auto-update from GitHub by default.
★ 5MITlast push verified
Rewritten in 2026 as a Rust coding agent (a fork of OpenAI's Codex) for open models. Fully local with --oss and Ollama or LM Studio as the provider.
★ 68.5kApache-2.0rust-v0.0.55 · 2026-09-30verified
Local-first memory for coding agents: a Rust CLI/TUI with project-scoped SQLite/FTS recall, audit and forgetting. No account, no cloud service.
★ 18MITv0.15.13 · 2026-09-15verified
Open-source (Apache-2.0) embedding database. Use it in-process with on-disk persistence (pip install chromadb) or as a local server (chroma run); Chroma Cloud is optional.
★ 29.4kApache-2.01.5.9 · 2026-05-05verified
Library for similarity search. The retrieval engine inside many of the others.
★ 41.1kMITv1.15.1 · 2026-09-16verified
Embedded vector database on the Lance columnar format. Runs in-process (Python, TypeScript, Rust) on local disk or object storage, with no server.
★ 11.6kApache-2.0v0.39.0 · 2026-09-17verified
High-performance vector DB. Self-host the open-source binary.
verified
Vector-search extension for SQLite that runs anywhere SQLite does (desktop, mobile, WASM), giving you a single-file local vector store. Pre-1.0 (alpha releases).
★ 8.2kApache-2.0v0.1.9 · 2026-03-31verified
Hybrid (vector + keyword) DB. Self-host the OSS distribution; cloud is optional.
verified
BAAI's BGE family. Strong English + multilingual variants. Run via llama.cpp, sentence-transformers, or fastembed.
★ 12.2kMITv1.4.2 · 2026-08-24verified
Lightweight CPU-friendly embedding library by Qdrant.
★ 3.2kApache-2.0v0.8.1 · 2026-09-22verified
Reference Python library for sentence + paragraph embeddings.
verified
Training toolkit with a web UI for LoRAs and fine-tunes of diffusion models (FLUX.1/2, Qwen-Image, Z-Image, SDXL, Wan 2.x, LTX-2 and more). Works on consumer GPUs; Apple Silicon is experimental.
★ 12.2kMITlast push verified
Config-driven fine-tuning framework. LoRA, QLoRA, full fine-tunes.
★ 12.5kApache-2.0v0.20.0 · 2026-09-30verified
Pipeline-parallel trainer for diffusion models. Multi-GPU LoRA on large image / video models.
★ 2kGPL-3.0last push verified
Apple's array/ML framework, native on Apple Silicon (unified memory, Metal), with CUDA and CPU backends on Linux. Train and infer without CUDA workarounds on M-series Macs.
★ 28.7kMITv0.32.3 · 2026-09-29verified
Open-source desktop app and Python library to run and fine-tune models locally (LLMs, vision, TTS, diffusion, embeddings). Training is about 2x faster with up to 70% less VRAM; NVIDIA, AMD and Intel GPUs or CPU.
verified
Self-hosted workspace tool with integrated RAG. Listed twice intentionally — strong both as a chat app and a RAG layer.
verified
Toolkit for building RAG pipelines. Works fully offline with local models + vector DBs.
★ 52.4kMITv0.14.25 · 2026-09-21verified
Open-source API layer for private AI apps on local models: messages API, document ingestion, RAG with citations, tools and MCP. Does not run models itself; point it at Ollama, llama.cpp or vLLM. Built-in workbench UI at /ui.
★ 57.6kApache-2.0v1.0.1 · 2026-06-18verified
Self-hosted meta-search engine. Pair with a local LLM for an offline Perplexity-style assistant.
★ 38kAGPL-3.0last push verified
Privacy-focused AI answering engine that runs entirely on your own hardware: SearXNG search plus a local LLM via Ollama.
★ 37kMITv1.12.2 · 2026-04-10verified
The two distros that come ready for local AI out of the box. The full, ranked comparison lives in awesome-linux-for-ai.
Arch-based desktop distro with a tuned kernel and recent NVIDIA / AMD drivers. Sane out-of-the-box for new GPUs (Blackwell, RDNA 4).
verified
System76's Ubuntu-based distro (24.04 LTS, COSMIC desktop) with dedicated NVIDIA ISOs, x86-64 and ARM64, that ship the proprietary driver for plug-and-play GPU work.
verified
Local AI server tuned by AMD engineers for Ryzen AI NPUs, Radeon and Strix Halo. Serves chat, coding, speech and image models over OpenAI, Anthropic and Ollama APIs, and also runs on other PCs. Apache-2.0.
★ 5.8kApache-2.0v2026.40.0 · 2026-09-30verified
Apple's LLM runtime on MLX for Apple Silicon. Generate, quantise, serve and LoRA/full-fine-tune Hugging Face models (pip install mlx-lm); the engine behind LM Studio's MLX backend.
★ 7.2kMITv0.31.3 · 2026-04-22verified
NVIDIA's open-source (Apache-2.0) LLM inference library for NVIDIA GPUs, Ampere to Blackwell, data-center and RTX. Linux only (x86_64 / aarch64); the fastest CUDA path for many models.
★ 14.8kv1.2.1 · 2026-04-20verified
Intel's inference toolkit. CPU, iGPU, dGPU (Arc), and NPU support for Intel laptops.
★ 10.9kApache-2.02026.4.1 · 2026-10-01verified
AMD GPU path for llama.cpp. Its HIP/ROCm backend ships prebuilt ROCm 10.0 binaries for Linux and Windows with every release; the Vulkan backend is the fallback for Radeon cards ROCm does not support.
★ 130.4kMITv0.5.0 · 2026-09-23verified
Maintained by Brethof AI. Companion to awesome-llms-txt and awesome-private-ai.
The phrase "local-first" has been stretched to mean almost anything in 2026. Lots of tools advertise on-device AI but then call out to remote APIs for "complex" prompts, sync your data to a "private" cloud bucket, or ship telemetry that re-introduces every privacy problem the local mode was meant to solve.
This list applies a strict rule: the tool must perform inference on hardware you own and not transmit prompts, embeddings, or audio off that hardware during normal use. Optional cloud features (like model download or update checks) are fine; mandatory ones disqualify.
Where a tool is partially-local (e.g. ships a fully-local mode but defaults to cloud), we say so explicitly. Every link is checked — a CI-style sweep cuts entries whose URL 404s.
To be listed:
New or small is fine: we don't turn a tool away for being young or little-known. Every entry is labelled honestly instead — 🆕 marks a listing from the last 60 days — and our weekly check removes anything that stops working or goes six months without a release, commit or merged PR. We decline only tools with nothing real to point at, misleading claims, or no local mode at all.
llms.txt for agent discovery.Open an issue with the tool name, repo or homepage URL, the category it
should land in, and one paragraph on why it's worth listing. Entries live
as one YAML file each under entries/; this README is generated from them,
so edit the YAML, not the list above. We do not list tools whose offline
mode is gated behind a paid plan.
MIT.
Maintained by Brethof AI — AI tools built for people who take their data seriously.
Local speech-to-text that learns your voice. Perpetual licence. Our flagship.
PAID · flagship
Memory for your AI agents — records, rules, the full history and the graph of decisions, already there when a session starts. It answers before you ask.
PAID · free tier
Print-ready digital models. GLB and OBJ included. Lifetime access.
PAID · digital catalog
Five printers and real capacity. No self-serve shop yet — tell us what you need and we quote it by email.
BY ARRANGEMENT · email us
Our YouTube channel. A cyber-tiger host walks through local AI tools and what they actually do.
CHANNEL · live
Curated GitHub lists for AI coding agents, MCP servers, local AI and Linux for AI. Every entry carries a source link.
FREE · curated
Long-form how-tos for local AI on Linux, Windows and macOS, with the configuration files included.
FREE · live
Negative-curation: practices and tools that waste your time, ranked. Receipts required.
FREE · live
Who we are, why we build privacy-first AI, and what we won't do.