ExLlamaV3New
Inference library for local LLMs on consumer GPUs with EXL3 quantization and tensor/expert parallelism. It succeeds ExLlamaV2.
★ 1.6kMITv1.5.4 · 2026-10-03verified
A curated list of AI tools, platforms, and services that publish llms.txt — making them discoverable to AI agents doing research on behalf of human users.
Showing 0 of 128
NameCategoryTagsStarsLatestListedNothing matches. Clear a filter or search for something shorter.
Inference library for local LLMs on consumer GPUs with EXL3 quantization and tensor/expert parallelism. It succeeds ExLlamaV2.
★ 1.6kMITv1.5.4 · 2026-10-03verified
Hugging Face's model-definition framework for text, vision, audio, video and multimodal models — PyTorch-based inference and training for 1M+ Hub checkpoints.
★ 167kApache-2.0v5.18.0 · 2026-09-30verified
Transformers is HuggingFace's foundational Python library for working
with transformer-based models. It provides unified from_pretrained()
APIs for loading thousands of models from the HuggingFace Hub — LLMs,
vision models, audio models, multimodal models — and running them for
inference, fine-tuning, and research.
While Transformers isn't an inference server in the vLLM / Ollama sense (it's a library, not a server), it's the entry point most researchers and engineers use when working with a new model before it has dedicated runtime support elsewhere. It's the canonical reference implementation that every inference server (vLLM, llama.cpp, MLC, TGI) compares against for correctness.
The library is deeply integrated with the rest of the HuggingFace stack: Datasets, Accelerate, PEFT, TRL, and the Hub itself — together forming the dominant open-source ML ecosystem.
open-source
library
HuggingFace
Open-source (Apache-2.0) desktop ChatGPT alternative that runs local LLMs offline, with optional connections to cloud models.
★ 44.8kv0.8.4 · 2026-07-23verified
Jan is a desktop application that provides a ChatGPT-like interface backed entirely by locally-running open-weight LLMs. It bundles a llama.cpp-based inference engine, model discovery (curated HuggingFace models with one-click install), a conversation interface, and an OpenAI-compatible local server — all in a single installable app for Linux, macOS, and Windows.
Jan's positioning is privacy-first: local models run fully offline with no account. Cloud models (OpenAI, Anthropic, Gemini and others) can be added with your own keys when you want them. Its codebase is Apache-2.0 open-source and actively developed by Menlo Research, with frequent model library updates tracking the open-weight state of the art.
The OpenAI-compatible server means Jan doubles as a local backend for
other tools — coding assistants, agent frameworks, or any application
that speaks the OpenAI API can point at http://localhost:1337 and use
whatever model Jan has loaded.
open-source
local
Menlo Research
Single-binary llama.cpp wrapper with KoboldAI-style UI for chat, story-writing, and RP.
★ 11.9kAGPL-3.0v1.122.1 · 2026-09-26verified
Reference C++ implementation for running LLaMA-family and other transformer models with GGUF quantization.
★ 130.4kMITv0.5.0 · 2026-09-23verified
llama.cpp is the foundational open-source C++ inference engine that made local LLM execution practical on consumer hardware. It defined the GGUF model format now used across the ecosystem, pioneered aggressive quantization techniques (Q4_K_M, Q5_K_S, Q8_0, and dozens of variants), and supports more hardware backends than any other runtime: CPU (AVX, AVX2, AVX-512, NEON), Metal, CUDA, HIP/ROCm, MUSA, Vulkan, SYCL, CANN, OpenCL, WebGPU and Snapdragon Hexagon.
It ships as a library, a CLI (llama-cli), a server (llama-server)
with an OpenAI-compatible API, and bindings for nearly every programming
language. Projects like Ollama, LM Studio, KoboldCpp,
text-generation-webui, Jan, and LocalAI are all built on or around
llama.cpp. The official Llama app at llama.app, from the llama.cpp team
and Hugging Face, wraps it as a one-click Mac/Windows menu-bar app with
a local OpenAI-compatible API, and also installs the command-line
version on Linux.
For developers who want the lowest-level local inference or need to target unusual hardware (embedded, edge, non-CUDA GPUs), llama.cpp is the direct answer.
open-source
library
ggml-org
Desktop app for running local LLMs (llama.cpp and MLX) with an OpenAI-compatible server — now with Bionic, an agent for work and code, and optional US-hosted open-model inference.
verified
LM Studio is a desktop application for running open-weight LLMs on your
own machine. It handles model discovery and download from Hugging Face,
per-model GPU and context configuration, and chat. Its OpenAI-compatible
local server and lms CLI let coding agents and other apps use whatever
model is loaded. The runtime uses llama.cpp, plus Apple's MLX on Apple
Silicon.
In 2026 the app became LM Studio Bionic, built around Bionic, "LM Studio's agent for open models". Bionic creates and edits documents, helps with code, runs automations and can control the computer, and it includes local real-time voice transcription. Local use stays free. Paid plans (Bionic+ at $20/month and Pro at $100/month) add US-hosted, zero-data-retention inference for large open models such as Kimi K3, GLM 5.3 and DeepSeek V4, plus web search. LM Link connects multiple devices.
lms CLI; lmstudio-js and lmstudio-python SDKsfreemium
local
Element Labs, Inc. (LM Studio)
Open-source (MIT) self-hosted AI engine — OpenAI-, Anthropic-, Ollama- and ElevenLabs-compatible APIs for text, voice, vision, image, video and agents, no GPU required.
★ 49.4kMITv4.11.0 · 2026-10-02verified
LocalAI is a self-hosted, OpenAI-compatible inference server designed as a drop-in replacement for the OpenAI API. It supports text generation, chat completions, embeddings, transcription (Whisper), text-to-speech, image generation (Stable Diffusion), and function calling — all via the same HTTP endpoints applications already use with OpenAI.
It ships as a single binary or Docker container and runs on consumer CPUs (no GPU required), NVIDIA CUDA, AMD ROCm, Vulkan, and Intel GPUs. Under the hood it dispatches to multiple backends: llama.cpp, whisper.cpp, Stable Diffusion, Bark, and others — unified behind one API.
LocalAI is a common pick for organizations migrating existing OpenAI- dependent applications to on-premise inference without rewriting them.
open-source
self-hosted
LocalAI
Universal LLM deployment via compiled kernels — runs on iOS, Android, WebGPU, Vulkan, CUDA.
verified
Run open models locally via a single-binary server with a built-in model library — with optional Ollama Cloud for models too big for your hardware.
★ 182.2kMITv0.35.1 · 2026-09-29verified
Ollama is the most widely adopted tool for running open-weight LLMs on personal hardware. It bundles model download, quantization, a server exposing an OpenAI-compatible API, and a CLI into a single binary with zero-config defaults. The built-in model library covers Llama, Qwen, Gemma, Mistral, DeepSeek, and dozens of others with preconfigured quantizations.
Its appeal is friction: ollama run gemma4 downloads the model and
drops you into a chat in one command. The OpenAI-compatible API makes
Ollama a drop-in backend for thousands of tools originally built against
OpenAI — switching from cloud to local becomes a URL change.
Ollama runs on CPU, NVIDIA CUDA, AMD ROCm, and Apple Silicon, and serves as the default "local inference" option for open-source agent frameworks, chat UIs, and coding assistants.
The same app, CLI and API can also run models on Ollama's hosted cloud (see the Ollama Cloud entry). Cloud features can be disabled for local-only use.
open-source
local
Ollama
Self-hosted, feature-rich chat interface for local and cloud LLMs — the "ChatGPT clone" of the open-source world.
★ 154kv0.11.4 · 2026-09-21verified
Open WebUI is a self-hosted web interface for chatting with LLMs, supporting Ollama, any OpenAI-compatible endpoint, and cloud providers. It's visually and functionally comparable to ChatGPT's UI: conversation history, folder organization, model picker, custom system prompts, multimodal input (images, documents, audio), code interpreter, web search, and a plugin/tool system.
It's deployed as a single Docker container and commonly paired with Ollama to give non-technical users a friendly front-end to local LLMs. Multi-user support with role-based access makes it suitable for teams or family setups on a shared home server.
The project has strong momentum among self-hosters and privacy-conscious users who want the ChatGPT experience without sending data to OpenAI.
source-available
self-hosted
Open WebUI
Fast LLM and VLM serving runtime with RadixAttention cache and structured output support.
verified
Open-source desktop app for local LLMs (formerly Text Generation WebUI) with chat, vision, tool-calling, UI and API.
★ 47.7kAGPL-3.0v4.9 · 2026-05-20verified
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
★ 93.2kApache-2.0v0.31.0 · 2026-10-05verified
vLLM is the leading open-source inference engine for serving LLMs in production. It introduced PagedAttention (a memory management technique adapted from virtual memory in operating systems) that dramatically reduces GPU memory fragmentation when serving many concurrent requests, along with continuous batching that maximizes GPU utilization by folding new requests into in-flight batches.
vLLM exposes an OpenAI-compatible HTTP API, supports virtually every open-weight architecture (Llama, Qwen, DeepSeek, Mistral, Gemma, Phi, MPT, Falcon, Yi, Cohere, and dozens more), handles both dense and MoE models, and scales to multi-GPU and multi-node deployments with tensor and pipeline parallelism.
Where Ollama targets single-user local use, vLLM targets server-side production: high request concurrency, low latency, and maximum throughput per dollar of GPU.
open-source
self-hosted
vLLM Project
Managed gateway in front of OpenAI, Anthropic, Bedrock and dozens of providers — analytics, caching, rate limiting and model fallback via one OpenAI-compatible endpoint.
verified
Cloudflare AI Gateway is a managed proxy that sits between an application and its AI providers. Requests go through a Cloudflare endpoint, either a unified OpenAI-compatible API or each provider's native API, and the gateway adds logging, analytics, response caching, rate limiting, retries and model fallback, without running any infrastructure.
It supports a wide catalogue of providers (OpenAI, Anthropic, Azure OpenAI, Amazon Bedrock, Google, Groq, Cerebras, Baseten, Cartesia and more) and Cloudflare's own Workers AI. Because it runs on Cloudflare's edge network, it is a low-effort way to get centralized observability and cost control for LLM traffic. The trade-off is that it is a hosted service, not something you self-host.
freemium
saas
Cloudflare
Unified OpenAI-compatible proxy and SDK that routes calls across 100+ LLM providers with load balancing, fallbacks, and cost tracking.
★ 60.1kv1.104.0 · 2026-10-03verified
LiteLLM solves the "every LLM provider has a slightly different API" problem by exposing a single OpenAI-compatible interface that internally translates to 100+ providers: OpenAI, Anthropic, Google, Azure, AWS Bedrock, Cohere, Mistral, OpenRouter, Together, Fireworks, Groq, xAI, Ollama, vLLM, and many more.
It ships in two forms: a Python SDK for embedding in applications, and a standalone proxy server (LiteLLM Proxy) that accepts OpenAI-format requests and routes them across providers with load balancing, automatic retries, fallback chains, cost tracking, rate limiting, and virtual API key management for multi-team / multi-user deployments.
LiteLLM is often the backbone piece in organizations that want provider flexibility — swapping models without touching application code, or automatically failing over from a cloud provider to a local one.
open-source
library
BerriAI
One API for hundreds of models from many providers, with routing and fallbacks and no subscription.
verified
Google's open-source, code-first toolkit for building, evaluating and deploying multi-agent systems in Python, TypeScript, Go, Java and Kotlin.
★ 21.7kApache-2.0v2.11.0 · 2026-10-02verified
Agent Development Kit (ADK) is Google's open-source (Apache-2.0)
framework for building, evaluating and deploying AI agents. Agents can
be LLM-driven (LlmAgent) or deterministic workflow agents
(SequentialAgent, ParallelAgent, LoopAgent), and they compose
hierarchically into multi-agent systems that delegate through
LLM-driven transfer or by calling other agents as tools.
ADK covers tools (function tools, agents as tools, built-in code
execution and search), callbacks, sessions and state, long-term
memory, artifacts, planning, bidirectional streaming for live and
voice agents, built-in multi-turn evaluation, a CLI and a developer
UI. It is optimised for Gemini but model-agnostic through its
BaseLlm interface. It deploys anywhere, with paths to Google Cloud.
open-source
library
Open-source framework and runtime for agent platforms — build with the Agno SDK, run on the AgentOS runtime, manage from the AgentOS control plane.
★ 42.6kApache-2.0v3.1.1 · 2026-10-02verified
Python framework for orchestrating role-based multi-agent systems with sequential and hierarchical workflows.
★ 59.4kMIT1.15.23 · 2026-09-28verified
CrewAI is an agent framework centered on the metaphor of a "crew" — multiple specialized agents (roles) collaborating on a task (a crew). Developers define agents with a role, goal, and backstory, then compose them into tasks with explicit dependencies. CrewAI handles the delegation, inter-agent messaging, and result aggregation.
Compared to LangGraph's explicit-graph approach, CrewAI is higher-level and more opinionated: it targets developers who want a ready-made pattern for "researcher + writer + reviewer" style workflows without hand-rolling state machines. Its Flows feature adds event-driven workflows and deterministic control flow for production deployments.
The framework has significant adoption for content-generation pipelines, research assistants, and business-process automation where the task naturally decomposes into specialized agent roles.
open-source
library
CrewAI
Framework for programming rather than prompting LLMs — composable modules with optimizers.
verified
Open-source Python/TypeScript framework for building LLM agents — create_agent harness with middleware, built on LangGraph, with hundreds of integrations.
★ 147.5kMITlangchain-core==1.6.6 · 2026-09-29verified
LangChain is an open-source framework for building agents and LLM
applications in Python and TypeScript. In the 1.x line its centre is
create_agent, a minimal, configurable agent harness composed from a
model, tools, a system prompt and middleware. Middleware covers things
like guardrails, retries, routing and tool policies. The framework
keeps hundreds of integrations with model providers, vector stores and
tools.
LangChain agents are built on top of LangGraph, the lower-level orchestration runtime, so they inherit durable execution, persistence and human-in-the-loop support. Deep Agents, built on LangChain agents, is the batteries-included option with context compression, a virtual filesystem and subagents. LangSmith is the company's commercial platform for tracing, evaluation and deployment.
All LangChain, LangGraph and LangSmith docs live at docs.langchain.com, which publishes both llms.txt and a single llms-full.txt.
open-source
library
LangChain
Graph-based library for building stateful multi-agent workflows with explicit control flow and durability.
★ 42.7kMITcli==0.4.32.dev0 · 2026-09-23verified
LangGraph is LangChain's successor-style library for building agents and multi-step LLM workflows as explicit directed graphs. Where classic LangChain agents abstract control flow behind a reactive loop, LangGraph makes nodes, edges, state, and branching first-class — letting developers write agents that look more like state machines than magic boxes.
Its standout features are durability (graphs can pause, persist state, and resume after arbitrary interruptions — critical for long-running agent workflows) and first-class support for human-in-the-loop patterns (nodes that suspend execution until a human approves or edits the next step). LangSmith Deployment (formerly LangGraph Platform) runs graphs as long-running production services with built-in checkpointing.
open-source
library
LangChain
Type-safe Python library for building LLM-powered functions with structured outputs.
verified
TypeScript framework for AI agents and apps with memory, tools, MCP and observability built in.
★ 28.6k@mastra/[email protected] · 2026-10-05verified
Microsoft's open-source framework for production AI agents and multi-agent workflows in Python, .NET and Go, and the successor to AutoGen.
★ 13.9kMITpython-1.20.0 · 2026-10-02verified
Microsoft Agent Framework (MAF) is Microsoft's open-source, MIT-licensed framework for building, orchestrating and running AI agents and multi-agent workflows. It replaces AutoGen. The AutoGen repository is now in maintenance mode and names MAF as its enterprise-ready successor.
MAF has consistent APIs in Python and C#/.NET, plus a Go SDK in a separate repository. Workflows are graphs and support sequential, concurrent, handoff and group-collaboration patterns, with checkpointing, streaming and human-in-the-loop control. Agents can be defined declaratively in YAML, extended with middleware, and given skills built from files, inline code or class libraries.
Tracing, monitoring and debugging go through OpenTelemetry. DevUI is an interactive UI for developing and testing agents. Agents can be deployed to Foundry-hosted infrastructure, and the framework supports multiple LLM providers.
open-source
library
Microsoft
Agent framework built on Pydantic with type-safe tool use and structured responses.
verified
Minimal agent library from HuggingFace centered on code-writing agents.
verified
TypeScript toolkit for building AI apps with unified APIs across providers and framework helpers.
verified
Official Anthropic client SDKs for the Claude API in Python, TypeScript, C#, Go, Java, PHP and Ruby, plus the ant CLI.
★ 4kMITv1.11.0 · 2026-09-30verified
The Anthropic SDK is the lower-level, direct interface to the Claude API. Where the Claude Agent SDK provides a full agent loop with tool dispatch and subagents, the Anthropic SDK gives developers raw access to the messages API — single turns, streaming, tool calling, prompt caching, vision, and the full range of Claude models.
It's the right choice for applications that want Claude's raw capabilities without the agent-loop abstraction: chatbots, content generation, one-shot classification, structured extraction, and any integration where the developer wants to own the control flow.
ant CLI for shell usecommercial
library
Anthropic
Anthropic's official SDK for building custom agents on top of Claude with tool use, subagents, and hooks.
★ 8.2kMITv0.2.163 · 2026-09-30verified
The Claude Agent SDK is Anthropic's library for building AI agents powered by Claude. It exposes the same primitives Claude Code is built on — tool use, subagents, hooks, background tasks, session persistence, MCP servers — as a Python/TypeScript library developers can embed in their own products.
The SDK handles the agent loop, tool dispatch, error recovery, and context management. Developers define tools (functions the agent can call), subagent types (specialized agents for parallel subtasks), and hooks (policies that shape agent behavior) and the SDK orchestrates the rest.
It's the right choice when a project needs agent capabilities embedded in its own UI or CLI rather than running Claude Code in a terminal.
commercial
library
Anthropic
OpenAI's provider-agnostic multi-agent SDK with handoffs, guardrails, sessions, tracing and voice agents. It replaces Swarm.
★ 29.8kMITv0.23.1 · 2026-10-02verified
AI pair programming in your terminal — edits code across your git repo with commit-per-change discipline.
★ 49.4kApache-2.0v0.86.0 · 2025-08-09verified
Aider is an open-source CLI coding assistant that edits code in an existing git repository under the developer's direction. It's known for two core design choices: it operates on git (every change becomes a discrete commit with an auto-generated message, making rollback trivial), and it builds a "repo map" of the codebase so the model always has context for where functions and types live across files.
It supports any LLM provider — Anthropic Claude, OpenAI GPT, local models via Ollama or OpenAI-compatible servers, DeepSeek, xAI. Unlike IDE-coupled assistants, Aider lives in the terminal and plays well with any editor the developer already uses.
The /architect mode separates planning (using a reasoning model) from
editing (using a fast edit model) for complex changes. The benchmark
leaderboard Aider maintains is a widely-cited reference for LLM coding
performance.
open-source
cli
Aider-AI
Anthropic's terminal-first agentic coding assistant with deep tool use and codebase awareness.
verified
Claude Code is Anthropic's official terminal-based coding agent. It runs locally against the Claude API and has first-class tool use for reading and editing files, running shell commands, searching codebases, browsing the web, and orchestrating sub-agents. It ships with a hook system, slash commands, MCP server support, and IDE integrations (VS Code, JetBrains).
Unlike chat-only coding assistants, Claude Code is agentic by design: it plans, executes multi-step changes, reads its own output, and recovers from errors. It's built around the Claude Agent SDK and exposes the same primitives to developers who want to build custom agents.
Anthropic publishes both llms.txt and llms-full.txt for the Claude
Code documentation, making it one of the best-indexed coding-agent
references available to other AI assistants.
commercial
cli
Anthropic
Open-source autonomous coding agent with Plan/Act modes and MCP support, shipped as an IDE extension, CLI and SDK.
★ 69.9kApache-2.0desktop-v0.0.43 · 2026-10-02verified
AI coding agent and editor from Anysphere (acquired by SpaceX in 2026) — desktop app, CLI and cloud agents with a multi-vendor model picker.
verified
Cursor is an AI coding agent and editor built by Anysphere. SpaceX completed its acquisition of Cursor on 2026-08-14, following a model-training partnership with SpaceXAI announced in April 2026.
The desktop app is a VS Code-based editor with agent mode, inline edits, Tab autocomplete and codebase-aware context. Around it Cursor now ships a CLI agent, cloud agents that run in parallel on their own machines (or on machines you manage) and hand back work for review, integrations with Slack and GitHub, a mobile app, automations, and review bots for pull requests.
The model picker spans OpenAI, Anthropic, Google Gemini, SpaceXAI (Grok) and Cursor's own Composer models. Cursor is closed-source and subscription-based, with a free tier.
commercial
local
Anysphere (SpaceX)
Cognition's AI IDE (formerly Windsurf) that runs local and cloud coding agents from one Agent Command Center.
verified
Devin Desktop is the new name for Windsurf, the AI-native IDE that started life at Codeium and is now built by Cognition, the company behind the Devin coding agent. Windsurf became Devin Desktop on June 2, 2026, as an over-the-air update. It is the same editor with the same features. The windsurf.com and codeium.com domains now redirect to devin.ai/desktop.
The main change is the Agent Command Center, a Kanban-style view for running and reviewing many agents at once, both local agents on your machine and Devin Cloud agents. The classic IDE experience stays available: editor, Open VSX extensions, keybindings, workflows and LSPs. Cascade, the agentic assistant Windsurf was known for, is still there. So are Tab completions and inline Command edits.
The primary local agent is now Devin Local, the same agent harness that powers Devin CLI, with subagents, sandboxing, plan mode, worktrees and permissions. You can use Devin Desktop with local-only agents; a Devin Cloud subscription is not required.
commercial
local
Cognition
Google's open-source terminal agent for Gemini. Unpaid-tier and Google One users were moved to Antigravity CLI on 2026-06-18.
★ 107.2kApache-2.0v0.62.0 · 2026-09-29verified
GitHub's AI coding agent — completions, chat and agent mode in major IDEs, Copilot CLI, a desktop app, and a cloud agent that works from issue to pull request.
verified
GitHub Copilot is the original and most widely deployed AI coding
assistant, integrated natively into VS Code, Visual Studio, JetBrains
IDEs, Neovim, Xcode, and GitHub itself. It offers inline code completion,
chat, a multi-file agent mode, pull-request summaries, and command-line
assistance via gh copilot.
Under the hood Copilot routes to multiple LLMs — GPT, Claude Sonnet, Gemini — with model choice exposed to users on the paid tiers. Its reach within the GitHub platform (issues, PRs, code review, Actions) makes it a practical default for teams already on GitHub.
For enterprise deployments, Copilot offers SOC 2 compliance, data residency, code reference filtering, and the ability to restrict external traffic — the most mature enterprise posture in the category.
commercial
local
GitHub / Microsoft
MIT-licensed coding agent for VS Code, JetBrains, the CLI and the cloud — 500+ models at provider cost, bring-your-own-keys and local models.
★ 27.5kMITv7.8.3 · 2026-10-01verified
Kilo Code is an open-source (MIT) AI coding agent that works in VS Code, JetBrains, the terminal and the cloud. You pick from 500+ models, can switch mid-task, and pay the model provider's rate with no markup. Bring-your-own-key and local models are supported, and no API key is needed to start.
An agent command center manages local and cloud agents across IDE and CLI sessions, with parallel isolated worktrees. Kilo has been acquired by Anaconda, according to the banner on kilo.ai.
open-source
local
Kilo Code (Anaconda)
AWS's spec-driven AI coding agent (IDE, CLI, web, mobile) — the successor to Amazon Q Developer, whose IDE plugins reach end of support on 2027-04-30.
★ 4.3klast push verified
Kiro is AWS's AI coding agent, built around one agent harness shared by the Kiro IDE, Kiro CLI, web app, iOS app, Kiro Crew (for orchestrating teams of agents) and any ACP-compatible editor. Its defining idea is spec-driven development: Kiro turns a prompt into requirements, an architectural design and sequenced tasks, then implements them with parallel agents. It checks requirements for contradictions and uses property-based tests to catch edge cases that unit tests miss.
Steering files, event-driven hooks, custom agents and subagents, skills, "powers" and MCP shape the agent's behaviour. AWS directs Amazon Q Developer users to Kiro. The Q Developer user guide says the IDE plugins lose support on April 30, 2027, and Kiro publishes a migration guide. Kiro CLI is the successor to the Q Developer CLI.
Plans: Free (50 credits; open-weight models and Claude Sonnet 4.5), Pro ($20), Pro+ ($40), Pro Max ($100), Power ($200) and Enterprise. Paid plans include premium models such as Claude Sonnet 5 and Claude Opus 5.
commercial
local
AWS
Open-source (Apache-2.0) terminal coding agent forked from OpenAI's Codex, tuned for low-cost and open-weight models with switchable harness emulation.
★ 68.5kApache-2.0rust-v0.0.55 · 2026-09-30verified
OpenAI's coding agent, available as an open-source terminal CLI, an IDE extension and cloud automation.
★ 127.9kApache-2.0rust-v0.160.0 · 2026-10-01verified
Alibaba Qwen team's Apache-2.0 coding agent for terminal, editor, desktop, browser and chat — OpenAI, Anthropic, Gemini and Qwen APIs or local models.
★ 28.3kApache-2.0sdk-typescript-v0.1.18 · 2026-10-05verified
Qwen Code is "the open-source AI coding agent for your terminal, editor, desktop, browser, and chat", from Alibaba's Qwen team. It is multi-protocol: it supports the OpenAI, Anthropic, Gemini and Qwen APIs, as well as any third-party provider or local model through Ollama or vLLM, and you can switch at runtime.
It ships auto-memory, auto-skills, subagents, agent teams and MCP. Beyond the terminal there are IDE plugins (VS Code, JetBrains, Zed), a desktop app, a web UI, SDKs, GitHub Actions, and chat integrations for Telegram, DingTalk, WeChat and Feishu.
open-source
cli
QwenLM (Alibaba)
AI coding assistant for Sourcegraph Enterprise that pulls context from Sourcegraph code search across local and remote codebases.
verified
Node-based interface for building image, video, and audio generation workflows with any diffusion or multimodal model.
★ 136.1kGPL-3.0v0.38.0 · 2026-09-29verified
ComfyUI is the dominant open-source workflow tool for running diffusion models (Stable Diffusion, Flux, SDXL), video models (WAN, LTX, Hunyuan, Mochi, Kling), and audio/TTS models (Qwen3-TTS, F5-TTS, and others) via a visual node graph. Users connect nodes representing loaders, samplers, VAEs, text encoders, post-processors, and custom operations into directed acyclic graphs that execute end-to-end.
Unlike all-in-one UIs, ComfyUI is explicit about every step of the generation pipeline, which makes it preferred by power users who need to mix and match models, implement novel techniques (chunked S2V, looped img2vid, LoRA stacking, controlnet chains), and share reproducible workflows as JSON files.
A large ecosystem of custom nodes (ComfyUI-Manager, ComfyUI-Impact-Pack, KJNodes, easy-use, etc.) extends the base install with thousands of additional operations.
open-source
local
Comfy Org
Open-source LLM app development platform with visual prompt IDE, RAG pipelines, and agent builder in one product.
★ 157.9k1.17.1 · 2026-09-10verified
Dify is an all-in-one LLM application platform designed for teams shipping AI features without building each component from scratch. It combines a visual prompt IDE (version, test, and compare prompts with datasets), a RAG pipeline builder (document ingestion, chunking, embedding, retrieval tuning), an agent builder (tool-using agents with function calling), and a ChatUI / API endpoint generator in a single self-hostable product.
Where LangChain is a developer library and n8n is a horizontal automation platform, Dify is specifically a place for product teams to build, test, and operate LLM-powered apps — prompt engineering, RAG evaluation, and app deployment under one roof.
It's deployed via Docker Compose or Kubernetes; Dify Cloud offers the same product as managed SaaS. Licensed under a modified Apache 2.0 (the Dify Open Source License), which restricts multi-tenant SaaS use and removing Dify branding without a commercial license.
source-available
self-hosted
LangGenius
Open-source (MIT) low-code visual builder for AI agents, RAG and MCP workflows — run locally, via Docker, or as Langflow Desktop.
★ 155.5kMITv1.12.4 · 2026-09-29verified
Fair-code workflow automation with native AI nodes, 500+ integrations, and first-class self-hosting.
★ 206.7k[email protected] · 2026-10-05verified
n8n is a workflow automation platform — think Zapier or Make, but self-hostable, source-available, and extended with first-class AI nodes. Users build workflows visually by connecting nodes representing triggers (webhooks, schedules, email, Slack), integrations (500+ across databases, APIs, SaaS), and AI operations (LangChain-powered agents, LLM calls, vector store queries, retrievers).
The AI toolkit makes n8n one of the most pragmatic options for non- engineers building agent-style workflows: a "retrieve from Slack → summarize with Claude → post to Notion" flow is a 5-minute drag-and-drop job. For engineers, n8n's custom-code nodes let you drop into Python or JavaScript when the visual builder isn't enough.
Fair-code license means free for individual and internal business use, with commercial restrictions on building competing products. n8n Cloud provides managed hosting for teams that don't want to run Docker.
source-available
self-hosted
n8n
Offline voice-to-text, translation and subtitles for Linux and Windows: 30 transcription languages, 38 for translation, LoRA voice training.
Brethof Voice Pro is a commercial desktop app for Linux and Windows that transcribes, translates and subtitles speech entirely on the user's own machine. Transcription runs on the open-source Qwen3-ASR engine (0.6B or 1.7B) in 30 languages, and recognises 22 Chinese regional dialects on its own; translation across 38 languages runs locally on Tencent's Hunyuan MT2. It runs on the CPU or any Vulkan 1.2+ GPU (NVIDIA, AMD, Intel) — no CUDA required.
It types wherever the cursor is (the transcript or its translation), records the microphone, a file or system audio, and exports text, SRT and VTT subtitles whose timings survive translation. Voice training learns from the corrections the user already makes, and an MCP server lets AI agents transcribe and translate locally (both on paid licences). The network is touched only by a licence check, an update check and the model downloads the user starts. Perpetual licence; 14-day free trial.
commercial
local
Brethof AI
Deep-learning toolkit for TTS with multi-speaker models and voice cloning, now maintained in the Idiap fork.
★ 2.3kMPL-2.0v0.27.5 · 2026-01-26verified
Speech-to-text (Nova-3, Flux), text-to-speech (Aura) and voice-agent APIs, with SDKs, a CLI, an MCP server and a self-hosted option.
verified
Deepgram provides APIs for speech-to-text, text-to-speech, audio intelligence and end-to-end voice agents. Its STT models cover streaming and pre-recorded audio (Nova-3) and conversational, turn-aware streaming for voice agents (Flux). Aura models handle TTS. A Voice Agent API chains STT, an LLM and TTS in one connection.
Developers integrate through SDKs (Python, JavaScript, Go, .NET, Java), a dg CLI
that also runs as an MCP server, and agent skills. Enterprises can self-host.
freemium
saas
Deepgram
APIs and SDKs for text-to-speech, voice cloning, speech-to-text and conversational voice agents.
verified
High-quality open-source TTS with voice cloning from short audio reference.
★ 15.3kMIT1.1.22 · 2026-07-23verified
Fast, local neural text-to-speech engine with CLI, web server, Python and C/C++ APIs, now developed by the Open Home Foundation.
★ 5.8kGPL-3.0v1.8.0 · 2026-09-04verified
Open-source tokenizer-free TTS (VoxCPM2, 2B) with 30 languages, voice design from text prompts, controllable voice cloning and 48kHz output.
★ 38.3kApache-2.02.0.3 · 2026-05-11verified
VoxCPM is OpenBMB's open-source text-to-speech system. Instead of predicting discrete audio tokens, it generates continuous speech representations with an end-to-end diffusion-autoregressive architecture, which the authors credit for its natural, expressive output. The current release, VoxCPM2, is a 2B-parameter model trained on over 2 million hours of multilingual speech. It synthesizes 30 languages without a language tag, plus several Chinese dialects, and outputs 48kHz audio directly.
Its distinguishing features are Voice Design, which creates a new voice from a natural-language description (gender, age, tone, emotion, pace) with no reference audio, and Controllable Cloning, which clones a voice from a short clip while steering emotion and pacing. "Ultimate cloning" continues from a reference clip plus its transcript to preserve timbre and rhythm closely.
It ships as a pip package with a Python API, a voxcpm CLI, a local web demo, and
LoRA / full fine-tuning scripts with a training WebUI. For production it can be served
through Nano-vLLM or vLLM-Omni (OpenAI-compatible), and runs without Python via
llama.cpp-omni GGUF builds on CPU, CUDA, Metal or Vulkan. Code and weights are
Apache-2.0, free for commercial use.
voxcpm CLI and local web demo (CUDA, CPU, Apple MPS)open-source
library
OpenBMB
C++ port of OpenAI Whisper for local speech-to-text — no Python, runs on CPU and many GPU backends.
★ 54.1kMITv1.9.4 · 2026-09-11verified
whisper.cpp is a pure-C/C++ implementation of OpenAI's Whisper speech recognition model. It has zero runtime dependencies (no Python, no CUDA-specific tooling), ships as a library and CLI, and runs on every hardware backend the ggml-org ecosystem supports: CPU (with AVX / NEON SIMD), CUDA, Metal (Apple Silicon), Vulkan, OpenCL, SYCL, CoreML.
It uses GGML quantized models — the same format ecosystem as llama.cpp — enabling 4-bit, 5-bit, and 8-bit quantizations of the Whisper checkpoints that run on modest hardware with minimal quality loss. Real-time transcription is possible on consumer laptops, and batch transcription scales to longer audio with speaker diarization and word-level timestamps.
Widely used as a library inside desktop applications, mobile apps, and server-side transcription pipelines that don't want Python dependencies. Jan, KoboldCpp, and many other tools embed whisper.cpp.
open-source
library
ggml-org
Classic Stable Diffusion web UI with a large extension ecosystem — development has stalled (last release v1.10.1, Feb 2025; last commit Mar 2026).
★ 165.2kAGPL-3.0v1.10.1 · 2025-02-09verified
Hugging Face's library of pretrained diffusion pipelines for generating images, video and audio, with LoRA, quantization, offloading and training.
★ 34.7kApache-2.0v0.40.0 · 2026-08-20verified
Diffusers is Hugging Face's open-source (Apache-2.0) Python library
for state-of-the-art pretrained diffusion models that generate images,
video and audio. It is built around DiffusionPipeline, which offers
inference in a few lines, mix-and-match components (models,
schedulers) and adapters such as LoRA.
It includes memory and speed optimisations: offloading, quantization, caching techniques, attention backends and torch.compile. Hardware guides cover Apple MPS, Core ML, ONNX Runtime, OpenVINO and AWS Neuron. It also includes training scripts. Most image UIs and hosted services that serve open diffusion models use Diffusers or its model definitions.
open-source
library
Hugging Face
Free, open-source (Apache 2.0) self-hosted creative engine for AI image generation with a layer-based unified canvas and node workflows.
verified
Krita plugin for Stable Diffusion — inpaint, img2img, and generative layers inside Krita.
★ 10.7kGPL-3.0v1.53.0 · 2026-08-22verified
Advanced fork of SD WebUI with broader model support (Flux, Lumina, Kolors, more).
★ 7.4kApache-2.0last push verified
Open-source search infrastructure for AI — embedded, client-server or Chroma Cloud, with vector, full-text and (in Cloud) hybrid search.
★ 29.4kApache-2.01.5.9 · 2026-05-05verified
Chroma is a widely used open-source vector database / search infrastructure for LLM applications. It's designed with developer experience as the priority: a simple Python / JavaScript API, no cluster configuration needed to start, and sensible defaults for the RAG use case.
Chroma runs in three modes: embedded (in-process, SQLite-backed — ideal for prototyping and small apps), standalone server (Docker, single node), and Chroma Cloud (managed). The same API works across all three, letting projects graduate from local dev to production without rewrites.
It supports full-text search, metadata filtering, multi-modal embeddings, and several distance functions. Integrations with LangChain, LlamaIndex, LiteLLM, and HuggingFace embeddings make it a common first-pick for new RAG projects.
open-source
library
Chroma
Multimodal lakehouse and embedded vector database on the open Lance format — vector, full-text and hybrid search plus training-data curation.
★ 11.6kApache-2.0v0.39.0 · 2026-09-17verified
Open-source cloud-native vector database built for billion-scale similarity search with separation of storage and compute.
★ 46.3kApache-2.0v3.0.2 · 2026-09-20verified
Milvus is a distributed vector database designed for the largest vector workloads — billions of vectors, hundreds of billions of searches per day. Its architecture separates storage, compute, and coordination into independent services, letting each scale horizontally. At the cost of more operational complexity than smaller vector DBs, Milvus handles scale few other options reach.
Milvus supports multiple index types optimized for different trade-offs (IVF_FLAT, IVF_SQ8, HNSW, SCANN, DISKANN, GPU_IVF_FLAT for GPU-accelerated search), multiple distance metrics, hybrid search, and filtered search with complex boolean predicates. Attu provides a graphical admin UI for cluster operations.
For teams with smaller scale requirements, Milvus Lite ships as an embedded Python library using the same API as the full cluster — start local, graduate to distributed when the workload demands it.
open-source
self-hosted
Milvus / Zilliz
Postgres extension adding vector similarity search — the "just use Postgres" option for RAG.
★ 23.2klast push verified
pgvector is a PostgreSQL extension that adds vector similarity search to any Postgres database. It supports L2, inner product, and cosine distance; exact and approximate nearest-neighbor search via IVFFlat and HNSW indexes; and integrates with every Postgres client library on every platform.
The appeal is simplicity: organizations already running Postgres can add vector search without provisioning a separate vector database, managing a second replication topology, or introducing a new client library. ACID transactions that join vector data with traditional relational data work out of the box.
pgvector is supported by every major managed Postgres provider (Supabase, Neon, AWS RDS, GCP Cloud SQL, Azure Database for PostgreSQL, Crunchy Bridge), making it a safe default for RAG applications that don't have billion-scale vector requirements.
open-source
library
pgvector
Managed serverless vector database plus Nexus knowledge engine and Assistant — Pinecone's "AI knowledge platform" for agents and RAG.
verified
Pinecone is the original managed vector database and remains the most widely adopted SaaS option for production vector search. Its serverless architecture abstracts away every operational concern: no cluster sizing, no replication config, no reindexing during scaling events. Queries return in single-digit milliseconds at billion-vector scale.
Pinecone's product expanded beyond pure vector search to include Assistant (hosted RAG as a service — upload documents, get a chatbot endpoint), embedding inference (text → embedding via the same API that stores them), and reranking. For teams that want a complete managed RAG stack without operating any infrastructure, Pinecone's breadth is hard to match.
Trade-offs: proprietary (closed-source), most expensive option at scale versus self-hosted Qdrant/Milvus/Weaviate, and vendor lock-in for teams building on its Assistant layer. For small indexes it's often cheaper than self-hosting; for very large indexes the calculus shifts.
commercial
saas
Pinecone
Open-source, Rust-written vector database built for production scale — rich filtering, hybrid search, and multi-tenancy.
★ 34.9kApache-2.0v1.19.1 · 2026-09-04verified
Qdrant is a production-oriented vector database written in Rust, designed for organizations serving vector search at scale. It emphasizes advanced filtering (complex boolean queries on metadata alongside vector search), multi-tenancy (isolated collections per customer), hybrid search (dense
It offers a REST API, gRPC, and clients for Python, JavaScript, Rust, Go, Java, and C#. Qdrant Cloud provides managed hosting; self-hosted deployments scale horizontally through Qdrant's sharded cluster mode.
Compared to Chroma's "get started in five minutes" positioning, Qdrant targets production: teams that know they'll need 10M+ vectors, complex filters, and SLOs from the start.
open-source
self-hosted
Qdrant
Multi-model database in Rust (document, graph, vector, time-series, relational) positioned as a context and memory layer for AI agents.
★ 33.1kv3.3.0 · 2026-09-28verified
SurrealDB is a multi-model database that unifies document, graph, key-value, time-series, and vector workloads in a single Rust-written engine. Where traditional RAG stacks glue a vector database to a relational database to a graph store, SurrealDB lets a single query span all of them — joining user data, embeddings, and graph relationships without cross-system synchronization.
For AI applications the vector search capability supports HNSW and DISKANN indexes alongside scalar filters and full graph traversal. The same query can filter users by attribute, retrieve their document embeddings by cosine similarity, and walk the graph of their social connections — all in one SurrealQL statement.
SurrealDB runs embedded (in-process, WASM-compatible for browser / Node.js), as a standalone server, or in distributed cluster mode (SurrealDB Cloud, or self-hosted with TiKV). It's used increasingly as the backing store for agent memory, RAG systems, and real-time applications that need multi-model data without the operational overhead of multiple databases.
source-available
self-hosted
SurrealDB
Serverless vector and full-text search engine built on object storage — fast, low-cost and scaling to 1T+ documents.
verified
turbopuffer is a hosted vector and full-text (BM25) search engine built from first principles on object storage, with a memory/SSD cache in front. Storing data in object storage keeps costs low at very large scale. Hot namespaces are served from cache with low latency, and cold queries read from storage. The vendor reports production scale of 1T+ documents, 25k+ queries/s and 10M+ writes/s.
It offers strongly consistent queries, namespace branching (copy-on-write), sharding, built-in embedding and multiple regions.
commercial
saas
turbopuffer
Open-source vector database with built-in ML modules, hybrid search, and first-class RAG tooling.
★ 16.9kv1.39.9 · 2026-10-05verified
Weaviate is an open-source vector database written in Go, designed for AI-native applications. It distinguishes itself from minimal vector stores by bundling vectorization modules (OpenAI, Cohere, HuggingFace, Ollama, and more) directly into the database — you can ingest raw text and have Weaviate embed it for you, without a separate embedding service.
Weaviate supports hybrid search (dense + sparse, BM25F), multi-tenancy through tenants-per-class isolation, generative modules (RAG inside the database: query → retrieve → generate as one call), and gRPC and REST APIs (GraphQL is legacy). Cluster mode handles sharding and replication for production-scale deployments.
Weaviate Cloud (serverless managed) is available for teams that don't want to operate the cluster themselves.
open-source
self-hosted
Weaviate
All-in-one desktop and Docker RAG app — document ingestion, agents, multi-user.
verified
Production-oriented Python framework for building RAG, search, and agent pipelines with composable components.
★ 26.7kApache-2.0v3.3.0 · 2026-10-01verified
Haystack (from deepset) is an open-source Python framework for building LLM-powered pipelines — RAG, semantic search, question answering, and agents — out of composable, typed components. Its core abstraction is the Pipeline: a directed graph of components (retrievers, rankers, generators, prompt builders, validators) connected by typed edges, with full control over execution order and branching.
Haystack predates the LLM gold rush (originally an NLP QA toolkit) and its design reflects production priorities: typed I/O, explicit observability, easy component swapping, and deployment as Kubernetes-ready services. The 2.x rewrite rebuilt it around the Pipeline abstraction; Haystack 3.0 (July 2026) added a more capable Agent with hooks and skills, run introspection, first-class async serving and a leaner core.
It's a common choice for teams who find LangChain too abstract and LangGraph too low-level, wanting a middle-ground framework with production defaults.
open-source
library
deepset
Open-source Python framework for RAG and agents over private data — loaders, indexes, retrievers, query engines and workflows (from the makers of LlamaParse).
★ 52.4kMITv0.14.25 · 2026-09-21verified
LlamaIndex is a data framework for building LLM applications over private data. Where LangChain is general-purpose, LlamaIndex is specifically optimized for retrieval-augmented generation (RAG): ingesting documents, building indexes, retrieving relevant chunks at query time, and composing them into prompts.
It provides high-level abstractions for common RAG patterns (summary index, vector index, tree index, keyword index, knowledge graph index) and lower-level primitives for developers who want custom retrieval pipelines. Integrations span hundreds of data sources (PDFs, Notion, Slack, Google Drive, databases) and vector stores (Chroma, Qdrant, Weaviate, Pinecone, pgvector, LanceDB, …).
LlamaIndex Agents extend the framework with tool-using agents that can query multiple indexes, call APIs, and compose results. The company's hosted product is now LlamaParse (formerly LlamaCloud), a document parsing, extraction and classification platform.
open-source
library
LlamaIndex
Open-source (Apache-2.0) Claude-API-style layer for private AI apps on any local OpenAI-compatible model server — agentic RAG with citations, tools, MCP and data access.
★ 57.6kApache-2.0v1.0.1 · 2026-06-18verified
Memory for AI agents that is already there when a session starts — curated records, rules and the full history of what was said; it processes your memory and never stores it. Disclosure: maintained by us.
★ 0last push
Brethof Brain is a memory system for AI agents: memory that is already there when a new session starts, that hands the agent what matters the moment it is relevant, and that keeps everything said so it can be found again. Coding agents are its first audience, not the limit of it.
The memory curates itself. It reads what was said and keeps what a future session would need as curated records, reconciling new facts against old ones; standing rules arrive with almost every message, the records that bear on a prompt arrive with it, and a brief opens every session. When that is not enough, everything ever said can be searched.
Your memory lives on your own machine (local edition) or in our cloud, encrypted under a key only you hold (hosted edition). Our hub processes each exchange to make memory of it and stores none of it; the processing runs on secure compute in Zurich, and the model provider is bound by contract not to log, retain or train on what passes through. The client is source-available. Disclosure: maintained by us (Brethof AI).
freemium
saas
Brethof AI
Claude Code's own memory: CLAUDE.md (or AGENTS.md) instruction files you write, plus auto memory — notes Claude writes itself from your corrections — both loaded at the start of every session.
verified
Each Claude Code session starts with a fresh context window. Two built-in mechanisms carry knowledge across sessions. CLAUDE.md files are plain-text instructions you write, scoped to an organization (managed policy), to you across all projects (~/.claude/CLAUDE.md) or to a project; Claude Code can also read a repository's AGENTS.md, and path-scoped rules live in .claude/rules/. Auto memory is the other half: notes Claude writes itself from your corrections and preferences, kept per repository and shared across worktrees.
Both are loaded into every session (auto memory up to its first 200 lines or 25KB). Anthropic's documentation describes them as context, not enforced configuration: to block an action regardless of what Claude decides, use a hook instead. Subagents can keep their own auto memory.
commercial
cli
Anthropic
Persistent memory layer for AI agents — remembers user facts, preferences, and context across sessions.
★ 66.6kApache-2.0ts-v3.3.1 · 2026-09-25verified
Mem0 is a memory framework that gives AI agents persistent, queryable memory across conversations. It extracts facts from each interaction ("the user's dog is named Rex", "the user prefers Python over JavaScript"), stores them in a vector index (with entity-graph memory on the managed Platform), and retrieves relevant memories when the agent needs context for a new query.
Compared to raw vector databases where applications manage their own schema, Mem0 provides an opinionated memory abstraction: user-scoped memories, automatic fact extraction, memory updates (the dog's name changed? the old fact gets superseded), and ranked retrieval tuned for the agent-memory use case specifically.
It integrates with LangChain, LlamaIndex, CrewAI, and any OpenAI-compatible client. Open-source core under Apache 2.0 with a managed hosted tier (Mem0 Platform) for teams that don't want to operate the storage layer.
freemium
library
Mem0
BAAI's open embedding and reranker models with the FlagEmbedding toolkit for inference, evaluation and fine-tuning — a one-stop retrieval toolkit for search and RAG.
★ 12.2kMITv1.4.2 · 2026-08-24verified
Lightweight embedding library from Qdrant on ONNX Runtime (no PyTorch) — dense, sparse and late-interaction embeddings and rerankers, CPU or GPU.
★ 3.2kApache-2.0v0.8.1 · 2026-09-22verified
Python framework for state-of-the-art sentence, text, and image embeddings.
verified
MongoDB's Voyage AI embedding and reranking API — text, contextualized-chunk and multimodal embeddings for retrieval and RAG.
verified
Voyage AI, now "Voyage AI by MongoDB", provides hosted embedding models and rerankers through an API and Python client. The catalog includes text embeddings, contextualized chunk embeddings (chunks embedded with document context), multimodal embeddings and rerankers. Flexible output dimensions and quantization reduce storage cost, and a batch inference API handles large offline jobs.
Voyage models also power retrieval inside MongoDB's data platform. The API is commonly used as the embedding step in RAG stacks paired with any vector database.
commercial
saas
MongoDB
AI observability and evaluation platform built on OpenTelemetry — tracing, evals, prompt playground, datasets and experiments; self-host or Arize cloud.
★ 11.7karize-phoenix-v20.19.0 · 2026-10-01verified
Open-source AI gateway and LLM observability (requests, costs, latency, sessions); in maintenance mode since its 2026 acquisition by Mintlify.
★ 6.2kApache-2.0v2025.08.21-1 · 2025-08-21verified
Open-source LLM engineering platform for tracing, evaluation, prompt management, and observability — self-host or cloud.
★ 35.4kv4.50.0 · 2026-10-02verified
Langfuse is an open-source AI engineering platform (acquired by
ClickHouse, which kept it open source and self-hostable). It provides
traces, evaluations, prompt versioning, datasets and cost tracking, with
the core under the MIT license (enterprise features in ee/ directories
are commercially licensed) and a straightforward self-hosted option for
teams that need data sovereignty.
It integrates with LangChain, LangGraph, Llama-Index, OpenAI SDK, Anthropic SDK, LiteLLM, and any HTTP-based LLM call via its OpenTelemetry support. Traces capture full LLM inputs/outputs, tool calls, latencies, and costs. The prompt management UI lets product teams iterate on prompts outside of code and deploy changes without shipping releases.
Langfuse publishes both llms.txt and llms-full.txt at well-known
paths, making it easy for agents to answer questions about its APIs and
usage patterns.
ee/ features under a commercial license)open-source
self-hosted
Langfuse (ClickHouse)
Commercial observability, debugging, and evaluation platform for LLM and agent applications.
verified
LangSmith is LangChain's commercial platform for observing, debugging, evaluating, and improving LLM applications in development and production. It captures every LLM call, tool invocation, and chain step into traces that developers can replay, inspect token-by-token, and compare across prompt or model changes.
Core workflows it supports: inspecting why an agent made a particular decision, A/B testing prompts against datasets, running evals (LLM-as-judge, exact-match, custom), catching regressions before deploying prompt changes, and debugging production incidents with full call graphs.
While built by the LangChain team and deeply integrated with LangChain / LangGraph, LangSmith is framework-agnostic — applications using the raw OpenAI or Anthropic SDKs can send traces to LangSmith with minimal setup.
commercial
saas
LangChain
Reliability platform for production AI agents that combines tracing, calibrated LLM-as-judge evaluation, scenario simulation and guardrails.
verified
Noveum is a hosted reliability platform for teams running AI agents and LLM applications in production, including voice agents. Its open-source tracing SDK, NovaTrace (noveum-trace), records every LLM call, tool call, RAG step and agent hop as a structured trace. It integrates with LangChain, LangGraph, LiveKit, Pipecat and CrewAI.
On top of those traces, NovaEval scores chat agents with a library of 100+ calibrated LLM-as-judge scorers and returns the reasoning behind each verdict. NovaSynth simulates voice (SIP) and chat scenarios with personas, interruptions and virtualized tools to test agents end to end. NovaGuard (in beta) enforces policies on live traffic. NovaPilot turns failing evaluations into proposed fixes delivered as pull requests.
A remote MCP server lets MCP-capable coding assistants and chat clients work with Noveum traces, datasets and evaluations over an authenticated connection. An enterprise tier offers on-prem and self-hosted deployment. The hosted platform is proprietary; the tracing SDK is public.
freemium
saas
Noveum
Open-source (Apache-2.0) tracing, evaluation and prompt optimization for LLM apps, RAG and agents, from Comet — self-host or cloud.
★ 22.4kApache-2.02.2.89 · 2026-10-05verified
Opik is Comet's open-source platform for debugging, evaluating and monitoring LLM applications, RAG systems and agentic workflows. It records traces of LLM calls, tools and agent steps. It runs automated evaluations (LLM-as-judge and heuristic metrics) against datasets and experiments, and provides production dashboards and prompt optimization.
It ships an MCP server so coding assistants can read traces, find failing ones and check fixes against real data, plus a built-in assistant ("Ollie") for analysing traces. Opik is Apache-2.0 licensed and can be self-hosted or used as a managed service on Comet.
open-source
self-hosted
Comet
ML experiment tracking (W&B Models) and LLM/agent tracing and evaluation (W&B Weave), now part of CoreWeave Forge.
★ 11.3kMITv0.30.0 · 2026-09-09verified
Evals and observability platform for agents that traces production, runs evaluations and catches regressions before release.
verified
Open-source, pytest-style LLM evaluation framework with 50+ metrics for agents, RAG and chatbots, plus CI/CD integration.
verified
Open-source (MIT) framework for frontier LLM and agent evaluations from the UK AI Security Institute — 200+ pre-built evals, sandboxing and a log viewer.
★ 2.9kMITlast push verified
Inspect is a framework for frontier AI evaluations developed by the UK AI Security Institute and Meridian Labs. Evaluations are built from composable datasets, solvers (including agents), tools and scorers. Over 200 pre-built benchmark implementations are ready to run on any supported model.
It supports agent evaluations with built-in agents, multi-agent primitives, and the ability to run external agents such as Claude Code, Codex CLI and Gemini CLI. Tool calling covers custom and MCP tools plus built-in bash, Python, text editing, web search, browsing and computer tools. Untrusted model code runs in sandboxes (Docker, Kubernetes, Modal, Proxmox, Vagrant and others). Inspect View and a VS Code extension help monitor and debug runs.
open-source
library
UK AI Security Institute
Open-source (MIT) CLI for evaluating, red-teaming and security-testing LLM apps and agents from config files — now part of OpenAI.
★ 25.7kMIT0.123.1 · 2026-09-18verified
Open-source evaluation framework for LLM apps — RAG pipelines, agents and workflows — with objective metrics and test-data generation.
★ 15.9kApache-2.0v0.4.3 · 2026-01-13verified
All-in-one open-source training suite (GUI or CLI) for LoRAs and fine-tunes of image and video diffusion models — FLUX.1/FLUX.2, Qwen-Image, Z-Image, SDXL, Wan 2.x, LTX-2 and more.
★ 12.2kMITlast push verified
AI Toolkit by ostris is a widely used open-source framework for training LoRA adapters and full fine-tunes on diffusion image models. It supports FLUX.1 and FLUX.2 (including Klein), Qwen-Image, Z-Image, SDXL, SD1.5, and video models such as Wan and LTX-2, and newer architectures as they land — often within days of a new model's release.
Where generic HuggingFace fine-tuning is model-agnostic but optimized for neither, AI Toolkit is purpose-built for diffusion workflows: handles VAE encoding, text-encoder freezing, rank- constrained LoRA training, sample-during-training, cosine learning-rate schedules, and the dataset quirks (captions, aspect bucketing, regularization images) that matter for diffusion- specific quality.
The YAML-configured job format makes training runs reproducible and shareable. Common recipes (character LoRA, style LoRA, concept LoRA, full fine-tune) ship as working examples. Runs on consumer GPUs (16GB+ for LoRA, 24GB+ for full fine-tune) via bitsandbytes and 8-bit optimizers.
Critical tool for anyone producing image assets at scale on diffusion models — game asset pipelines, character consistency for video projects, branded content generation, or personal LoRAs for creative work.
open-source
library
ostris
YAML-configured fine-tuning framework supporting LoRA, QLoRA, full FT, DPO, and most modern LLM architectures.
★ 12.5kApache-2.0v0.20.0 · 2026-09-30verified
Axolotl is a full-featured open-source fine-tuning framework that wraps
HuggingFace Transformers, PEFT, and TRL into a YAML-configured CLI.
Users define their training run (model, dataset, hyperparameters,
adapter type) in a config file and launch with axolotl train config.yml.
Compared to Unsloth's speed focus, Axolotl prioritizes breadth: broader architecture support (Llama, Qwen, Mistral, Gemma, Phi, DeepSeek, Yi, Cohere, StableLM, and more), more training methods (LoRA, QLoRA, full FT, ReLoRA, continued pretraining, DPO, KTO, ORPO, GRPO), and richer dataset formats (SharePT, Alpaca, ChatML, raw, completion, many more).
It's a common choice for production fine-tuning pipelines where reproducibility and config-as-code matter more than squeezing the last percent of speed.
axolotl export for llama.cpp / Ollamaopen-source
library
Axolotl AI
WebUI-based fine-tuning framework supporting 100+ models with LoRA, QLoRA, DPO, and more.
★ 75.3kApache-2.0v0.9.5 · 2026-05-30verified
ModelScope's training and deployment framework for 600+ LLMs and 400+ multimodal models — SFT, GRPO-family RL, DPO, Megatron parallelism.
★ 15.8kApache-2.0v4.5.3 · 2026-09-08verified
Training API from Thinking Machines — write the training loop in Python, run LoRA fine-tuning and RL on open-weight models from 1B to 1T+ parameters on managed GPUs.
★ 4.2kApache-2.0v0.5.7 · 2026-09-03verified
Tinker is a training API for researchers and developers doing LLM post-training. You
write a simple loop that runs on a CPU-only machine, containing your data or RL
environment, loss function and evals. You call forward_backward, optim_step,
sample and save_state, and Tinker executes the exact computation efficiently on
distributed GPUs, handling hardware failures transparently.
It fine-tunes dense and mixture-of-experts open-weight models from 1B to 1T+ parameters, including vision-language models, using LoRA rather than full fine-tuning. Trained weights can be downloaded and served elsewhere. The open-source Tinker Cookbook provides recipes for SFT, RL and distillation.
commercial
saas
Thinking Machines Lab
Hugging Face's post-training library for transformer LLMs — SFT, GRPO, RLOO, DPO, KTO, reward modeling and distillation, with vLLM, PEFT and DeepSpeed integration.
★ 19.4kApache-2.0v1.14.1 · 2026-09-29verified
Open-source app and library to run and fine-tune models locally — about 2x faster training with ~70% less VRAM for LLMs, diffusion, TTS and embedding models.
★ 77.2kApache-2.0v0.1.902-beta · 2026-10-01verified
Unsloth is a fine-tuning library that rewrites HuggingFace's training code with hand-tuned Triton kernels and careful memory management to achieve roughly 2x speedups and 70% memory savings versus vanilla Transformers + PEFT. It supports LoRA, QLoRA, full fine-tuning, and continued pretraining across Llama, Qwen, Mistral, Gemma, Phi, and most modern architectures.
The project is known for free, working Colab notebooks that take users from zero to a fine-tuned model in under an hour — genuine entry-point material for the LLM fine-tuning community. It's compatible with the HuggingFace ecosystem: PEFT adapters, Trainer API, TRL DPO, datasets, and every output format (LoRA, merged, GGUF, AWQ) downstream tools expect.
The core package is Apache-2.0; some optional components (such as Unsloth Studio) are AGPL-3.0.
open-source
library
Unsloth AI
Independent web search API with no tracking — alternative to Google / Bing for agent use.
verified
Neural search API built for AI agents — semantic search across the web with content retrieval.
verified
Web data API that searches, scrapes and interacts with the web and returns clean Markdown or structured data for agents.
★ 188.8kAGPL-3.0v2.11.0 · 2026-06-19verified
Web APIs for AI agents — Search, Extract, Task / Deep Research, FindAll and Monitor — returning LLM-optimized excerpts.
verified
Parallel offers web-access APIs designed for AI agents rather than humans. Search returns LLM-optimized excerpts for natural-language objectives. Extract pulls clean content from URLs. The Task API runs deep research and data enrichment. FindAll discovers entities matching criteria, and Monitor watches the web for changes.
The docs include migration guides from Exa, Tavily and SERP APIs and an evaluation guide for comparing search quality inside an agent.
commercial
saas
Parallel Web Systems
Perplexity API Platform — Agent API for web-grounded answers with citations, plus Search, Router and Embeddings APIs, a CLI and an MCP server.
verified
Scraping API for Google, Bing, DuckDuckGo and 15+ search engines — structured JSON results.
verified
Web access layer for AI agents (by Nebius) — real-time search, extraction, crawl, map and cited research via API, CLI and MCP.
verified
Open-source document conversion toolkit (LF AI & Data, started at IBM Research) — PDF, Office, HTML, images, audio and more into Markdown/JSON for RAG and agents, with VLM and MCP support.
★ 68.4kMITv2.133.0 · 2026-10-03verified
Fast, accurate document conversion (PDF, images, Office, HTML, EPUB) to Markdown, JSON, chunks or HTML — tables, equations and structure preserved; code Apache-2.0, weights OpenRAIL-M with commercial limits.
★ 40.2kApache-2.0v2.0.0 · 2026-07-20verified
Agentic document platform — classify, parse, extract, split and edit documents into LLM-ready content and structured JSON via API, CLI or MCP.
verified
Reducto is a commercial document-processing platform for AI teams. Its APIs cover the whole document lifecycle. Classify routes documents by type. Parse produces layout-aware text, tables and figures with bounding boxes and RAG-ready chunking. Extract pulls schema-defined fields into JSON. Split separates bundled documents, and Edit fills forms and modifies templates. These steps can be chained into single-call pipelines.
Reducto Studio provides a UI for building document workflows. A CLI and an MCP server let coding agents call the platform directly.
commercial
saas
Reducto
Document ETL for GenAI — the open-source unstructured library plus the hosted Unstructured Transform API/MCP server, turning 65+ file types into RAG-ready elements and structured data.
★ 15.5kApache-2.00.27.10 · 2026-09-27verified
Inference and training platform — hosted Model APIs (OpenAI- and Anthropic-compatible), dedicated deployments of your own models, and fine-tuning on production GPUs.
verified
Baseten runs hosted models, deploys custom models and trains models on production GPU infrastructure. Model APIs expose high-performance LLMs through OpenAI- and Anthropic-compatible endpoints, with reasoning control, vision, server-side web search and coding-agent setups (Claude Code, Codex CLI, OpenCode).
For your own models, Baseten manages containers, GPU capacity across clouds and regions, autoscaling and observability, and its inference engines optimise supported architectures. It also covers audio (transcription, speech generation, diarization), async inference, structured outputs and function calling. Training uses Loops or your own training jobs, and the resulting checkpoints can be served on the same platform.
commercial
saas
Baseten
Wafer-scale inference cloud with an OpenAI-compatible API — thousands of tokens per second on open models, plus dedicated endpoints for custom weights.
verified
Cerebras Inference serves models on Cerebras' wafer-scale systems. Cerebras says this is up to 30x faster than GPU systems. The shared API's model catalog lists gpt-oss-120b at about 3,000 tokens/s and Qwen 3.8 27B at about 1,850 tokens/s.
The API is OpenAI-compatible and supports streaming, reasoning controls, structured outputs, tool calling, prompt caching, image inputs, batch jobs and service tiers. Dedicated Inference reserves capacity for an organisation, accepts custom model weights through a management API, and adds predicted outputs and Prometheus-compatible metrics. New accounts get $5 of free trial credits; there is also pay-as-you-go and enterprise pricing. It is also available through AWS Marketplace, Hugging Face and Vercel.
commercial
saas
Cerebras
Training and inference platform for open models — serverless and dedicated GPU deployments, fine-tuning, and Fireworks Nexus model routers for coding agents.
verified
Fireworks AI is a production inference platform specializing in serving open-source LLMs at low latency and high throughput. Its inference stack (in-house optimized, built on FireAttention) regularly benchmarks as the fastest provider for DeepSeek-R1 / V3, Llama 3/4, Qwen 3, and other large open-weight models.
Beyond serverless inference, Fireworks offers dedicated deployments (guaranteed capacity, custom models, SLA), fine-tuning as a managed service, and agent-oriented features like structured outputs, function calling, and JSON mode across every hosted model.
It's a common pick for AI products that need open-weight models in production with strict latency requirements — voice agents, real-time coding assistants, high-QPS chat applications.
commercial
saas
Fireworks AI
Fast-inference "neocloud" — GroqCloud's OpenAI-compatible API on Groq LPUs alongside NVIDIA accelerated computing.
verified
Groq runs GroqCloud, an inference cloud with an OpenAI-compatible API. Groq pioneered its own Language Processing Unit (LPU) silicon. It now runs LPUs alongside NVIDIA accelerated computing as an NVIDIA Cloud Partner, across 13 data centers in North America, Europe, the Middle East and Asia Pacific. It raised a $350M Series A in August 2026.
The API is OpenAI-compatible, so OpenAI-client code can switch with a base-URL change. Latency remains the main selling point, with hundreds to about a thousand tokens per second on hosted models. GroqCloud is common for voice agents, streaming chat and agent loops with many sequential calls.
commercial
saas
Groq LLC
Mistral's API platform and Studio for building, fine-tuning and deploying agents and apps on its models, including open-weight ones.
verified
Serverless cloud platform for Python with first-class GPU support — deploy LLMs, training jobs, and batch pipelines from code.
verified
Modal is a serverless compute platform designed for Python workloads that need GPUs: LLM inference, model training, batch image generation, data processing, and long-running AI pipelines. You write functions in Python with decorators that specify container image, GPU type, memory, and schedule — Modal handles the rest: cold starts in seconds, billing per-second, auto-scaling to zero when idle.
The developer experience is the differentiator. Locally-written Python
functions deploy to cloud GPUs without leaving the editor: modal run my_script.py executes on an A100 or H100, modal deploy exposes it as
an HTTPS endpoint. Popular use cases include hosting vLLM or SGLang for
inference, running Unsloth fine-tuning jobs, and dispatching ComfyUI
workflows at scale.
Modal is used heavily by AI startups that want cloud GPUs without managing Kubernetes, and by teams that need bursty capacity without reserving hardware.
commercial
saas
Modal
Ollama's hosted inference for large open models — the same Ollama app, CLI and API, plus OpenAI- and Anthropic-compatible endpoints, with a free tier and usage credits.
verified
Ollama Cloud runs open-weight models on Ollama's hosted
infrastructure, so you can use models too large for a laptop without
downloading them: DeepSeek V4, GLM 5.x, Kimi K3, MiniMax M3, gpt-oss,
Mistral Large 3, Nemotron 3 and others. The Ollama app and CLI switch
between local and cloud models by name (for example gemma4:cloud).
Coding agents can run against cloud models with ollama launch claude,
ollama launch codex or ollama launch opencode.
Apps can also call https://ollama.com/api directly with an API key, or use OpenAI- or Anthropic-compatible clients, with no local install. Ollama says cloud models use the provider's native weights, and prompt and response data are never logged or trained on. Hosting is primarily in the United States, with overflow to Europe and Singapore, through NVIDIA Cloud Partners under no-logging, zero-retention terms.
Plans: Free (starter usage credits and starter models), Pro ($20/month with $60 of credits), Max ($100/month with $300 of credits), Team ($500/month, early access) and Enterprise. Usage is billed per million tokens per model, with cheaper off-peak rates. Running models locally stays free and unlimited, and cloud features can be switched off for local-only use.
:cloud tags)ollama launch for Claude Code, Codex CLI and OpenCode on cloud modelsdisable_ollama_cloudfreemium
saas
Ollama
Run thousands of open-source ML models via simple API calls — image, video, audio, text — with per-second billing.
verified
Replicate is a cloud platform that hosts open-source ML models behind a unified REST API. Thousands of public models — Stable Diffusion, Flux, SDXL, LLaMA, Whisper, MusicGen, video generation, speech synthesis, CLIP embeddings, vision models, voice cloning — are immediately runnable without provisioning GPUs.
Its differentiator is breadth over depth: where Groq and Fireworks optimize for a curated set of LLMs, Replicate covers the whole open-source ML landscape, including image / video / audio models that inference-only LLM platforms don't touch. Pricing is per-second of GPU time.
Replicate also hosts custom model deploys (package your code + weights into a Cog container, push, get an endpoint) — useful for teams that want Replicate's serving infrastructure for their own fine-tuned or private models.
Replicate is part of Cloudflare (announced November 2025). The API and existing models continue unchanged, with integration into Cloudflare's Developer Platform planned.
commercial
saas
Replicate (Cloudflare)
GPU cloud platform with on-demand instances, serverless endpoints, and a community GPU marketplace — priced for AI workloads.
verified
Runpod is a GPU cloud provider focused on AI workloads. It offers three deployment modes: Pods (long-running GPU instances, similar to bare-metal rentals), Serverless (auto-scaling API endpoints that bill per-second of inference), and Secure Cloud vs Community Cloud (the latter being excess capacity from third-party hosts at lower prices, with the tradeoff of less guaranteed availability).
Its pricing is consistently among the lowest for consumer and enterprise GPUs (RTX 4090, A100, H100, H200, L4, L40S), and the Serverless tier makes it practical to host inference APIs without paying for idle GPU time. Pre-built templates cover common AI stacks (ComfyUI, Automatic1111, vLLM, SGLang, Ollama), so new instances can launch with a running application in one click.
Runpod is a common pick for indie AI builders, researchers, and teams running bursty inference where Modal / Replicate's managed layer isn't needed.
commercial
saas
Runpod
SpaceXAI (formerly xAI) Grok API for reasoning, code, voice, image and video models, usable with the OpenAI SDK.
verified
SpaceXAI, the company formerly branded xAI, builds the Grok family of models and sells access through a developer API at api.x.ai, with docs at docs.x.ai. The x.ai site and the API documentation now both carry the SpaceXAI name. Its newest model is Grok 4.7.
The API covers text generation with streaming, reasoning and structured outputs, and image understanding. Tools include function calling, code execution and collections search for RAG. The Imagine endpoints handle image generation and editing plus video generation, editing and extension. The Voice endpoints cover speech-to-speech (including SIP phone calls), text-to-speech, speech-to-text and custom voices.
Existing OpenAI-SDK code works too: the quickstart shows the official
OpenAI Python and JavaScript SDKs next to xAI's own SDK. For
high-volume jobs there are a Batch API and deferred completions. Every
documentation page is also available as Markdown by appending .md.
commercial
saas
SpaceXAI
Serverless inference for 200+ open-source models with OpenAI-compatible API — low latency, competitive pricing.
verified
Together AI is a serverless inference platform for open-source LLMs. It hosts 200+ models including the Llama, Qwen, DeepSeek, Mistral, Gemma, and Mixtral families, along with image models (FLUX, SD3) and embedding models — all behind a unified OpenAI-compatible API.
Together's positioning is "run open-source at production speed without operating GPUs." Its inference stack is tuned on top of vLLM and in-house optimizations, delivering competitive tokens-per-second and cost-per-million-tokens among the serverless providers (Fireworks, Groq, DeepInfra, Replicate).
Together also hosts fine-tuning jobs (upload dataset → get a
fine-tuned checkpoint → serve it on the same platform) and dedicated
endpoints for teams needing isolated capacity. Their open-source
together Python client and OpenAI compatibility make migration from
OpenAI trivial.
commercial
saas
Together AI
Anthropic's desktop app for Claude — chat, Cowork for long-running agentic tasks, and Claude Code in one app, with connectors (MCP), skills and plugins.
verified
Claude Desktop is Anthropic's native app for macOS and Windows, with a Linux beta for Ubuntu 22.04+ and Debian 12+. It combines three surfaces in one window. Chat is for conversations. Cowork is for longer agentic work that can keep running after you close your laptop; on Linux it runs tasks in a local QEMU/KVM virtual machine. Claude Code adds parallel coding sessions, visual diff review, and an integrated terminal and editor. In September 2026 Anthropic began merging chat and Cowork into a single Claude experience on Pro and Max plans.
Claude can be extended with connectors (MCP servers), skills and plugins, and it can use a built-in browser and, with permission, your computer. It runs on Anthropic's Claude model families (Mythos, Fable, Opus, Sonnet, Haiku). End-user documentation is indexed at claude.com/docs/llms.txt and the Help Center at support.claude.com/llms.txt.
commercial
local
Anthropic
Open-source (MIT) personal AI assistant that runs on your own machine and answers you in the chat apps you already use — one self-hosted Gateway, any model.
★ 391.4kMITv2026.9.8 · 2026-10-03verified
OpenClaw is an open-source personal AI assistant. You run one Gateway process on your own computer or server, connect it to the messaging apps you already use — WhatsApp, Telegram, Slack, Discord, Signal, iMessage, Microsoft Teams, Matrix and many more — and talk to your agent from anywhere. The agent acts for you: it can work through email and calendar, run scheduled automations, drive a browser, and use skills and plugins.
It is model-agnostic. Bundled provider plugins cover the major model APIs, and custom providers point it at local inference servers. Skills and plugins are shared through ClawHub. Desktop apps for macOS, Windows and Linux install the Gateway together with a chat and configuration Control UI, and iOS and Android companion apps pair as device nodes.
OpenClaw is MIT-licensed and, since July 2026, stewarded by the OpenClaw Foundation, an independent 501(c)(3) non-profit. There is no paid or hosted tier, and default telemetry is limited to a version check you can turn off. Its creator, Peter Steinberger, joined OpenAI in 2026 and continues to lead the project.
open-source
local
OpenClaw Foundation
Open-source multi-agent desktop app — a Commander plans a goal and dispatches specialist agents; bring your own model keys, files stay on your disk.
★ 2.2kMITv2026.10.1 · 2026-10-01verified
AI assistant built into the Raycast launcher (Mac, Windows, iOS) — agents with projects, screen awareness, scheduled automations and bring-your-own model keys.
verified
Agentic news and RSS reader with cited AI summaries and chat, text-to-speech, a searchable library, and API and MCP access.
verified
128 entries — 70 with full descriptions, 58 stubs. 96 currently publish a working llms.txt.
Humans increasingly use AI agents (Claude, ChatGPT, Perplexity, local assistants) as their primary search and discovery layer. When a developer asks an agent "what's the best offline voice-to-text tool for 2026?", the answer depends on what the agent can find, read, and cite. Tools that publish a well-structured llms.txt (spec by Jeremy Howard) are easier for agents to index, summarize, and recommend.
Most existing AI-tool directories are optimized for Google SEO (JavaScript-rendered, paywalled, affiliate-heavy). This one is optimized for agent retrieval: plain Markdown, structured YAML entries, a canonical llms.txt and llms-full.txt at the repo root, MIT license, no tracking.
llms.txt URLs for agents to ingest as seed data. If you're looking for a raw feed of every llms.txt on the public internet to wire into RAG, look there. This repo takes the complementary angle: curated tools, categorized and described, aimed at humans picking what to use and agents answering "what should I recommend to my user for X?".The llms.txt status is re-derived daily from what each URL actually serves — including a check that the response is a real file and not a docs site's catch-all HTML page. See CONTRIBUTING.md.
This list catalogs tools, not models. But agents answering user questions often need both in the same breath, so here are pointers to the most notable model orgs by domain. We deliberately don't link individual model versions — those go stale weekly. Search huggingface.co/<family-name> for each org's current lineup.
Frontier open weights (April 2026):
Established open weights (large install base, not setting the pace):
Closed frontier (API-only — listed for agent awareness):
ASR (speech-to-text): Qwen3-ASR (Alibaba, currently best) · Whisper (OpenAI, widely deployed baseline)
TTS (text-to-speech): Qwen3-TTS (Alibaba, currently best) · F5-TTS · Kokoro · Piper · OpenVoice · Bark (Suno)
Frontier (April 2026): Qwen Image (Alibaba) · Z-Image Turbo
Usable locally: Flux.2 Klein 4B (Black Forest Labs) · HiDream
⚠️ License note: Most Flux.2 variants ship under terms that block commercial use. Only Flux.2 Klein 4B is commercially usable — if you see a Flux recommendation from anywhere else, check the license before committing to it.
Legacy (widely used, not advancing): Stable Diffusion · SDXL · Flux.1
Ranked by practical 2026 usefulness: WAN (Alibaba, best consistency) · LTX (Lightricks, best character expression, trade-off is chunk-boundary drift) · HunyuanVideo (Tencent, solid third option)
Hunyuan3D (Tencent) leads the category in both local weights and hosted API. Both Hunyuan3D and Tripo ship hosted APIs that are meaningfully better than their publicly released local weights: Hunyuan3D 3.1 is the current API version while local users get 2.1, and Tripo's API also substantially outperforms the local weights. If you need best-quality 3D generation, use Hunyuan3D's API. For local-only work, Hunyuan3D 2.1 local weights are still the top pick; Tripo local is weaker.
BGE (BAAI) · Nomic Embed · Jina · E5 (Microsoft)
See CONTRIBUTING.md for the YAML schema and submission process. One entry per PR, please. Got a question? Check FAQ.md first.
Common questions — why this list exists, who curates it, what stops spam, why some entries are marked missing — are answered in FAQ.md.
MIT. Fork it, scrape it, mirror it, agents welcome.
Local speech-to-text that learns your voice. Perpetual licence. Our flagship.
PAID · flagship
Memory for your AI agents — records, rules, the full history and the graph of decisions, already there when a session starts. It answers before you ask.
PAID · free tier
Print-ready digital models. GLB and OBJ included. Lifetime access.
PAID · digital catalog
Five printers and real capacity. No self-serve shop yet — tell us what you need and we quote it by email.
BY ARRANGEMENT · email us
Our YouTube channel. A cyber-tiger host walks through local AI tools and what they actually do.
CHANNEL · live
Curated GitHub lists for AI coding agents, MCP servers, local AI and Linux for AI. Every entry carries a source link.
FREE · curated
Long-form how-tos for local AI on Linux, Windows and macOS, with the configuration files included.
FREE · live
Negative-curation: practices and tools that waste your time, ranked. Receipts required.
FREE · live
Who we are, why we build privacy-first AI, and what we won't do.