Awesome list · free and open on GitHub

llms.txt

A curated list of AI tools, platforms, and services that publish llms.txt — making them discoverable to AI agents doing research on behalf of human users.

128 entries

20 categories

96 publish a working llms.txt

Every link checked 2026-10-05

Inference Runtimes 13

ExLlamaV3New

Inference library for local LLMs on consumer GPUs with EXL3 quantization and tensor/expert parallelism. It succeeds ExLlamaV2.

no llms.txt yetstub

★ 1.6kMITv1.5.4 · 2026-10-03verified

Details
License

open-source

Deployment

library

HuggingFace Transformers

Hugging Face's model-definition framework for text, vision, audio, video and multimodal models — PyTorch-based inference and training for 1M+ Hub checkpoints.

llms.txttransformerspytorchhuggingfaceinferencetraininglibrary

★ 167kApache-2.0v5.18.0 · 2026-09-30verified

Details
About

Transformers is HuggingFace's foundational Python library for working with transformer-based models. It provides unified from_pretrained() APIs for loading thousands of models from the HuggingFace Hub — LLMs, vision models, audio models, multimodal models — and running them for inference, fine-tuning, and research.

While Transformers isn't an inference server in the vLLM / Ollama sense (it's a library, not a server), it's the entry point most researchers and engineers use when working with a new model before it has dedicated runtime support elsewhere. It's the canonical reference implementation that every inference server (vLLM, llama.cpp, MLC, TGI) compares against for correctness.

The library is deeply integrated with the rest of the HuggingFace stack: Datasets, Accelerate, PEFT, TRL, and the Hub itself — together forming the dominant open-source ML ecosystem.

Features
  • Unified API for 1M+ model checkpoints on the Hugging Face Hub
  • PyTorch-based (v5 dropped TensorFlow and JAX); Python 3.10+, PyTorch 2.5+
  • Pipeline API, generate API and Trainer
  • Tokenizers, processors and feature extractors included
  • Model definitions reused by vLLM, SGLang, llama.cpp, MLX, Axolotl, Unsloth and others
  • Quantization via bitsandbytes, GPTQ, AWQ and more
  • Frequent releases with day-zero support for new architectures
Best for
  • Research and experimentation with new model architectures
  • Fine-tuning LLMs and vision models on custom data
  • Building inference pipelines before moving to a dedicated runtime
  • Teaching / learning transformer internals
  • Any ML workflow that starts with "download this model from the Hub"
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

HuggingFace

Jan

Open-source (Apache-2.0) desktop ChatGPT alternative that runs local LLMs offline, with optional connections to cloud models.

no llms.txt yetinferencechat-uijanlocalprivacyllama-cpp

★ 44.8kv0.8.4 · 2026-07-23verified

Details
About

Jan is a desktop application that provides a ChatGPT-like interface backed entirely by locally-running open-weight LLMs. It bundles a llama.cpp-based inference engine, model discovery (curated HuggingFace models with one-click install), a conversation interface, and an OpenAI-compatible local server — all in a single installable app for Linux, macOS, and Windows.

Jan's positioning is privacy-first: local models run fully offline with no account. Cloud models (OpenAI, Anthropic, Gemini and others) can be added with your own keys when you want them. Its codebase is Apache-2.0 open-source and actively developed by Menlo Research, with frequent model library updates tracking the open-weight state of the art.

The OpenAI-compatible server means Jan doubles as a local backend for other tools — coding assistants, agent frameworks, or any application that speaks the OpenAI API can point at http://localhost:1337 and use whatever model Jan has loaded.

Features
  • ChatGPT-style desktop UI with conversation history
  • llama.cpp backend with GGUF model support
  • Curated model library with one-click download
  • Built-in OpenAI-compatible local server
  • NVIDIA CUDA, AMD ROCm, Apple Metal, CPU backends
  • Extensions system for custom integrations
  • Fully offline with local models; optional cloud model providers
Best for
  • Privacy-conscious users wanting ChatGPT UX without cloud
  • Non-technical users running local LLMs without CLI
  • Teams piloting local LLMs on consumer hardware
  • Local backend for coding assistants and agent frameworks
  • Offline work environments
License

open-source

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

Menlo Research

KoboldCpp

Single-binary llama.cpp wrapper with KoboldAI-style UI for chat, story-writing, and RP.

no llms.txt yetstub

★ 11.9kAGPL-3.0v1.122.1 · 2026-09-26verified

Details
License

open-source

Deployment

local

llama.cpp

Reference C++ implementation for running LLaMA-family and other transformer models with GGUF quantization.

no llms.txt yetinferenceggufcppquantizationllamavulkancudarocmmetal

★ 130.4kMITv0.5.0 · 2026-09-23verified

Details
About

llama.cpp is the foundational open-source C++ inference engine that made local LLM execution practical on consumer hardware. It defined the GGUF model format now used across the ecosystem, pioneered aggressive quantization techniques (Q4_K_M, Q5_K_S, Q8_0, and dozens of variants), and supports more hardware backends than any other runtime: CPU (AVX, AVX2, AVX-512, NEON), Metal, CUDA, HIP/ROCm, MUSA, Vulkan, SYCL, CANN, OpenCL, WebGPU and Snapdragon Hexagon.

It ships as a library, a CLI (llama-cli), a server (llama-server) with an OpenAI-compatible API, and bindings for nearly every programming language. Projects like Ollama, LM Studio, KoboldCpp, text-generation-webui, Jan, and LocalAI are all built on or around llama.cpp. The official Llama app at llama.app, from the llama.cpp team and Hugging Face, wraps it as a one-click Mac/Windows menu-bar app with a local OpenAI-compatible API, and also installs the command-line version on Linux.

For developers who want the lowest-level local inference or need to target unusual hardware (embedded, edge, non-CUDA GPUs), llama.cpp is the direct answer.

Features
  • GGUF format (the de facto standard for quantized open-weight models)
  • 20+ quantization variants for quality/size tradeoffs
  • Backends for CPU, Metal, CUDA, HIP/ROCm, MUSA, Vulkan, SYCL, CANN, OpenCL, WebGPU, Hexagon
  • Built-in OpenAI-compatible HTTP server
  • Speculative decoding, prompt caching, continuous batching
  • Grammar-constrained generation (JSON Schema, custom grammars)
  • Language bindings for Python, Node.js, Go, Rust, Java, C#, and more
Best for
  • Embedding local LLM inference in applications (the library path)
  • Running LLMs on hardware without CUDA (Vulkan, ROCm, Metal)
  • Resource-constrained deployments (edge, embedded, low-VRAM)
  • Building higher-level tools on top of a battle-tested inference core
  • Quantizing new models and publishing them in GGUF format
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

ggml-org

LM Studio

Desktop app for running local LLMs (llama.cpp and MLX) with an OpenAI-compatible server — now with Bionic, an agent for work and code, and optional US-hosted open-model inference.

llms.txtinferencellmlocaldesktopguillama-cppmlx

verified

Details
About

LM Studio is a desktop application for running open-weight LLMs on your own machine. It handles model discovery and download from Hugging Face, per-model GPU and context configuration, and chat. Its OpenAI-compatible local server and lms CLI let coding agents and other apps use whatever model is loaded. The runtime uses llama.cpp, plus Apple's MLX on Apple Silicon.

In 2026 the app became LM Studio Bionic, built around Bionic, "LM Studio's agent for open models". Bionic creates and edits documents, helps with code, runs automations and can control the computer, and it includes local real-time voice transcription. Local use stays free. Paid plans (Bionic+ at $20/month and Pro at $100/month) add US-hosted, zero-data-retention inference for large open models such as Kimi K3, GLM 5.3 and DeepSeek V4, plus web search. LM Link connects multiple devices.

Features
  • Desktop GUI for Hugging Face model browsing, download and configuration
  • Bionic agent for documents, coding, automations and computer control
  • OpenAI-compatible local server and lms CLI; lmstudio-js and lmstudio-python SDKs
  • llama.cpp and MLX backends
  • Local real-time voice transcription
  • Optional US-hosted, zero-data-retention cloud inference for large open models (paid plans)
  • LM Link to share models across devices
Best for
  • Non-CLI users who want to run local LLMs with a GUI
  • Evaluating models side-by-side before committing to one
  • Private chat assistants for knowledge work
  • Local backend for coding assistants (Continue, Cursor via OpenAI proxy)
  • Teams piloting local LLMs before deploying server-side
License

freemium

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

Element Labs, Inc. (LM Studio)

LocalAI

Open-source (MIT) self-hosted AI engine — OpenAI-, Anthropic-, Ollama- and ElevenLabs-compatible APIs for text, voice, vision, image, video and agents, no GPU required.

no llms.txt yetinferenceopenai-compatibleself-hostedmulti-modalwhisperstable-diffusion

★ 49.4kMITv4.11.0 · 2026-10-02verified

Details
About

LocalAI is a self-hosted, OpenAI-compatible inference server designed as a drop-in replacement for the OpenAI API. It supports text generation, chat completions, embeddings, transcription (Whisper), text-to-speech, image generation (Stable Diffusion), and function calling — all via the same HTTP endpoints applications already use with OpenAI.

It ships as a single binary or Docker container and runs on consumer CPUs (no GPU required), NVIDIA CUDA, AMD ROCm, Vulkan, and Intel GPUs. Under the hood it dispatches to multiple backends: llama.cpp, whisper.cpp, Stable Diffusion, Bark, and others — unified behind one API.

LocalAI is a common pick for organizations migrating existing OpenAI- dependent applications to on-premise inference without rewriting them.

Features
  • Drop-in APIs — OpenAI, Anthropic, Ollama and ElevenLabs compatible; realtime over WebRTC
  • Per-model swappable engines (llama.cpp, vLLM, SGLang, MLX, diffusers, whisper, parakeet, …)
  • Text, embeddings, transcription and diarization, TTS and voice cloning, vision, image/video/music generation
  • Agents, MCP and skills built in
  • Model gallery with one-click installs (1,000+ models)
  • CPU, CUDA, ROCm, Vulkan, Intel and Apple backends
  • Single binary or Docker / Kubernetes deployment
Best for
  • Drop-in replacement for OpenAI API in existing applications
  • Self-hosted AI infrastructure for organizations with data constraints
  • Unified inference server covering text, image, and audio models
  • Edge and on-prem deployments without cloud dependencies
  • CPU-only environments needing local inference
License

open-source

Deployment

self-hosted

Platforms
  • linux
  • macos
  • windows
Maintainer

LocalAI

MLC LLM

Universal LLM deployment via compiled kernels — runs on iOS, Android, WebGPU, Vulkan, CUDA.

no llms.txt yetstub

verified

Details
License

open-source

Deployment

library

Ollama

Run open models locally via a single-binary server with a built-in model library — with optional Ollama Cloud for models too big for your hardware.

llms.txtinferencellmlocalquantizationollamacudarocmmetal

★ 182.2kMITv0.35.1 · 2026-09-29verified

Details
About

Ollama is the most widely adopted tool for running open-weight LLMs on personal hardware. It bundles model download, quantization, a server exposing an OpenAI-compatible API, and a CLI into a single binary with zero-config defaults. The built-in model library covers Llama, Qwen, Gemma, Mistral, DeepSeek, and dozens of others with preconfigured quantizations.

Its appeal is friction: ollama run gemma4 downloads the model and drops you into a chat in one command. The OpenAI-compatible API makes Ollama a drop-in backend for thousands of tools originally built against OpenAI — switching from cloud to local becomes a URL change.

Ollama runs on CPU, NVIDIA CUDA, AMD ROCm, and Apple Silicon, and serves as the default "local inference" option for open-source agent frameworks, chat UIs, and coding assistants.

The same app, CLI and API can also run models on Ollama's hosted cloud (see the Ollama Cloud entry). Cloud features can be disabled for local-only use.

Features
  • Single-binary install with built-in model library
  • OpenAI-compatible HTTP API (/v1/chat/completions)
  • Automatic GPU detection and offload (CUDA, ROCm, Metal)
  • Quantization support (Q4, Q5, Q8, FP16)
  • Modelfile system for customizing prompts and parameters
  • Multimodal support for vision models
  • macOS, Linux, Windows native builds
  • OpenAI- and Anthropic-compatible APIs
  • ollama launch for Claude Code, Codex and OpenCode
Best for
  • Local development against open-weight LLMs without cloud costs
  • Air-gapped or privacy-constrained deployments
  • Dev teams that need shared internal inference with consistent config
  • Backend for local coding assistants and agent frameworks
  • Experimenting with new open-weight models as they release
License

open-source

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

Ollama

Open WebUI

Self-hosted, feature-rich chat interface for local and cloud LLMs — the "ChatGPT clone" of the open-source world.

llms.txtchat-uiself-hostedopen-webuiollamamulti-userprivacy

★ 154kv0.11.4 · 2026-09-21verified

Details
About

Open WebUI is a self-hosted web interface for chatting with LLMs, supporting Ollama, any OpenAI-compatible endpoint, and cloud providers. It's visually and functionally comparable to ChatGPT's UI: conversation history, folder organization, model picker, custom system prompts, multimodal input (images, documents, audio), code interpreter, web search, and a plugin/tool system.

It's deployed as a single Docker container and commonly paired with Ollama to give non-technical users a friendly front-end to local LLMs. Multi-user support with role-based access makes it suitable for teams or family setups on a shared home server.

The project has strong momentum among self-hosters and privacy-conscious users who want the ChatGPT experience without sending data to OpenAI.

Features
  • ChatGPT-like web UI with conversation history
  • Multi-user with role-based access control (admin, user, pending)
  • Works with Ollama, any OpenAI-compatible API, or cloud providers
  • {'Multimodal input': 'images, documents, audio'}
  • RAG with local document uploads and web search
  • Custom models (system prompts) via the UI
  • Plugin/tool system (Functions, Pipelines)
  • Docker Compose deployment
Best for
  • Self-hosted ChatGPT replacement for home or team use
  • Privacy-conscious individuals running local LLMs daily
  • Non-technical users accessing local inference via a web UI
  • Multi-user deployments where each person gets isolated history
  • Research environments comparing many models with the same UI
License

source-available

Deployment

self-hosted

Platforms
  • linux
Maintainer

Open WebUI

SGLang

Fast LLM and VLM serving runtime with RadixAttention cache and structured output support.

llms.txtstub

verified

Details
License

open-source

Deployment

self-hosted

TextGen

Open-source desktop app for local LLMs (formerly Text Generation WebUI) with chat, vision, tool-calling, UI and API.

no llms.txt yetstub

★ 47.7kAGPL-3.0v4.9 · 2026-05-20verified

Details
License

open-source

Deployment

local

vLLM

High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.

llms.txtinferenceservingpaged-attentionvllmproductioncudarocm

★ 93.2kApache-2.0v0.31.0 · 2026-10-05verified

Details
About

vLLM is the leading open-source inference engine for serving LLMs in production. It introduced PagedAttention (a memory management technique adapted from virtual memory in operating systems) that dramatically reduces GPU memory fragmentation when serving many concurrent requests, along with continuous batching that maximizes GPU utilization by folding new requests into in-flight batches.

vLLM exposes an OpenAI-compatible HTTP API, supports virtually every open-weight architecture (Llama, Qwen, DeepSeek, Mistral, Gemma, Phi, MPT, Falcon, Yi, Cohere, and dozens more), handles both dense and MoE models, and scales to multi-GPU and multi-node deployments with tensor and pipeline parallelism.

Where Ollama targets single-user local use, vLLM targets server-side production: high request concurrency, low latency, and maximum throughput per dollar of GPU.

Features
  • PagedAttention for efficient KV-cache memory management
  • Continuous batching for high request concurrency
  • OpenAI-compatible API with streaming and tool calling
  • Tensor and pipeline parallelism for multi-GPU serving
  • FP16, BF16, FP8, INT8, AWQ, GPTQ quantization support
  • Speculative decoding and prefix caching
  • Extensive model support (Llama, Qwen, DeepSeek, Mistral, Gemma, …)
Best for
  • Serving open-weight LLMs in production behind a load balancer
  • High-QPS inference for internal developer tools or customer products
  • Multi-GPU deployments of large models (70B+, 405B)
  • Cost-optimized inference infrastructure replacing closed APIs
  • Research throughput benchmarking
License

open-source

Deployment

self-hosted

Platforms
  • linux
Maintainer

vLLM Project

LLM Gateways 3

Cloudflare AI GatewayNew

Managed gateway in front of OpenAI, Anthropic, Bedrock and dozens of providers — analytics, caching, rate limiting and model fallback via one OpenAI-compatible endpoint.

llms.txtgatewayproxycachingrate-limitingobservabilitycloudflare

verified

Details
About

Cloudflare AI Gateway is a managed proxy that sits between an application and its AI providers. Requests go through a Cloudflare endpoint, either a unified OpenAI-compatible API or each provider's native API, and the gateway adds logging, analytics, response caching, rate limiting, retries and model fallback, without running any infrastructure.

It supports a wide catalogue of providers (OpenAI, Anthropic, Azure OpenAI, Amazon Bedrock, Google, Groq, Cerebras, Baseten, Cartesia and more) and Cloudflare's own Workers AI. Because it runs on Cloudflare's edge network, it is a low-effort way to get centralized observability and cost control for LLM traffic. The trade-off is that it is a hosted service, not something you self-host.

Features
  • Unified OpenAI-compatible endpoint plus provider-native and WebSocket endpoints
  • Request logging and analytics (tokens, cost, latency, errors)
  • Response caching and rate limiting
  • Retries and model fallback across providers
  • Supports many providers and Cloudflare Workers AI
Best for
  • Adding observability and cost control to existing OpenAI/Anthropic calls
  • Failing over between providers without code changes
  • Caching repeated prompts to cut spend and latency
License

freemium

Deployment

saas

Platforms
  • web
Maintainer

Cloudflare

LiteLLM

Unified OpenAI-compatible proxy and SDK that routes calls across 100+ LLM providers with load balancing, fallbacks, and cost tracking.

llms.txtproxygatewayllmopenai-compatibleload-balancingmulti-provider

★ 60.1kv1.104.0 · 2026-10-03verified

Details
About

LiteLLM solves the "every LLM provider has a slightly different API" problem by exposing a single OpenAI-compatible interface that internally translates to 100+ providers: OpenAI, Anthropic, Google, Azure, AWS Bedrock, Cohere, Mistral, OpenRouter, Together, Fireworks, Groq, xAI, Ollama, vLLM, and many more.

It ships in two forms: a Python SDK for embedding in applications, and a standalone proxy server (LiteLLM Proxy) that accepts OpenAI-format requests and routes them across providers with load balancing, automatic retries, fallback chains, cost tracking, rate limiting, and virtual API key management for multi-team / multi-user deployments.

LiteLLM is often the backbone piece in organizations that want provider flexibility — swapping models without touching application code, or automatically failing over from a cloud provider to a local one.

Features
  • Unified OpenAI-compatible API across 100+ providers
  • Proxy server with load balancing and automatic fallbacks
  • Cost tracking per user / team / virtual API key
  • Rate limiting and budget enforcement
  • Caching (Redis, in-memory) for identical requests
  • Embedding, chat completion, and image generation support
  • Integration with LangFuse, LangSmith, Helicone for tracing
Best for
  • Multi-provider LLM applications with automatic fallback
  • Centralized LLM gateway for organizations with multiple dev teams
  • Cost attribution and budget control across many consumers
  • Migrating between LLM providers without code changes
  • Hybrid deployments routing between cloud and local inference
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

BerriAI

OpenRouterNew

One API for hundreds of models from many providers, with routing and fallbacks and no subscription.

llms.txtstub

verified

Details
License

commercial

Deployment

saas

Agent Frameworks 12

Agent Development Kit (ADK)New

Google's open-source, code-first toolkit for building, evaluating and deploying multi-agent systems in Python, TypeScript, Go, Java and Kotlin.

llms.txtagent-frameworkgooglegeminimulti-agentpythontypescriptgojava

★ 21.7kApache-2.0v2.11.0 · 2026-10-02verified

Details
About

Agent Development Kit (ADK) is Google's open-source (Apache-2.0) framework for building, evaluating and deploying AI agents. Agents can be LLM-driven (LlmAgent) or deterministic workflow agents (SequentialAgent, ParallelAgent, LoopAgent), and they compose hierarchically into multi-agent systems that delegate through LLM-driven transfer or by calling other agents as tools.

ADK covers tools (function tools, agents as tools, built-in code execution and search), callbacks, sessions and state, long-term memory, artifacts, planning, bidirectional streaming for live and voice agents, built-in multi-turn evaluation, a CLI and a developer UI. It is optimised for Gemini but model-agnostic through its BaseLlm interface. It deploys anywhere, with paths to Google Cloud.

Features
  • LLM agents plus Sequential, Parallel and Loop workflow agents
  • Hierarchical multi-agent systems with delegation and AgentTool
  • Tools, callbacks, sessions/state, memory and artifacts
  • Live and voice agents with bidirectional streaming
  • Built-in agent evaluation, CLI and developer UI
  • SDKs for Python, TypeScript, Go, Java and Kotlin
  • Gemini-optimised, other models via BaseLlm
Best for
  • Multi-agent applications with a mix of deterministic and LLM-driven steps
  • Voice and real-time agents on Gemini Live
  • Teams deploying agents to Google Cloud
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

Google

Agno

Open-source framework and runtime for agent platforms — build with the Agno SDK, run on the AgentOS runtime, manage from the AgentOS control plane.

llms.txtstub

★ 42.6kApache-2.0v3.1.1 · 2026-10-02verified

Details
License

open-source

Deployment

library

CrewAI

Python framework for orchestrating role-based multi-agent systems with sequential and hierarchical workflows.

llms.txtmulti-agentcrewaipythonrole-basedorchestration

★ 59.4kMIT1.15.23 · 2026-09-28verified

Details
About

CrewAI is an agent framework centered on the metaphor of a "crew" — multiple specialized agents (roles) collaborating on a task (a crew). Developers define agents with a role, goal, and backstory, then compose them into tasks with explicit dependencies. CrewAI handles the delegation, inter-agent messaging, and result aggregation.

Compared to LangGraph's explicit-graph approach, CrewAI is higher-level and more opinionated: it targets developers who want a ready-made pattern for "researcher + writer + reviewer" style workflows without hand-rolling state machines. Its Flows feature adds event-driven workflows and deterministic control flow for production deployments.

The framework has significant adoption for content-generation pipelines, research assistants, and business-process automation where the task naturally decomposes into specialized agent roles.

Features
  • Role-based agents with goals, backstories, and tools
  • Sequential and hierarchical process modes
  • Flows API for event-driven and conditional workflows
  • Tool integration (web search, code execution, custom tools)
  • {'Memory': 'short-term, long-term, entity-level'}
  • Provider-agnostic — native provider integrations, LiteLLM optional
  • CrewAI AMP platform for building, deploying and tracing crews and flows
Best for
  • Multi-step content generation (research → outline → draft → edit)
  • Business workflow automation with role separation
  • Research assistants that decompose questions across specialists
  • Customer-facing agents with internal supervisor oversight
  • Rapid prototyping of multi-agent patterns
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

CrewAI

DSPy

Framework for programming rather than prompting LLMs — composable modules with optimizers.

llms.txtstub

verified

Details
License

open-source

Deployment

library

LangChain

Open-source Python/TypeScript framework for building LLM agents — create_agent harness with middleware, built on LangGraph, with hundreds of integrations.

llms.txtframeworkragagentslangchainpythontypescriptllm

★ 147.5kMITlangchain-core==1.6.6 · 2026-09-29verified

Details
About

LangChain is an open-source framework for building agents and LLM applications in Python and TypeScript. In the 1.x line its centre is create_agent, a minimal, configurable agent harness composed from a model, tools, a system prompt and middleware. Middleware covers things like guardrails, retries, routing and tool policies. The framework keeps hundreds of integrations with model providers, vector stores and tools.

LangChain agents are built on top of LangGraph, the lower-level orchestration runtime, so they inherit durable execution, persistence and human-in-the-loop support. Deep Agents, built on LangChain agents, is the batteries-included option with context compression, a virtual filesystem and subagents. LangSmith is the company's commercial platform for tracing, evaluation and deployment.

All LangChain, LangGraph and LangSmith docs live at docs.langchain.com, which publishes both llms.txt and a single llms-full.txt.

Features
  • create_agent harness with composable middleware
  • Hundreds of integrations (OpenAI, Anthropic, Google, local models, vector DBs, tools)
  • Built on LangGraph (durable execution, persistence, human-in-the-loop)
  • Deep Agents for batteries-included long-running agents
  • RAG and structured-output primitives
  • LangSmith tracing and evaluation hooks
  • Python and TypeScript implementations
Best for
  • RAG applications over internal document collections
  • Chatbots with multi-turn memory and tool use
  • Structured data extraction from unstructured text
  • Orchestrating multiple LLM calls with parsing and validation between
  • Rapid prototyping of LLM-powered features
  • Teaching / learning LLM application patterns
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

LangChain

LangGraph

Graph-based library for building stateful multi-agent workflows with explicit control flow and durability.

llms.txtagent-frameworkgraphstate-machinelanggraphdurablehitl

★ 42.7kMITcli==0.4.32.dev0 · 2026-09-23verified

Details
About

LangGraph is LangChain's successor-style library for building agents and multi-step LLM workflows as explicit directed graphs. Where classic LangChain agents abstract control flow behind a reactive loop, LangGraph makes nodes, edges, state, and branching first-class — letting developers write agents that look more like state machines than magic boxes.

Its standout features are durability (graphs can pause, persist state, and resume after arbitrary interruptions — critical for long-running agent workflows) and first-class support for human-in-the-loop patterns (nodes that suspend execution until a human approves or edits the next step). LangSmith Deployment (formerly LangGraph Platform) runs graphs as long-running production services with built-in checkpointing.

Features
  • Nodes and edges define explicit agent control flow
  • Durable state with pluggable checkpointers (SQLite, Postgres, Redis)
  • Pause/resume for human-in-the-loop and async workflows
  • Streaming of intermediate state to the client
  • Multi-agent supervisor and swarm patterns built-in
  • Integrates with any LLM provider (via LangChain's model abstractions)
  • LangSmith Deployment (formerly LangGraph Platform) for production hosting
Best for
  • Complex agent workflows with explicit branching and retries
  • Long-running agent tasks that need to survive process restarts
  • Human-in-the-loop review gates in agent pipelines
  • Multi-agent systems (supervisor, swarm, hierarchical)
  • Replacing ad-hoc LangChain AgentExecutor setups
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

LangChain

Magentic

Type-safe Python library for building LLM-powered functions with structured outputs.

no llms.txt yetstub

verified

Details
License

open-source

Deployment

library

MastraNew

TypeScript framework for AI agents and apps with memory, tools, MCP and observability built in.

llms.txtstub

★ 28.6k@mastra/[email protected] · 2026-10-05verified

Details
License

open-source

Deployment

library

Microsoft Agent FrameworkNew

Microsoft's open-source framework for production AI agents and multi-agent workflows in Python, .NET and Go, and the successor to AutoGen.

no llms.txt yetmulti-agentmicrosoftautogen-successorworkflowspythondotnetopentelemetry

★ 13.9kMITpython-1.20.0 · 2026-10-02verified

Details
About

Microsoft Agent Framework (MAF) is Microsoft's open-source, MIT-licensed framework for building, orchestrating and running AI agents and multi-agent workflows. It replaces AutoGen. The AutoGen repository is now in maintenance mode and names MAF as its enterprise-ready successor.

MAF has consistent APIs in Python and C#/.NET, plus a Go SDK in a separate repository. Workflows are graphs and support sequential, concurrent, handoff and group-collaboration patterns, with checkpointing, streaming and human-in-the-loop control. Agents can be defined declaratively in YAML, extended with middleware, and given skills built from files, inline code or class libraries.

Tracing, monitoring and debugging go through OpenTelemetry. DevUI is an interactive UI for developing and testing agents. Agents can be deployed to Foundry-hosted infrastructure, and the framework supports multiple LLM providers.

Features
  • Python and C#/.NET implementations with consistent APIs; Go SDK available
  • Graph-based workflows with sequential, concurrent, handoff and group patterns
  • Checkpointing, streaming and human-in-the-loop control
  • Declarative agents defined in YAML
  • Middleware pipeline for requests, responses and exceptions
  • Built-in OpenTelemetry tracing and monitoring
  • DevUI for interactive agent development and debugging
  • Multiple LLM provider support
Best for
  • Moving an AutoGen project onto its maintained successor
  • Production multi-agent systems that need durability and restartability
  • .NET shops that want first-class agent tooling
  • Teams that need governance and observability around agent workflows
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

Microsoft

Pydantic AI

Agent framework built on Pydantic with type-safe tool use and structured responses.

llms.txtstub

verified

Details
License

open-source

Deployment

library

smolagents

Minimal agent library from HuggingFace centered on code-writing agents.

llms.txtstub

verified

Details
License

open-source

Deployment

library

Vercel AI SDK

TypeScript toolkit for building AI apps with unified APIs across providers and framework helpers.

llms.txtstub

verified

Details
License

open-source

Deployment

library

Agent SDKs 3

Anthropic SDK

Official Anthropic client SDKs for the Claude API in Python, TypeScript, C#, Go, Java, PHP and Ruby, plus the ant CLI.

llms.txtapianthropicclaudesdkpythontypescriptllm

★ 4kMITv1.11.0 · 2026-09-30verified

Details
About

The Anthropic SDK is the lower-level, direct interface to the Claude API. Where the Claude Agent SDK provides a full agent loop with tool dispatch and subagents, the Anthropic SDK gives developers raw access to the messages API — single turns, streaming, tool calling, prompt caching, vision, and the full range of Claude models.

It's the right choice for applications that want Claude's raw capabilities without the agent-loop abstraction: chatbots, content generation, one-shot classification, structured extraction, and any integration where the developer wants to own the control flow.

Features
  • Messages API with streaming, tool use and vision support
  • Prompt caching (cache reads at a fraction of the base input price)
  • Extended thinking for reasoning models
  • Batch API for high-volume asynchronous workloads (50% off)
  • Claude Mythos, Fable, Opus, Sonnet and Haiku model families
  • SDKs for Python, TypeScript, C#, Go, Java, PHP and Ruby; ant CLI for shell use
  • Built-in error handling and retries
Best for
  • Production chatbots and conversational interfaces
  • Structured data extraction and classification at scale
  • Content generation pipelines (summarization, translation, copy)
  • Document and image analysis via the vision API
  • Tool-calling integrations where the developer owns the orchestration
  • High-throughput batch processing via the Batch API
License

commercial

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

Anthropic

Claude Agent SDK

Anthropic's official SDK for building custom agents on top of Claude with tool use, subagents, and hooks.

llms.txtagent-sdkanthropicclaudepythontypescriptmcptool-use

★ 8.2kMITv0.2.163 · 2026-09-30verified

Details
About

The Claude Agent SDK is Anthropic's library for building AI agents powered by Claude. It exposes the same primitives Claude Code is built on — tool use, subagents, hooks, background tasks, session persistence, MCP servers — as a Python/TypeScript library developers can embed in their own products.

The SDK handles the agent loop, tool dispatch, error recovery, and context management. Developers define tools (functions the agent can call), subagent types (specialized agents for parallel subtasks), and hooks (policies that shape agent behavior) and the SDK orchestrates the rest.

It's the right choice when a project needs agent capabilities embedded in its own UI or CLI rather than running Claude Code in a terminal.

Features
  • Tool-use loop with automatic retries and error recovery
  • Subagent system for parallel specialized workers
  • Hook system for pre/post-tool-call behavior shaping
  • MCP (Model Context Protocol) server integration
  • Background task support with async completion notifications
  • Session persistence and context window management
  • Python and TypeScript implementations
Best for
  • Embedding an AI assistant in an existing product's CLI or web UI
  • Building specialized agents for narrow domains (legal, medical, code review)
  • Orchestrating multi-agent workflows within a single application
  • Adding Claude-powered automation to internal developer tools
  • Replacing ad-hoc Claude API loops with a structured agent framework
License

commercial

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

Anthropic

OpenAI Agents SDKNew

OpenAI's provider-agnostic multi-agent SDK with handoffs, guardrails, sessions, tracing and voice agents. It replaces Swarm.

llms.txtstub

★ 29.8kMITv0.23.1 · 2026-10-02verified

Details
License

open-source

Deployment

library

Coding Agents 13

Aider

AI pair programming in your terminal — edits code across your git repo with commit-per-change discipline.

no llms.txt yetcoding-agentcliaidergitopen-sourcepair-programming

★ 49.4kApache-2.0v0.86.0 · 2025-08-09verified

Details
About

Aider is an open-source CLI coding assistant that edits code in an existing git repository under the developer's direction. It's known for two core design choices: it operates on git (every change becomes a discrete commit with an auto-generated message, making rollback trivial), and it builds a "repo map" of the codebase so the model always has context for where functions and types live across files.

It supports any LLM provider — Anthropic Claude, OpenAI GPT, local models via Ollama or OpenAI-compatible servers, DeepSeek, xAI. Unlike IDE-coupled assistants, Aider lives in the terminal and plays well with any editor the developer already uses.

The /architect mode separates planning (using a reasoning model) from editing (using a fast edit model) for complex changes. The benchmark leaderboard Aider maintains is a widely-cited reference for LLM coding performance.

Features
  • Git-first workflow with automatic commits per change
  • Repo-map context for large codebase awareness
  • Any LLM provider (Claude, GPT, local, DeepSeek, xAI)
  • {'Architect mode': 'reasoning model plans, edit model executes'}
  • Voice input support
  • In-chat commands for file management, linting, testing
  • Language-agnostic (Python, JS, Rust, Go, Java, C++, … )
Best for
  • Pair programming in the terminal without IDE lock-in
  • Refactoring across multi-file changes with git discipline
  • Adding features to existing codebases with reviewable commits
  • Running local LLMs as coding assistants
  • Developers who prefer vim / emacs over IDE-integrated assistants
License

open-source

Deployment

cli

Platforms
  • linux
  • macos
  • windows
Maintainer

Aider-AI

Claude Code

Anthropic's terminal-first agentic coding assistant with deep tool use and codebase awareness.

llms.txtcoding-agentclianthropicclaudeagenticmcp

verified

Details
About

Claude Code is Anthropic's official terminal-based coding agent. It runs locally against the Claude API and has first-class tool use for reading and editing files, running shell commands, searching codebases, browsing the web, and orchestrating sub-agents. It ships with a hook system, slash commands, MCP server support, and IDE integrations (VS Code, JetBrains).

Unlike chat-only coding assistants, Claude Code is agentic by design: it plans, executes multi-step changes, reads its own output, and recovers from errors. It's built around the Claude Agent SDK and exposes the same primitives to developers who want to build custom agents.

Anthropic publishes both llms.txt and llms-full.txt for the Claude Code documentation, making it one of the best-indexed coding-agent references available to other AI assistants.

Features
  • Terminal CLI with native tool use (Bash, Read, Edit, Write, Grep)
  • Hook system for shaping agent behavior per-project
  • Slash commands and user-defined skills via SKILL.md
  • MCP server support for connecting external tools and data
  • IDE integrations for VS Code and JetBrains
  • Background task support and session persistence
  • Built on the Claude Agent SDK, exposed for custom agent builders
Best for
  • Refactoring and migration across large codebases
  • Debugging and fixing issues with full shell + git access
  • Writing new features with reviewed commits and CI awareness
  • Onboarding to unfamiliar repos via guided exploration
  • Building and testing in one loop without leaving the terminal
License

commercial

Deployment

cli

Platforms
  • linux
  • macos
  • windows
Maintainer

Anthropic

ClineNew

Open-source autonomous coding agent with Plan/Act modes and MCP support, shipped as an IDE extension, CLI and SDK.

llms.txtstub

★ 69.9kApache-2.0desktop-v0.0.43 · 2026-10-02verified

Details
License

open-source

Deployment

local

Cursor

AI coding agent and editor from Anysphere (acquired by SpaceX in 2026) — desktop app, CLI and cloud agents with a multi-vendor model picker.

llms.txtcoding-agentcursoridevscode-forkagent-modecommercial

verified

Details
About

Cursor is an AI coding agent and editor built by Anysphere. SpaceX completed its acquisition of Cursor on 2026-08-14, following a model-training partnership with SpaceXAI announced in April 2026.

The desktop app is a VS Code-based editor with agent mode, inline edits, Tab autocomplete and codebase-aware context. Around it Cursor now ships a CLI agent, cloud agents that run in parallel on their own machines (or on machines you manage) and hand back work for review, integrations with Slack and GitHub, a mobile app, automations, and review bots for pull requests.

The model picker spans OpenAI, Anthropic, Google Gemini, SpaceXAI (Grok) and Cursor's own Composer models. Cursor is closed-source and subscription-based, with a free tier.

Features
  • Desktop editor (VS Code-based) with agent mode, inline edits and Tab autocomplete
  • Cursor CLI agent
  • Cloud agents running in parallel, including on self-managed machines
  • Slack and GitHub integrations, review bots and automations
  • Model choice across OpenAI, Anthropic, Gemini, SpaceXAI (Grok) and Cursor's Composer models
  • Codebase-wide semantic search and @-context references
Best for
  • Professional developers who want maximum AI integration
  • Teams that want a consistent IDE across members with built-in AI
  • Large refactors that benefit from multi-file agent coordination
  • Developers migrating from pure VS Code who want more AI than extensions provide
License

commercial

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

Anysphere (SpaceX)

Devin Desktop

Cognition's AI IDE (formerly Windsurf) that runs local and cloud coding agents from one Agent Command Center.

llms.txtcoding-agentdevinwindsurfcognitionideagent-modecommercial

verified

Details
About

Devin Desktop is the new name for Windsurf, the AI-native IDE that started life at Codeium and is now built by Cognition, the company behind the Devin coding agent. Windsurf became Devin Desktop on June 2, 2026, as an over-the-air update. It is the same editor with the same features. The windsurf.com and codeium.com domains now redirect to devin.ai/desktop.

The main change is the Agent Command Center, a Kanban-style view for running and reviewing many agents at once, both local agents on your machine and Devin Cloud agents. The classic IDE experience stays available: editor, Open VSX extensions, keybindings, workflows and LSPs. Cascade, the agentic assistant Windsurf was known for, is still there. So are Tab completions and inline Command edits.

The primary local agent is now Devin Local, the same agent harness that powers Devin CLI, with subagents, sandboxing, plan mode, worktrees and permissions. You can use Devin Desktop with local-only agents; a Devin Cloud subscription is not required.

Features
  • Agent Command Center to manage local and cloud agents in one view
  • Devin Local agent (shared with Devin CLI) with subagents, plan mode and worktrees
  • Cascade agentic assistant with memories, rules, skills, workflows and hooks
  • Tab completions and inline Command edits
  • MCP server support and AGENTS.md directory instructions
  • Adaptive model router that picks a model per task
  • SSH, Dev Containers and WSL support
  • Installers for macOS, Windows and Linux (apt and yum repositories)
Best for
  • Developers who want an AI IDE that can also dispatch and review cloud agents
  • Running several agent tasks in parallel in isolated git worktrees
  • Existing Windsurf users, whose plans and work carry over unchanged
  • Teams standardizing on one vendor for IDE, CLI and cloud coding agents
License

commercial

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

Cognition

Gemini CLINew

Google's open-source terminal agent for Gemini. Unpaid-tier and Google One users were moved to Antigravity CLI on 2026-06-18.

llms.txtstub

★ 107.2kApache-2.0v0.62.0 · 2026-09-29verified

Details
License

open-source

Deployment

cli

GitHub Copilot

GitHub's AI coding agent — completions, chat and agent mode in major IDEs, Copilot CLI, a desktop app, and a cloud agent that works from issue to pull request.

llms.txtcoding-agentgithubcopilotideenterprisecommercial

verified

Details
About

GitHub Copilot is the original and most widely deployed AI coding assistant, integrated natively into VS Code, Visual Studio, JetBrains IDEs, Neovim, Xcode, and GitHub itself. It offers inline code completion, chat, a multi-file agent mode, pull-request summaries, and command-line assistance via gh copilot.

Under the hood Copilot routes to multiple LLMs — GPT, Claude Sonnet, Gemini — with model choice exposed to users on the paid tiers. Its reach within the GitHub platform (issues, PRs, code review, Actions) makes it a practical default for teams already on GitHub.

For enterprise deployments, Copilot offers SOC 2 compliance, data residency, code reference filtering, and the ability to restrict external traffic — the most mature enterprise posture in the category.

Features
  • VS Code, Visual Studio, JetBrains, Neovim, Xcode integrations
  • Inline autocomplete, chat and agent mode in the IDE
  • Copilot CLI for agentic work in the terminal
  • GitHub Copilot desktop app for managing agent-driven work
  • Cloud agent that plans, codes on a branch and opens pull requests; third-party agents (Claude, Codex) assignable
  • Pull-request code review and MCP server support (including the GitHub MCP Server)
  • Free plan for everyone; Student plan for verified students; usage-based AI Credits on paid plans
Best for
  • Professional development at organizations already on GitHub
  • Teams needing enterprise compliance posture out of the box
  • Delegating issues to a cloud agent and reviewing the resulting PRs
  • Cross-IDE consistency (one subscription, many editors)
  • Students and individual developers on the Free or Student plan
License

commercial

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

GitHub / Microsoft

Kilo CodeNew

MIT-licensed coding agent for VS Code, JetBrains, the CLI and the cloud — 500+ models at provider cost, bring-your-own-keys and local models.

llms.txtcoding-agentvscodejetbrainsclikiloopen-sourcebyoklocal-llm

★ 27.5kMITv7.8.3 · 2026-10-01verified

Details
About

Kilo Code is an open-source (MIT) AI coding agent that works in VS Code, JetBrains, the terminal and the cloud. You pick from 500+ models, can switch mid-task, and pay the model provider's rate with no markup. Bring-your-own-key and local models are supported, and no API key is needed to start.

An agent command center manages local and cloud agents across IDE and CLI sessions, with parallel isolated worktrees. Kilo has been acquired by Anaconda, according to the banner on kilo.ai.

Features
  • VS Code and JetBrains extensions, CLI and cloud agents
  • 500+ models with zero inference markup; mid-task model switching
  • Bring your own keys; local models supported
  • Agent command center for local and cloud agents; parallel isolated worktrees
  • MIT-licensed source you can audit and fork
Best for
  • One open-source agent across VS Code, JetBrains and the terminal
  • Teams that want BYOK and model choice without vendor markup
  • Local-model coding workflows
License

open-source

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

Kilo Code (Anaconda)

KiroNew

AWS's spec-driven AI coding agent (IDE, CLI, web, mobile) — the successor to Amazon Q Developer, whose IDE plugins reach end of support on 2027-04-30.

llms.txtcoding-agentkiroawsamazon-qspec-drivenmcpcommercial

★ 4.3klast push verified

Details
About

Kiro is AWS's AI coding agent, built around one agent harness shared by the Kiro IDE, Kiro CLI, web app, iOS app, Kiro Crew (for orchestrating teams of agents) and any ACP-compatible editor. Its defining idea is spec-driven development: Kiro turns a prompt into requirements, an architectural design and sequenced tasks, then implements them with parallel agents. It checks requirements for contradictions and uses property-based tests to catch edge cases that unit tests miss.

Steering files, event-driven hooks, custom agents and subagents, skills, "powers" and MCP shape the agent's behaviour. AWS directs Amazon Q Developer users to Kiro. The Q Developer user guide says the IDE plugins lose support on April 30, 2027, and Kiro publishes a migration guide. Kiro CLI is the successor to the Q Developer CLI.

Plans: Free (50 credits; open-weight models and Claude Sonnet 4.5), Pro ($20), Pro+ ($40), Pro Max ($100), Power ($200) and Enterprise. Paid plans include premium models such as Claude Sonnet 5 and Claude Opus 5.

Features
  • Spec-driven development (requirements → design → tasks) with parallel agents
  • Requirement checking and property-based tests
  • Steering files, event-driven hooks, custom agents, subagents, skills and powers
  • MCP support
  • Kiro IDE, Kiro CLI (successor to Q Developer CLI), web, iOS, Kiro Crew; ACP for other editors
  • Free, Pro, Pro+, Pro Max, Power and Enterprise plans
Best for
  • Amazon Q Developer users migrating before the 2027-04-30 end of support
  • Teams that want a written spec before the agent touches code
  • AWS-centric organisations standardising on an AWS-supported agent
License

commercial

Deployment

local

Platforms
  • linux
  • macos
  • windows
  • web
Maintainer

AWS

Open Interpreter

Open-source (Apache-2.0) terminal coding agent forked from OpenAI's Codex, tuned for low-cost and open-weight models with switchable harness emulation.

no llms.txt yetstub

★ 68.5kApache-2.0rust-v0.0.55 · 2026-09-30verified

Details
License

open-source

Deployment

cli

OpenAI CodexNew

OpenAI's coding agent, available as an open-source terminal CLI, an IDE extension and cloud automation.

llms.txtstub

★ 127.9kApache-2.0rust-v0.160.0 · 2026-10-01verified

Details
License

open-source

Deployment

cli

Qwen CodeNew

Alibaba Qwen team's Apache-2.0 coding agent for terminal, editor, desktop, browser and chat — OpenAI, Anthropic, Gemini and Qwen APIs or local models.

llms.txtcoding-agentcliqwenalibabaopen-sourcemcplocal-llm

★ 28.3kApache-2.0sdk-typescript-v0.1.18 · 2026-10-05verified

Details
About

Qwen Code is "the open-source AI coding agent for your terminal, editor, desktop, browser, and chat", from Alibaba's Qwen team. It is multi-protocol: it supports the OpenAI, Anthropic, Gemini and Qwen APIs, as well as any third-party provider or local model through Ollama or vLLM, and you can switch at runtime.

It ships auto-memory, auto-skills, subagents, agent teams and MCP. Beyond the terminal there are IDE plugins (VS Code, JetBrains, Zed), a desktop app, a web UI, SDKs, GitHub Actions, and chat integrations for Telegram, DingTalk, WeChat and Feishu.

Features
  • Terminal agent, Apache-2.0
  • OpenAI, Anthropic, Gemini and Qwen API protocols; local models via Ollama / vLLM
  • Auto-memory, skills, subagents and agent teams
  • MCP support and sandboxed execution
  • VS Code, JetBrains and Zed integrations; desktop app, web UI, SDKs, GitHub Actions
Best for
  • Qwen-model users who want a first-party agent
  • An open-source agent with local-model support
  • Teams mixing Chinese and Western model providers
License

open-source

Deployment

cli

Platforms
  • linux
  • macos
  • windows
Maintainer

QwenLM (Alibaba)

Sourcegraph Cody

AI coding assistant for Sourcegraph Enterprise that pulls context from Sourcegraph code search across local and remote codebases.

no llms.txt yetstub

verified

Details
License

commercial

Deployment

local

Workflow Tools 4

ComfyUI

Node-based interface for building image, video, and audio generation workflows with any diffusion or multimodal model.

llms.txtdiffusionstable-diffusionfluxvideo-genimage-genworkflownode-graph

★ 136.1kGPL-3.0v0.38.0 · 2026-09-29verified

Details
About

ComfyUI is the dominant open-source workflow tool for running diffusion models (Stable Diffusion, Flux, SDXL), video models (WAN, LTX, Hunyuan, Mochi, Kling), and audio/TTS models (Qwen3-TTS, F5-TTS, and others) via a visual node graph. Users connect nodes representing loaders, samplers, VAEs, text encoders, post-processors, and custom operations into directed acyclic graphs that execute end-to-end.

Unlike all-in-one UIs, ComfyUI is explicit about every step of the generation pipeline, which makes it preferred by power users who need to mix and match models, implement novel techniques (chunked S2V, looped img2vid, LoRA stacking, controlnet chains), and share reproducible workflows as JSON files.

A large ecosystem of custom nodes (ComfyUI-Manager, ComfyUI-Impact-Pack, KJNodes, easy-use, etc.) extends the base install with thousands of additional operations.

Features
  • Visual node graph with typed inputs/outputs
  • Saves and loads complete workflows as JSON
  • Subgraph support for encapsulating reusable sub-pipelines
  • For-loop and conditional nodes for programmatic workflows
  • Thousands of community-contributed custom nodes
  • Runs on NVIDIA CUDA, AMD ROCm, Apple MPS, Intel XPU
  • API and queue-based server mode for headless execution
  • Manager UI for discovering and installing custom nodes
Best for
  • Image generation with fine-grained control (SDXL, Flux, controlnet)
  • Video generation workflows (text-to-video, image-to-video, audio-driven)
  • AI character animation pipelines (image + audio → talking head video)
  • Model research and technique prototyping
  • Batch content generation for marketing, game assets, stock imagery
  • Building custom pipelines that chain multiple model types
License

open-source

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

Comfy Org

Dify

Open-source LLM app development platform with visual prompt IDE, RAG pipelines, and agent builder in one product.

llms.txtllm-platformdifyragagent-builderself-hostedvisual-ide

★ 157.9k1.17.1 · 2026-09-10verified

Details
About

Dify is an all-in-one LLM application platform designed for teams shipping AI features without building each component from scratch. It combines a visual prompt IDE (version, test, and compare prompts with datasets), a RAG pipeline builder (document ingestion, chunking, embedding, retrieval tuning), an agent builder (tool-using agents with function calling), and a ChatUI / API endpoint generator in a single self-hostable product.

Where LangChain is a developer library and n8n is a horizontal automation platform, Dify is specifically a place for product teams to build, test, and operate LLM-powered apps — prompt engineering, RAG evaluation, and app deployment under one roof.

It's deployed via Docker Compose or Kubernetes; Dify Cloud offers the same product as managed SaaS. Licensed under a modified Apache 2.0 (the Dify Open Source License), which restricts multi-tenant SaaS use and removing Dify branding without a commercial license.

Features
  • Visual prompt IDE with versioning and comparison
  • RAG pipeline builder with document ingestion and retrieval tuning
  • Agent builder with function calling and tool orchestration
  • ChatUI templates and auto-generated API endpoints
  • Dataset-based prompt evaluation and annotation
  • Supports OpenAI, Anthropic, local LLMs via any OpenAI-compatible endpoint
  • Multi-user workspaces with role-based access
  • Docker Compose / Kubernetes self-hosting
Best for
  • Product teams shipping LLM-powered features end-to-end
  • Non-engineers prototyping RAG applications
  • Prompt engineering and evaluation workflows
  • Companies wanting an internal "ChatGPT-for-our-docs" platform
  • Teams avoiding the lock-in of OpenAI GPTs / Anthropic Projects
License

source-available

Deployment

self-hosted

Platforms
  • linux
Maintainer

LangGenius

Langflow

Open-source (MIT) low-code visual builder for AI agents, RAG and MCP workflows — run locally, via Docker, or as Langflow Desktop.

llms.txtstub

★ 155.5kMITv1.12.4 · 2026-09-29verified

Details
License

open-source

Deployment

self-hosted

n8n

Fair-code workflow automation with native AI nodes, 500+ integrations, and first-class self-hosting.

llms.txtworkflowautomationn8nlangchainself-hostedlow-code

★ 206.7k[email protected] · 2026-10-05verified

Details
About

n8n is a workflow automation platform — think Zapier or Make, but self-hostable, source-available, and extended with first-class AI nodes. Users build workflows visually by connecting nodes representing triggers (webhooks, schedules, email, Slack), integrations (500+ across databases, APIs, SaaS), and AI operations (LangChain-powered agents, LLM calls, vector store queries, retrievers).

The AI toolkit makes n8n one of the most pragmatic options for non- engineers building agent-style workflows: a "retrieve from Slack → summarize with Claude → post to Notion" flow is a 5-minute drag-and-drop job. For engineers, n8n's custom-code nodes let you drop into Python or JavaScript when the visual builder isn't enough.

Fair-code license means free for individual and internal business use, with commercial restrictions on building competing products. n8n Cloud provides managed hosting for teams that don't want to run Docker.

Features
  • Visual workflow builder with 500+ integration nodes
  • {'Native AI nodes': 'LangChain agents, LLM calls, embeddings, vector stores'}
  • Webhook, schedule, and event triggers
  • Custom Python / JavaScript nodes for bespoke logic
  • Self-host via Docker or use n8n Cloud
  • Fair-code license (free for personal and internal commercial use)
  • Active community-contributed node catalog
Best for
  • Agent workflows for non-engineers (Slack + Claude + Notion)
  • Internal business automation with AI steps woven in
  • ETL pipelines with LLM-powered data transformation
  • RAG applications without writing Python
  • Replacing Zapier with a self-hosted, AI-aware equivalent
License

source-available

Deployment

self-hosted

Platforms
  • linux
Maintainer

n8n

Voice (STT / TTS) 8

Brethof Voice Pro

Offline voice-to-text, translation and subtitles for Linux and Windows: 30 transcription languages, 38 for translation, LoRA voice training.

llms.txtasrtranslationsubtitlesqwen3hunyuan-mtloraofflinevulkanmcpspeech-to-text
Details
About

Brethof Voice Pro is a commercial desktop app for Linux and Windows that transcribes, translates and subtitles speech entirely on the user's own machine. Transcription runs on the open-source Qwen3-ASR engine (0.6B or 1.7B) in 30 languages, and recognises 22 Chinese regional dialects on its own; translation across 38 languages runs locally on Tencent's Hunyuan MT2. It runs on the CPU or any Vulkan 1.2+ GPU (NVIDIA, AMD, Intel) — no CUDA required.

It types wherever the cursor is (the transcript or its translation), records the microphone, a file or system audio, and exports text, SRT and VTT subtitles whose timings survive translation. Voice training learns from the corrections the user already makes, and an MCP server lets AI agents transcribe and translate locally (both on paid licences). The network is touched only by a licence check, an update check and the model downloads the user starts. Perpetual licence; 14-day free trial.

Features
  • Transcription in 30 languages plus 22 Chinese dialects (Qwen3-ASR 0.6B / 1.7B)
  • Offline translation across 38 languages (Tencent Hunyuan MT2), multi-target and bilingual output
  • Text, SRT and VTT subtitles; translating a subtitle file keeps every timing
  • Voice keyboard: types the transcript or its translation into any app
  • Microphone, audio/video file, or system-audio input
  • LoRA voice training from the user's own corrections; hotwords steer ASR and translation
  • MCP server for AI agents (paid licence)
  • Vulkan GPU inference (NVIDIA / AMD / Intel) or CPU-only; optional noise reduction
Best for
  • Private dictation into any app
  • Meeting and interview transcripts where cloud upload is unacceptable
  • Subtitling and translating your own videos without per-minute fees
  • Live translation while typing: speak one language, type another
  • Giving an AI agent local transcription and translation through MCP
License

commercial

Deployment

local

Platforms
  • linux
  • windows
Maintainer

Brethof AI

Coqui TTS

Deep-learning toolkit for TTS with multi-speaker models and voice cloning, now maintained in the Idiap fork.

no llms.txt yetstub

★ 2.3kMPL-2.0v0.27.5 · 2026-01-26verified

Details
License

open-source

Deployment

library

DeepgramNew

Speech-to-text (Nova-3, Flux), text-to-speech (Aura) and voice-agent APIs, with SDKs, a CLI, an MCP server and a self-hosted option.

llms.txtasrttsspeech-to-textvoice-agentsmcp

verified

Details
About

Deepgram provides APIs for speech-to-text, text-to-speech, audio intelligence and end-to-end voice agents. Its STT models cover streaming and pre-recorded audio (Nova-3) and conversational, turn-aware streaming for voice agents (Flux). Aura models handle TTS. A Voice Agent API chains STT, an LLM and TTS in one connection.

Developers integrate through SDKs (Python, JavaScript, Go, .NET, Java), a dg CLI that also runs as an MCP server, and agent skills. Enterprises can self-host.

Features
  • Streaming and batch speech-to-text (Nova-3) and conversational STT (Flux)
  • Text-to-speech (Aura) with single-request and streaming modes
  • End-to-end Voice Agent API
  • Audio and text intelligence (sentiment, topics, summaries, intents)
  • SDKs for Python, JavaScript, Go, .NET and Java; CLI and MCP server
  • Self-hosted deployment option
Best for
  • Real-time transcription for calls, meetings and voice agents
  • Building low-latency voice agents
  • Giving a coding agent speech tools via MCP
License

freemium

Deployment

saas

Platforms
  • web
Maintainer

Deepgram

ElevenLabsNew

APIs and SDKs for text-to-speech, voice cloning, speech-to-text and conversational voice agents.

llms.txtstub

verified

Details
License

freemium

Deployment

saas

F5-TTS

High-quality open-source TTS with voice cloning from short audio reference.

no llms.txt yetstub

★ 15.3kMIT1.1.22 · 2026-07-23verified

Details
License

open-source

Deployment

library

Piper

Fast, local neural text-to-speech engine with CLI, web server, Python and C/C++ APIs, now developed by the Open Home Foundation.

no llms.txt yetstub

★ 5.8kGPL-3.0v1.8.0 · 2026-09-04verified

Details
License

open-source

Deployment

library

VoxCPMNew

Open-source tokenizer-free TTS (VoxCPM2, 2B) with 30 languages, voice design from text prompts, controllable voice cloning and 48kHz output.

no llms.txt yetttsvoice-cloningvoice-designmultilingualopenbmbloraapache-2.0

★ 38.3kApache-2.02.0.3 · 2026-05-11verified

Details
About

VoxCPM is OpenBMB's open-source text-to-speech system. Instead of predicting discrete audio tokens, it generates continuous speech representations with an end-to-end diffusion-autoregressive architecture, which the authors credit for its natural, expressive output. The current release, VoxCPM2, is a 2B-parameter model trained on over 2 million hours of multilingual speech. It synthesizes 30 languages without a language tag, plus several Chinese dialects, and outputs 48kHz audio directly.

Its distinguishing features are Voice Design, which creates a new voice from a natural-language description (gender, age, tone, emotion, pace) with no reference audio, and Controllable Cloning, which clones a voice from a short clip while steering emotion and pacing. "Ultimate cloning" continues from a reference clip plus its transcript to preserve timbre and rhythm closely.

It ships as a pip package with a Python API, a voxcpm CLI, a local web demo, and LoRA / full fine-tuning scripts with a training WebUI. For production it can be served through Nano-vLLM or vLLM-Omni (OpenAI-compatible), and runs without Python via llama.cpp-omni GGUF builds on CPU, CUDA, Metal or Vulkan. Code and weights are Apache-2.0, free for commercial use.

Features
  • 2B-parameter VoxCPM2 model, 30 languages plus Chinese dialects, 48kHz output
  • Voice Design — create a voice from a text description, no reference audio
  • Controllable and "ultimate" voice cloning from short reference clips
  • Real-time streaming (RTF ~0.3 on RTX 4090, ~0.13 with Nano-vLLM)
  • Python API, voxcpm CLI and local web demo (CUDA, CPU, Apple MPS)
  • LoRA and full fine-tuning from 5–10 minutes of audio, with a training WebUI
  • Serving via Nano-vLLM / vLLM-Omni; GGUF inference via llama.cpp-omni
  • Apache-2.0 code and weights
Best for
  • Local, commercial-friendly multilingual TTS for apps and agents
  • Designing brand or character voices from a written description
  • Voice cloning for narration, dubbing and accessibility, with consent
  • Fine-tuning a custom speaker or domain voice on a small dataset
  • Self-hosted OpenAI-compatible TTS endpoints
License

open-source

Deployment

library

Platforms
  • linux
  • macos
Maintainer

OpenBMB

whisper.cpp

C++ port of OpenAI Whisper for local speech-to-text — no Python, runs on CPU and many GPU backends.

no llms.txt yetasrwhispercpptranscriptionggmloffline

★ 54.1kMITv1.9.4 · 2026-09-11verified

Details
About

whisper.cpp is a pure-C/C++ implementation of OpenAI's Whisper speech recognition model. It has zero runtime dependencies (no Python, no CUDA-specific tooling), ships as a library and CLI, and runs on every hardware backend the ggml-org ecosystem supports: CPU (with AVX / NEON SIMD), CUDA, Metal (Apple Silicon), Vulkan, OpenCL, SYCL, CoreML.

It uses GGML quantized models — the same format ecosystem as llama.cpp — enabling 4-bit, 5-bit, and 8-bit quantizations of the Whisper checkpoints that run on modest hardware with minimal quality loss. Real-time transcription is possible on consumer laptops, and batch transcription scales to longer audio with speaker diarization and word-level timestamps.

Widely used as a library inside desktop applications, mobile apps, and server-side transcription pipelines that don't want Python dependencies. Jan, KoboldCpp, and many other tools embed whisper.cpp.

Features
  • Pure C/C++ with no runtime dependencies
  • GGML quantization (4-bit, 5-bit, 8-bit Whisper variants)
  • {'Multiple backends': 'CPU, CUDA, Metal, Vulkan, OpenCL, SYCL, CoreML'}
  • Word-level timestamps and speaker diarization
  • Real-time transcription on consumer hardware
  • CLI, library, and language bindings (Python, Rust, Go, Node.js, …)
  • Supports Whisper tiny, base, small, medium, large, turbo
Best for
  • Embedding local transcription in desktop / mobile apps
  • Offline transcription pipelines without Python
  • Real-time transcription on resource-constrained hardware
  • Edge deployments (Raspberry Pi, mobile, embedded)
  • Server-side transcription with predictable resource usage
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
  • ios
  • android
Maintainer

ggml-org

Image Generation 5

AUTOMATIC1111

Classic Stable Diffusion web UI with a large extension ecosystem — development has stalled (last release v1.10.1, Feb 2025; last commit Mar 2026).

no llms.txt yetstub

★ 165.2kAGPL-3.0v1.10.1 · 2025-02-09verified

Details
License

open-source

Deployment

self-hosted

DiffusersNew

Hugging Face's library of pretrained diffusion pipelines for generating images, video and audio, with LoRA, quantization, offloading and training.

llms.txtdiffusionimage-generationvideo-generationhuggingfacepytorchlora

★ 34.7kApache-2.0v0.40.0 · 2026-08-20verified

Details
About

Diffusers is Hugging Face's open-source (Apache-2.0) Python library for state-of-the-art pretrained diffusion models that generate images, video and audio. It is built around DiffusionPipeline, which offers inference in a few lines, mix-and-match components (models, schedulers) and adapters such as LoRA.

It includes memory and speed optimisations: offloading, quantization, caching techniques, attention backends and torch.compile. Hardware guides cover Apple MPS, Core ML, ONNX Runtime, OpenVINO and AWS Neuron. It also includes training scripts. Most image UIs and hosted services that serve open diffusion models use Diffusers or its model definitions.

Features
  • DiffusionPipeline for text-to-image, image-to-image, video and audio generation
  • Interchangeable models and schedulers; AutoPipeline
  • LoRA and other adapters
  • Offloading, quantization, caching, attention backends, torch.compile
  • Apple MPS, Core ML, ONNX Runtime, OpenVINO and AWS Neuron guides
  • Training and fine-tuning scripts
Best for
  • Building image or video generation into Python applications
  • Running new open diffusion models from the Hub on day one
  • Fine-tuning diffusion models or LoRAs
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

Hugging Face

InvokeAI

Free, open-source (Apache 2.0) self-hosted creative engine for AI image generation with a layer-based unified canvas and node workflows.

llms.txtstub

verified

Details
License

open-source

Deployment

local

Krita AI Diffusion

Krita plugin for Stable Diffusion — inpaint, img2img, and generative layers inside Krita.

no llms.txt yetstub

★ 10.7kGPL-3.0v1.53.0 · 2026-08-22verified

Details
License

open-source

Deployment

local

SD.Next

Advanced fork of SD WebUI with broader model support (Flux, Lumina, Kolors, more).

no llms.txt yetstub

★ 7.4kApache-2.0last push verified

Details
License

open-source

Deployment

self-hosted

Vector Databases 9

Chroma

Open-source search infrastructure for AI — embedded, client-server or Chroma Cloud, with vector, full-text and (in Cloud) hybrid search.

llms.txtvector-dbchromaragembeddedembeddingssqlite

★ 29.4kApache-2.01.5.9 · 2026-05-05verified

Details
About

Chroma is a widely used open-source vector database / search infrastructure for LLM applications. It's designed with developer experience as the priority: a simple Python / JavaScript API, no cluster configuration needed to start, and sensible defaults for the RAG use case.

Chroma runs in three modes: embedded (in-process, SQLite-backed — ideal for prototyping and small apps), standalone server (Docker, single node), and Chroma Cloud (managed). The same API works across all three, letting projects graduate from local dev to production without rewrites.

It supports full-text search, metadata filtering, multi-modal embeddings, and several distance functions. Integrations with LangChain, LlamaIndex, LiteLLM, and HuggingFace embeddings make it a common first-pick for new RAG projects.

Features
  • Embedded (in-process), client-server, and Chroma Cloud modes with one API
  • Official Python, TypeScript/JavaScript and Rust clients
  • Full-text and regex search alongside vector search
  • Metadata filtering with boolean and comparison operators
  • Multiple distance metrics (cosine, L2, inner product)
  • Built-in embedding functions (OpenAI, HF, ONNX, and more)
  • Chroma CLI for running a local server, browsing and copying collections
  • Chroma Cloud Search API with sparse vectors and hybrid search (RRF)
Best for
  • Prototyping RAG applications quickly without infra overhead
  • Local / embedded vector search in desktop or mobile apps
  • Small-to-medium RAG deployments (up to millions of vectors)
  • Teaching / learning RAG patterns
  • Applications graduating from embedded to cloud without rewrites
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

Chroma

LanceDB

Multimodal lakehouse and embedded vector database on the open Lance format — vector, full-text and hybrid search plus training-data curation.

llms.txtstub

★ 11.6kApache-2.0v0.39.0 · 2026-09-17verified

Details
License

open-source

Deployment

library

Milvus

Open-source cloud-native vector database built for billion-scale similarity search with separation of storage and compute.

llms.txtvector-dbmilvusdistributedbillion-scalegpurag

★ 46.3kApache-2.0v3.0.2 · 2026-09-20verified

Details
About

Milvus is a distributed vector database designed for the largest vector workloads — billions of vectors, hundreds of billions of searches per day. Its architecture separates storage, compute, and coordination into independent services, letting each scale horizontally. At the cost of more operational complexity than smaller vector DBs, Milvus handles scale few other options reach.

Milvus supports multiple index types optimized for different trade-offs (IVF_FLAT, IVF_SQ8, HNSW, SCANN, DISKANN, GPU_IVF_FLAT for GPU-accelerated search), multiple distance metrics, hybrid search, and filtered search with complex boolean predicates. Attu provides a graphical admin UI for cluster operations.

For teams with smaller scale requirements, Milvus Lite ships as an embedded Python library using the same API as the full cluster — start local, graduate to distributed when the workload demands it.

Features
  • {'Distributed architecture': 'separated storage, compute, coordination'}
  • Multiple index types (IVF, HNSW, SCANN, DISKANN, GPU-accelerated)
  • {'Hybrid search': 'dense + sparse vectors + filters'}
  • Multi-tenancy through partitions and databases
  • Milvus Lite for embedded / local use with same API
  • GPU search support for large-scale low-latency workloads
  • Attu admin UI and CLI tools
Best for
  • Vector search at 1B+ scale
  • High-QPS inference products with sub-50ms retrieval SLOs
  • Multi-tenant vector search in a single deployment
  • GPU-accelerated similarity search
  • Teams graduating from Milvus Lite (local) to production cluster
License

open-source

Deployment

self-hosted

Platforms
  • linux
Maintainer

Milvus / Zilliz

pgvector

Postgres extension adding vector similarity search — the "just use Postgres" option for RAG.

no llms.txt yetvector-dbpostgrespgvectorextensionraghnsw

★ 23.2klast push verified

Details
About

pgvector is a PostgreSQL extension that adds vector similarity search to any Postgres database. It supports L2, inner product, and cosine distance; exact and approximate nearest-neighbor search via IVFFlat and HNSW indexes; and integrates with every Postgres client library on every platform.

The appeal is simplicity: organizations already running Postgres can add vector search without provisioning a separate vector database, managing a second replication topology, or introducing a new client library. ACID transactions that join vector data with traditional relational data work out of the box.

pgvector is supported by every major managed Postgres provider (Supabase, Neon, AWS RDS, GCP Cloud SQL, Azure Database for PostgreSQL, Crunchy Bridge), making it a safe default for RAG applications that don't have billion-scale vector requirements.

Features
  • Native Postgres extension — vectors as a first-class column type
  • Exact and approximate search (IVFFlat, HNSW indexes)
  • L2, inner product, and cosine distance
  • Integrates with every Postgres driver and ORM
  • ACID transactions joining vector and relational data
  • Supported by all major managed Postgres providers
  • Up to 16,000 dimensions per vector
Best for
  • RAG applications where Postgres is already in the stack
  • Teams avoiding the operational burden of a dedicated vector DB
  • Applications mixing relational queries with vector search
  • Small-to-medium RAG (up to low-millions of vectors per table)
  • Rapid prototyping using free-tier Supabase / Neon
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

pgvector

Pinecone

Managed serverless vector database plus Nexus knowledge engine and Assistant — Pinecone's "AI knowledge platform" for agents and RAG.

llms.txtvector-dbpineconesaasmanagedserverlessrag

verified

Details
About

Pinecone is the original managed vector database and remains the most widely adopted SaaS option for production vector search. Its serverless architecture abstracts away every operational concern: no cluster sizing, no replication config, no reindexing during scaling events. Queries return in single-digit milliseconds at billion-vector scale.

Pinecone's product expanded beyond pure vector search to include Assistant (hosted RAG as a service — upload documents, get a chatbot endpoint), embedding inference (text → embedding via the same API that stores them), and reranking. For teams that want a complete managed RAG stack without operating any infrastructure, Pinecone's breadth is hard to match.

Trade-offs: proprietary (closed-source), most expensive option at scale versus self-hosted Qdrant/Milvus/Weaviate, and vendor lock-in for teams building on its Assistant layer. For small indexes it's often cheaper than self-hosting; for very large indexes the calculus shifts.

Features
  • Serverless — no cluster management, automatic scaling
  • Sub-10ms query latency at billion-vector scale
  • Metadata filtering with complex boolean queries
  • Hybrid search (dense + sparse vectors)
  • {'Pinecone Assistant': 'managed RAG as an API'}
  • Built-in embedding inference and reranking
  • Multi-region deployment for latency
  • Clients in Python, JavaScript, Go, Java, .NET, Rust
  • Pinecone Nexus — compiles sources into queryable knowledge and serves cited answers to agents
  • Bring Your Own Cloud (BYOC) deployment
Best for
  • Production RAG without operating vector-DB infrastructure
  • Applications needing sub-10ms vector query SLOs
  • Teams that want a complete managed RAG stack (DB + embeddings + RAG)
  • Companies where ops time is more expensive than storage cost
  • Multi-region vector search for global applications
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Pinecone

Qdrant

Open-source, Rust-written vector database built for production scale — rich filtering, hybrid search, and multi-tenancy.

llms.txtvector-dbqdrantrustproductionhybrid-searchself-hosted

★ 34.9kApache-2.0v1.19.1 · 2026-09-04verified

Details
About

Qdrant is a production-oriented vector database written in Rust, designed for organizations serving vector search at scale. It emphasizes advanced filtering (complex boolean queries on metadata alongside vector search), multi-tenancy (isolated collections per customer), hybrid search (dense

  • sparse), and memory efficiency (quantization on-disk).

It offers a REST API, gRPC, and clients for Python, JavaScript, Rust, Go, Java, and C#. Qdrant Cloud provides managed hosting; self-hosted deployments scale horizontally through Qdrant's sharded cluster mode.

Compared to Chroma's "get started in five minutes" positioning, Qdrant targets production: teams that know they'll need 10M+ vectors, complex filters, and SLOs from the start.

Features
  • Rust core for performance and memory safety
  • Advanced payload filtering with complex boolean logic
  • Hybrid search (dense + sparse vectors)
  • On-disk storage with scalar / product / binary quantization
  • Distributed mode with sharding and replication
  • REST API, gRPC, and clients in 6+ languages
  • Multi-tenancy through isolated collections
  • Snapshot and restore for backup
Best for
  • Production vector search at 10M+ vectors
  • Multi-tenant SaaS needing per-customer isolation
  • Hybrid retrieval combining semantic and keyword search
  • Memory-constrained deployments using quantization
  • RAG systems with complex metadata filtering requirements
License

open-source

Deployment

self-hosted

Platforms
  • linux
  • macos
  • windows
Maintainer

Qdrant

SurrealDB

Multi-model database in Rust (document, graph, vector, time-series, relational) positioned as a context and memory layer for AI agents.

llms.txtdatabasemulti-modelvectorgraphdocumentsurrealdbrust

★ 33.1kv3.3.0 · 2026-09-28verified

Details
About

SurrealDB is a multi-model database that unifies document, graph, key-value, time-series, and vector workloads in a single Rust-written engine. Where traditional RAG stacks glue a vector database to a relational database to a graph store, SurrealDB lets a single query span all of them — joining user data, embeddings, and graph relationships without cross-system synchronization.

For AI applications the vector search capability supports HNSW and DISKANN indexes alongside scalar filters and full graph traversal. The same query can filter users by attribute, retrieve their document embeddings by cosine similarity, and walk the graph of their social connections — all in one SurrealQL statement.

SurrealDB runs embedded (in-process, WASM-compatible for browser / Node.js), as a standalone server, or in distributed cluster mode (SurrealDB Cloud, or self-hosted with TiKV). It's used increasingly as the backing store for agent memory, RAG systems, and real-time applications that need multi-model data without the operational overhead of multiple databases.

Features
  • Document, graph, key-value, time-series, and vector in one engine
  • HNSW and DISKANN vector indexes with cosine / Euclidean / Manhattan distance
  • SurrealQL — one query language across all models
  • Embedded (in-process) or standalone server modes
  • Live queries with real-time change notifications
  • Native graph traversal with SELECT on edges
  • WebAssembly builds for browser and Node.js
  • Schemaless or schemafull tables, per your preference
  • Agent Memory layer with provenance and fact supersession
Best for
  • Agent memory stores combining facts, relationships, and embeddings
  • RAG applications that need relational context alongside vectors
  • Real-time applications with live-query subscriptions
  • Graph-heavy AI workflows (knowledge graphs, entity linking)
  • Replacing a vector-DB + Postgres + Redis stack with one system
  • Embedded database for desktop AI apps (WASM / single-binary)
License

source-available

Deployment

self-hosted

Platforms
  • linux
  • macos
  • windows
  • web
Maintainer

SurrealDB

turbopufferNew

Serverless vector and full-text search engine built on object storage — fast, low-cost and scaling to 1T+ documents.

llms.txtvector-dbfull-text-searchserverlessobject-storage

verified

Details
About

turbopuffer is a hosted vector and full-text (BM25) search engine built from first principles on object storage, with a memory/SSD cache in front. Storing data in object storage keeps costs low at very large scale. Hot namespaces are served from cache with low latency, and cold queries read from storage. The vendor reports production scale of 1T+ documents, 25k+ queries/s and 10M+ writes/s.

It offers strongly consistent queries, namespace branching (copy-on-write), sharding, built-in embedding and multiple regions.

Features
  • Vector and BM25 full-text search with filtering
  • Object-storage-native architecture with memory/SSD caching
  • Strongly consistent reads after durable upserts
  • Copy-on-write namespace branching and sharding
  • Built-in embedding; multiple cloud regions
Best for
  • Very large, multi-tenant vector search at low storage cost
  • Hybrid keyword + semantic retrieval for RAG and code search
License

commercial

Deployment

saas

Platforms
  • web
Maintainer

turbopuffer

Weaviate

Open-source vector database with built-in ML modules, hybrid search, and first-class RAG tooling.

llms.txtvector-dbweaviatehybrid-searchmulti-tenancyraggo

★ 16.9kv1.39.9 · 2026-10-05verified

Details
About

Weaviate is an open-source vector database written in Go, designed for AI-native applications. It distinguishes itself from minimal vector stores by bundling vectorization modules (OpenAI, Cohere, HuggingFace, Ollama, and more) directly into the database — you can ingest raw text and have Weaviate embed it for you, without a separate embedding service.

Weaviate supports hybrid search (dense + sparse, BM25F), multi-tenancy through tenants-per-class isolation, generative modules (RAG inside the database: query → retrieve → generate as one call), and gRPC and REST APIs (GraphQL is legacy). Cluster mode handles sharding and replication for production-scale deployments.

Weaviate Cloud (serverless managed) is available for teams that don't want to operate the cluster themselves.

Features
  • Built-in vectorization modules (no separate embedding service needed)
  • Hybrid search combining dense vectors and BM25F sparse
  • Generative modules — RAG as a single database call
  • Multi-tenancy with per-tenant data isolation
  • Cluster mode with sharding and replication
  • gRPC (data operations) and REST APIs; legacy GraphQL API still available
  • Official Python, TypeScript, Java and C# clients (Go client being updated)
  • Weaviate Cloud agents — Query Agent (managed RAG) and Engram agent memory (preview)
Best for
  • Production RAG systems with built-in embedding pipelines
  • Multi-tenant SaaS where each customer needs isolated vector search
  • Applications combining keyword and semantic search
  • Teams wanting RAG-in-the-database rather than separate components
  • Hybrid search use cases requiring BM25F + vectors
License

open-source

Deployment

self-hosted

Platforms
  • linux
Maintainer

Weaviate

RAG Frameworks 4

AnythingLLM

All-in-one desktop and Docker RAG app — document ingestion, agents, multi-user.

no llms.txt yetstub

verified

Details
License

open-source

Deployment

self-hosted

Haystack

Production-oriented Python framework for building RAG, search, and agent pipelines with composable components.

llms.txtraghaystackdeepsetpipelinesemantic-searchpython

★ 26.7kApache-2.0v3.3.0 · 2026-10-01verified

Details
About

Haystack (from deepset) is an open-source Python framework for building LLM-powered pipelines — RAG, semantic search, question answering, and agents — out of composable, typed components. Its core abstraction is the Pipeline: a directed graph of components (retrievers, rankers, generators, prompt builders, validators) connected by typed edges, with full control over execution order and branching.

Haystack predates the LLM gold rush (originally an NLP QA toolkit) and its design reflects production priorities: typed I/O, explicit observability, easy component swapping, and deployment as Kubernetes-ready services. The 2.x rewrite rebuilt it around the Pipeline abstraction; Haystack 3.0 (July 2026) added a more capable Agent with hooks and skills, run introspection, first-class async serving and a leaner core.

It's a common choice for teams who find LangChain too abstract and LangGraph too low-level, wanting a middle-ground framework with production defaults.

Features
  • Composable Pipeline abstraction with typed components
  • 100+ integrations (LLMs, vector stores, retrievers, rankers)
  • Branching, looping, and conditional pipelines
  • Evaluation harness built into the framework
  • Haystack Enterprise Platform (deepset) for hosted, governed deployment
  • Hayhooks to serve pipelines and agents as REST APIs or MCP tools
  • Python-only (no JavaScript port)
Best for
  • Production RAG systems with strict latency / cost requirements
  • Semantic search over large document collections
  • Hybrid retrieval combining dense and sparse methods
  • Agent pipelines with clear, inspectable control flow
  • Teams wanting a more disciplined framework than LangChain
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

deepset

LlamaIndex

Open-source Python framework for RAG and agents over private data — loaders, indexes, retrievers, query engines and workflows (from the makers of LlamaParse).

llms.txtragllamaindexretrievalindexespythontypescriptknowledge

★ 52.4kMITv0.14.25 · 2026-09-21verified

Details
About

LlamaIndex is a data framework for building LLM applications over private data. Where LangChain is general-purpose, LlamaIndex is specifically optimized for retrieval-augmented generation (RAG): ingesting documents, building indexes, retrieving relevant chunks at query time, and composing them into prompts.

It provides high-level abstractions for common RAG patterns (summary index, vector index, tree index, keyword index, knowledge graph index) and lower-level primitives for developers who want custom retrieval pipelines. Integrations span hundreds of data sources (PDFs, Notion, Slack, Google Drive, databases) and vector stores (Chroma, Qdrant, Weaviate, Pinecone, pgvector, LanceDB, …).

LlamaIndex Agents extend the framework with tool-using agents that can query multiple indexes, call APIs, and compose results. The company's hosted product is now LlamaParse (formerly LlamaCloud), a document parsing, extraction and classification platform.

Features
  • Document loaders for 200+ data sources
  • Multiple index types (vector, summary, tree, keyword, knowledge graph)
  • Query engines with advanced retrieval (HyDE, re-ranking, query decomposition)
  • Agent and event-driven Workflows framework built on top of indexes and tools
  • Document parsing and extraction via the companion LlamaParse platform
  • Integrations with all major vector stores and LLM providers
  • Python library (the TypeScript port, LlamaIndex.TS, was deprecated and archived in April 2026)
Best for
  • RAG applications over corporate document collections
  • Chatbots answering questions from manuals, wikis, support articles
  • Structured data extraction from PDFs, scanned documents
  • Agents querying multiple private knowledge sources
  • Research assistants over academic paper collections
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

LlamaIndex

PrivateGPT

Open-source (Apache-2.0) Claude-API-style layer for private AI apps on any local OpenAI-compatible model server — agentic RAG with citations, tools, MCP and data access.

llms.txtstub

★ 57.6kApache-2.0v1.0.1 · 2026-06-18verified

Details
License

open-source

Deployment

self-hosted

Agent Memory 3

Brethof BrainNew

Memory for AI agents that is already there when a session starts — curated records, rules and the full history of what was said; it processes your memory and never stores it. Disclosure: maintained by us.

llms.txtmemoryagentsmcpclaude-codelong-term-memorybrethof

★ 0last push

Details
About

Brethof Brain is a memory system for AI agents: memory that is already there when a new session starts, that hands the agent what matters the moment it is relevant, and that keeps everything said so it can be found again. Coding agents are its first audience, not the limit of it.

The memory curates itself. It reads what was said and keeps what a future session would need as curated records, reconciling new facts against old ones; standing rules arrive with almost every message, the records that bear on a prompt arrive with it, and a brief opens every session. When that is not enough, everything ever said can be searched.

Your memory lives on your own machine (local edition) or in our cloud, encrypted under a key only you hold (hosted edition). Our hub processes each exchange to make memory of it and stores none of it; the processing runs on secure compute in Zurich, and the model provider is bound by contract not to log, retain or train on what passes through. The client is source-available. Disclosure: maintained by us (Brethof AI).

Features
  • Curated records per project, reconciled as facts change
  • Standing rules and a session brief delivered automatically
  • Full searchable archive of every conversation
  • Playbooks — the agent's own procedures, kept and reused
  • Local edition (memory on your machine) or hosted (encrypted under your key)
  • Proven on 20 agent harnesses, including Claude Code, Codex, OpenCode, Cline, Gemini CLI and Hermes Agent
  • MCP server and plugins for each harness; free plan to start
Best for
  • Coding agents that remember decisions and context across sessions and machines
  • Several agents and sessions sharing one project memory
  • Replacing hand-kept memory files with curated, searchable memory
  • Applications that keep a separate memory per end-user or topic
License

freemium

Deployment

saas

Platforms
  • linux
  • windows
Maintainer

Brethof AI

Claude Code memory (built in)New

Claude Code's own memory: CLAUDE.md (or AGENTS.md) instruction files you write, plus auto memory — notes Claude writes itself from your corrections — both loaded at the start of every session.

llms.txtmemoryclaude-codeclaude-mdagents-mdanthropicbuilt-in

verified

Details
About

Each Claude Code session starts with a fresh context window. Two built-in mechanisms carry knowledge across sessions. CLAUDE.md files are plain-text instructions you write, scoped to an organization (managed policy), to you across all projects (~/.claude/CLAUDE.md) or to a project; Claude Code can also read a repository's AGENTS.md, and path-scoped rules live in .claude/rules/. Auto memory is the other half: notes Claude writes itself from your corrections and preferences, kept per repository and shared across worktrees.

Both are loaded into every session (auto memory up to its first 200 lines or 25KB). Anthropic's documentation describes them as context, not enforced configuration: to block an action regardless of what Claude decides, use a hook instead. Subagents can keep their own auto memory.

Features
  • CLAUDE.md files at organization, user and project scope
  • Reads AGENTS.md, alone or alongside CLAUDE.md
  • Path-scoped rules in .claude/rules/
  • Auto memory written by Claude from corrections, per repository
  • Loaded at the start of every session
  • Persistent memory for subagents
Best for
  • Coding standards, build commands and project layout Claude should always know
  • Remembering corrections so the same mistake is not repeated
  • Organization-wide instructions managed by IT
License

commercial

Deployment

cli

Platforms
  • linux
  • macos
  • windows
Maintainer

Anthropic

Mem0

Persistent memory layer for AI agents — remembers user facts, preferences, and context across sessions.

llms.txtmemoryagentsmem0personalizationraglong-term-memory

★ 66.6kApache-2.0ts-v3.3.1 · 2026-09-25verified

Details
About

Mem0 is a memory framework that gives AI agents persistent, queryable memory across conversations. It extracts facts from each interaction ("the user's dog is named Rex", "the user prefers Python over JavaScript"), stores them in a vector index (with entity-graph memory on the managed Platform), and retrieves relevant memories when the agent needs context for a new query.

Compared to raw vector databases where applications manage their own schema, Mem0 provides an opinionated memory abstraction: user-scoped memories, automatic fact extraction, memory updates (the dog's name changed? the old fact gets superseded), and ranked retrieval tuned for the agent-memory use case specifically.

It integrates with LangChain, LlamaIndex, CrewAI, and any OpenAI-compatible client. Open-source core under Apache 2.0 with a managed hosted tier (Mem0 Platform) for teams that don't want to operate the storage layer.

Features
  • Automatic fact extraction from conversation turns
  • User-scoped memories with updates and deletions
  • Vector retrieval with optional reranking (OSS); entity graph memory on Mem0 Platform
  • Integrations with LangChain, LangGraph, LlamaIndex, CrewAI, raw OpenAI SDK
  • Many vector-store backends (Qdrant, Weaviate, pgvector, Chroma, and more)
  • Open-source core (Apache 2.0) + managed Mem0 Platform
  • Python and TypeScript SDKs, CLI, and a hosted MCP server
Best for
  • Personalized chat assistants that remember user preferences
  • Customer support agents that track interaction history
  • Long-running agents that accumulate project-specific context
  • Multi-session agent workflows without bloating each prompt
  • Replacing ad-hoc "memory in the system prompt" with structured storage
License

freemium

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

Mem0

Embeddings 4

BGE / FlagEmbedding

BAAI's open embedding and reranker models with the FlagEmbedding toolkit for inference, evaluation and fine-tuning — a one-stop retrieval toolkit for search and RAG.

no llms.txt yetstub

★ 12.2kMITv1.4.2 · 2026-08-24verified

Details
License

open-source

Deployment

library

FastEmbed

Lightweight embedding library from Qdrant on ONNX Runtime (no PyTorch) — dense, sparse and late-interaction embeddings and rerankers, CPU or GPU.

no llms.txt yetstub

★ 3.2kApache-2.0v0.8.1 · 2026-09-22verified

Details
License

open-source

Deployment

library

Sentence Transformers

Python framework for state-of-the-art sentence, text, and image embeddings.

no llms.txt yetstub

verified

Details
License

open-source

Deployment

library

Voyage AINew

MongoDB's Voyage AI embedding and reranking API — text, contextualized-chunk and multimodal embeddings for retrieval and RAG.

llms.txtembeddingsrerankerragapimongodbmultimodal

verified

Details
About

Voyage AI, now "Voyage AI by MongoDB", provides hosted embedding models and rerankers through an API and Python client. The catalog includes text embeddings, contextualized chunk embeddings (chunks embedded with document context), multimodal embeddings and rerankers. Flexible output dimensions and quantization reduce storage cost, and a batch inference API handles large offline jobs.

Voyage models also power retrieval inside MongoDB's data platform. The API is commonly used as the embedding step in RAG stacks paired with any vector database.

Features
  • Text, contextualized-chunk and multimodal embedding models
  • Rerankers
  • Flexible dimensions and quantized outputs
  • Batch inference API
  • Python client; organizations, projects and published SLOs
Best for
  • Embedding and reranking for RAG and semantic search
  • Document retrieval where chunk context matters
  • Teams already on MongoDB Atlas
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

MongoDB

Observability 7

Arize Phoenix

AI observability and evaluation platform built on OpenTelemetry — tracing, evals, prompt playground, datasets and experiments; self-host or Arize cloud.

llms.txtstub

★ 11.7karize-phoenix-v20.19.0 · 2026-10-01verified

Details
License

source-available

Deployment

self-hosted

Helicone

Open-source AI gateway and LLM observability (requests, costs, latency, sessions); in maintenance mode since its 2026 acquisition by Mintlify.

llms.txtstub

★ 6.2kApache-2.0v2025.08.21-1 · 2025-08-21verified

Details
License

open-source

Deployment

saas

Maintainer

Helicone (Mintlify)

Langfuse

Open-source LLM engineering platform for tracing, evaluation, prompt management, and observability — self-host or cloud.

llms.txtobservabilitytracingopen-sourceself-hostedevaluationprompt-management

★ 35.4kv4.50.0 · 2026-10-02verified

Details
About

Langfuse is an open-source AI engineering platform (acquired by ClickHouse, which kept it open source and self-hostable). It provides traces, evaluations, prompt versioning, datasets and cost tracking, with the core under the MIT license (enterprise features in ee/ directories are commercially licensed) and a straightforward self-hosted option for teams that need data sovereignty.

It integrates with LangChain, LangGraph, Llama-Index, OpenAI SDK, Anthropic SDK, LiteLLM, and any HTTP-based LLM call via its OpenTelemetry support. Traces capture full LLM inputs/outputs, tool calls, latencies, and costs. The prompt management UI lets product teams iterate on prompts outside of code and deploy changes without shipping releases.

Langfuse publishes both llms.txt and llms-full.txt at well-known paths, making it easy for agents to answer questions about its APIs and usage patterns.

Features
  • Full trace capture (LLM calls, tool calls, sessions, users)
  • Prompt management with versioning, labels, and A/B rollouts
  • Dataset and evaluation runs (LLM-as-judge, custom, exact-match)
  • Cost and latency analytics dashboards
  • OpenTelemetry-based instrumentation; framework-agnostic
  • Self-host (Docker Compose, Kubernetes) or Langfuse Cloud
  • MIT-licensed core (enterprise ee/ features under a commercial license)
Best for
  • Open-source observability for LLM apps
  • Self-hosted deployments where data can't leave the org
  • Prompt A/B testing and versioned rollouts
  • Cost monitoring across many agents / features / users
  • Replacing paid observability with a self-operated stack
License

open-source

Deployment

self-hosted

Platforms
  • linux
Maintainer

Langfuse (ClickHouse)

LangSmith

Commercial observability, debugging, and evaluation platform for LLM and agent applications.

llms.txtobservabilitytracingevaluationlangchainmonitoringdebugging

verified

Details
About

LangSmith is LangChain's commercial platform for observing, debugging, evaluating, and improving LLM applications in development and production. It captures every LLM call, tool invocation, and chain step into traces that developers can replay, inspect token-by-token, and compare across prompt or model changes.

Core workflows it supports: inspecting why an agent made a particular decision, A/B testing prompts against datasets, running evals (LLM-as-judge, exact-match, custom), catching regressions before deploying prompt changes, and debugging production incidents with full call graphs.

While built by the LangChain team and deeply integrated with LangChain / LangGraph, LangSmith is framework-agnostic — applications using the raw OpenAI or Anthropic SDKs can send traces to LangSmith with minimal setup.

Features
  • Full trace capture for LLM calls, tool calls, and agent steps
  • Replay and inspect individual traces token-by-token
  • Dataset management and evaluation runs
  • LLM-as-judge and custom evaluator support
  • Production monitoring with latency, cost, and error dashboards
  • A/B comparison across prompt or model variants
  • SDKs for Python, TypeScript, and any OpenAI-compatible client
Best for
  • Debugging why an agent produced an unexpected output
  • Regression testing prompts before shipping a new version
  • Evaluating model/prompt changes against a held-out dataset
  • Monitoring production LLM applications for drift or failures
  • Sharing traces with teammates for collaborative debugging
  • Cost attribution across agents, users, or feature flags
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

LangChain

NoveumNew

Reliability platform for production AI agents that combines tracing, calibrated LLM-as-judge evaluation, scenario simulation and guardrails.

llms.txtobservabilitytracingevaluationagentsvoice-agentsmcp

verified

Details
About

Noveum is a hosted reliability platform for teams running AI agents and LLM applications in production, including voice agents. Its open-source tracing SDK, NovaTrace (noveum-trace), records every LLM call, tool call, RAG step and agent hop as a structured trace. It integrates with LangChain, LangGraph, LiveKit, Pipecat and CrewAI.

On top of those traces, NovaEval scores chat agents with a library of 100+ calibrated LLM-as-judge scorers and returns the reasoning behind each verdict. NovaSynth simulates voice (SIP) and chat scenarios with personas, interruptions and virtualized tools to test agents end to end. NovaGuard (in beta) enforces policies on live traffic. NovaPilot turns failing evaluations into proposed fixes delivered as pull requests.

A remote MCP server lets MCP-capable coding assistants and chat clients work with Noveum traces, datasets and evaluations over an authenticated connection. An enterprise tier offers on-prem and self-hosted deployment. The hosted platform is proprietary; the tracing SDK is public.

Features
  • Open-source tracing SDK for LLM calls, tools, RAG steps and agent hops
  • Integrations with LangChain, LangGraph, LiveKit, Pipecat and CrewAI
  • Library of 100+ calibrated LLM-as-judge scorers with per-verdict reasoning
  • Voice (SIP) and chat scenario simulation with tool virtualization
  • Real-time guardrails on live traffic (beta)
  • Remote MCP server for AI clients
  • On-prem and self-hosted enterprise deployment
Best for
  • Investigating failed or slow agent runs from production traces
  • Regression-testing chat and voice agents before a release
  • Monitoring voice agents built on LiveKit or Pipecat
  • Letting a coding assistant query traces and evaluations over MCP
License

freemium

Deployment

saas

Platforms
  • web
Maintainer

Noveum

OpikNew

Open-source (Apache-2.0) tracing, evaluation and prompt optimization for LLM apps, RAG and agents, from Comet — self-host or cloud.

llms.txtobservabilitytracingevaluationopen-sourcemcp

★ 22.4kApache-2.02.2.89 · 2026-10-05verified

Details
About

Opik is Comet's open-source platform for debugging, evaluating and monitoring LLM applications, RAG systems and agentic workflows. It records traces of LLM calls, tools and agent steps. It runs automated evaluations (LLM-as-judge and heuristic metrics) against datasets and experiments, and provides production dashboards and prompt optimization.

It ships an MCP server so coding assistants can read traces, find failing ones and check fixes against real data, plus a built-in assistant ("Ollie") for analysing traces. Opik is Apache-2.0 licensed and can be self-hosted or used as a managed service on Comet.

Features
  • Tracing for LLM calls, tools, RAG steps and agents
  • Automated evaluations, datasets and experiments
  • Production monitoring dashboards
  • Prompt management and optimization
  • MCP server for coding assistants
  • Self-hosted (Apache-2.0) or Comet-hosted
Best for
  • Open-source alternative to hosted LLM observability
  • Regression-testing prompts and agents against datasets
  • Letting a coding agent debug from production traces over MCP
License

open-source

Deployment

self-hosted

Platforms
  • linux
  • macos
  • windows
  • web
Maintainer

Comet

Weights & Biases

ML experiment tracking (W&B Models) and LLM/agent tracing and evaluation (W&B Weave), now part of CoreWeave Forge.

llms.txtstub

★ 11.3kMITv0.30.0 · 2026-09-09verified

Details
License

freemium

Deployment

saas

Maintainer

CoreWeave

Evaluation 5

BraintrustNew

Evals and observability platform for agents that traces production, runs evaluations and catches regressions before release.

llms.txtstub

verified

Details
License

commercial

Deployment

saas

DeepEval

Open-source, pytest-style LLM evaluation framework with 50+ metrics for agents, RAG and chatbots, plus CI/CD integration.

llms.txtstub

verified

Details
License

open-source

Deployment

library

InspectNew

Open-source (MIT) framework for frontier LLM and agent evaluations from the UK AI Security Institute — 200+ pre-built evals, sandboxing and a log viewer.

llms.txtevaluationevalsagentsbenchmarkssandboxaisipython

★ 2.9kMITlast push verified

Details
About

Inspect is a framework for frontier AI evaluations developed by the UK AI Security Institute and Meridian Labs. Evaluations are built from composable datasets, solvers (including agents), tools and scorers. Over 200 pre-built benchmark implementations are ready to run on any supported model.

It supports agent evaluations with built-in agents, multi-agent primitives, and the ability to run external agents such as Claude Code, Codex CLI and Gemini CLI. Tool calling covers custom and MCP tools plus built-in bash, Python, text editing, web search, browsing and computer tools. Untrusted model code runs in sandboxes (Docker, Kubernetes, Modal, Proxmox, Vagrant and others). Inspect View and a VS Code extension help monitor and debug runs.

Features
  • Composable datasets, solvers, tools and scorers
  • 200+ pre-built evaluations
  • Agent evals, including external agents (Claude Code, Codex CLI, Gemini CLI)
  • MCP and built-in bash, python, web, browser and computer tools
  • Sandboxing via Docker, Kubernetes, Modal, Proxmox, Vagrant and extensions
  • Inspect View log viewer and VS Code extension
Best for
  • Safety and capability evaluations of frontier models
  • Benchmarking coding and agentic tasks in sandboxes
  • Reproducible eval suites shared across teams
License

open-source

Deployment

library

Platforms
  • linux
  • macos
  • windows
Maintainer

UK AI Security Institute

Promptfoo

Open-source (MIT) CLI for evaluating, red-teaming and security-testing LLM apps and agents from config files — now part of OpenAI.

llms.txtstub

★ 25.7kMIT0.123.1 · 2026-09-18verified

Details
License

open-source

Deployment

cli

Maintainer

Promptfoo (OpenAI)

Ragas

Open-source evaluation framework for LLM apps — RAG pipelines, agents and workflows — with objective metrics and test-data generation.

llms.txtstub

★ 15.9kApache-2.0v0.4.3 · 2026-01-13verified

Details
License

open-source

Deployment

library

Training & Fine-tuning 7

AI Toolkit (Ostris)

All-in-one open-source training suite (GUI or CLI) for LoRAs and fine-tunes of image and video diffusion models — FLUX.1/FLUX.2, Qwen-Image, Z-Image, SDXL, Wan 2.x, LTX-2 and more.

no llms.txt yettrainingloradiffusionfine-tuningfluxsdxlostrisai-toolkit

★ 12.2kMITlast push verified

Details
About

AI Toolkit by ostris is a widely used open-source framework for training LoRA adapters and full fine-tunes on diffusion image models. It supports FLUX.1 and FLUX.2 (including Klein), Qwen-Image, Z-Image, SDXL, SD1.5, and video models such as Wan and LTX-2, and newer architectures as they land — often within days of a new model's release.

Where generic HuggingFace fine-tuning is model-agnostic but optimized for neither, AI Toolkit is purpose-built for diffusion workflows: handles VAE encoding, text-encoder freezing, rank- constrained LoRA training, sample-during-training, cosine learning-rate schedules, and the dataset quirks (captions, aspect bucketing, regularization images) that matter for diffusion- specific quality.

The YAML-configured job format makes training runs reproducible and shareable. Common recipes (character LoRA, style LoRA, concept LoRA, full fine-tune) ship as working examples. Runs on consumer GPUs (16GB+ for LoRA, 24GB+ for full fine-tune) via bitsandbytes and 8-bit optimizers.

Critical tool for anyone producing image assets at scale on diffusion models — game asset pipelines, character consistency for video projects, branded content generation, or personal LoRAs for creative work.

Features
  • LoRA, LoHA, LoCon, and full fine-tuning support
  • Covers FLUX.1/FLUX.2 (incl. Klein), Qwen-Image/Edit, Z-Image, HiDream, Chroma, SDXL, SD1.5
  • YAML-configured training jobs for reproducibility
  • Day-one support for new model architectures
  • Sample-during-training for real-time quality monitoring
  • Aspect-ratio bucketing and caption handling
  • 8-bit optimizers for consumer GPU training
  • Active community and maintainer
  • Video model training (Wan 2.1/2.2, LTX-2.x) and a web GUI alongside the CLI
Best for
  • Training character LoRAs for consistent figures across generations
  • Style LoRAs matching brand or art-direction requirements
  • Full fine-tunes on curated datasets for specialized domains
  • Training LoRAs for video generation workflows (WAN, LTX character consistency)
  • Reproducing published research that depends on diffusion fine-tuning
License

open-source

Deployment

library

Platforms
  • linux
  • windows
  • macos
Maintainer

ostris

Axolotl

YAML-configured fine-tuning framework supporting LoRA, QLoRA, full FT, DPO, and most modern LLM architectures.

no llms.txt yetfine-tuningtrainingaxolotlloradpoyamlpython

★ 12.5kApache-2.0v0.20.0 · 2026-09-30verified

Details
About

Axolotl is a full-featured open-source fine-tuning framework that wraps HuggingFace Transformers, PEFT, and TRL into a YAML-configured CLI. Users define their training run (model, dataset, hyperparameters, adapter type) in a config file and launch with axolotl train config.yml.

Compared to Unsloth's speed focus, Axolotl prioritizes breadth: broader architecture support (Llama, Qwen, Mistral, Gemma, Phi, DeepSeek, Yi, Cohere, StableLM, and more), more training methods (LoRA, QLoRA, full FT, ReLoRA, continued pretraining, DPO, KTO, ORPO, GRPO), and richer dataset formats (SharePT, Alpaca, ChatML, raw, completion, many more).

It's a common choice for production fine-tuning pipelines where reproducibility and config-as-code matter more than squeezing the last percent of speed.

Features
  • YAML-configured training runs (no code required for common recipes)
  • Supports 20+ base model architectures
  • LoRA, QLoRA, full FT, ReLoRA, DPO, KTO, ORPO, GRPO
  • Multi-GPU / multi-node via DeepSpeed, FSDP2 and N-D (tensor, context, expert) parallelism
  • FlashAttention and sample packing for efficiency
  • Direct integration with WandB and MLflow for tracking
  • Docker images for reproducible training environments
  • GGUF export via axolotl export for llama.cpp / Ollama
Best for
  • Production fine-tuning pipelines with config-as-code
  • Multi-GPU / multi-node training runs
  • Preference tuning via DPO, KTO, ORPO, GRPO
  • Continued pretraining on domain corpora
  • Teams wanting reproducible training via version-controlled YAMLs
License

open-source

Deployment

library

Platforms
  • linux
Maintainer

Axolotl AI

LlamaFactory

WebUI-based fine-tuning framework supporting 100+ models with LoRA, QLoRA, DPO, and more.

no llms.txt yetstub

★ 75.3kApache-2.0v0.9.5 · 2026-05-30verified

Details
License

open-source

Deployment

library

MS-Swift

ModelScope's training and deployment framework for 600+ LLMs and 400+ multimodal models — SFT, GRPO-family RL, DPO, Megatron parallelism.

no llms.txt yetstub

★ 15.8kApache-2.0v4.5.3 · 2026-09-08verified

Details
License

open-source

Deployment

library

TinkerNew

Training API from Thinking Machines — write the training loop in Python, run LoRA fine-tuning and RL on open-weight models from 1B to 1T+ parameters on managed GPUs.

llms.txttrainingfine-tuningrlloraapi

★ 4.2kApache-2.0v0.5.7 · 2026-09-03verified

Details
About

Tinker is a training API for researchers and developers doing LLM post-training. You write a simple loop that runs on a CPU-only machine, containing your data or RL environment, loss function and evals. You call forward_backward, optim_step, sample and save_state, and Tinker executes the exact computation efficiently on distributed GPUs, handling hardware failures transparently.

It fine-tunes dense and mixture-of-experts open-weight models from 1B to 1T+ parameters, including vision-language models, using LoRA rather than full fine-tuning. Trained weights can be downloaded and served elsewhere. The open-source Tinker Cookbook provides recipes for SFT, RL and distillation.

Features
  • Low-level primitives (forward_backward, optim_step, sample) with full control of the loop
  • Built-in and custom losses (cross-entropy, importance sampling, PPO, CISPO, DRO)
  • LoRA fine-tuning of dense and MoE models from 1B to 1T+ parameters, including VLMs
  • Managed distributed training with fault tolerance
  • Downloadable trained weights
  • Open-source cookbook of recipes
Best for
  • RL and SFT research without managing GPU clusters
  • Custom post-training algorithms on large open-weight models
License

commercial

Deployment

saas

Platforms
  • web
Maintainer

Thinking Machines Lab

TRL (HuggingFace)

Hugging Face's post-training library for transformer LLMs — SFT, GRPO, RLOO, DPO, KTO, reward modeling and distillation, with vLLM, PEFT and DeepSpeed integration.

llms.txtstub

★ 19.4kApache-2.0v1.14.1 · 2026-09-29verified

Details
License

open-source

Deployment

library

Unsloth

Open-source app and library to run and fine-tune models locally — about 2x faster training with ~70% less VRAM for LLMs, diffusion, TTS and embedding models.

llms.txtfine-tuningtrainingloraqloraunslothtritonpython

★ 77.2kApache-2.0v0.1.902-beta · 2026-10-01verified

Details
About

Unsloth is a fine-tuning library that rewrites HuggingFace's training code with hand-tuned Triton kernels and careful memory management to achieve roughly 2x speedups and 70% memory savings versus vanilla Transformers + PEFT. It supports LoRA, QLoRA, full fine-tuning, and continued pretraining across Llama, Qwen, Mistral, Gemma, Phi, and most modern architectures.

The project is known for free, working Colab notebooks that take users from zero to a fine-tuned model in under an hour — genuine entry-point material for the LLM fine-tuning community. It's compatible with the HuggingFace ecosystem: PEFT adapters, Trainer API, TRL DPO, datasets, and every output format (LoRA, merged, GGUF, AWQ) downstream tools expect.

The core package is Apache-2.0; some optional components (such as Unsloth Studio) are AGPL-3.0.

Features
  • 2x faster LoRA, QLoRA and full fine-tuning vs vanilla Transformers, ~70% less VRAM
  • Unsloth Desktop (native app), Unsloth Studio (web UI) and Unsloth Core (Python library)
  • Trains LLMs, vision, diffusion, TTS and embedding models; RL via GRPO and TRL integration
  • Multi-GPU and NVIDIA / AMD / Intel GPU, CPU and Vulkan support
  • Export to GGUF, NVFP4, FP8, merged or LoRA-only formats
  • Data Recipes for building datasets from PDFs, CSVs and DOCX files
  • Free Colab notebooks for common fine-tuning recipes
Best for
  • Fine-tuning LLMs on consumer GPUs (24GB and below)
  • Cost-effective domain adaptation for internal applications
  • Teaching / learning fine-tuning with working starter notebooks
  • Rapid iteration on fine-tunes before scaling to multi-GPU
  • Reducing training costs on cloud GPU providers
License

open-source

Deployment

library

Platforms
  • linux
  • windows
  • macos
Maintainer

Unsloth AI

Web Search for Agents 7

Exa

Neural search API built for AI agents — semantic search across the web with content retrieval.

llms.txtstub

verified

Details
License

freemium

Deployment

saas

FirecrawlNew

Web data API that searches, scrapes and interacts with the web and returns clean Markdown or structured data for agents.

llms.txtstub

★ 188.8kAGPL-3.0v2.11.0 · 2026-06-19verified

Details
License

open-source

Deployment

saas

ParallelNew

Web APIs for AI agents — Search, Extract, Task / Deep Research, FindAll and Monitor — returning LLM-optimized excerpts.

llms.txtsearch-apiweb-searchdeep-researchagents

verified

Details
About

Parallel offers web-access APIs designed for AI agents rather than humans. Search returns LLM-optimized excerpts for natural-language objectives. Extract pulls clean content from URLs. The Task API runs deep research and data enrichment. FindAll discovers entities matching criteria, and Monitor watches the web for changes.

The docs include migration guides from Exa, Tavily and SERP APIs and an evaluation guide for comparing search quality inside an agent.

Features
  • Search API with LLM-optimized excerpts and search modes
  • Extract API for page content
  • Task API for deep research and enrichment
  • FindAll entity discovery and Monitor change tracking
  • Responses / chat-style endpoint
Best for
  • Grounding agents in fresh web results
  • Automated deep research and enrichment pipelines
  • Monitoring sites or topics for changes
License

commercial

Deployment

saas

Platforms
  • web
Maintainer

Parallel Web Systems

Perplexity API

Perplexity API Platform — Agent API for web-grounded answers with citations, plus Search, Router and Embeddings APIs, a CLI and an MCP server.

llms.txtstub

verified

Details
License

commercial

Deployment

saas

SerpAPI

Scraping API for Google, Bing, DuckDuckGo and 15+ search engines — structured JSON results.

llms.txtstub

verified

Details
License

commercial

Deployment

saas

Tavily

Web access layer for AI agents (by Nebius) — real-time search, extraction, crawl, map and cited research via API, CLI and MCP.

llms.txtstub

verified

Details
License

freemium

Deployment

saas

Maintainer

Tavily (Nebius)

OCR & Document Parsing 4

Docling

Open-source document conversion toolkit (LF AI & Data, started at IBM Research) — PDF, Office, HTML, images, audio and more into Markdown/JSON for RAG and agents, with VLM and MCP support.

no llms.txt yetstub

★ 68.4kMITv2.133.0 · 2026-10-03verified

Details
License

open-source

Deployment

library

Marker

Fast, accurate document conversion (PDF, images, Office, HTML, EPUB) to Markdown, JSON, chunks or HTML — tables, equations and structure preserved; code Apache-2.0, weights OpenRAIL-M with commercial limits.

no llms.txt yetstub

★ 40.2kApache-2.0v2.0.0 · 2026-07-20verified

Details
License

open-source

Deployment

library

Maintainer

Datalab

ReductoNew

Agentic document platform — classify, parse, extract, split and edit documents into LLM-ready content and structured JSON via API, CLI or MCP.

llms.txtocrdocument-parsingextractionragmcp

verified

Details
About

Reducto is a commercial document-processing platform for AI teams. Its APIs cover the whole document lifecycle. Classify routes documents by type. Parse produces layout-aware text, tables and figures with bounding boxes and RAG-ready chunking. Extract pulls schema-defined fields into JSON. Split separates bundled documents, and Edit fills forms and modifies templates. These steps can be chained into single-call pipelines.

Reducto Studio provides a UI for building document workflows. A CLI and an MCP server let coding agents call the platform directly.

Features
  • Layout-aware parsing with tables, figures and bounding boxes
  • Schema-based structured extraction
  • Classification, splitting and document editing / form filling
  • Composable multi-step pipelines
  • Reducto Studio, CLI and MCP server
  • Large-file upload (up to 5GB)
Best for
  • Intake pipelines for invoices, contracts and medical records
  • RAG ingestion of complex PDFs
  • Automated form filling
License

commercial

Deployment

saas

Platforms
  • web
Maintainer

Reducto

Unstructured

Document ETL for GenAI — the open-source unstructured library plus the hosted Unstructured Transform API/MCP server, turning 65+ file types into RAG-ready elements and structured data.

llms.txtstub

★ 15.5kApache-2.00.27.10 · 2026-09-27verified

Details
License

freemium

Deployment

library

Deployment & Hosting 11

BasetenNew

Inference and training platform — hosted Model APIs (OpenAI- and Anthropic-compatible), dedicated deployments of your own models, and fine-tuning on production GPUs.

llms.txtinferencedeploymentgputrainingopenai-compatibleanthropic-compatiblesaas

verified

Details
About

Baseten runs hosted models, deploys custom models and trains models on production GPU infrastructure. Model APIs expose high-performance LLMs through OpenAI- and Anthropic-compatible endpoints, with reasoning control, vision, server-side web search and coding-agent setups (Claude Code, Codex CLI, OpenCode).

For your own models, Baseten manages containers, GPU capacity across clouds and regions, autoscaling and observability, and its inference engines optimise supported architectures. It also covers audio (transcription, speech generation, diarization), async inference, structured outputs and function calling. Training uses Loops or your own training jobs, and the resulting checkpoints can be served on the same platform.

Features
  • Model APIs with OpenAI- and Anthropic-compatible endpoints
  • Dedicated deployments of open-source, fine-tuned or custom models
  • Autoscaling across clouds and regions; SSH into deployments
  • Training via Loops and Training Jobs, then serve the checkpoints
  • Speech and audio inference; async inference; structured outputs and tool calling
  • Coding-agent integrations; docs MCP server and agent skill
Best for
  • Serving custom or fine-tuned models in production without running GPU infrastructure
  • Pointing coding agents at hosted open models
  • Low-latency voice and transcription pipelines
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Baseten

Cerebras InferenceNew

Wafer-scale inference cloud with an OpenAI-compatible API — thousands of tokens per second on open models, plus dedicated endpoints for custom weights.

llms.txtinferencecerebraswafer-scalelow-latencyopenai-compatiblesaas

verified

Details
About

Cerebras Inference serves models on Cerebras' wafer-scale systems. Cerebras says this is up to 30x faster than GPU systems. The shared API's model catalog lists gpt-oss-120b at about 3,000 tokens/s and Qwen 3.8 27B at about 1,850 tokens/s.

The API is OpenAI-compatible and supports streaming, reasoning controls, structured outputs, tool calling, prompt caching, image inputs, batch jobs and service tiers. Dedicated Inference reserves capacity for an organisation, accepts custom model weights through a management API, and adds predicted outputs and Prometheus-compatible metrics. New accounts get $5 of free trial credits; there is also pay-as-you-go and enterprise pricing. It is also available through AWS Marketplace, Hugging Face and Vercel.

Features
  • OpenAI-compatible API on wafer-scale hardware
  • About 3,000 tok/s on gpt-oss-120b, about 1,850 tok/s on Qwen 3.8 27B (shared tier)
  • Reasoning, structured outputs, tool calling, prompt caching, image inputs
  • Batch API and service tiers
  • Dedicated Inference with custom weights, predicted outputs and metrics
  • Free trial credits, pay-as-you-go and enterprise plans
Best for
  • Agentic coding and agent loops where latency compounds
  • Real-time voice and interactive apps
  • Serving custom open-weight models on reserved capacity
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Cerebras

Fireworks AI

Training and inference platform for open models — serverless and dedicated GPU deployments, fine-tuning, and Fireworks Nexus model routers for coding agents.

llms.txtinferencefireworksserverlessfastopenai-compatiblefine-tuning

verified

Details
About

Fireworks AI is a production inference platform specializing in serving open-source LLMs at low latency and high throughput. Its inference stack (in-house optimized, built on FireAttention) regularly benchmarks as the fastest provider for DeepSeek-R1 / V3, Llama 3/4, Qwen 3, and other large open-weight models.

Beyond serverless inference, Fireworks offers dedicated deployments (guaranteed capacity, custom models, SLA), fine-tuning as a managed service, and agent-oriented features like structured outputs, function calling, and JSON mode across every hosted model.

It's a common pick for AI products that need open-weight models in production with strict latency requirements — voice agents, real-time coding assistants, high-QPS chat applications.

Features
  • Serverless inference with Standard, Priority and Fast modes; reserved throughput with SLA; US-only option
  • On-demand deployments on dedicated GPUs with autoscaling; custom model upload
  • Training and fine-tuning with managed serving of the results
  • Fireworks Nexus — FireConnect for coding harnesses, FireRouter model routers, per-user spend caps
  • OpenAI-compatible API with streaming, tools, JSON mode, prompt caching
  • Text, vision and embedding models
Best for
  • Production applications with strict latency requirements
  • Real-time AI products (voice agents, live coding, streaming chat)
  • Serving fine-tuned DeepSeek or Llama variants at scale
  • Migrating off OpenAI / Anthropic for cost reasons
  • Multi-model inference behind a single bill and dashboard
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Fireworks AI

Groq

Fast-inference "neocloud" — GroqCloud's OpenAI-compatible API on Groq LPUs alongside NVIDIA accelerated computing.

llms.txtinferencelpugroqsaaslow-latencyopenai-compatible

verified

Details
About

Groq runs GroqCloud, an inference cloud with an OpenAI-compatible API. Groq pioneered its own Language Processing Unit (LPU) silicon. It now runs LPUs alongside NVIDIA accelerated computing as an NVIDIA Cloud Partner, across 13 data centers in North America, Europe, the Middle East and Asia Pacific. It raised a $350M Series A in August 2026.

The API is OpenAI-compatible, so OpenAI-client code can switch with a base-URL change. Latency remains the main selling point, with hundreds to about a thousand tokens per second on hosted models. GroqCloud is common for voice agents, streaming chat and agent loops with many sequential calls.

Features
  • OpenAI-compatible API with streaming, tool use and structured outputs
  • Roughly 280–1,000 tokens/sec on production models
  • Production models include gpt-oss 120B/20B, Llama 3.x and Whisper large-v3 / turbo
  • Preview models such as Qwen 3.8 27B, MiniMax M2.7 and Orpheus TTS
  • Speech-to-text, text-to-speech and vision
  • Batch API for async high-volume jobs
  • Free tier for prototyping
Best for
  • Real-time voice agents and conversational AI
  • Streaming chat UIs where time-to-first-token matters
  • Agent loops with many sequential LLM calls (speed compounds)
  • Live coding assistants (inline autocomplete, fast chat)
  • High-throughput inference pipelines with latency SLOs
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Groq LLC

Mistral AINew

Mistral's API platform and Studio for building, fine-tuning and deploying agents and apps on its models, including open-weight ones.

llms.txtstub

verified

Details
License

commercial

Deployment

saas

Modal

Serverless cloud platform for Python with first-class GPU support — deploy LLMs, training jobs, and batch pipelines from code.

llms.txtserverlessgpupythonmodalinferencetrainingsaas

verified

Details
About

Modal is a serverless compute platform designed for Python workloads that need GPUs: LLM inference, model training, batch image generation, data processing, and long-running AI pipelines. You write functions in Python with decorators that specify container image, GPU type, memory, and schedule — Modal handles the rest: cold starts in seconds, billing per-second, auto-scaling to zero when idle.

The developer experience is the differentiator. Locally-written Python functions deploy to cloud GPUs without leaving the editor: modal run my_script.py executes on an A100 or H100, modal deploy exposes it as an HTTPS endpoint. Popular use cases include hosting vLLM or SGLang for inference, running Unsloth fine-tuning jobs, and dispatching ComfyUI workflows at scale.

Modal is used heavily by AI startups that want cloud GPUs without managing Kubernetes, and by teams that need bursty capacity without reserving hardware.

Features
  • Serverless Python functions on cloud GPUs (A100, H100, L4, L40S, …)
  • Sub-10-second cold starts with container snapshotting
  • {'Auto-scaling': 'scale to zero when idle, scale up under load'}
  • Per-second billing, no reserved capacity required
  • Built-in web endpoints, scheduled jobs, and WebSocket support
  • Volumes, secrets, and distributed dicts for state
  • CPU-only functions also supported for pipeline orchestration
Best for
  • Serverless LLM inference hosting (vLLM, SGLang, TGI)
  • Fine-tuning jobs that need a GPU for a few hours, then shut down
  • Batch image / video generation with ComfyUI or diffusers
  • AI application backends where traffic is spiky
  • Replacing Kubernetes + GPU nodes for AI workloads
  • Webhooks, scheduled jobs, and event-driven AI pipelines
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Modal

Ollama CloudNew

Ollama's hosted inference for large open models — the same Ollama app, CLI and API, plus OpenAI- and Anthropic-compatible endpoints, with a free tier and usage credits.

llms.txtinferenceollamacloudopen-modelsopenai-compatibleanthropic-compatiblesaas

verified

Details
About

Ollama Cloud runs open-weight models on Ollama's hosted infrastructure, so you can use models too large for a laptop without downloading them: DeepSeek V4, GLM 5.x, Kimi K3, MiniMax M3, gpt-oss, Mistral Large 3, Nemotron 3 and others. The Ollama app and CLI switch between local and cloud models by name (for example gemma4:cloud). Coding agents can run against cloud models with ollama launch claude, ollama launch codex or ollama launch opencode.

Apps can also call https://ollama.com/api directly with an API key, or use OpenAI- or Anthropic-compatible clients, with no local install. Ollama says cloud models use the provider's native weights, and prompt and response data are never logged or trained on. Hosting is primarily in the United States, with overflow to Europe and Singapore, through NVIDIA Cloud Partners under no-logging, zero-retention terms.

Plans: Free (starter usage credits and starter models), Pro ($20/month with $60 of credits), Max ($100/month with $300 of credits), Team ($500/month, early access) and Enterprise. Usage is billed per million tokens per model, with cheaper off-peak rates. Running models locally stays free and unlimited, and cloud features can be switched off for local-only use.

Features
  • Hosted open-weight models — DeepSeek, GLM, Kimi, MiniMax, gpt-oss, Gemma, Mistral, Nemotron
  • Same Ollama app, CLI and API for local and cloud models (:cloud tags)
  • Direct API at ollama.com/api with API keys; OpenAI- and Anthropic-compatible clients
  • ollama launch for Claude Code, Codex CLI and OpenCode on cloud models
  • Tool calling tested on supported cloud models; web search
  • Per-million-token pricing with off-peak discounts; Free, Pro, Max, Team and Enterprise plans
  • No logging or training on prompts; US-primary hosting with zero-retention partners
  • Local-only mode via disable_ollama_cloud
Best for
  • Running models too large for local hardware with the same Ollama workflow
  • Coding agents on open models without owning a GPU
  • Mixing local and hosted open models behind one API
  • Teams that want shared billing and model access controls for open models
License

freemium

Deployment

saas

Platforms
  • linux
  • macos
  • windows
  • web
Maintainer

Ollama

Replicate

Run thousands of open-source ML models via simple API calls — image, video, audio, text — with per-second billing.

llms.txtinferencereplicateapiopen-source-modelssaascog

verified

Details
About

Replicate is a cloud platform that hosts open-source ML models behind a unified REST API. Thousands of public models — Stable Diffusion, Flux, SDXL, LLaMA, Whisper, MusicGen, video generation, speech synthesis, CLIP embeddings, vision models, voice cloning — are immediately runnable without provisioning GPUs.

Its differentiator is breadth over depth: where Groq and Fireworks optimize for a curated set of LLMs, Replicate covers the whole open-source ML landscape, including image / video / audio models that inference-only LLM platforms don't touch. Pricing is per-second of GPU time.

Replicate also hosts custom model deploys (package your code + weights into a Cog container, push, get an endpoint) — useful for teams that want Replicate's serving infrastructure for their own fine-tuned or private models.

Replicate is part of Cloudflare (announced November 2025). The API and existing models continue unchanged, with integration into Cloudflare's Developer Platform planned.

Features
  • Unified REST API across thousands of open-source models
  • {'Coverage': 'image, video, audio, text, vision, embeddings'}
  • {'Cog': 'containerize your own model and deploy to Replicate'}
  • Per-second GPU billing, scale-to-zero
  • Webhooks for long-running inference
  • Python, JavaScript, Go, Elixir, Swift clients
  • Integrations with LangChain, LlamaIndex, Vercel AI SDK
Best for
  • Building apps that need multiple model types (text + image + audio)
  • Creative tools — image generation, voice synthesis, video
  • Hosting custom fine-tuned models without infrastructure
  • Prototyping with many different models before committing to one
  • Applications that can tolerate cold-start latencies (~10-30 seconds)
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Replicate (Cloudflare)

Runpod

GPU cloud platform with on-demand instances, serverless endpoints, and a community GPU marketplace — priced for AI workloads.

llms.txtgpu-cloudrunpodserverlessinferencetrainingdeployment

verified

Details
About

Runpod is a GPU cloud provider focused on AI workloads. It offers three deployment modes: Pods (long-running GPU instances, similar to bare-metal rentals), Serverless (auto-scaling API endpoints that bill per-second of inference), and Secure Cloud vs Community Cloud (the latter being excess capacity from third-party hosts at lower prices, with the tradeoff of less guaranteed availability).

Its pricing is consistently among the lowest for consumer and enterprise GPUs (RTX 4090, A100, H100, H200, L4, L40S), and the Serverless tier makes it practical to host inference APIs without paying for idle GPU time. Pre-built templates cover common AI stacks (ComfyUI, Automatic1111, vLLM, SGLang, Ollama), so new instances can launch with a running application in one click.

Runpod is a common pick for indie AI builders, researchers, and teams running bursty inference where Modal / Replicate's managed layer isn't needed.

Features
  • Pods (long-running GPU instances) and Serverless endpoints
  • Community Cloud for lowest-cost GPU rentals
  • {'Pre-built templates': 'ComfyUI, vLLM, Ollama, SGLang, Stable Diffusion'}
  • Consumer (RTX 3090/4090/5090) and enterprise (A100/H100/H200) GPUs
  • Per-second billing on Serverless
  • Network volume storage persistent across pods
  • API and CLI for programmatic pod / endpoint management
  • Global datacenters for latency-sensitive workloads
Best for
  • GPU workloads priced lower than AWS / GCP / Azure
  • Serverless hosting for custom fine-tuned models
  • Running ComfyUI or Stable Diffusion at scale without local GPU
  • Research and experimentation on latest consumer/enterprise GPUs
  • Burst capacity for training or batch inference jobs
  • Deploying vLLM / SGLang behind a stable endpoint
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Runpod

SpaceXAI (xAI)

SpaceXAI (formerly xAI) Grok API for reasoning, code, voice, image and video models, usable with the OpenAI SDK.

llms.txtllm-providerxaispacexaigrokapiopenai-compatiblevoicereasoning

verified

Details
About

SpaceXAI, the company formerly branded xAI, builds the Grok family of models and sells access through a developer API at api.x.ai, with docs at docs.x.ai. The x.ai site and the API documentation now both carry the SpaceXAI name. Its newest model is Grok 4.7.

The API covers text generation with streaming, reasoning and structured outputs, and image understanding. Tools include function calling, code execution and collections search for RAG. The Imagine endpoints handle image generation and editing plus video generation, editing and extension. The Voice endpoints cover speech-to-speech (including SIP phone calls), text-to-speech, speech-to-text and custom voices.

Existing OpenAI-SDK code works too: the quickstart shows the official OpenAI Python and JavaScript SDKs next to xAI's own SDK. For high-volume jobs there are a Batch API and deferred completions. Every documentation page is also available as Markdown by appending .md.

Features
  • Grok model family, newest model Grok 4.7
  • Works with the OpenAI SDKs as well as the native xAI SDK
  • Streaming, reasoning and structured outputs
  • Tools including function calling, code execution and collections search (RAG)
  • Image and video generation and editing (Imagine)
  • Voice APIs covering speech-to-speech, text-to-speech, speech-to-text and custom voices
  • Batch API and deferred completions for asynchronous workloads
  • Markdown versions of every docs page for agent ingestion
Best for
  • Agent and coding workloads that want a Grok model as a primary or fallback provider
  • Real-time voice agents, including phone calls over SIP
  • Multimodal apps that need text, image and video generation from one provider
  • Multi-provider stacks that already use the OpenAI SDK
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

SpaceXAI

Together AI

Serverless inference for 200+ open-source models with OpenAI-compatible API — low latency, competitive pricing.

llms.txtinferencetogetherserverlessopenai-compatiblefine-tuning

verified

Details
About

Together AI is a serverless inference platform for open-source LLMs. It hosts 200+ models including the Llama, Qwen, DeepSeek, Mistral, Gemma, and Mixtral families, along with image models (FLUX, SD3) and embedding models — all behind a unified OpenAI-compatible API.

Together's positioning is "run open-source at production speed without operating GPUs." Its inference stack is tuned on top of vLLM and in-house optimizations, delivering competitive tokens-per-second and cost-per-million-tokens among the serverless providers (Fireworks, Groq, DeepInfra, Replicate).

Together also hosts fine-tuning jobs (upload dataset → get a fine-tuned checkpoint → serve it on the same platform) and dedicated endpoints for teams needing isolated capacity. Their open-source together Python client and OpenAI compatibility make migration from OpenAI trivial.

Features
  • Serverless inference for 200+ open-source models
  • OpenAI-compatible API with streaming, tool use, structured outputs
  • Per-million-tokens pricing with tiered rate limits
  • Managed fine-tuning (LoRA and full FT)
  • Dedicated endpoints for enterprise isolation
  • Image generation (FLUX, Stable Diffusion 3)
  • Embedding and reranking models
Best for
  • Production applications needing open-source LLMs without GPU ops
  • Cost-optimized inference as an OpenAI alternative
  • Fine-tuning pipelines with managed training + serving
  • Multi-model architectures (chat + embeddings + image in one bill)
  • Teams wanting open-weight model access without self-hosting
License

commercial

Deployment

saas

Platforms
  • linux
  • macos
  • windows
Maintainer

Together AI

Desktop Applications 5

Claude Desktop

Anthropic's desktop app for Claude — chat, Cowork for long-running agentic tasks, and Claude Code in one app, with connectors (MCP), skills and plugins.

llms.txtdesktop-appclaudeanthropicmcpskillsagentassistant

verified

Details
About

Claude Desktop is Anthropic's native app for macOS and Windows, with a Linux beta for Ubuntu 22.04+ and Debian 12+. It combines three surfaces in one window. Chat is for conversations. Cowork is for longer agentic work that can keep running after you close your laptop; on Linux it runs tasks in a local QEMU/KVM virtual machine. Claude Code adds parallel coding sessions, visual diff review, and an integrated terminal and editor. In September 2026 Anthropic began merging chat and Cowork into a single Claude experience on Pro and Max plans.

Claude can be extended with connectors (MCP servers), skills and plugins, and it can use a built-in browser and, with permission, your computer. It runs on Anthropic's Claude model families (Mythos, Fable, Opus, Sonnet, Haiku). End-user documentation is indexed at claude.com/docs/llms.txt and the Help Center at support.claude.com/llms.txt.

Features
  • Native apps for macOS and Windows; Linux beta (Ubuntu/Debian, x86_64 and arm64)
  • Chat, Cowork (long-running agentic tasks, scheduled tasks) and Claude Code in one app
  • Connectors (MCP servers), skills and plugins
  • Built-in browser and computer use in Cowork
  • Claude Code desktop — parallel sessions, diff review, integrated terminal and editor
  • Claude Docs, Slides and Design created inside conversations (beta, paid plans)
  • Claude Mythos, Fable, Opus, Sonnet and Haiku model families
Best for
  • Daily AI assistant with access to local files and tools via MCP
  • Power-user chat with full tool use (browser, shell, APIs)
  • Workflow automation via custom skills
  • Research, writing, and coding assistance at the desktop
  • Users who want Claude's capabilities beyond the web UI
License

commercial

Deployment

local

Platforms
  • linux
  • macos
  • windows
Maintainer

Anthropic

OpenClaw

Open-source (MIT) personal AI assistant that runs on your own machine and answers you in the chat apps you already use — one self-hosted Gateway, any model.

llms.txtpersonal-assistantagentself-hostedgatewaychatwhatsapptelegrammcpskills

★ 391.4kMITv2026.9.8 · 2026-10-03verified

Details
About

OpenClaw is an open-source personal AI assistant. You run one Gateway process on your own computer or server, connect it to the messaging apps you already use — WhatsApp, Telegram, Slack, Discord, Signal, iMessage, Microsoft Teams, Matrix and many more — and talk to your agent from anywhere. The agent acts for you: it can work through email and calendar, run scheduled automations, drive a browser, and use skills and plugins.

It is model-agnostic. Bundled provider plugins cover the major model APIs, and custom providers point it at local inference servers. Skills and plugins are shared through ClawHub. Desktop apps for macOS, Windows and Linux install the Gateway together with a chat and configuration Control UI, and iOS and Android companion apps pair as device nodes.

OpenClaw is MIT-licensed and, since July 2026, stewarded by the OpenClaw Foundation, an independent 501(c)(3) non-profit. There is no paid or hosted tier, and default telemetry is limited to a version check you can turn off. Its creator, Peter Steinberger, joined OpenAI in 2026 and continues to lead the project.

Features
  • Self-hosted Gateway with 20+ chat channels (WhatsApp, Telegram, Slack, Discord, Signal, iMessage, Teams, Matrix, …)
  • Any model provider via bundled provider plugins, custom endpoints and local inference servers
  • Skills and plugins distributed through the ClawHub marketplace
  • Automations — scheduled jobs, webhooks, Gmail Pub/Sub and IMAP triggers, standing orders
  • Long-term memory, multi-agent routing and subagents
  • Browser control, voice, and paired device nodes (camera, screen, location)
  • MCP servers and an A2A channel for external agents
  • Desktop apps for macOS, Windows and Linux; iOS and Android companion apps; browser Control UI
Best for
  • A personal assistant you message from your phone that runs on hardware you control
  • Recurring chores (inbox, calendar, reports) handled by scheduled automations
  • A shared team assistant in Slack, Teams or Discord without a hosted vendor in the middle
  • Running a personal agent on local or self-chosen models
License

open-source

Deployment

local

Platforms
  • linux
  • macos
  • windows
  • ios
  • android
Maintainer

OpenClaw Foundation

OrkasNew

Open-source multi-agent desktop app — a Commander plans a goal and dispatches specialist agents; bring your own model keys, files stay on your disk.

llms.txtstub

★ 2.2kMITv2026.10.1 · 2026-10-01verified

Details
License

open-source

Deployment

local

Raycast AI

AI assistant built into the Raycast launcher (Mac, Windows, iOS) — agents with projects, screen awareness, scheduled automations and bring-your-own model keys.

llms.txtstub

verified

Details
License

commercial

Deployment

local

smryNew

Agentic news and RSS reader with cited AI summaries and chat, text-to-speech, a searchable library, and API and MCP access.

llms.txtstub

verified

Details
License

freemium

Deployment

saas

Shell Tools 1

ShellGPT

Command-line productivity tool powered by LLMs — generate shell commands, code, and configs.

no llms.txt yetstub

★ 12.3kMIT1.5.1 · 2026-05-06verified

Details
License

open-source

Deployment

cli

128 entries — 70 with full descriptions, 58 stubs. 96 currently publish a working llms.txt.

Why this list exists

Humans increasingly use AI agents (Claude, ChatGPT, Perplexity, local assistants) as their primary search and discovery layer. When a developer asks an agent "what's the best offline voice-to-text tool for 2026?", the answer depends on what the agent can find, read, and cite. Tools that publish a well-structured llms.txt (spec by Jeremy Howard) are easier for agents to index, summarize, and recommend.

Most existing AI-tool directories are optimized for Google SEO (JavaScript-rendered, paywalled, affiliate-heavy). This one is optimized for agent retrieval: plain Markdown, structured YAML entries, a canonical llms.txt and llms-full.txt at the repo root, MIT license, no tracking.

Related work

  • SecretiveShell/Awesome-llms-txt — index of llms.txt URLs for agents to ingest as seed data. If you're looking for a raw feed of every llms.txt on the public internet to wire into RAG, look there. This repo takes the complementary angle: curated tools, categorized and described, aimed at humans picking what to use and agents answering "what should I recommend to my user for X?".
  • llmstxt.org — the spec itself, by Jeremy Howard.

The llms.txt status is re-derived daily from what each URL actually serves — including a check that the response is a real file and not a docs site's catch-all HTML page. See CONTRIBUTING.md.

Notable model families

This list catalogs tools, not models. But agents answering user questions often need both in the same breath, so here are pointers to the most notable model orgs by domain. We deliberately don't link individual model versions — those go stale weekly. Search huggingface.co/<family-name> for each org's current lineup.

Text LLMs

Frontier open weights (April 2026):

  • Qwen (Alibaba) — https://huggingface.co/Qwen
  • DeepSeek — https://huggingface.co/deepseek-ai
  • Kimi (Moonshot) — https://huggingface.co/moonshotai
  • GLM (Zhipu) — https://huggingface.co/THUDM

Established open weights (large install base, not setting the pace):

  • Llama (Meta) — https://huggingface.co/meta-llama
  • Mistral — https://huggingface.co/mistralai
  • Gemma (Google) — https://huggingface.co/google
  • Phi (Microsoft) — https://huggingface.co/microsoft

Closed frontier (API-only — listed for agent awareness):

  • Claude (Anthropic) — https://claude.com
  • Gemini (Google) — https://gemini.google.com
  • Grok (xAI) — https://x.ai

Speech

ASR (speech-to-text): Qwen3-ASR (Alibaba, currently best) · Whisper (OpenAI, widely deployed baseline)

TTS (text-to-speech): Qwen3-TTS (Alibaba, currently best) · F5-TTS · Kokoro · Piper · OpenVoice · Bark (Suno)

Image generation

Frontier (April 2026): Qwen Image (Alibaba) · Z-Image Turbo

Usable locally: Flux.2 Klein 4B (Black Forest Labs) · HiDream

⚠️ License note: Most Flux.2 variants ship under terms that block commercial use. Only Flux.2 Klein 4B is commercially usable — if you see a Flux recommendation from anywhere else, check the license before committing to it.

Legacy (widely used, not advancing): Stable Diffusion · SDXL · Flux.1

Video generation

Ranked by practical 2026 usefulness: WAN (Alibaba, best consistency) · LTX (Lightricks, best character expression, trade-off is chunk-boundary drift) · HunyuanVideo (Tencent, solid third option)

3D generation

Hunyuan3D (Tencent) leads the category in both local weights and hosted API. Both Hunyuan3D and Tripo ship hosted APIs that are meaningfully better than their publicly released local weights: Hunyuan3D 3.1 is the current API version while local users get 2.1, and Tripo's API also substantially outperforms the local weights. If you need best-quality 3D generation, use Hunyuan3D's API. For local-only work, Hunyuan3D 2.1 local weights are still the top pick; Tripo local is weaker.

Embeddings

BGE (BAAI) · Nomic Embed · Jina · E5 (Microsoft)

Contributing

See CONTRIBUTING.md for the YAML schema and submission process. One entry per PR, please. Got a question? Check FAQ.md first.

FAQ

Common questions — why this list exists, who curates it, what stops spam, why some entries are marked missing — are answered in FAQ.md.

License

MIT. Fork it, scrape it, mirror it, agents welcome.

Hear it when it ships

New releases, real benchmarks and the occasional deep-dive. No spam, unsubscribe in one click.

Everything we build

External:   YouTube · GitHub