Alternatives
Self-hosted alternatives to Together AI
Projects in the registry that can run on your own hardware and serve at least one of the same use cases as Together AI.
- Source
- Registry relations over fetched repository records — the evidence sits on each repository page
- Verified
- Evidence not verified
- Confidence
- High
01How this list was built
- Compared against
- Together AI
- Delivery
- Cloud platforms
- Shared use cases
- Choosing a model, Local LLM, Document Q&A, Agent workflow
- Selection rule
- open source, or self-hostable, and shares a use case
- Candidates found
- 17
02Alternatives
AnythingLLM
Hybrid — hosted or self-hosted
Desktop and server application that turns a document set into a chat workspace, with per-workspace embeddings, multiple model back ends and a built-in vector store.
- Repository
- Mintplex-Labs/anything-llm
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document qa · Local llm
Chroma
Hybrid — hosted or self-hosted
Embedded and server-mode vector store with a small API surface, often used for prototypes and single-node retrieval before a larger engine is justified.
- Repository
- chroma-core/chroma
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document qa
Continue
Hybrid — hosted or self-hosted
Open-source IDE extension for VS Code and JetBrains that connects completion and chat to any model back end, including a local server, under a checked-in configuration file.
- Repository
- continuedev/continue
- Stars
- 35,622
- Licence
- Apache-2.0
- In stacks
- 1
Shares: Local llm
Dify
Hybrid — hosted or self-hosted
Platform for building LLM applications: visual workflow editor, retrieval pipelines, agent tools, prompt management and an API layer. Available self-hosted or as a managed cloud service.
- Repository
- langgenius/dify
- Stars
- 153,448
- Licence
- Other
- In stacks
- 1
Shares: Document qa
Docling
Libraries
Document conversion toolkit that parses PDF, Office and image files into structured Markdown or JSON, preserving reading order, tables and figures for downstream retrieval.
- Repository
- docling-project/docling
- Stars
- 65,519
- Licence
- MIT
- In stacks
- 5
Shares: Document qa
Flowise
Hybrid — hosted or self-hosted
Visual builder for LLM chains and agents, with retrieval nodes, tool calling and an API or embeddable chat widget for the finished flow.
- Repository
- FlowiseAI/Flowise
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document qa
Jan
Hybrid — hosted or self-hosted
Open-source desktop assistant that runs open-weight models locally and can also call remote endpoints. The open alternative to LM Studio for a single-machine deployment.
- Repository
- janhq/jan
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Local llm · Document qa
Kotaemon
Self-hosted projects
Document question-answering application with a configurable retrieval pipeline, inline citations that highlight the source passage, and support for local or hosted models.
- Repository
- Cinnamon/kotaemon
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document qa
LangChain
Frameworks
Library for composing model calls, retrieval and tools into applications, with adapters for most providers and vector stores. A building block, not a deployable product.
- Repository
- langchain-ai/langchain
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document qa
Langflow
Hybrid — hosted or self-hosted
Visual environment for composing LLM pipelines and agents from components, with export to a Python application or an API endpoint.
- Repository
- langflow-ai/langflow
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document qa
LangGraph
Frameworks
Graph-structured agent framework with explicit state, checkpoints and interrupts — the mechanism behind a human approval step rather than a prompt asking for one.
- Repository
- langchain-ai/langgraph
- Stars
- —
- Licence
- unknown
- In stacks
- 1
Shares: Agent workflow
LibreChat
Self-hosted projects
Self-hosted multi-model chat application with authentication, per-conversation model switching, plugins, file upload and an admin configuration file. Familiar interface for staff moving off consumer tools.
- Repository
- danny-avila/LibreChat
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document qa
LiteLLM
Cloud platforms
Gateway that presents one OpenAI-compatible API in front of many providers and local servers, with per-team keys, budgets, rate limits, fallbacks and request logging.
- Repository
- BerriAI/litellm
- Stars
- 57,216
- Licence
- Other
- In stacks
- 3
Shares: Local llm
llama.cpp
Libraries
C++ inference engine for quantised models on CPU, Apple Silicon and GPUs, with a bundled HTTP server. The engine underneath many desktop runtimes and the GGUF quantisation format.
- Repository
- ggml-org/llama.cpp
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Local llm
LlamaIndex
Frameworks
Data framework for retrieval applications: loaders for many document types, indexing and query pipelines, and evaluation helpers for retrieval quality.
- Repository
- run-llama/llama_index
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document qa
LMDeploy
Model servers
Serving toolkit from the InternLM team with its own inference engine and quantisation tooling, and strong coverage of Chinese open-weight model families.
- Repository
- InternLM/lmdeploy
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Local llm
LM Studio
Hybrid — hosted or self-hosted
Desktop application for downloading and running open-weight models locally, with a chat interface and a local OpenAI-compatible server. Windows, macOS and Linux.
- Repository
- none linked
- Stars
- —
- Licence
- unknown
- In stacks
- 1
Shares: Local llm · Document qa
03Repository facts
| Repository | Stars | Health |
|---|---|---|
| Mintplex-Labs/anything-llm | — | Health not measuredunknown |
| chroma-core/chroma | — | Health not measuredunknown |
| continuedev/continueopen-source coding agent | 36k | Maintenance 100%, Releases 60%, Contributors 50%, Issues 74%, Community 91%80 |
| langgenius/difyBuild Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack. | 153k | Maintenance 100%, Releases 60%, Contributors 50%, Issues 94%, Community 100%84 |
| docling-project/doclingGet your documents ready for gen AI | 66k | Maintenance 100%, Releases 80%, Contributors 50%, Issues 85%, Community 96%86 |
| FlowiseAI/Flowise | — | Health not measuredunknown |
| janhq/jan | — | Health not measuredunknown |
| Cinnamon/kotaemon | — | Health not measuredunknown |
| langchain-ai/langchain | — | Health not measuredunknown |
| langflow-ai/langflow | — | Health not measuredunknown |
| langchain-ai/langgraph | — | Health not measuredunknown |
| danny-avila/LibreChat | — | Health not measuredunknown |
| BerriAI/litellmThe fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM] | 57k | Maintenance 100%, Releases 100%, Contributors 50%, Issues 14%, Community 95%83 |
| ggml-org/llama.cpp | — | Health not measuredunknown |
| run-llama/llama_index | — | Health not measuredunknown |
| InternLM/lmdeploy | — | Health not measuredunknown |
04Ask about your own case
Which alternative is right depends on how many people will use it, what the data is and what you can operate. Ask and the answer is researched for your organisation.