Apps / AnythingLLM
Managed AnythingLLM hosting.
One server, no app limits.
Chat with your own documents. A private RAG workspace that ingests PDFs and sites, pairs with Ollama or any API model, and keeps data on your server.
What it is
AnythingLLM turns your documents into something you can talk to: a private RAG workspace that ingests PDFs, docs, and websites, embeds them, and answers questions grounded in your own material, with citations.
Minimum RAM
4 GB
Fits on SC-4GB; SC-8GB is the comfortable pick with room to stack.
What we install
Our playbook deploys the official Docker image with persistent storage for your documents and vector data, HTTPS issued and forced, and multi-user mode ready. Point it at the model you prefer: your API key or an Ollama on the same server.
What managed covers
Server updates, snapshots, dual offsite backups that include your document store, monitoring, firewall, and humans on call. Your documents never leave your server unless you choose a cloud model for inference.
The app is $0. The managed server is the only bill, and everything above is part of it.
What can I actually do with AnythingLLM?
Feed it contracts, manuals, research, or an exported wiki and ask questions in plain language; answers cite the source passages. Teams use workspaces to separate clients or projects, each with its own documents and chat history.
Does my data stay private?
Documents and embeddings live on your server, inside our dual offsite backups. If you pair it with Ollama on the same Cube, even inference stays local and nothing ever leaves the machine; with a commercial API key, only the retrieved passages travel to the provider.
How much RAM does AnythingLLM need?
4 GB runs the app and its embedded vector database for serious document sets. Add Ollama for local inference and the pair wants 16 GB or more depending on the model; the cart meter does that math for you.
Which models work with it?
OpenAI, Anthropic, Google, and any OpenAI-compatible endpoint, plus local models through Ollama. Embeddings can run locally either way, which keeps ingestion free and private.
Stack it
AnythingLLM runs well with.
Same server, no extra bill. These are the companions our team installs next to AnythingLLM most often.
Ollama
+16 GBRun open LLMs on your own server. Honest requirement: real models want real RAM, and the meter will route you to the right plan.
LiteLLM
+1 GBOne API gateway for 100+ LLM providers: unified OpenAI-format calls, spend tracking, rate limits, and failover, self-hosted.
Nextcloud
+2 GBYour own file, calendar, and office suite. The self-hosted office anchor, with backups included instead of sold separately.
Frequently asked questions
Is AnythingLLM good for a small team?
Yes: multi-user mode with per-workspace permissions is built into the Docker deployment we ship. One server holds separate workspaces per client or department cleanly.
AnythingLLM or a custom RAG build?
AnythingLLM gets you a working, private document-chat system the day the server provisions. A custom Flowise pipeline wins when you need exotic retrieval logic; plenty of teams run both, and both are $0 here.
What does it cost to run?
The app is $0; a 4 GB managed Cube starts at $40/mo on annual billing. Local embeddings and Ollama inference add no per-token cost; commercial API models bill with their provider.
Sizing an app stack? The VPS sizing calculator recommends a plan from your apps and traffic, with the math shown.
Shared CPU servers
The right home for AnythingLLM and most stacks: 3.0+ GHz vCPU, from $10/mo.
See Shared CPU →Dedicated CPU servers
For CPU-hungry stacks and busy databases: cores that are physically yours.
See Dedicated CPU →All plans and pricing
Every plan, both families, one honest price with everything included.
See pricing →Your server runs. You sleep.
Fully managed hosting from people who have been doing this since 2001.