RemarkableCloud

Apps / LiteLLM

Managed LiteLLM hosting.
One server, no app limits.

One API gateway for 100+ LLM providers: unified OpenAI-format calls, spend tracking, rate limits, and failover, self-hosted.

What it is

LiteLLM is the traffic controller for AI APIs: one self-hosted gateway that speaks OpenAI format and routes to 100+ providers, with per-key spend tracking, rate limits, retries, and failover. Point every app at one URL and manage the chaos in one place.

Minimum RAM

1 GB

Fits on SC-1GB; SC-2GB is the comfortable pick with room to stack.

What we install

Our playbook deploys the LiteLLM proxy with a PostgreSQL-backed key store, HTTPS issued and forced, and an admin key generated for you. Add provider keys, mint virtual keys per app or teammate, done.

What managed covers

Server updates, snapshots, dual offsite backups, monitoring, firewall, and humans on call. Your provider keys stay in your database on your server.

The app is $0. The managed server is the only bill, and everything above is part of it.

Why put a gateway in front of AI providers?

Because five apps with five hardcoded API keys is how spend surprises and outages happen. LiteLLM gives each app a virtual key with its own budget and rate limit, tracks spend per key, and fails over to a second provider when the first one has a bad day.

What does OpenAI-compatible actually mean?

Any tool that can talk to OpenAI's API can talk to your LiteLLM URL unchanged, while the gateway translates to Anthropic, Google, Mistral, Ollama, or a hundred others behind the scenes. Switching providers becomes a config line, not a code change.

How much server does LiteLLM need?

It is the lightest app in our AI shelf: 1 GB runs the proxy for personal and small-team traffic. It almost always shares a server with the apps that call it, which is exactly what the cart meter is for.

Does LiteLLM see my prompts?

Requests pass through your gateway on your server on the way to the provider you chose; nothing is logged beyond what you enable. That is the point of self-hosting the control plane.

Stack it

LiteLLM runs well with.

Same server, no extra bill. These are the companions our team installs next to LiteLLM most often.

LibreChat

+4 GB

A private ChatGPT-style interface for every major model provider. One clean UI, your API keys, your conversation history on your server.

Flowise

+2 GB

Drag-and-drop builder for LLM apps and agents. Prototype chatbots and pipelines visually, then serve them as APIs from your own server.

PostgreSQL

+1 GB

The engineers’ database. Installed, tuned for your RAM, and backed up twice offsite.

Frequently asked questions

Who should run LiteLLM?

Anyone with more than one AI-consuming app or more than one teammate holding raw provider keys. It is the natural companion to LibreChat, Flowise, and OpenClaw on the same server, all $0 here.

Does the gateway add latency?

Single-digit milliseconds on the same server, which disappears next to model inference time measured in seconds. Failover and retries typically save far more time than the hop costs.

What does LiteLLM hosting cost?

The app is $0; a 1 GB managed Cube starts at $20/mo on annual billing and runs it with room to spare. Provider usage bills with each provider through your own keys.

Sizing an app stack? The VPS sizing calculator recommends a plan from your apps and traffic, with the math shown.

Your server runs. You sleep.

Fully managed hosting from people who have been doing this since 2001.