RemarkableCloud

Apps / Ollama

Managed Ollama hosting.
One server, no app limits.

Run open LLMs on your own server. Honest requirement: real models want real RAM, and the meter will route you to the right plan.

What it is

Ollama runs large language models on your own server: Llama, Mistral, Qwen, and the open-model world behind one clean local API, with your prompts never leaving the machine.

Minimum RAM

16 GB

Fits on SC-16GB; SC-48GB is the comfortable pick with room to stack.

What we install

The playbook deploys the current Ollama with the API bound privately by default, your chosen starter models pulled, memory sized to the plan, and the service supervised. Open WebUI beside it is one more checkbox.

What managed covers

Platform and server updates, daily snapshots plus dual offsite backups, monitoring on memory and the service, and the managed firewall keeping the API off the public internet unless you decide otherwise.

The app is $0. The managed server is the only bill, and everything above is part of it.

What does managed Ollama hosting include?

A working local-model server sized honestly, kept updated, and secured by default, at flat pricing with no per-token anything. The API is yours; the meter does not exist.

What can a CPU server realistically run?

The honest answer: small and mid-size models. 7B-class models run acceptably for chat and drafting on 16 GB; quantized 13B fits with patience; large models want GPUs we do not sell, and we will say so rather than sell you disappointment. For summaries, extraction, and private chat, CPU-served 7B is genuinely useful.

Why run models locally at all?

Privacy and cost shape. Client documents, contracts, and internal data can be processed without leaving your server, which turns entire compliance conversations into a shrug. And a flat server price beats per-token billing the moment usage becomes routine.

What pairs with Ollama?

Open WebUI for a ChatGPT-style interface on top, and n8n for automation that calls the models: summarize inbound email, draft replies, classify tickets, all over localhost, all private. The Private AI kit bundles exactly this.

Stack it

Ollama runs well with.

Same server, no extra bill. These are the companions our team installs next to Ollama most often.

Open WebUI

+2 GB

The chat interface for Ollama. Together they are a private ChatGPT on hardware nobody shares.

n8n

+2 GB

Self-hosted workflow automation: connect your apps and let the flows run 24/7 on a server that is monitored for you.

Frequently asked questions

How much RAM does Ollama need?

16 GB is the honest minimum for comfortable 7B-class use, and the meter enforces it. More RAM means bigger or less-quantized models.

Which models can I pull?

Anything in the Ollama library or importable as GGUF: Llama, Mistral, Qwen, Gemma, Phi, and the rest of the open world.

Is the API exposed?

Not by default: it binds privately, apps on the same server reach it over localhost, and remote access happens through WireGuard or a deliberate firewall grant.

Sizing an app stack? The VPS sizing calculator recommends a plan from your apps and traffic, with the math shown.

Your server runs. You sleep.

Fully managed hosting from people who have been doing this since 2001.