How to Run Ollama and Open WebUI on a VPS (Private AI Chat Anywhere)

Run your own private ChatGPT-style assistant on a VPS with Ollama and Open WebUI. Which models a CPU server can handle, setup, security and updates.

A chat bubble on top of a server on an AI tools gradient background

This post explains how to run Ollama and Open WebUI on a VPS. Ollama runs open AI models such as Llama, Qwen and Gemma, and Open WebUI puts a ChatGPT-style chat interface on top of it. On your own computer, they only work while that computer is on. On a VPS, your private assistant is available from any browser, phone or app, at any time.

In our list of free software to run LLMs locally, Ollama and Open WebUI are the best pair for exactly this setup. This guide shows how to set them up on a Hostinger VPS with its ready-made Ollama template, which models a CPU-only server can realistically run, and how to keep it secure.

This post contains affiliate links. See our affiliate disclaimer.

Highlights:

  • Private AI chat that you can open from any device.
  • An OpenAI-compatible API for your own apps and automations, with no per-token fees.
  • Ollama and Open WebUI pre-installed with Hostinger’s template.
  • Honest limits: a CPU-only VPS runs small models, slowly.

First, Set Your Expectations

A regular VPS has no GPU. Ollama runs models on the CPU instead, which works but is much slower than a gaming PC or a Mac with Apple silicon, and far slower than ChatGPT or Claude. The size of model you can run is limited by the server’s RAM, because the whole model has to fit in memory.

Model size Examples RAM the model needs (approx.) Good VPS size
1B–4B Llama 3.2 3B, Gemma 3 4B, Qwen 3 4B 2–4GB 8GB RAM (KVM 2)
7B–8B Llama 3.1 8B, Qwen 3 8B 5–6GB 16GB RAM (KVM 4)
14B and up Qwen 3 14B, Gemma 3 12B+ 9GB+ Not practical without a GPU

These figures are for the default 4-bit (Q4) versions that Ollama downloads. Leave a few GB free for the operating system and Open WebUI. Hostinger itself recommends 16GB of RAM for its Ollama template.

When this setup makes sense:

  • You want a private assistant for notes, drafts, summaries or documents that shouldn’t go to a big AI company.
  • You want an always-on API for automations, for example an n8n workflow that classifies emails or summarizes RSS feeds, without paying per token.
  • You want to learn and experiment with open models without keeping your own computer running.

When it doesn’t: if you want the smartest possible answers or fast coding help, a cloud model will be better and, for light use, often cheaper. Open WebUI can connect to cloud APIs too, so you can mix both (see Step 5).

What You Need

  • A VPS with 8GB of RAM or more. Hostinger’s KVM 2 (2 vCPU, 8GB RAM, 100GB NVMe) handles 1B–4B models. For 7B–8B models, KVM 4 (4 vCPU, 16GB RAM, 200GB NVMe) is the comfortable choice, and the extra CPU cores also make replies faster.
  • Disk space for models. Each small model takes 2–5GB, so 100GB is plenty.
  • A domain name (optional), if you want a nice address and HTTPS for the chat interface.

At the time of writing, KVM 2 costs $8.99/month and KVM 4 $12.99/month on a 24-month term, renewing at $14.99 and $28.99/month.

See Hostinger VPS plans

Step 1: Install the Ollama Template

  1. Buy a VPS plan and, when asked for an operating system, choose the Ubuntu 24.04 with Ollama template. It comes with Ollama and Open WebUI already installed.
  2. Pick a server location close to you. Unlike automations, chat is interactive, so lower latency feels better.
  3. Set a strong root password and finish setup.

If you already have a Hostinger VPS, go to VPS, click Manage, open OS & Panel > Operating System, search for Ollama and click Change OS. This erases everything on the server.

Step 2: Open Open WebUI and Create the Admin Account

  1. Open the chat interface from hPanel .
  2. Click Sign up and create an account. The first account created becomes the administrator, so do this right away, before anyone else can find the page.
  3. Go to Admin Panel > Settings and turn off new sign-ups, unless you want others to create accounts. You can still invite users yourself.

Step 3: Download a Model

You can download models from Open WebUI’s model selector or from the command line. To use the command line, connect over SSH (or use the Browser terminal in hPanel) and run:

ollama pull llama3.2:3b
ollama list

Good first choices for a CPU server are llama3.2:3b (fast, general chat), qwen3:4b (strong reasoning for its size) and gemma3:4b (good writing, understands images). On a 16GB server, try llama3.1:8b or qwen3:8b for better answers at lower speed.

To see how fast a model runs on your server, use:

ollama run llama3.2:3b --verbose

After each answer, the eval rate shows the speed in tokens per second. Around 5 tokens per second feels slow but usable for chat; below that, pick a smaller model.

Step 4: Add HTTPS and Your Own Domain

Logging in over plain HTTP sends your password and chats unencrypted, so set up HTTPS before regular use.

  1. Add an A record for a subdomain, such as ai.yourdomain.com, pointing to your VPS’s IP address.
  2. Put a reverse proxy with a free Let’s Encrypt certificate in front of Open WebUI.
  3. Once HTTPS works, block direct access to Open WebUI’s port so that the domain is the only way in (see the security checklist).

Step 5: Use It From Other Apps (Optional)

Ollama provides an OpenAI-compatible API, so many tools that support “custom OpenAI endpoints” can use your models. There’s a catch: Ollama’s API (port 11434) has no authentication, so never expose it directly to the internet.

Safer options:

  • Use Open WebUI’s API instead. Create an API key in your Open WebUI account settings and point apps at Open WebUI, which checks the key before passing requests to Ollama.
  • Run the app on the same server. An n8n instance on the same VPS can call Ollama at http://localhost:11434 without opening any port.
  • Use an SSH tunnel from your own computer: ssh -L 11434:localhost:11434 root@your-server-ip, then use http://localhost:11434 locally.

Open WebUI can also add cloud models (OpenAI, Anthropic through a compatible gateway, OpenRouter and others) under Admin Panel > Settings > Connections, so you can use a local model for private tasks and a cloud model when you need more power, in one interface.

Step 6: Keep Everything Updated

Both projects release updates frequently.

  • Ollama: if Ollama is installed as a system service, rerun the official install script, which upgrades it in place:

    curl -fsSL https://ollama.com/install.sh | sh
  • Open WebUI: it runs in Docker, so pull the new image and recreate the container:

    cd /path/to/compose/directory
    docker compose pull
    docker compose up -d

Take a VPS snapshot in hPanel before updating, so you can roll back if something breaks.

Security Checklist

  • Create the admin account first, then disable open sign-ups.
  • Use HTTPS for Open WebUI; never log in over plain HTTP.
  • Never expose Ollama’s port 11434 to the internet.
  • In Hostinger’s VPS Firewall, allow only SSH, HTTP and HTTPS.
  • Use SSH keys instead of the root password.

Is Hostinger a Good Place for This?

What works well: the Ollama template skips the Docker and Open WebUI setup entirely, and Hostinger’s plans give a lot of RAM for the price, which is exactly what CPU inference needs.

What to watch out for: there are no GPU plans, so larger models are out of reach; the low prices need a 24-month prepayment and renew higher; and support is chat-based with mixed reviews. If you need to run 14B+ models quickly, look at GPU cloud providers instead, and expect to pay much more.

See Hostinger VPS plans

Final Thoughts:

Ollama and Open WebUI on a VPS give you a private, always-on AI assistant and an API for your own automations for a fixed monthly price. Choose a server with enough RAM for the models you want (8GB for 3B–4B models, 16GB for 8B models), set up the admin account and HTTPS right away, and keep Ollama’s API off the public internet. Just don’t expect cloud-level speed or intelligence from a CPU server; use it for what it does well, which is privacy and always-on availability.