How to Self-Host Ollama and Open WebUI with NVIDIA GPU on Ubuntu 22.04/24.04

Overview

This step-by-step guide shows you how to self-host Ollama with Open WebUI on Ubuntu 22.04/24.04 and use your NVIDIA GPU for fast, private large language model (LLM) inference. You will install the correct NVIDIA drivers, Docker, and NVIDIA Container Toolkit, then deploy Ollama and Open WebUI with Docker Compose. The tutorial also covers updating, backing up models, and troubleshooting common errors such as GPU visibility and port conflicts.

Prerequisites

Before you begin, make sure you have: (1) Ubuntu 22.04 or 24.04 with sudo access, (2) an NVIDIA GPU with at least 6 GB VRAM for medium models (smaller models can work with less), (3) a stable internet connection, and (4) at least 20 GB free disk space for images and model files.

Step 1: Install NVIDIA Driver and Verify CUDA

Use Ubuntu’s built-in tool to install a matching proprietary driver. If Secure Boot is enabled, you may need to enroll a Machine Owner Key (MOK) during installation to load the NVIDIA kernel module.

sudo apt update
sudo ubuntu-drivers install
sudo reboot

After the reboot, confirm the driver is active:

nvidia-smi

You should see a table with your GPU and driver version. If you get an error, check Secure Boot (disable it or enroll the NVIDIA module), then repeat the install.

Step 2: Install Docker Engine and NVIDIA Container Toolkit

Install Docker from the official repository and add your user to the docker group so you can run containers without sudo.

sudo apt-get install -y ca-certificates curl gnupg
sudo install -m 0755 -d /etc/apt/keyrings
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg
echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] \
https://download.docker.com/linux/ubuntu $(. /etc/os-release && echo $VERSION_CODENAME) stable" | \
sudo tee /etc/apt/sources.list.d/docker.list > /dev/null
sudo apt-get update
sudo apt-get install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
sudo usermod -aG docker $USER
newgrp docker

Add NVIDIA Container Toolkit so containers can use your GPU:

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit.gpg
distribution=$(. /etc/os-release; echo $ID$VERSION_ID)
curl -fsSL https://nvidia.github.io/libnvidia-container/stable/$distribution/libnvidia-container.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

Verify Docker can see your GPU:

docker run --rm --gpus all nvidia/cuda:12.5.0-base-ubuntu22.04 nvidia-smi

Step 3: Deploy Ollama and Open WebUI with Docker Compose

We will bind both services to localhost for safety. You can put a reverse proxy in front later for remote access.

mkdir -p ~/ollama-stack && cd ~/ollama-stack
nano compose.yaml

Paste the following compose file (save and exit):

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "127.0.0.1:11434:11434"
    environment:
      - OLLAMA_HOST=0.0.0.0
      - OLLAMA_NUM_PARALLEL=1
    volumes:
      - ollama:/root/.ollama
    gpus: all

  open-webui:
    image: ghcr.io/open-webui/open-webui:latest
    container_name: open-webui
    restart: unless-stopped
    depends_on:
      - ollama
    environment:
      - OLLAMA_API_BASE=http://ollama:11434
    ports:
      - "127.0.0.1:3000:8080"
    volumes:
      - openwebui:/app/backend/data

volumes:
  ollama:
  openwebui:

Start the stack:

docker compose up -d

Pull a model into Ollama (example: a small, fast model):

docker exec -it ollama ollama pull llama3.2:3b

Quick API test:

curl http://127.0.0.1:11434/api/generate \
  -H "Content-Type: application/json" \
  -d '{"model":"llama3.2:3b","prompt":"Say hello in one sentence."}'

Open your browser at http://127.0.0.1:3000 and select Ollama as the provider. Choose the model you pulled and start chatting.

Step 4: Updates and Backups

To update Ollama and Open WebUI to the latest images while keeping your models and data, run:

cd ~/ollama-stack
docker compose pull
docker compose up -d

Back up volumes (models and WebUI data) with a simple tar archive:

docker stop open-webui ollama
docker run --rm -v ollama:/data -v "$PWD":/backup alpine \
  sh -c 'tar czf /backup/ollama-vol.tar.gz -C /data .'
docker run --rm -v openwebui:/data -v "$PWD":/backup alpine \
  sh -c 'tar czf /backup/openwebui-vol.tar.gz -C /data .'
docker start ollama open-webui

Troubleshooting

No CUDA-capable device detected: Ensure the NVIDIA driver is loaded (nvidia-smi works on the host). If Secure Boot is on, enroll the MOK or disable Secure Boot. Confirm the container sees the GPU with the CUDA test image. Re-run: sudo nvidia-ctk runtime configure --runtime=docker and restart Docker.

Compose error: unknown field "gpus": Your Docker Compose is outdated. Update Docker or use: docker run --gpus all ... Alternatively, in compose.yaml, remove gpus: all and start Ollama with: docker run -d --gpus all -p 127.0.0.1:11434:11434 -v ollama:/root/.ollama --name ollama ollama/ollama:latest

Port already in use: Change the host ports in compose.yaml (for example, 127.0.0.1:11435:11434 and 127.0.0.1:3001:8080) and re-run docker compose up -d.

Out-of-memory or slow responses: Choose a smaller or more quantized model (e.g., llama3.2:1b or a Q4 version if available). Limit parallel requests with OLLAMA_NUM_PARALLEL=1. Ensure you have adequate swap configured on the host for large models.

Security Tips

Keep services bound to 127.0.0.1 and place a reverse proxy with TLS in front (Caddy, Traefik, or Nginx) if you need remote access. For Open WebUI, enable authentication in its settings. Restrict firewall rules to only allow your reverse proxy and management IPs. Regularly update images and prune unused layers with docker system prune -af.

Clean Uninstall

To remove the stack and its volumes (this deletes downloaded models and chat data), run:

cd ~/ollama-stack
docker compose down -v

Conclusion

You have a fully private, GPU-accelerated local AI setup with Ollama and Open WebUI running on Ubuntu. This stack is easy to update, simple to back up, and flexible: you can try multiple models, script against the API, or place it behind a secure reverse proxy for team access. With one machine and an NVIDIA GPU, you now own your LLM workflow end to end.

Self-Host Ollama + Open WebUI with NVIDIA GPU on Ubuntu (Docker Compose Guide)

Overview

This guide shows you how to self-host Ollama and Open WebUI on Ubuntu using Docker Compose with NVIDIA GPU acceleration. Ollama makes it easy to run popular local LLMs (like Llama 3, Mistral, Phi, and Qwen), while Open WebUI provides a clean, multi-user chat interface, prompt management, and model switching. By the end, you will have a persistent, GPU-enabled AI stack reachable in your browser, suitable for personal use or a small team.

Prerequisites

- Ubuntu 22.04 or 24.04 (server or desktop), 16 GB RAM recommended.

- An NVIDIA GPU with recent drivers (8 GB VRAM or more recommended for 7B/8B models).

- Docker Engine and the Docker Compose plugin.

- A user with sudo privileges and outbound internet access.

Step 1 — Install Docker and Docker Compose

If Docker is not installed, run:

sudo apt update && sudo apt install -y ca-certificates curl gnupg

sudo install -m 0755 -d /etc/apt/keyrings

curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg

echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] https://download.docker.com/linux/ubuntu $(. /etc/os-release && echo $UBUNTU_CODENAME) stable" | sudo tee /etc/apt/sources.list.d/docker.list > /dev/null

sudo apt update && sudo apt install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin

Verify Docker works: docker version and docker compose version.

Step 2 — Enable NVIDIA GPU in Containers

Install the NVIDIA Container Toolkit so Docker can access your GPU. First, ensure the NVIDIA driver is installed and nvidia-smi works on the host. Then run:

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

curl -fsSL https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \

sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

sudo apt update && sudo apt install -y nvidia-container-toolkit

sudo nvidia-ctk runtime configure --runtime=docker

sudo systemctl restart docker

Test inside a container: docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi. You should see your GPU listed.

Step 3 — Create the Docker Compose file

Make a new folder for the stack and create compose.yml in it:

mkdir -p ~/ai-stack && cd ~/ai-stack

Use this minimal Compose configuration (Ollama + Open WebUI, GPU-enabled, with persistent volumes):

services:
ollama:
image: ollama/ollama:latest
container_name: ollama
ports:
- "11434:11434"
volumes:
- ollama:/root/.ollama
environment:
- OLLAMA_KEEP_ALIVE=24h
gpus: all
restart: unless-stopped

open-webui:
image: ghcr.io/open-webui/open-webui:latest
container_name: open-webui
depends_on:
- ollama
ports:
- "3000:8080"
environment:
- OLLAMA_BASE_URL=http://ollama:11434
- WEBUI_AUTH=True
volumes:
- openwebui:/app/backend/data
restart: unless-stopped

volumes:
ollama:
openwebui:

This setup exposes Ollama on port 11434 (API) and Open WebUI on 3000 (web). Data persists in Docker volumes, so updates do not erase models or chats.

Step 4 — Launch the stack and pull a model

Start both services in the background:

docker compose up -d

Check logs to confirm GPU access and healthy startup:

docker logs -f ollama and docker logs -f open-webui

Pull your first model (example: Llama 3.1 8B) and verify inference:

docker exec -it ollama ollama pull llama3.1:8b

docker exec -it ollama ollama run llama3.1:8b

Open a browser to http://<your_server_ip>:3000, create your admin account, choose the pulled model, and start chatting.

Step 5 — Secure access and basic hardening

Open WebUI has built-in auth. The Compose file sets WEBUI_AUTH=True, which prompts for signup on first visit. After creating the admin user, disable new registrations by adding ENABLE_SIGNUP=False under the open-webui environment and redeploy with docker compose up -d.

If you will expose the UI on the internet, place it behind a reverse proxy with HTTPS. For example, with Caddy on the same host, you can proxy to port 3000 and get automatic TLS:

my-ai.example.com {
reverse_proxy 127.0.0.1:3000
}

Alternatively, use Nginx and a free TLS certificate from Let's Encrypt. Restrict access with IP allowlists or SSO if available.

Step 6 — Useful environment options

- OLLAMA_KEEP_ALIVE: Keeps models warm for faster first-token latency (e.g., 24h).

- WEBUI_AUTH and ENABLE_SIGNUP: Enable auth and control who can create accounts.

- OLLAMA_NUM_PARALLEL: Limit concurrent requests to protect VRAM.

- OPENAI_API_BASE_URL (Open WebUI): Point tools or plugins to Ollama if needed for compatibility layers.

Step 7 — Backup and update strategy

Your chats and models live in Docker volumes (ollama and openwebui). To back them up quickly, stop the stack and archive the volumes:

docker compose down

docker run --rm -v ollama:/data -v $(pwd):/backup busybox tar czf /backup/ollama-vol.tar.gz -C /data .

docker run --rm -v openwebui:/data -v $(pwd):/backup busybox tar czf /backup/openwebui-vol.tar.gz -C /data .

To update, pull the latest images and redeploy:

docker compose pull && docker compose up -d

Troubleshooting

- No GPU visible in containers: confirm nvidia-smi works on the host, that the NVIDIA Container Toolkit is installed, and that gpus: all is present under the Ollama service.

- Port in use: change 11434 or 3000 in compose.yml if conflicts arise.

- Out of memory (VRAM): try a smaller model variant (e.g., 7B/8B quantized like Q4_K_M), or reduce parallel requests. Example pull: ollama pull llama3.1:8b-instruct-q4_K_M.

- Slow first response: increase OLLAMA_KEEP_ALIVE or keep frequently used models loaded.

What you built

You now have a modern, GPU-accelerated local AI stack with Ollama and Open WebUI running on Docker Compose. It is easy to manage, fast to update, and simple to secure behind HTTPS. Add more models, enable extensions, or integrate with automation tools to turn this into a private, production-ready assistant for your workstation or team.

3.

Popular Posts

Install Ollama and Open WebUI on Ubuntu 24.04 with NVIDIA GPU Acceleration (Step-by-Step)

Install Ollama + Open WebUI on Ubuntu 24.04 with NVIDIA GPU Acceleration (Step-by-Step)

Install a Local AI Chatbot on Ubuntu 24.04 with Ollama and Open WebUI (Step-by-Step)

Trending Now

Recovering from Btrfs Boot Failures Using GUI Tools on Fedora

By the end of this guide the reader will be able to identify a Btrfs‑based Fedora installation, boot from a live USB, list and restore snapshots using the graphical utilities btrfs‑assistant and snapper, and verify that the system returns to a functional state without resorting to the command line. Understanding the Btrfs Layout Used by Fedora Fedora Workstation and Fedora KDE install the root filesystem as a single Btrfs partition that contains two default sub‑volumes. One sub‑volume holds the traditional “/” hierarchy, while the second is dedicated to /var/lib/machines . The latter exists to keep container images out of snapshot operations; it remains empty on systems that do not run virtual machines. Because Btrfs stores data in sub‑volumes rather than separate partitions, a snapshot captures the state of an entire sub‑volume at a point in time. The installer (Anaconda) automatically registers these sub‑volumes with the snapper service. Snapper maintains a series of read‑only ...