KoboldCpp on a VPS: how to run a local LLM on CPU or GPU
Contabo has published a step-by-step guide to installing KoboldCpp on a Linux VPS and serving GGUF models via web interface and API.
UptimeMag editorial team · 7 October 2026 · 2 min read

Contabo has published on its blog an operational guide for installing KoboldCpp on a Linux VPS and turning it into a private LLM server, complete with a web interface and an OpenAI-compatible API. It's aimed at anyone wanting to test open-weight models without relying on an external API, keeping their data on their own server.
KoboldCpp is a single-file program that loads models in GGUF format and is built on llama.cpp, the C++ engine behind many local inference tools. The project is maintained by a developer known as Concedo (LostRuins on GitHub) under the AGPLv3 licence. The latest stable release is 1.122.1, published on 26 September 2026, which removes some superfluous logging and introduces a built-in agent with nine tools and a system prompt of roughly 2,000 tokens, according to the release notes on GitHub.
Installation in brief
Contabo's guide uses a test environment with Ubuntu 24.04, 6 vCPUs and 11 GiB of RAM. The steps are: SSH connection, checking resources with free -h && df -h / && nproc, installing curl and tmux, downloading the koboldcpp-linux-x64-nocuda binary (around 130 MB, a CUDA-free build for servers without an NVIDIA GPU), downloading a test model (Qwen2.5-0.5B-Instruct in Q4_K_M quantisation) and launching it inside a tmux session with --usecpu --host 127.0.0.1 --port 5001 --contextsize 2048. Access is then made via an SSH tunnel on port 5001.
How much RAM you need
For models quantised to 4-bit with a 2048-token context, the KoboldCpp wiki indicates the following minimum RAM thresholds:
| Model parameters | Approximate minimum RAM |
|---|---|
| 7 billion | 8 GB |
| 13 billion | 16 GB |
| 30 billion | 32 GB |
| 65 billion | 64 GB |
On CPU, generation speed remains low: the guide describes it as suitable for a single user or a small team, not as a chat service for many users in production. For larger models or faster responses, Contabo offers its own GPU VPS with an NVIDIA RTX PRO 6000 card with 96 GB of VRAM paired with 18 vCPUs and 96 GB of RAM, where a 70-billion-parameter model in 4-bit (roughly 40 GB of weights) fits entirely into VRAM.
The guide recommends always setting --host explicitly: in version 1.122.1, starting the server without this flag would still accept connections on the server's network address, despite showing "localhost" in the log.
Written with the help of artificial intelligence and checked by the editors (EU AI Act, art. 50).