Guides

SGLang on GPU VPS: self-hosting LLMs with Docker

Contabo publishes a guide to installing SGLang via Docker on a GPU VPS: continuous batching, RadixAttention and an OpenAI-compatible API.

UptimeMag editorial team · 5 October 2026 · 2 min read

SGLang su GPU VPS: come servire LLM in self-hosting con Docker

Contabo has published an operational guide on its blog for installing SGLang, an open-source inference server for large language models, on a GPU VPS instance. It's aimed at anyone who needs to serve an LLM to multiple users or applications simultaneously without going through a pay-per-token API.

What SGLang solves

A model loaded with a basic script handles one request at a time: with multiple concurrent users, requests queue up and the GPU remains under-utilised. SGLang is an open-source serving framework that runs large language models quickly on your own GPU, handling batching, caching and structured output, and exposing an OpenAI-compatible API.

The core mechanism is continuous batching: new requests enter an already-running batch as soon as space frees up, instead of waiting for the slowest request to finish. The standout feature is RadixAttention: the KV cache of shared prompts (same system prompt, same document) is stored in a tree structure and reused rather than recomputed every time. The server also supports quantised models, tensor parallelism across multiple GPUs, speculative decoding and open model families such as Llama, Qwen and DeepSeek.

Installation in brief

The guide proceeds step by step with Docker: pulling the official lmsysorg/sglang image, starting the container with docker run --gpus all pointing to a Hugging Face model (the example uses Qwen2.5-7B-Instruct), checking with curl http://localhost:30000/health, moving to Docker Compose with an API key, and exposing HTTPS via a Caddy reverse proxy, keeping the firewall closed on port 30000 and open only on 80/443.

Hardware sizing

The starting constraint is VRAM: a model's weight is estimated by multiplying the number of parameters by the bytes per parameter. A 70-billion-parameter model requires roughly 140 GB at 16-bit, roughly 70 GB at 8-bit and roughly 35 GB at 4-bit. More free VRAM means a larger cache, and therefore more concurrent requests and longer contexts.

The product used as a reference

For this scenario, Contabo offers its own GPU VPS with a dedicated NVIDIA RTX PRO 6000 Blackwell Server Edition. On the official price list, checked on 5 October 2026, the configuration includes 96 GB of GDDR7 VRAM, 24,064 CUDA cores, 120 TFLOPS FP32 and 1,597 GB/s of bandwidth, paired with 18 vCPUs, 96 GB of RAM and 900 GB of NVMe storage. The list price is €999.00 per month excluding VAT (€799.20 per month including VAT for the first 24 months under the two-year plan shown on the homepage), with unlimited traffic subject to fair use and availability in the Europe and US Central regions. Instances already ship with Ubuntu 24.04, CUDA and nvidia-container-toolkit pre-installed.

1PrimeCDN and other Prime Software products are not mentioned in this article.

Written with the help of artificial intelligence and checked by the editors (EU AI Act, art. 50).