Yes, you can run Ollama on a VPS without a GPU. Ollama switches to CPU-only mode by itself, and any model whose file fits in RAM, with room left for its context, loads and answers. What you give up is speed: on a CPU every new token means reading the model’s weights from memory again, so smaller model files answer faster, and the largest models suit overnight batch jobs better than chat.
This Ollama VPS guide is written for CPU servers, because that is what we sell: HourlyVPS has no GPU plans. It sizes models by RAM from the file sizes in Ollama’s own library, installs Ollama 0.35.1 (the current release on October 3, 2026) on Ubuntu 24.04 with the official script, keeps the API off the internet, and shows how to measure tokens per second on your own server instead of trusting a number from someone else’s hardware.
Key takeaways
- Ollama runs on a CPU-only VPS without extra setup, and ollama ps shows 100% CPU; it is slower than a GPU because every generated token reads the model’s weights again, so smaller files answer faster.
- Budget RAM as model file + context cache + about 1 GB: an 8B model’s default 4-bit file is 4.9–5.2 GB and fits 8 GB at Ollama’s 4,096-token CPU default, while 64,000 tokens of context adds about 8.4 GB for Llama 3.1 8B.
- Install with the official script after reading it: it creates an ollama service user and a systemd unit and leaves the API on 127.0.0.1:11434.
- Ollama’s local API has no authentication, so never set OLLAMA_HOST=0.0.0.0 on a public IP; use an SSH tunnel, or Caddy with basic_auth and header_up Host.
- Measure eval rate yourself with ollama run --verbose or eval_count ÷ eval_duration on a server billed by the hour; for all-day use keep a Chrono server running, and its charges stop at the plan’s monthly price each billing period.
Can you run Ollama on a VPS without a GPU?
Yes. When the Linux installer finds no NVIDIA or AMD GPU, it finishes anyway and says so: “No NVIDIA/AMD GPU detected. Ollama will run in CPU-only mode.” Models then load into system RAM, and ollama ps shows 100% CPU in its PROCESSOR column, which Ollama’s FAQ defines as “the model was loaded entirely in system memory”.
The reason CPU inference is slow fits in one sentence from Hugging Face’s inference guide: “For each input token, the model weights are loaded each time during the forward pass, which is slow and cumbersome when a model has billions of parameters.” GPUs are built around very fast memory; a VPS has a few vCPUs and ordinary server RAM. Two practical rules follow:
- A smaller file is a faster model. Pick the smallest model that does the job well enough, in its default 4-bit build.
- Mixture-of-experts (MoE) models read only part of their weights per token. Ollama’s library describes Gemma 4 26B as a “Mixture of Experts model with 4B active parameters”. The whole file still has to fit in RAM, but each token touches a fraction of it.
So match the job to the hardware before you deploy anything:
| Job | On a CPU-only VPS | Why |
|---|---|---|
| Embeddings for search or RAG | Good fit | Embedding models are small (nomic-embed-text is 274 MB) and inputs are short. |
| Classification, tagging, extraction, summaries in a queue | Good fit | Nobody waits on each token; a batch can run overnight. |
| Testing prompts and app code against the Ollama API | Good fit | The API is the same on CPU and GPU; only the speed changes. |
| Private drafting or chat for one person with a 3–9B model | Workable | Measure the eval rate first (see below) and decide if it is fast enough for you. |
| A chat interface for a team, or coding agents that need 64,000 tokens of context | Poor fit | Long contexts need a lot of RAM and long prompts take time to process; a GPU host or a hosted API suits these. |
| Dense 70B models | Batch jobs only | Each token reads a 43 GB file. |
Note: HourlyVPS sells CPU servers only: Quartz plans with shared vCPU and Chrono plans with dedicated vCPU, up to 16 vCPU and 64 GB of RAM. If the speed you measure is too slow for your use, the answer is usually a smaller model, a GPU host or a hosted API. Our guide to running AI agents on a VPS covers when a GPU is worth it.
How much RAM does Ollama need?
At Ollama’s default 4,096-token context, a 2 GB server runs embedding models and sub-1B models, 4 GB runs 1–4B models (4B ones only just), 8 GB runs 7–8B models, 16 GB runs 12–14B models, 32 GB runs 27–35B models and 64 GB runs 70B models, each in its default 4-bit build. Longer contexts need more. Those figures come from the budget below, applied to the file sizes in Ollama’s library.
Ollama publishes no RAM formula, so this guide uses a three-part budget. Each part is either sourced or labeled as our reasoning:
RAM needed ≈ model file + context cache + about 1 GB for the system
- Model file. The size on the model’s tags page in the Ollama library. Default tags are usually 4-bit builds: the llama3.1:8b page lists “quantization Q4_K_M” at 4.9 GB.
- Context cache. It grows with the context window. Ollama’s context length page says a larger context “will increase the amount of memory required”, and the FAQ says required RAM scales with
OLLAMA_NUM_PARALLEL×OLLAMA_CONTEXT_LENGTH. Ollama sets the default by VRAM, and below 24 GiB of VRAM it is 4k, so a CPU-only server starts at 4,096 tokens. - System and headroom. About 1 GB for Ubuntu 24.04, SSH, the Ollama server and a small safety margin. That is our reasoning, not a measurement; add more if Docker or Open WebUI run on the same server. Swap guards against an out-of-memory kill, but it is not extra room for a model: weights that land in swap are read from disk on every token.
Then check the real figure. After a model has answered once, ollama ps shows what it occupies (SIZE), where it runs (PROCESSOR) and the window it loaded with (CONTEXT), and free -h shows what is left for everything else.
Quantization: why the default 4-bit tag suits a CPU
Most models come in several builds. The tag without a suffix is usually Q4_K_M, a 4-bit build; q8_0 and fp16 builds keep more precision and need much more RAM. Qwen3 8B, as listed on its tags page:
| Tag | Build | File size |
|---|---|---|
qwen3: (default) | Q4_K_M, 4-bit | 5.2 GB |
qwen3: | 8-bit | 8.9 GB |
qwen3: | 16-bit | 16 GB |
Start with the default build. It needs about a third of the 16-bit file’s RAM, and because every token reads the weights, a smaller file should also generate faster on the same CPU. That follows from the Hugging Face quote above; check it with the measurement steps below before you rely on it.
Context length costs RAM too
The context cache stores a key and a value vector per layer for every token in the window. The model metadata Ollama publishes for Llama 3.1 8B lists 32 layers, 32 attention heads, 8 key-value heads and an embedding length of 4,096, so each head has 128 values. With the default 16-bit cache that is 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes (128 KiB) per token. Illustrative arithmetic, which ollama ps will confirm or correct on your server:
| Context (tokens) | Cache at f16 | Cache at q8_0 | 4.9 GB model + f16 cache |
|---|---|---|---|
| 4,096 (CPU default) | 0.54 GB | 0.27 GB | about 5.4 GB |
| 8,192 | 1.07 GB | 0.54 GB | about 6.0 GB |
| 32,768 | 4.29 GB | 2.15 GB | about 9.2 GB |
| 64,000 (Ollama’s advice for agents and coding tools) | 8.39 GB | 4.19 GB | about 13.3 GB |
| 131,072 (the model’s maximum) | 17.18 GB | 8.59 GB | about 22.1 GB |
ollama ps.Two settings shrink the cache. Keep the context at the 4,096-token default unless a job needs more, and consider OLLAMA_KV_CACHE_TYPE=q8_0, which the FAQ says uses “approximately 1/2 the memory of f16” with “a very small loss in precision”; it only applies when Flash Attention is on. Parallel requests multiply the cache: the FAQ’s example is a 2K context with 4 parallel requests becoming an 8K context. This table is also why our AI agent sizing guide puts an 8B model on 16 GB: an agent’s 64,000-token context needs more RAM than the model itself.
Which Ollama VPS plan fits which model?
This table applies the budget above at the default 4,096-token context. Model sizes are the default tags in the Ollama library on October 3, 2026, except where a tag is named in full; the plan choice is our reasoning, not a benchmark.
| Plan | RAM and vCPU | Model files up to about | Examples from the Ollama library | Per hour; monthly cap |
|---|---|---|---|---|
| Quartz Q2 | 2 GB, 1 shared | 0.6 GB | gemma3: (292 MB), nomic-embed-text (274 MB), qwen3: (523 MB), embeddinggemma (622 MB) | $0.02/hour; $10.00/month cap |
| Quartz Q4 | 4 GB, 2 shared | 2.5 GB | llama3. (1.3 GB), qwen3: (1.4 GB), llama3. (2.0 GB), granite4: (2.1 GB), phi4-mini (2.5 GB), qwen3: (2.5 GB) | $0.03/hour; $15.00/month cap |
| Quartz Q8 or Chrono C8 | 8 GB, 4 shared or 2 dedicated | 5.5 GB | gemma3: (3.3 GB), mistral (4.4 GB), llama3. (4.9 GB), qwen3: (5.2 GB), deepseek-r1: (5.2 GB) | Q8 $0.05/hour; $25.00/month cap C8 $0.07/hour; $35.00/month cap |
| Quartz Q16 or Chrono C16 | 16 GB, 8 shared or 4 dedicated | 12 GB | qwen3. (6.6 GB), gemma4: (6.6 GB), gemma3: (8.1 GB), deepseek-r1: (9.0 GB), qwen3: (9.3 GB) | Q16 $0.09/hour; $45.00/month cap C16 $0.13/hour; $65.00/month cap |
| Chrono C32 | 32 GB, 8 dedicated | 26 GB | gpt-oss: (14 GB, MoE), gemma3: (17 GB), qwen3. (17 GB), gemma4: (18 GB, MoE), qwen3: (19 GB, MoE), qwen3. (24 GB, MoE) | $0.25/hour; $125.00/month cap |
| Chrono C64 | 64 GB, 16 dedicated | 56 GB | llama3. (43 GB), deepseek-r1: (43 GB) | $0.48/hour; $240.00/month cap |
| None of our plans | — | — | gpt-oss: (65 GB) and qwen3. (81 GB) are larger than 64 GB of RAM. | — |
Quartz or Chrono? Inference keeps every thread it uses at full load for as long as it generates. Our acceptable use policy says that running shared Quartz vCPU at full load for long periods can slow down other customers on the same host and may lead us to limit the server’s CPU, while Chrono plans have dedicated vCPU and are “designed for sustained CPU work”. Use Quartz for short tests and light, bursty use; use Chrono for batch jobs and anything that answers requests all day. If the eval rate swings between identical runs on a Quartz plan, check steal time, the st value in top, as our VPS benchmark guide explains.
Disk is rarely the limit: every plan’s NVMe disk holds several models of the sizes it can run. Ollama keeps them in /usr/share/ollama/.ollama/models on Linux; ollama ls lists them and ollama rm deletes one. Disk sizes for every plan are on the pricing page.
Chrono C16
KVM · 25 Gbps port · Istanbul · initial credit $5
- vCPU
- 4 dedicated
- RAM
- 16 GB
- NVMe
- 200 GB
- Traffic
- Istanbul: 8 TB/month
- Per hour$0.13/hour
- Per day (24 h)$3.12/day
- Monthly cap$65.00/month
How to install Ollama on Ubuntu 24.04
Start from a server you can reach over SSH as a sudo user, with UFW on and only SSH allowed. If that is not done yet, follow connecting to your VPS with SSH and our VPS security checklist first. The steps below follow Ollama’s Linux docs.
Step 1: Install zstd
Ollama’s Linux builds ship as .tar.zst archives, and the install script stops with an error if the zstd tool is missing. Installing it first is harmless if it is already there:
sudo apt update
sudo apt install -y zstd
Step 2: Download the install script and read it
The docs pipe the script straight into sh. On a server, download it first and read it; it is about 450 lines of shell:
curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh
On a CPU-only Ubuntu server, the Linux part of the script does this:
| Part | What it does |
|---|---|
| Binaries | Installs ollama in /usr/ and its libraries in /usr/, removing an older lib/ first. |
| Service user | Creates a system user ollama with no login shell and the home directory /usr/, where models are stored. |
| Your user | Adds the user who ran the script to the ollama group. |
| systemd unit | Writes /etc/, which runs ollama serve as ollama with Restart=, then enables and starts it. |
| GPU check | Finds no GPU and prints a CPU-only warning, or asks for lspci or lshw if neither is installed. Both are expected on a VPS. |
| Service start | Restarts the service as the script exits. The API then listens on 127., the default the FAQ documents. |
Step 3: Run it as your sudo user
Run it without sudo: the script calls sudo itself where it needs root, and it adds whoever runs it to the ollama group, which should be you rather than root.
sh ollama-install.sh
To pin a version for a reproducible setup, the docs use the OLLAMA_VERSION variable, for example OLLAMA_VERSION=0.35.1 sh ollama-install.sh. To update later, run the current script again. If you prefer no script at all, the Linux page also documents a manual install from the ollama-linux-amd64.tar.zst archive with a systemd unit you write yourself.
Step 4: Check the service and the listening address
systemctl status ollama --no-pager
ollama -v
sudo ss -ltnp | grep 11434
The last line must show 127.0.0.1:11434. If it shows 0.0.0.0:11434 or [::]:11434, something has set OLLAMA_HOST; fix that before you go further (see remote access).
Prefer Docker? Publish the API on 127.0.0.1 only
Ollama also ships an official image, which replaces steps 1 to 4 on a server that already runs Docker. Its Docker page gives a CPU-only command that publishes the API with -p 11434:11434. On a VPS, change that one flag: a mapping without a host address listens on every address, and Docker’s published ports skip UFW’s rules, as our Docker guide explains. Bind it to loopback instead:
docker run -d -v ollama:/root/.ollama -p 127.0.0.1:11434:11434 --name ollama ollama/ollama
Run models with docker exec -it ollama ollama run llama3.2:3b, pass settings as -e flags (for example -e OLLAMA_CONTEXT_LENGTH=8192) instead of the systemd override below, and add --restart unless-stopped if the container should start again after a reboot. The API is still on 127.0.0.1:11434, so the curl, tunnel and Caddy steps below work the same way.
Tune Ollama for a CPU-only server
Ollama reads its settings from environment variables, and ollama serve --help lists the main ones. For the systemd service, the Linux docs put them in an override file, a drop-in that our systemd service guide explains. This one suits a CPU server running one model:
sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf <<'EOF'
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_KEEP_ALIVE=30m"
Environment="OLLAMA_MAX_LOADED_MODELS=1"
Environment="OLLAMA_NO_CLOUD=1"
EOF
Reload systemd, restart Ollama and confirm the variables arrived:
sudo systemctl daemon-reload
sudo systemctl restart ollama
systemctl show ollama --property=Environment
| Variable | Value here | Why |
|---|---|---|
OLLAMA_ | 8192 | The default without a GPU is 4,096 tokens. Raise it only as far as your prompts need; each step costs RAM (table above). Leave the line out to keep 4,096. |
OLLAMA_ | 30m | By default a model unloads after 5 minutes idle, and the next request waits while it loads again (the API reports this as load_). |
OLLAMA_ | 1 | The FAQ’s default for CPU inference is 3 models at once if they fit. One keeps memory predictable on a small plan. |
OLLAMA_ | not set (default 1) | Each parallel slot adds another context’s worth of cache. Keep 1 unless you have measured RAM to spare. |
OLLAMA_ | 1 | Local-only mode: turns off Ollama’s cloud models and web search, and the log then shows Ollama cloud disabled: true. The FAQ says prompts to local models are not seen by Ollama either way. |
For long contexts, you can also add Environment="OLLAMA_FLASH_ATTENTION=1" and Environment="OLLAMA_KV_CACHE_TYPE=q8_0" to halve the cache. Compare SIZE in ollama ps before and after, so you know it took effect on your CPU.
How many CPU threads does Ollama use?
Ollama leaves the thread count to its runner. In version 0.35.1 every GGUF model runs in llama.cpp’s llama-server, and Ollama’s source passes a thread count only when you set num_thread. Otherwise llama.cpp decides, and on x86 Linux it counts physical cores by reading each CPU’s thread_siblings list. On a VPS, “physical” means whatever CPU topology the hypervisor shows the guest, so check it:
lscpu | grep -E '^CPU\(s\)|Thread\(s\) per core|Core\(s\) per socket'
- Thread(s) per core: 1 means the default already uses every vCPU.
- Thread(s) per core: 2 means the default uses half of them. Try
num_threadset to your vCPU count and compare the eval rate both ways. - Watch
topduring a generation (press 1 for one line per CPU) to see how many vCPUs are actually busy.
num_thread is a load-time setting: Ollama’s API types list it among the “Runner options which must be set when the model is loaded into memory”, so changing it reloads the model. Set it per request in options (example in the API section below). Going above your vCPU count gains nothing, and on a shared Quartz plan the fair-use rule above applies however many threads you use.
The CPU’s instruction sets matter too. Ollama’s troubleshooting page says its AVX2 CPU library is the fastest, followed by AVX, with plain cpu the slowest but most compatible. Its command shows which flags your server’s CPU exposes; look for avx2 and avx512f in the list:
cat /proc/cpuinfo| grep flags | head -1
Pull a model and measure its speed
Start small. Llama 3.2 3B is a 2.0 GB download with a 128K context window, and it fits a 4 GB plan:
ollama pull llama3.2:3b
Run it with timings switched on:
ollama run llama3.2:3b --verbose
Type a prompt typical of your job. The --verbose flag (“Show timings for response” in the CLI source) prints timings after each answer, including prompt eval rate, how fast your prompt was read, and eval rate, how fast the answer was generated, both in tokens per second. Eval rate is the speed you feel. Type /bye to leave. Then look at the loaded model from the shell:
ollama ps
Read three columns: SIZE (what the model really occupies, cache included), PROCESSOR (100% CPU on a CPU VPS) and CONTEXT (the window it loaded with). Run the same prompt against two or three candidate models and keep the smallest whose answers are good enough. Write the numbers down: they are the benchmark for your workload on this plan, and the only speed figures worth planning with.
Tip: Size the plan by measuring, not guessing. Deploy one size up, pull your candidate models, note eval rate and SIZE for each, then delete the server, which stops billing. A three-hour session on Chrono C16 costs $0.39.
Call the Ollama API with curl
Everything the CLI does goes through the HTTP API on 127.0.0.1:11434. This non-streaming request to /api/generate also sets the context window and thread count for the request, using options from Ollama’s API reference:
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2:3b",
"prompt": "Summarize in one sentence: Ollama runs open models on your own server.",
"stream": false,
"keep_alive": "30m",
"options": {
"num_ctx": 4096,
"num_thread": 2
}
}'
Set num_thread to your plan’s vCPU count, or drop the line to keep the default. The JSON answer carries the text in response plus usage metrics, and the API usage docs say all timings are in nanoseconds. So eval_count ÷ eval_duration × 109 is tokens per second. This prints only the answer and the speed, with the python3 that Ubuntu 24.04 already has:
curl -s http://localhost:11434/api/generate \
-d '{"model": "llama3.2:3b", "prompt": "Name three uses for a VPS.", "stream": false}' \
| python3 -c 'import json, sys; r = json.load(sys.stdin); print(r["response"]); print(round(r["eval_count"] / r["eval_duration"] * 1e9, 1), "tokens/s")'
Embeddings use /api/embed with model and input. Pull nomic-embed-text first; at 274 MB it runs on the smallest plans:
curl http://localhost:11434/api/embed -d '{
"model": "nomic-embed-text",
"input": "Ollama runs on CPU-only servers."
}'
Apps written for the OpenAI API can point at http://localhost:11434/v1; Ollama’s quickstart shows /v1/chat/completions working against a local server. Tools on the same server, such as a self-hosted n8n, can call the API on localhost without any of the exposure below.
How do you reach Ollama remotely without exposing it?
Start with what Ollama’s own API docs say:
The local API at http://localhost:11434 does not require authentication.
Ollama docs, “Authentication”
That is fine on a laptop and dangerous on a public IP. Anyone who can reach port 11434 can use every endpoint: run models on CPU time you pay for, pull new models onto your disk, or delete yours. The FAQ shows how to bind other addresses with OLLAMA_HOST=0.0.0.0:11434; on a VPS, don’t. Keep the default 127.0.0.1 and choose one of these:
| Method | Suited to | Who gets in | What listens publicly |
|---|---|---|---|
| SSH tunnel | You, from your own computer | Holders of your SSH key | Nothing new |
Caddy reverse proxy with basic_ | Apps and people on other machines | Holders of the password, over HTTPS | Ports 80 and 443 |
| Private network (WireGuard or Tailscale) | Several machines of your own | Members of the network | The VPN’s port only |
OLLAMA_ on a public IP | Nobody | Everyone on the internet | Port 11434, unauthenticated |
SSH tunnel: the simplest safe option
On your own computer, forward a local port to the server’s loopback address:
ssh -N -L 11434:localhost:11434 user@SERVER_IP
Leave it running. In a second terminal on your computer, http://localhost:11434 now reaches the server’s Ollama, so curl http://localhost:11434/api/tags lists its models. If Ollama also runs on your computer, pick another local port, such as -L 11435:localhost:11434. Our SSH port forwarding guide explains local, remote and dynamic tunnels, and an SSH config alias shortens the command.
Caddy reverse proxy with a password
For apps on other machines, put Caddy in front with HTTPS and a password. Point a DNS name at the server, install Caddy and open ports 80 and 443 as in our Caddy reverse proxy guide, create a password hash with caddy hash-password --algorithm argon2id as in its access control section, then use this site block in /etc/caddy/Caddyfile:
ollama.example.com {
basic_auth argon2id {
alice PASTE_THE_HASH_HERE
}
reverse_proxy localhost:11434 {
header_up Host {upstream_hostport}
}
}
The header_up line is not optional. An Ollama bound to loopback answers 403 Forbidden to requests whose Host header names a public domain; in version 0.35.1 the check is allowedHostsMiddleware in server/routes.go. That is why the FAQ’s Nginx example sets Host localhost:11434, and Caddy’s {upstream_hostport} does the same. Clients then send the password with every request, for example curl -u alice https://ollama.example.com/api/tags. Some clients built for the OpenAI API only send bearer tokens; for those, a tunnel or a private network is simpler.
Open WebUI: a chat interface in front of Ollama
Open WebUI is a separate open-source web interface for Ollama. It runs in Docker (see our Docker install guide). With Ollama on the same server listening on 127.0.0.1, Open WebUI’s connection troubleshooting page gives this command:
docker run -d --network=host -v open-webui:/app/backend/data \
-e OLLAMA_BASE_URL=http://127.0.0.1:11434 --name open-webui \
--restart always ghcr.io/open-webui/open-webui:main
With --network=host the interface listens on port 8080 of the server, and its start script binds every interface (0.0.0.0) by default. Because no port is published with -p, Docker adds no rule that routes around UFW, the problem our Docker guide describes, so UFW’s default deny keeps 8080 closed to the internet. Reach it through a tunnel: ssh -N -L 8080:localhost:8080 user@SERVER_IP, then open http://localhost:8080 on your computer.
Warning: Open WebUI’s quick start says the first account created is the administrator. Before you start the container, check that UFW is active with sudo ufw status. Once it runs, test from your own computer: curl -m 5 http://SERVER_IP:8080 should time out. Then create the admin account straight away, through the tunnel.
What does an Ollama VPS cost? Hourly tests vs always-on
Most Ollama projects start with a question: which model is good enough, on which plan? Answering it takes a few hours on an hourly VPS that you delete afterward. If the answer is “this model, all day”, keep the server on: it is billed the same way, and its charges stop at the plan’s monthly price each billing period (one month from your order date). One plan across the whole range:
| Duration | Hours on the meter | Cost $0.13 | Note |
|---|---|---|---|
| 1 hour | 1 | $0.13 | |
| 3 hours | 3 | $0.39 | |
| 8 hours | 8 | $1.04 | |
| 1 day | 24 | $3.12 | |
| 7 days | 168 | $21.84 | |
| 30 days | 720 | $65.00 | Capped at the monthly price |
| Job | Plan | Duration | Cost |
|---|---|---|---|
| Compare three small models for a bot | Quartz Q4 | 2 hours | $0.06 |
| Pick an 8B–14B model and measure its eval rate | Chrono C16 | 3 hours | $0.39 |
| See whether a 70B model is worth it before buying hardware | Chrono C64 | 4 hours | $1.92 |
| Overnight batch: summarize or embed a document archive | Chrono C32 | 12 hours | $3.00 |
| A private API for an internal app, always on | Chrono C16 | 30 days | $65.00 |
How you pay: Every server is billed by the hour: the plan’s hourly rate is deducted from the server’s prepaid balance for every hour it exists, powered on or off, until you delete it. Billing is by the hour: every hour a server exists is charged at the plan’s hourly rate. The price tapes and cost tables on this site count every started hour as a full hour, so they show the most a duration can cost; the cost calculator charges a partial hour to the nearest cent, as the bill does. Every plan is capped at 500 hours per billing period: the monthly price is 500 times the hourly rate, and after 500 billed hours (about 20.8 days) the rest of the period is free, so a server left on all month pays exactly the monthly price and never more. A billing period is one month from your order date, not the calendar month, and the cap resets every period. Worked examples are in hourly VPS billing explained.
Delete, don’t stop: A stopped server is still billed, because its vCPU, memory, disk and IP addresses stay reserved for you; only deleting the server stops billing. When a test is done, copy off anything you need and delete the server; the delete VPS checklist covers what to save and revoke first.
Keeping a test server for production takes no new order. Each plan is one product, so there is no billing mode to pick or switch: the same server is billed by the hour whether you keep it for an hour, a day or all month, and the monthly cap applies by itself. A private API that answers requests around the clock is what a monthly VPS means here, with no contract and no prepaid term. A batch that runs over a weekend is 48 to 72 hours of hourly billing, which is all a daily VPS is. The hourly vs monthly VPS guide compares the cap with providers that bill every hour, and the VPS cost calculator prices any other schedule. Servers deploy in Istanbul today; New York is coming soon.
Troubleshooting Ollama on a CPU VPS
Ollama’s troubleshooting docs read the service log on systemd systems with:
journalctl -u ollama --no-pager --follow --pager-end
| Symptom | Likely cause | Fix |
|---|---|---|
| Install stops: “This version requires zstd for extraction” | The zstd tool is missing | sudo apt install zstd, then run the script again. |
| Installer warns “No NVIDIA/AMD GPU detected” or “Unable to detect NVIDIA/AMD GPU” | Normal on a CPU VPS | Nothing to fix. Ollama runs in CPU-only mode. |
| A request fails while a large model loads, or the service restarts | Not enough RAM for the model plus its context | Check sudo journalctl -k | grep -i "out of memory". Use a smaller model, a shorter context or the next plan up. |
| Long prompts or chats lose their beginning | The context window (4,096 tokens by default on CPU) is shorter than the input | Raise OLLAMA_ or num_; check CONTEXT in ollama ps. |
| 403 Forbidden through a reverse proxy | The proxy forwards a public Host header | header_ in Caddy, or proxy_ in Nginx. |
A Docker container cannot reach host. | Ollama listens on 127.0.0.1 only, by design | Use --network= and http: as above. Do not rebind Ollama to 0.0.0.0. |
Only half the vCPUs are busy in top | llama.cpp counted physical cores, and the guest shows 2 threads per core | Try num_ equal to your vCPU count and compare eval rates. |
| Long jobs slow down on a Quartz plan | Sustained full load on shared vCPU may be limited (AUP section 8), or other VMs take CPU time (st in top) | Move the job to a Chrono plan. |
| The disk fills up | Models accumulate in /usr/ | ollama ls, then ollama rm the ones you no longer use. |
Ollama has had serious security bugs before; CVE-2024-37032, fixed in version 0.1.34, was a path traversal through model digests the server did not validate. Re-run the install script now and then to update, and keep port 11434 private so the next bug is not reachable from the internet. For the rest of the server, keep the automatic security updates from our checklist switched on.
Deploy this setup
Test Ollama models on a CPU-only server
Chrono C16 · 4 dedicated vCPU · 16 GB RAM · 200 GB NVMe · Istanbul
- Per hour$0.13/hourFor this job
- Per day (24 h)$3.12/day
- Monthly cap$65.00/month
Starts with a $5 initial credit, which goes into the server’s balance and pays for its hours.
Billed by the hour, never more than $65.00 per billing period. Delete the server and billing stops.



