Ollama VPS Guide: Run LLMs on a CPU-Only Server (2026)

Run Ollama on a CPU-only VPS: RAM by model from Ollama’s library, which plan fits which model, a reviewed install on Ubuntu 24.04, thread tuning and remote access without an open port.

Title card reading “Ollama on a CPU VPS” for the HourlyVPS guide to running Ollama models on a CPU-only server

Yes, you can run Ollama on a VPS without a GPU. Ollama switches to CPU-only mode by itself, and any model whose file fits in RAM, with room left for its context, loads and answers. What you give up is speed: on a CPU every new token means reading the model’s weights from memory again, so smaller model files answer faster, and the largest models suit overnight batch jobs better than chat.

This Ollama VPS guide is written for CPU servers, because that is what we sell: HourlyVPS has no GPU plans. It sizes models by RAM from the file sizes in Ollama’s own library, installs Ollama 0.35.1 (the current release on October 3, 2026) on Ubuntu 24.04 with the official script, keeps the API off the internet, and shows how to measure tokens per second on your own server instead of trusting a number from someone else’s hardware.

Key takeaways

  • Ollama runs on a CPU-only VPS without extra setup, and ollama ps shows 100% CPU; it is slower than a GPU because every generated token reads the model’s weights again, so smaller files answer faster.
  • Budget RAM as model file + context cache + about 1 GB: an 8B model’s default 4-bit file is 4.9–5.2 GB and fits 8 GB at Ollama’s 4,096-token CPU default, while 64,000 tokens of context adds about 8.4 GB for Llama 3.1 8B.
  • Install with the official script after reading it: it creates an ollama service user and a systemd unit and leaves the API on 127.0.0.1:11434.
  • Ollama’s local API has no authentication, so never set OLLAMA_HOST=0.0.0.0 on a public IP; use an SSH tunnel, or Caddy with basic_auth and header_up Host.
  • Measure eval rate yourself with ollama run --verbose or eval_count ÷ eval_duration on a server billed by the hour; for all-day use keep a Chrono server running, and its charges stop at the plan’s monthly price each billing period.

Can you run Ollama on a VPS without a GPU?

Yes. When the Linux installer finds no NVIDIA or AMD GPU, it finishes anyway and says so: “No NVIDIA/AMD GPU detected. Ollama will run in CPU-only mode.” Models then load into system RAM, and ollama ps shows 100% CPU in its PROCESSOR column, which Ollama’s FAQ defines as “the model was loaded entirely in system memory”.

The reason CPU inference is slow fits in one sentence from Hugging Face’s inference guide: “For each input token, the model weights are loaded each time during the forward pass, which is slow and cumbersome when a model has billions of parameters.” GPUs are built around very fast memory; a VPS has a few vCPUs and ordinary server RAM. Two practical rules follow:

  • A smaller file is a faster model. Pick the smallest model that does the job well enough, in its default 4-bit build.
  • Mixture-of-experts (MoE) models read only part of their weights per token. Ollama’s library describes Gemma 4 26B as a “Mixture of Experts model with 4B active parameters”. The whole file still has to fit in RAM, but each token touches a fraction of it.

So match the job to the hardware before you deploy anything:

JobOn a CPU-only VPSWhy
Embeddings for search or RAGGood fitEmbedding models are small (nomic-embed-text is 274 MB) and inputs are short.
Classification, tagging, extraction, summaries in a queueGood fitNobody waits on each token; a batch can run overnight.
Testing prompts and app code against the Ollama APIGood fitThe API is the same on CPU and GPU; only the speed changes.
Private drafting or chat for one person with a 3–9B modelWorkableMeasure the eval rate first (see below) and decide if it is fast enough for you.
A chat interface for a team, or coding agents that need 64,000 tokens of contextPoor fitLong contexts need a lot of RAM and long prompts take time to process; a GPU host or a hosted API suits these.
Dense 70B modelsBatch jobs onlyEach token reads a 43 GB file.
Fit is our reasoning from the sizes and docs cited in this guide, not a benchmark.

Note: HourlyVPS sells CPU servers only: Quartz plans with shared vCPU and Chrono plans with dedicated vCPU, up to 16 vCPU and 64 GB of RAM. If the speed you measure is too slow for your use, the answer is usually a smaller model, a GPU host or a hosted API. Our guide to running AI agents on a VPS covers when a GPU is worth it.

How much RAM does Ollama need?

At Ollama’s default 4,096-token context, a 2 GB server runs embedding models and sub-1B models, 4 GB runs 1–4B models (4B ones only just), 8 GB runs 7–8B models, 16 GB runs 12–14B models, 32 GB runs 27–35B models and 64 GB runs 70B models, each in its default 4-bit build. Longer contexts need more. Those figures come from the budget below, applied to the file sizes in Ollama’s library.

Ollama publishes no RAM formula, so this guide uses a three-part budget. Each part is either sourced or labeled as our reasoning:

RAM needed ≈ model file + context cache + about 1 GB for the system

  • Model file. The size on the model’s tags page in the Ollama library. Default tags are usually 4-bit builds: the llama3.1:8b page lists “quantization Q4_K_M” at 4.9 GB.
  • Context cache. It grows with the context window. Ollama’s context length page says a larger context “will increase the amount of memory required”, and the FAQ says required RAM scales with OLLAMA_NUM_PARALLEL × OLLAMA_CONTEXT_LENGTH. Ollama sets the default by VRAM, and below 24 GiB of VRAM it is 4k, so a CPU-only server starts at 4,096 tokens.
  • System and headroom. About 1 GB for Ubuntu 24.04, SSH, the Ollama server and a small safety margin. That is our reasoning, not a measurement; add more if Docker or Open WebUI run on the same server. Swap guards against an out-of-memory kill, but it is not extra room for a model: weights that land in swap are read from disk on every token.

Then check the real figure. After a model has answered once, ollama ps shows what it occupies (SIZE), where it runs (PROCESSOR) and the window it loaded with (CONTEXT), and free -h shows what is left for everything else.

Quantization: why the default 4-bit tag suits a CPU

Most models come in several builds. The tag without a suffix is usually Q4_K_M, a 4-bit build; q8_0 and fp16 builds keep more precision and need much more RAM. Qwen3 8B, as listed on its tags page:

TagBuildFile size
qwen3:8b (default)Q4_K_M, 4-bit5.2 GB
qwen3:8b-q8_08-bit8.9 GB
qwen3:8b-fp1616-bit16 GB
File sizes from the Ollama library, checked October 3, 2026.

Start with the default build. It needs about a third of the 16-bit file’s RAM, and because every token reads the weights, a smaller file should also generate faster on the same CPU. That follows from the Hugging Face quote above; check it with the measurement steps below before you rely on it.

Context length costs RAM too

The context cache stores a key and a value vector per layer for every token in the window. The model metadata Ollama publishes for Llama 3.1 8B lists 32 layers, 32 attention heads, 8 key-value heads and an embedding length of 4,096, so each head has 128 values. With the default 16-bit cache that is 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes (128 KiB) per token. Illustrative arithmetic, which ollama ps will confirm or correct on your server:

Context (tokens)Cache at f16Cache at q8_04.9 GB model + f16 cache
4,096 (CPU default)0.54 GB0.27 GBabout 5.4 GB
8,1921.07 GB0.54 GBabout 6.0 GB
32,7684.29 GB2.15 GBabout 9.2 GB
64,000 (Ollama’s advice for agents and coding tools)8.39 GB4.19 GBabout 13.3 GB
131,072 (the model’s maximum)17.18 GB8.59 GBabout 22.1 GB
Llama 3.1 8B, our arithmetic from its published metadata; 1 GB = 109 bytes. Other architectures differ, so read SIZE in ollama ps.

Two settings shrink the cache. Keep the context at the 4,096-token default unless a job needs more, and consider OLLAMA_KV_CACHE_TYPE=q8_0, which the FAQ says uses “approximately 1/2 the memory of f16” with “a very small loss in precision”; it only applies when Flash Attention is on. Parallel requests multiply the cache: the FAQ’s example is a 2K context with 4 parallel requests becoming an 8K context. This table is also why our AI agent sizing guide puts an 8B model on 16 GB: an agent’s 64,000-token context needs more RAM than the model itself.

Which Ollama VPS plan fits which model?

This table applies the budget above at the default 4,096-token context. Model sizes are the default tags in the Ollama library on October 3, 2026, except where a tag is named in full; the plan choice is our reasoning, not a benchmark.

PlanRAM and vCPUModel files up to aboutExamples from the Ollama libraryPer hour; monthly cap
Quartz Q22 GB, 1 shared0.6 GBgemma3:270m (292 MB), nomic-embed-text (274 MB), qwen3:0.6b (523 MB), embeddinggemma (622 MB)$0.02/hour; $10.00/month cap
Quartz Q44 GB, 2 shared2.5 GBllama3.2:1b (1.3 GB), qwen3:1.7b (1.4 GB), llama3.2:3b (2.0 GB), granite4:micro (2.1 GB), phi4-mini (2.5 GB), qwen3:4b (2.5 GB)$0.03/hour; $15.00/month cap
Quartz Q8 or Chrono C88 GB, 4 shared or 2 dedicated5.5 GBgemma3:4b (3.3 GB), mistral (4.4 GB), llama3.1:8b (4.9 GB), qwen3:8b (5.2 GB), deepseek-r1:8b (5.2 GB)Q8 $0.05/hour; $25.00/month cap
C8 $0.07/hour; $35.00/month cap
Quartz Q16 or Chrono C1616 GB, 8 shared or 4 dedicated12 GBqwen3.5:9b (6.6 GB), gemma4:e4b-it-q4_K_M (6.6 GB), gemma3:12b (8.1 GB), deepseek-r1:14b (9.0 GB), qwen3:14b (9.3 GB)Q16 $0.09/hour; $45.00/month cap
C16 $0.13/hour; $65.00/month cap
Chrono C3232 GB, 8 dedicated26 GBgpt-oss:20b (14 GB, MoE), gemma3:27b (17 GB), qwen3.5:27b (17 GB), gemma4:26b-a4b-it-q4_K_M (18 GB, MoE), qwen3:30b (19 GB, MoE), qwen3.6:35b-a3b-q4_K_M (24 GB, MoE)$0.25/hour; $125.00/month cap
Chrono C6464 GB, 16 dedicated56 GBllama3.1:70b (43 GB), deepseek-r1:70b (43 GB)$0.48/hour; $240.00/month cap
None of our plans——gpt-oss:120b (65 GB) and qwen3.5:122b (81 GB) are larger than 64 GB of RAM.—
Sizes from ollama.com/library tags pages, checked October 3, 2026. “Up to about” leaves room for the 4,096-token context cache and about 1 GB for the system; longer contexts need the next row up. The gpt-oss library page says the smaller model runs “on systems with as little as 16GB memory”, which leaves almost nothing for context on a 16 GB plan. The monthly cap is the most a server left on costs in a billing period.

Quartz or Chrono? Inference keeps every thread it uses at full load for as long as it generates. Our acceptable use policy says that running shared Quartz vCPU at full load for long periods can slow down other customers on the same host and may lead us to limit the server’s CPU, while Chrono plans have dedicated vCPU and are “designed for sustained CPU work”. Use Quartz for short tests and light, bursty use; use Chrono for batch jobs and anything that answers requests all day. If the eval rate swings between identical runs on a Quartz plan, check steal time, the st value in top, as our VPS benchmark guide explains.

Disk is rarely the limit: every plan’s NVMe disk holds several models of the sizes it can run. Ollama keeps them in /usr/share/ollama/.ollama/models on Linux; ollama ls lists them and ollama rm deletes one. Disk sizes for every plan are on the pricing page.

Chrono C16

KVM · 25 Gbps port · Istanbul · initial credit $5

vCPU
4 dedicated
RAM
16 GB
NVMe
200 GB
Traffic
Istanbul: 8 TB/month
  • Per hour$0.13/hour
  • Per day (24 h)$3.12/day
  • Monthly cap$65.00/month

Deploy Chrono C16 Plan details for Chrono C16

How to install Ollama on Ubuntu 24.04

Start from a server you can reach over SSH as a sudo user, with UFW on and only SSH allowed. If that is not done yet, follow connecting to your VPS with SSH and our VPS security checklist first. The steps below follow Ollama’s Linux docs.

Step 1: Install zstd

Ollama’s Linux builds ship as .tar.zst archives, and the install script stops with an error if the zstd tool is missing. Installing it first is harmless if it is already there:

sudo apt update
sudo apt install -y zstd

Step 2: Download the install script and read it

The docs pipe the script straight into sh. On a server, download it first and read it; it is about 450 lines of shell:

curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh

On a CPU-only Ubuntu server, the Linux part of the script does this:

PartWhat it does
BinariesInstalls ollama in /usr/local/bin and its libraries in /usr/local/lib/ollama, removing an older lib/ollama first.
Service userCreates a system user ollama with no login shell and the home directory /usr/share/ollama, where models are stored.
Your userAdds the user who ran the script to the ollama group.
systemd unitWrites /etc/systemd/system/ollama.service, which runs ollama serve as ollama with Restart=always, then enables and starts it.
GPU checkFinds no GPU and prints a CPU-only warning, or asks for lspci or lshw if neither is installed. Both are expected on a VPS.
Service startRestarts the service as the script exits. The API then listens on 127.0.0.1:11434, the default the FAQ documents.
From install.sh as served on October 3, 2026.

Step 3: Run it as your sudo user

Run it without sudo: the script calls sudo itself where it needs root, and it adds whoever runs it to the ollama group, which should be you rather than root.

sh ollama-install.sh

To pin a version for a reproducible setup, the docs use the OLLAMA_VERSION variable, for example OLLAMA_VERSION=0.35.1 sh ollama-install.sh. To update later, run the current script again. If you prefer no script at all, the Linux page also documents a manual install from the ollama-linux-amd64.tar.zst archive with a systemd unit you write yourself.

Step 4: Check the service and the listening address

systemctl status ollama --no-pager
ollama -v
sudo ss -ltnp | grep 11434

The last line must show 127.0.0.1:11434. If it shows 0.0.0.0:11434 or [::]:11434, something has set OLLAMA_HOST; fix that before you go further (see remote access).

Prefer Docker? Publish the API on 127.0.0.1 only

Ollama also ships an official image, which replaces steps 1 to 4 on a server that already runs Docker. Its Docker page gives a CPU-only command that publishes the API with -p 11434:11434. On a VPS, change that one flag: a mapping without a host address listens on every address, and Docker’s published ports skip UFW’s rules, as our Docker guide explains. Bind it to loopback instead:

docker run -d -v ollama:/root/.ollama -p 127.0.0.1:11434:11434 --name ollama ollama/ollama

Run models with docker exec -it ollama ollama run llama3.2:3b, pass settings as -e flags (for example -e OLLAMA_CONTEXT_LENGTH=8192) instead of the systemd override below, and add --restart unless-stopped if the container should start again after a reboot. The API is still on 127.0.0.1:11434, so the curl, tunnel and Caddy steps below work the same way.

Tune Ollama for a CPU-only server

Ollama reads its settings from environment variables, and ollama serve --help lists the main ones. For the systemd service, the Linux docs put them in an override file, a drop-in that our systemd service guide explains. This one suits a CPU server running one model:

sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf <<'EOF'
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_KEEP_ALIVE=30m"
Environment="OLLAMA_MAX_LOADED_MODELS=1"
Environment="OLLAMA_NO_CLOUD=1"
EOF

Reload systemd, restart Ollama and confirm the variables arrived:

sudo systemctl daemon-reload
sudo systemctl restart ollama
systemctl show ollama --property=Environment
VariableValue hereWhy
OLLAMA_CONTEXT_LENGTH8192The default without a GPU is 4,096 tokens. Raise it only as far as your prompts need; each step costs RAM (table above). Leave the line out to keep 4,096.
OLLAMA_KEEP_ALIVE30mBy default a model unloads after 5 minutes idle, and the next request waits while it loads again (the API reports this as load_duration).
OLLAMA_MAX_LOADED_MODELS1The FAQ’s default for CPU inference is 3 models at once if they fit. One keeps memory predictable on a small plan.
OLLAMA_NUM_PARALLELnot set (default 1)Each parallel slot adds another context’s worth of cache. Keep 1 unless you have measured RAM to spare.
OLLAMA_NO_CLOUD1Local-only mode: turns off Ollama’s cloud models and web search, and the log then shows Ollama cloud disabled: true. The FAQ says prompts to local models are not seen by Ollama either way.
Variables and defaults from Ollama’s FAQ and context length docs, checked October 3, 2026.

For long contexts, you can also add Environment="OLLAMA_FLASH_ATTENTION=1" and Environment="OLLAMA_KV_CACHE_TYPE=q8_0" to halve the cache. Compare SIZE in ollama ps before and after, so you know it took effect on your CPU.

How many CPU threads does Ollama use?

Ollama leaves the thread count to its runner. In version 0.35.1 every GGUF model runs in llama.cpp’s llama-server, and Ollama’s source passes a thread count only when you set num_thread. Otherwise llama.cpp decides, and on x86 Linux it counts physical cores by reading each CPU’s thread_siblings list. On a VPS, “physical” means whatever CPU topology the hypervisor shows the guest, so check it:

lscpu | grep -E '^CPU\(s\)|Thread\(s\) per core|Core\(s\) per socket'
  • Thread(s) per core: 1 means the default already uses every vCPU.
  • Thread(s) per core: 2 means the default uses half of them. Try num_thread set to your vCPU count and compare the eval rate both ways.
  • Watch top during a generation (press 1 for one line per CPU) to see how many vCPUs are actually busy.

num_thread is a load-time setting: Ollama’s API types list it among the “Runner options which must be set when the model is loaded into memory”, so changing it reloads the model. Set it per request in options (example in the API section below). Going above your vCPU count gains nothing, and on a shared Quartz plan the fair-use rule above applies however many threads you use.

The CPU’s instruction sets matter too. Ollama’s troubleshooting page says its AVX2 CPU library is the fastest, followed by AVX, with plain cpu the slowest but most compatible. Its command shows which flags your server’s CPU exposes; look for avx2 and avx512f in the list:

cat /proc/cpuinfo| grep flags | head -1

Pull a model and measure its speed

Start small. Llama 3.2 3B is a 2.0 GB download with a 128K context window, and it fits a 4 GB plan:

ollama pull llama3.2:3b

Run it with timings switched on:

ollama run llama3.2:3b --verbose

Type a prompt typical of your job. The --verbose flag (“Show timings for response” in the CLI source) prints timings after each answer, including prompt eval rate, how fast your prompt was read, and eval rate, how fast the answer was generated, both in tokens per second. Eval rate is the speed you feel. Type /bye to leave. Then look at the loaded model from the shell:

ollama ps

Read three columns: SIZE (what the model really occupies, cache included), PROCESSOR (100% CPU on a CPU VPS) and CONTEXT (the window it loaded with). Run the same prompt against two or three candidate models and keep the smallest whose answers are good enough. Write the numbers down: they are the benchmark for your workload on this plan, and the only speed figures worth planning with.

Tip: Size the plan by measuring, not guessing. Deploy one size up, pull your candidate models, note eval rate and SIZE for each, then delete the server, which stops billing. A three-hour session on Chrono C16 costs $0.39.

Call the Ollama API with curl

Everything the CLI does goes through the HTTP API on 127.0.0.1:11434. This non-streaming request to /api/generate also sets the context window and thread count for the request, using options from Ollama’s API reference:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2:3b",
  "prompt": "Summarize in one sentence: Ollama runs open models on your own server.",
  "stream": false,
  "keep_alive": "30m",
  "options": {
    "num_ctx": 4096,
    "num_thread": 2
  }
}'

Set num_thread to your plan’s vCPU count, or drop the line to keep the default. The JSON answer carries the text in response plus usage metrics, and the API usage docs say all timings are in nanoseconds. So eval_count ÷ eval_duration × 109 is tokens per second. This prints only the answer and the speed, with the python3 that Ubuntu 24.04 already has:

curl -s http://localhost:11434/api/generate \
  -d '{"model": "llama3.2:3b", "prompt": "Name three uses for a VPS.", "stream": false}' \
  | python3 -c 'import json, sys; r = json.load(sys.stdin); print(r["response"]); print(round(r["eval_count"] / r["eval_duration"] * 1e9, 1), "tokens/s")'

Embeddings use /api/embed with model and input. Pull nomic-embed-text first; at 274 MB it runs on the smallest plans:

curl http://localhost:11434/api/embed -d '{
  "model": "nomic-embed-text",
  "input": "Ollama runs on CPU-only servers."
}'

Apps written for the OpenAI API can point at http://localhost:11434/v1; Ollama’s quickstart shows /v1/chat/completions working against a local server. Tools on the same server, such as a self-hosted n8n, can call the API on localhost without any of the exposure below.

How do you reach Ollama remotely without exposing it?

Start with what Ollama’s own API docs say:

The local API at http://localhost:11434 does not require authentication.

Ollama docs, “Authentication”

That is fine on a laptop and dangerous on a public IP. Anyone who can reach port 11434 can use every endpoint: run models on CPU time you pay for, pull new models onto your disk, or delete yours. The FAQ shows how to bind other addresses with OLLAMA_HOST=0.0.0.0:11434; on a VPS, don’t. Keep the default 127.0.0.1 and choose one of these:

MethodSuited toWho gets inWhat listens publicly
SSH tunnelYou, from your own computerHolders of your SSH keyNothing new
Caddy reverse proxy with basic_authApps and people on other machinesHolders of the password, over HTTPSPorts 80 and 443
Private network (WireGuard or Tailscale)Several machines of your ownMembers of the networkThe VPN’s port only
OLLAMA_HOST=0.0.0.0 on a public IPNobodyEveryone on the internetPort 11434, unauthenticated

SSH tunnel: the simplest safe option

On your own computer, forward a local port to the server’s loopback address:

ssh -N -L 11434:localhost:11434 user@SERVER_IP

Leave it running. In a second terminal on your computer, http://localhost:11434 now reaches the server’s Ollama, so curl http://localhost:11434/api/tags lists its models. If Ollama also runs on your computer, pick another local port, such as -L 11435:localhost:11434. Our SSH port forwarding guide explains local, remote and dynamic tunnels, and an SSH config alias shortens the command.

Caddy reverse proxy with a password

For apps on other machines, put Caddy in front with HTTPS and a password. Point a DNS name at the server, install Caddy and open ports 80 and 443 as in our Caddy reverse proxy guide, create a password hash with caddy hash-password --algorithm argon2id as in its access control section, then use this site block in /etc/caddy/Caddyfile:

ollama.example.com {
	basic_auth argon2id {
		alice PASTE_THE_HASH_HERE
	}
	reverse_proxy localhost:11434 {
		header_up Host {upstream_hostport}
	}
}

The header_up line is not optional. An Ollama bound to loopback answers 403 Forbidden to requests whose Host header names a public domain; in version 0.35.1 the check is allowedHostsMiddleware in server/routes.go. That is why the FAQ’s Nginx example sets Host localhost:11434, and Caddy’s {upstream_hostport} does the same. Clients then send the password with every request, for example curl -u alice https://ollama.example.com/api/tags. Some clients built for the OpenAI API only send bearer tokens; for those, a tunnel or a private network is simpler.

Open WebUI: a chat interface in front of Ollama

Open WebUI is a separate open-source web interface for Ollama. It runs in Docker (see our Docker install guide). With Ollama on the same server listening on 127.0.0.1, Open WebUI’s connection troubleshooting page gives this command:

docker run -d --network=host -v open-webui:/app/backend/data \
  -e OLLAMA_BASE_URL=http://127.0.0.1:11434 --name open-webui \
  --restart always ghcr.io/open-webui/open-webui:main

With --network=host the interface listens on port 8080 of the server, and its start script binds every interface (0.0.0.0) by default. Because no port is published with -p, Docker adds no rule that routes around UFW, the problem our Docker guide describes, so UFW’s default deny keeps 8080 closed to the internet. Reach it through a tunnel: ssh -N -L 8080:localhost:8080 user@SERVER_IP, then open http://localhost:8080 on your computer.

Warning: Open WebUI’s quick start says the first account created is the administrator. Before you start the container, check that UFW is active with sudo ufw status. Once it runs, test from your own computer: curl -m 5 http://SERVER_IP:8080 should time out. Then create the admin account straight away, through the tunnel.

What does an Ollama VPS cost? Hourly tests vs always-on

Most Ollama projects start with a question: which model is good enough, on which plan? Answering it takes a few hours on an hourly VPS that you delete afterward. If the answer is “this model, all day”, keep the server on: it is billed the same way, and its charges stop at the plan’s monthly price each billing period (one month from your order date). One plan across the whole range:

Chrono C16 (4 dedicated vCPU, 16 GB): an Ollama test, a working day, a week and a month, billed by the hour with the monthly cap applied
DurationHours on the meterCost $0.13/hour · cap $65.00/monthNote
1 hour1$0.13
3 hours3$0.39
8 hours8$1.04
1 day24$3.12
7 days168$21.84
30 days720$65.00Capped at the monthly price
Charges stop at $65.00 after 500 hours (about 20.8 days) in a billing period; the rest of that period is free.
JobPlanDurationCost
Compare three small models for a botQuartz Q42 hours$0.06
Pick an 8B–14B model and measure its eval rateChrono C163 hours$0.39
See whether a 70B model is worth it before buying hardwareChrono C644 hours$1.92
Overnight batch: summarize or embed a document archiveChrono C3212 hours$3.00
A private API for an internal app, always onChrono C1630 days$65.00
Calculated from current plan prices: each hour at the plan’s hourly price, with a started hour counted as a full hour, never more than its monthly price in a billing period.

How you pay: Every server is billed by the hour: the plan’s hourly rate is deducted from the server’s prepaid balance for every hour it exists, powered on or off, until you delete it. Billing is by the hour: every hour a server exists is charged at the plan’s hourly rate. The price tapes and cost tables on this site count every started hour as a full hour, so they show the most a duration can cost; the cost calculator charges a partial hour to the nearest cent, as the bill does. Every plan is capped at 500 hours per billing period: the monthly price is 500 times the hourly rate, and after 500 billed hours (about 20.8 days) the rest of the period is free, so a server left on all month pays exactly the monthly price and never more. A billing period is one month from your order date, not the calendar month, and the cap resets every period. Worked examples are in hourly VPS billing explained.

Delete, don’t stop: A stopped server is still billed, because its vCPU, memory, disk and IP addresses stay reserved for you; only deleting the server stops billing. When a test is done, copy off anything you need and delete the server; the delete VPS checklist covers what to save and revoke first.

Keeping a test server for production takes no new order. Each plan is one product, so there is no billing mode to pick or switch: the same server is billed by the hour whether you keep it for an hour, a day or all month, and the monthly cap applies by itself. A private API that answers requests around the clock is what a monthly VPS means here, with no contract and no prepaid term. A batch that runs over a weekend is 48 to 72 hours of hourly billing, which is all a daily VPS is. The hourly vs monthly VPS guide compares the cap with providers that bill every hour, and the VPS cost calculator prices any other schedule. Servers deploy in Istanbul today; New York is coming soon.

Troubleshooting Ollama on a CPU VPS

Ollama’s troubleshooting docs read the service log on systemd systems with:

journalctl -u ollama --no-pager --follow --pager-end
SymptomLikely causeFix
Install stops: “This version requires zstd for extraction”The zstd tool is missingsudo apt install zstd, then run the script again.
Installer warns “No NVIDIA/AMD GPU detected” or “Unable to detect NVIDIA/AMD GPU”Normal on a CPU VPSNothing to fix. Ollama runs in CPU-only mode.
A request fails while a large model loads, or the service restartsNot enough RAM for the model plus its contextCheck sudo journalctl -k | grep -i "out of memory". Use a smaller model, a shorter context or the next plan up.
Long prompts or chats lose their beginningThe context window (4,096 tokens by default on CPU) is shorter than the inputRaise OLLAMA_CONTEXT_LENGTH or num_ctx; check CONTEXT in ollama ps.
403 Forbidden through a reverse proxyThe proxy forwards a public Host headerheader_up Host {upstream_hostport} in Caddy, or proxy_set_header Host localhost:11434 in Nginx.
A Docker container cannot reach host.docker.internal:11434Ollama listens on 127.0.0.1 only, by designUse --network=host and http://127.0.0.1:11434 as above. Do not rebind Ollama to 0.0.0.0.
Only half the vCPUs are busy in topllama.cpp counted physical cores, and the guest shows 2 threads per coreTry num_thread equal to your vCPU count and compare eval rates.
Long jobs slow down on a Quartz planSustained full load on shared vCPU may be limited (AUP section 8), or other VMs take CPU time (st in top)Move the job to a Chrono plan.
The disk fills upModels accumulate in /usr/share/ollama/.ollama/modelsollama ls, then ollama rm the ones you no longer use.
Causes and fixes from Ollama’s docs, its v0.35.1 source and the sections above.

Ollama has had serious security bugs before; CVE-2024-37032, fixed in version 0.1.34, was a path traversal through model digests the server did not validate. Re-run the install script now and then to update, and keep port 11434 private so the next bug is not reachable from the internet. For the rest of the server, keep the automatic security updates from our checklist switched on.

Deploy this setup

Test Ollama models on a CPU-only server

Chrono C16 · 4 dedicated vCPU · 16 GB RAM · 200 GB NVMe · Istanbul

  • Per hour$0.13/hourFor this job
  • Per day (24 h)$3.12/day
  • Monthly cap$65.00/month
Deploy Chrono C16

Starts with a $5 initial credit, which goes into the server’s balance and pays for its hours.

Billed by the hour, never more than $65.00 per billing period. Delete the server and billing stops.

FAQ

Can Ollama run without a GPU?

Yes. On a server with no supported GPU, Ollama runs in CPU-only mode and loads models into system RAM; ollama ps shows 100% CPU in the PROCESSOR column. Generation is slower than on a GPU, so small models and batch jobs suit it.

How much RAM does Ollama need?

Plan for the model file, plus the context cache, plus about 1 GB for the system. An 8B model’s default 4-bit download is 4.9–5.2 GB, so it fits an 8 GB server at the 4,096-token default context; long contexts for agents need 16 GB or more.

Is it safe to set OLLAMA_HOST=0.0.0.0 on a VPS?

No. Ollama’s docs say the local API does not require authentication, so anyone who reaches port 11434 can run, pull or delete models. Keep it on 127.0.0.1 and connect through an SSH tunnel, a VPN or a reverse proxy that adds a password.

How many CPU threads does Ollama use?

Unless you set num_thread, Ollama 0.35.1 lets llama.cpp choose, and on Linux it uses the number of physical cores it sees. If lscpu shows 2 threads per core, try num_thread equal to your vCPU count and compare the eval rate.

Which Ollama model is best for a CPU-only server?

The smallest one that does your job well enough, in its default 4-bit build: 1–4B models such as llama3.2:3b or qwen3:4b for tagging and drafts, 8B models such as llama3.1:8b for better answers. Test two or three with the same prompt and compare quality and eval rate.

How do I check Ollama’s speed on my server?

Run ollama run MODEL --verbose and read the eval rate printed after each answer, in tokens per second. Through the API, divide eval_count by eval_duration, which is in nanoseconds, and multiply by one billion.

Does Ollama send my prompts to ollama.com?

Not for local models: Ollama’s FAQ says it does not see your prompts or data when you run locally. Cloud models are processed by Ollama’s service; set OLLAMA_NO_CLOUD=1 to turn cloud features off on a server.

Sources

  1. Linux install, service and updatesOllama · docs.ollama.com · checked
  2. FAQ: server configuration, OLLAMA_HOST, concurrency and K/V cacheOllama · docs.ollama.com · checked
  3. Context lengthOllama · docs.ollama.com · checked
  4. API authenticationOllama · docs.ollama.com · checked
  5. Docker: CPU-only containerOllama · docs.ollama.com · checked
  6. API reference: generate, options and embeddingsOllama (GitHub) · github.com · checked
  7. Troubleshooting: logs and CPU librariesOllama · docs.ollama.com · checked
  8. Qwen3 tags and file sizes (Ollama library)Ollama · ollama.com · checked
  9. Open WebUI: server connection errorOpen WebUI · docs.openwebui.com · checked
  10. Optimizing inferenceHugging Face Transformers docs · huggingface.co · checked
All posts