A VPS for AI agents is a small, always-on Linux server that runs the agent itself (its loop, tools, memory and sandbox) while the language model usually runs at an API provider such as Anthropic or OpenAI. Because the model runs elsewhere, most agents need CPU and RAM, not a GPU: OpenClaw documents 1 vCPU and 1 GB of RAM as its absolute minimum, and the terminal coding agents Claude Code and Codex CLI document 4 GB.
This guide sizes four kinds of agents from their own documentation, shows where a local model changes the math, and gives a copy-paste security model for agents that run shell commands, with the cost of an evening’s trial on an hourly VPS and of an always-on month.
Key takeaways
- API-based AI agents run their loop on the VPS and the model at the provider, so they need CPU and RAM, not a GPU.
- Documented minimums: OpenClaw 1 vCPU and 1 GB (2 GB+ recommended), Claude Code 4 GB, Codex CLI 4 GB (8 GB recommended), Dify 2 cores and 4 GiB, and 1 GiB per Claude Agent SDK agent.
- Local LLMs change the math: the model must fit in RAM with at least 64,000 tokens of context (Ollama defaults to 4,096 on a CPU server), and CPU speed is capped by memory bandwidth.
- Sandbox shell-capable agents in layers: their own VPS, a user without sudo, systemd limits, outbound firewall rules, scoped and spend-limited keys, and a snapshot before each new task.
- Billing is hourly: a trial or size check costs only its hours, and an agent left on 24/7 never costs more than the monthly price in a billing period.
What does a VPS for AI agents actually run?
An agent is a loop: it sends context to a model, gets back an answer or a tool call, runs the tool, and repeats until the task is done. Anthropic’s hosting guide for its Agent SDK describes every running agent as “a long-lived process tied to local state”: a process that owns a shell, a working directory and session files on disk. That process, not the model, is what you host.
- Runtime: Node.js, Python or a single binary that runs the loop.
- Tools: the shell, git, package managers, a headless browser, sometimes Docker.
- State: session transcripts, memory files, a database or a vector store.
- Front ends: a Telegram or Discord bot, a Slack app, webhooks, a web UI or plain SSH.
With an API-based agent, each reasoning step is an HTTPS request to the model provider, so the server spends much of its time waiting on the network. That is why a modest VPS is enough for most agents, and why a GPU on the VPS would sit idle.
A VPS also stays on when your laptop doesn’t. OpenClaw’s own FAQ answers “laptop or VPS?” with “Want 24/7 reliability? Use a VPS”, and lists sleep, network drops and OS reboots as the cost of running the agent locally.
Four kinds of AI agents you can host on a VPS
Every requirement below comes from the project’s own docs, checked on October 3, 2026. Where a project publishes no minimum, we say so.
Coding agents
Terminal coding agents read a repository, edit files and run your build and test commands. On a VPS they work on a copy of the code, away from your personal machine, and a session keeps going after you close the laptop.
- Claude Code (Anthropic): 4 GB+ RAM, an x64 or ARM64 CPU, Ubuntu 20.04+ or Debian 10+.
- Codex CLI (OpenAI): 4 GB RAM minimum, 8 GB recommended, Ubuntu 20.04+ or Debian 10+.
- Gemini CLI (Google): its installation guide recommends 4 GB+ RAM for short sessions and 16 GB+ for long sessions on large codebases, with Node.js 20+ and Ubuntu 20.04+.
- opencode: an open-source terminal coding agent; its docs list no RAM minimum.
- OpenHands Agent Canvas (beta): a self-hosted control center that runs OpenHands, Claude Code, Codex and other agents. It needs Node.js 24+ and uv; its most isolated mode runs each new conversation in its own Docker container, which needs Docker Engine on the server (see how to install Docker on a VPS).
The agent is rarely the heaviest process. The test suite, the TypeScript compiler or the Docker build it starts can need more memory than the agent itself, so size for the heaviest command it will run. Our Claude Code on a VPS guide walks through a full remote dev box with tmux and SSH.
Workflow automation with AI steps
Workflow tools put an LLM step inside a pipeline: a webhook arrives, an AI node classifies or drafts a reply, and the next node files the ticket.
- n8n: an idle n8n Cloud instance needs about 100 MB, and n8n’s illustrative sizing gives 320 MB to 2 GB of memory. Its Docker Compose guide asks for at least 4 GB of RAM and 2 vCPUs once you add the sandbox that runs n8n Assistant’s AI-generated code. From n8n 3.0, launching October 2026, new installs are Docker-only.
- Dify: at least 2 CPU cores and 4 GiB of RAM, per its README; the quick start runs it with Docker Compose.
- Activepieces: an open-source automation tool with AI agent and MCP features; no minimum is listed in its README.
Step-by-step setup with Docker, Postgres and HTTPS is in how to self-host n8n on a VPS.
Chat and messaging agents
These agents live in the chat apps you already use: they read your messages, run tools and reply.
- OpenClaw: a self-hosted gateway that connects Telegram, WhatsApp, Discord, Slack, Signal and other chat apps to AI agents. Absolute minimum 1 vCPU, 1 GB RAM and about 500 MB of disk; recommended 1–2 vCPU and 2 GB+ RAM. It needs Node 24.16+ or 26.1+, and its recommended OS is Ubuntu LTS.
- Hermes Agent (Nous Research): one gateway process for Telegram, Discord, Slack, WhatsApp, Signal, email and more. Its docs list no RAM minimum, but they require a model with at least 64,000 tokens of context.
- Your own bot: a Discord or Telegram bot that calls an LLM API is an agent with a chat front end. See hosting a Discord bot on a VPS and running a Telegram bot on a VPS.
Agent frameworks and SDKs
If you are building your own agent, the framework is a library inside your app, so the footprint is your code plus whatever you attach to it.
- Claude Agent SDK (Python and TypeScript): Anthropic suggests 1 GiB RAM, 5 GiB disk and 1 CPU per agent as a starting point, and notes that memory grows with session length and tool activity.
- OpenAI Agents SDK, LangGraph and CrewAI: frameworks for single-agent and multi-agent apps.
- Microsoft Agent Framework (Python and .NET): AutoGen’s README says AutoGen is now in maintenance mode and that new users should start with this framework.
Add Postgres, Redis, a vector database or a headless browser, and those, not the framework, set your RAM budget.
Do AI agents need a GPU?
No, not when the model runs at an API provider, which is how every agent above runs by default. You need much more memory, and usually a GPU, only when the model itself runs on your server.
We sell CPU servers, so here is the honest version of the local-model math. Three things decide whether a local model is usable:
- Size. The model has to fit in RAM next to everything else. Ollama’s library lists the default Llama 3.1 8B download at 4.9 GB, gpt-oss-20b at 14 GB, Llama 3.1 70B at 43 GB and gpt-oss-120b at 65 GB.
- Context. Ollama’s docs say agents and coding tools should use at least 64,000 tokens of context, and that more context needs more memory. Ollama sets its default by available VRAM, and below 24 GiB that default is 4,096 tokens, so on a CPU server set
OLLAMA_CONTEXT_LENGTH=64000and confirm withollama ps. Check the model’s own window too: Ollama lists Qwen3 8B at 40K tokens, below both that advice and Hermes Agent’s 64,000-token minimum. - Speed. On a CPU, generating each token means reading roughly all of a dense model’s weights from memory, so memory bandwidth caps the speed. Illustrative arithmetic, not a measurement: a 4.9 GB model on a system that streams 20 GB/s tops out near 4 tokens per second, before any other overhead. Mixture-of-experts models read less per token: gpt-oss-20b has 21B parameters, of which 3.6B are active.
| Model (Ollama default download) | Size | Smallest plan with headroom | Realistic use |
|---|---|---|---|
| Llama 3.1 8B | 4.9 GB | Chrono C16 (4 dedicated vCPU, 16 GB) | Batch jobs: summaries, tagging, classification |
| gpt-oss-20b | 14 GB | Chrono C32 (8 dedicated vCPU, 32 GB) | Batch jobs and private drafting where waiting is fine |
| Llama 3.1 70B | 43 GB | Chrono C64 (16 dedicated vCPU, 64 GB) | Fits in RAM; too slow for an interactive agent loop |
| gpt-oss-120b | 65 GB | None (larger than 64 GB) | OpenAI sizes it for a single 80 GB GPU |
Why Chrono and not Quartz? Our acceptable use policy notes that long full load on shared Quartz vCPU can slow down other customers and may be limited, while Chrono’s dedicated vCPU is designed for sustained CPU work. For the full setup, with RAM sizing per model and an API that listens only on localhost, see Ollama on a CPU VPS.
Note: If your agent needs a local model to answer interactively, keep the agent on a VPS and run the model on a GPU host or through an API. HourlyVPS does not offer GPU servers.
How much RAM and CPU does an AI agent need?
Memory runs out before CPU for almost every agent. The table starts from each project’s documented minimum and adds headroom for what runs next to it; the last column gives the reasoning, so you can adjust it.
| Workload | Documented minimum | Start on | Per hour / monthly cap | Why this size |
|---|---|---|---|---|
| Chat agent or bot gateway (OpenClaw, Hermes Agent, an API-calling Discord or Telegram bot) | OpenClaw: 1 vCPU, 1 GB; 2 GB+ recommended. Hermes Agent: none published | Quartz Q2 (1 vCPU, 2 GB) | $0.02/hour $10.00/month cap | OpenClaw’s recommended 2 GB leaves room for logs, media and several channels. A plain bot without a gateway can start on Q1. |
| One Claude Agent SDK app | 1 GiB RAM, 5 GiB disk, 1 CPU per agent | Q2 for one agent; Q4 for two or three | $0.02/hour $10.00/month cap | Memory grows with session length; keep about 1 GB free for the OS. |
| Terminal coding agent on one repo (Claude Code, Codex CLI) | Claude Code: 4 GB. Codex CLI: 4 GB minimum, 8 GB recommended. Gemini CLI: 4 GB+, 16 GB+ for long sessions | Quartz Q4 (2 vCPU, 4 GB); Q8 for heavy builds or long sessions | $0.03/hour $15.00/month cap | The builds and test suites it runs often need more memory than the agent. |
| Workflow automation (n8n, Dify) | n8n: 4 GB, 2 vCPU with the Assistant sandbox. Dify: 2 cores, 4 GiB | Q4; Q8 with Postgres, Redis and queue workers | $0.03/hour $15.00/month cap | Docker stacks run several containers at once. |
| Browser-driving agent (headless Chromium) | No per-browser figure published | Quartz Q8 (4 vCPU, 8 GB) | $0.05/hour $25.00/month cap | Each headless browser is a full Chromium process tree; measure one run, multiply by concurrency. |
| Several agents 24/7 with a database or vector store | Sum of the parts | Q8, or Chrono C8 (2 dedicated vCPU, 8 GB) when CPU stays busy | $0.05/hour $25.00/month cap | Dedicated vCPU suits sustained load such as embeddings, parsing and builds. |
| Local LLM on CPU (Ollama, llama.cpp) | Model size + context + OS | Chrono C16 (4 dedicated vCPU, 16 GB) and up | $0.13/hour $65.00/month cap | See the GPU section above. |
Several agents on one server: Anthropic’s Agent SDK hosting guide sizes hosts as agents per host = (host RAM − overhead) / per-session RAM ceiling, where the ceiling is the peak memory of one representative session, and calls its 1 GiB starting point “a floor, not the ceiling”. Agents that read untrusted content still belong on separate servers (layer 1 below).
Operating system: Ubuntu 24.04 LTS meets every requirement above. Claude Code, Codex CLI and Gemini CLI list Ubuntu 20.04+, OpenClaw recommends Ubuntu LTS, and the commands below assume 24.04.
Every plan, with its disk and traffic allowance, is on the pricing page. Check your guess on the first day instead of trusting any table, ours included: run the agent through a typical task, then look at memory and per-service usage from a second SSH session.
free -h
systemd-cgtop
If available memory sits near zero or the kernel log shows the OOM killer at work, move up one plan:
sudo journalctl -k | grep -i "out of memory"
Tip: Run the first day one size up, measure, then deploy the size you actually need and leave that one running. Billing is hourly, so a wrong guess costs you hours, not a month.
How to sandbox an AI agent that has shell access
An agent with a shell can do whatever its Linux user can do, and the instructions it follows can come from the content it reads. OWASP ranks this first among LLM risks (LLM01:2025 Prompt Injection) and lists giving an agent more power than it needs as a risk of its own (LLM06:2025 Excessive Agency). Anthropic’s deployment guide gives a concrete example:
If a repository’s README contains unusual instructions, Claude Code might incorporate those into its actions in ways the operator didn’t anticipate.
Anthropic, “Securely deploying AI agents”
The answer is layers, so that one mistake doesn’t turn into a breach. Apply them in this order.
1. One agent, one VPS
A KVM VPS is a full virtual machine. In Anthropic’s comparison of isolation technologies, virtual machines rate “Excellent (with correct setup)”, while plain Docker containers rate “Setup dependent”. Give the agent its own VPS and your laptop, your SSH keys and your production servers start outside its reach.
Then apply the baseline from our new VPS security checklist: SSH keys only, no root login, a firewall and automatic security updates.
2. A dedicated user without sudo
Create a user for the agent to run as, and don’t add it to the sudo group. From your normal sudo account, run:
sudo useradd --create-home --shell /bin/bash agent
Close your own home directory to other users. Ubuntu has made new home directories private (750) since 21.04, so on a fresh image this changes nothing:
chmod 750 ~
If the agent installs its own user service, as OpenClaw’s openclaw onboard --install-daemon and Hermes Agent’s hermes gateway install do on Linux, turn on lingering first. Hermes Agent’s docs use it so that a user service starts at boot and keeps running after you log out.
sudo loginctl enable-linger agent
Do setup work, such as installing the agent and signing in to the model provider, in a login session as that user. sudo -iu agent is fine for installs, but systemctl --user needs the clean login session that machinectl shell opens. It ships in the systemd-container package:
sudo apt install systemd-container
Open the session:
sudo machinectl shell agent@
3. A systemd sandbox with memory and CPU caps
For an agent you write or wrap yourself (a bot, an SDK app, a scheduled worker), let systemd start it as the agent user and fence it in. Keep the API key out of environment variables: the systemd.exec manual for Ubuntu 24.04 warns that environment variables “are not suitable for passing secrets” and recommends LoadCredential= instead. Create a root-only file for the key:
sudo install -d -m 700 /etc/credstore
sudo install -m 600 /dev/null /etc/credstore/anthropic-key
Paste the key into it with sudo nano /etc/credstore/anthropic-key. Then save this unit as /etc/systemd/system/agent.service, replacing ExecStart with your agent’s start command:
[Unit]
Description=AI agent (sandboxed)
Wants=network-online.target
After=network-online.target
[Service]
User=agent
Group=agent
WorkingDirectory=/home/agent/app
ExecStart=/home/agent/app/.venv/bin/python agent.py
LoadCredential=anthropic-key:/etc/credstore/anthropic-key
Restart=on-failure
RestartSec=10
NoNewPrivileges=yes
ProtectSystem=strict
ReadWritePaths=/home/agent
PrivateTmp=yes
MemoryHigh=1200M
MemoryMax=1500M
CPUQuota=80%
TasksMax=512
[Install]
WantedBy=multi-user.target
NoNewPrivileges=yes: the process and all its children can never gain new privileges, so setuid tools such as sudo can’t raise its privileges.ProtectSystem=strictwithReadWritePaths=/home/agent: the whole file system is read-only for the agent except its own home, andPrivateTmp=yesgives it a private /tmp.MemoryHighandMemoryMax: the values shown suit a 2 GB plan; setMemoryMaxto about three quarters of your RAM. If the agent can’t stay under it, the OOM killer acts inside this unit, not on your SSH session.CPUQuota=80%caps CPU time (100% is one vCPU); keep it below your vCPU count so SSH stays responsive.TasksMaxstops runaway process spawning.LoadCredential=: systemd reads the root-only file and exposes a copy only to the agent’s user, in the directory named byCREDENTIALS_DIRECTORY.
Read the key in the agent from that directory. In Python:
import os
from pathlib import Path
def read_credential(name: str) -> str:
"""Return a secret that systemd passed in with LoadCredential=."""
path = Path(os.environ["CREDENTIALS_DIRECTORY"]) / name
return path.read_text(encoding="utf-8").strip()
api_key = read_credential("anthropic-key")
# Pass api_key to your model client explicitly instead of exporting it.
Load the unit and start it:
sudo systemctl daemon-reload
sudo systemctl enable --now agent.service
Follow its log:
sudo journalctl -u agent.service -f
4. Close inbound ports and limit outbound traffic
Keep agent dashboards and gateways off the public internet. OpenClaw’s VPS guide calls loopback plus an SSH tunnel or Tailscale Serve the “secure default”, and its security docs note that the Gateway binds to loopback on a regular host install.
To open a loopback-only dashboard, tunnel it over SSH from your laptop and browse to http://127.0.0.1:18789/. This is the command from the OpenClaw Linux docs, for its default port:
ssh -N -L 18789:127.0.0.1:18789 you@your-server-ip
Outbound traffic matters even more for agents: it decides where a hijacked agent can send data, and whether it can spam or scan others from your IP address, which is your responsibility under our acceptable use policy. Outbound SMTP on port 25 is already closed on new HourlyVPS servers. The rules below close everything else except DNS, HTTP, HTTPS and time sync.
Keep a second SSH session open while you apply them; if you lock yourself out, the VNC console in the client portal still works.
sudo ufw allow 22/tcp
sudo ufw default deny incoming
sudo ufw default deny outgoing
sudo ufw allow out 53
sudo ufw allow out 80/tcp
sudo ufw allow out 443/tcp
sudo ufw allow out 123/udp
sudo ufw enable
If your server gets its IP address over DHCP, also run sudo ufw allow out 67/udp before enabling, and if SSH listens on another port, allow that port instead of 22 (syntax: the ufw manual for Ubuntu 24.04). Git over SSH stops working too, so let the agent push over HTTPS with the scoped token from layer 5. Two limits to know:
- Ports, not destinations. The agent can still reach any HTTPS host. To restrict domains, put a proxy in the path: Claude Code’s sandbox checks each host against an allowlist that starts empty, and Anthropic’s Agent SDK guide recommends “an egress proxy that enforces domain allowlists, injects credentials, and logs requests”.
- Docker bypasses ufw. Docker’s docs state that “Docker and ufw use firewall rules in ways that make them incompatible with each other”: container traffic is diverted before ufw’s rules see it. If the agent’s tools run in containers, filter them in Docker’s DOCKER-USER chain, or give them no network with
--network none, as Anthropic’s hardened container example does.
5. Secrets the agent can’t abuse
Assume the agent can read every secret it can use. The goal is to make each secret worth little if it leaks.
- Model API key: one key per agent, in its own Claude Console workspace with a monthly spend limit (Settings, Workspaces, Spend limits). The Default workspace can’t take limits, so create a new one. On another provider, look for the equivalent per-project limit.
- Code access: a GitHub fine-grained personal access token limited to one repository, the permissions the task needs and an expiry date. GitHub recommends fine-grained tokens over classic ones.
- No personal keys on the box: Anthropic’s dev-container docs say to “avoid mounting host secrets such as ~/.ssh or cloud credential files” and to “prefer repository-scoped or short-lived tokens”. The same goes for a VPS.
- Who can give orders: a chat agent obeys whoever messages it. On most channels OpenClaw answers unknown senders with a pairing code and takes numeric Telegram user IDs in its allowlist. Hermes Agent’s gateway “denies all users who are not in an allowlist or paired via DM” by default. Keep those defaults, and run
openclaw security auditafter OpenClaw config changes.
6. Snapshots and a kill switch
Take a snapshot before you hand the agent a new task, a new tool or a new repository. If it breaks the system, restore instead of debugging. As our security page puts it, snapshots are a rollback tool, not a backup, so keep copies of anything you need somewhere else.
Know your kill switch before you need it:
- Stop the agent:
sudo systemctl stop agent.service, or its own stop command, such ashermes gateway stop. - Revoke or rotate its model key and tokens at each provider.
- Delete the server if you no longer trust it. A stopped server is still billed, because its vCPU, memory, disk and IP addresses stay reserved for you; only deleting the server stops billing. Our checklist before you delete a VPS covers what to copy off first.
Using Claude Code? Turn on its sandbox too
Claude Code’s built-in sandbox is off by default; run /sandbox in a session to turn it on. Sandboxed shell commands can write only to the working directory and a temp directory, and their traffic goes through a proxy that checks each host against an allowlist that starts empty. On Linux it needs two packages:
sudo apt-get install bubblewrap socat
On Ubuntu 24.04 and later, AppArmor can block bubblewrap until you add a small profile; our Claude Code VPS guide has the exact steps. Two caveats from Anthropic’s docs: by default, sandboxed commands can still read most of the machine, including ~/.ssh, and the CLI rejects --dangerously-skip-permissions when launched as root.
Hourly for experiments, a monthly cap for always-on agents
Agent work comes in two shapes. Trials are short: install a new agent, connect a chat app, try it for an evening, decide. Production agents run around the clock. Both use the same hourly billing: a trial costs only its hours, and an always-on agent stops adding charges once it reaches the plan’s monthly price in a billing period (one month from your order date).
- An evening with OpenClaw or Hermes Agent on Quartz Q2 (4 hours): $0.08
- A working day with a coding agent on Quartz Q4 (8 hours): $0.24
- A weekend with a browser agent on Quartz Q8 (2 days): $2.40
- An always-on Telegram agent on Quartz Q2 (one billing period, the monthly cap): $10.00
| Duration | Hours on the meter | Cost $0.03 | Note |
|---|---|---|---|
| 1 hour | 1 | $0.03 | |
| 8 hours | 8 | $0.24 | |
| 1 day | 24 | $0.72 | |
| 2 days | 48 | $1.44 | |
| 7 days | 168 | $5.04 | |
| 30 days | 720 | $15.00 | Capped at the monthly price |
Billing is by the hour: every hour a server exists is charged at the plan’s hourly rate. The price tapes and cost tables on this site count every started hour as a full hour, so they show the most a duration can cost; the cost calculator charges a partial hour to the nearest cent, as the bill does. Every server is billed by the hour: the plan’s hourly rate is deducted from the server’s prepaid balance for every hour it exists, powered on or off, until you delete it.
For a trial, add prepaid credit (minimum top-up $5), deploy, and delete the server when you’re done; ordering a server takes an initial credit, prepaid and used for that server’s hours. A block of days is 24 hours of hourly billing per day, as the daily VPS page shows, and a 24/7 agent is simply a server you leave on: a monthly VPS with an automatic cap and no contract. The hourly vs monthly guide explains why the cap means hourly never costs more than monthly, and the VPS cost calculator prices any duration.
Deploy Quartz Q4 hourly for an agent trialThe server is only half the bill. Anthropic’s hosting guide says “Anthropic token cost typically dominates container infrastructure cost by an order of magnitude or more”, so set the spend limit from layer 5 before you leave an agent running.
Choose the location by where you connect from and what the agent talks to, as our VPS location guide explains; HourlyVPS deploys in Istanbul today, with New York coming soon (see locations). Istanbul plans include the monthly traffic allowance listed for each plan, prorated for a server that exists for part of a billing period.
Browsing agents should expect some websites to challenge or block datacenter IP addresses. Respect each site’s terms and rate limits: an agent that hammers websites can get your address onto a blocklist, and our acceptable use policy makes that your responsibility to stop.
Checklist: before you leave an AI agent running on a VPS
- The model runs at an API provider, or you sized a Chrono plan for a local model and accept CPU speed.
- The plan comes from the sizing table, with one size of headroom on day one.
- Baseline hardening is done: SSH keys, no root login, firewall, automatic updates.
- The agent runs as its own user, without sudo, under a systemd unit with memory and CPU caps.
- Inbound ports are closed, outbound is limited, and a domain allowlist sits in front of agents that read untrusted content.
- The model key is scoped and spend-limited, the GitHub token is repository-scoped, and no personal SSH keys are on the server.
- Only your own user IDs can message a chat agent.
- A snapshot exists, and you know the three-step kill switch.
- Trial servers you no longer need are deleted, not just stopped: a stopped server is still billed.
Deploy this setup
Run an API-based AI agent 24/7
Quartz Q2 · 1 shared vCPU · 2 GB RAM · 50 GB NVMe · Istanbul
- Per hour$0.02/hour
- Per day (24 h)$0.48/day
- Monthly cap$10.00/monthFor this job
Starts with a $5 initial credit, which goes into the server’s balance and pays for its hours.
Billed by the hour, never more than $10.00 per billing period. Delete the server and billing stops.



