Load testing with k6 from a VPS takes five steps: deploy a short-lived Linux server, install Grafana k6 from its official apt repository, run a script whose stages shape the load and whose thresholds decide pass or fail, read the summary, then copy the results off and delete the server. Run it only against systems you own or have written permission to test.
A VPS gives the load generator datacenter bandwidth, CPU that your browser and IDE don’t share, a fixed IP address the target’s team can allowlist, and a bill that stops when you delete it. This guide uses k6 v2.3.0 (released September 21, 2026) on Ubuntu 24.04 LTS, with commands from the official k6 documentation as of October 3, 2026.
Key takeaways
- Generate load from a separate server, never from the one you are testing, and keep the load generator under 80% CPU and 90% memory, as the k6 documentation recommends.
- Install k6 v2.3.0 on Ubuntu 24.04 from Grafana's apt repository, smoke-test every script with k6 run --once, and set K6_WEB_DASHBOARD_HOST=127.0.0.1 because the dashboard otherwise listens on every interface.
- Stages shape the load and thresholds decide pass or fail; k6 exits with code 99 when a threshold fails. Measure latency on successful responses only and add an abortOnFail error-rate threshold as a kill switch.
- Load test only systems you own or have written permission to test, from a fixed IP address the owner knows about.
- A three-hour session is billed as three hours at the plan's hourly rate: copy the HTML reports and JSON off the server, then delete it, because a stopped server is still billed.
Why do load testing with k6 from a VPS instead of your laptop?
k6 runs anywhere. What matters is whether the machine that generates the load distorts what you measure. The k6 documentation asks you to keep the load generator at or below 80% CPU: when k6 uses all of the CPU, it throttles itself and can report response times much longer than they really are. A laptop that is also running a browser, an IDE and a video call can cross that line without you noticing.
| Your laptop | A short-lived VPS | |
|---|---|---|
| CPU and memory | Shared with everything else you have open | Used by k6 alone; Chrono plans have dedicated vCPUs |
| Network path | Wi-Fi plus a home or office uplink | Datacenter uplink; you can watch it with iftop |
| Source IP | Changes with your network | One dedicated IPv4 address plus IPv6 for the whole session, easy to allowlist |
| Location | Wherever you are sitting | Where your host has data centers (for us, Istanbul today; New York is coming soon), near your users or near the server |
| Long runs | Stop when the lid closes or the Wi-Fi drops | Keep running in tmux after you disconnect |
| Cost | Nothing extra, but noisy results | Billed by the hour until you delete it |
Don’t run k6 on the server you’re testing, either. It would compete with your app for the same CPU and memory, and the numbers would describe neither of them.
Choose the location by the question you’re asking. To measure raw capacity without internet latency, generate load from the same region as the target. To see what users in a region experience under load, generate it from near them. Our guide to measuring latency before you pick a VPS location shows how to check the round-trip time first; HourlyVPS deploys in Istanbul today, and New York is coming soon.
Only test what you own: permission, scope and a kill switch
A load test and a denial-of-service attack send the same kind of traffic. The difference is authorization. Without it, the same test can be an offense under computer-misuse laws such as the US Computer Fraud and Abuse Act (18 U.S.C. § 1030) or section 3 of the UK Computer Misuse Act 1990, which covers unauthorized acts that impair a computer’s operation. This is not legal advice; ask a lawyer if you’re unsure.
Our rule: The HourlyVPS acceptable use policy forbids flooding any network or system and says load tests against systems outside our network need the written permission of the system’s owner and must not affect our network. Keep that permission where you can show it if we ask. A developer’s “sure” in chat does not cover infrastructure that a client, a hosting company or a CDN owns.
Before the first request, write these down and get them signed off:
- Targets: the exact hostnames and IP addresses in scope, and nothing else.
- Owner approval: from whoever owns the infrastructure. On shared hosting that is the hosting company, because your neighbors share the server. If the target runs on a cloud platform, read its testing policy too; AWS, for example, publishes an Amazon EC2 Testing Policy.
- Window and ceiling: start and end time in UTC, plus the maximum VUs or requests per second.
- Source: the VPS’s IPv4 and IPv6 addresses, so the team can recognize or allowlist them.
- Third parties: payment, email, SMS and maps APIs that your app calls receive the load too. Stub them or use their sandbox environments; their terms decide what’s allowed.
- Stop rule: who watches the target, how they reach you, and the error rate at which you stop.
Build the stop rule into the script as well. In the script below, a threshold with abortOnFail ends the run on its own when more than 10% of requests fail, and a custom userAgent tells whoever reads the access logs who is calling. You can stop a run at any time with Ctrl+C. Prefer a staging copy sized like production; if you must test production, pick a low-traffic window and keep the on-call team watching.
What size VPS does k6 need?
The k6 documentation gives three rules for the load-generator machine. Keep CPU at or below 80%. Keep memory below 90%, with no swapping. Expect simple scripts to use about 1–5 MB of RAM per VU, and tens of MB per VU for scripts that upload files or load large JavaScript modules. One k6 process uses all CPU cores, and the docs say a single instance can run 30,000–40,000 VUs when the machine has the resources and follows their tuning guidance.
Memory is the part you can calculate. The table divides 90% of each plan’s RAM by the documented 1–5 MB per VU, before the operating system takes its share. CPU is the part you have to measure, because it depends on your script, on TLS and on response sizes.
| Plan | vCPU | RAM | VU ceiling by memory (5 MB to 1 MB per VU) | Good fit |
|---|---|---|---|---|
| Quartz Q2 | 1 shared | 2 GB | ~370 to ~1,840 | Writing scripts, smoke tests |
| Quartz Q4 | 2 shared | 4 GB | ~740 to ~3,690 | Short average-load tests of a small site |
| Quartz Q8 | 4 shared | 8 GB | ~1,470 to ~7,370 | Short tests that need more VUs |
| Chrono C8 | 2 dedicated | 8 GB | ~1,470 to ~7,370 | Stress and soak tests where steady CPU matters |
| Chrono C16 | 4 dedicated | 16 GB | ~2,950 to ~14,750 | Larger stress tests, TLS-heavy scripts |
| Chrono C32 | 8 dedicated | 32 GB | ~5,900 to ~29,490 | Close to the single-instance range in the k6 docs |
Why Chrono for heavy tests: Quartz plans share their vCPUs, and our acceptable use policy says a server that runs shared vCPU at full load for long periods may have its CPU allocation limited. A limit that kicks in halfway through a stress test would skew every latency number after it. Chrono plans have dedicated, pinned vCPUs built for sustained CPU work, which is what a stress or soak test is. To check a server’s CPU, disk and network before you trust its numbers, run the tests in our VPS benchmark guide first. Compare all sizes on the pricing page.
Calibrate instead of guessing: The k6 docs’ method: run your script with 100 VUs, note k6’s memory use in htop, and multiply by your target VU count divided by 100. If CPU passes 80% before you reach the target, deploy a larger plan or split the load across two servers.
How many VUs do you need?
With stages, each VU loops: send a request, wait for the response, sleep, repeat. Throughput is therefore about VUs ÷ iteration duration. To reach 100 requests per second with one request per iteration, 1 s of sleep and roughly 0.2 s of response time, you need about 100 × 1.2 = 120 VUs. When the target slows down, iterations get longer and the request rate falls. The k6 docs call this the closed model. If you need a fixed request rate whatever the response time, use the ramping-arrival-rate executor instead.
Estimate the transfer too. After the smoke test, data_received divided by iterations gives bytes per iteration. At 200 KB per iteration, 50 VUs and 1.2 s per iteration, that is about 8.3 MB/s, or roughly 30 GB per hour. Istanbul plans include the monthly traffic allowance listed for each plan, prorated for a server that exists for part of a billing period. The same bytes leave the target as outbound traffic, so on a cloud that bills egress, the test costs the target’s owner money as well.
Set up the VPS and install k6 on Ubuntu 24.04
Deploy an Ubuntu 24.04 server, connect to it over SSH and work as a sudo user, not root. k6 needs no inbound ports, so the firewall only has to allow SSH:
sudo ufw allow OpenSSH
sudo ufw enableThe full baseline, with SSH keys only and root login off, is in our VPS security checklist. Since k6 v2.0.0, the REST API server that older versions started on localhost:6565 stays off unless you pass --address.
Step 1. Add Grafana’s signing key and the k6 apt repository, exactly as the k6 install docs give them:
curl -fsSL https://dl.k6.io/key.gpg | sudo gpg --dearmor -o /usr/share/keyrings/k6-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/k6-archive-keyring.gpg] https://dl.k6.io/deb stable main" | sudo tee /etc/apt/sources.list.d/k6.listStep 2. Install k6, the monitoring tools the k6 docs recommend (htop for CPU and memory, iftop for network) and tmux, which keeps a test running if SSH drops:
sudo apt-get update
sudo apt-get install -y k6 htop iftop tmuxStep 3. Check the version. It should report v2.3.0 or later, the version in the repository on October 3, 2026:
k6 versionDebian 12, ARM and Docker: Debian 12 uses the same commands; if a minimal image has no gpg, run sudo apt-get install -y gnupg first. The k6 apt repository publishes amd64 packages only (checked October 3, 2026), which covers every HourlyVPS plan (Intel Xeon, x86-64); on ARM machines, use the binary from the k6 GitHub releases page. With Docker, the Running k6 page pipes the script in on stdin: docker run --rm -i grafana/k6 run -e BASE_URL=https://staging.example.com - <load-test.js. Our Docker install guide covers the setup.
Only for thousands of VUs: raise the OS limits
The k6 docs say you can skip this step until you see a “too many open files” error. For large tests, apply their network settings:
sudo sysctl -w net.ipv4.ip_local_port_range="1024 65535"
sudo sysctl -w net.ipv4.tcp_tw_reuse=1
sudo sysctl -w net.ipv4.tcp_timestamps=1Then raise the open-files limit in the same shell that will run k6:
ulimit -n 250000The ulimit value applies to that shell only, and the sysctl values reset on reboot, which suits a server you will delete anyway. If ulimit fails with “Operation not permitted”, the hard limit is lower; check it with ulimit -Hn and follow the k6 Fine-tune OS guide. Check for swap with swapon --show as well: the docs recommend disabling it, because a swapping load generator produces results you can’t trust.
Write a k6 script with stages and thresholds
Create a working folder and open a new file:
mkdir -p ~/k6 && cd ~/k6
nano load-test.jsPaste this average-load test. It ramps to 50 VUs over 2 minutes, holds for 10 and ramps down over 2, and its thresholds turn the run into a pass or a fail:
import http from 'k6/http';
import { check, sleep } from 'k6';
// Pass the target on the command line: k6 run -e BASE_URL=https://staging.example.com load-test.js
const BASE_URL = __ENV.BASE_URL;
if (!BASE_URL) {
throw new Error('Set BASE_URL, for example: -e BASE_URL=https://staging.example.com');
}
export const options = {
stages: [
{ duration: '2m', target: 50 }, // ramp up to 50 VUs
{ duration: '10m', target: 50 }, // hold at 50 VUs
{ duration: '2m', target: 0 }, // ramp down
],
thresholds: {
// Pass/fail: fewer than 1% failed requests; stop early if more than 10% fail
http_req_failed: [
'rate<0.01',
{ threshold: 'rate<0.10', abortOnFail: true, delayAbortEval: '1m' },
],
// Pass/fail: 95% of successful requests under 500 ms, 99% under 1.5 s
'http_req_duration{expected_response:true}': ['p(95)<500', 'p(99)<1500'],
checks: ['rate>0.99'],
},
discardResponseBodies: true,
userAgent: 'acme-loadtest/1.0 ([email protected])',
};
export default function () {
const res = http.get(`${BASE_URL}/`, { tags: { name: 'home' } });
check(res, {
'status is 200': (r) => r.status === 200,
});
sleep(1);
}What each part does:
stagesis a shortcut for one scenario with the ramping-vus executor: k6 moves the VU count in a straight line to eachtargetover itsduration.http_req_failedcounts a request as failed when its status is outside 200–399 or no response arrives, which is the k6 default.rate<0.01allows less than 1% failures.- The
abortOnFailthreshold is the kill switch.delayAbortEvalwaits a minute so warm-up noise can’t end the test early. http_req_duration{expected_response:true}measures only successful responses. Plainhttp_req_durationincludes failures too, so a server answering 503 in a few milliseconds can make a failing test look fast. Percentiles are in milliseconds.checks: a failed check alone never fails the run; the threshold on thechecksmetric does.discardResponseBodiessaves memory on the load generator. If one request needs its body, setresponseType: 'text'on that request.userAgentreplaces the defaultGrafana k6/<version>with a name and contact the target’s team will recognize.BASE_URLcomes from-eon the command line. The script refuses to start without it, so it can never hit a default host by mistake.sleep(1)is think time. Without it, every VU fires requests back to back.
Add the pages and API calls your users really hit, each with its own tags: { name: '...' }, so you can set a threshold per endpoint. For URLs with IDs in them, the name tag also stops k6 from creating a new time series for every unique URL. If a CDN or page cache serves the home page, a test of / measures the cache, not your application, so include requests the cache can’t answer: search, a logged-in page with a test account, or an API call.
Run it: smoke test, then average load, then stress
The k6 docs name six test types and say to start with a smoke test before any higher load. Choose the type by the question you need answered:
| Test type | Question it answers | Load and duration (k6 docs) | Load generator |
|---|---|---|---|
| Smoke | Does the script work at all? | 2–20 VUs, 30 seconds to 3 minutes | Any plan; Quartz Q2 is enough |
| Average-load | Does it hold normal traffic? | Average production load, 5–60 minutes | Quartz Q4 for small sites, Chrono C8 above that |
| Stress | What happens above normal? | Above average, 5–60 minutes | Chrono C8 or C16 |
| Spike | Does it survive a sudden surge? | Very high, a few minutes | Chrono, sized for the peak VUs |
| Soak | Does it degrade over hours? | Average load, hours | Chrono; a full day is 24 billed hours (see the cost section) |
| Breakpoint | Where does it break? | Rises until something fails | Chrono, watching k6’s own CPU |
Spike and breakpoint tests are meant to push the target past its limits, so name them in the sign-off. Start a tmux session so the test survives a dropped SSH connection:
tmux new -s k6Step 1: smoke test. One VU, one iteration. The --once flag is new in k6 v2.3.0 and keeps the script’s other settings:
k6 run --once -e BASE_URL=https://staging.example.com load-test.jsThen a short low-load run; the k6 smoke-testing guide suggests 2–20 VUs for 30 seconds to 3 minutes. The --vus and --duration flags replace the stages for this run:
k6 run --vus 3 --duration 1m -e BASE_URL=https://staging.example.com load-test.jsStep 2: average load. Turn on the live web dashboard and bind it to the server itself; the dashboard section below explains why the host setting matters:
export K6_WEB_DASHBOARD=true K6_WEB_DASHBOARD_HOST=127.0.0.1Run the script as written (14 minutes), with an HTML report at the end and every data point saved as gzipped JSON:
K6_WEB_DASHBOARD_EXPORT=average.html k6 run --out json=average.json.gz -e BASE_URL=https://staging.example.com load-test.jsStep 3: stress. Override the stages from the command line. This is the profile from the k6 stress-testing guide: up to 200 VUs over 10 minutes, hold for 30, down over 5:
K6_WEB_DASHBOARD_EXPORT=stress.html k6 run --stage 10m:200 --stage 30m:200 --stage 5m:0 --out json=stress.json.gz -e BASE_URL=https://staging.example.com load-test.jsDetach from tmux with Ctrl+B, then D, and reattach later with tmux attach -t k6. In a second SSH session, watch the load generator with htop and sudo iftop while the test runs. Ask the target’s team to watch their side at the same time (CPU, database connections, error logs), so each stage can be matched to the resource that runs out first.
Watch it live: the web dashboard over an SSH tunnel
With K6_WEB_DASHBOARD=true, k6 serves a live dashboard on port 5665. The k6 docs list localhost as the default host, but in the k6 v2.3.0 source code the default host is empty, which means every network interface, even though the startup banner prints http://127.0.0.1:5665. That is why Step 2 sets K6_WEB_DASHBOARD_HOST=127.0.0.1; the firewall rule above, which allows only SSH, is a second layer. Forward the port over SSH from your laptop instead of opening 5665:
ssh -L 5665:127.0.0.1:5665 [email protected]Then open http://localhost:5665 in your laptop’s browser. Two details from the docs: k6 waits to exit while a dashboard tab is still open, so close it when the test ends; and the exported HTML report includes graphs only when the test runs longer than three times the refresh period, which is 10 seconds by default.
Split the load across two servers
One k6 process already uses every core, so a second server is for more load than one machine can produce, or for a second location (until New York opens, that means a server at another provider). Split the test with execution segments: run the first command on one server and the second on the other, so each generates half of the load:
k6 run --execution-segment "0:1/2" --execution-segment-sequence "0,1/2,1" -e BASE_URL=https://staging.example.com load-test.js
k6 run --execution-segment "1/2:1" --execution-segment-sequence "0,1/2,1" -e BASE_URL=https://staging.example.com load-test.jsThe docs list the limits. No primary instance coordinates the others, so start both at the same moment. Each instance evaluates thresholds on its own half only. You also have to combine the metrics yourself, for example from the two JSON files.
How to read k6 results
When the run ends, k6 prints a summary in compact mode by default; --summary-mode=full adds the connection-timing metrics and per-group and per-scenario results. Read it from the top:
| Summary line | What it means | What to look for |
|---|---|---|
| THRESHOLDS ✓ / ✗ | Pass or fail for each rule you set | Any ✗ fails the run, and k6 exits with code 99 |
checks_ | Share of checks that passed | Below 100%: find the failing check by name |
http_ avg, med, p(90), p(95), max | Latency across all requests; the percentiles show the slow tail | Compare p(95) with your target; an average can hide a slow tail. Add p(99) with --summary-trend-stats |
{ expected_ | The same latency, successful requests only | A big gap from the line above means fast errors |
http_ | Share of requests outside 200–399 or with no response | Rising as VUs rise: you found a capacity limit |
http_ | Total requests and requests per second | The throughput you actually reached |
http_ (full) | Time to first byte: the server’s processing time | Grows with load: the app or database is the bottleneck |
http_, http_ (full) | Time to open TCP and TLS connections | Spikes: connection limits or TLS cost on the target |
http_ (full) | Time waiting for a free TCP connection slot before sending | Usually near zero; if it grows, check connection limits on both machines |
vus, iterations | VU count and completed loops | Iterations slowing at the same VUs: the target is slowing down |
data_, data_ | Bytes transferred | Feed your traffic estimate |
The exit code tells you the verdict without reading anything: k6 returns 0 when every threshold passes and 99 when a threshold fails, so echo $? right after the run, or a CI job, can act on it. The same script can gate deployments from a self-hosted GitHub Actions runner. The summary is printed only once, so keep the HTML report and the JSON file; for a summary file, export a handleSummary() function as shown in the k6 custom-summary docs.
Is the bottleneck the target or the load generator?
Before you blame the application, rule out the machine that generates the load:
| Symptom | Most likely side | What to do |
|---|---|---|
read: connection reset by peer | Target: the server or load balancer can’t handle the traffic | Note the VU level where it starts; check the target’s connection limits |
context deadline exceeded | Target: no response within k6’s default 60-second timeout | Read the app and database logs at that timestamp |
dial tcp ...: i/ | Target: it can’t accept new TCP connections | Look for a firewall, WAF or full connection backlog |
socket: too many open files | Load generator: file-descriptor limit | Raise ulimit -n as shown above |
k6 above 80% CPU in htop | Load generator | Use a larger Chrono plan, or split across two servers. The docs also suggest fewer checks, custom metrics and abortOnFail thresholds at large scale |
| Memory above 90%, or swapping | Load generator | Keep discardResponseBodies, share test data with SharedArray, or size up |
iftop flat at the port’s top speed | The network between them | You are measuring bandwidth, not the app |
| 403 or 429 responses after a few seconds | A CDN, WAF or rate limiter in front of the target | Have the owner allowlist the VPS IP for the window, or test the origin with their approval |
The docs also put errors in proportion: at large scale some errors always appear, and 100 failures in 50 million requests is generally a good result. Decide on your acceptable error rate before the test and write it into the http_req_failed threshold.
What does a 3-hour load test cost?
A typical session fits in three hours: setup, a smoke test, the 14-minute average-load run, the 45-minute stress profile, time to read and rerun, and clean-up.
| Time | Step |
|---|---|
| 0:00–0:15 | Deploy, set up the firewall, install k6 |
| 0:15–0:30 | Smoke test and a 100-VU calibration run |
| 0:30–0:50 | Average-load test (14 minutes) |
| 0:50–1:40 | Stress test (45 minutes) |
| 1:40–2:40 | Read the results, fix the script or the app, rerun |
| 2:40–3:00 | Copy the reports to your laptop and delete the server |
Three hours: On Chrono C8, this session costs $0.21. A lighter test on Quartz Q4 costs $0.09. Billing is by the hour: every hour a server exists is charged at the plan’s hourly rate. The price tapes and cost tables on this site count every started hour as a full hour, so they show the most a duration can cost; the cost calculator charges a partial hour to the nearest cent, as the bill does. Charges are deducted in whole cents: a full hour costs exactly the hourly rate, and a partial first or last hour is rounded to the nearest cent, never more than a full hour.
| Duration | Hours on the meter | Cost $0.07 | Note |
|---|---|---|---|
| 1 hour | 1 | $0.07 | |
| 3 hours | 3 | $0.21 | |
| 8 hours | 8 | $0.56 | |
| 1 day | 24 | $1.68 |
There is no billing mode to choose. Every session is metered by the hour from the server’s prepaid balance, a full day of soak testing is simply 24 hours at the hourly rate, and a server kept all month never costs more than the plan’s monthly price in a billing period (one month from your order date). How hourly VPS billing works explains the monthly cap and the initial credit each new server is ordered with, and the VPS cost calculator prices your own schedule.
Is Grafana Cloud k6 cheaper than a VPS?
For small, occasional tests it can be. Grafana’s managed service is the main alternative, and as of October 3, 2026, its pricing page lists a free tier of 500 virtual-user hours (VUh) a month, and a Pro tier from $0.15 per VUh with a $19 monthly platform fee that includes 500 VUh. Grafana’s Performance Testing pricing docs count VUh as maximum VUs × test minutes ÷ 60, with a minimum of 1 VUh per test. The 45-minute, 200-VU stress test above therefore uses 150 VUh, so the free tier covers three such runs a month; the 14-minute, 50-VU average-load test uses about 12 VUh.
| Grafana Cloud k6 | k6 on your own VPS | |
|---|---|---|
| Billing unit | VUh: maximum VUs × minutes ÷ 60 | Server hours, however many VUs you run |
| Free allowance | 500 VUh a month | None; prepaid credit from $5 |
| Where the load comes from | Grafana’s load zones | Your server’s own IP, in Istanbul (New York coming soon) |
| Results | Stored and shared in Grafana Cloud | Terminal summary, HTML report and JSON that you keep |
| Setup per session | None | About 15 minutes |
Grafana Cloud k6 is the better fit for small, occasional tests that stay inside 500 VUh, for load from many regions at once, and for teams that want results stored and shared without managing servers. A VPS fits when tests are larger or more frequent, when the target accepts traffic only from an allowlisted IP, or when you want the load to come from a specific city. The two also combine: k6 cloud run --local-execution runs the test on your server and streams the results to Grafana Cloud, which bills locally run tests at 25% fewer VUh.
Copy the results, then delete the server
Deleting a server is immediate and can’t be undone, and our terms say deleted data, snapshots included, can’t be recovered. Copy the results first. This command, run on your laptop, copies the whole ~/k6 folder with the script, the HTML reports and the JSON files:
scp -r [email protected]:k6 ./k6-resultsThen delete the server in the portal. Stopping it is not enough: A stopped server is still billed, because its vCPU, memory, disk and IP addresses stay reserved for you; only deleting the server stops billing. Run through the checklist before you delete a VPS for anything else worth keeping, such as shell history or a tuned sysctl setting you want to reuse.
Remove the allowlist entry: Ask the target’s team to remove the VPS IP from every allowlist the day you delete the server. Our terms of service say a deleted server’s IP addresses return to the pool and may be assigned to another customer. An entry left behind would trust a stranger.
Deploy this setup
Run a 3-hour k6 load test
Chrono C8 · 2 dedicated vCPU · 8 GB RAM · 100 GB NVMe · Istanbul
- Per hour$0.07/hourFor this job
- Per day (24 h)$1.68/day
- Monthly cap$35.00/month
Starts with a $5 initial credit, which goes into the server’s balance and pays for its hours.
Billed by the hour, never more than $35.00 per billing period. Delete the server and billing stops.



