Two readings, then the quote
A few seconds of pain at the top of the hour is ordinary on a shared host. Moving for that spike buys a machine you will not fill. Sample the business peak, one-second or ten-second intervals, for at least five weekdays. Keep the overnight backup in a separate file. If two readings line up with p95 latency on those peaks, put the monthly price next to the current cloud bill. If only one reading hits, resize the VM and repeat the same window next week.
A single-tenant machine removes the neighbor from the scheduler. Slow SQL, missing indexes, and lock waits stay where they are. Distance to the facility stays too. Confirm the slowness is CPU and disk before you price the move.
A noisy neighbor shows up as steal
On a KVM-style cloud server the guest can see steal: the vCPU was runnable, and the host gave that time to someone else.
vmstat prints it as st. top prints %st. In Prometheus, rate node_cpu_seconds_total in mode steal, then average across cores.
A dedicated server versus a cloud server, on this column, is whether another tenant can take that time.
Bursts of a few seconds are a neighbor’s spike. Leave the VM there. The reading that matters is steal climbing to about 10% or more, several times, holding for minutes, while your own user and system time are not the whole story and API p95 moves with it. Adding vCPUs on a packed host often lengthens the queue. More cores, same neighbors.
Split the other peak. User plus system near saturation and steal near zero means you used the cores yourself. Raise the cloud size first. If steal returns on the larger size during the same business hours, the host is the constraint. The 10% line is a reading aid. Tighten it when the latency budget is tight. Loosen it for batch work that can wait. The pair that has to hold is several business days, plus p95 moving the same way.
A VPS needs one extra check before you treat it like a cloud VM. A virtualized VPS uses the same steal column, usually with a lower ceiling.
A container VPS has no st. Read nr_throttled and throttled_usec from the CPU controller’s cpu.stat.
Sustained throttling at the peak means the quota is gone. Once that quota is already at the plan’s top and throttling remains, judge it the way you judge steal.
That is the practical split between a VPS and a dedicated server: the VPS is still a slice. The dedicated box hands the scheduler to one tenant.
On the host OS of a whole machine, steal should sit near zero, because no other customer is queued on that CPU. If you then run your own guests, those guests can still show steal. That steal is your split of the 32 cores.
The working set hit the VM ceiling
The working set is the pages the workload keeps touching. Linux will spend idle RAM on cache.
Low MemAvailable with steady latency is usually cache. A “90% memory” chart is a weak reason to move.
Write the reading down when these show up together during business hours. vmstat si and so stay above zero.
Major faults rise with database latency. The kernel log shows an OOM. The buffer pool you want is already larger than the VM’s memory, and the provider’s size list has no larger step you can use.
Ballooning can pressure a guest that does not look full, because the host took pages back. Check the balloon before you call it a working set.
A working set around 32 GB, with no swap and no OOM, is a larger cloud size. When the working set is already at the largest memory that cloud will sell you, the next step is a whole machine with two fixed sizes: 128 GB or 256 GB of DDR4-3200 ECC RDIMM. You pick the size on the order. Sticks are not added later. Size it from the peak working set plus headroom for the OS and page cache. If 128 GB cannot hold the set, the 256 GB machine is $592 a month. If neither size holds it, this hardware will not hold it either.
Low available memory and stable latency: treat it as cache. Sustained paging, an OOM, or a buffer pool larger than the VM: that is the memory reading.
The disk queue is sitting on commits
When the database slows, read the data volume’s queue before you read the NIC.
iostat -xz 1 shows aqu-sz, the average queue, and w_await, write wait.
OLTP cares about fsync at commit time. A backup’s sequential megabytes per second do not explain the afternoon.
At the peak, aqu-sz often above 1, w_await into the tens of milliseconds, and commit time in the slow log moving with it: the queue is holding the database.
Throughput in MB/s can stay modest the whole time.
Network volumes add IOPS credits. When the credits empty, wait jumps, then recovers hours later. Same hour every day, and a process restart changes nothing: that shape is the queue.
If you already bought the top IOPS tier and the wait still repeats on business days, the limit is the shared disk path.
A full disk at night and a calm w_await by day is a backup window. Move the backup.
Low wait with lock time or full scans in the slow log is SQL. Fix the query.
On NVMe, %util near 100 is a poor saturation signal. Keep using queue depth and write wait after you move, so the new disks are not misread.
| At the peak | Stay on the cloud | Price a whole machine |
|---|---|---|
| Steal flickers for seconds, p95 stays put | Keep the size | Reading does not hold |
| Steal hits about 10% for minutes, p95 follows, and extra vCPUs still see it | The resize was tried | A neighbor is keeping the cores |
| User plus system is full, steal is near zero | Add cloud CPU | Wait |
| Container throttle climbs at the peak and the quota is already maxed | Same judgment as steal | Quota is gone and throttle remains |
| Available memory is low, no paging, latency is steady | Call it cache | Reading does not hold |
| Paging or OOM at the peak, buffer pool larger than the VM | No larger cloud memory | Pick 128 GB or 256 GB from the working set |
| Disk is busy only in the backup window | Move the backup | Reading does not hold |
| Data-disk queue often above 1, write wait in the tens of ms, commits slow, top IOPS already bought | The volume tier was changed | The queue is the shared storage |
Bare metal versus a cloud server is the same table. The cloud columns move because another tenant, a credit pool, or a balloon can change them. On a single-tenant machine those three inputs are yours.
Pull the week from the VM you have
You can finish this before anyone orders hardware. The commands sample. The judgment sits outside them.
-
1
Cover five weekday peaks
Run
vmstat 1and keepus,sy,st,si, andso. Export p95 for the same window. Store the backup hour apart from the peak. -
2
Separate cache from the working set
Read
MemAvailablenext to the database buffer pool and the peak RSS. Keep the OOM line from the kernel log. On a container, also read the memory cap. -
3
Record only the data disk
Run
iostat -xz 1. Keepaqu-szandw_awaitfor the data volume, beside commit time in the slow log. Leave the backup disk out of that row. -
4
Three rows, then a decision
Each row is one reading: present or not, which days, whether p95 moved with it. Two yes-rows go to the order page. One yes-row means resize the VM and repeat next week.
vmstat 1
iostat -xz 1
In vmstat, us is user time, sy is system, id is idle, wa is I/O wait, and st is steal. si and so are swap in and out.
On a container, add the delta of nr_throttled beside st. Skip the CPU conclusion if that column is missing.
Latency that is mostly distance will still be there after steal disappears. The machines are in Tokyo, Japan. A path to Japan and a path to a far region are different problems. Locks, missing indexes, and cross-region calls are outside these three readings.
Where the curves should sit after the move
Once two readings hold, check that the hardware can absorb them. The CPU is one AMD EPYC 7543P, 32 cores / 64 threads, 2.8–3.7 GHz. One tenant uses that processor. On the host OS, steal should stay near zero. Guests you create yourself will show steal only for the way you divide those 32 cores.
Memory is 128 GB at $465 a month or 256 GB at $592, both DDR4-3200 ECC RDIMM. Pick from the peak working set. If the set fits in 128 GB, take 128 GB.
Disks ship as 2 × 1 TB NVMe, no hardware RAID, each disk attached directly. Read usable capacity as 2 TB. A queue that came from a network volume usually drops once writes hit local NVMe. If it does not drop, look at SQL, or at whether one disk is short. You can add up to two 2 TB disks at $83 a month each. Those are direct-attached as well. They do not join an array on their own. Software RAID, if you want disk redundancy, brings usable space back to about 1 TB. Keep backups off the machine.
Bandwidth starts at 250 Mbps, unmetered. A disk queue is storage wait. Raise bandwidth to 1 Gbps (+$251 a month) or 2 Gbps (+$749) when the 250 Mbps port itself is full, not because fsync is slow. One IPv4 is included. Extra addresses are $2 a month each, up to 256 on the server, sold one by one. Ubuntu 24.04 is a free choice. Windows Server Standard adds $21 a month. Datacenter adds $27. Linux is SSH. Windows is RDP.
Monthly total is the chassis price, plus disk count times 83, plus the bandwidth step, plus extra IPv4 count times 2, plus the OS add-on. The amount due is that monthly total times the number of months. Monthly, quarterly, half-year, and yearly terms all use that multiple. There is no setup fee. The servers are in Tokyo, Japan. If only the disk queue matched, and CPU and memory were quiet, take the cloud volume to its top tier before you compare it with a machine that starts at $465 a month. A light working set and a short queue usually stay cheaper on the VM. When two readings hold, open the order page and pick 128 GB or 256 GB.
Readings matched. Pick the memory.
128 GB or 256 GB. Disks are 2×1 TB, direct-attached. You can order before you sign in.