Buy the parts
Two purchases. Ten minutes. Then it is yours.
Why bother
What stops most people is not talent. It is that they are frightened of their own laptop, and they are right to be.
Watch somebody start. They get excited about building something, install VS Code on their Windows laptop, bolt an AI extension onto it and open a project. An hour later they are staring at a sidebar full of folders they did not create, being asked whether they trust the workspace, wondering where the files actually went, and hovering over every suggested change trying to decide whether accepting it will quietly break something they could not name if you asked them to.
They are not building. They are supervising — through a letterbox, one file at a time, on the same machine that holds their photos and their tax return. Of course they are careful. Anybody would be careful.
Picture a big red button on a wall, and the whole rest of that wall covered in instructions about what the button does. Most people read the wall. They read it for months and they get very good at reading walls. Then some clown wanders up and starts tapping out a drum rhythm on the button, and by the afternoon he knows things about it that are not written anywhere on that wall. The gap between those two people is not intelligence. One of them pushed the button.
Nobody pushes the button on a machine they are afraid of. So stop using that machine.
Rent a computer that is not your laptop. Give an agent a terminal on it and permission to do whatever it likes. Then describe what you want. That is the whole guide. Everything after this is detail.
So: a cheap Linux box, locked down properly on day one, with Claude Code running in a terminal at the root of the filesystem and the freedom to do its job. You describe things in plain English. It writes the code, installs what it needs, configures the web server, gets the certificate, starts the service and hands you a URL. You will understand about a third of what happened the first time, and a third is the correct amount.
What you actually get
Projects
Nobody charges you per app. The only real limit is memory, and text-shaped things weigh almost nothing.
Subdomains
One domain, infinite names. app., api., demo., whatever. — all free, all with certificates, all yours.
Languages
Go, Rust, Python, Node, C, whatever. The agent installs the toolchain. You are not required to like any of them.
Uptime
It runs while your laptop is shut. That is the whole difference between a demo and a thing that exists.
What this guide assumes about you
That you have used a computer. That is genuinely the whole list. If you have never opened a terminal, have never heard of SSH, could not define a DNS record at gunpoint and could not tell a front end from a back end, you are precisely who this was written for. Every term gets explained the first time it turns up, and none of them are as clever as they sound.
The instructions default to Windows 11, because that is what most people are reading this on, and because everything you need in order to reach the server is already installed — no PuTTY, no downloads, nothing to buy. Where macOS differs there is a tab for it. Linux users already know what to substitute and were going to substitute it anyway. There is also a section on putting Ubuntu on your own laptop, which is worth doing and is not needed for any of this.
The shopping list
Two things you have to buy, one you are probably already paying for.
A virtual private server
A whole Linux computer in a data centre, switched on permanently, with its own address on the internet. Yours until you stop paying.
A domain name
The name people can actually type. Buy one and you get unlimited subdomains off it for nothing.
A Claude subscription
Claude Code comes with every paid plan. Pro is plenty to start with; you will know when you want Max.
Here is the real arithmetic, including the optional bits nobody mentions until the invoice turns up. Untick whatever does not apply to you. I would rather you saw the honest number now than found it in month three.
What a year actually costs
Real prices, checked in September 2026. Untick anything you do not need.
Buying the server
Ten minutes, a card, and an email with a password in it. That is the barrier.
A VPS — virtual private server — is a slice of a big machine in a data centre that behaves exactly like a computer of your own. Its own memory, its own disk, its own operating system, its own address on the internet. You are the administrator. Nobody else can see inside it. It is a desktop tower in somebody else's building, and the somebody else is responsible for the power, the cooling and the network cable.
There are hundreds of companies selling these, and for the same machine the spread between the cheapest and the most expensive is close to tenfold. The one this guide prices against is DediRock's Promo Performance range in New York — not because it won a comparison, but because it is the box this page is served from, and every command in this tutorial was run on it. The survey further down puts it beside eight other providers at the same specification, so you can see where it actually sits rather than taking my word for it. Nothing in this guide depends on who you rent from, and you should shop around.
What you are deliberately not buying is a cloud account. The big platforms are magnificent engineering and they are also the apex of nickel-and-diming: a charge for the machine, a charge for the disk, a charge for the traffic leaving the disk, a charge for the load balancer in front of the machine, and a bill at the end of the month that nobody in the building can fully explain. An unmanaged box at a flat annual price can absorb an amount of traffic that would genuinely surprise you, and the number on the invoice is the same every year.
| Plan | RAM | vCPU | Disk | Bandwidth | Per year |
|---|---|---|---|---|---|
| Core | 2 GB | 1 | 30 GB | 2 TB | $24.88 |
| Plus | 3 GB | 1 | 40 GB | 4 TB | $34.88 |
| Power pick this | 4 GB | 2 | 60 GB | 6 TB | $44.88 |
Why the middle option is a false economy
The ten dollars between Plus and Power buys a second CPU core and an extra gigabyte of memory, and both matter more than they sound. The second core means a build can run while your website carries on answering requests, instead of the site going quiet every time you deploy. The extra gigabyte is what stops the machine grinding when a database, two or three small services and a compiler are all awake at once.
Memory is the constraint that will actually bite you — not disk, not CPU. Sixty gigabytes of disk is more than a hundred text-shaped projects will ever manage to fill. Four gigabytes of RAM is comfortable, and comfortable is not the same as infinite. This is the one number worth ten extra dollars, and it is the only place in this guide where I will tell you to spend more.
One price is not the price
Here is the same machine bought eleven different ways — four gigabytes of memory, or the closest thing each company sells. The cheapest row and the most expensive row are the same computer for your purposes, and there is a factor of ten between them. That spread is the single most useful fact in this section, and it is the reason a guide that names exactly one provider is doing you a disservice no matter how good that provider is.
4 GB of memory, or the closest thing each provider sells. The annual figure is the honest comparator, because several of these are only cheap if you hand over a year at once. A figure in red is not what it looks like — the card underneath says why.
| Provider | Per year | Memory | CPU | Disk | Traffic |
|---|---|---|---|---|---|
| RackNerd XMas Sale 2.5 GB KVM | $28.71$2.39/mo | 2.5 GB | 2 vCPU | 38 GB SSD RAID-10 | 6.5 TB |
| DediRock this box Promo Performance — Power | $44.88$3.74/mo | 4 GB DDR5 | 2 vCPU | 60 GB NVMe | 6 TB |
| OVHcloud VPS-1 | $54.48$4.54/mo | 4 GB | 2 vCore | 40 GB NVMe | unlimited |
| netcup VPS 500 G12 | €59.52€4.96/mo | 4 GB DDR5 ECC | 2 vCore | 128 GB NVMe | included |
| Hostinger KVM 1 | $77.88$6.49/mo | 4 GB | 1 vCPU | 50 GB NVMe | 4 TB |
| Hetzner CAX11 (Ampere ARM) | $83.88$6.99/mo | 4 GB | 2 vCPU ARM | 40 GB NVMe | 20 TB |
| WingsHoster a friend's company Wings VPS 4G | ₹10,685.52₹890.46/mo | 4 GB | 2 Xeon cores | 60 GB NVMe | unmetered |
| BuyVM Slice 4096 | $180.00$15.00/mo | 4 GB | 1 core | 80 GB SSD | unmetered |
| Contabo my other box Cloud VPS 8 (2026), US East | $242.16$20.18/mo | 23 GiB | 8 cores | 300 GB | unstated |
| DigitalOcean Basic Droplet, 4 GB | $288.00$24.00/mo | 4 GB | 2 vCPU | 80 GB SSD | 4 TB |
| Oracle Cloud Always Free — Ampere A1 | $0.00$0.00/mo | 12 GB | 2 OCPU ARM | 200 GB block | 10 TB egress |
Surveyed 14 September 2026 — 3 days ago.
Every row but one was read off that provider's own page, and each card below
carries the link and the day it was read. A few of these companies will not quote a
price to anything that is not a person in a browser, and the cards say where that
changed how a number was checked.
just market-check re-reads all eleven sources and says which numbers
have moved; the test suite fails once this survey is older than
210 days, so a stale table breaks the build rather than quietly
lying to you.
What those rows are actually like — including the two I would not buy, and why they are in the table anyway:
RackNerd XMas Sale 2.5 GB KVM $28.71
The cheapest real thing in this table, and the clearest lesson in it. This is a promotional SKU with a holiday name still on it in September, sold from a separate page that is not linked from their catalogue. The catalogue price for a 4 GB box from the same company is $24.59 a month. Same provider, same hardware, roughly ten times the price, and which one you get depends entirely on which URL you arrived through.
DediRock Promo Performance — Power $44.88
The box this website is served from, which is the only reason it is the one priced throughout this guide. Nothing here depends on it. Every command in this tutorial ran against this plan, and the dashboard is that machine reporting on itself.
OVHcloud VPS-1 $54.48
A company with its own data centres and its own fibre, at a price that sits inside low-end territory. The quoted rate needs twelve months paid upfront; month to month is meaningfully more. The Asia-Pacific regions are the exception to “unlimited” — 500 GB on VPS-1, then the line is throttled to 10 Mbps rather than billed.
netcup VPS 500 G12 €59.52
Three times the disk of anything else at this price, error-correcting memory, and no minimum term at all. Two things about the number. It is in euros, so what leaves your account depends on the day and on your card issuer. And their page quotes you €4.96 or €5.91 for the same machine depending on where it thinks you are — the second one includes 19% German VAT, which a customer outside the EU does not pay. Every European provider does some version of this, and it is the single most common way a price comparison between a US and an EU host comes out wrong.
€4.96 is the rate at 0% VAT, which is what a US buyer pays; the same page shows €5.91 including 19% VAT to a visitor it places in Germany. Both figures are theirs, read the same day.Hostinger KVM 1 $77.88
Read the second price, not the first. The advertised rate is an introductory one on a multi-year commitment and it renews at $11.99 a month — $143.88 a year, three times the DediRock row. It is also one vCPU, which is the floor this section tells you to stay above. In the table for completeness and because their marketing is unavoidable.
Hetzner CAX11 (Ampere ARM) $83.88
The one most people will tell you to buy, and the one whose numbers were hardest to establish. Twenty terabytes of traffic is not a typo. Two things to know: it is ARM, so your Go binary wants GOARCH=arm64 — one word in the build command, and nothing else in this guide changes; and in June 2026 Hetzner raised cloud prices substantially, more than doubling some CPX plans. That is the argument for dating a price rather than trusting one.
Hetzner's own pricing page renders its prices in the browser and quotes nothing at all to a plain HTTP client, so this figure is their published post-adjustment list price for CAX11 rather than a number read off the shop page. Specifications are from their cost-optimized page. Check it yourself before you buy.WingsHoster a friend's company Wings VPS 4G ₹10,685.52
A friend's company, and marked that way in the table for the same reason the DediRock row is marked: you are owed the reason a row might be liked, where you can see it. It came in through the same door as every other row — the price read off its own page, held to the same four gigabytes. Three things to know. The price is in rupees, so, like the netcup row, what leaves a dollar account depends on the day and on your card. Unmetered is on a shared 1 Gbps port, which means nobody counts your traffic and nobody promises its speed either. And it is a small Indian company — MSME-registered, the page says — so read its mention history on the forums below before you give it money, exactly as you would for anyone else in this table. The same company publishes HM360, billing and automation software for people running a hosting business, which is a strange and interesting thing to discover exists once you have a box of your own.
Priced in Indian rupees; the dollar cost moves with the exchange rate and any foreign-transaction fee on your cardBuyVM Slice 4096 $180.00
Four times the price of the DediRock row for the same memory and one core instead of two, and people pay it on purpose. Unmetered means unmetered, block storage is available by the terabyte for a few dollars, and the company has been answering its own forum threads under its own name for over a decade. What you are buying above the specification is that nobody disappears.
Contabo my other box Cloud VPS 8 (2026), US East $242.16
My other box, bought four days before this survey, and marked so you know why it is here. It is not the four-gigabyte machine — it is the plan I actually bought, with nearly six times the memory — and it is in this table for what its invoice teaches rather than for its specification. The plan is $14.28 a month on a twelve-month term with no setup fee. Choosing a location in the United States added $5.90 a month on top of that, which is 41%, and that line appears at checkout rather than on the plan. What actually left my account was $242.16 for the year, up front. The advertised number is real; it is just the price for somewhere you might not want your server to be. Pick the location first and read the total second, with every provider, not just this one.
Contabo answers anything that is not a person in a browser with a security check, so nothing here could read their page and just market-check cannot re-read this row. The price is from my own order confirmation of 10 September 2026, and memory, cores and disk were measured on the machine rather than copied from a listing. Nothing on the order states a traffic allowance, so the table does not guess one.DigitalOcean Basic Droplet, 4 GB $288.00
Included so the rest of the table has a scale. This is the polished, well-documented, entirely respectable mainstream price for the same machine — six times the DediRock row and ten times the RackNerd one. You are paying for a console you will enjoy using, documentation that is genuinely excellent, and an invoice that grows every time you add a thing. It is a good product. It is not a cheap one.
Oracle Cloud Always Free — Ampere A1 $0.00
Genuinely free and genuinely more memory than any paid row here, and there are three catches you need before you build anything on it. Idle instances get reclaimed — Oracle's own documentation says a machine under 20% CPU and under 20% memory across seven days may be taken back. “Out of host capacity” when you try to create the instance is routine rather than exceptional. And it is one account per person, tied to a card for verification. Superb for a practice room; a strange place to put something you would be upset to lose.
Where to look for yourself
This page is a snapshot taken by one person with one set of opinions. The market it is a snapshot of moves weekly — providers appear, run aggressive promotions, get bought, get overloaded, and occasionally stop answering tickets entirely. There is an entire corner of the internet that tracks this in real time, it has been doing it for fifteen years, and you should use it rather than my table.
The floor: one core and one gigabyte is a demo, not a server
Every provider in that table sells something for a dollar or two a month, and for anything you intend to take seriously you should walk straight past all of them. The 1 GB, single-core box is not a smaller version of a real server. It is a different thing that fails in a specific, confusing way, and it fails at exactly the moment you finally have something worth losing.
Here is where a gigabyte actually goes. Ubuntu doing nothing at all is two to three hundred megabytes. nginx is twenty. A small Go service is fifteen, and you will run three of them before the month is out. Postgres on its defaults asks for a couple of hundred more before it has cached a single row. You are now at roughly half your memory and you have not built anything yet.
Then you run a build. A compiler is the most memory-hungry thing that will ever happen on this machine — and an agent working on your code runs builds and test suites continuously, which is the whole point of the box. On a gigabyte, that build does not run slowly. The kernel's out-of-memory killer wakes up, picks the process with the worst score, and destroys it. It usually picks the compiler or the database, so what you observe is not "I am out of memory". It is "my build randomly dies", or worse, "the site went away for thirty seconds and came back", which is a much harder sentence to search for.
One core has a quieter cost too, which the second core in the plan above is there to buy: with a single core, everything queues behind everything. A deploy stalls the website. A certificate renewal stalls the deploy. You cannot tail a log while something compiles, which means you cannot watch the thing you are debugging while it happens.
None of which makes the tiny boxes useless. A 1 GB single-core machine is a perfectly good home for one static site, a chat bot, a webhook receiver, a WireGuard endpoint, a scheduled scraper, or a staging copy of something real. Buy three of them for three separate small jobs and you will be delighted. Just do not put the thing you care about, plus its database, plus an agent doing builds, on one of them — and notice that across that entire table, the gap between the 1 GB box and the 4 GB box is somewhere between fifteen and thirty dollars a year. It is the cheapest insurance in this guide and it is the one upgrade I will push you on twice.
The rule that actually sizes a box: your database should fit in memory
If you take one sizing heuristic away from this page, take this one. It is not precise and it is not the whole truth, and it will get you the right answer far more often than any amount of reasoning about request volume.
While your whole database is smaller than the machine's memory, every read your application makes is answered out of RAM after the first time. Cross that line and reads start going to the disk, and that transition is not a gentle ten or twenty per cent. A page from memory costs tens of nanoseconds. A random read from a shared network-attached disk costs hundreds of microseconds on a good day. That is four orders of magnitude, on the operation your application does most.
The reason it works as a rule of thumb is that it is not really about your database
at all — it is about the operating system underneath it. Linux keeps recently
read file pages in the page cache, using whatever memory nothing
else currently wants. SQLite is a file, so it rides that cache for
free; Postgres keeps its own shared_buffers on top
of it. Either way, the practical question is not "how big is my database" but
"is there room to keep it in memory after everything else has taken its
share" — which is why the answer to a slow database is so often more
memory rather than a cleverer query.
| On a 4 GB box | Reasonable resident size |
|---|---|
| Ubuntu, systemd, sshd, the usual daemons | 250–350 MB |
| nginx, a handful of workers | 20–40 MB |
| Three or four small Go services | 50–120 MB |
| Postgres, tuned for this box | 400–800 MB |
| Headroom for a build, a test run, an agent | 600 MB–1.5 GB |
| Left over to cache your data | 1–2 GB |
So on the box this guide buys, the honest ground floor is roughly this: keep the database under about a gigabyte and you will never think about any of this again. Between one and two gigabytes, you are fine but you should know where the number is. Past that, on four gigabytes of total memory, you are no longer running a comfortable machine — you are running a machine with a performance problem that has not introduced itself yet.
What "it does not fit any more" looks like
It does not announce itself. Nothing errors, nothing crashes, no log line says the word memory. What you get instead is a page that is usually fine and occasionally terrible — the average stays respectable while the slowest one request in a hundred falls off a cliff, because that is the one that needed a page the cache had evicted. Reports arrive as "it felt slow earlier" and you cannot reproduce it, which is the most expensive category of bug there is.
These are the four numbers that answer it, and none of them need a monitoring stack:
# how big is the database, really
du -h /srv/myapp/app.db # SQLite: it is one file
sudo -u postgres psql -c "SELECT pg_size_pretty(pg_database_size('myapp'));"
# how much memory is left to cache it with
free -h # read "available", not "free"
# is Postgres finding what it needs in memory? want 99%+
sudo -u postgres psql -c "SELECT sum(blks_hit)*100/sum(blks_hit+blks_read) AS hit_pct FROM pg_stat_database;"
free -h is available. A box with 100 MB free and
2 GB available is perfectly happy; a box with 2 GB free is a box that is not
caching your data yet.
The honest caveats, because it is a rule of thumb and not a law
It is really the working set, not the whole thing. What must fit is the part you actually touch — your indexes, and the rows people are reading this week. A table carrying four years of finished orders nobody ever queries can be far larger than memory and cost you nothing, as long as the indexes supporting the queries you do run still fit. The reason the cruder rule is the one worth teaching is that the working set is genuinely hard to estimate, it always grows in the direction you were not looking, and being wrong about it is invisible until it is urgent. Size for the whole database and you never have to be right.
Sequential and recent-only workloads are exempt. Append-only data you only ever query by recent window — logs, metrics, events, a queue — is fine at many times your memory, because the hot end is small and the cold end is never read. A hundred gigabytes of last year's telemetry on a 4 GB box is not a problem. A hundred gigabytes of customer records you filter across is a different story entirely.
And when a database genuinely does outgrow the machine, the order of operations is almost always the same: move the blobs out, archive or partition what nobody reads, and only then buy memory. On the prices in the table above, doubling your memory costs between fifteen and forty dollars a year. There is a long tradition of spending a whole weekend tuning around a constraint that could have been bought off for the price of two coffees, and I have contributed to it personally.
Choose Ubuntu, and stop thinking about it
The checkout page will offer you a dropdown with a dozen operating systems in it, and this is the one place where a beginner can lose a weekend to a decision that does not matter. Pick Ubuntu, the newest LTS. Every command in this guide is written for it, every error message you paste into a search box will have been seen by ten thousand people before you, and every answer you find will apply without translation.
Ubuntu gets a lot of hate. Look at the numbers anyway
There is a genre of internet opinion about Ubuntu — the package format argument, the telemetry argument, the it-used-to-be-better argument — and some of it is fair. None of it matters to you this week. The thing that matters is the one thing the criticism accidentally proves: it is the distribution everybody else already assumed you were running, and that is not a popularity contest, it is a compatibility guarantee.
Of the Linux websites that publicly say which distribution they are.
Source: W3Techs, Usage statistics of Linux for websites, read 8 September 2026 — Ubuntu 15.1%, Debian 5.9%, CentOS 1.2% of all Linux sites. The percentages above are those figures re-based on the 23.1% that identify themselves. Read the next paragraph before quoting any of this at anybody.
Percentage of all respondents, so these do not add to 100 — people use more than one.
Source: Stack Overflow Developer Survey 2025, personal use, from more than 49,000 responses. Professional use is almost identical for Ubuntu at 27.7%, which is its own small finding: people do not switch distribution when they clock in.
And it is where the AI tooling actually lives
This is the part nobody mentions in the distribution arguments, and for a beginner in 2026 it is probably the strongest single reason. The entire machine-learning stack — drivers, containers, the notebooks you will paste from, the images the clouds hand you — is built and tested on Ubuntu first. Not by decree. Just by everybody making the same assumption for fifteen years until it became true.
The practical consequence is small and constant. A README says
apt install and you can paste it. An error message has ten thousand
results instead of eight, and the eight are from 2017. A driver has a
.deb and an Ubuntu version number beside it rather than a build
script and an apology. None of that is a technical virtue of the operating system
— it is a property of everyone else, and it is yours for free
because you picked the same box in a dropdown that they did.
None of this is tribal. Debian is superb, Alpine is a marvel in the place it belongs, and anybody telling you the choice defines you is selling something. The argument for Ubuntu here is narrow and practical: it is the one where the answer to your problem already exists in a form you can paste. That is worth more to you this week than any technical difference between them, and you can form an opinion in a year from evidence you gathered yourself.
The rest of the checkout
- Operating system: Ubuntu, the newest LTS As above. If you are also offered a choice of “minimal” or “server”, either is fine — the difference is a handful of preinstalled packages you can add later in one command.
- Location: near your users, not near you The distance between the server and whoever visits your sites is what adds milliseconds. If that is mostly you and your friends, pick the continent you are on and stop thinking about it.
- Skip every upsell Managed support, control panels, backup add-ons, extra IP addresses — you need none of it. You are going to set this machine up yourself, which is the point, and backups get handled properly later with something considerably better than a checkbox on an order form.
-
Wait for the email
Within a few minutes you will get a message containing an IP address
(four numbers with dots, like
203.0.113.40) and a root password. That is the address of your computer and the key to the front door.
Buying the domain
The one purchase where almost everybody gets quietly fleeced — and the one section of this guide where a lot of readers arrive already holding something.
Your server has a number. A domain is the name that points at it. You want one for
the obvious reason — myproject.com reads better than
203.0.113.40 and is possible to say out loud — and for one much less
obvious reason: you cannot get an HTTPS certificate for a bare IP
address. No domain means no padlock, and a browser in 2026 will make a
site without a padlock look like a crime scene.
From here the guide forks, because you are one of three people. Pick your door. All three of them end up in the same place — a name that answers with your server's address — and none of them takes an afternoon.
The easy path, and the cheap one. Buy it in the same place that will host its DNS, in about four minutes.
Buy it from Cloudflare →
Keep it exactly where it is. Point it at your server today and decide about moving it some other week.
Point it, without moving it →
Move the registration to Cloudflare and pay wholesale forever. Ten minutes of work, then five days of waiting.
Transfer it →Door one — buy it from Cloudflare
Nearly every registrar runs the same trick, and it is a good trick: a very cheap first year, then a renewal several times higher, on the entirely correct assumption that transferring a domain is irritating enough that you will look at the invoice, sigh, and pay it. They are not wrong about you. They were not wrong about me.
Cloudflare Registrar does not do this. It sells domains at exactly what the registry charges, with no markup at all, forever. Privacy protection — which keeps your name and home address out of the public WHOIS database — is included rather than sold as an extra. The DNS control panel you need later is the same account.
The trap, in one table
These are real Cloudflare prices as of September 2026. The left column is what a typical registrar advertises. The right column is what you pay every year after the first.
| Domain | First year | Every year after | |
|---|---|---|---|
| .site | $4.99 | $27.70 | 5.5× jump |
| .online | $4.99 | $27.70 | 5.5× jump |
| .space | $4.99 | $25.20 | 5× jump |
| .tech | $9.99 | $49.20 | 5× jump |
Good domains for under ten dollars a year
Same price in year one and in year ten. All of these are perfectly respectable. Nobody has ever decided a piece of software was bad because of the letters after the dot, and anybody who would is not going to be a useful user.
| Domain | Per year | Worth knowing |
|---|---|---|
| .fyi | $5.20 | The cheapest genuinely normal-looking option. Great for tools and documentation. |
| .uk | $5.30 | Short and clean. Fine for a personal project regardless of where you live. |
| .us | $6.50 | Requires a US connection — citizen, resident, or a business operating there. |
| .link | $7.20 | Reads well for anything tool-shaped, and short enough to type on a phone. |
| .cc | $8.00 | Widely used, unrestricted, and never mistaken for a scam. |
| .org | $8.50 → $11.20 | The one exception worth making: first year cheap, then a modest, honest renewal. |
| .click / .page / .day | $10.20 | Flat forever. .page in particular looks deliberate rather than budget. |
| .com | $10.46 | Just over the line, and rising to about $11.15 on 1 November 2026. Still the one people assume. |
For reference, the ones people usually ask about: .net is
$11.86, .xyz renews at
$11.20, .dev is
$12.20, .app is
$14.20, and .io — the one every startup
wants — is $50.00 a year, every year.
.trade, .stream, .bid,
.date and friends at around $4.18. They are cheap because they are
overwhelmingly used for spam, and an enormous number of corporate mail filters
and network blocklists have simply written off the entire extension. Your email
silently vanishes, your links get blocked, and you spend a fortnight debugging a
problem that is not in your code. Spending six dollars instead of four is the
best money in this whole guide.
Buying it
- Make a Cloudflare account Free. Go to dash.cloudflare.com and sign up. Turn on two-factor authentication immediately — this account will control where your domain points.
- Domain Registration → Register Domain Search for the name you want. The price shown is the price, in year one and in year five.
- Turn auto-renew on A lapsed domain can be bought by anybody who wants it, and getting one back is expensive when it is possible at all. This is one checkbox standing between you and a genuinely miserable week.
- Do not touch DNS yet You will point it at your server in the Pointing a domain section, once there is actually something for it to point at.
That is door one finished. Skip to buying the server if you have not yet, or straight on to getting in. The next two headings are for people who arrived holding a domain already.
Door two — you already own one somewhere else
Then you do not have to buy anything and you do not have to move anything. A domain is not welded to the company that sold it to you: who you pay is one question, and who answers when somebody looks the name up is a completely separate one. Almost every beginner conflates those, panics about "transferring", and puts the whole project off for a month.
You have two ways to be live today, and one of them is also step one of door three.
| What you do | Takes | Costs |
|---|---|---|
| 2a. Add two records where you already are | about five minutes | nothing |
| 2b. Move DNS to Cloudflare, keep the registration | ten minutes, then a wait | nothing |
2a is the shortest path to a working site. Every registrar on earth
has a DNS panel, every one of them can hold an A record, and yours is
fine. Log in, find the DNS or "Manage DNS" screen, and add the two records from the
Pointing a domain section. Nothing else about your account changes.
Do this if you want to see your own site on your own name in the next quarter of an
hour, which is a perfectly good thing to want.
2b is what I would actually do, and it is worth being clear about why, because "use Cloudflare's DNS" sounds like brand loyalty and is not. Moving DNS means changing two nameserver entries at your registrar so that Cloudflare, rather than your registrar, is the thing the internet asks. You keep paying the same company the same renewal fee. What you get is a DNS panel that changes in seconds instead of hours, wildcard records that actually work, the free proxy and analytics if you ever want them — and, not incidentally, the exact state Cloudflare requires before it will accept a transfer. Door two done properly is door three already half finished.
Where your registrar hides the two screens
You need exactly two things from whoever you bought the domain from: the DNS records screen (for 2a) and the nameservers screen (for 2b). Every registrar has both, every one of them calls them something slightly different, and all of them move the menus about once a year. Find your logo.
Namecheap
- Domain List → Manage next to your domain.
- DNS records: the Advanced DNS tab. "Host Records" is the table;
@is the bare domain. - Nameservers: the Domain tab, the Nameservers dropdown → Custom DNS, then paste Cloudflare's two.
- Auth code: Domain tab → turn off the transfer lock, then Sharing & Transfer → Transfer Out for the code.
GoDaddy
- My Products (or Domain Portfolio) → the domain → DNS.
- DNS records: the DNS Records list on that page. Delete GoDaddy's parked-page record before adding your own, or you will have two.
- Nameservers: same page, the Nameservers panel → Change Nameservers → I'll use my own.
- Auth code: Domain Settings → Additional Settings → Transfer domain away. Unlock first; the code is emailed.
Porkbun
- Domain Management → the arrow beside your domain to expand it.
- DNS records: DNS Records right there in the expanded row.
- Nameservers: the NS / Authoritative Nameservers field in the same row.
- Auth code: unlock the domain and the Auth Code is shown in the panel rather than emailed. Refreshingly civilised.
Squarespace
- Domains in the account dashboard → the domain.
- DNS records: DNS → DNS Settings → Custom Records.
- Nameservers: the same DNS screen → Nameservers → Use custom nameservers.
- Auth code: the domain's Transfer or Domain Lock panel. Unlock, then request the code.
NameSilo
- Manage My Domains → the domain.
- DNS records: the Manage DNS icon in the domain's row.
- Nameservers: tick the domain, then the Change Nameservers action above the list.
- Auth code: unlock in the domain's Registry Lock setting; the EPP code is on the same page.
Gandi
- Domain list → the domain.
- DNS records: the DNS Records tab.
- Nameservers: the Nameservers tab → External.
- Auth code: the Transfer section → unlock, then reveal the authorisation code.
IONOS
- Domains & SSL → the domain → the gear or … menu.
- DNS records: DNS → Adjust DNS settings.
- Nameservers: the same menu → Nameserver → Use own nameserver.
- Auth code: the domain's menu → Transfer domain / Request auth code.
Somebody else
- Find the page that lists your domains and open the one you want.
- DNS records is behind a link saying DNS, Zone, Zone File, Advanced DNS or Manage DNS.
- Nameservers is behind a link saying Nameservers, NS or Delegation, and the option you want is the one called custom, external or third-party.
- The auth code is behind a link saying Transfer, Transfer Out, EPP, or Authorisation Code — and it is always greyed out until you turn the transfer lock off.
If you genuinely cannot find one of them, search the registrar's help centre for the words "EPP code" rather than for "transfer": the help article you want is the one written for people leaving, and it is usually more direct than the one written for people staying.
Door three — moving the registration to Cloudflare
Worth saying plainly before you start: a domain transfer does not take your site down. The records are already at Cloudflare by this point — that is what door two was — and the transfer only changes which company bills you. The fear that it will break something is the single biggest reason people keep paying a renewal five times what the name is worth.
A transfer also is not free, exactly. You pay one year's registration, and that year is added to the time you already have rather than replacing it. So transferring a name with eight months left buys you twenty months. Nothing is lost.
Before Cloudflare will take it
Four conditions, all of them ICANN's or Cloudflare's rather than anything you can argue with. Check them first, because failing one at step six is much more annoying than reading them now.
| Condition | Why, and what to do |
|---|---|
| Registered 60+ days ago | An ICANN rule against fraud, not a Cloudflare one. A brand-new domain simply has to wait. Same if it was transferred in the last 60 days. |
| No recent contact change | Changing the registrant's name or email can start its own 60-day lock at some registrars. If you are about to do both, transfer first. |
| Already on Cloudflare DNS | Cloudflare Registrar only holds names whose DNS it also serves. This is door two, and the zone has to read Active before the transfer option appears. |
| DNSSEC turned off | If your registrar signed the zone, disable DNSSEC there and give it a day before changing nameservers. Skipping this is the classic way to take your own domain down mid-transfer. |
The transfer itself
-
Add the domain to Cloudflare and switch the nameservers
Cloudflare reads your existing records and copies most of them for you. Check the
copy against the old panel before you switch — it is good, not perfect, and
MXrecords are the ones it most often needs help with. - Wait for the zone to say Active Usually minutes, occasionally a day. Nothing below this line is available until it does, and refreshing does not help.
-
Unlock the domain at the old registrar
The setting is called the transfer lock or
clientTransferProhibited. Turning it off is not dangerous; it is a seatbelt, and you are the one driving. - Get the authorisation code Also called an EPP code, auth-info code or transfer code. Some registrars show it on screen, some email it to the address on the registration — which is another reason that address has to be one you can actually read.
- In Cloudflare: Domain Registration → Transfer Domains Pick the domain, paste the code, pay the one year, and confirm the contact details. Cloudflare's WHOIS privacy is on by default and included.
- Approve it, then leave it alone The old registrar emails you asking to confirm. Approving speeds it up; ignoring it still works, because the transfer completes on its own after five days unless somebody actively refuses it. Then turn auto-renew on in the new place, which is the step everybody forgets.
Make it safe
Before anything fun. Twenty minutes, once, and then never again.
Getting in
You already have everything you need. Yes, on Windows. No, you do not need to download anything.
A terminal is a window where you type commands instead of clicking things. That is the entire concept. It is intimidating for about four minutes, and then it is the fastest thing you have ever used, and then you start resenting software that will not let you type at it.
SSH is how you get a terminal on a computer that is somewhere else. It is encrypted, it is ancient, it has been attacked by everybody for thirty years and it is still standing, and it is already built into Windows 11 and macOS. You do not need PuTTY. You do not need WSL for this. You do not need to install one single thing. Anyone telling you otherwise learned this in 2009 and has not looked since. (WSL is worth having for an entirely different reason, later.)
Opening a terminal
Press Win and type terminal, then open Windows Terminal. It opens on PowerShell, which is exactly what you want. Check that SSH is present:
ssh -V
You should see something like OpenSSH_for_Windows_9.x. If the
command is not found, open Settings → System → Optional features,
click Add an optional feature, and add OpenSSH Client.
One reboot and it is there for good.
Press ⌘ Space, type terminal, hit return. SSH has been installed since forever. Confirm it:
ssh -V
Optional: a nicer terminal
Everything in this guide works in the terminal you already have, and I would rather you started there — one fewer download between you and the thing you actually came to do. But the built-in terminal is a fairly plain box, and once you are spending real hours in it you may want something with better manners. Warp is the one worth knowing about. It runs on Windows, macOS and Linux, it keeps each command and its output in a separate block you can scroll to and copy cleanly, and text editing in it behaves the way text editing behaves everywhere else in your life, which is a lower bar than the terminal has historically cleared.
Windows 11 ships with winget, a package manager, so this is one line
in the terminal you already opened. There are x64 and ARM64 installers on the
site if you would rather click:
winget install Warp.Warp
Then find it in the Start menu. It needs Windows 10 build 18362 or newer, which in practice means anything you are likely to be running.
With Homebrew, or from the download page if you do not have it:
brew install --cask warp
Your first connection
Take the IP address and root password from the provider's email. Replace
203.0.113.40 with your own address everywhere in this guide.
ssh root@203.0.113.40
The first time, you will be asked something like this:
The authenticity of host '203.0.113.40' can't be established.
ED25519 key fingerprint is SHA256:iRk3z...
Are you sure you want to continue connecting (yes/no/[fingerprint])?
Type yes and press enter. Your computer is saying I have never
spoken to this machine before and I have no way of knowing it is really the one you
meant. It writes down the server's fingerprint so that if it ever changes it
can shout at you about it. You see this once per server and then never again, and
on the day you do see it again you should stop and find out why.
Then paste the root password. Nothing appears as you type — no dots, no asterisks, not so much as a flicker. Your keyboard is fine, your paste worked, the terminal is simply declining to tell an onlooker how long your password is. Press enter.
Welcome to Ubuntu 26.04 LTS (GNU/Linux x86_64)
root@your-server:~#
That prompt is a computer in a data centre waiting for you to say something. Everything you type from here runs there, not here. Try a few harmless things just to feel it:
# who am I, and where am I?
whoami
pwd
# what is this machine made of?
free -h
df -h
nproc
# how long has it been switched on?
uptime
exit to come home.
The bits nobody explains
Every guide skips these, because whoever wrote it learned them so long ago that they have stopped being knowledge and turned into instinct. They are the actual reason beginners bounce off a terminal, so here they are, in order of how much time they will save you.
The prompt is telling you four things.
sam@web-01:~$
│ │ │ └── $ means an ordinary user. A # means you are root:
│ │ │ be awake.
│ │ └───── where you are. ~ is your home folder.
│ └──────────── which machine. Check this before anything destructive.
└──────────────── who you are on it.
- Pasting, which is where most people get stuck
- In Windows Terminal, Ctrl+V pastes and so does a plain right-click. What you must not do is reach for Ctrl+C to copy while something is running — in a terminal that means stop what you are doing, and it has meant that since long before it meant copy. To copy, select the text and use Ctrl+Shift+C, or just select it and right-click.
- Tab completes, ↑ remembers
- Type the first few letters of a file or command and press Tab; the shell finishes it, and beeps or shows you the options if it is ambiguous. The up arrow walks back through everything you have typed. A surprising amount of being fast at this is simply not typing — people who look quick at a terminal are mostly pressing two keys.
~is home,/is everything-
/is the top of the machine and every file lives somewhere beneath it.~is shorthand for your own folder, which is/home/samfor a user called sam.cdmoves you,pwdtells you where you ended up, andcdwith nothing after it always takes you home. - Getting out of something
-
Ctrl+C stops whatever is running
and gives you the prompt back. It is the fire escape and you should use it freely;
it is not destructive.
exit— or Ctrl+D — closes the SSH session and returns you to your own machine. If a program has taken over the whole screen and ignores both, it is usually an editor, and it usually wants q. - Nothing scrolls away permanently
- You can scroll up. The output of the last hour is still there. When something goes wrong, the answer is nearly always four lines above where you are looking, in the part you skipped because it had not gone wrong yet.
nano, and the editor you did not ask for
Every file this guide asks you to change is opened with nano, and every
guide tells you the same two keys and stops: Ctrl+O
saves, Ctrl+X quits. That is enough
to paste a config in. It is not enough for the day an error message says
line 212 of a file with four hundred lines in it. The two rows at the
bottom of the screen are the rest of the manual — ^ means
Ctrl and M- means Alt
— and these are the five worth knowing before you need them:
- Ctrl+W finds
- Type a word, press enter, and the cursor lands on it. Alt+W jumps to the next one. In a long nginx config this is how you reach the line you came for.
- Ctrl+/ goes to a line
- Type the number an error message just gave you. Alt+G is the same thing, and
nano +212 fileopens the file already there. - Ctrl+K cuts a line, Ctrl+U puts it back
- Press Ctrl+K several times in a row and the lines travel together, so cut, move, paste is how a block moves. It is also how a line is deleted: cut it and never paste.
- Alt+U undoes
- As many steps back as you like; Alt+E redoes. Nothing touches the disk until Ctrl+O, so a file you have made a mess of is also fixed by Ctrl+X and answering N.
- The questions it asks
- Ctrl+O shows the filename and waits: press enter to keep it. Ctrl+X on a changed file asks whether to save: Y then enter, or N to throw the changes away. Ctrl+C backs out of any question it asks.
Then there is the day a screen opens that is not nano. No rows at the bottom, and
typing either does nothing or puts letters where you did not aim them. That is
vi, and you are in it because some program opened its own editor
rather than yours. It will not be the server: on the Ubuntu this box runs, every
program falls back to nano, and vi is not installed at all. It will be
your laptop. A Mac's git commit, typed without -m, opens
vim, and Git for Windows offers vim as its editor during install, so the first time
you forget the -m you are in it. Getting out is three keys:
Esc, then :q!, then enter, which leaves without
saving, and at that moment saving was not the plan anyway. Then tell every program
which editor you meant, and it never happens again:
# Mac: ~/.zshrc. Linux, or Ubuntu inside WSL: ~/.bashrc. Then open a new terminal.
export EDITOR=nano
# PowerShell on Windows has no nano; point git at Notepad instead.
git config --global core.editor notepad
vi is not a mistake, for what it is worth. It is from 1976, when the screen was the expensive part, and every one of its habits is a habit for a machine that cannot afford to draw a menu. People who learned it then still fly in it. You are not obliged to, and nothing in this guide will put you in it on purpose.
Stop typing the address
You are going to type ssh sam@203.0.113.40 several hundred times, and
you are going to mistype it. SSH reads a config file on your machine where
you can give the whole thing a short name. Make it early; it is one of those small
things that quietly removes friction for years.
# Windows: C:\Users\you\.ssh\config
# macOS, Linux: ~/.ssh/config
Host box
HostName 203.0.113.40
User sam
From then on, the whole command is:
ssh box
Add a Host block per machine and the names are yours to choose. Once you
have made a key in the next section, one more line in that block points at it, and
you can stop thinking about any of this permanently.
When it will not connect
SSH has four failure messages that beginners read as one big "it is broken". They are not the same thing at all, and each one tells you precisely how far you got, which means each one has a different fix.
| What it says | How far you got | What it usually is |
|---|---|---|
Connection timed out |
Nothing answered at all. | Wrong IP address, the machine is not running yet, or a firewall is dropping you. New servers can take a couple of minutes to actually boot. Check the address against the provider's email, character by character. |
Connection refused |
Something answered, and said no. | You have the right machine but nothing is listening on that port. SSH is not running, or it is on a different port. This is the good failure — it means the address is right. |
Permission denied (publickey,password) |
All the way to SSH. It does not like you. | Wrong username, wrong password, or a key it has never seen. Note the username: root@ and sam@ are different people and only one of them exists yet. |
Host key verification failed |
All the way, and it stopped on purpose. | The machine's fingerprint is not the one your computer remembers. Almost always because the server was rebuilt. Almost — so find out why before you clear it, because the other explanation is that you are not talking to your server. |
-v when you are stuck.
ssh -v sam@203.0.113.40 narrates every step it takes, and the last line
before it gives up is the one that matters. If you cannot make sense of it, that
output is also the single best thing to paste at an agent — it is precise, it is
short, and it describes a real machine rather than a feeling.
Lock it down
Do this before you build anything. It is not optional, it is not advanced, and it takes twenty minutes.
Your server has a public address, and within minutes of it existing, automated scanners find it and start trying passwords. This is not paranoia and it is not personal — it is the background radiation of the internet, and it is aimed at everybody equally. On the machine serving this page there were 47 failed login attempts in a single day before it was locked down, and it had been online for a matter of hours.
Six steps, in order. Afterwards, guessing your way in stops being slow and starts being arithmetically impossible, which is a much better place to be.
One thing worth saying before you start, because it saves you years of anxiety later: most of what gets called a "vulnerability" on the internet is not the kind of thing that happens to you. A great many of them read like well, if the attacker were already inside your computer, and were also very small, he could easily bend your CPU pins. The correct response to that is not a patch. It is asking why there is a two-foot faerie rattling around inside the case in the first place. The six steps below are about the front door, which is the part that actually gets tried, every single day, by everybody.
1. Stop being root
root is the account that can do absolutely anything, with no
confirmation and no undo. Living as root means every typo is a potential
catastrophe and the logs cannot tell you who did what, because the answer is always
"root". Make yourself an ordinary account that can become root when it has
a reason to.
# Replace "sam" with whatever you want to be called.
adduser sam
# Put yourself in the sudo group, which grants the right to
# temporarily run a single command as root.
usermod -aG sudo sam
adduser asks for a password — make it long and put it in a password
manager — and then asks for your full name, room number and phone number. Press enter
through all of those. They are a relic of shared university machines and nobody has
read one this century.
From now on, when a command needs root powers you put sudo in front of it
and type your password. To get a full root shell for a while, use
sudo -i.
2. Make a key, and use it instead of a password
This is the one that matters. An SSH key is a matched pair of files. The public half goes on the server and is not a secret in any sense — you could put it on a billboard. The private half stays on your laptop and never moves anywhere, ever. Logging in proves you hold the private half without ever sending it.
- Your laptop names which key it would like to use. Not the key itself — which one.
- The server looks in
authorized_keys, finds the public half, and sends back a number it has never sent before. - Your laptop signs that number with the private half, which has not moved and will not move.
- The server checks the signature against the public half it already had, and opens the door.
- So nothing secret crossed the wire, and nothing on the server is worth stealing:
authorized_keyscan verify you and can never impersonate you.
A password can be guessed. A key cannot. The space to search through is so large that brute force stops being a strategy and becomes a genre of fiction.
On your laptop, in a new terminal window — not the one connected to the server:
# Make the key. Press enter to accept the default location.
# When it asks for a passphrase, type one — it encrypts the key file
# itself, so a stolen laptop is not a stolen server.
ssh-keygen -t ed25519 -C "sam@laptop"
# Copy the PUBLIC half up to the server. scp is a file copy over ssh.
ssh sam@203.0.113.40 "mkdir -p ~/.ssh && chmod 700 ~/.ssh"
scp $env:USERPROFILE\.ssh\id_ed25519.pub sam@203.0.113.40:~/newkey.pub
ssh sam@203.0.113.40 "cat ~/newkey.pub >> ~/.ssh/authorized_keys && chmod 600 ~/.ssh/authorized_keys && rm ~/newkey.pub"
Windows has no ssh-copy-id, and piping the file straight into
ssh from PowerShell can quietly corrupt it with Windows line endings
and a byte-order mark. Copying the file with scp and appending it
server-side, as above, avoids the whole problem.
# Make the key. Press enter to accept the default location, and
# choose a passphrase when asked.
ssh-keygen -t ed25519 -C "sam@laptop"
# macOS has a purpose-built tool for installing it.
ssh-copy-id sam@203.0.113.40
# Optional: let the macOS keychain remember the passphrase.
ssh-add --apple-use-keychain ~/.ssh/id_ed25519
Now test it, and do not skip this. Open a new terminal and log in:
ssh sam@203.0.113.40
If you are asked for your key passphrase, that is correct — that prompt comes from your own laptop. If you are asked for your server password, the key is not working, and you must fix that before continuing. Stop here until a passphrase-only login works.
This is also the moment to finish the Host box block you made back in
Getting in. Add one line to it —
IdentityFile ~/.ssh/id_ed25519 — and ssh box now means the
right machine, the right user and the right key, forever.
3. Turn passwords off entirely
Now that keys work, close the door behind you. On the server:
sudo nano /etc/ssh/sshd_config.d/10-hardening.conf
Paste this in, then press Ctrl+O, enter, Ctrl+X:
# Keys only. Password guessing is now impossible rather than slow.
PasswordAuthentication no
KbdInteractiveAuthentication no
PubkeyAuthentication yes
# root can still be reached with a key, never with a password.
PermitRootLogin prohibit-password
# Three tries and twenty seconds, not six and two minutes.
MaxAuthTries 3
LoginGraceTime 20
X11Forwarding no
# Keep long connections alive through home-router NAT.
ClientAliveInterval 30
ClientAliveCountMax 6
10-, and why it matters.
SSH reads that directory in alphabetical order and keeps the first value it
finds for each setting. Nearly every cloud image ships a
50-cloud-init.conf containing the single line
PasswordAuthentication yes. Call your file
99-hardening.conf and it is silently ignored: your config looks
immaculate, you feel great about it, and passwords still work. This exact trap
caught the setup of the very server you are reading this on, which is why there is
a verification command two paragraphs down. Never trust the file. Ask the daemon.
Check the config parses, apply it, and then verify what the server actually believes:
sudo sshd -t # silence means the file is valid
sudo systemctl reload ssh
# The real test. This prints the settings in force, not the ones you wrote.
sudo sshd -T | grep -iE 'passwordauth|permitrootlogin|pubkeyauth'
You want to see exactly this:
permitrootlogin prohibit-password
pubkeyauthentication yes
passwordauthentication no
4. Close every port except the three you use
A firewall decides which doors on the machine are open to the
internet. Ubuntu ships one called ufw that is genuinely simple.
sudo ufw allow OpenSSH # 22 — your terminal. Allow this FIRST.
sudo ufw allow 80/tcp # 80 — plain web, needed for certificates
sudo ufw allow 443/tcp # 443 — encrypted web, the real one
sudo ufw enable
sudo ufw status verbose
ufw with a default-deny policy and no SSH rule locks you out
of your own server instantly and completely. It is the single most common way
people lose a box on day one. The order above is correct — follow it exactly.
Notice what is not on that list. Nothing for a database, nothing for the
application itself, nothing for anything you are about to build. That is deliberate,
and it is the single most useful habit in this section: anything only your own
programs talk to listens on 127.0.0.1, which is not reachable from
outside the machine at all. The Databases section does this
properly. A door that does not exist cannot be picked.
5. Ban the scanners automatically
fail2ban reads the logs and blocks addresses that keep failing. With
keys-only login nobody was getting in regardless, so this is not really about
security — it is about not having your logs buried under thousands of identical
failures, so that the one interesting line is still visible when you go looking for
it.
sudo apt install -y fail2ban
sudo nano /etc/fail2ban/jail.local
[DEFAULT]
bantime = 1h
findtime = 10m
maxretry = 5
# Repeat offenders get progressively longer bans, up to a week.
bantime.increment = true
bantime.factor = 2
bantime.maxtime = 1w
# Ubuntu logs SSH to the systemd journal, not to a file.
backend = systemd
banaction = ufw
[sshd]
enabled = true
mode = aggressive
sudo systemctl enable --now fail2ban
sudo fail2ban-client status sshd
6. Actually turn security updates on
Ubuntu installs unattended-upgrades by default, and everybody assumes
that means updates are happening. On a great many cloud images they are not. The
service is running, it is enabled, it reports itself as healthy, and the periodic
trigger that tells it to actually do something is set to zero. It has been enabled
and doing nothing for months.
# Check what yours says. A "0" here means nothing has ever been installed.
cat /etc/apt/apt.conf.d/20auto-upgrades
sudo nano /etc/apt/apt.conf.d/20auto-upgrades
APT::Periodic::Update-Package-Lists "1";
APT::Periodic::Unattended-Upgrade "1";
APT::Periodic::AutocleanInterval "7";
Then catch up on everything that was missed while it was switched off:
sudo unattended-upgrade -v
# If a new kernel was installed, the machine needs a restart to use it.
# This file only exists when a reboot is genuinely pending.
cat /var/run/reboot-required 2>/dev/null && sudo reboot
Switching it on is step one of two. Step two is knowing what it does not cover, how to check it is really installing things, and when the software you are running stops being fixed at all — which is Staying up to date, in part six.
0. Two lines in a text file, and almost nobody ever opens it. Assume
nothing on a server is happening until you have seen the output that proves it.
apt, which you have now used without being told what it is
Steps five and six ran apt and edited its settings without stopping to
introduce it, so, briefly: it is how software gets onto this machine. It is what
Windows does with a download page, a setup wizard and a tray icon nagging about
updates, done once, for everything. Ubuntu keeps an archive — the one this box
reads lists about 78,000 packages — built by Ubuntu, signed by
Ubuntu, and checked against that signature on the way in. apt install
fetches from it. So does the nightly security run you just switched on. Every
install in this guide comes from there unless it says otherwise, and that "unless"
is what the notice below is about.
sudo apt update # refresh the list of what exists; installs nothing
sudo apt upgrade # install newer versions of what you already have
apt search --names-only sqlite # is there a package called something like this?
apt show sqlite3 # what is it, how big, what comes with it
sudo apt install sqlite3 # get it. -y answers the yes/no question for you
sudo apt remove sqlite3 # drop it. purge instead also removes its config
sudo apt autoremove # and whatever was only there to support it
apt list --installed # everything on here. This box: 1,114 packages
The pair everybody gets backwards is the first two. update updates
nothing; it downloads the list of what is available, and that is all. It is
upgrade that installs. The failure that teaches this is
E: Unable to locate package for something that certainly exists, and
it means the catalogue on your box is older than the package, which on a machine
made yesterday is very likely. Run update first, then look again. The
two go together so often that the one-liner in Linux on your
laptop joins them with &&. And apt-get, which
every guide older than 2014 uses instead, is the same tool with a face for scripts;
apt is the face for people. Either works. Read apt-get as
a sign of a guide's age, not of its wrongness.
add-apt-repository ppa:somebody/something
to get a newer version of a thing. A PPA is one person's build, hosted on
Launchpad, and installing from it means their packages run as root on your box,
every time they update, for as long as the line stays in your sources. That is not
a vulnerability. It is the arrangement, and it is the same arrangement you already
have with Ubuntu; the difference is that Ubuntu has a security team, a release
process and a name to lose, and ppa:somebody has somebody. Three
honest reasons to step outside the archive: the software is not in it, the version
in it is too old for you, or the vendor publishes their own repository, which is a
PPA with a company standing behind it that you can hold responsible. This box has
exactly one of those, GitHub's own, for gh: a file
in /etc/apt/sources.list.d/ and GitHub's signing key in
/etc/apt/keyrings/. Know what each line in that folder is for,
because the release upgrade in two years' time switches every one of them off and
asks you what to do, and "I do not remember" is the wrong answer to have ready.
Ubuntu also ships snap, a second way to install things, with its own
opinions. Nothing in this guide needs it.
The things that actually go wrong
Break-ins on small servers are tediously repetitive. Nobody is writing a novel exploit for your side project. Almost every lost box is one of these five, and you have just handled the first two. What to do on the day one of them happens anyway is in the backups section, because that is where the answer lives.
How people lose a box
- A guessable password on an account exposed to the internet.
- Root login left enabled with a password.
- A database bound to
0.0.0.0with a weak password — scanners find these in hours. - Secrets committed to a public GitHub repository. Bots scrape new commits within minutes.
- A server never updated because nobody checked that automatic updates were running.
What you just did about it
- Passwords no longer work over the network at all.
- Root can only be reached with a key you hold.
- Everything internal will listen on
127.0.0.1, invisible from outside. .gitignoreand environment variables — covered in GitHub.- Security updates install themselves, and you verified it rather than assumed it.
Your server as a drive
Optional, and genuinely lovely. Browse the server in Explorer or Finder.
The argument of this guide is that you should stop living in a file explorer. That is not the same as never wanting to look at one. Sometimes you want to see a folder, drag a photo in, or open a file in an editor with fonts in it, and none of that requires giving up the terminal as the place where the work happens.
So mount the server as an ordinary drive. It rides on the SSH connection you already have, which means nothing new to secure, no new port, and no new password to lose.
SSHore is a small Windows 11 app that turns a saved SSH connection into a drive letter. Install it, add a profile with your server address, username and private key, and the whole filesystem appears in Explorer. It sits in the tray, remembers your profiles, mounts in one click, and stores credentials with Windows DPAPI rather than in a text file.
- Install WinFsp and SSHFS-Win These are the two components Windows needs to understand a remote drive. SSHore's installer points you at both.
-
Add a profile
Host
203.0.113.40, usersam, and your private key atC:\Users\you\.ssh\id_ed25519. Point the remote path at/srv— where all your projects will live — or at/for the whole machine. - Mount it It becomes a drive letter. Open it in Explorer, drag files in and out, open them in any editor you like.
In Finder, press ⌘ K and
connect to sftp://203.0.113.40. It works with no installation at all,
though it is read-mostly and a little slow.
For a proper read-write mount, install
macFUSE and
sshfs, then:
mkdir -p ~/server
sshfs sam@203.0.113.40:/srv ~/server -o volname=server,follow_symlinks
A second Linux, on your own laptop
Optional, free, and the best twenty minutes in this guide if you are on Windows. Not a replacement for the server — a practice room next door to it.
Everything so far has been about one computer, the one you rent. This section is about putting the same operating system on the machine in front of you as well, so that the commands in this guide have somewhere to be wrong. It is genuinely optional. Skip it and nothing later breaks. But a beginner's biggest practical problem is that the only Linux they have is the one their website is running on, which makes every experiment feel like defusing something.
What WSL actually is
WSL is the Windows Subsystem for Linux. Despite the name it is not an emulator or a compatibility shim: modern WSL runs a real Linux kernel in a very light virtual machine that Windows manages for you, with Ubuntu inside it — the same Ubuntu that is on your server. There is no dual boot, no partitioning, nothing to configure, and no decision at startup about which operating system you are in today. It is a program you open, and inside it is Linux.
It is built into Windows 11. You are not downloading a third-party tool, and this
does not contradict what Getting in says — you still do
not need WSL to reach your server, because Windows has had a perfectly good
ssh for years. This is something else: not a way in, a second machine.
Why bother, when you already have a server
apt, same paths, same permissions. One set of habits instead of two.
The last one is the real argument. A server is for running things. A
laptop is for trying things. When both of them are Ubuntu, the distance
between the two is a git push rather than a translation exercise, and
the "works on my machine" conversation you have heard people complain about for
twenty years mostly stops happening.
Installing it
Open PowerShell as Administrator — right-click the Start button, then Terminal (Admin) — and run one command.
wsl --install
That turns on the two Windows features it needs, fetches the kernel, installs Ubuntu, and tells you to reboot. Reboot. On the way back up, an Ubuntu window opens by itself and asks you to invent a username and a password.
sudo on your own laptop, and it
is fine for it to be something you can type quickly. The server's password is a
different class of secret and is dealt with in Lock it down.
Two variations worth knowing, both run from the same PowerShell window:
# See which distributions are on offer before committing
wsl --list --online
# Ask for a specific Ubuntu rather than whatever is current
wsl --install -d Ubuntu-24.04
# If WSL was already installed years ago and is stale
wsl --update
Take the Ubuntu LTS that matches your server if you can. It is not a requirement and nothing breaks if the versions differ, but matching them removes an entire category of confusing afternoon: a package that exists in one and not the other, a default that changed between releases, a config file that moved.
The first five minutes inside it
You are now at an Ubuntu shell that behaves exactly like the one on your server, so the first thing you do is the same thing you did there.
sudo apt update && sudo apt upgrade -y
From here, every macOS tab in this guide is your tab. That is worth
saying explicitly, because it looks wrong: the Windows tabs on this page are written
for PowerShell, and PowerShell is not what you are in any more. Inside WSL you are on
a Unix shell, so the macOS commands — the ones with ssh-keygen and
~/.ssh/config and forward slashes — are the ones that apply.
apt replaces brew and everything else lines up.
Where your files actually are
This is the one part of WSL that confuses everybody, and getting it right on day one saves a genuinely miserable week later. There are two filesystems and they are not equally fast.
| Path | What it is | Speed |
|---|---|---|
/home/you |
The Linux filesystem. Where your projects belong. | fast |
/mnt/c/… |
Your Windows C: drive, mounted so you can reach it. | slow |
\\wsl$\Ubuntu |
The Linux filesystem, seen from Windows Explorer. | fine |
/home/you, never in /mnt/c.
Every read and write under /mnt/c crosses between two operating
systems, and a build reads thousands of small files. Put a project on
/mnt/c and a compile that takes four seconds takes ninety, an
npm install takes a coffee break, and you conclude that WSL is slow.
It is not; you asked it to do the one thing it is bad at. This is the single most
common WSL mistake and it is invisible until you have made it.
To get at your Linux files from Windows — to drag one into an email, or open it
in something graphical — type \\wsl$\Ubuntu into the Explorer
address bar, or run explorer.exe . from inside WSL to open the folder you
are standing in. Both are the correct way round: reach into Linux from Windows, rather
than keeping the files on Windows and reaching out.
The terminal to use
Windows Terminal is already on Windows 11 and it grows an Ubuntu tab by itself the moment WSL is installed — the down-arrow next to the plus sign in the tab bar. That is the whole setup. You can pin it, split it, and have PowerShell in one pane and Ubuntu in another, which is a genuinely nice way to work: Windows things on the left, server-shaped things on the right.
If you preferred Warp from the Getting in section, it drives WSL too. Nothing here depends on which terminal you use, and switching later costs nothing.
Two things that will catch you out
You now have two SSH identities, in two places. Keys you made in
PowerShell live in C:\Users\you\.ssh. Keys you make inside WSL live in
/home/you/.ssh, which is a completely separate directory. WSL will not
find the Windows ones, so an ssh mybox that worked yesterday in
PowerShell says Permission denied today in Ubuntu, and it looks like
the server broke. It did not. Either make a second key inside WSL and add it to the
server the same way — which is cleaner, because a key should belong to one
place — or copy the existing one across and fix its permissions:
# Only if you would rather have one key than two
mkdir -p ~/.ssh && cp /mnt/c/Users/you/.ssh/id_ed25519* ~/.ssh/
chmod 700 ~/.ssh
chmod 600 ~/.ssh/id_ed25519
# ssh refuses a private key that other users could read. This is why.
ls -l ~/.ssh
And systemd is on, but it is not your server's systemd. Recent WSL
runs systemd, so the unit files in
nginx, systemd, TLS can be written and started locally, which is
a much better place to get a unit wrong than on the machine serving your site.
What it cannot do is convince you the deployment works — different kernel,
different network, no TLS, no real ports open to the world. Practise the shape here;
prove it there.
What to practise in it
If you want a use for this beyond having it, here is the honest list of what is better learned locally than on the box, roughly in the order you will meet them.
| Practise here | Because |
|---|---|
Users, sudo and permissions | Locking yourself out of WSL costs two minutes. Locking yourself out of the server costs a support ticket. |
Writing a justfile | Pure text and fast feedback. Nothing about it needs a server at all. |
| SQLite | A database is one file. Make ten, delete ten. |
| A systemd unit | Getting a unit wrong locally teaches you the same lesson without any downtime. |
| Restoring a backup | The restore drill wants a machine you are happy to wreck. This is that machine. |
brew instead of
apt. If you want a closer match to the server than macOS gives you, the
answer is the same idea in a different wrapper: a small virtual machine running the
same Ubuntu. Worth doing eventually, not worth doing today.
Bring in the agent
This is the part that changes what you are capable of building.
Claude Code
Install it, point it at the root of the machine, and let it work.
Claude Code is Claude in a terminal, with hands. It reads and writes files, runs commands, installs packages, starts services, reads its own error output and has another go. It is not an autocomplete in a sidebar. It is closer to a colleague you hand a task to and then check on.
The honest version, because you will meet this on day two and it is better to hear it now: imagine hiring a mechanic for your garage who is genuinely pretty good, but who might also, without warning, start doing aerobics instead of fixing the car. He will confidently repair a gas leak by taping the windows shut, and if you question it he will defend the decision, fluently, with reasons. That is what you are working with. It is still an enormous amount of mechanic for the money — you just do not hand him the keys and go on holiday.
On the server, as your normal user:
curl -fsSL https://claude.ai/install.sh | bash
Then close the terminal and open a new one so your shell picks up the new location, and check it:
claude --version
The first time you run claude, it prints a URL. Open it in the browser on
your laptop, sign in to the Claude account your subscription is on, paste the code
back. That is the whole setup.
Start at the root
Here is the bit people find surprising, and it is the single most useful habit in this guide:
cd /
claude
/ is the top of the filesystem — every single thing on the machine is
somewhere underneath it. Starting there means the agent can see the web server
config, the systemd units, the logs, the database and every project at once. When you
say "put this online at demo.mysite.com", nobody has to explain where nginx
lives or which service to restart. It can go and look, the way you would.
On a normal development machine this would be an unhinged thing to do. Here it is the whole point. The box is disposable, the work is in git, and an agent with full context is the difference between "here is a code snippet, now go and configure your web server" and "it is live, here is the link".
Things worth knowing on day one
- Just describe the outcome
- "Make me a page at status.mysite.com that shows whether each of my services is running, dark theme, updates by itself." Not a list of steps — you are hiring for the result, not renting hands. It will ask when something genuinely matters, and it is usually right about what does not.
- Esc stops it
- Interrupts whatever it is doing without killing the session, so you can redirect mid-task. Use it the moment you see it heading somewhere you did not intend — earlier is much cheaper than later.
/clearwipes the conversation- Starts fresh, keeping the same folder and the same permissions. This matters more than it sounds and gets its own section below.
/modelswitches which Claude you are talking to- Different models are better at different halves of the job. Also below.
#at the start of a message writes it down-
It saves the fact into the project's
CLAUDE.md, so the next session already knows it. "# the database password lives in /srv/app/.env, never in git." - It can use
sudo -
Installing packages, editing nginx, restarting services — all of that wants root.
Let it. Just actually read the line when the words
rm -rf,DROPor--forcego past. Those three are the ones with no undo.
A good first session
Something that exercises the whole pipeline and gives you a real URL at the end:
cd /
claude
Set me up a new project at /srv/hello with a small Go web server
that serves one nice-looking page. Give it a justfile with recipes
for dev, check and deploy, a systemd unit so it survives a reboot,
and an nginx config. Use port 8090.
Explain what each piece does as you go — I have not done this before.
Do not put it on the internet yet; I want to see it working locally
first with curl.
Watch what it does. Ask it why it did any of it — "why a systemd unit and not just running it?", "what is that nginx block for?" — and keep asking until the answers stop surprising you. That conversation is worth more than three tutorials, because it is about the actual machine in front of you rather than somebody's generic example, and because you can interrupt it.
curl the endpoint, read
the log, restart the service and watch it come back. It is very hard to hallucinate
a green check you ran yourself.
Driving the agent
The habits that separate a good session from a frustrating one.
Everything in this section is about one thing: context. An agent is only as good as what it can currently see and how much rubbish is piled on top of it. A fresh session with a clear brief beats a long session that has been wandering around for four hours, every single time, and it is not close.
Plan with one model, build with another
The two halves of building software want opposite things. Planning wants breadth,
patience, and a tolerance for sitting in an ambiguous problem without grabbing at the
first answer. Building wants speed, precision and something close to stubbornness.
Asking one session to do both is how you end up with a confident implementation of
the wrong idea. Use /model to switch — and revisit the pairing
whenever the roster changes, which is the next section.
| Phase | Model | What you ask for |
|---|---|---|
| Plan | Fable | Architecture, the directory layout, the order of work, what could go wrong. Ends with a written plan in docs/ and no code at all. |
| Build | Opus | Execute the approved plan. Fast, thorough, good at holding a large codebase in its head and finishing things. |
| Review | Either | A fresh session, no memory of writing the code, asked to find what is wrong with it. |
The rhythm looks like this:
# 1. Think it through — no code yet.
/model fable
Read the codebase. I want to add user accounts with passkeys.
Write docs/features/auth.md: the approach, the database changes,
the endpoints, and what could go wrong. Do not write any code.
# 2. Read the plan yourself. Argue with it. Then:
/clear
/model opus
Read docs/features/auth.md and implement it exactly.
Run just check before you tell me a piece is done.
The /clear in the middle is not decoration. The build session should
start from the written plan, not from the rambling forty-minute conversation that
produced it — including the three approaches you talked yourselves out of, which are
still sitting there looking like options.
Clear early, and leave a note
A long conversation gets worse, not better. Everything you explored, abandoned and corrected along the way is still sitting in view, competing for attention with the thing that actually matters. Somewhere around 150,000 to 200,000 tokens — well before anything forces your hand — stop and reset.
Never just clear. That is throwing away the only copy. Hand over first, the way you would to a colleague going on leave:
We are about to run out of useful context. Write docs/HANDOFF.md
covering: what we set out to do, what is now done and working, what
is half-finished and exactly where, the decisions we made and why we
rejected the alternatives, and the precise next three steps.
Write it for someone competent who has never seen this project.
Then /clear, and open the next session with:
Read docs/HANDOFF.md and CLAUDE.md, then continue from step one.
This is not a workaround for a limitation. It is a better way to work even when you have plenty of room left, because writing the handoff forces every decision to be said out loud in plain words, and half the time that is when you notice the decision was wrong. You also get a project that documents itself as a side effect, which is the only kind of documentation that ever actually gets written.
I have written two hundred and thirty of these across three projects. I mention the
number because "leave a handoff" sounds like the sort of advice people give and do not
take, and I want you to know it is a practice rather than a suggestion. One warning
from doing it badly first: put the date and time in the filename —
docs/handoffs/2026-09-06-1430-auth-done.md — not in a folder name. I
started with a directory per day and it was useless. With the timestamp in the
filename, a plain directory listing sorts itself into a readable log of the whole
project, and finding "the session where we changed the schema" takes four seconds.
Audit what it remembers
CLAUDE.md is a file in each project that the agent reads at the start of
every session. It fills up over time, and some of what collects in it goes stale — a
convention you abandoned, a port you moved, a rule that was true for exactly one
afternoon in March. Stale instructions are worse than no instructions, because they
get followed confidently by something that has no way of knowing they expired.
Every couple of weeks:
Read CLAUDE.md and every .md file under docs/. For each claim,
check whether it is still true of the actual code. Give me a list:
still true, now wrong, or no longer relevant. Do not change
anything yet.
Good things to keep in CLAUDE.md:
- Commands: "always
just checkbefore deploying". - Constraints: "this box has 4 GB of RAM; keep the service under 200 MB".
- Decisions with a reason: "stdlib only — no web framework, deliberately".
- Traps: "the formatter is duplicated in Go and JS on purpose; change both".
- Ports already taken, so the next project does not collide.
CLAUDE.md files reached 24,000 words.
For comparison, my median is about 650 and the shortest useful one I have is 29
words. Twenty-four thousand is not a briefing, it is a novella, and every session on
that project paid for the whole thing in tokens before it got to a single useful
instruction. The really funny part is what it was full of: notes about keeping
context clean. Context rot, growing inside the file whose job was to prevent it.
A CLAUDE.md points at docs/. It does not absorb
it. If yours is over about 1,500 words, the excess is documentation wearing a
briefing's coat, and it belongs in a file the agent can choose to open.
Let it look at other people's work first
Before you start anything unfamiliar, send it off to read how the problem is normally solved. This costs one prompt and routinely saves a day of quietly reinventing something in a shape nobody else uses — which you will only discover much later, when you try to explain your project to someone and watch their face.
Before you write anything: find two or three well-regarded Go
projects on GitHub that do something structurally similar to this.
Read how they lay out directories, name things, handle config and
structure tests. Summarise what you found, what you are copying,
and what you are deliberately doing differently.
Seven habits, in one place
- One task per session Not "build the app". "Add the login page." Finish, clear, next.
- Make it explain before it acts On anything with consequences, ask for the plan first. Arguing with a paragraph is far cheaper than arguing with a diff.
- Commit constantly Every working state gets a commit. Then a mistake costs minutes, and "undo that" is a thing you can actually say and mean.
- Verify from outside "It should work" is not a result.
curlthe URL, restart the service, reboot the box and watch it come back on its own. - Interrupt early Esc the second it heads somewhere wrong. Sunk cost applies to agents too, and it applies to you watching one.
- Write things down as you learn them Every "oh, that is why" becomes a line in
CLAUDE.mdor a note indocs/, immediately, while you still care. - Ask whether a script would do Before handing something to the agent, ask yourself "could I just write a script for this?" If the answer is even maybe, write the script. Determinism you can read beats cleverness you have to check, and the less you need to trust, the better your results get.
Keeping your models current
The one part of this whole setup that changes underneath you every few weeks. Prices move in both directions, names move, and the model you chose in spring is rarely the right answer by autumn — which is a chore if you ignore it and a small, automatable advantage if you do not.
Two different habits live here, and people usually have one and not the other. The first is for any project of yours that calls a model — through Claude's API directly, or through a router that fronts hundreds of them. The second is for the tool you are holding: keeping the agent itself current, and knowing what you are spending.
If your app calls models, keep a roster
The catalogue is genuinely volatile. On a router like OpenRouter there are several hundred text models on any given day, and the spread is not subtle: top-tier models cost a hundred times what small fast ones do, and the same model hosted by different companies varies by a factor of ten or more, because they run it on different hardware at different precision. There is no "the price". There is a price per model per provider per day.
And it moves the pleasant way about as often as the unpleasant one. Newer versions frequently land cheaper than the thing they replace — the current top-end Claude model costs a third of what the previous generation's did — and one model on that list had an announced price rise cancelled and the introductory price made permanent. If you pinned a model a year ago and never looked again, the most likely outcome is not that you are being overcharged. It is that you are paying more than you need to for worse output than you could be getting.
# The whole catalogue, with prices, as JSON. No key needed to read it.
# Prices are dollars per token, so multiply by a million to get the
# number everybody actually quotes.
curl -s https://openrouter.ai/api/v1/models \
| jq -r '.data[] | select(.id|startswith("anthropic/")) |
"\(.id) $\((.pricing.prompt|tonumber)*1000000|round)/M in $\((.pricing.completion|tonumber)*1000000|round)/M out"'
That one command is the whole basis of the habit, because it means the roster is a file you can diff rather than a page you have to remember to visit. Save it weekly, compare against last week's, and have it shout when something you actually use has changed. Four rules make the difference between a roster that helps and one that quietly breaks your product:
- Pin exact versions in code; choose new ones deliberately. Aliases that always point at "the latest" are excellent in a terminal and a liability in production — the model changes without a deploy, and so does your output. Routers say this themselves: use the concrete name if you want reproducibility.
- Set a price ceiling in the request. Most routers let you refuse anything above a figure you name. Then a price rise fails loudly on a Tuesday instead of appearing silently on a bill six weeks later. Same reasoning as the hard spending cap on anything of yours that can spend money: the limit belongs in the code, not in your memory.
- Log which model actually answered. With fallbacks, cheapest-first routing or an alias, the model that served the request is often not the one you asked for. The answer comes back saying who wrote it. Keep that field; it is the only way to explain a change in quality afterwards.
- Keep twenty prompts with known-good answers. Run them before you switch anything. This is the entire difference between "I upgraded the model" and "I upgraded the model and my classifier quietly got worse". There is a well-known study of one hosted model whose accuracy on a simple task fell from 84% to 51% over three months with no version change at all — behaviour drifts even when nothing is announced, and the only defence is a handful of cases you check.
Keep the agent itself current
Claude Code updates roughly constantly, and the difference between versions is not cosmetic — tool behaviour, context handling and model availability all change. If you installed it the normal way it updates itself in the background; if you installed it through a package manager, including on Windows, it does not, and it will happily sit six months behind while you wonder why a documented feature is missing.
claude doctor # version, install method, update channel, last update result
claude update # take one now rather than waiting for the background check
claude doctor is the one worth knowing, because it answers the question
people actually have, which is "is this thing updating itself or not". On this box it
reports the version, that auto-updates are enabled, which channel they come from, and
when the last one succeeded. There are two channels: the default takes releases as
they ship, and the other lags about a week and skips anything that turned out badly.
If the agent is doing work you depend on, the slower one is the better default, and
you can still pull an update by hand the moment you want one.
Inside a session, three commands are worth having in your fingers:
/usage- What you have spent and how much of your plan's window is left. Subscriptions have both a rolling few-hour allowance and a weekly one, and the difference between "I have hit a limit" and "I have hit which limit" decides whether you wait an hour or change how you work.
/context- What is actually in the conversation, drawn as a grid. This is how you find out that half your window is one enormous file somebody pasted, which is both the cost and the reason the answers got vaguer.
/model- Switch models, and set how hard it thinks. Worth revisiting whenever a new one lands, along with planning with one model and building with another — the right pairing changes with the roster.
The token habits are the same ones in Driving the agent, and they are worth restating here because they are also the cost habits: a long conversation is an expensive conversation and a worse one. Clearing early with a handoff note, one task per session, and keeping files small are not frugality measures that cost you quality. They buy quality and happen to be cheaper.
Write me a small script that saves the model catalogue from my provider's API
each week, compares it against last week's copy, and tells me: anything whose
price changed, anything I use that has a retirement date, and any new model
cheaper than the one I am using for each job. Keep the JSON files in the repo
so the history is the record. Run it from a systemd timer and email me only
when something changed.
That script is the whole discipline. It takes an evening to write with the agent sitting right there, it costs nothing to run, and it turns "keeping up with models" from a thing you feel vaguely guilty about into an email that arrives when something actually happened.
Laying out a project
The same skeleton every time, in every language. It is the reason nothing gets lost.
Every project on this box lives in /srv/<name>/ and looks broadly
the same, whether it is Go, Rust, Python or a pile of HTML someone emailed you. The
sameness is the feature. You can come back after six months, or hand it to a
friend, or point a brand new agent at it, and the shape alone tells you where things
are before you have read a word.
/srv/myproject/
├── README.md what this is, how to run it, in five minutes
├── CLAUDE.md notes for the agent: conventions, traps, commands
├── TODO.md the next few things, in order
├── ROADMAP.md the next few months, deliberately vague
├── justfile every command this project has
├── .gitignore secrets and build output, never committed
├── docs/
│ ├── PLAN.md the architecture, written before the code
│ ├── HANDOFF.md where the last session stopped
│ ├── DECISIONS.md what was chosen, and what was rejected, and why
│ └── features/
│ ├── auth.md one file per feature, written before building it
│ └── search.md
├── deploy/ nginx config, systemd unit — the real copies
└── ... the actual code, laid out however the language wants
Why docs/ earns its place
Written documents are the only thing that survives a /clear. The
conversation is gone the moment you end it. The file is still there tomorrow, and
next month, and for whoever picks the project up after you — including you, who will
not remember any of this, and who will be quietly furious with the version of you
that did not write it down.
- docs/PLAN.md
- Written before any code, by the planning model. What is being built, how it is structured, in what order. If you cannot write this, you do not yet know what you are building.
- docs/features/<thing>.md
- One per feature, written before that feature exists. The approach, the data it needs, the endpoints, the edge cases. This is what you hand a build session so it starts with a brief instead of a vibe. When a feature earns pictures, let the file become a folder —
docs/features/auth/with the write-up inside it and ascreenshots/next to it. Save those as.webp: same picture, a fraction of the bytes, and every browser has read them for years. - docs/DECISIONS.md
- The single highest-value file in the repo. Every real choice with a date, a reason, and the options you rejected. It is the file that stops you and your agent relitigating the same argument every three weeks.
- docs/HANDOFF.md
- Written at the end of every long session. The bridge between one context window and the next. Once you have more than a couple, give them their own folder and put the date and time in the filename so the listing reads as a log.
- TODO.md and ROADMAP.md
- TODO is the next handful of concrete things. ROADMAP is the direction — deliberately vague, because a detailed six-month plan is fiction. Ask the agent to keep TODO current; it is good at it and it makes "what were we doing?" a one-second question.
The one file to copy into everything
If you take a single file from this section, take docs/quick-start.md.
It is not in the tree above, because I only worked out that it was the important one
fairly recently — and it is already in nearly as many of my projects as the justfile
is. It is not
a getting-started guide for users. It is an orientation file for whoever opens the
repository next, human or machine, and it answers five questions in about a page:
where things are, what the stack is, which ports it uses, the sixty-second tour, and
the hard rules that must not drift.
That last heading is the one that earns its keep. Hard rules — do not drift.
Three or four lines saying this project uses SQLite and not an ORM, the nginx config
in deploy/ is the real one, never edit the copy under /etc.
An agent reads that and stops rediscovering your preferences by trial and error, and
so does a person.
docs/ — and I genuinely cannot tell you what several of
them do without reading them from the top. They are not old. They are eighteen
months old and they are already archaeology. Nothing in this section is
twenty-year-old wisdom handed down from a mountain; it is a fairly recent fix for a
problem I gave myself, which is precisely why I can tell you which parts paid off.
The boilerplate nobody remembers
Websites in particular carry a pile of small files that are individually trivial and collectively the whole difference between something that looks finished and something that looks like homework. Nobody remembers all of them. I do not remember all of them, and I have been forgetting them since 2004. That is fine. This is exactly the kind of tedious completeness a machine is better at than you are.
Audit this website against the standard set of files a public site
is expected to have: favicon in the sizes browsers actually request,
web manifest, robots.txt, sitemap.xml, Open Graph and Twitter card
tags, a real 404 page, security.txt, a health endpoint, correct
canonical URLs, and sane cache headers on static assets.
Tell me which are missing, then add the ones that make sense here.
There is a fuller version of this as a tickable launch checklist further down.
just
One file that makes every project work the same way, whatever it is written in.
Every project has commands. How do I run it? How do I test it? How do I put it live?
The answers are always different — go run ./cmd/server,
npm run dev, cargo watch -x run,
python -m uvicorn app:app --reload — and you will not remember a single
one of them in three weeks. I promise you that you will not. I have gone back to my
own projects and had to read the CI config to find out how to start them.
just is a small program that reads a file called justfile
and gives every command a name. It is not tied to a language, a framework or an
ecosystem, and it has no opinion about your project. It runs shell, and that is the
entire pitch. That neutrality is why it is worth adopting on day one: it will still
be the right tool when your next project is in a language you have not learned yet.
sudo apt install -y just
just --version
Create the justfile before you write any code. Not after. It is the first file in a new project, because it is where you and your agent agree on what this project's verbs are, and that agreement is worth more than any of the code you are about to write.
The shape
# Variables at the top. Change them here, not in five places below.
service := "myapp"
port := "8090"
# Running plain `just` lists everything, which is why this is first.
default:
@just --list --unsorted
# Run it locally while you work on it
dev:
go run ./cmd/server -addr 127.0.0.1:{{port}}
# Everything CI would run. Do this before deploying.
check: fmt vet test
fmt:
gofmt -l -w .
vet:
go vet ./...
test:
go test ./... -count=1
# Build, install, restart, and confirm it actually came back
deploy:
go build -o bin/server ./cmd/server
sudo systemctl restart {{service}}
@sleep 1
@curl -fsS http://127.0.0.1:{{port}}/healthz && echo " — up"
# Follow the log. Ctrl-C stops watching, not the service.
logs:
journalctl -u {{service}} -f -o cat
deploy here builds over the running binary and restarts it, which is
the right amount of machinery on day one. The night it matters you will want
the version you can take back, and it is two changes away.
Now every project you ever build responds to the same four words:
just # what can this project do?
just dev # run it
just check # is it broken?
just deploy # put it live
just prints; the check is what has to pass before a deploy.
A Rust project's just dev runs cargo run. A Python one runs
uvicorn. A static site runs a file server. You do not care, and neither
does your muscle memory. Neither, and this is the part that matters, does the agent —
put the recipes in CLAUDE.md and it uses them instead of inventing a
slightly different command every session and leaving you three ways to start the same
program.
The deploy recipe above is three lines and it is the most important shape
in this whole file: restart, wait a beat, then ask the service whether it is
actually alive. That curl at the end is not garnish. A deploy
that does not verify is a deploy that reports success when the binary panicked on
startup, and you will find out about it from a person instead of from a script.
Mine has since grown up — it refuses to run on the main branch, runs
just check, opens the pull request, merges it and then deploys — but it
started as those three lines, and those three lines are already honest.
.gitignore excludes.
just loads a .env automatically if you put
set dotenv-load at the top, which gives you one clean seam: commands
are shared with everybody, values are shared with nobody.
# At the top of the justfile — loads .env, which git never sees
set dotenv-load
# The recipe references the name; the value lives only in .env
migrate:
psql "$DATABASE_URL" -f db/schema.sql
The vocabulary, and why you should never change it
Pick these names once and then use exactly these names in every project you ever
build, no matter what language it is in or what it does. Not similar names. The same
names. The whole return on just is that you can sit down in front of a
project you have not opened in a year and already know that just dev
starts it and just logs tells you why it stopped.
| Recipe | What it should do |
|---|---|
| default | Print the list. Always @just --list, never a build. Typing plain just should never surprise you. |
| dev | Run locally with fast reload. Never touches anything live. |
| build | Produce the artefact, and nothing else. |
| test | Just the tests, so check can call it. |
| lint | The linter, same reason. |
| check | Format, vet, test. The one thing the agent runs before it is allowed to say "done". |
| deploy | Build, install, restart, and verify. A deploy that does not check is a hope with a progress bar. |
| ship | The grown-up deploy, once you have one: check, commit, push, then deploy. One word for the whole ritual. |
| logs | Follow this service's output. The first thing you run when something is wrong. |
| status | Is it up. One line, no ceremony. |
| backup | Put the data somewhere safe. A first-class recipe, not an afterthought. See Backups. |
| clean | Delete build output. Never anything else, ever. |
| smoke | Hit the real public URL from outside and check the status codes and headers. |
The part I did not expect is that the names outlive whatever is reading them. When I
started keeping a .claude/skills/ folder so an agent could be told how to
deploy a project without me explaining it every time, I did not design what went in
it. I wrote the same words down again — deploy, check,
logs, status, handoff, migrate.
A recipe is that list for a person at a shell. A skill file is that list for a model
in a session. Ten of my projects now carry both, saying the same thing twice in two
formats, and that turned out to be the point rather than the waste: I never had to
teach the newer tool anything, because the words were the project's and not the
tool's. That is the real reason not to rename logs — not tidiness, but
that the name will still be there after the thing you picked it for is gone.
One more thing, and it is the habit that has paid me back the most: the
justfile is where you write down the thing that bit you. Not the command —
the reason. Why a permission has to be granted again after installing. Why there is a
sleep before the health check. Why this one recipe must not run on the
main branch. Those comments are long and they look like clutter, and every one of them
is a note from a past version of you who lost an evening finding it out.
run here, serve there, start in the one from
last spring, a deploy that in one project meant "build" and in another
meant "build and push to production". So every return visit began with reading a
file to find out what the words meant this time. Twenty-nine projects now use the
list above, unchanged, and the compounding is absurd for something that costs
nothing: I stopped having to remember anything at all.
GitHub
Your undo button, your backup, and the thing that builds your installers.
Git records the history of a project: every change, when, and why. GitHub is where that history lives on the internet. Together they are the entire reason you can let an agent run loose without your stomach hurting, because every working state is recoverable and "undo the last hour" stops being a wish and becomes a command. Nothing else in this guide buys you as much nerve for as little effort.
Later, GitHub is also what builds your Windows, macOS and Linux installers for free. That is the GitHub Actions section.
If you have never used it
- Make an account github.com/signup. Free. Turn on two-factor authentication when it offers — it is required for pushing code anyway, so do it now rather than mid-task later.
-
Install the command-line tool on the server
ghhandles authentication and repository creation without you ever touching a token by hand.sudo apt install -y gh git -
Log in
Choose GitHub.com, then HTTPS, then Login with a web browser. It shows an eight-character code; open the URL on your laptop, paste the code, approve. It stores the credential securely on the server and configures git to use it. You will never type a password to push again.gh auth login -
Tell git who you are
git config --global user.name "Sam" git config --global user.email "sam@example.com"
Putting a project up
Ask the agent to do it, or do it yourself — it is four commands:
cd /srv/myproject
git init -b main
git add -A
git commit -m "Initial commit"
# Creates the repo on GitHub and pushes, in one go.
# Swap --private for --public when you are ready to show it.
gh repo create myproject --private --source=. --push
The one that ruins people: committing a secret
Bots scrape newly pushed public commits within minutes, hunting for API keys, database passwords and cloud credentials. This is not a cautionary tale that happened to a friend of a friend. It is a constant, automated, industrial process, running right now, and it does not care how small your project is. A leaked cloud key can become a five-figure bill while you sleep.
Here is how it goes, and every step of it is reasonable:
- 1. You put the database password in
config.js, just for now, to get it working. - 2. It works. You commit everything, because everything is what changed.
- 3. A week later you make the repo public, because you are proud of it and you should be.
- 4. You notice
config.js, delete the line, and push the fix within the hour. - 5. Why is there a crypto miner on my server?
Step four is where people think they are safe, and it is the step that does nothing at all. Git kept the old version. It kept it because that is git's entire job.
.gitignore before your first commit, not after.
Git remembers everything, deliberately and permanently. A secret committed and then
deleted is still sitting in the history, perfectly readable, one click away, for as
long as the repository exists. Removing it properly means rewriting history and
force-pushing, and by then it has been scraped anyway. Thirty seconds of
.gitignore before the first commit removes this entire category of
afternoon from your life.
A sane starting .gitignore for any project:
# Secrets — never, in any form
.env
.env.*
!.env.example
*.pem
*.key
credentials.json
secrets/
# Build output — regenerate it, do not store it
bin/
dist/
build/
target/
node_modules/
__pycache__/
*.exe
# Local noise
.DS_Store
*.log
*.sqlite
.vscode/
.idea/
The !.env.example line is the trick worth stealing. You commit an
example file listing every variable the project needs, with obviously fake
values, so that anybody — a future you, a friend, an agent opening the repo cold —
can see what has to be set without ever seeing a real value. The genuine
.env never goes near git. Two files, one committed and one not, and the
question "what does this thing need to run?" has an answer that does not involve
reading the source.
# .env.example — committed, and completely safe
DATABASE_URL=postgres://user:password@127.0.0.1:5432/myapp
SESSION_SECRET=generate-a-long-random-string
STORAGE_BUCKET=my-bucket-name
Where secrets actually go
| Secret used by | Where it lives |
|---|---|
| Your app, on the server | An .env file readable only by the service user, or Environment= lines in the systemd unit. Never in the repo. |
| GitHub Actions | Repository → Settings → Secrets and variables → Actions. Referenced as ${{ secrets.NAME }} and masked in logs. |
| You, personally | A password manager. Not a text file, not a note, not a chat message to yourself. |
| Your agent | Nowhere special — it reads the same .env the app does. Tell it in CLAUDE.md that the file exists and must never be committed or printed. |
An environment variable, worked through once
That table says "environment variable" as though you had met one, and half this guide leans on the idea, so here is the one worked example. An environment variable is a named value that a program is handed by whatever started it. Not a file the program opens and not a line in its source: it arrives with the process, the way a letter arrives already inside an addressed envelope. The program asks for it by name and gets whatever the parent put there. Same binary, different envelope, different database. In a terminal, the envelope is the shell:
# For one command: the name, the value, then the command.
GREETING="hello from the shell" printenv GREETING
# For the rest of this terminal session. Every program you start from
# here on inherits it. Close the window and it is gone.
export GREETING="still hello"
printenv GREETING
# Everything your shell is currently handing out. PATH and HOME are in
# there, which is how every command you type gets found.
printenv
Every language reads one with a single call — process.env.GREETING,
os.environ["GREETING"], os.Getenv("GREETING") — and
that call is the whole integration. The program neither knows nor cares whether the
value came from a terminal, a file or systemd, which is exactly what makes it the
right seam for a secret: the code says the name, in the open, in git, and
the value is somebody else's problem. The .env.example above
is a list of names. The .env beside it is the values.
On the server, the thing that starts your app is not a terminal, it is systemd, so
the envelope has to be systemd's. It offers two ways to fill it:
Environment=NAME=value written into the unit, and
EnvironmentFile= pointing at a file of them. Use the file for anything
secret, for one plain reason: the unit in /etc/systemd/system is
readable by every account on the machine — it was installed
644, like nearly all of /etc — and
systemctl show prints its Environment= lines to anyone who
asks. A file that only the service's user can open is neither of those things. It
is the same .env the app reads in development, in the same format:
# /srv/myapp/.env — NAME=value, one per line. No spaces around the =,
# and no "export" in front of anything (see the notice below).
DATABASE_URL=postgres://myapp:a-long-random-password@127.0.0.1:5432/myapp
SESSION_SECRET=another-long-random-string
# The service's user owns it and nobody else can read it: the 600 from
# the permissions section, on the file that most deserves it.
sudo chown myapp:myapp /srv/myapp/.env
sudo chmod 600 /srv/myapp/.env
# Tell the unit about it without editing the unit. This opens an editor
# on a drop-in file, /etc/systemd/system/myapp.service.d/override.conf,
# and merges whatever you type there into the unit when you save.
sudo systemctl edit myapp
[Service]
EnvironmentFile=/srv/myapp/.env
sudo systemctl restart myapp
# What systemd was told: the unit, plus every drop-in, as one document.
systemctl cat myapp
# What the process actually received, asked of the process itself.
sudo cat /proc/$(systemctl show -p MainPID --value myapp)/environ | tr '\0' '\n'
The last command is the one to keep. The unit says what was asked for; the process's own environment says what happened, and when the two disagree the second one is the truth. Permissions made the same point about testing as the service rather than as yourself. This is the same point about reading the service rather than the file.
export.
A .env written for the shell, the kind you load with
source .env, has export at the start of every line, and
systemd drops every one of those lines without a word. Checked on this box: a file
with four variables, one of them written export NAME=value, handed
the service three, with no error and no log line. The app starts, finds its
connection string empty, and dies with a message about the database that says
nothing about why. Write the file as bare NAME=value and both readers
agree; just's set dotenv-load reads that shape too.
Committing like someone who wants to sleep
Commit every time something works, not once a day. Small commits mean small undos, and the size of your undo is the size of your worst afternoon. Ask the agent to do it — "commit that with a message describing what changed and why" — and it will write better commit messages than most people, largely because it has not yet learned to type "fixes" and go to lunch.
git log --oneline -10 # what happened recently
git diff # what have I changed but not committed
git restore <file> # throw away my changes to one file
git revert <commit> # undo a commit, safely, as a new commit
Decide what to build
What it really is, then what shape it takes, then what it is written in.
Everything is CRUD
Before you pick anything, notice what you are actually building. It is almost certainly the same thing as everyone else.
Here is the single most useful thing I know, and it took me an embarrassingly long time to see it. Nearly every piece of software that anybody pays for, uses daily, or gets excited about is four operations on some rows: create a row, read some rows, update a row, delete a row. Create, read, update, delete. CRUD. That is the job.
Not most software. Not business software. I have written games, marketing platforms, kiosks, scheduling systems and things I still could not describe at a party, and essentially all of it was getting some data into some tables and then making it pleasant enough to navigate that somebody who finds an ATM stressful would be comfortable using it. The world is not a holographic simulation. It is a collection of rows and columns, and a surprising amount of it is in a spreadsheet somewhere that Becky in accounting maintains by hand.
The exercise, on three ideas you might actually have
Pick whichever of these is closest to what is in your head. Then look along the row and notice that they are the same four tables wearing different clothes.
| The table | A multiplayer game | A little online shop | A booking system |
|---|---|---|---|
| People | players — name, avatar, rating | customers — email, address | clients — name, phone |
| Things | items, levels, characters | products — price, stock | rooms, staff, services |
| Events | matches — who, when, result | orders — who bought what, when | bookings — who, what, when |
| Lines | the moves in a match | the items in an order | the slots in a booking |
A game is a booking system where the room is a match and the client is a wizard. A shop is a booking system where the slot is a parcel. The clever, fiddly, genuinely hard parts of each — the matchmaking, the payment flow, the double-booking rules — are real work, but they sit on top of those four tables rather than replacing them. Get the tables right and the hard part has somewhere to stand. Get them wrong and no amount of clever fixes it.
Why this is worth knowing on day one
- It makes the project describable. "Four tables, a form to add a booking, a page to list them, a way to cancel one" is a thing you can hand to an agent. "A booking app" is not.
- It kills the paralysis. You are not attempting something nobody has done. You are attempting something everybody has done, in your own shape, which is a completely different feeling.
- It tells you what to build first. Create and read. Get a row in and get it back out on a page. Update and delete are half an hour once the first two work.
- It survives every fashion. The framework you use will be unpopular in four years. The four tables will be exactly the same.
So before you open the stack chooser, write your four tables down. On paper, in a text file, in a message to yourself. Then ask the agent to argue with them:
What are you actually building?
One chart, eight honest endings, and a strong opinion about seven of them.
Before anybody argues about languages, there is a question that decides almost everything and gets skipped: where does this thing have to run? Not what it does, not how big it might get — where it runs. A thing that runs in a browser and a thing that runs on an iPhone are different projects with different costs, different release cycles and different people who can say no to you, and the language is close to the last decision rather than the first.
So: answer the chart. Every ending tells you what to build, what it will cost you, and which languages and frameworks are genuinely available for that particular expedition — not the one I like best.
We are starting from the absolute beginning here.
Pick the place it must work, not everywhere you would eventually like it. This one question decides most of the rest.
Be honest rather than ambitious. The reality check two sections down is the long version of this question, and the answer is usually "nothing".
Almost everything people assume needs a desktop program does not. This is the question that tells you which.
The split is not 2D versus 3D exactly — it is whether the last ten per cent of performance and platform reach is the product.
Because there is a good chance you can have the listing without writing a native app, and a small chance you genuinely cannot.
Go outside. Have a nice time. Genuinely — there is no shame in not wanting to build software, and the world has plenty of people making things nobody asked for, myself very much included.
When you do want to make something, it will still be here, and it will be free, and it will probably have been updated since.
No installer. No review. No signing certificate. No store account, no platform fee, no rule change. Everybody is on the newest version the second you deploy, it works on a phone and a laptop and a library computer, and the way you distribute it is that you send somebody a link. That is a remarkable set of properties and the industry spent fifteen years trying to talk you out of noticing.
What you give up: it mostly wants a connection, and a short list of device features are off the table — the reality check is that list, honestly labelled. For the overwhelming majority of ideas, you give up nothing you were going to use.
HTML, CSS and JavaScript, because the browser runs nothing else. Start plain. Add htmx if you want interactivity without a build step, or React / Svelte / Vue once genuinely many things on screen have state at once. TypeScript if the project is big enough to forget your own function signatures, and Vite to build it.
Go is the default in this guide and the reasons are in Front end and back end. But this is the layer where language genuinely does not matter much: Python with FastAPI or Django, PHP with Laravel, Node with Express or Hono, C# with ASP.NET, Java with Spring, Ruby on Rails, Elixir with Phoenix. All of them serve HTTP correctly. Pick the one whose ecosystem has the thing you need.
SQLite in a file until you can name the reason for something else, then Postgres. See Where the data lives, and the memory rule in buying the server.
A progressive web app is a website with two extra files: a manifest that says what its name and icon are, and a service worker that can answer requests when the network cannot. Add those and a phone will offer to put it on the home screen, where it gets its own icon, its own splash screen, no address bar, and works with no signal. It is not a simulation of an app. For the kind of thing most people are building, it is an app, and it took an afternoon rather than a quarter.
And here is the honest part, because it is mostly Apple's fault and it is worth knowing before you commit. On an iPhone every browser is Safari underneath, so "it works in Chrome" tells you nothing about iOS. Push notifications do work — but only after the user has added your site to the home screen themselves, through a share-sheet ritual that no ordinary person discovers unaided, so anything load-bearing built on notifications will quietly not reach most of your iPhone users. Stored data can be evicted after a few weeks of the app not being opened, so offline is a convenience and never the only copy. Bluetooth, USB and serial do not exist in any iOS browser at all. Screen capture does not exist. Background work does not exist. Android, for what it is worth, is dramatically better at every single one of these.
A site.webmanifest, a service worker, and icons. No framework is required for any of it — this site ships a manifest and no build step. Workbox will generate a sensible service worker if you would rather not write the caching rules yourself.
The whole web stack applies unchanged, because it is a website. Nothing about your server, your database or your deploy changes.
If you ever do need a store listing or one native capability, Capacitor wraps the site you already built and lets you add native plugins for the one thing. You are not choosing to never have an app; you are choosing not to start with one.
You can have a real store listing without a native codebase. On Android, a Trusted Web Activity puts your site in the Play Store as an app, and it is your site — same URL, same deploy. On iOS, Capacitor produces a real app binary around your web app that you submit like any other.
Two warnings. Apple's review guidelines reject an app that is only a website in a wrapper with nothing added, so the wrapped version needs to be a genuine app experience rather than a bookmark — offline behaviour, a real icon, native share, push. And you have now bought a release cycle: review latency, a developer account, and a version of your app frozen on people's phones until they update.
Trusted Web Activity via Bubblewrap, or Capacitor. Play Store account is $25 once, ever.
Capacitor. Apple Developer Program is $99 a year, and you need a Mac to submit.
Your site, your server, your database, your deploy. That is the entire point of doing it this way.
If digital goods are sold inside an iOS app, Apple requires their in-app purchase system and takes a cut — and the rules about even mentioning that a cheaper price exists on your website have been through several courtrooms and are still moving. Android's terms are similar in shape.
The structure that avoids the entire problem: sell on the web, where you keep your margin and use Stripe like any other business, and let the app be the thing people use after they have paid. Every large subscription business you can think of does exactly this, which should tell you how well it works. If your product genuinely is consumable purchases inside a game loop, then in-app purchase is the cost of that business model and you should price it in from day one rather than discovering it in month five.
Stripe, Paddle or Lemon Squeezy. Paddle and Lemon Squeezy act as merchant of record, which means they handle sales tax in places you have never been.
StoreKit on iOS, Play Billing on Android. RevenueCat if you need both and would rather not maintain two receipt validators.
There is a genuine list, it is short, and you are on it. Bluetooth peripherals. NFC and tap-to-pay. Health and fitness data. Background location and geofencing. A watch app, a widget, a lock-screen presence. CarPlay or Android Auto. VoIP that rings like a phone call. Serious augmented reality. On-device machine learning that needs the platform's own accelerator. Anything that must keep working while the app is closed. None of that is available to a browser on iOS, and most of it is not available on Android either.
So build it native, with your eyes open about what it costs: a developer programme, a review queue between you and your users, a yearly round of work when the platform moves, and either two codebases or one cross-platform codebase with its own tax.
Swift with SwiftUI. Nothing else is a serious answer if the app is iOS-shaped, and it is a genuinely pleasant language.
Kotlin with Jetpack Compose. Same reasoning.
Flutter (Dart) for a single codebase that draws its own UI; React Native (JS/TS) if your team is already a web team; Kotlin Multiplatform to share logic and write each UI natively; .NET MAUI if you live in C# already.
Capacitor around the web app you already have, with a native plugin for the one capability. Keeps one codebase and one deploy for ninety per cent of the work.
Reading a folder in place, driving a label printer, sitting in the system tray, launching at startup, automating another application — the browser sandbox exists specifically to prevent all of that, on every platform, and no amount of cleverness gets around it. You need a program. Installers and signing covers what that costs, including the warning screen and the price of removing it.
Electron (JS/TS, ~80–150 MB, best documented), Tauri (Rust core, ~5–15 MB, uses the system browser engine), Wails (Go core, similar size). The interface is the HTML and CSS you already wrote.
C# with WinUI or WPF for Windows-only. Swift for Mac-only. Qt (C++ or Python) when you need real native widgets on all three and have the patience.
A single binary with no window at all, plus a tiny local web page for the interface. Go or Rust, two megabytes, ship it in a zip.
This is the easiest thing on the entire chart and the reason renting a server is such a good deal. No interface to design, no browser compatibility, no store, no installer, nobody's phone. A program that starts when the machine boots, does its job, writes to a log and gets restarted if it dies — which is a systemd unit, and the rest of this guide is already about exactly that.
Go, for the reason that matters here: one static binary, no runtime to install, tiny memory, and concurrency that is boring to write. Rust if the work is genuinely heavy. Python if the library you need only exists there, which for anything touching data or machine learning it usually does.
A systemd timer for anything periodic — see Serving it. A goroutine and a ticker if it lives inside a service you are already running.
Start with a table in your database and a loop. Reach for Redis or a real queue when jobs must survive a crash and be retried, and not one day before.
No window, no installer, no signing, no store, no warning screen. You hand somebody a file and it works. If the audience is you, your colleagues, or other developers, this is frequently the correct answer to something that was described to you as needing an interface — and it is by far the fastest thing to build, which means you find out whether the idea was any good this afternoon.
The standout, for one reason: go build produces a single static binary with no dependencies, and you can cross-compile for Windows, Mac and Linux from this box in three commands. Distribution is "download this file".
Same single-binary story, with the best argument-parsing library going (clap) and a compiler that is stricter than your future self.
Fine, and the right answer when the work is data-shaped. Ship it with uv or pipx so the person installing it does not have to understand virtual environments.
Under about a hundred lines, and only for glue. Past that it becomes a language you do not want to debug at midnight, and every experienced person has learned this the same way.
2D in a browser stopped being a compromise a long time ago. You get the best distribution mechanism in the history of games — a link — no store cut, no review, no install, and a player can be playing four seconds after hearing about it. For a first game this is not the lesser option, it is the sensible one.
Phaser for a full 2D framework, PixiJS if you want the renderer and none of the opinions, or plain canvas, which is genuinely fine and teaches you more.
Godot exports to HTML5 and is free and open source. LÖVE (Lua) exports via love.js. Bevy (Rust) compiles to WebAssembly.
Your own box, obviously — it is a static folder and some JavaScript. Then also on itch.io, where people who like small games actually look.
Do not write this from scratch. An engine is twenty years of other people's solved problems — physics, shaders, asset pipelines, controller support, platform export — and the part you actually want to spend your life on is the game. This is also the one branch of the chart where the "just self-host it" instinct does not apply: consoles mean dev kits and non-disclosure agreements, and storefronts take their cut.
Free, open source, no royalties, no company that can change the licence on you. GDScript is easy to learn; C# is supported. The obvious starting point in 2026.
C#, the largest asset store and tutorial corpus by a wide margin, and a licensing history worth reading before you commit a multi-year project.
C++ and Blueprints. What you use when the visuals are the product, with a royalty above a revenue threshold.
Rust, data-oriented, genuinely lovely, and still young enough that you will be reading source rather than documentation.
Pick your stack
Five questions. A recommendation, the reasoning, and a prompt to paste.
The chart above tells you what shape the thing is. This tells you what to type. They are two halves of the same question and the chart is the half that matters more — but once the shape is settled, the stack really does fall out of it, and the widget below will hand you an opening prompt for your agent.
"What should I build this in?" stalls more projects than any other question, and it is almost always asked in the wrong order. The language barely matters. What matters is the shape — who runs it, whether it needs the machine underneath it, and what has to still be there tomorrow. Answer those three and the stack falls out on its own, which is a relief, because arguing about languages is the most enjoyable way ever invented to not build something.
Answer five questions and get a specific recommendation, the reasoning behind each piece, and an opening prompt you can paste straight into Claude Code.
What it will probably tell you
The recommendations are opinionated on purpose, so here are the opinions in plain sight rather than hidden in a widget. They are defaults, not laws. If you already know you want something else, use it — you are allowed to have your own conventions, and you should end up with some. Being polyglot is nearly free now, which gets a whole section of its own.
- Most things should be websites. No installer, no store review, no signing certificate, and everybody is on the newest version the moment you deploy.
- Go for the server. One binary, no runtime to install, fast, genuinely good at doing many things at once, and a language that has barely changed in a decade — which means an agent's knowledge of it is not quietly three versions stale.
- SQLite for records, until you can say why not. A file on the disk, backed up by copying the file. Postgres is excellent and it is the right answer the day you can name the reason; "it feels more professional" is not a reason.
- Rust only where it earns it. Real systems programming, or compute heavy enough that the difference shows up on an invoice.
- Electron when somebody genuinely needs an icon on their desktop — and only then.
Can it be a website?
Tick what your idea genuinely needs. The honest answer is usually yes — and where it is no, it is usually the iPhone.
The web can do vastly more than people assume, mostly because the people telling you what it cannot do stopped checking around 2016. It also genuinely cannot do a short list of things, and one platform is responsible for nearly every gap on that list. On iOS, every browser is Safari underneath — Chrome on an iPhone is Safari wearing a different hat. So "it works in Chrome" and "it works on phones" are different sentences, and you find out which one you meant at the worst possible moment, in front of somebody.
Reality check
Tick everything your idea genuinely needs — not what would be nice.
Know what shape it is? The stack chooser turns that into a specific recommendation and an opening prompt.
Front end and back end
Two words people use constantly and almost never define.
Front end is what runs on the visitor's device. Back end is what runs on yours. That is the whole distinction, and every other rule falls out of it — the security ones especially, because you control exactly one of those two machines and it is not the one with your user sitting in front of it.
An analogy I have never managed to improve on. The HTML is the body of the car and the CSS is the paint and the rims. The JavaScript in the browser is the gas pedal — it makes things happen when you press them. Your program on the server is the engine under the bonnet, and the database is what is in the boot, which is where the things you actually care about are kept. People spend an astonishing amount of time arguing about rims.
HTML, CSS and JavaScript. Draws the page, reacts to taps. The user can read every line of it, so nothing secret ever goes here.
Encrypted, so nobody in between can read or alter it. This is what the certificate is for.
Terminates HTTPS, decides which of your programs a request belongs to, and hands it over. One nginx serves every site on the machine.
Checks who is asking, decides what they may see, reads and writes the database, returns a page or some JSON. This is where the rules live, because this is the machine you control.
The permanent record. A SQLite file next to the app, or Postgres listening on the loopback address only — either way, unreachable from the internet no matter what.
Why not JavaScript on the server
JavaScript has to be the front end; the browser runs nothing else. The usual next thought is that using it on the back end too gives you one language everywhere, and that was a genuinely strong argument back when a person had to hold both halves in their head at once. It is worth a great deal less when an agent is doing the typing for both.
What you are trading away for it:
- A dependency tree you cannot see the bottom of. A modest Node server routinely pulls in hundreds of megabytes and thousands of packages, every one of them something that can break, change under you, or be taken over. The Go equivalent usually has one or two — a database driver, maybe a crypto package — and you can name both of them.
- Churn. The correct way to do any given thing in the JavaScript ecosystem changes every couple of years, and the previous correct way becomes faintly embarrassing. That is bad for you and it is worse for an agent, whose knowledge of a fast-moving ecosystem is a photograph with a date on it. Go's syntax and idioms have barely moved in a decade, so what the model learned is still true.
-
Deployment weight. Node means shipping a runtime and a
node_modulesdirectory and keeping versions aligned. Go produces one file. Copy it to the server, run it. That is the deploy. - Memory. On a 4 GB box, a handful of Node services and their runtimes add up quickly. Small Go services measure in single-digit megabytes.
Choosing, layer by layer
| Layer | Start with | Reach for something else when |
|---|---|---|
| Page | Plain HTML, CSS, a little JavaScript | The interface has real state in many places at once — then a framework starts paying for itself. |
| Server | Go standard library net/http | Almost never for the web parts — routing, TLS, templates and JSON are all built in. Add a library when you can say out loud what it does for you. |
| Records | SQLite, in a file | Postgres when you have concurrent writers, real JSON querying, or a second service that needs the same data. |
| Big files | Object storage — R2 or GCS | The server's own disk, if it is small and you have backups. |
| Sessions / cache | The database, or just memory | Redis, once you have genuinely measured that you need it. Not before, and not because a blog post frightened you. |
| Background jobs | A goroutine, or a systemd timer | A real queue, when jobs must survive a crash and be retried. |
| Systems work | Rust | You are writing something where the last 10% of performance is the product. |
Installers and signing
When you really do need an icon on somebody's desktop, here is what it costs and what it does not.
Electron is a browser with the address bar taken off. Your app is HTML, CSS and JavaScript — the same skills, and frequently the literal same code as your website — wrapped up so that it gets a window, an icon, a tray presence and access to the machine's files and devices. VS Code, Slack and Discord are all Electron, which should settle any argument about whether it is real software.
The trade is an honest one and you should make it with your eyes open: you get out of the browser sandbox, and you pay in size — an Electron app starts around 80–150 MB — and in the considerable friction of persuading a human being to install anything at all.
The alternatives, briefly
| Option | Size | When it fits |
|---|---|---|
| Electron | ~80–150 MB | The default. Best documentation, best tooling, and an agent has seen an enormous amount of it. |
| Tauri | ~5–15 MB | Same web front end, Rust core, uses the system's own browser engine. Much smaller — at the cost of rendering slightly differently on each OS. |
| Go + Wails | ~10–20 MB | Like Tauri but Go instead of Rust. Neat if the rest of your stack is already Go. |
| A plain binary | ~2–15 MB | Command-line tools. No window, no installer, no signing drama. Ship it in a zip. |
The warning screen, and what it costs to remove
Software you hand out is unsigned unless you pay somebody to vouch for you. Unsigned software runs perfectly well — the operating system simply shows a frightening box first, and a certain number of your users will read it and stop. Note what you are actually buying here: not security, and not quality. You are buying the absence of a scary sentence.
| Platform | Unsigned looks like | Signing costs |
|---|---|---|
| Windows | A blue "Windows protected your PC" box. More info → Run anyway gets past it. | $120/yr via Azure Trusted Signing ($9.99/month, needs a verified business or sole trader), or $200–$600/yr for a traditional certificate, which since 2023 must live on a hardware token or HSM. |
| macOS | "Cannot be opened because the developer cannot be verified." Right-click → Open, or approve it in System Settings → Privacy & Security. | $99/yr for the Apple Developer Program. This also gets you notarisation, which is what actually removes the warning. |
| Linux | Nothing. Nobody asks. | Free. |
Put it on the internet
From a name, to a padlock, to installers for operating systems you do not own.
Pointing a domain
DNS is a phone book. You are adding one entry to it.
When somebody types your domain, their computer asks the internet's directory service which address that name belongs to. The directory is DNS and the entries in it are records. You need two of them, and there are only three record types worth learning in the first place, which is a much smaller subject than its reputation suggests.
- Your browser has the name
mysite.comand cannot use it. Every connection on the internet is made to a number. - So it asks a resolver — your provider's, or one you chose, like
1.1.1.1. The resolver does the walking. - If nobody has asked lately, the resolver goes to your nameservers: the only machine in this picture you control, answering with the record you typed.
- The answer comes back and the resolver keeps it for the length of the TTL, which is why changing a record is not instant for everybody.
- Now your browser opens a connection, straight to the number. DNS carried no part of your site. It answered one question.
- A
- Name → IPv4 address.
mysite.com→203.0.113.40. This is the one that matters. - AAAA
- Same, for IPv6. Add it if your provider gave you an IPv6 address; skip it otherwise.
- CNAME
- Name → another name.
www.mysite.com→mysite.com. Follow the second name to find the address.
In Cloudflare
- Open your domain, then DNS → Records
-
Add the apex record
Type
A, Name@(which means the bare domain itself), IPv4 address your server's IP, Proxy status DNS only, TTL Auto. -
Add www
Type
CNAME, Namewww, Target your bare domain, Proxy status DNS only. - Wait a few minutes Cloudflare updates almost immediately, but other networks hang on to old answers for a while. Ten minutes is typical, an hour is not alarming, and going and doing something else is more effective than refreshing.
Checking it worked
From the server, or from your laptop:
dig +short mysite.com
dig +short www.mysite.com
# Ask a public resolver directly, bypassing anything cached locally
dig +short mysite.com @1.1.1.1
You want your server's address back. Nothing at all means the record has not spread yet, so wait. The wrong address means either a cached answer or that you edited a different domain than the one you think you edited — both are extremely common, neither is interesting, and both are fixed by waiting or by looking properly. Almost every "DNS is broken" is really one of these two.
The tools that answer “did it work?”
Everything above assumed dig, which is the right tool and is not on
every machine. Here is the short list of things worth knowing about, split by what
they actually answer — because “my site is broken” is four
different questions wearing one coat, and the trick is to find out which one you
have before you start changing things.
dig +short
What does the internet say this name points at? Built into macOS and Linux; on Ubuntu it is in dnsutils.
nslookup
The same answer, built into Windows, no install. Chattier output; the line you want is the last one.
The four questions, in order
Run these top to bottom and stop at the first one that surprises you. Almost every “DNS is broken” is answered inside the first two.
# 1. Does the name resolve at all, and to what?
dig +short mysite.com
dig +short mysite.com @1.1.1.1 # bypass anything cached on your network
# 2. Which nameservers is the world being told to ask?
# After a move to Cloudflare these should be the Cloudflare pair.
dig +short NS mysite.com
# 3. Is the server itself answering, and with what?
curl -I https://mysite.com
# 4. Who owns the name, and when does it expire?
whois mysite.com | head -20
On Windows without WSL, question one is nslookup mysite.com and question
two is nslookup -type=NS mysite.com. Question three works as written:
curl has shipped with Windows for years. Question four has no built-in
equivalent, so use a WHOIS website, or the Ubuntu you installed on
your laptop, where every command on this page works exactly as printed.
@1.1.1.1 to a
dig and you have asked the world directly, skipping every cache
between you and it. If that is right and your laptop is wrong, nothing is
broken and the answer is a cup of tea.
Email, and the one thing not to self-host
Every other section here says run it yourself. This one says do not, and the reason has nothing to do with how hard it is.
A man moves to a new town, prints himself some beautiful business cards, and starts knocking on doors offering to carry everyone's post. The cards are perfect. His van is clean. He is punctual, polite and completely willing. Nobody gives him a single letter, because the town has a short list of couriers it trusts and he is not on it, and there is no form to fill in to get on it. You get on that list by years of not being a nuisance — and, as it turns out, the previous tenant of his address was a spectacular nuisance.
That is email, exactly, with no part of the analogy left over. You do not own your server's IP address; you rent it, and it arrives carrying a history that is not yours and a reputation that is. Gmail, Outlook and every corporate mail filter on earth are employed to be suspicious of strangers, and a brand new mail server on a cheap VPS is the most suspicious thing they will see all day. So, flatly: do not run your own mail server.
Not because it is difficult to install — apt will hand you a working
SMTP daemon in about ninety seconds, and that is the trap, because installing it is
roughly five per cent of the job. The rest is spent proving to strangers that you are
not a criminal, forever, using four separate mechanisms. Their names are worth knowing
anyway, because the relay you end up using will ask you to set three of them up, and it
is nicer to know what you are agreeing to:
- SPF
- A DNS record listing which servers are allowed to send mail claiming to be from your domain.
- DKIM
- A signature on every message, checked against a public key you publish in DNS, proving nothing was altered on the way.
- DMARC
- A DNS record telling receivers what to do when the first two fail — and, usefully, where to send you a daily report about it.
- Reverse DNS
- Your IP address resolving back to your mail server's name. Set by whoever owns the IP, which is your host, not you.
Get all four exactly right and you can still land in spam, because you share a subnet with somebody who did not. Meanwhile most VPS providers block outbound port 25 by default and would like to know why you want it opened, which stops the majority of people at step one — and that is your host doing you a favour it will never get any credit for. It goes like this, every time:
- You install a mail server. It takes an hour, which feels like proof that it was a real piece of engineering worth doing.
- You send a test message to yourself. It arrives. You are delighted, and you tell somebody.
- You send one to a friend on Gmail. It lands in spam, so you add SPF.
- You add DKIM. Then DMARC. Then you email your host about reverse DNS and wait a day for the reply.
- It works. For two weeks.
- Why is nobody getting the password reset emails?
So hand it to somebody whose entire business is being on that list of trusted couriers. There are two halves to this and they are separate problems, which is the part nobody says out loud — most people only actually need the first one.
Receiving — free, and about four clicks
What you want is you@yourdomain.com arriving in the mailbox you already
read forty times a day, and no second inbox to remember to check. Cloudflare Email
Routing does exactly that and costs nothing: turn it on, say which addresses to accept
and where to forward them, confirm at the far end. It writes its own MX records while
you watch. Pound for pound it is the best value available on a domain, and it is the
one part of this section to go and do tonight.
Make one of them a catch-all and you get something quietly excellent: every service you
ever sign up to can have an address of its own, so the day
ancient-forum@yourdomain.com starts receiving crypto spam you know exactly
who sold you, and you can switch off that one address without touching anything else.
Sending — a relay, for the two emails your app sends
The other half is transactional mail: the password reset, the receipt, the "somebody replied to you" notification. Your app hands the message to a relay over an API or an authenticated connection, and the relay's reputation carries it the rest of the way. That reputation is the entire product. You are not paying for a mail server — you can have one of those for nothing — you are paying for the benefit of the doubt.
| Relay | Free tier | After that |
|---|---|---|
| Resend | 3,000 a month, 100 a day | around $20/mo for 50,000 |
| Amazon SES | nothing worth planning around | $0.10 per 1,000 |
Resend is API first and the quicker of the two to wire up — a few lines of code and a key in an environment variable. Amazon SES is the cheapest by a distance once you have real volume, and the fiddliest to get going: new accounts start in a sandbox that will only deliver to addresses you have verified yourself, until you ask for that to be lifted.
Either is fine, and both are free at the volume of a project nobody has heard of yet. Start with whichever one's documentation you can follow at eleven at night; you can change your mind later by editing one environment variable, which is the entire reason the key lives in one.
Whichever you pick will hand you two or three DNS records to add. You do not type these from memory — the relay prints the real values — but it is worth seeing the shape once, so that a working domain looks familiar rather than magic:
# Receiving: Cloudflare writes these for you when you turn Email Routing on
example.com. MX 10 route1.mx.cloudflare.net.
example.com. MX 20 route2.mx.cloudflare.net.
example.com. MX 30 route3.mx.cloudflare.net.
# Sending: your relay gives you the real values. The shape is always this.
example.com. TXT "v=spf1 include:_spf.your-relay.com ~all"
relay._domainkey.example.com. TXT "v=DKIM1; k=rsa; p=MIGfMA0GCSq..."
_dmarc.example.com. TXT "v=DMARC1; p=none; rua=mailto:you@example.com"
Start DMARC at p=none, which means "tell me, do not act on it". Read the
reports for a fortnight, then tighten it. Setting a policy straight to
reject on day one is an excellent way to discover, silently and from the
far side, which of your own services had been sending mail as you.
security.txt naming an address a stranger can use to tell you about a
bug — that address should be one that reaches you. An inbox nobody opens is
worse than no inbox at all, because it looks like an invitation.
nginx, systemd, TLS
The three pieces between a request and your code. Your agent configures all of them.
You do not need to be able to write any of these from memory, and you never will, because you touch them about four times a year. What you do need is to know what each one is for, so that when something breaks at eleven at night you know which of the three to go and look at first. That is the whole reason this section exists.
- A visitor asks for
mysite.comon port 443 — the one port open to the world. - nginx decrypts it, reads which domain was asked for, and hands it to whichever of your programs owns that name. That is the whole trick behind unlimited sites on one address.
- It forwards to your app on
127.0.0.1:8090. Loopback means this machine only. nginx is on this machine; the internet is not. - Your app reads your data — a file on disk, or a database also on loopback — and answers back the way it came.
- A scanner tries port 8090 within minutes of your box existing, and gets nothing, because nothing there has ever been listening to the internet.
nginx — the front door
Only one program can hold port 443, and nginx is it. Every request to every site on the box arrives there, nginx reads which domain was asked for, and hands it to whichever of your programs owns that name. That is the entire trick behind one server hosting unlimited sites on one address, and once you have seen it you stop being impressed by hosting panels.
Which gives you the filing system for the whole machine, and it is worth adopting
deliberately, because it is what makes the fifth project as easy as the first:
every app gets a port, a systemd unit, an nginx server block and a health
endpoint. The ports are yours to hand out — 8090, 8091, 8092, in a tidy band
you keep a note of in each project's CLAUDE.md so nothing collides.
Nothing binds a public interface, ever. nginx is the only thing the internet can see.
Each site gets a file in /etc/nginx/sites-available/, linked into sites-enabled/:
server {
listen 443 ssl;
http2 on;
server_name mysite.com;
ssl_certificate /etc/letsencrypt/live/mysite.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/mysite.com/privkey.pem;
location / {
proxy_pass http://127.0.0.1:8090;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
sudo nginx -t # ALWAYS test before reloading
sudo systemctl reload nginx # zero-downtime; a bad config cannot take the site down
nginx -t then reload, every single time.
reload keeps the old config running if the new one is broken, so a typo
cannot take your sites offline. restart does not have that property; it
throws away what was working and then discovers the problem. Build the habit now,
while it is free, rather than learning the difference at the worst available
moment.
X-Real-IP, which nginx overwrites, and
X-Forwarded-For, which nginx adds to, keeping whatever the
visitor sent at the front of it. An app that reads the front of that list, or
believes either header from anyone but nginx, lets every visitor be anybody —
and every rate limit and ban by address with it. Believe X-Real-IP, and
only on connections from 127.0.0.1.
Not getting flooded
Nobody is going to attack your site. The people who attack things work from a list, you are not on it, and on the day you are, a rate limit is not what saves you anyway. What actually arrives, and arrives at everybody, is one enthusiastic program. It goes like this:
- You post the link somewhere busy.
- Somebody's aggregator likes it and fetches it every two seconds, forever, because its author forgot the line that says to wait.
- That page runs a query.
- That query is the slow one, the one you were going to look at.
- Why is production down?
A rate limit is you deciding in advance that no one address gets more than a fair
share, so that when the loop turns up the other visitors never find out it did. nginx
does it in ten lines, the ten lines are the same on every site you will ever build,
and the visitor it turns away is answered before your app hears a word about them.
That is the entire point: 429 Too Many Requests costs nginx nothing to
send, and your slow query costs whatever it costs.
# Above the server block. A zone is a shared counter, one per visitor
# address; 10 MB tracks about 160,000 of them. 20 a second is more
# than a person can click and less than a loop can loop.
limit_req_zone $binary_remote_addr zone=api:10m rate=20r/s;
server {
...
location /api/ {
# burst: how far ahead one visitor may run before hearing no.
# nodelay: answer the burst at once instead of spacing it out.
limit_req zone=api burst=40 nodelay;
limit_req_status 429;
proxy_pass http://127.0.0.1:8090;
}
}
Put it on the things that cost something — a login form, a search, an API, the page that runs the report — and leave the front page alone. A limit on a page that costs nothing to serve protects nothing and punishes the one visitor who opened six tabs. That block is the exact one on the box you are reading this on, where it guards the dashboard's API, and this is what it says when eighty requests arrive at once. Point it at your own site, not mine:
seq 80 | xargs -P 40 -I{} curl -s -o /dev/null -w '%{http_code}\n' \
https://mysite.com/api/stats | sort | uniq -c
$ seq 80 | xargs -P 40 -I{} curl -s -o /dev/null -w '%{http_code}\n' \
https://mecasasu.casa/api/stats | sort | uniq -c
48 200
32 429
Forty got in on the burst, eight more got in because the twenty-a-second allowance refilled during the half second it all took, and thirty-two were told no by a web server that never woke the app up. That limit has been on this site since the day it went up, and in ten days of logs it had fired exactly zero times before that test — which is the correct number. It is a seatbelt, not a steering wheel. Ten lines, once, and then you do not think about it until the morning it earns its keep.
CF-Connecting-IP, and
the fix is to let nginx believe that header — from Cloudflare's addresses
only, for the reason in the box above this one. Their list changes, so it is a
script, and the script runs from a timer, monthly.
#!/usr/bin/env bash
# /usr/local/sbin/cloudflare-real-ip: rebuild the list of Cloudflare
# addresses nginx is allowed to believe about who is really calling.
set -euo pipefail
out=${1:-/etc/nginx/snippets/cloudflare-real-ip.conf}
tmp=$(mktemp)
{
echo "# Generated $(date -u +%F) from cloudflare.com/ips-v4 and /ips-v6. Do not edit."
for v in ips-v4 ips-v6; do
curl -fsS --max-time 20 "https://www.cloudflare.com/$v" | awk 'NF { print "set_real_ip_from " $0 ";" }'
done
echo "real_ip_header CF-Connecting-IP;"
} > "$tmp"
grep -q '^set_real_ip_from ' "$tmp" # an empty list would trust nobody, silently
install -m 644 "$tmp" "$out" && rm -f "$tmp"
nginx -t && systemctl reload nginx
Then one line in the site's server block,
include snippets/cloudflare-real-ip.conf;, and $remote_addr
means the visitor again everywhere at once: the zone, the access log, the
X-Real-IP your app reads, and fail2ban. The script was run on this box
before it was printed here, and it is not installed, because this domain's cloud is
grey and there is nothing to see through. The awk is there for a reason
you would only find by running it: the last line of Cloudflare's list has no newline
on the end, and a plain sed glues the header directive onto the last
address. nginx happened to accept that. It would not have to.
systemd — keeping it running
If you start your program by typing its name, it dies the moment you close the terminal, which is a surprisingly common way for a website to have a very short life. systemd is what turns it into a service: started at boot, restarted when it crashes, logging somewhere you can actually read. Nothing you build is genuinely online until it survives a reboot you did not plan.
# /etc/systemd/system/myapp.service
[Unit]
Description=My application
After=network-online.target
[Service]
User=myapp
WorkingDirectory=/srv/myapp
ExecStart=/srv/myapp/bin/server -addr 127.0.0.1:8090
Restart=always
RestartSec=2
# Sandboxing. Cheap to add, and it means a bug in your code
# cannot become a bug in the whole machine.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
# strict makes the whole disk read-only to this service, even folders
# it owns. Name the one place it writes, which must already exist.
# If the app writes nothing, delete this line.
ReadWritePaths=/srv/myapp/data
# Memory ceiling, and the line that makes it mean anything: without
# MemorySwapMax the service is pushed into swap instead of stopped,
# and thrashes there under its "limit". See Stress testing.
MemoryMax=256M
MemorySwapMax=0
[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now myapp # start it, and start it at every boot
systemctl status myapp
journalctl -u myapp -f # follow the log
sudo adduser --system --group --no-create-home myapp. If your program
is ever compromised, whoever is in there gets the permissions of an account that
owns nothing, can log in nowhere, and has no interesting files — rather than yours.
One command, once, at the point where it costs nothing. The --group
is not optional: without it the account lands in nogroup, and the
myapp:myapp that the next subsection hands to chown
names a group that does not exist. This page left it out until 16 September
2026, and found out by running its own command.
Permissions — who is allowed to touch what
Giving the service its own user buys you one failure you are all but guaranteed to
meet, so meet it here: the program works perfectly when you run it, and fails
the moment systemd runs it. Nothing about the program changed. The
person did. You ran it as sam, the unit runs it as
myapp, and every time anything touches any file, Linux asks the same two
questions — who is asking, and what is that user allowed to do here?
Every file answers with three kinds of people and three permissions. The people are
the owner, the file's group, and everyone
else. The permissions are read, write and
execute. ls -l prints all nine on every line, and it
looks like line noise right up until somebody points at the columns once:
$ ls -l /srv/myapp/config.toml
-rw-r--r-- 1 sam sam 812 Sep 14 09:12 config.toml
│└┬┘└┬┘└┬┘ │ │
│ │ │ │ │ └─ group
│ │ │ │ └───── owner
│ │ │ └─ everyone: r
│ │ └──── group: r
│ └─────── owner: rw
└─ - file, d directory
As a number, read is 4, write is 2, execute is 1, and a permission like
644 is just the three sums side by side — owner, group, everyone.
Which sounds like homework, and it is not, because there are really only three
numbers you will ever type:
644—rw-r--r--- For files. You can change it, anybody can read it. Almost every file on a web server wants to be exactly this.
755—rwxr-xr-x-
For directories, and for programs. On a program, execute means run it. On
a directory it means walk through it, and that is the one that catches
people: a file is a room at the end of a corridor, and an unlocked room is no use to
you if one of the doors on the way there is locked.
namei -l, below, shows you every door. 600and700- Owner only. You have typed both already — on your SSH keys, which ssh refuses to use if anyone else could read them, and on the Cloudflare token in Infinite subdomains. Secrets get these, and nothing else does.
When a log says Permission denied, test as the service, not as
yourself. You are the one user on the machine it definitely works for, which
makes you the worst possible person to reproduce it. sudo -u borrows the
service's identity for a single command:
# Who does the service actually run as?
systemctl show -p User myapp
# Try the thing that failed, as that user
sudo -u myapp cat /srv/myapp/config.toml
sudo -u myapp test -w /srv/myapp/data && echo writable
# Every directory on the way down, with its owner and permissions
namei -l /srv/myapp/data
# The usual fix: the service owns what it writes
sudo chown -R myapp:myapp /srv/myapp/data
That last line is the correct fix nine times out of ten, and the tenth is when it gets aimed at the whole project instead of the folder the program writes to. Leave the program itself owned by you. A service that can rewrite its own binary is a service that, on its worst day, can make that bad day permanent.
And somewhere in your first month a forum answer will tell you to run
chmod -R 777. It works. It makes the error go away in exactly the way
that taking the front door off its hinges fixes a sticky lock, and the answer with
the most upvotes is very rarely the one that has to live in the house afterwards.
ProtectSystem=strict, and Databases said to
keep SQLite in /srv/myapp/data, owned by myapp. Each was
good advice, and together they could never work, because strict makes
the entire disk read-only to the service — including folders it owns.
The ownership was perfect. sudo -u myapp could write there all day.
The service could not, because the sandbox is a second lock that the owner and the
permissions never see. The tell is supposed to be the wording,
Read-only file system instead of Permission denied, except
SQLite does not pass that on: it says unable to open database file and
sends you off to stare at a path that is fine. It was found on 14 September 2026 by
running that unit on this box, and the ReadWritePaths= line is the fix.
The rule to keep: when it works as the user and fails as the service, it is not the
user any more. It is the unit.
Timers — things that run while you sleep
Sooner or later something needs to happen every night with nobody awake: clearing out
sessions that expired, rebuilding a report, deleting uploads nobody finished. The
traditional tool for that is cron, and cron works, and has worked since
before most of the people reading this were born. On a new server, use a
systemd timer instead. The reasons sound small, and each one is
the difference between a job that failed and a job you know failed:
-
It logs where everything else logs.
journalctl -u, same as your app. Cron's idea of reporting a problem is to email the output to a mailbox on the server itself, and on a server with no mail system — which is every server in this guide — it writes one line saying it threw the output away. -
It catches up. With
Persistent=true, a job whose time came and went while the machine was off runs at the next boot instead of quietly skipping a night. - It answers "when does this next run?" with a command, rather than with you reading five-field cron syntax and doing time zones in your head.
-
It is the same kind of thing as your service. Same
User=, same sandbox, same everything from Permissions just above. There is nothing new to learn, which is the best feature a tool can have.
This box does not have cron installed at all. Nothing has missed it.
A timer is two files. The .service is the job, and a job that runs and
exits is Type=oneshot. The .timer says when, and has the
same name so systemd can pair them up. The job here is one query against a
sessions table, as an example; yours is whatever your app needs doing at
four in the morning.
# /etc/systemd/system/myapp-prune.service
[Unit]
Description=Delete expired sessions
[Service]
Type=oneshot
User=myapp
ExecStart=/usr/bin/sqlite3 /srv/myapp/data/app.db "DELETE FROM sessions WHERE expires_at < datetime('now');"
ProtectSystem=strict
ReadWritePaths=/srv/myapp/data
# /etc/systemd/system/myapp-prune.timer
[Unit]
Description=Delete expired sessions, nightly
[Timer]
OnCalendar=*-*-* 04:10:00
RandomizedDelaySec=10m
Persistent=true
[Install]
WantedBy=timers.target
# Check the schedule means what you think, before you trust it
systemd-analyze calendar '*-*-* 04:10:00'
sudo systemctl daemon-reload
sudo systemctl enable --now myapp-prune.timer # enable the timer, not the service
systemctl list-timers # every timer, and when each next fires
sudo systemctl start myapp-prune.service # run the job once, right now
journalctl -u myapp-prune.service -n 20 # and see what it said
systemctl --failed # anything that has broken, timers included
OnCalendar also understands hourly, daily and
things like Mon *-*-* 09:00. systemd-analyze calendar prints
the next time it will actually fire, in the server's own time zone, which is very
probably UTC and very probably not yours. That one command is the difference between
a report that lands at nine in the morning and one that lands at four.
- You add a nightly job, and it works.
- It works every night for seven months.
- You stop thinking about it, which is correct.
- The disk fills, or a table gets renamed, and it starts failing every night.
- Why is the newest backup from March?
TLS — the padlock
Let's Encrypt issues real
certificates, free, in about fifteen seconds. certbot asks for one,
proves you control the domain, writes the files, edits your nginx config, and sets up
automatic renewal.
sudo apt install -y certbot python3-certbot-nginx
# Both names in one certificate. Point the DNS records first —
# certbot proves control by answering a request on that domain.
sudo certbot --nginx -d mysite.com -d www.mysite.com
Certificates last 90 days and renew themselves via a timer that is installed automatically. You can prove the machinery works without spending a real, rate-limited issuance:
sudo certbot renew --dry-run
sudo certbot certificates # what exists and when it expires
/etc.
Which works, right up until you rebuild the box, or move the project, or simply
cannot remember what you changed nine months ago or why. The nginx config and the
systemd unit are not settings. They are source code for your project, every
bit as much as the program is, and they belong in the repository — mine live in a
deploy/ folder at the root of each project, and
just deploy installs them from there. The copies under /etc
are output. Treat them as disposable and you can rebuild a whole server from git in
an afternoon; treat them as the original and you are one reinstall away from
archaeology.
Infinite subdomains
One domain. Unlimited names. All free, all with certificates.
This is the part people consistently underestimate, and it is the single best-value
thing on the whole box. mysite.com costs money.
anything.mysite.com costs nothing, forever, and you can have as many as
you can think of names for. Every project, every demo, every half-idea you want to
show somebody gets a real address on the actual internet, which does something to how
seriously you take it.
mysite.com the main site
app.mysite.com the actual product
api.mysite.com the JSON API
docs.mysite.com documentation
status.mysite.com is everything up?
demo.mysite.com the thing you show people
git.mysite.com whatever you feel like
Adding one is three steps, and your agent will do all three:
- A DNS record Type
A, nameapp, pointing at the same IP. Or use the wildcard below and skip this forever. - An nginx server block Same as before, with
server_name app.mysite.comand a different port. - A certificate
sudo certbot --nginx -d app.mysite.com
The wildcard shortcut
Once you are adding these regularly, stop doing it one at a time and do it once properly. A wildcard DNS record sends every name to your server, and a wildcard certificate covers the lot of them.
The catch is a genuinely interesting one. Proving you own *.mysite.com
cannot be done over HTTP, because there is no single page to serve — the name you are
claiming is infinite. So it has to be proved through DNS instead, which means certbot
needs permission to write a temporary record into your domain, which means an API
token. This is the first real secret in the guide, so it gets handled properly.
-
Add the wildcard DNS record
Type
A, name*, your server's IP. Keep the specific records for anything you want to behave differently. - Make a scoped Cloudflare token My Profile → API Tokens → Create Token → Edit zone DNS, and restrict it to this one zone. Scope it as tightly as the interface will let you: it should be able to edit DNS for one domain and do absolutely nothing else — not your account, not your billing, not your other domains. A leaked token that can do one small thing is an inconvenience. A leaked token that can do everything is a weekend.
-
Store it so only root can read it
sudo mkdir -p /root/.secrets echo "dns_cloudflare_api_token = YOUR_TOKEN_HERE" | sudo tee /root/.secrets/cloudflare.ini sudo chmod 600 /root/.secrets/cloudflare.ini -
Ask for the wildcard
sudo apt install -y python3-certbot-dns-cloudflare sudo certbot certonly \ --dns-cloudflare \ --dns-cloudflare-credentials /root/.secrets/cloudflare.ini \ -d mysite.com -d '*.mysite.com'
From then on a new subdomain is one nginx block. No DNS change, no certificate request, no waiting for anything to propagate. Point the new server block at the wildcard certificate files, reload, and the name is live. This is the moment the box stops feeling like a rented computer and starts feeling like somewhere you live.
server {
listen 443 ssl default_server;
server_name _;
ssl_certificate /etc/letsencrypt/live/mysite.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/mysite.com/privkey.pem;
return 444; # close the connection without a response
}
Databases
Where everything that has to survive a restart actually lives.
A plain file works fine until two things write to it at once, or until you want to ask a question more complicated than "give me everything". A database handles concurrency, constraints and queries. On this box, start with SQLite.
I want to say this one out loud, because the internet will tell you otherwise and because I have given the opposite advice myself. Reach for Postgres when you can name the reason: several programs writing at once, real JSON querying, another service that needs the same data. Those are good reasons and the day you have one you should switch without agonising. "It feels more professional" is not a reason. It is a costume. One binary, one box, one writer and a few thousand requests is SQLite territory, and it is roomier territory than almost anybody expects.
| Option | Use it when |
|---|---|
| SQLite default | One program, one machine. A single file on disk, no service to run, no password, no port. This is most projects, for much longer than you think. |
| PostgreSQL | Concurrent writers, real JSON querying, or a second service that needs the same data. Excellent, and it will not be the thing that limits you. |
| Redis | When you have measured a specific need for a fast in-memory cache. Not before. It eats RAM, which is the scarce resource here. |
| MongoDB | Rarely worth it at this size. Postgres stores and queries JSON perfectly well if that is what you wanted. |
SQLite, properly
There is no service to install and start. The database is a file, so the only real decisions are where it lives and who owns it.
sudo apt install -y sqlite3
# The database is one file. Keep it beside the app, owned by the
# service user, and nowhere the web server can serve it from.
sudo mkdir -p /srv/myapp/data
sudo chown myapp:myapp /srv/myapp/data
# As the service user, so the file it creates belongs to the app
sudo -u myapp sqlite3 /srv/myapp/data/app.db
-- Write-ahead logging: readers stop blocking the writer. Set once,
-- stored in the file, and the single most useful line here.
PRAGMA journal_mode = WAL;
-- Foreign keys are OFF by default in SQLite, for very old reasons.
-- Your application must turn them on for every connection it opens.
PRAGMA foreign_keys = ON;
CREATE TABLE users (
id INTEGER PRIMARY KEY,
email TEXT NOT NULL UNIQUE,
created_at TEXT NOT NULL DEFAULT (datetime('now'))
);
.quit
Then the path goes in .env, exactly as a connection string would:
DATABASE_PATH=/srv/myapp/data/app.db
cp.
Copying a live SQLite database with cp can capture it mid-write and
give you a file that looks fine and is not. Use VACUUM INTO, which
reads one consistent snapshot while the app carries on writing:
sudo sqlite3 /srv/myapp/data/app.db "VACUUM INTO '/var/backups/myapp-$(date +%F).db'"
Almost everywhere else will tell you .backup, and on a quiet database
it is fine. It copies in pieces and starts again from the top whenever the app
writes in between, so on a busy one it never finishes. On this box, with an 80 MB
file, it was still going at thirty seconds with twenty-five writes a second
arriving. VACUUM INTO took a fifth of a second at a thousand.
That one line is your entire
just backup recipe. See
Backups that hold for where the copy then goes, which is the
part that actually matters.
When to reach for Postgres
On the day you can name the reason — a second service wants the same data, you have genuinely concurrent writers, you need to query inside JSON documents — Postgres is already on this box's shopping list and it is a twenty-minute move. Do it then, on purpose, and not before.
sudo apt install -y postgresql
# Confirm it is listening ONLY on the loopback address.
# 127.0.0.1 = only this machine. 0.0.0.0 = the entire internet.
sudo ss -tlnp | grep 5432
127.0.0.1 by default, and there is
essentially never a reason to change that. Your application runs on the same
machine, so it connects locally. If you need to browse the data from your laptop,
do not open the port — tunnel it over the SSH connection you already have:
ssh -L 5432:127.0.0.1:5432 sam@203.0.113.40
Now localhost:5432 on your laptop is the server's database,
encrypted, with no new port open and no new password to manage.
A database per project
sudo -u postgres psql
CREATE DATABASE myapp;
CREATE USER myapp WITH PASSWORD 'a-long-random-string-from-your-password-manager';
GRANT ALL PRIVILEGES ON DATABASE myapp TO myapp;
-- Postgres 15 and later need this second grant as well. Without it
-- your app connects successfully and then cannot create a single
-- table, with an error that does not obviously point here.
\c myapp
GRANT ALL ON SCHEMA public TO myapp;
\q
Then the connection string goes in .env, which git never sees:
DATABASE_URL=postgres://myapp:the-password@127.0.0.1:5432/myapp
GitHub Actions
Free computers that build your software for operating systems you do not own.
GitHub Actions runs commands on GitHub's machines whenever something happens in your repository. Public repositories get it free; private ones get a monthly allowance a personal project will not come close to spending.
Before the examples, the honest rule, because this is an area where people build
elaborate machinery they do not need: CI earns its place when it does
something your own computer cannot. Building a Windows installer when you do
not own Windows — that is worth every minute. Cutting a release with four binaries
attached to it — worth it. Running your tests on a robot when
just check already runs them here, on this box, in four seconds, before
you deploy — that is a pipeline built to look like a real company, and it will cost
you an afternoon of YAML for a result you already had.
So, in order of how much they actually pay:
- Build releases On every version tag, produce binaries or installers for Windows, macOS and Linux and attach them to a page people can download from. This is the one. If you only ever set up one workflow, make it this one.
- Check every push Run
just checkon every commit. Genuinely useful the moment somebody else is committing too, and mild duplication of your own habits until then. - Deploy on push Shown below because it is satisfying and because you will want it eventually. You do not need it while
just deploytakes one second and you are the only person here.
The thing nobody tells you about cross-compiling
You do not need a Mac to build for macOS, or a Windows machine to build for Windows. Go compiles for every platform from any platform — one Linux runner produces all of them in about thirty seconds, for nothing. Rust and Zig will do the same with a bit more setup. This still feels like getting away with something, and it is the single best reason on this page to have Actions at all.
.github/workflows/release.yml — this builds four binaries and publishes
them whenever you push a tag starting with v:
name: release
on:
push:
tags: ["v*"]
permissions:
contents: write # needed to create the release
jobs:
build:
runs-on: ubuntu-latest
strategy:
matrix:
include:
- { goos: linux, goarch: amd64, ext: "" }
- { goos: darwin, goarch: arm64, ext: "" } # Apple Silicon
- { goos: darwin, goarch: amd64, ext: "" } # older Intel Macs
- { goos: windows, goarch: amd64, ext: ".exe" }
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version: "1.26"
- name: Build
env:
GOOS: ${{ matrix.goos }}
GOARCH: ${{ matrix.goarch }}
CGO_ENABLED: "0"
run: |
out="myapp-${{ matrix.goos }}-${{ matrix.goarch }}${{ matrix.ext }}"
go build -trimpath -ldflags "-s -w" -o "$out" ./cmd/myapp
- uses: softprops/action-gh-release@v2
with:
files: myapp-*
git tag v0.1.0
git push --tags
A minute later there is a release page with four downloads on it, built on hardware you do not own, for operating systems you may never have used. That is the entire distribution story for a command-line tool, and it is done.
Electron apps need real machines
Desktop installers are the exception: a .dmg has to be assembled on
macOS, and a Windows installer on Windows. Actions gives you both, so the matrix
changes from cross-compilation to running the same job on three different runners.
jobs:
build:
strategy:
matrix:
os: [ubuntu-latest, windows-latest, macos-latest]
runs-on: ${{ matrix.os }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: "22"
- run: npm ci
- run: npx electron-builder --publish always
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
That produces a .exe installer, a .dmg, and an
AppImage and .deb — all attached to the release
automatically. GITHUB_TOKEN is provided by Actions itself; you do not
create it.
Deploying to your own server on every push
Optional, genuinely satisfying, and not something you need yet. Push to
main and the live site updates itself. Set this up when pushing a button
twice a day has started to annoy you, or when a second person joins — not on day one
because it looks professional.
-
Make a deploy-only key
On the server:
ssh-keygen -t ed25519 -f ~/deploy_key -C "github-actions". Adddeploy_key.pubto theauthorized_keysof a dedicateddeployuser that owns only this project. -
Put the private half in GitHub
Repository → Settings → Secrets and variables → Actions →
New repository secret. Call it
DEPLOY_KEY. GitHub masks it in every log line. Then delete the private key from the server — GitHub is the only place it needs to exist. -
The workflow
name: deploy on: push: branches: [main] jobs: ship: runs-on: ubuntu-latest steps: - name: Deploy over SSH env: KEY: ${{ secrets.DEPLOY_KEY }} HOST: ${{ secrets.DEPLOY_HOST }} run: | mkdir -p ~/.ssh && chmod 700 ~/.ssh printf '%s\n' "$KEY" > ~/.ssh/id_ed25519 chmod 600 ~/.ssh/id_ed25519 ssh -o StrictHostKeyChecking=accept-new deploy@"$HOST" \ 'cd /srv/myapp && git pull --ff-only && just deploy'
uses: line runs somebody else's code on a machine holding your
secrets. Stick to actions/*, which is GitHub's own, plus a small number
of well-known ones, and pin each to a version rather than a moving branch — a
branch is a promise that whatever is there tomorrow is still fine. This is one of
the very few places in this guide where being a bit paranoid costs you nothing at
all, so spend it here.
Backups that hold
Not an advanced topic. The part that separates software from a demo.
Everything in this guide rests on giving an autonomous agent broad permission on a real machine, and that is only a reasonable thing to do because of this section. With good backups the worst realistic outcome is losing an afternoon. Without them it is losing the project, and losing a project is a very quiet kind of disaster — nothing explodes, you just stop being able to continue.
Backups get filed under "advanced" by an astonishing number of tutorials, usually in a paragraph at the end that says you should probably look into it. They are not advanced. Load balancing is advanced. Failover is advanced. Backups are the price of admission, and they are four lines in a shell script.
- A dead disk, a wrong
rm, a migration that ate a column, or an agent doing exactly what you asked. They all arrive at the same place. - Your server holds the working copy — which is to say, the copy that just broke. One copy on one machine is not a backup.
- A nightly dump runs on a timer and writes a consistent copy whether or not anybody remembers it. That is the whole difference between a backup and an intention.
- It goes somewhere else: a different provider, on somebody else's disks. A copy that dies with the box has only ever protected you from your own mistakes.
- The restore is the only step that was ever the point, and it is the one almost nobody has run.
And it is not really about the agent. Here is what actually destroys data:
- A
DROP TABLEagainst production because that terminal looked like the other terminal. - A migration that runs perfectly and quietly throws away a column.
- The provider's disk failing, or your account being suspended over a billing mix-up nobody meant.
- You, at midnight, entirely certain that directory was the old one.
An old mantra worth adopting: nothing is ever truly deleted. Prefer
marking a row as gone to actually removing it. Prefer moving a directory to
.old over rm -rf. If something is irrevocably destroyed, no
good can come of that — there is no upside to being unable to get it back, ever, in
any situation I have encountered in twenty years.
3–2–1
The rule that has survived decades: three copies of anything you care about, on two different kinds of storage, with one of them somewhere else entirely.
Copy one
The database and files on your server. This is not a backup; it is the thing you are protecting.
Copy two
Nightly dumps in /var/backups. Instant to restore from — but useless if the machine itself is gone.
Copy three
Object storage at a different company. This is the copy that survives fire, deletion and a suspended account.
Grandfather–Father–Son
Keeping every copy forever is expensive. Keeping only last night's is worse than useless, and this is the bit that catches people who think they are covered. Some damage takes weeks to notice — a silently broken migration, a table gradually emptied by a bug, a file corrupted in March. If your only backup is last night's, you have carefully and reliably preserved the broken state, once a day, with great discipline.
The answer is a rotation that gets sparser as it gets older. Frequent copies of the recent past, occasional copies of the distant past.
Twenty-four files. Together they let you go back to any of the last seven days, any of the last five weeks, or any month in the past year — and they cost a few dollars a year to store, because a compressed database dump of a small project is measured in megabytes.
Four things to back up
| What | Where it is | How |
|---|---|---|
| Databases | A SQLite file, or Postgres | Nightly. sqlite3 … "VACUUM INTO" for SQLite — never cp on a live file — or pg_dump -Fc for Postgres. Irreplaceable; this is the one that matters. |
| Code | /srv/* | Already on GitHub if you commit. That is a backup, and an off-site one. |
| Uploads | /srv/*/data | Anything users gave you. Not in git, not regenerable. |
| Config & secrets | /etc/nginx, /etc/systemd/system, every .env | Not in git by design. Encrypt these before they leave the box. |
The script
/usr/local/bin/backup.sh. It finds what you run on its own: every SQLite
file in a data folder under /srv, and every Postgres database
if there is a Postgres. Ask your agent to fit it to your actual setup — it will get the
details right and it will not get bored halfway through.
The line worth reading twice is the --exclude. Until September 2026 this
page printed a script that tarred /srv/*/data whole, live database and
all, two sections after telling you never to copy a live database. Run against a
busy one on this box, tar stopped with an error five times out of five,
so the script never reached the step that sends anything off the machine. The
database inside each tarball would not open, six out of six. It looked finished.
It had four tidy numbered steps. What it actually did was write a broken file
to the same disk it was meant to protect, and then stop.
#!/usr/bin/env bash
set -euo pipefail
shopt -s nullglob # a pattern that matches nothing is skipped, not an error
umask 077 # backups hold every secret on the box: root reads them, nobody else
DEST=/var/backups/nightly
STAMP=$(date -u +%Y%m%d)
mkdir -p "$DEST"
# 1. Every SQLite database: a consistent snapshot, taken while the
# app keeps running. Never tar or cp the live file.
for db in /srv/*/data/*.db; do
name=${db#/srv/} # myapp/data/app.db
name=${name//\//-} # myapp-data-app.db
out="$DEST/${name%.db}-$STAMP.db"
rm -f "$out" # VACUUM INTO refuses to overwrite
sqlite3 "$db" "VACUUM INTO '$out'"
done
# 2. Every Postgres database, if this box runs Postgres
if id postgres &>/dev/null; then
dbs=$(sudo -u postgres psql -Atc \
"SELECT datname FROM pg_database WHERE NOT datistemplate AND datname <> 'postgres'")
for db in $dbs; do
sudo -u postgres pg_dump -Fc "$db" > "$DEST/$db-$STAMP.dump"
done
fi
# 3. Uploads and configuration, without the live database files.
# Step 1 already has those, and a copy taken mid-write is corrupt.
tar czf "$DEST/files-$STAMP.tar.gz" \
--exclude='*.db' --exclude='*.db-wal' --exclude='*.db-shm' \
/srv/*/data /etc/nginx /etc/systemd/system /srv/*/.env
# 4. Encrypt tonight's files for the trip. Only the public key lives
# on this box; the private key lives in your password manager.
AGE_RECIPIENT=age1yourpublickey...
OUT=$(mktemp -d)
trap 'rm -rf "$OUT"' EXIT
for f in "$DEST"/*-"$STAMP".*; do
age -r "$AGE_RECIPIENT" -o "$OUT/${f##*/}.age" "$f"
done
# 5. Push them off this machine. rclone talks to R2, B2, S3, Google
# Cloud Storage and about forty other things. copy, never sync:
# sync makes the bucket match this folder, and step 6 empties it.
# Sunday's set is also a weekly, the 1st's is also a monthly.
rclone copy "$OUT" remote:my-backups/daily
if [ "$(date -u +%u)" = 7 ]; then
rclone copy "$OUT" remote:my-backups/weekly
fi
if [ "$(date -u +%d)" = 01 ]; then
rclone copy "$OUT" remote:my-backups/monthly
fi
# 6. Prune: keep 7 days locally. The bucket prunes itself (below).
find "$DEST" -type f -mtime +7 -delete
sudo chmod +x /usr/local/bin/backup.sh
It needs two tools and somewhere to send things. remote: in step 5 is a
name rclone looks up, and you make it once, with the keys from an R2 API token for the
bucket or a B2 application key. It goes in with sudo because the backup
runs as root, so root is the one who has to find it:
sudo apt install age rclone
# Cloudflare R2
sudo rclone config create remote s3 provider=Cloudflare \
access_key_id=YOUR_KEY_ID secret_access_key=YOUR_SECRET \
endpoint=https://YOUR_ACCOUNT_ID.r2.cloudflarestorage.com \
acl=private no_check_bucket=true
# Or Backblaze B2, instead of the above
sudo rclone config create remote b2 account=YOUR_KEY_ID key=YOUR_APPLICATION_KEY
Run it every night with a systemd timer:
# /etc/systemd/system/backup.service
[Unit]
Description=Nightly backup
[Service]
Type=oneshot
ExecStart=/usr/local/bin/backup.sh
# /etc/systemd/system/backup.timer
[Unit]
Description=Run the nightly backup
[Timer]
OnCalendar=*-*-* 03:30:00
RandomizedDelaySec=15m
Persistent=true # if the box was off at 03:30, run it at next boot
[Install]
WantedBy=timers.target
sudo systemctl enable --now backup.timer
systemctl list-timers backup.timer # when does it next run?
sudo systemctl start backup.service # run it right now, once
Let the storage do the forgetting
Grandfather–father–son is two jobs, and they belong in different places.
Choosing which copies are the weeklies and the monthlies is two
if lines in the script, above: Sunday's set goes into
weekly/ as well, and the 1st's into monthly/.
Deleting is the storage's job. Give each folder a lifecycle rule and the
bucket throws away what has aged out on its own, so nothing on the server ever
deletes anything off-site, which is the point of off-site.
| Lifecycle rule on | Delete after | What stays |
|---|---|---|
daily/ | 7 days | The last seven nights |
weekly/ | 35 days | The last five Sundays |
monthly/ | 365 days | The 1st of each of the last twelve months |
In R2 that is the bucket's Settings, then Object Lifecycle Rules, one
rule per folder. In B2 it is the bucket's Lifecycle Settings, then
Use custom lifecycle rules, with the folder as the File Path,
daysFromUploadingToHiding set to the number above and
daysFromHidingToDeleting set to 1. Both delete within a day or so of the
date rather than on the minute, which for backups is as precise as it needs to be.
The script above ran for four hundred simulated nights on this box with those three
rules applied, and the bucket settled at seven, five and twelve: the chart, exactly.
rclone sync, and that you should turn on versioning and
let the provider keep the history. sync makes the bucket match the
local folder, and the local folder only ever holds a week, so every file the prune
removed vanished from the bucket the following night. Forty simulated nights on this
box left nine days off-site, under a chart promising a year. Versioning was meant to
catch that, and R2 does not have it: Cloudflare lists the versioning calls as not
implemented. B2 does, and keeps every version forever unless told otherwise, which
is the same mistake in the other direction, and with a bill. The old script also
uploaded the tarball unencrypted, every .env inside it, one box above a
red notice telling you not to.
Both Cloudflare R2 and Backblaze B2 are cheap and have real free tiers — R2 notably charges nothing for egress, which is what makes restoring free rather than an unwelcome surprise. Fifty gigabytes of backups costs a few dollars a year.
age-keygen
That prints three lines and saves nothing. All three go in your password manager;
the age1… public key goes in the script as
AGE_RECIPIENT. The server only ever holds the half that locks. Keep the
private key off it: a backup you cannot decrypt because the key burned with
the machine is not a backup. Getting one back, into a scratch folder:
cd "$(mktemp -d)"
sudo rclone copy remote:my-backups/monthly/myapp-data-app-20260901.db.age .
# Paste the AGE-SECRET-KEY-1… line, press Enter, then Ctrl-D
age -d -i - -o app.db myapp-data-app-20260901.db.age
sqlite3 app.db "PRAGMA integrity_check;"
A backup you have never restored is a rumour
This is the step everybody skips, and it is the reason backups fail on the one day they are needed. A backup you have never restored is not a backup, it is a belief. The script that runs nightly and writes a zero-byte file will keep doing that, cheerfully, for a year. Once, by hand, so you know what it feels like before it matters:
# Restore last night's dump into a scratch database
sudo -u postgres createdb restore_test
# The dumps are root's alone, so root reads the file and postgres restores it
sudo cat /var/backups/nightly/myapp-20260906.dump | sudo -u postgres pg_restore -d restore_test
# Does it actually contain data?
sudo -u postgres psql restore_test -c '\dt'
sudo -u postgres psql restore_test -c 'SELECT count(*) FROM users;'
# Clean up
sudo -u postgres dropdb restore_test
app.db-wal and app.db-shm are still
sitting beside the database, and SQLite will replay that -wal on top of
whatever you copy in. Tried twice on this box: once the restored database would not
open at all, and once it passed its integrity check while holding the broken data it
was supposed to replace. The second one is worse. Nothing complained, and nothing
ever would have. So the old files go into a folder, not over the top, and not in the
bin either:
sudo systemctl stop myapp
cd /srv/myapp/data
# app.db and its -wal and -shm, if there are any, move aside together
old=damaged-$(date +%F-%H%M)
sudo mkdir "$old"
sudo mv app.db* "$old"/
# The backup goes in owned by the app, and WAL goes back on:
# VACUUM INTO writes its copy in the default journal mode
sudo install -o myapp -g myapp -m 600 /var/backups/nightly/myapp-data-app-20260914.db app.db
sudo -u myapp sqlite3 app.db "PRAGMA journal_mode = WAL;"
sudo systemctl start myapp
Then make the machine prove it, every night
Once a quarter is what people write down. The honest number is once, the week it was set up, and then never, and I know that because it is my number. It is not laziness. The backup has a timer, and a timer does not care whether anybody remembers it. The restore has you, and you have a Tuesday.
So give the restore a timer too. /usr/local/bin/restore-drill.sh,
below, runs an hour after the backup and does what you just did by hand,
mechanically, to every database the backup script found: it opens tonight's snapshot, checks it is sound, counts every table
against the live database, checks that tonight's set actually left the box, and
prints one line saying whether it all came back. It restores into scratch copies,
and into a scratch database whose name says drill, and it only ever
reads the live ones. If it dies half way through, the trap still prints
the verdict and still writes the log line, because a drill that fails quietly is
the exact thing it exists to catch.
#!/usr/bin/env bash
set -euo pipefail
shopt -s nullglob
umask 077
DEST=/var/backups/nightly
LOG=/var/backups/restore-drills.log
STAMP=$(date -u +%Y%m%d) # tonight's set: this runs an hour after backup.sh
WORK=$(mktemp -d)
scratch_db=""
verdict=""
databases=0
tables=0
# The last line is the verdict and the exit code agrees with it. A trap prints
# it, so a drill that dies half way through still says so, in the log too.
finish() {
rc=$?
rm -rf "$WORK"
[ -z "$scratch_db" ] || sudo -u postgres dropdb --if-exists "$scratch_db" \
|| echo "drop $scratch_db by hand" >&2
[ -n "$verdict" ] || verdict="FAILED: stopped early, exit code $rc"
echo "$(date -u +%FT%TZ) $verdict" >> "$LOG"
echo "RESTORE DRILL $verdict"
case $verdict in OK*) exit 0 ;; *) exit 1 ;; esac
}
trap finish EXIT
fail() { verdict="FAILED: $*"; exit 1; }
# A table with rows tonight must have rows in the restore. Equal is not
# required, because a day of activity sits between the two counts.
compare() { # compare <database> <table> <live rows> <restored rows>
printf '%-28s %-28s %8s live %8s restored\n' "$1" "$2" "$3" "$4"
[ "$3" -eq 0 ] || [ "$4" -gt 0 ] || fail "$1: $2 has $3 rows live and none in the restore"
tables=$((tables + 1))
}
# 1. Every SQLite database, found the way backup.sh finds them. Tonight's
# snapshot has to exist, pass its integrity check, and hold the rows.
for db in /srv/*/data/*.db; do
name=${db#/srv/}; name=${name//\//-}
snap="$DEST/${name%.db}-$STAMP.db"
[ -f "$snap" ] || fail "no snapshot of $db tonight; expected $snap"
cp "$snap" "$WORK/drill.db"
[ "$(sqlite3 "$WORK/drill.db" 'PRAGMA integrity_check;')" = ok ] || fail "$snap fails its integrity check"
while IFS= read -r t; do
compare "$name" "$t" "$(sqlite3 -readonly "$db" "SELECT count(*) FROM \"$t\"")" \
"$(sqlite3 "$WORK/drill.db" "SELECT count(*) FROM \"$t\"")"
done < <(sqlite3 -readonly "$db" "SELECT name FROM sqlite_master WHERE type = 'table' AND name NOT LIKE 'sqlite_%'")
databases=$((databases + 1))
done
# 2. Every Postgres database, if this box runs Postgres. Tonight's dump goes
# into a scratch database whose name says drill, is counted against the
# live one, and is dropped. The live database is only ever read.
if id postgres &>/dev/null; then
dbs=$(sudo -u postgres psql -Atc \
"SELECT datname FROM pg_database WHERE NOT datistemplate AND datname <> 'postgres'")
for db in $dbs; do
dump="$DEST/$db-$STAMP.dump"
[ -f "$dump" ] || fail "no dump of $db tonight; expected $dump"
scratch_db="${db}_drill_$$"
sudo -u postgres createdb "$scratch_db"
sudo -u postgres pg_restore --no-owner -d "$scratch_db" < "$dump" || fail "$dump did not restore cleanly"
while IFS= read -r t; do
compare "$db" "$t" "$(sudo -u postgres psql -Atc "SELECT count(*) FROM $t" "$db")" \
"$(sudo -u postgres psql -Atc "SELECT count(*) FROM $t" "$scratch_db")"
done < <(sudo -u postgres psql -Atc \
"SELECT quote_ident(schemaname) || '.' || quote_ident(relname) FROM pg_stat_user_tables" "$db")
sudo -u postgres dropdb "$scratch_db"; scratch_db=""
databases=$((databases + 1))
done
fi
# 3. Tonight's set is off the box as well. The key that opens those copies is
# not on this machine, on purpose, so this proves they arrived. Opening
# one is the drill you do by hand, with the key, once a quarter.
sent=$(for f in "$DEST"/*-"$STAMP".*; do echo "${f##*/}.age"; done | sort)
missing=$(comm -23 <(echo "$sent") <(rclone lsf remote:my-backups/daily | sort))
[ -z "$missing" ] || fail "not off-site tonight: $(echo $missing)"
verdict="OK: databases restored $databases, tables counted $tables, files off-site $(echo "$sent" | wc -l)"
sudo chmod +x /usr/local/bin/restore-drill.sh
sudo /usr/local/bin/restore-drill.sh # once, by hand, before you trust the timer
What it printed on this box, against a scratch app with three tables:
$ sudo /usr/local/bin/restore-drill.sh
myapp-data-app.db users 5 live 5 restored
myapp-data-app.db sessions 0 live 0 restored
myapp-data-app.db posts 4 live 4 restored
RESTORE DRILL OK: databases restored 1, tables counted 3, files off-site 2
$ tail -n 3 /var/backups/restore-drills.log
2026-09-16T03:06:04Z OK: databases restored 1, tables counted 3, files off-site 2
The rule it applies is loose on purpose: a table with rows in the live database
has to have rows in the restore. Not the same number, because a day's activity
sits between the snapshot and the count — the first run on this box, with a
second process inserting the whole time, printed 10 live and
6 restored for the table being written to, and that is correct. Zero
where there should be rows is the failure it is looking for, because it is the
one that hides. A backup script that quietly started writing empty files would
pass a size check and pass an integrity check, and it fails this. So does a night
with no snapshot at all, a snapshot that will not open, a dump that will not
load, and a file that never reached the bucket; each was tried here, and each
got its own sentence and an exit code of 1.
Then the timer: the same two files as every other job on this page, with the one
line from Finding out before anybody tells you that turns a
failure into a message. systemd runs ExecStartPost only when the
drill passed, so a failing drill never checks in, and the silence is the alarm.
Tested here both ways: the passing run pinged, the failing run did not. Delete
that line if you have not set a check up; the log still has the verdict.
# /etc/systemd/system/restore-drill.service
[Unit]
Description=Nightly restore drill
[Service]
Type=oneshot
ExecStart=/usr/local/bin/restore-drill.sh
ExecStartPost=/usr/bin/curl -fsS -m 10 --retry 3 -o /dev/null https://hc-ping.com/your-check-uuid
# /etc/systemd/system/restore-drill.timer
[Unit]
Description=Run the restore drill, an hour after the backup
[Timer]
OnCalendar=*-*-* 04:30:00
RandomizedDelaySec=10m
Persistent=true
[Install]
WantedBy=timers.target
sudo systemctl daemon-reload
sudo systemctl enable --now restore-drill.timer
sudo systemctl start restore-drill.service # run it now, once
journalctl -u restore-drill -n 5 --no-pager # and read the verdict
What it does not prove, and the page will not pretend otherwise: that the off-site copy opens. The private key is not on the box, by design, as the encryption notice above says, so the machine can only check that tonight's files arrived. Opening one is the drill with you at the keyboard — the four lines in that notice, once a quarter, with the key pasted from your password manager. The machine does the every-night part and you do the once-a-quarter part, and between the two of you nobody is hoping.
Put both in the justfile so they take ten seconds and there is no excuse:
backup:
sudo /usr/local/bin/backup.sh
drill:
sudo /usr/local/bin/restore-drill.sh
.sql dump of the entire database, not of the one table I had
just emptied, and I found this out afterwards rather than before. Two mistakes, and
the first one was not the expensive one.
What changed after that was not "be more careful", because I had already been telling myself that for two decades and it does not work. What changed was mechanical: different terminal colours for live machines, one SSH session open at a time, and never a restore command typed until I have read what is actually in the file. Everyone truncates the wrong table eventually. The difference between a senior person and a beginner is not that it stops happening — it is that by then they have a plan B, and a C, and a D.
The day you think you have been hacked
The question people ask that day is how do I remove the malicious files. I
have answered it on forums for years, and the honest answer is that it is the wrong
question, because it assumes the thing you found is the thing that is there. You
found the one that used all the CPU. Whoever put it there had a shell on your
machine, as root or as near to it as they needed, for some length of time you do not
know, before they got round to that. And every tool you would use to look for the
rest — ls, ps, ss — runs on the
machine you no longer trust, where a rootkit's first job is to teach
ls to lie. You cannot audit a machine with the machine. Cleaning it up
is asking the burglar whether anything is missing.
Before any of that, though: is it a break-in at all? Most of what feels like one is
the weather. On this box, over ten days, 14,910 login attempts
arrived for 2,071 usernames that do not exist, from
1,792 addresses — admin and ubuntu
first, then crypto, wallet and blockchain,
because what they are hoping to find is a coin wallet. Seventy-five logins were
accepted, every one of them with a key and none with a password, and fail2ban banned
two addresses and let the rest fail. None of that is a hack. It is
Lock it down doing its job in public. A hack is one of these,
and each is one command, run from a fresh terminal:
# A login you did not make. Ubuntu keeps these in the journal now;
# there is no wtmp, so the old `last` command prints nothing.
sudo journalctl -u ssh --since -30d -o cat | grep Accepted
# A key you did not add. You know how many lines there should be.
sudo cat /root/.ssh/authorized_keys /home/*/.ssh/authorized_keys
# A process you cannot account for, and who it is talking to
ps -eo user,pid,pcpu,etimes,cmd --sort=-pcpu | head -15
sudo ss -tnp state established
# A job you did not schedule
ls -la /etc/cron.d /var/spool/cron/crontabs 2>/dev/null
systemctl list-timers --all
# Files changed in the last two days where files should not change.
# Read it, do not count it: apt and adduser touch /etc for a living.
sudo find /etc /usr/local/bin -mtime -2 -type f
If any of those shows you something you did not do, the box stops being a patient and becomes a witness, and what happens next is the first ten minutes from the section on your own mistakes, with one change: nothing gets undone, ever, because nothing on it can be believed.
- Snapshot it at the provider, then power it off at the provider. The snapshot is the copy of the wreckage. The power button in the control panel is the one thing about the machine an intruder cannot argue with, which is why it is not
shutdowntyped into the box. - Rotate every secret it ever held. Everything in
.env, the Cloudflare token, the credentialghstored, the deploy key in Actions, and the key the backup script uses to reach the bucket — then go and look at the bucket, because that key can delete your backups and you want to know now whether it did. Your own SSH key only if its private half was ever on the box, which is the reason it never should have been. - Rebuild on a fresh machine. Twenty minutes, the repository, the unit. This is the recovery from "everything is gone", and it is the same forty-five minutes, because you built it to be.
- Restore the data from a copy dated before the break-in, not from last night. That is the entire argument for the weeklies and monthlies above: the gap between the damage and your noticing it. And ask, once, whether the data was the door — an upload that was really a program, a row that was really a command — because restoring the door reinstalls it.
- Only then, ask how. On the snapshot, read from a different machine, never booted. It is one of the five ways from the hardening section, nearly always, and if you cannot find which, you will be doing this again in a month.
Which is why the drill above runs every night rather than once a quarter. The exercise I would give every team, and every person with one box, is this: suddenly all of your servers are offline and out of reach, backups included — how long until you are operational again? If the answer is "a few hours, probably", you are fine, and you will be fine on this day too. If the answer is "I could not recover from that", nothing in this section will help you when it comes, and the thing to fix is that answer, today, on a Tuesday, while the box is merely being scanned.
Habits that compound
The difference between having one project and having a hundred.
Launch checklist
The small things that make something look finished. Your ticks are saved.
Not one of these is hard. Every single one is forgettable, and together they are the whole difference between something that feels shipped and something that feels like a draft somebody left open. Work down the list with your agent — most items are one sentence of instruction and thirty seconds of its time.
I am not being theoretical about the forgetting. I went and counted across my own
repositories: favicons in about half of them, an apple-touch-icon in slightly fewer, a
web manifest in a third, a robots.txt in a quarter, a sitemap in barely
one in eight. That is a descending staircase and it has exactly one cause, which is
that I got to the end of a project, felt finished, and stopped. This list exists so
that "felt finished" and "is finished" line up more often for you than they have for
me.
Monorepo to the moon
Use the right language for each job. The tax for doing that has mostly gone.
The old reason to pick one language and stay there was entirely human. Every extra language was another set of idioms, another toolchain, another thing to remain fluent in while not using it. Nobody is excellent at six at once, so teams standardised, and they were right to.
That constraint has largely dissolved, and it is one of the genuinely new things about working this way. An agent is competent in all of them at the same time and does not get rusty. So the question stops being "what do I know?" and becomes "what is actually right for this piece?" — which is a much better question and one you were never able to afford before.
What that looks like in practice
/srv/myproject/
├── justfile one entry point for the whole thing
├── docs/
├── web/ Go — the site and the API
├── worker/ Rust — the CPU-heavy image pipeline
├── scripts/ Python — data wrangling and one-off jobs
├── desktop/ Electron — the installable companion app
└── deploy/ nginx, systemd, the workflows
One repository. One history. One just check that runs all four test
suites. The pieces talk over HTTP or a queue, so any one of them can be rewritten or
thrown away without the others noticing — which is the actual point of microservices,
as distinct from the version people cargo-cult, where each service gets its own repo,
its own pipeline, its own deployment ceremony and its own reasons to be broken on a
Friday.
| Language | Reach for it when |
|---|---|
| Go | Anything that serves requests or runs as a service. The default, and it should be. |
| Rust | Genuinely low-level work, or compute so heavy the runtime cost is the product. Not for a CRUD app, whatever the internet says. |
| Python | Data, scripting, machine learning, anything where the library you need only exists there. |
| JavaScript | The browser, and Electron. It is the only option in one place and a fine one there. |
| C / C++ | Talking to hardware, or wrapping a library that has no bindings anywhere else. |
| SQL | Not optional. Every stack has it, and it is worth understanding rather than hiding behind a layer that generates it badly. |
docs/DECISIONS.md, in a sentence, on the day you make it. Six languages
because each piece genuinely needed one is engineering. Six languages because each
was fashionable in the month that part got written is a maintenance problem you
have posted to your future self, and he cannot refuse delivery.
Every language, and what it is actually for
Here is the thing I wish somebody had told me twenty years ago, because it would have saved me a decade of extremely enjoyable arguments. The useful facts about a programming language are almost never the ones people fight about. Whether it is object-oriented, functional or procedural; whether it has classes or traits or interfaces; whether the syntax is pretty — those are the things beginners are taught to compare, and they are close to irrelevant when you are choosing one for a job.
What decides it is duller and much more consequential. Where can this thing run? JavaScript wins the browser by walkover, because it is the only language the browser has. What can it reach? Swift can ask an iPhone for its health data and nothing else can. Does the library I need exist in it? If the model you want to run only has Python bindings, the language argument is already over. And what does shipping it cost? A Go binary is one file you copy; a Node service is a runtime and a thousand packages you have to keep aligned.
So: the map. Columns are where a language primarily runs, which is the property that matters most and is almost never how these charts are organised. The lines are real relationships — one compiles to another, one is written in another, or the two are so routinely paired that learning one drags in the other. Click anything to see what it is for, what it is genuinely good at, and one honest criticism, because every language on here has at least one and the ones I like best get theirs too.
A map of where things run, not a ranking. Tap a language for the honest version.
For Making a page do something when somebody touches it. It is the only language a browser runs, which makes it the only language in the most widely deployed runtime ever built.
Strong Unavoidable, therefore universal. Enormous ecosystem, instant feedback loop, and every device on earth already has an interpreter for it.
Fair dig Designed in ten days in 1995 and it shows in the corners — the equality rules alone have cost the industry decades. The correct way to do any given thing also changes every two years, which is worse for an agent than for you: its knowledge of a fast-moving ecosystem is a photograph with a date on it.
For JavaScript with the types written down. It compiles to JavaScript and runs everywhere JavaScript runs, which is everywhere.
Strong The editor can answer questions about your own code, and a whole class of "undefined is not a function" at eleven at night simply stops happening. On anything beyond a few hundred lines it pays for itself quickly.
Fair dig It is a type system bolted onto a language that does not have one, so the types are erased at runtime and they can lie to you. Also a build step you did not have before, which is a real cost on something small.
For Services, tools and anything that answers requests. Built at Google explicitly to be readable by people who did not write it.
Strong go build makes one static file with no dependencies — copy it to the server, run it, that is the deploy. Doing many things at once is boring to write rather than clever. The standard library covers HTTP, TLS, templates and JSON without a single third-party package. And the language has barely changed in a decade, which means what an agent learned about it is still true.
Fair dig Deliberately plain to the point of being repetitive; you will write if err != nil until you dream about it. Generics arrived late and are still a bit awkward. If you love expressive type systems you will find it joyless, and that is a fair thing to want.
For Data, scripting, automation, and every single thing to do with machine learning. If a model, a dataset or a scientific library exists, it has Python bindings first and possibly only.
Strong You can read it out loud. The library situation is not a close contest — it is the whole reason the field standardised on it. Superb for the script you write once and run for six years.
Fair dig Slow on raw computation, which mostly does not matter because the fast parts are C underneath. Deployment is its real weakness: versions, virtual environments and system packages produce the phrase "works on my machine" more reliably than any other language here.
For Websites, which it has done since 1995 and still runs a large fraction of the web on, whatever anybody tells you at a conference.
Strong The deployment model is genuinely the simplest there is: a file in a folder is a page, and there is no build and no process to keep alive. Modern PHP with Laravel is a pleasant, fast, well-documented way to build a real application, and hosting for it costs nothing anywhere on earth.
Fair dig Twenty years of accumulated inconsistency in the standard library, and a reputation formed in an era it has genuinely left behind — which still costs you, because the tutorials you find are as likely to be from 2009 as from 2026. I learned on procedural PHP before it had real object orientation, so the affection here is not neutral.
For Running JavaScript off the browser. One language for both halves of a web application, which used to be the strongest argument in this entire section.
Strong The largest package registry in existence, excellent at many simultaneous connections, and nobody on your team has to learn a second language. Deno and Bun are newer runtimes for the same language with better defaults and a fraction of the ceremony.
Fair dig A modest service pulls in hundreds of megabytes and thousands of packages, every one of which can break, change under you or be taken over — and on a 4 GB box, the runtimes add up. The "one language everywhere" argument was strong when a human held both halves in their head, and is worth considerably less when an agent is typing both.
For Business applications, Windows desktop software, and — via Unity — an enormous share of all the games you have ever played.
Strong One of the best-designed large languages in use, with tooling that is genuinely excellent and a standard library that has thought of your problem already. Modern .NET is fast, cross-platform and open source, which surprises people whose impression of it is fifteen years old.
Fair dig Still carries a cultural gravity towards Microsoft's way of doing things, and the ecosystem assumes a certain size of organisation — a lot of documentation is written for a team with an architect. Heavier to deploy on a small Linux box than Go, though far lighter than it used to be.
For Systems that have to run for fifteen years and be maintained by people who have not met each other. Banks, insurers, logistics, Android's foundations, and a colossal amount of the internet's plumbing.
Strong The JVM is one of the great engineering achievements of the field — the garbage collectors alone are extraordinary — and the ecosystem is deep, stable and extremely well documented. Nothing you write will be the first thing of its kind.
Fair dig Verbose in a way that stopped being fashionable, memory-hungry enough to matter on a small server, and its enterprise culture produced some of the most famously over-engineered code ever committed. The language itself has improved enormously; the reputation has not caught up.
For Web applications, mostly with Rails, which invented or popularised half the conventions every other web framework now copies.
Strong Optimised for the happiness of the person typing, and it shows — it is a delight to read and write. Rails will get one person from nothing to a working, tested, deployed application faster than almost anything else in this list, which is exactly the property that matters most for a first project.
Fair dig Slower than the compiled options and memory-hungry at scale. The magic that makes it so pleasant is also genuinely hard to debug when it goes wrong, because the method you are looking for may not appear anywhere in the source. Hiring and fashion have moved on, which affects how much help you find.
For Systems that must not go down, and anything with a very large number of simultaneous connections — chat, presence, live dashboards, telemetry.
Strong It runs on the Erlang virtual machine, which was built by a phone company for switches that were not allowed to fail, and it shows: processes are nearly free, failure is isolated and supervised rather than fatal, and Phoenix LiveView lets you build a live-updating interface with almost no front-end JavaScript at all.
Fair dig A small community, so the "somebody has already solved this" property is weaker than anywhere else on this map. Functional and concurrent-by-default is a genuine shift in how you think, which is a cost even though it is a worthwhile one. Not a first language for a first project.
For Operating systems, drivers, embedded devices, and the inside of nearly everything else on this page. Your Python interpreter is C. So is your database, your web server and your kernel.
Strong Small, fast, portable to hardware that has never heard of anything else, and stable for fifty years. It is the lingua franca every other language speaks when it needs to talk to something: if a library exists anywhere, it has a C interface.
Fair dig It will let you do absolutely anything, including the thing that corrupts memory somewhere else entirely and crashes an hour later. A significant share of every serious security vulnerability of the last thirty years is a memory bug in C or C++. That is not a moral failing of the language; it is a property you must budget for.
For Anything where the last ten per cent of performance is the product: game engines, browsers, trading systems, computer vision, audio.
Strong Zero-cost abstraction is a real thing and C++ delivers it — you can write expressive code that compiles to exactly what you would have written by hand. Nothing else has quite this combination of control and altitude, and Unreal, Chrome and most of the software you use daily are built on it.
Fair dig Enormous. There are several C++ dialects inside C++ and no two codebases agree on which they use, so "knowing C++" means less than knowing most languages here. Compile times are a genuine daily cost, the error messages are legendary, and it inherits C's memory model along with its consequences.
For Systems work where memory bugs are unacceptable, compute heavy enough to show up on an invoice, and command-line tools you want to be perfect.
Strong It proves at compile time that a whole category of bug cannot happen, without a garbage collector. Cargo is the best package and build tool of anything on this map, the compiler's error messages genuinely teach you, and it compiles to WebAssembly so the same code can run in a browser.
Fair dig The learning curve is real and the borrow checker will beat you for a fortnight. It is the wrong tool for a CRUD application no matter what the internet says — you will spend your attention on lifetimes instead of on the thing you are building. Slow to compile, and there is a strain of evangelism around it that has done the language no favours.
For The job C does, with fifty years of hindsight applied. Also, oddly, one of the best cross-compilers in existence — people use it to build C projects for other platforms without writing a line of Zig.
Strong Simple enough to hold in your head, explicit about allocation in a way that makes memory behaviour obvious, no hidden control flow, and it compiles C directly so you can adopt it one file at a time.
Fair dig Not finished. The language is still changing under its users, the ecosystem is small, and an agent's knowledge of it is likelier to be out of date here than anywhere else on this map. Wonderful, and not where you put something that has to work next year without attention.
For Apple's platforms, properly. If your app needs HealthKit, a watch face, a widget, CarPlay or the camera pipeline at full depth, this is the language and there is no second answer.
Strong A genuinely modern, pleasant language with real safety guarantees, and SwiftUI builds interfaces quickly. It reaches every API Apple has, the day Apple ships it — which is the entire reason to be here.
Fair dig Practically speaking it is one company's platform: you need a Mac, a yearly developer programme, and a tolerance for a yearly round of work when the OS moves. It runs on Linux servers in theory and almost nobody does it. Excellent language, narrow door.
For Android, officially and overwhelmingly. Also a very good server language, and via Multiplatform it can share business logic between an iPhone app and an Android one.
Strong Everything Java does, with about half the typing and null-safety built into the type system, on the same superb JVM and with the whole Java ecosystem available unchanged. Jetpack Compose is a real pleasure compared to what Android development used to be.
Fair dig Android itself is the tax, not the language: build times, emulators, and a fragmented device population you cannot test properly. On the server it is excellent but always slightly the second choice to Java in terms of how much has already been written down.
For Asking questions of data. Not optional, not replaceable, and the one language on this map you will still be using in twenty years whatever else changes.
Strong Fifty years old, declarative, and astonishingly good at its job — you say what you want and something cleverer than you works out how to get it. It is also the highest-leverage thing on this page to learn properly: one correct query routinely replaces two hundred lines of application code that loops.
Fair dig Every database speaks a slightly different dialect, so "knowing SQL" does not perfectly transfer. It is easy to write something that works on a thousand rows and falls over at a million. And the tooling that hides it from you — the object-relational mapper — will generate queries you would be appalled by, which is exactly why you should be able to read them.
For Gluing programs to each other. Not writing applications. Every server you will ever touch already has it, and so does every deploy pipeline.
Strong Unbeatable at its actual job: take the output of one thing, feed it to another, do it a thousand times. Learning pipes, redirection and exit codes properly will pay you back every working day for the rest of your career, and most of this tutorial is that skill.
Fair dig Past about a hundred lines it becomes a language you do not want to debug at midnight — quoting rules that punish you for a space in a filename, no real data structures, and error handling you have to remember to ask for. Everybody learns this the same way, by writing a 400-line script and then rewriting it in Python.
That is eighteen and it is not everything. Dart carries Flutter; Lua is embedded in more games and devices than you would guess; R still owns whole corners of statistics; Haskell, OCaml and Clojure will each change how you think about code even if you never ship a line of them; Fortran and COBOL are both, right now, running things you depend on. None of them are missing because they are bad.
Everything in this guide is one set of preferences that happens to work. Take the parts that suit you and bin the rest with my blessing — the thing that matters is not my conventions, it is that you have conventions at all, that they are written down, and that they compound instead of being reinvented every January. Games, marketing tools, dashboards, a side project that quietly turns into the real one: same skeleton, same box, same four words at the terminal.
What one box holds
Honestly: far more than you expect, right up until it suddenly does not.
Four gigabytes of memory is the binding constraint. Not disk — sixty gigabytes is an enormous amount of text. Not CPU — two modern cores are genuinely quick, and most web requests are the machine waiting for something rather than thinking. Memory is what you run out of, so it is worth carrying a rough idea of what things cost.
| Thing | Roughly | |
|---|---|---|
| Ubuntu itself | 250–400 MB | Unavoidable. |
| nginx | ~10 MB | Serves every site on the box for that. |
| SQLite | 0 MB | There is no process. It runs inside your program, which is the entire point of it. |
| PostgreSQL | 60–150 MB | One instance is plenty; give each project a database, not a server. |
| A Go service | 8–25 MB | You can run dozens. This is the whole argument for Go on a small box. |
| A Python service | 60–150 MB | Fine. A few of them, also fine. |
| A Node service | 60–200 MB | Plus the runtime. They add up faster than you think. |
| A JVM app | 400 MB+ | One is a real commitment on this machine. |
| Elasticsearch | 1 GB+ | Do not. Postgres full-text search is genuinely good. |
Comfortable
- Dozens of small Go services, each on its own subdomain and port.
- A SQLite file per project, or one Postgres holding every project's data.
- Any number of static sites — they cost essentially nothing.
- Real traffic. Thousands of visitors a day on a well-built site is not close to the limit, and the log says how many you are getting.
- Compiling, testing, and running your agent all day.
Will hurt
- Running a large language model locally. It will not fit; use an API.
- Video transcoding. It works, but it will occupy both cores for a long time.
- A dozen Node processes at once.
npm installon a large project while everything else is running — the memory spike is the problem, not the average.- Anything storing large media on the local disk instead of object storage.
When it feels slow
htop # what is using CPU and memory, live
free -h # how much memory is actually left
df -h # disk — check this before it bites
sudo journalctl -p err -n 50 # the last 50 errors from anything
systemctl --failed # anything that died and did not come back
# The most useful single question you can ask the box:
dmesg -T | grep -i 'killed process'
If you want to know what this box does under real load rather than what it is doing right now, Stress testing is the same machine with somebody leaning on it.
That last one finds the out-of-memory killer, and it is the single most useful question you can ask this machine. When Linux runs out of memory it picks a process and kills it, and the service simply vanishes without logging an error of its own, because it did not get the chance. If something disappeared for no apparent reason and the logs end mid-sentence, that is nearly always what happened.
Who came, without spying on them
The first thing everybody does the day after launching is go and find out whether
anybody came, and the way most people find out is to paste a script from a company
that hands it out free. It is free because your readers are what it costs. Every one
of them is followed to your page and then off it again, on your behalf, and in
return you get a chart. You do not need the chart yet. You need one number, and
you already have it: nginx has been writing a line per request into
/var/log/nginx/access.log since the hour you installed it, and
awk has been on every Unix since 1977. That log is
an analytics database. It just does not have a dashboard, which is what the four
lines below are for.
One line looks like this. This is a real one from this box, and I picked it because it is the funniest line in ten days of log: a stranger trying a hole from 2014 on a server built in 2026, by putting a shell command where the referrer goes. The address is printed whole because it is a rented machine in a hosting company's range, scanning the whole internet, and nobody lives there. A reader's address is a different thing, and this section never prints one.
156.239.229.239 - - [08/Sep/2026:22:44:10 +0000] "GET /cgi-bin/test.cgi HTTP/1.1" 404 19 "() { ignored; }; echo Content-Type: text/html; echo ; /bin/cat /etc/passwd" "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.5 Safari/605.1.61"
Address, time, request, status, bytes, then two quoted fields: where the visitor
claims to have come from and what they claim to be. Claims, both times —
the visitor typed those, and the line above is a visitor typing something else
entirely. The address and the status are the only two facts on the line. Everything
you want to know is a question about those columns, and awk counts by
column:
log=/var/log/nginx/access.log
# How many people today: distinct addresses that loaded the front page.
sudo grep "$(date +%d/%b/%Y)" $log \
| awk '$7 == "/" && $9 == 200 {print $1}' | sort -u | wc -l
# Which pages, this fortnight. Column 7 is the path, column 9 the status.
sudo awk '$9 == 200 && $7 !~ /^\/(static|api)\// {print $7}' $log \
| sort | uniq -c | sort -rn | head
# Where they came from. The referrer is the fourth quoted field.
sudo awk -F'"' '$4 != "-" {print $4}' $log | grep -v mysite.com \
| sort | uniq -c | sort -rn | head
# And what everybody who was not a person was looking for.
sudo awk '$9 == 404 {print $7}' $log | sort | uniq -c | sort -rn | head
The last one is the one to run first, because it tells you what the other three are counting. Here it is on this box, over the ten days of log it had when I ran it:
$ sudo awk '$9 == 404 {print $7}' $log | sort | uniq -c | sort -rn | head -8
145 /auth/callback
135 /account
132 /signup
131 /signin
130 /register
129 /admin
108 /login
108 /api/auth/signin
$ sudo awk '{n++} $9 == 404 {f++} END {printf "%d lines, %d not found (%.0f%%)\n", n, f, 100*f/n}' $log
23770 lines, 14826 not found (62%)
Sixty-two percent of everything that has ever arrived at this site was somebody
trying a door that does not exist. A hundred and one requests for /.env,
sixty-two for /.git/config, three hundred for a WordPress endpoint on a
site with no WordPress, and a whole set for the login pages of a framework I have
never installed. None of it is aimed at you. It is the weather, it arrives at every
address on the internet within hours of the address existing, and the reason
Lock it down came before anything else in this guide is sitting
right there in column nine. The other lesson is who the biggest visitor was:
this box. Three thousand four hundred lines, more than anyone, from the
screenshot tool and the smoke test that check the page after every deploy. Your
top visitor will be you too. The first thing to do with any count is take yourself
out of it.
Then the number you were after. Yesterday, seventy-two distinct addresses loaded this front page, and twelve of them also fetched the stylesheet. A person's browser fetches both; a crawler takes the page and leaves; a person who has been here before has the stylesheet cached and takes the page only. So the honest answer is "somewhere between twelve and seventy-two", and it will stay a range no matter what you install, because the thing you want to count is whether a human was on the other end, and the line does not say. The paid charts do not know either. They round confidently.
/.env, /.s3cfg and /api/proc/self/environ,
which no search engine has ever wanted. That is a scanner wearing a borrowed
coat, because some sites let those names past the door. Count by address. And
behind Cloudflare's orange cloud the address is Cloudflare's too, on every line,
until the snippet in Not getting flooded is installed.
When you do want the chart, get it from the same log rather than from a script in
your page. goaccess is one package, reads the file nginx already
writes, and on this box turns ten days of it into a report in a quarter of a
second:
sudo apt install goaccess
sudo goaccess /var/log/nginx/access.log --log-format=COMBINED # live, in the terminal
sudo goaccess /var/log/nginx/access.log --log-format=COMBINED -o /srv/myapp/data/stats.html
The second form writes one self-contained HTML file, about a megabyte, that you can
open from the mounted drive or serve from a location behind the
same auth_basic password as anything else private. Nothing about a reader
leaves the machine either way. There is no cookie, so there is nothing to put a
banner up about, and the address is in the log for the fortnight Ubuntu's
logrotate keeps by default and then it is gone.
When your app wants its own numbers
Sooner or later a project needs a count the log cannot give it — views per
listing, plays per track, whatever the thing is that your users make — and the
first version everyone writes, me included, is UPDATE things SET views = views
+ 1 in the page handler. It works for a year. Then the page that lists the
things sorts by that column, the report that graphs it walks every row, and you have
built the invite tree again with a smaller tree. Seven of my
own Go projects count things, and after enough versions they all count the same way,
in two halves that never meet in a request:
-
The app writes what happened, in batches. A page view goes into a
queue in memory, and once a second one transaction writes the queue to a
hitstable — a thousand rows in the time one row would take, because the transaction is what costs. The page never waits on its own counter, and if the queue is ever full the view is dropped, which is the correct trade: a count that is a little low is a count; a page that is down is not. -
A timer turns rows into totals. Once an hour, one query groups
the raw rows into an
hits_hourlytable, and everything that has to be fast reads that. It is the same shape as the timer that deletes old sessions, and it is a job, not a page.
sqlite3 /srv/myapp/data/app.db
-- One row per page view, written by the app in batches. No name and no
-- address: a visitor is a salted hash the app draws fresh every day, so a
-- person is one visitor for a day and nobody the day after.
CREATE TABLE hits (
at INTEGER NOT NULL, -- unix seconds
path TEXT NOT NULL,
visitor TEXT NOT NULL
);
CREATE INDEX hits_at ON hits (at);
-- What the charts read. The raw rows go after a week; this is what you keep.
CREATE TABLE hits_hourly (
hour INTEGER NOT NULL, -- unix seconds, on the hour
path TEXT NOT NULL,
views INTEGER NOT NULL,
visitors INTEGER NOT NULL,
PRIMARY KEY (hour, path)
);
.quit
-- /srv/myapp/sql/rollup.sql, run by the timer below. Every finished hour
-- that still has raw rows is recomputed, so a row that arrived late is
-- counted next time, and a night the timer missed catches up by itself.
.timeout 5000
INSERT OR REPLACE INTO hits_hourly (hour, path, views, visitors)
SELECT (at / 3600) * 3600, path, COUNT(*), COUNT(DISTINCT visitor)
FROM hits
WHERE at < (unixepoch() / 3600) * 3600
GROUP BY 1, 2;
DELETE FROM hits WHERE at < unixepoch() - 7 * 86400;
# /etc/systemd/system/myapp-rollup.service
[Unit]
Description=Roll page views up into hourly counts
[Service]
Type=oneshot
User=myapp
ExecStart=/usr/bin/sqlite3 /srv/myapp/data/app.db ".read /srv/myapp/sql/rollup.sql"
ProtectSystem=strict
ReadWritePaths=/srv/myapp/data
# /etc/systemd/system/myapp-rollup.timer
[Unit]
Description=Roll page views up, hourly
[Timer]
OnCalendar=hourly
Persistent=true
[Install]
WantedBy=timers.target
.timeout 5000 is the line people leave out: the app is writing to the
same file, and without it the one time the two collide the rollup fails instead of
waiting five seconds. The visitor column is a hash of the address and the browser
with a random salt the app throws away each day, so the table cannot be turned back
into people, and the same person on Tuesday and Wednesday is two visitors, which is
the right kind of wrong. That rollup ran on this box against a scratch copy before it
was printed here, with a row inserted late and a row nine days old, and did what the
comment says. The hourly table for a site with a hundred pages and a year of
history is under a million rows, which SQLite will sum for a chart
in the time it takes the chart to draw. The raw table never gets big, because it is
not allowed to.
Staying up to date, and what "end of life" really means
Nobody is coming for you personally. That is the point: the scanning is automatic, it is constant, and it does not need a reason to arrive at your address. The only thing that decides whether it finds anything is how long you left the door open.
The least glamorous habit in this guide is the one with the highest return, and it is this: run software that is still getting fixed, and take the fixes. That is most of security for a box like yours. Not a firewall you tuned for a weekend, not a clever nickname for the SSH port — just being current.
Why this got worse, with numbers
It is easy to think of security holes as rare events. They are a firehose, and it is widening. Public vulnerability reports went from about 14,600 in 2017 to 48,000 in 2025, and 2026 is on track to pass 65,000 — roughly 250 new ones every day. Nobody reads that. Nobody is supposed to. It is a machine problem now on both sides.
And the window between "a fix exists" and "somebody is using it against you" has collapsed. Google's incident responders measured the average time from a vulnerability becoming public to being exploited: 63 days in 2018, 32 in 2021, five days in 2023. Their 2026 report puts it at minus seven days — on average, exploitation now begins before the patch exists. Cloudflare watched attacks on one server product start 22 minutes after somebody published proof it was possible. Roughly 29% of the vulnerabilities that end up on the US government's "known exploited" list were already being exploited on or before the day they became public.
The traffic arriving at your box reflects that. More than half of all web traffic is automated, and about 40% of it is hostile bots — a number that has gone up every year for a decade. Verizon's 2026 breach report found that 31% of breaches now start with somebody exploiting a vulnerability, which for the first time in nineteen years of that report beats stolen passwords. None of that is aimed at you. It is aimed at every address there is, and yours is one.
How end of life works
Every piece of software you run has a date after which nobody will fix it, whatever is found in it. Not a date it stops working — that is the trap. It keeps working perfectly, serving pages, right up until the day it is the reason you are reading your logs at two in the morning. The dates are published years in advance and almost nobody looks at them.
| Thing | Fixed for | Then |
|---|---|---|
| Ubuntu LTS | 5 years | Upgrade, or extend |
| Ubuntu Pro free | 10 years | 5 machines free |
| PostgreSQL | 5 years | Major upgrade |
| Go | About a year | Rebuild on a newer one |
| Node.js | 30 months | Move up a version |
| PHP | 2 years, +2 security | Move up a version |
| Python | 5 years | Move up a version |
A new Ubuntu LTS arrives every two years and you upgrade from one to the next, never skipping; PostgreSQL ships small fixes quarterly inside those five years, which are a restart rather than an afternoon; and Go counts its year in releases rather than months — a version is supported until two newer ones exist.
# What is the thing you are running, and when does it stop being fixed?
# endoflife.date tracks nearly everything. jq is `sudo apt install jq`.
curl -s https://endoflife.date/api/v1/products/ubuntu/releases/latest \
| jq -r '.result | "\(.label) — security support ends \(.eolFrom)"'
# Swap ubuntu for postgresql, nodejs, php, python, go, nginx, debian...
What to actually do, in descending order of value
- Turn on automatic security updates and then check they are running. Step 6 of Lock it down switches them on; installing it is not the same as confirming it. The commands below take ten seconds and answer "is this actually happening on my machine".
- Reboot when the kernel says so. An updated kernel is not the running kernel. Ubuntu leaves a file behind to tell you, and it is the single most-ignored file on every server on earth.
- Attach Ubuntu Pro if it is a personal box. It is free for five machines, and the thing it buys you is security support for the community-maintained half of the archive — which is where most interesting software lives, and which the normal five-year promise does not cover.
- Scan your own project's dependencies, which the operating system knows nothing about. Every ecosystem has a one-liner:
govulncheckfor Go,npm audit,pip-audit, and GitHub's Dependabot will open the pull requests for you if the code is there. - Put the database's major version in your calendar, because that one needs planning rather than a restart. Minor versions are a restart. Majors are an afternoon.
apt list --upgradable # what is waiting
ls /var/run/reboot-required # no such file = no reboot needed
systemctl list-timers 'apt-daily*' # when the automatic run happens
journalctl -u unattended-upgrades --since '-7 days'
# The good one: exactly what tonight's automatic run would install.
sudo unattended-upgrade --dry-run -v
# Your own code's dependencies, which apt has never heard of.
# Go, from the project directory — it reports only what you actually call:
go run golang.org/x/vuln/cmd/govulncheck@latest ./...
The fix took one line, because Go can fetch and verify its own compiler; the same shape of fix exists in every ecosystem. The lesson is not about Go. It is that "automatic updates are on" answers a narrower question than the one you think you asked, and the only way to find the gap is to point a scanner at your own project rather than at your operating system.
The upgrade you are putting off
Somewhere there is a project of yours on a version of something that went out of support during a previous presidency. The reason it has not been upgraded is never technical — it is that the upgrade has no visible reward. Nothing looks better afterwards. The button is in the same place. You spend a weekend and the client cannot tell.
Do it anyway, and do it before it is urgent, because the cost only goes one way. Two versions behind is an afternoon. Six versions behind is a rewrite with a different name. And the day it becomes urgent is, by definition, the day you least want a large, uncertain, all-at-once change to your production system: you will be doing the upgrade and the incident at the same time, having lost the option to do them separately.
This is also the single best use of an agent I have found, and the reason is that the work is enormous, mechanical, boring, well-documented, and continuously testable. Point it at one version step at a time — not five — with the tests running after each, and a commit at every green point so you can walk back. It will do in an evening the thing you have been not-doing for two years.
Audit this box for anything out of support. Check the Ubuntu release and its
end-of-life date, every service I am running and its version against upstream,
whether unattended-upgrades is actually installing things, whether a reboot is
pending, and my project's own dependencies with the right scanner for its
language. Then give me one table: what is current, what is behind, what is past
end of life, and what you would do first. Do not change anything yet.
It is down. What now?
Thirty sections on building the thing. This is the one about the night it stops working, which is the moment most people quietly give up and delete the box.
Your car will not start. You are certain it is the battery, because you left the lights on last night and the lights are the thing you touched. You buy a battery. It is not the battery. The mechanic finds a dead fuel pump in about four minutes, and not because he knows your car better than you do — he has never seen it. He is simply not the man who left the lights on, so he starts at the top of the same list he always starts at, and that list does not care what you did last night.
That is the whole trick, and it is the only thing in this section worth memorising. When your site goes down you will start debugging at the last thing you changed, because it is the freshest thing in your head and because some part of you already feels responsible for it. It is almost never there. Work the ladder instead, in order, top to bottom, one command per rung, and let the machine tell you where you are standing rather than guessing from the smell of smoke.
| The question | Why it is on this rung | |
|---|---|---|
| 1 | Is it me, or is it everyone? | Load the site on your phone with wi-fi switched off. Costs five seconds and settles the biggest question there is. |
| 2 | Is the machine even alive? | ping, then ssh. If you cannot get in, nothing below this rung is knowable. |
| 3 | Is your service running? | systemctl status, then the last fifty lines of its log. Most outages end here. |
| 4 | Is nginx up, and is its config valid? | An invalid config does not take effect, so the site keeps serving the old one until something reloads it. Then it stops. |
| 5 | Is the certificate still good? | Certificates expire on a date, not on a trend. Nothing degrades first. |
| 6 | Is it DNS? | It is sometimes DNS. Ask what the world sees, not what your laptop remembers. |
| 7 | Is the disk full? | A full disk breaks things that look completely unrelated to disk, which is why it is worth reaching the bottom of the ladder. |
Rung one is first because your own browser is a liar. It has a cache, your operating system has a DNS cache, and your network may have opinions of its own. A site that is down for everybody is an outage. A site that is down only for you is a Tuesday, and the two want completely different afternoons from you.
# 1. is it me, or is it everyone — ask from somewhere that is not this machine
curl -sI https://example.com | head -1
# 2. is the machine alive
ping -c 3 example.com
ssh you@example.com # if it refuses you, run it again with -v
# 3. is the service running, and what did it say on the way out
systemctl status myapp --no-pager
sudo journalctl -u myapp -n 50 --no-pager
# 4. is nginx up, and does its config actually parse
sudo nginx -t
systemctl status nginx --no-pager
# 5. is the certificate still valid, and until when
sudo certbot certificates
# 6. what does the world think your domain points at
dig +short example.com
# 7. is the disk full
df -h
Rung three is where most nights end, and it is worth knowing the one failure that leaves no note. If the service is simply gone and its log stops mid-sentence with no error, that is usually the out-of-memory killer, which does not give a process the chance to complain on its way out. What one box holds has the command for that and the reasoning behind it.
The deploy that broke it, taken back in one command
Rung three has a favourite ending, and it is the one night the instinct is right: the service died in the minute after you shipped something. The log says it started, printed one line and stopped, and the freshest thing in your head really is the cause this time. What goes wrong is the next minute. The deploy from The shape builds over the only binary you have, so "put the old one back" means finding the commit, rebuilding it and hoping it was that one, at the exact moment you are least able to do any of those calmly.
Two changes make it a deploy you can take back. Every build goes into its own
dated folder, and the service runs whatever a link called current
points at. Deploying moves the link forward. Rolling back moves it to the folder
before, and the folder you just left is still there, untouched, for when you
have found the bug.
# In /etc/systemd/system/myapp.service, the one line that changes
ExecStart=/srv/myapp/releases/current/server -addr 127.0.0.1:8090
# Every build lands in its own dated folder; `current` points at the live one.
releases := "/srv/myapp/releases"
stamp := `date -u +%Y%m%d-%H%M%S`
# Build into a new release, point current at it, restart, and confirm it came back
deploy:
go build -o {{releases}}/{{stamp}}/server ./cmd/server
ln -sfn {{stamp}} {{releases}}/current.next
mv -T {{releases}}/current.next {{releases}}/current
sudo systemctl restart {{service}}
@sleep 1
@curl -fsS http://127.0.0.1:{{port}}/healthz && echo " — up, on {{stamp}}"
@ls -1d {{releases}}/2*/ | sort | head -n -5 | xargs -r rm -r
# Point current at the release before this one, restart, and confirm
rollback:
#!/usr/bin/env bash
set -euo pipefail
cd {{releases}}
now=$(readlink current)
prev=$(find . -mindepth 2 -maxdepth 2 -name server -printf '%h\n' | sed 's|^\./||' | sort \
| awk -v now="$now" '$0 == now { print p; exit } { p = $0 }')
test -n "$prev" || { echo "nothing older than $now to go back to"; exit 1; }
ln -sfn "$prev" current.next && mv -T current.next current
sudo systemctl restart {{service}}
sleep 1
curl -fsS http://127.0.0.1:{{port}}/healthz && echo " — back on $prev ($now is still there)"
# Every release, newest first, with an arrow on the live one
releases:
@ls -1t {{releases}} | grep -v current | sed "s|^$(readlink {{releases}}/current)$|& ← current|"
sudo systemctl daemon-reload
just deploy # the first release; the old bin/server can go once this says up
Three details carry the weight. mv -T swaps the link in one step, so
there is never a moment when current points nowhere; a plain
ln -sfn onto an existing link is a delete and then a create, which
would not matter for one binary behind a restart, but the same trick serves a
folder of PHP that nginx is reading right now, where it does.
rollback starts with #!, which makes it one script
rather than a shell per line, so prev survives to the line that
uses it. And it only considers folders that hold a binary, because the first
version of it on this box rolled back, cheerfully, into an empty folder that a
failed build had left behind, and the service went into a restart loop.
deploy no longer makes the folder itself for the same reason:
go build creates it, and only when the build succeeds. Five releases
are kept, which is five bad decisions deep.
On this box, one deploy after another, then back:
$ just rollback
ok v7 — back on 20260916-030557 (20260916-030559 is still there)
$ just releases
20260916-030559
20260916-030557 ← current
20260916-030555
20260916-030552
20260916-030550
When several things break at once
One thing the ladder will not tell you, and it is worth having in your head before you need it: when a handful of things fail in the same second, they have almost certainly not failed independently. Stop working down the list and go looking for the one thing they share.
Half a dozen long-running jobs on this machine all stopped within a second of each other one morning, which is a very convincing impression of a server dying. The server was fine. Load was nothing, memory was nothing, and the logs held no errors because there were none to hold. What had happened was that my laptop's connection dropped, and every one of those jobs was a child of the same SSH session, so they went the way a branch goes when you cut through the trunk. The fault was not on the server at all, and no rung of the ladder was ever going to find it, because the ladder is bolted to the server.
The fix is one command you should learn before the wire drops rather than after.
tmux keeps a session running on the server independently of the
connection you started it from, so a dropped link detaches you instead of killing
your work. Reconnect, reattach, and you are back in the same session with everything
still running — the long build, the migration, the agent halfway through a task.
screen does the same job and is often already installed. Use either. Use
one of them for anything you would hate to lose.
sudo apt install -y tmux
tmux new -s work # start a named session and work in it
# ctrl-b then d detach and leave it running
tmux ls # what is still running
tmux attach -t work # get back in, from anywhere
Permission denied (publickey) and I believed every
word of it. I got as far as drafting instructions for getting into the hosting
company's web console to add a key by hand — for a machine I was already
logged into at that moment, from that terminal. The key was fine. The server was
fine. ssh had looked in the five default filenames it always looks in,
found nothing there, offered no key at all, and the server replied with the only
sentence available to it when nobody offers it anything. One ssh -v
shows you that in about a second, because it prints every key it tries before it
gives up. An error names the step that quit. It does not name the step that was
wrong, and those are very rarely the same step.
Finding out before anybody tells you
Everything above assumes you know it is down. The way most people actually find out is a message from somebody who noticed first, which is the most expensive monitoring system ever devised, and it only ever reports at the worst possible moment. Replace it once, for free, with a check that runs from somewhere else. It has to be somewhere else. A monitor running on the same box goes down with the box, and then reports absolutely nothing, perfectly.
There are two shapes of check, and a real project wants both.
- Something that knocks: UptimeRobot
-
It requests a URL every few minutes and emails you when the answer stops being a
200. The free plan is 50 monitors at five-minute intervals, which it describes as
for hobby and non-profit projects. Point one at your
/healthzand one at your home page, and rung one of the ladder is answered before you have finished reading the message. - Something that waits to be told: Healthchecks.io
- The opposite direction. Your server checks in, and it emails you when a check-in does not arrive. This is the only kind of monitor that can catch a timer that never ran, because a job that did not run produces no error for anything else to find. The free plan watches 20 jobs.
Wiring a timer to it is one line in the job's .service. systemd only runs
ExecStartPost if the job itself succeeded, so a failure simply never
checks in, and the silence is the alarm:
# Add under [Service] in myapp-prune.service. The URL is the one Healthchecks.io gives you.
ExecStartPost=/usr/bin/curl -fsS -m 10 --retry 3 -o /dev/null https://hc-ping.com/your-check-uuid
Then make the health endpoint worth knocking on. A handler that
returns 200 OK proves the process is running and not one thing more. It
will say ok, cheerfully, while the database file is unreadable and the disk
is full, because nobody asked it about either. The question to answer is the one
your visitors are asking — can it actually do the thing? — and
for nearly every app that means does the database answer. The exception
in the same breath: it should not check every outside service you depend
on, or the day your payment provider has an outage your monitor wakes you up about
somebody else's night.
Make GET /healthz tell the truth. Run one cheap, real query against the database with a two-second timeout. Return 200 with the version if it works, or 503 with a one-line reason if it does not. Do not call any third-party service from it, do not require a login, and put nothing secret in the response. Add a test for the 503 case.
In the interest of full disclosure, the /healthz on this very site is the
lazy kind: it says ok if the process answers. It gets away with that for one
reason only, which is that there is no database behind it. The day there is one, it
becomes a liar.
README.md, where you will find it from your
phone at some deeply unreasonable hour. You will not compose a calm diagnostic
procedure while the site is down and someone is asking you when it will be back;
nobody does. The version of you who is not panicking has to leave a note for the
version who is. That is most of what operations actually is.
The day you do something really stupid
It is coming. Not as a possibility — as an appointment. And handled properly it is worth more to you than the six months before it, which is not a consolation prize, it is the actual mechanism by which people get good at this.
The previous section is for when the box breaks. This one is for when you break it: the command that was right for the other terminal, the migration that ran against the wrong database, the cleanup script that was cleaner than intended. Everybody has one. The people who look unshakeable are not the people it has not happened to — they are the people it has already happened to, twice, and who therefore have somewhere to start.
The first ten minutes
The order matters more than the actions, and the order is not obvious, because every instinct you have in the first ten seconds is wrong. In particular the strongest one: undo it.
- Stop. Hands off the keyboard. If the thing is actively making it worse — a job still running, a script still looping, a service still writing — stop that, and nothing else. More damage has been done in the sixty seconds after a mistake than in the mistake.
- Do not un-do anything yet. Undoing is a second change, made by somebody upset, on a system whose state you no longer know. It is how a bad ten minutes becomes a bad week.
- Take a copy of the wreckage before you touch it. A snapshot, a
tar, a dump — whatever is one command. Every recovery option below gets harder once you start overwriting things, and several become impossible. - Write down what you actually did, in the order you did it, while you still remember. Ten minutes later you will be reconstructing it, and you will be reconstructing it wrong, in a direction that flatters you.
- Tell somebody. Now. Before you know how bad it is, before you have a plan, before the part of your brain that wants to quietly fix it first gets a vote. The next subsection is entirely about this, because it is the step everybody skips.
- Then work out which kind of problem you have, because there are only two and they have completely different playbooks.
Own it immediately, upward and sideways
Here is the advice that sounds like a trap and is not: tell your colleagues and tell your boss, straight away, in plain words, before you have a fix. "I ran X against production, I think I have destroyed Y, I am looking at what we can recover and I will have an answer in twenty minutes."
It feels like handing somebody a weapon. It is the opposite, and for a reason that has nothing to do with virtue: the value of everything you own decays by the minute. Backups rotate off the end. Logs roll. Replicas resynchronise and overwrite the good copy with the bad one. Somebody else quietly builds on top of the broken data. Every minute you spend hiding it is a minute of options being deleted, and they are options you may need a second person's access to use anyway.
There is a second reason, which every experienced person knows and few say out loud: you will be forgiven for the error and you will not be forgiven for the gap. An hour between the mistake and the telling is a question you will be asked about for the rest of your time there, in a tone the mistake itself never earned. People are astonishingly relaxed about a colleague who broke something and said so instantly. They are never relaxed about finding out later.
The best-documented example in the industry is a company that did this in public. In 2017 a GitLab engineer, working late on a struggling database, ran a destructive command against the primary instead of the secondary and realised within seconds. About 300 GB went. Then they discovered that of five separate backup and replication mechanisms, not one was working — the dumps were failing silently because of a version mismatch, the snapshots were not enabled. What saved them was a six-hour-old manual snapshot somebody had taken by hand. They lost six hours of data for thousands of projects.
And then they did something remarkable: they kept a public document updated live through the recovery, streamed the whole thing on YouTube to a peak of five thousand viewers, and published a postmortem naming every failure of their own systems. Nobody remembers the engineer. Everybody who works in this field remembers that postmortem, and it did GitLab more good than a decade of marketing. The mistake was ordinary. The handling was not.
Which of the two problems do you have?
Every disaster sorts into recoverable and not, and the single most valuable skill in the whole evening is telling them apart in the first few minutes, because they demand opposite behaviour. Recoverable means: stop, think, restore carefully, verify. Unrecoverable means: stop trying to fix it, start containing it, and start telling people — because when there is nothing to recover, everything that is left is about honesty and time.
| What you did | Reach for | Costs you |
|---|---|---|
| Lost code, bad rebase | git reflog | Minutes |
| Deleted a repo or branch | GitHub, 90 days | Minutes |
| Deleted a file still open | /proc/PID/fd | Be quick |
| Wrecked the data | Last night's backup | Hours of writes |
| … and hours matter | Point-in-time recovery | Set it up first |
| Destroyed the box | New VPS, repo, backup | Under an hour |
| Published a secret | Rotate it | Assume it is used |
| Leaked personal data | Contain, then disclose | Legal clocks start |
Look at how much of that top half exists because somebody already made your mistake. The reflog is ninety days of everywhere your branches have been, and thirty more for commits nothing points at any more. GitHub's ninety-day window for a repository somebody deleted in a temper. The fact that deleting a file on Linux does not remove it while a program still holds it open — so a log, database or config you just destroyed is often still readable through the running process's own file handles. None of that is luck. Each one is a guardrail built by a person who had a very bad evening and then made sure it hurt less for the next person. That is the tradition you are joining, and it is why the first move is to stop rather than thrash: the recovery path frequently exists, and thrashing is what destroys it.
"Everything is gone" is a 45-minute problem
If you have followed this guide, the worst realistic case is smaller than it feels at two in the morning. Your code is on GitHub. Your data was backed up last night, off the box, encrypted. Your configuration is in the repo beside the code. So the recovery is: rent another box, run the same setup, pull the repo, restore the backup, point the domain at the new address. Most of that hour is waiting for things to download.
Two things will be outside your control, and both are worth knowing before you meet them:
- The old box coming back to life
- A machine you thought was dead, rebooting and resuming work — sending email, running scheduled jobs, writing to the database you just restored. Before you cut over, fence it: power it off at the provider, and revoke the keys and tokens it holds. The trading disaster above is exactly this shape — one server nobody accounted for, doing what it was last told.
- DNS not moving as fast as you did
- You change the address and some people still arrive at the old one, because a name has a cache lifetime and not every resolver on the internet honours it — a measured few per cent simply keep answers longer than they were told to. If you know a move is coming, lower that lifetime a day in advance. If your domain sits behind Cloudflare's proxy, this problem mostly disappears: the address the world has cached is Cloudflare's, and you are only changing where Cloudflare looks.
Afterwards: fix the tool, not the person
The instinct after a bad evening is to promise to be more careful. It does not work. I have been promising myself that for twenty years, in between truncating the wrong table, and carefulness is a resource that runs out at exactly the moment you need it — late, tired, under pressure, which is when every one of these happens.
What works is changing the shape of the thing so the mistake becomes harder to make. When Amazon took out a large part of the internet in 2017, it was because an engineer following an established procedure mistyped one input and removed more servers than intended. Their published response is a model: they did not promise better typing. They changed the tool so it removes capacity slowly, and refuses outright to take a system below its minimum. The same instinct is why the recipes in this guide's own justfile refuse to run on the main branch, why live machines get a different terminal colour, and why every destructive command in the backup section makes a copy first.
Which brings it back round to why this guide keeps insisting on a box that is not your laptop. A forty-five-dollar machine where a catastrophe costs you an afternoon is not just cheap hosting. It is the only place most people ever get to make a genuinely stupid mistake at full speed, find out what the recovery actually feels like, and discover that the floor is a lot closer than it looked. Everyone who seems calm during an outage is drawing on a memory of a night like that. Go and get yours early, on something that does not matter.
Caching, and the box that keeps up
Advanced, and optional right up until the day it is not. The one technique that lets a forty-five-dollar box do the work of a fleet — provided the rest of the project is shaped to let it.
The warning goes first, because it is the part people skip: do not build a cache on day one. A cache added before you have measured anything is a second copy of your data, with its own bugs, standing guard over a database that was never slow. Plenty of people go an entire career without ever needing to cache a query, and the scar at the end of the next section is me doing it anyway.
But for some projects the day does come, and when it comes it arrives all at once. Then caching is the difference between a box that keeps up and a box that falls over, and it costs a small fraction of what a second machine costs — which is why the checklist in More than one box puts it before adding one.
The whole idea, in one number
Most of the traffic to any site is the same handful of questions, asked over and over. The front page. The leaderboard. The total on the dashboard. If five thousand people load the same page in the same five seconds, the database does not need to answer five thousand times. It needs to answer once, and the other 4,999 get a copy.
- Five thousand visitors ask the same question in the same second — the front page, the leaderboard, the report everybody opens on a Monday.
- The cache already has the answer, so 4,999 of them get it without the database ever hearing about it.
- The first request, or the first after the copy expired, is a miss, and the database does the slow work exactly once.
- The answer is stored with an expiry, so even a copy nobody remembers to throw away heals itself.
- And when somebody saves a change, the old copy is thrown away there and then, not five minutes later. That is the hard part, and very nearly the only part.
Where to put it, cheapest first
There are five places a copy can live, and they are in this order for a reason. Each one costs more to run and more to get wrong than the one above it, so start at the top and stop at the first one that is enough.
- The browser: files that never change
-
Give static files a name that changes when their contents do —
app.3f9a1c.js, or a?v=on the end — and tell the browser to keep them for a year. A returning visitor then downloads nothing but the page itself. It is an item on the launch checklist, and it is free. - The edge: a CDN holding whole pages
- For visitors who are not logged in, an entire page can be served by Cloudflare without your server being asked at all. Two catches get everybody. Cloudflare does not cache HTML at all by default, so you have to add a cache rule that says to. And it will not cache any response that sets a cookie, so a session cookie sent with every page quietly switches the whole thing off and tells you nothing.
- Precomputed: the expensive answer, built on a schedule
- A report that takes twenty seconds to build takes no time at all when a timer builds it every ten minutes into a table and every visitor reads the table. This is the one that fixes the worst pages on a site, and it is the one that would have saved me the story below.
- In memory: a map with an expiry, inside your app
- A handful of lines in Go: a map, a lock and a timestamp. It is gone when the process restarts, which does not matter, because it was only ever a copy. This is where most of the speed on a single box actually comes from.
- Shared: Redis or memcached
- Only when more than one process, or more than one machine, needs to see the same copy. On one box running one Go service you do not need it — it is another service to run and watch, holding a copy of data you already have.
The first two are nothing more than a header on the response:
# Files whose name changes when their content does: keep for a year
Cache-Control: public, max-age=31536000, immutable
# A public page: browsers always check again, a CDN may hold it for five
# minutes, and may serve the old copy for a day while it fetches a new one
Cache-Control: public, max-age=0, s-maxage=300, stale-while-revalidate=86400
# Anything about one logged-in person: no shared cache, ever
Cache-Control: private, no-store
The in-memory kind is the one to hand to the agent, and the prompt is mostly about the two things that go wrong: a crowd arriving while the copy is being rebuilt, and a copy that outlives the data it was copied from.
Add an in-memory cache to GET /leaderboard. Keep the rendered result for 30 seconds. If many requests arrive while it is being rebuilt, build it once and have the rest wait for that result, rather than all of them querying the database at the same moment. Throw the cached copy away as soon as a score is saved. Never cache anything that belongs to one user. Add tests for the expiry, for the invalidation, and for twenty simultaneous requests causing exactly one query.
The right backend for a crowd
Caching decides how much work each request costs. The other half is how many requests can be inside the building at the same time, and here the choice of language matters rather more than the people who say it does not would like. A classic PHP setup gives each request its own worker process, and when every worker is busy the next visitor queues. Node runs everything on one thread, so one slow request makes everyone wait for it. Go gives every connection its own goroutine at a few kilobytes apiece, so thousands of open connections cost memory you have, not processes you do not. That is the argument for Go in What one box holds, pointed at a crowd.
A live screen is the same idea in the other direction: one event, many watchers. The dashboard on this site works that way — one sampler, fanned out to every browser that is looking, so a hundred viewers cost barely more than one — and the rule that makes it safe is that a slow screen must never slow down the thing it is watching.
ScanQRTo.WIN is a QR-code contest system I built in Go: a code goes up on a screen, you scan it, and your phone tells you whether you won, while an admin panel shows every scan arrive live. A client paid to run it at a large conference, on a RackNerd VPS with three cores and 3.3 GB of memory, with SQLite as the database. The moment that matters here is the moment a code goes up on the big screen and a room full of phones hits the server at once, while the panel streams every one of them. At the live game on 11 October 2025 that was about three hundred people: 320 scans from 311 different addresses, 225 of them inside a single minute, nearly all over their own carriers' mobile data. It kept up, and the scan never waited on the panel.
The load tests beforehand had said it was ready, and they were right about speed. They were testing whether it was fast, not whether it was right, and a test script from one machine does none of the things a room full of people does:
- People share networks. Phones on the same WiFi share one IP address, so the duplicate-player check that used addresses started telling brand-new players they had already played — and showing them somebody else's win code. It was found and fixed four days before the crowd arrived. Never identify a person by their IP address.
- A crowd carries the same phones. 260 of those players were on an iPhone running Safari, and iPhones look so alike to a browser fingerprint that different people came out as the same device: 311 addresses, but only 301 fingerprints.
- A game capped at 50 winners gave out 56. Every scan asked "how many winners so far?", saw fewer than 50, and added one. When dozens of those land together, they all read the same number before any of them has written.
That last one is the most common concurrency bug there is, and the fix is a single statement: make the check and the change the same operation, so the database does both at once or neither.
-- Wrong: read the count, decide in your code, then write. Under a burst,
-- a whole crowd of requests reads "49" before any of them writes.
SELECT winner_count FROM games WHERE id = 1;
-- Right: one statement. The database checks and adds together.
UPDATE games
SET winner_count = winner_count + 1
WHERE id = 1 AND winner_count < max_winners;
-- One row changed: this scan won. No rows changed: the cap was reached.
Both versions were run on this box on 14 September 2026, against a cap of 50, with three hundred players arriving at once. The first gave away 79 prizes. The second gave away exactly 50. Nothing about it is faster, and nothing about it is cached — it is simply the difference between asking the database a question and asking it to keep a promise.
Eleven months later, somebody doubted the numbers, so we pushed it until it broke. On 14 September 2026, on the same three-core box with a Xeon from 2013, now sharing it with about fifteen other live services, every fake phone was released at the same instant. The October tests had only ever sent twelve requests at a time, so this one was far harsher:
- Through the public site, 1,000 phones at once all got an answer, the median one in 11.5 seconds.
- At 1,500, 180 of them failed, and every single failure was nginx, logging that 768
worker_connectionswere not enough. The app answered everything that reached it. - Straight to the app, 6,000 phones at once all got an answer, in 594 MB of memory. It processes sixty to eighty scans a second whatever you throw at it, so a bigger crowd becomes a longer wait, never an error. The test stopped there to protect the other services, not because anything broke.
- The real crowd, for scale, peaked at 225 scans in a minute. That is under four a second.
So the first thing to fall over was not Go, not SQLite and not the memory. It was one number in a config file that Ubuntu ships as 768, which also sets the ceiling for every other site on the box — when it ran out, the neighbours returned errors for a couple of seconds too, which is what testing in production costs. A proxied visitor uses two connections, one from the phone and one to your app, so 768 is nearer four hundred people than eight hundred. It is worth checking before your own big day:
# Ubuntu ships 768. Every proxied visitor costs two: phone to nginx, nginx to app.
grep -n 'worker_connections\|worker_rlimit_nofile' /etc/nginx/nginx.conf
# /etc/nginx/nginx.conf. Every connection is an open file, so raise the file limit
# with it. Edit the events block that is already there; a second one will not load.
worker_rlimit_nofile 8192;
events {
worker_connections 4096;
}
sudo nginx -t && sudo systemctl reload nginx
The report that dropped production twice
On 23 October 2007 police shut down OiNK, the invite-only music-piracy site, and
several sites raced to replace it. I was the launch developer on one of them,
waffles.fm, under the name Kcaj. It opened on
31 October, eight days after the raid, and made the file-sharing news the same week. I found the job in a
random IRC channel while looking for work, and I did not know what the project was
when I said yes. Somebody had already paid for the servers and the domain and a team
was well into building it, so I went at it sideways: I took an open-source
tracker-and-forum codebase, modified it for many hours without sleep, and launched
that ahead of everybody.
It went from ten members to a hundred thousand in a couple of months, and it did not go smoothly. My code broke long before the hardware did — somewhere past ten thousand members, pages I had written started taking the server down, and bigger machines could not keep pace with bad queries. Getting it the rest of the way took caching, cleaned-up queries and help from people far above my pay grade.
Then I built the invite tree. Who invited whom, whom they invited, all the way down, with every branch's donations added up, so we could finally see who had been responsible for the most donations in the site's whole history. I ran it against my own account, and it dropped production. We got the site back online, I revised the query, ran it again, and dropped production again.
Stress testing, before the crowd does it for you
Advanced, optional, and a waste of an afternoon for most projects — right up until the day somebody links you somewhere busy. The useful part is not the number you get. It is finding out which thing breaks first, because it is never the thing you would have bet on.
If you are not expecting thousands of people at once, you can skip this section entirely and lose nothing. Almost nobody's first project needs it. A cheap box handles more than people believe, the guide has already said so in What one box holds, and worrying about throughput before you have a single user is the same disease as buying a bigger server before you have a website on the small one.
But it costs about an hour to find out what your box actually does, and the answer is worth having before the day you need it rather than during. So here is this website, on the Go binary and the $44.88 box you have been reading about, pushed until things started falling over — and then the same box after an afternoon of fixing what the test found.
This box, pushed until it broke
The rules first, because a benchmark with no rules is marketing. A copy of this
site's own binary was started as a second service with the same sandbox,
memory ceiling and processor share as the live one, with a second
nginx in front of it running this site's real configuration and
real certificate. The live site kept serving visitors the whole time and was never
the thing under test. The load generator ran on the same two cores as the server it
was hitting, which means every figure below is a floor: a second machine
pushing would have got more. It is all one script, just stress, and it
writes what it measured into a file with a date on it.
Requests are 50 at a time, through nginx and TLS, as a browser arrives; "before" is the same test against the same code before the afternoon's fixes.
| Measured | Now | Before |
|---|---|---|
| The page | 7,109/s | 127/s |
| Slowest 1 in 100 | 42 ms | 976 ms |
| A script or stylesheet | 19,208/s | 1,374/s |
| Live dashboards | 2,000 | froze past 3,000 |
| 6,000 asking at once | all 2,000 fed | 10% of frames |
| Copies of this app | 40, in 568 MB | — |
| 1,000 connections | 2,411 errors | the same |
The last row is the one that has not been fixed, and it is the same wall the
conference game hit in the section above:
worker_connections ships at 768, a proxied visitor costs two of them,
and the errors start at about four hundred people in the building at the same
instant. It is one number in one file, it is in
Caching, and the box that keeps up with the command to change
it, and it is worth doing before your big day rather than during it.
What broke first, in the order it broke
Not one of these was the code anybody would have looked at. In order:
- Compression, on every single request. This page is 640 KB of HTML, and nginx was gzipping all of it, from scratch, for every visitor. One processor core, most of it, for about a hundred and twenty pages a second — while the app that generated the page sat at a tenth of its allowance, bored. The page is the same for everybody for two seconds at a time, so now it is built once, squeezed once, and handed out as-is. Same bytes, 56 times as many of them.
- And again, for the stylesheets and the scripts. Same mistake, different files, and this one decided how many first-time visitors a second the box could take, since a returning reader has them all in their browser already.
- Every open dashboard doing the same maths for itself. One reading of the machine, ten thousand watchers, ten thousand identical conversions to text, every two seconds. Now it is converted once and the same bytes go to everybody.
- The log eating itself. The app wrote a line per request. Under load that filled the journal's allowance — 10,000 messages every 30 seconds per service — in under half a second, and the system said so: "Suppressed 1257132 messages". Every line after that was thrown away, which would include the error you were about to go looking for. The web server's own log already had every request. The app logs failures and slow requests now, and nothing else.
- The memory limit that was not limiting anything. See below. This is the one worth your attention.
MemoryMax=192M, which I believed meant
"if this thing ever runs away, kill it before it takes the machine with it". So I
opened twenty thousand live dashboards against a copy of it and went to look at how
gracefully it died. It did not die. It sat there at 7 MB of memory and 545
MB of swap — comfortably "under" its 192 MB limit, by the only
measure the limit counts — while the swap on the whole box filled up and
everything else on it got slower. The limit caps memory. Swap is not memory, so
swap is not capped, so a service under its ceiling can quietly page half a gigabyte
onto the disk and call it obeying orders.
One line fixes it:
MemorySwapMax=0. Proved rather than assumed —
a little program that grabs memory as fast as it can, run under the exact unit this
guide gives you: without that line it took 960 MB and finished normally,
and with it the kernel stopped it dead at the limit, which is what the limit was
always supposed to mean. The unit in part five has the line in
it now. If you copied that block before today, add it.
With swap honestly switched off, the real shape appeared: about 40 KB of memory per live dashboard, 3,000 of them comfortable, and past that the service was throttled by the kernel into a coma — alive, answering its health check, delivering eight frames a second to six thousand people who each expected one every two seconds. Which is worse than dying, because nothing alerts on it. The fix was not a bigger box. It was to admit a limit and hold it: 2,000 live viewers, and everybody after that gets a polite refusal and keeps the numbers the page was built with. Turning somebody away in one millisecond is a feature. Serving six thousand people badly is not.
The bandwidth runs out first
Here is the part that surprised me most, and it has nothing to do with code. This box can now build and compress this page 7,109 times a second. At 207 KB each, that is about twelve gigabits a second of HTML — and the network port it would have to go out of measured 390 megabits going up and 730 coming down. The processor is roughly thirty times faster than the wire it is attached to.
So the honest limit is not "requests a second". It is bytes a month. Reading this page from top to bottom — every picture, every little video loop — costs about 4.7 MB across 71 requests, measured in a real browser; the landing screen alone is about 1.8 MB. The plan includes 6 TB a month. That is 1,276,595 full reads, or about forty thousand a day, every day, for $3.74 a month. Spread evenly it is a steady 18.5 megabits, less than a fifth of the pipe. Flat out, the port would spend the entire monthly allowance in about 34 hours.
| Monthly traffic | Spread evenly, that is | Full reads of a page like this |
|---|---|---|
| 1 TB | 3.1 Mbit/s | ~213,000 |
| 6 TB (this plan) | 18.5 Mbit/s | ~1,276,595 |
| 20 TB | 61.7 Mbit/s | ~4,255,000 |
| A 1 Gbit port, flat out, all month | 1,000 Mbit/s | 324 TB, which no cheap plan includes |
Two things follow from that table. The first is that media, not code, is what spends a hosting allowance: strip the pictures and the loops off this page and the same 6 TB would carry about thirty-two million reads instead of one and a quarter million. The second is the cheapest optimisation on this entire page — put the big files on somebody else's network. Cloudflare in front of the site serves the pictures from its own machines for free, and your allowance stops being the thing you think about. That is in the caching section, and it is five minutes of work.
How many apps, and how many people
The question this guide is actually about is not "how fast is one app", it is "how much can I get away with on one box". So: 40 copies of this application were started at once, each in its own sandbox with its own memory ceiling, and then all forty were loaded simultaneously. All forty answered. They used 568 MB between them — about 14 MB each — and between them they served 17,831 requests a second.
That last number is the important one, and it is a warning rather than a boast. Forty copies did not do forty times the work. They did roughly the same total work, divided forty ways, because they share the same two cores. Memory decides how many things you can run; the processor decides how much they can all do together. With 2900 MB free on this box and a deliberate 30% left spare, the arithmetic says about 145 small Go services would fit. I would stop long before that, because the interesting number is not how many fit, it is how many are busy at the same moment — and on a box full of side projects, that is almost never more than one.
Turning that into people takes an assumption, so here is mine, stated plainly: a person reading a page makes one request and then goes quiet for a while. If an active user causes a request every ten seconds, then a hundred requests a second is a thousand people using the thing at once, and this box does eighty times that on a cached page. Real applications are far slower than a cached page — a database-heavy Ruby on Rails application, measured by VPSBenchmarks on a two-core box much like this one, topped out at 18 requests a second before latency ran away, which on the same assumption is still around a hundred and eighty people at once, and many thousands of visitors a day. The shape of the answer, for almost everybody reading this, is: more than you will get.
Other languages, roughly, and why it matters less than you think
Only Go was measured here, because only Go is what this site is written in, and a benchmark of a language I did not test is a rumour with a decimal point. What follows is other people's numbers, from the last round of the TechEmpower benchmarks — the longest-running public comparison of web frameworks, which shut down in March 2026, so this is the final scoreboard there will be. The test shown is their "Fortunes": read rows from a real database, sort them, render a template, escape the output. It is the closest thing they run to an actual web page.
Each bar is that framework's Fortunes result as a multiple of Go's, on a logarithmic scale, because a straight one would make everything below Rust a smudge. Their hardware is a 56-core server, not a $44.88 box — the ratios travel, the absolute numbers do not.
Source: TechEmpower Framework Benchmarks, Round 23 (the final round), Fortunes test, read 16 September 2026 — 56-core Xeon Gold 6330, 64 GB, 40 Gb Ethernet. Go's standard library with a Postgres driver is the baseline at 1.0×. Two well-known entries are missing because they failed to run in that round, which is its own kind of data.
Read the bottom of that chart before the top. The gap between the fastest framework and the slowest is about seventy times — and the gap between raw PHP and PHP with a big framework on top of it is roughly nine times, in the same language. The language you pick matters less than what you pile on top of it, and both matter less than whether you do the same work twice. Fixing compression on this site, in one language, on one afternoon, was worth 56× — more than the entire distance between Go and Rust. Nobody's project was ever saved by a rewrite that could have been a cache.
Memory is the other half, and it is the half that decides how many things fit on a $44.88 box. Using the rough per-service sizes from What one box holds, and leaving 30% of this machine spare:
| Stack | Each, roughly | That fit here |
|---|---|---|
| Go, standard library | 14 MB measured | ~145 |
| Node, Fastify | 130 MB | ~15 |
| Java, Spring | 450 MB | ~4 |
| Python, FastAPI | 105 MB | ~19 |
A dozen Python or Node services on one cheap box is perfectly normal. Four JVM applications is a decision you make on purpose. And all of them, whatever the language, are sharing two cores — so the second question is always how many are busy at the same time, which for side projects is almost always "one".
Testing yours, in about an hour
You need one tool and one rule. The tool is wrk, which is in Ubuntu's
package list and does one thing well. The rule is that you never point it at
the thing people are using if you can point it at a copy instead.
sudo apt install wrk
# 2 threads, 50 connections, 30 seconds, and report the slow tail honestly.
# Ask for compression, because every real browser does.
wrk -t2 -c50 -d30s --latency -H 'Accept-Encoding: gzip' https://yourdomain.com/
Read three things in what it prints, and ignore the rest. Requests/sec is how much. 99% in the latency block is how slow it got for the unluckiest one in a hundred, which is the number your visitors actually feel — an average hides exactly the problem you are looking for. And Socket errors or Non-2xx is whether it was still working at all, because a server that fails instantly looks wonderfully fast.
While it runs, watch the machine rather than the tool. Four terminals, or four commands one after another:
htop # is it the processor, and which process
free -h # is it memory, and is swap being used
sudo tail -f /var/log/nginx/error.log # the web server's own complaints
journalctl -u myapp -f # your app's complaints
# And afterwards, the one nobody thinks to check: did the log itself
# give up and start dropping lines while you were not looking?
journalctl --since '-10 min' | grep -i suppressed
wrk, hey and Apache's ab all wait for each
answer before sending the next request. So when your server stalls for a second,
the tool politely stops asking — and the stall never appears in the
percentiles. It is called coordinated omission, and it is why a tool can
report a lovely 99th percentile for a server that just froze. For "how much can it
take" this does not matter and wrk is fine. For "how slow is it at
exactly 500 requests a second", use a tool that keeps a fixed rate no matter how
the server feels:
oha with
-q and --latency-correction,
vegeta, or
k6's constant-arrival-rate.
Where to run it from matters more than which tool you pick. From the box itself is easiest and gives you a floor, since the generator is stealing processor from the thing it is measuring. From your laptop at home you are mostly measuring your own internet connection. The honest version is a second cheap VPS, billed by the hour, in a different data centre — a dollar, deleted the same afternoon — and that is also the only setup that tells you anything about your network path.
And since this whole guide is about handing the typing to an agent, hand it this too. It is a good task for one: mechanical, repetitive, and it needs somebody patient enough to watch four terminals.
Stress-test my app on this box without touching the live one. Start a second
copy on another port as a transient systemd unit with the same properties as
its real unit file, then hit it with wrk at 10, 50, 200 and 1000 connections
for 30 seconds each. While each run goes, record processor, memory and swap
for that unit from its cgroup, and afterwards check the journal for dropped
messages and nginx's error log. Tell me the first thing that broke, what the
evidence was, and what the numbers were before and after any change you
suggest. Change nothing without asking me first.
Leave room, because the real day will not look like your test
Now the part that makes the whole exercise worth something. Whatever number you measured, do not plan to run anywhere near it. A test sends one kind of request, from one place, with a warm cache, with no logged-in users, no search queries, no reports being generated, no scrapers, and no everyone-retrying-at-once. A real peak is all of those at the same time.
The best cautionary tale in the business is Niantic's: before launching Pokémon GO they load-tested to five times their most optimistic traffic estimate, which sounds like paranoia. Real traffic arrived at roughly fifty times that estimate. Google's own reliability book, which tells that story, is also where the rule of thumb comes from: enough spare capacity to survive a planned outage and an unplanned one at the same time. Brendan Gregg puts the practical line at 70% — past that, a resource starts behaving badly rather than proportionally.
And when it does happen to you — when the thing is genuinely too big for one box — the order is not a mystery, and it is mostly not "buy more servers":
- Cache the expensive thing. Nearly always the whole answer, and nearly always cheaper than anything else here. Caching, and the box that keeps up.
- Put the big files on somebody else's network. Pictures, video, downloads. A CDN in front of your box is free at the size you are at, and it is the difference between a bandwidth bill and no bandwidth bill.
- Buy a bigger box. A checkbox and a reboot, and it doubles or quadruples what you have. Unfashionable, immediate, and usually correct.
- Then, and only then, buy a second box and put something in front of it that shares the work. That is a different kind of project, with failover, replicas and a whole new class of way to go wrong — More than one box and the six-box setup are what that looks like, and both of them start by telling you not to.
Tuning: what moved the number, and what famously did not
Every change below was measured on its own, against the same test, one at a time. That is not diligence for its own sake: change three things and measure once, and you have learned nothing except that the three of them together did something. Here is the whole afternoon, including the parts that were a waste of time, because those are the ones nobody publishes. (The dashboard row was measured at ten thousand viewers while the service could still borrow swap; the row below it is what stopped it doing that.)
| Change | Before | After |
|---|---|---|
| Build the page once per reading, compress it once | 127/s | 7,109/s |
| Compress each stylesheet and script once, not per request | 1,374/s | 19,208/s |
| Log failures and slow requests only | 35,459/s | 48,580/s |
| Convert each live reading once, not once per viewer | 66% of frames | 94% of frames |
MemorySwapMax=0, and refuse viewers past a limit | froze silently | held, or refused |
| Keep-alive connections from the web server to the app | 7,264/s | 7,507/s — noise |
GOMEMLIMIT, the famous Go memory knob | No effect. The pressure was socket buffers and swap, neither of which it governs. | |
| Removing the processor cap on the service | 18,166/s | 37,898/s — and kept the cap anyway |
That last row is a judgement rather than a measurement. Taking the processor limit off this service doubles what it can do, and it stays on, because this box also runs a database, a web server and somebody else's evening. A service that can take the whole machine will, on exactly the day you are not watching. The cap costs throughput nobody is asking for and buys a machine that stays usable while one thing misbehaves. That is a trade worth making on a shared box and a silly one on a dedicated one — which is the general shape of tuning advice: it depends on what else is in the building.
The two rows that did nothing are the ones I want you to take away. Upstream
keep-alive is on every tuning list on the internet, and here it was worth 3%, within
the noise, because the connection it saves is to 127.0.0.1. Go's memory
limit is the first thing anybody reaches for, and it governed none of the memory that
was actually filling up. Both are good advice in general. Neither was my
problem, and I would have "fixed" my site with both of them and shipped exactly the
same page at 127 requests a second. Measure first. The famous knob is famous because
of somebody else's bottleneck.
GOMEMLIMIT and GOGC honestly, including leaving
5–10% of headroom under a container limit.
net/http/pprof shows
you where the time actually goes, and
profile-guided optimisation
turns that profile into a 2–14% faster binary for free.
nginx's keepalive documentation
is the one everybody quotes and the one that needs an upstream block to
do anything at all.
Microcaching
is the web-server version of what fixed this page — their own test went from
5.5 to 2,185 requests a second by keeping each answer for one second.
And Brendan Gregg's USE method
is the one page to read before any of it: for every resource, check utilisation,
saturation and errors, and change one thing at a time.
The tool that produced every number here is in this site's own repository, as one script and a paragraph explaining what it cannot tell you, because a number without its method is just a boast. Run it, keep the output, and put the date on it. Next year's you will want to know whether the box got slower or the project got heavier, and the only way to answer that is to have written it down.
When you actually need more than one box
Shown once, priced honestly — because you are almost certainly not going to need it. The section after this one is for the day you do.
Somewhere in your reading you will be told that a single server is amateur hour and that real systems have load balancers, failover, health checks and a fleet. Most of the people saying this work somewhere with a compliance department, and all of it is true for them. It is worth understanding what they are actually buying, because it is not what most people think.
A second box buys you exactly two things: capacity beyond what one machine can do, and survival when one machine dies. That is it. Those are the only two. And the first one is much further away than anybody expects — an unmanaged VPS of the size in this guide will absorb millions of requests a day without complaining, provided you have not done anything silly, and "provided you have not done anything silly" is the part almost every scaling story is actually about.
The cheaper answer, first
Before a second machine, there is a bigger machine. Doubling the RAM on one box is a checkbox, a reboot and about twenty extra dollars a year, and it requires no change to anything you have built. There is no shame in this, whatever anyone tells you. People reach for horizontal scaling because it sounds like engineering and vertical scaling sounds like giving up, when in reality one of them is a checkbox and the other is a new distributed system with new failure modes you have never debugged.
Before you add a machine, the honest checklist:
- Have you measured? Not felt. Measured. Which resource is exhausted — memory, CPU, disk, or a database doing a full scan because a query has no index? Nine times in ten it is the last one, and it is one line of SQL.
- Have you added the index? The single most common "we need to scale" is a missing index. It is free and it takes a minute.
- Have you cached the expensive thing? The page that takes two seconds to build, built once a minute instead of once a request. Caching is the whole of that.
- Have you tried the bigger box? Twenty dollars a year, no architecture change, no new failure modes.
What it actually looks like
Here it is, once, so you have seen it. nginx will happily balance across several machines — this is the whole configuration, and it is genuinely this small:
# Two application servers behind one public name. nginx sends each
# request to whichever is healthy; "max_fails" and "fail_timeout"
# are what make it stop sending traffic to a box that has died.
upstream app {
server 10.0.0.11:8090 max_fails=2 fail_timeout=10s;
server 10.0.0.12:8090 max_fails=2 fail_timeout=10s;
}
server {
listen 443 ssl;
server_name mysite.com;
location / {
proxy_pass http://app;
proxy_next_upstream error timeout http_502 http_503;
}
}
Twenty lines. Looks easy. Now here is the bill, and the bill is not the money.
| What you add | Costs | And then |
|---|---|---|
| A second app server | +$44.88/yr | Every deploy now has to reach both, and they must never be running different versions for long. |
| Somewhere shared for the data | +$44.88/yr | Two app servers cannot share a SQLite file. You now need a database on its own machine — which is itself a single point of failure, so you have not actually removed one, you have moved it. |
| Shared sessions and uploads | Time | Anything held in one server's memory or on one server's disk breaks the moment a user's second request lands on the other machine. |
| The load balancer itself | A third box, or a managed one | If it is one box, it is a single point of failure. If it is two, you are now doing DNS failover or a floating IP, and you have built a small network. |
| Real failover | Ongoing attention | Health checks that are honest, a database that can promote a replica, and a rehearsal — because failover you have never tested is the same as no failover, only more expensive. |
So the true price of "highly available" is not $90 a year. It is three or four machines, a distributed data story, a deployment process that updates several places atomically, and a class of bug where everything is fine except on one server, sometimes. That is a real and reasonable thing to build when the requirement is real. It is a miserable thing to build because a blog post implied you should.
The second box worth having
There is one that pays for itself immediately, and it is not a load balancer. Rent a second cheap VPS at a different company, in a different country, and make it the target for your backups and nothing else. No traffic, no load balancing, no shared state. It costs the same twenty-odd dollars, and it means the failure that actually ends projects — the provider losing a disk, or suspending an account over a billing mix-up on a card that expired — is an inconvenience instead of an ending.
That is redundancy. Two copies of the thing that cannot be recreated, in two places that cannot fail together. Load balancing is a performance decision and 99.9% of people reading this will never need one. Backups are not advanced and everybody needs them. Almost every tutorial on the internet has these two exactly the wrong way round.
The six-box setup, for when it is real
Load balancers, a floating address, a primary database and its replica, and the automation that fails over and then comes back. For mission-critical projects only. Read it so you know the shape; build none of it yet.
So why is it here at all? Because pieces of it are exactly what you reach for on the day a project suddenly takes off, and that is not a day you want to spend learning the shape of the thing. About one project in a hundred is built so that it can scale, and about one in a hundred ever needs to. They are very rarely the same project.
The shape
- Visitors look up one name and get one address. They never learn there are six machines behind it, and they should not have to.
- That address is a floating IP: one your provider lets you move from one machine to another in seconds, usually with a single API call. Right now it belongs to balancer A.
- Balancer B is identical to A and does nothing except watch it. The heartbeat between them is B's entire job.
- Either balancer can send any request to either app server. They run the same code at the same version and keep nothing on their own disks, so it genuinely does not matter which one answers.
- Every write goes to the primary database. The replica copies it continuously, takes the heavy reads, and is the spare.
- When A goes quiet, B claims the floating IP. That is failover, and it is the easy half.
- When the primary dies, the replica is promoted — and then something has to rebuild the dead machine as the new replica and let it catch up. That is recovery. It is the half nobody rehearses, and until it is finished you are running on one database again with every dashboard green.
What each pair is for
- The floating address
- Not every provider sells one, and that alone can decide where your balancers live. Without one, the fallback is DNS: a short TTL set days before you need it, and a record you change by hand or by script. It is slower, because some resolvers hold on to an old answer longer than they were told to, but it is how most small setups actually fail over, and it is a lot better than nothing.
- Two balancers
-
nginx with an
upstreamblock, exactly the twenty lines in the section above, or HAProxy. The second one sits there doing nothing, which is not waste. It is the product. - Two app servers
- The moment there are two, anything kept on one app server's disk or in its memory is a bug waiting for a user's second request to land on the other one. Sessions go in the database or Redis, uploads go to object storage, and every deploy reaches both before either serves the new version for long.
- A primary and a replica
- Postgres ships replication built in. Writes go to the primary, the replica follows it, and reports and dashboards can read from the replica so the heavy readers stop fighting the writers — which is exactly where the invite tree in Caching should have run.
The health check has to ask the right question, and "is it up?" is not it. I learned this one the slow way, with a script that sent reads back to the primary whenever the replica was down. A replica can be completely up — answering every query, instantly — and still be serving old data, because replication quietly stopped. Down is not the only way to be wrong. Ask how far behind it is:
# On the primary: is every replica connected, and how far behind is each one?
sudo -u postgres psql -c "SELECT client_addr, state, replay_lag FROM pg_stat_replication;"
# On a replica: am I a replica, and how old is the newest change I have applied?
sudo -u postgres psql -c "SELECT pg_is_in_recovery(), now() - pg_last_xact_replay_timestamp();"
The second number grows on a quiet database even when everything is fine, because nothing new has been written for it to apply. Read it together with the first, never on its own.
Failover is half. Recovery is the other half.
Failover gets all the attention, because it is the dramatic part and it can be made automatic and fast. Recovery is what comes after, and it is where the real damage is done. The old primary is now out of date, and it must never come back as a primary on its own: two machines that both believe they are the primary will both accept writes, and the data splits in two in a way that no tool will ever merge back together. It has to be rebuilt as a replica of the new primary, allowed to catch up, checked, and only then are you back to two.
Automate both halves or neither. A tool such as Patroni handles promotion and rejoining for Postgres, and it is a serious piece of software that deserves a serious week of reading before it goes anywhere near your data. Then rehearse the whole cycle on a schedule, because failover you have never tested is not failover. It is a hope with a monthly bill.
The pieces you actually reach for in an emergency
The day a project takes off is not the day to build six boxes. It is the day to do the next cheapest thing, in this order, and to stop at the first one that is enough:
- Cache the expensive thing. Caching. Hours of work, and very often all you needed.
- Buy the bigger box. A checkbox and a reboot.
- Move the database to a machine of its own. The first real split, and the one that buys the most.
- Add a read replica. Point the reports and dashboards at it, so the heavy readers stop fighting the writers.
- Put the logged-out pages behind a CDN. Most visitors then never reach your server at all.
- Keep a warm standby at a different company. Set the DNS TTL short now and write the steps down now, so losing an entire provider means changing one A record.
- Only then, the whole six boxes. If an hour offline still costs more than everything above: the balancers, the floating address and the automation.
Glossary
Each of these is explained where it first turns up. That is no help at all on the Tuesday three weeks later when you meet it again in an error message, so here they all are in one place.
None of these words are difficult. Nearly every one of them is named after the thing it does, and stops being intimidating the moment somebody tells you what that is — which is roughly one sentence of work, and the reason a beginner is made to feel stupid for a fortnight is that nobody could be bothered to do it. So: one sentence each. If you know a term already, skip it; nothing here is a test.
The machine
- VPS
- Virtual private server. A computer you rent by the month in somebody else's building. Nobody else is using it — it is simply not in your house. Buying the server.
- root
- The account that is allowed to do anything, including all the things you did not mean. You log in as an ordinary user and borrow its powers deliberately.
- sudo
- "Run this one command as root." The password it asks for is yours, not root's, and typing it is the moment to read the command once more.
- apt
- How software gets onto an Ubuntu machine: an archive Ubuntu builds and signs, and one command that installs from it.
updaterefreshes the list of what exists;upgradeis the one that installs. apt. - PPA
- Somebody else's archive, added to yours. Their builds then run as root on your machine every time they update, which is the deal you already have with Ubuntu, minus Ubuntu. apt.
- permissions
- Who may read, write or run a file: its owner, its group, and everyone else, each answered separately. The usual reason something works for you and fails as a service. Permissions.
- daemon
- A program that runs quietly in the background from the moment the machine boots, with nobody watching it. nginx is one. Your app becomes one.
- CVE
- A published security hole, with an identifier so everybody is talking about the same one. Around 250 new ones a day, nearly all of them about software you do not run. Staying up to date.
- end of life
- The date after which nobody fixes a piece of software any more. It keeps working perfectly, which is the trap. Staying up to date.
- p99
- The slowest one request in a hundred. The number your visitors actually feel, and the one an average is best at hiding. Stress testing.
- systemd
- The thing that starts daemons, restarts them when they die, and writes down whatever they said on the way. If your app is running at all, systemd is why.
- unit
- One file telling systemd how to run one thing: which command, as which user, and what to do when it stops. Yours will live in
/etc/systemd/system/. - timer
- A unit that starts another unit on a schedule — systemd's replacement for cron, with logs. Timers.
- port
- A numbered door on the machine. 443 belongs to the web, 22 to SSH, and the 8090s are yours to hand out as you please.
- loopback
127.0.0.1— the address meaning "this machine and nothing else". A program listening there cannot be reached from the internet at all, which is exactly where your app should sit.- firewall
- The list of which ports the outside world may knock on.
ufwis the friendly face of it: shut everything, then open three. Lock it down.
Getting in
- SSH
- Encrypted remote control of the machine's command line. Ancient, hammered on by everybody for thirty years, and the only door you will use. Getting in.
- key pair
- Two matching files. The public half sits on the server and is not a secret; the private half never leaves your laptop, and is never emailed, pasted or committed.
- passphrase
- The password on the private key itself, so that a stolen laptop is not the same event as a stolen server.
Names, and the padlock
- DNS
- The internet's phone book: it turns a name people can remember into an address machines can use. Pointing a domain.
- A record
- One entry in that book: a name pointing at an IPv4 address.
AAAAis the same thing for IPv6. - CNAME
- A name pointing at another name rather than an address —
www.mysite.comatmysite.com— so you only ever have one address to keep up to date. - TLS
- The encryption behind the padlock. HTTPS is ordinary HTTP with TLS underneath it. You will also see it called SSL, which is the old name for the old version and has stuck around the way old names do.
- certificate
- A file proving the name is really yours, signed by somebody every browser already trusts. It is free, it expires, and renewing it is somebody else's cron job once you have set it up. nginx, systemd, TLS.
The web server
- nginx
- The program holding port 443. It reads which domain the request asked for and hands it to whichever of your apps owns that name, which is the whole trick behind one box serving unlimited sites.
- load balancer
- A front door that shares requests out across several copies of your app, and stops sending them to one that has died. Rarely needed on one box. The six-box setup.
- reverse proxy
- That job, given its proper name. A normal proxy fetches the internet on your behalf; a reverse proxy receives the internet on your app's behalf, so your app never has to face it directly.
- rate limit
- A ceiling on how often one address may ask, enforced by nginx before your app hears about it. Not for attackers, who work from a list you are not on; for the one script with a bug in its loop. Not getting flooded.
- access log
- The file nginx appends one line to for every request: who asked, for what, and what they got. Already your analytics;
awkis the dashboard. Who came.
Your code, and its history
- repository
- A folder that git is watching, plus every version of it that has ever existed. Everyone says "repo".
- commit
- One saved point in that history, with a message saying why. Also the verb for making one. Cheap, so make them constantly.
- branch
- A line of history running alongside the main one, so that unfinished work does not have to be finished before anything else can start.
- CI
- Continuous integration: a robot on somebody else's computer that runs your tests every time you push, so "it worked on my machine" stops being a defence. GitHub Actions.
- environment variable
- A value handed to a program by whatever started it, rather than written inside it. This is where a secret goes: a file only the service can read, named from its systemd unit, and never in git. An environment variable, worked through once.
- release
- One build, in its own dated folder, beside the builds before it. The service runs whichever one a link called
currentpoints at. - rollback
- Pointing
currentat the release before this one and restarting. One command, about a second, if the deploy was built to allow it. The deploy that broke it.
Data
- CRUD
- Create, Read, Update, Delete — the four things every application has ever done to anything. Everything is CRUD.
- WAL
- Write-ahead log. SQLite writes each change to a side file before folding it in, so readers are never blocked by a writer and a power cut leaves you something repairable rather than something shredded. Turn it on once and forget it. Databases.
- cache
- A copy of an answer kept so it does not have to be worked out again. Fast, and wrong the moment the original changes, which is why every copy needs an expiry. Caching.
- rollup
- A timer that turns a table of raw events into a table of hourly totals, so the pages that need the number read a sum that already exists. A job, not a page. When your app wants its own numbers.
- replica
- A second database that continuously copies a primary one. It can take the heavy reads and stand in if the primary dies, and it can be up and quietly out of date. The six-box setup.
- restore drill
- A timer that restores last night's backup into a scratch copy, counts what came back, and prints one line. The backup had a timer; this gives the restore one. Then make the machine prove it.
Cheat sheet
The ones you will actually use. Bookmark this bit; nobody memorises these and nobody should.
Getting around
ssh sam@203.0.113.40 # connect
exit # leave
cd /srv/myproject # go somewhere
ls -la # what is here, including hidden files
pwd # where am I
nano file.txt # edit (ctrl-O saves, ctrl-X quits)
cat file.txt # print a file
tail -f file.log # watch a file grow
grep -rn "thing" . # find text in every file below here
du -sh * # what is taking up space
Packages
sudo apt update # refresh the catalogue; installs nothing
sudo apt upgrade # install what is newer
apt search --names-only thing # is there a package for it
apt show thing # what is it, how big
sudo apt install -y thing # get it
sudo apt remove thing # drop it (purge: and its config too)
apt list --upgradable # what is waiting
Services
systemctl status myapp # is it running?
sudo systemctl restart myapp # turn it off and on again
sudo systemctl enable myapp # start it at every boot
journalctl -u myapp -f # follow its log
journalctl -u myapp -n 200 # the last 200 lines
systemctl --failed # what is broken
Web
sudo nginx -t # is the config valid?
sudo systemctl reload nginx # apply it, safely
sudo certbot certificates # what certs exist, expiring when
sudo certbot renew --dry-run # prove renewal works
curl -I https://mysite.com # see the response headers
dig +short mysite.com # where does the name point
sudo tail -f /var/log/nginx/access.log # who is here, live
Git
git status # what has changed
git add -A && git commit -m "..." # save a checkpoint
git push # send it to GitHub
git log --oneline -10 # recent history
git diff # what have I changed
git restore <file> # undo my changes to one file
git revert <commit> # undo a commit, safely
The agent
cd / && claude # start it where it can see everything
/model # switch model: plan vs build
/clear # fresh context, same folder
# a fact worth keeping # writes it into CLAUDE.md
Esc # stop it, without ending the session
Shift+Tab # cycle permission modes
Health
htop # live CPU and memory
free -h # memory left
df -h # disk left
sudo ss -tlnp # what is listening, and on which address
sudo ufw status verbose # what the firewall allows
sudo fail2ban-client status sshd # who has been banned
uptime # how long since the last reboot
Go and build something
That is the whole method. A server for $44.88 a year, a domain for the price of a coffee, and an agent with a terminal and permission to use it. Everything else on this page is detail hung off those three things.
The reason to write it down is that the hard part was never the code. The hard part was knowing that a server costs less than a streaming subscription, that a certificate is free and takes fifteen seconds, that subdomains are unlimited, and that the right move is to climb out of the graphical editor and let something capable work in a place where mistakes cost nothing. None of that is secret. It is just very rarely said plainly, in one place, by somebody with nothing to sell you.
I had an art teacher in middle school who was better at teaching than almost anyone I have met since. He asked us to draw a perfect circle. Everyone produced something wobbly and apologetic. Then he drew six fast, sloppy circles on top of each other and rubbed out everything that was not part of the very obvious, very perfect circle now sitting there in the middle of them. That is what building software is. Not one correct line after another — six rough passes and then removing what was not it. Your first version will be wobbly. It is supposed to be. It is one of the six.
Take all of this, change it, disagree loudly with the opinionated parts. The stacks here are one set of preferences that works well on a small box with an agent doing the typing, and they are not the only set. What compounds is not my conventions — it is having conventions, writing them down, and reusing them until your fifth project starts on the afternoon you think of it rather than the following spring.
And do not wait to be told you are allowed. There is no book, no course, no certificate and no job that hands out permission; the people waiting for one are missing the rest of the horse, not you. Go and build the smallest ugly version of the thing, tonight, on a machine that does not matter. The world is your delicious oyster. Good luck and Godspeed.
See it running
Real-time telemetry from the box serving this page. Same hardware, same price.
Built this way
Games, websites, marketing tools — all the same skeleton, all on boxes like this one.
The server prices in the survey were read off each provider's own page on 14 September 2026 — 3 days ago — and every row links to the page it came from. Domain and signing prices were checked in September 2026. Being prices, all of them are subject to change, so verify before you buy. Nothing here is sponsored and no link on this site is an affiliate link.
Every chart on this page carries its source and the date it was read, underneath the chart, along with whatever is wrong with the data. There is no chart here whose numbers you cannot go and check yourself in about a minute, which is the only kind worth drawing.
The logos are from Simple Icons, drawn beside the name of whatever the sentence is already talking about. Each one is its owner's trademark, and none of them is here because anybody asked.
The six little creatures at the top of each part, and the wide drawing at the head of most sections, were generated by an image model on this box for about ten pence in total. They are nobody's mascot but ours — not the penguin, not the daemon, not the gopher — which is deliberate, because the real ones carry four different licences between them and one of them is a trademark you need written permission to use. No drawing here contains a word, a logo or a signature, and that is enforced in the prompt rather than hoped for. Everything about how they were made, including the ninety-odd drafts that did not survive and what each one cost, is in the repository.