the complete beginner's guide

Your own server

$44.88 for a year · plus a domain · plus your Claude subscription

A real computer on the internet, running whatever you want, for less than four dollars a month. Unlimited projects. Unlimited subdomains. Any language. No experience needed, because the point is that you are not going to be typing most of this. ⁠Claude Code is.

This page is being served by one of them right now — 2 cores, 3.82 GiB of RAM, 58.8 GiB of disk. Watch it work →
processor, the last three minutes
cpu
0.3%
memory
15.9%
up
1d 17h
measured just now

Stress-tested on 16 September 2026, on this box

pages a second
7,109
live viewers at once
2,000
copies of this app, all loaded
40
full reads before 6 TB runs out
1,276,595

Room for about 145 small Go services, or 19 in Python, 15 in Node, 4 on the JVM (estimated). The bandwidth runs out long before the processor does. What broke first, and how to test yours →

No VS Code needed Written for Windows 11 Hardening first Unlimited projects Any language
01

Buy the parts

Two purchases. Ten minutes. Then it is yours.

A small green creature with a CRT monitor for a head, unpacking a server from a cardboard box.

Why bother

What stops most people is not talent. It is that they are frightened of their own laptop, and they are right to be.

A small figure looking up at a tall tower of stacked boxes with one lit door open at its base.
A rented computer, switched on permanently, with a door you hold the only key to. That is the entire product.

Watch somebody start. They get excited about building something, install VS Code on their Windows laptop, bolt an AI extension onto it and open a project. An hour later they are staring at a sidebar full of folders they did not create, being asked whether they trust the workspace, wondering where the files actually went, and hovering over every suggested change trying to decide whether accepting it will quietly break something they could not name if you asked them to.

They are not building. They are supervising — through a letterbox, one file at a time, on the same machine that holds their photos and their tax return. Of course they are careful. Anybody would be careful.

Picture a big red button on a wall, and the whole rest of that wall covered in instructions about what the button does. Most people read the wall. They read it for months and they get very good at reading walls. Then some clown wanders up and starts tapping out a drum rhythm on the button, and by the afternoon he knows things about it that are not written anywhere on that wall. The gap between those two people is not intelligence. One of them pushed the button.

Nobody pushes the button on a machine they are afraid of. So stop using that machine.

Rent a computer that is not your laptop. Give an agent a terminal on it and permission to do whatever it likes. Then describe what you want. That is the whole guide. Everything after this is detail.

The whole trick is that it is not your machine. Every hesitation you feel in an editor — what if it deletes something, what if it installs something weird, what if I break my setup — costs $3.74 a month to make irrelevant. Break it badly enough and you throw it away and have a new one before the coffee goes cold. Nothing you own is on it. You stop supervising and start directing, and that is a different job.

So: a cheap Linux box, locked down properly on day one, with ⁠Claude Code running in a terminal at the root of the filesystem and the freedom to do its job. You describe things in plain English. It writes the code, installs what it needs, configures the web server, gets the certificate, starts the service and hands you a URL. You will understand about a third of what happened the first time, and a third is the correct amount.

What you actually get

Projects
Unlimited

Nobody charges you per app. The only real limit is memory, and text-shaped things weigh almost nothing.

Subdomains
Unlimited

One domain, infinite names. app., api., demo., whatever. — all free, all with certificates, all yours.

Languages
All of them

Go, ⁠Rust, ⁠Python, ⁠Node, C, whatever. The agent installs the toolchain. You are not required to like any of them.

Uptime
Always

It runs while your laptop is shut. That is the whole difference between a demo and a thing that exists.

The thing I got wrong: for years I only had one computer. Everything I built, I built on the machine I also lived on, so I was careful, and careful is slow. I never tried the stupid version of an idea first, because the stupid version ran in the same place as everything else I owned and I could not afford to find out what it did. A second machine that costs less than a sandwich is not a luxury upgrade. It is permission to be reckless somewhere it does not matter, which turns out to be most of how anybody learns anything.

What this guide assumes about you

That you have used a computer. That is genuinely the whole list. If you have never opened a terminal, have never heard of SSH, could not define a DNS record at gunpoint and could not tell a front end from a back end, you are precisely who this was written for. Every term gets explained the first time it turns up, and none of them are as clever as they sound.

The instructions default to Windows 11, because that is what most people are reading this on, and because everything you need in order to reach the server is already installed — no PuTTY, no downloads, nothing to buy. Where ⁠macOS differs there is a tab for it. Linux users already know what to substitute and were going to substitute it anyway. There is also a section on putting ⁠Ubuntu on your own laptop, which is worth doing and is not needed for any of this.

Who is writing this

Written September 2026. Be suspicious of anybody who tells you this is easy, and start with me.

My name is Jack. Online I am saintpetejackboy, which is where the code is, and Deadend Deafchild, which is where the music is. About twenty years of building software, most of it self-taught, most of it line-of-business CRUD for companies whose names would mean nothing to you and which I would not tell you anyway.

The path was LAMP, then procedural ⁠PHP — learned before PHP had real object orientation, which is the kind of sentence that dates a person — then a long stretch of proprietary work for other people, and now Go for back ends, ⁠Rust when something is genuinely low-level, and ⁠TypeScript when the browser leaves me no choice. I learned by running a web server as a kid, on hardware measured in megabytes, taking payment by money order in the mail.

The other half of that childhood was Gaming World, the RPG Maker community at gamingw.net, where I was Kazeuri — Kaze, to most people — from about as far back as it goes. I was born in 1987, so “a kid” means a kid. At one point I got myself banned. So, naturally, I had a couple of people I knew on the site's IRC network change my hostname to look like it came from Squaresoft, introduced myself as “Hironobu”, and talked my way onto the staff reviewing games, at a moment when the only two people who could have called it were both, conveniently, away. A few people have claimed since that they knew all along. That is how I ended up running the mailbag, and it is the first security lesson I ever taught anybody, which the ⁠nginx section puts to use.

gamingworld.uk is my homage to that era now: more than a hundred game previews, and a lot of my own games live there. Everything else I build that I can show you is in the portfolio, including both of the projects this guide tells stories about.

It is not an unbroken run, and the gap in the middle is the part I should say plainly rather than leave you to work out. I went to federal prison for a long time — years, not months — for importing chemicals from China. That is the break. I am not going to perform remorse at you in a hosting tutorial, and I am not going to be coy about it either: it is a matter of public record and I have talked about it at length in public. There is a Reddit AMA a great many people read. There is an interview on the Live Struggle Legacy podcast. There is a spot on the Toby Gribben show, where I was billed as Dead End Deaf Child. There is a documentary interview I did with Lawrence for the Kratom Advocacy Podcast, and another for the Global Kratom Coalition that I never did find out the fate of. If you want that story, it is out there and I have never hidden from it.

The other three are named and not linked, for now. The Toby Gribben show and both kratom documentaries live on platforms I do not control, one of them I have genuinely lost track of, and a dead link on a page like this is worse than no link. They get cards here when I have collected the URLs myself. Until then, the names above are enough to find them.

What that has to do with this page is narrow, and it is the honest reason the whole thing is free. You can take a great many things from a person. You cannot take what they know how to do. I came back to programming after more than half a decade away and was mainly surprised by how little had changed — the fundamentals I learned as a kid on a box measured in megabytes were still the fundamentals, and everything built on top of them was a fashion. I also came back without the things that normally get you hired, which turns out to matter less than you would think when you can rent a whole computer for the price of a couple of sandwiches and put whatever you like on it. Nobody's permission is required. Nobody can take the box away because of a form. That is not an abstract argument about self-hosting for me, and it is why this guide exists at all.

The other reason is a person. My best friend Julio and I met in high school and we still see each other constantly, twenty-five-odd years later — me and Julio, down by the schoolyard, and everybody who knows us already knows that. He is on this machine too; it is not only mine. Most of what is written here is what I would tell him, in the order I would tell him, because that is who I am actually writing for — and once you have explained something properly to one friend you may as well write it down for everybody, since the second time costs nothing.

So: no advertising, no affiliate links, no course at the end, no newsletter, no email address to surrender. This is a living document. It was last revised in September 2026, the prices are dated and checked against their sources, and I update it when something changes or when somebody points out that I am wrong — which happens, and is welcome.

Which brings us to the one caveat that covers everything below. I am a person with opinions, not a standards body. This is one way to do it, arrived at by one guy over twenty years, and it works — the website you are reading is running on the exact setup it describes. It is not the only way, several of the choices are matters of taste that I will happily defend and you are free to ignore, and there is more than one way to skin a cat. Take what is useful, bin the rest with my blessing.

So the honest summary: most of my code would make a grown neckbeard cry himself to sleep for a solid month. What I can also tell you is that I finish things, and that almost nobody who is intimidating you online does.

The thing I got wrong: I did not do half of this a year ago. The justfile, the docs/ folder, the notes I leave for the agent — every project I started in 2024 has none of them. I am not going to dress eleven-month habits up as twenty-year discipline. If anything it makes them more useful to you, because I still remember perfectly well what it was like without them, and I can tell you which ones actually paid.

What twenty years does buy is a sense of what to skip. Most of the advice aimed at beginners is really advice for a company with forty engineers and a compliance department, handed down by people who have never had to pay for their own server. You are going to be told you need containers, a pipeline, a staging environment, an orchestrator and a monitoring stack before you are allowed to have a website. You do not. You need one cheap box you are not frightened of, and the discipline to back it up.

The first time you do any of this, something will break, and the thing that breaks will be dull — a typo in a config file, a DNS record that has not propagated, a permission on a folder. It will not be dull at the time. That is normal, it happens to everyone at every level of experience, and it is the entire job. Anybody who presents this as frictionless is selling you something.

The shopping list

Two things you have to buy, one you are probably already paying for.

A virtual private server

A whole Linux computer in a data centre, switched on permanently, with its own address on the internet. Yours until you stop paying.

$44.88per year

A domain name

The name people can actually type. Buy one and you get unlimited subdomains off it for nothing.

$5–$11per year

A Claude subscription

⁠Claude Code comes with every paid plan. Pro is plenty to start with; you will know when you want Max.

$20+per month

Here is the real arithmetic, including the optional bits nobody mentions until the invoice turns up. Untick whatever does not apply to you. I would rather you saw the honest number now than found it in month three.

What a year actually costs

Real prices, checked in September 2026. Untick anything you do not need.

VPS — 4 GB RAM, 2 vCPU, 60 GB NVMeDediRock Promo Performance NY Power, billed annually
$44.88
Your first year $0.00
Only the first two lines are infrastructure. A server and a domain come to about $55 for a whole year. Everything under that is your AI subscription and two optional extras that most projects never need and that you can buy later, on the day you need them, rather than today.

Buying the server

Ten minutes, a card, and an email with a password in it. That is the barrier.

Three tall server cabinets standing apart, a small figure at the left holding up an oversized key.
You are not buying a service. You are renting a computer — and the difference turns up on the invoice every month for the rest of your life.

A VPS — virtual private server — is a slice of a big machine in a data centre that behaves exactly like a computer of your own. Its own memory, its own disk, its own operating system, its own address on the internet. You are the administrator. Nobody else can see inside it. It is a desktop tower in somebody else's building, and the somebody else is responsible for the power, the cooling and the network cable.

There are hundreds of companies selling these, and for the same machine the spread between the cheapest and the most expensive is close to tenfold. The one this guide prices against is DediRock's Promo Performance range in New York — not because it won a comparison, but because it is the box this page is served from, and every command in this tutorial was run on it. The survey further down puts it beside eight other providers at the same specification, so you can see where it actually sits rather than taking my word for it. Nothing in this guide depends on who you rent from, and you should shop around.

What you are deliberately not buying is a cloud account. The big platforms are magnificent engineering and they are also the apex of nickel-and-diming: a charge for the machine, a charge for the disk, a charge for the traffic leaving the disk, a charge for the load balancer in front of the machine, and a bill at the end of the month that nobody in the building can fully explain. An unmanaged box at a flat annual price can absorb an amount of traffic that would genuinely surprise you, and the number on the invoice is the same every year.

PlanRAMvCPUDiskBandwidthPer year
Core2 GB130 GB2 TB$24.88
Plus3 GB140 GB4 TB$34.88
Power pick this4 GB260 GB6 TB$44.88

Why the middle option is a false economy

The ten dollars between Plus and Power buys a second CPU core and an extra gigabyte of memory, and both matter more than they sound. The second core means a build can run while your website carries on answering requests, instead of the site going quiet every time you deploy. The extra gigabyte is what stops the machine grinding when a database, two or three small services and a compiler are all awake at once.

Memory is the constraint that will actually bite you — not disk, not CPU. Sixty gigabytes of disk is more than a hundred text-shaped projects will ever manage to fill. Four gigabytes of RAM is comfortable, and comfortable is not the same as infinite. This is the one number worth ten extra dollars, and it is the only place in this guide where I will tell you to spend more.

One price is not the price

Here is the same machine bought eleven different ways — four gigabytes of memory, or the closest thing each company sells. The cheapest row and the most expensive row are the same computer for your purposes, and there is a factor of ten between them. That spread is the single most useful fact in this section, and it is the reason a guide that names exactly one provider is doing you a disservice no matter how good that provider is.

Nobody paid to be in this table and there are no referral links on this site. Not here, not in the domain section, nowhere — a hosting guide with affiliate codes in it is an advertisement wearing a tutorial's clothes, and you would have no way to tell which rows were moved by money. One row is marked this box, because this website is served from it. One is a friend's company, because I know the people who run it and it is here because they asked, and one is my other box, because I pay for it. All three are conflicts of interest, and they belong in the open, in the table, rather than in a disclaimer nobody scrolls to. Those rows were held to exactly the same rules as the rest.
Four gigabytes, eleven ways

4 GB of memory, or the closest thing each provider sells. The annual figure is the honest comparator, because several of these are only cheap if you hand over a year at once. A figure in red is not what it looks like — the card underneath says why.

ProviderPer yearMemoryCPUDiskTraffic
RackNerd XMas Sale 2.5 GB KVM $28.71$2.39/mo 2.5 GB 2 vCPU 38 GB SSD RAID-10 6.5 TB
DediRock this box Promo Performance — Power $44.88$3.74/mo 4 GB DDR5 2 vCPU 60 GB NVMe 6 TB
OVHcloud VPS-1 $54.48$4.54/mo 4 GB 2 vCore 40 GB NVMe unlimited
netcup VPS 500 G12 €59.52€4.96/mo 4 GB DDR5 ECC 2 vCore 128 GB NVMe included
Hostinger KVM 1 $77.88$6.49/mo 4 GB 1 vCPU 50 GB NVMe 4 TB
Hetzner CAX11 (Ampere ARM) $83.88$6.99/mo 4 GB 2 vCPU ARM 40 GB NVMe 20 TB
WingsHoster a friend's company Wings VPS 4G ₹10,685.52₹890.46/mo 4 GB 2 Xeon cores 60 GB NVMe unmetered
BuyVM Slice 4096 $180.00$15.00/mo 4 GB 1 core 80 GB SSD unmetered
Contabo my other box Cloud VPS 8 (2026), US East $242.16$20.18/mo 23 GiB 8 cores 300 GB unstated
DigitalOcean Basic Droplet, 4 GB $288.00$24.00/mo 4 GB 2 vCPU 80 GB SSD 4 TB
Oracle Cloud Always Free — Ampere A1 $0.00$0.00/mo 12 GB 2 OCPU ARM 200 GB block 10 TB egress

Surveyed 14 September 2026 — 3 days ago. Every row but one was read off that provider's own page, and each card below carries the link and the day it was read. A few of these companies will not quote a price to anything that is not a person in a browser, and the cards say where that changed how a number was checked. just market-check re-reads all eleven sources and says which numbers have moved; the test suite fails once this survey is older than 210 days, so a stale table breaks the build rather than quietly lying to you.

What those rows are actually like — including the two I would not buy, and why they are in the table anyway:

RackNerd XMas Sale 2.5 GB KVM $28.71

The cheapest real thing in this table, and the clearest lesson in it. This is a promotional SKU with a holiday name still on it in September, sold from a separate page that is not linked from their catalogue. The catalogue price for a 4 GB box from the same company is $24.59 a month. Same provider, same hardware, roughly ten times the price, and which one you get depends entirely on which URL you arrived through.

$28.71 a year — $2.39 a month, annual, prepaid

9 US cities, Amsterdam, France, Dublin, Toronto

my.racknerd.com — read 12 September 2026

DediRock Promo Performance — Power $44.88

The box this website is served from, which is the only reason it is the one priced throughout this guide. Nothing here depends on it. Every command in this tutorial ran against this plan, and the dashboard is that machine reporting on itself.

$44.88 a year — $3.74 a month, annual, prepaid

New York

billing.dedirock.com — read 12 September 2026

OVHcloud VPS-1 $54.48

A company with its own data centres and its own fibre, at a price that sits inside low-end territory. The quoted rate needs twelve months paid upfront; month to month is meaningfully more. The Asia-Pacific regions are the exception to “unlimited” — 500 GB on VPS-1, then the line is throttled to 10 Mbps rather than billed.

$54.48 a year — $4.54 a month, 12 months upfront

Canada, US, France, Germany, UK, Poland, Singapore, Sydney, Mumbai

ovhcloud.com — read 12 September 2026

netcup VPS 500 G12 €59.52

Three times the disk of anything else at this price, error-correcting memory, and no minimum term at all. Two things about the number. It is in euros, so what leaves your account depends on the day and on your card issuer. And their page quotes you €4.96 or €5.91 for the same machine depending on where it thinks you are — the second one includes 19% German VAT, which a customer outside the EU does not pay. Every European provider does some version of this, and it is the single most common way a price comparison between a US and an EU host comes out wrong.

€4.96 is the rate at 0% VAT, which is what a US buyer pays; the same page shows €5.91 including 19% VAT to a visitor it places in Germany. Both figures are theirs, read the same day.

€59.52 a year — €4.96 a month, no minimum, plus VAT in the EU

Germany, Austria, US

netcup.com — read 12 September 2026

Hostinger KVM 1 $77.88

Read the second price, not the first. The advertised rate is an introductory one on a multi-year commitment and it renews at $11.99 a month — $143.88 a year, three times the DediRock row. It is also one vCPU, which is the floor this section tells you to stay above. In the table for completeness and because their marketing is unavoidable.

$77.88 a year — $6.49 a month, long intro term, then $11.99/mo

US, UK, Netherlands, Lithuania, India, Singapore, Brazil

hostinger.com — read 12 September 2026

Hetzner CAX11 (Ampere ARM) $83.88

The one most people will tell you to buy, and the one whose numbers were hardest to establish. Twenty terabytes of traffic is not a typo. Two things to know: it is ARM, so your Go binary wants GOARCH=arm64 — one word in the build command, and nothing else in this guide changes; and in June 2026 Hetzner raised cloud prices substantially, more than doubling some CPX plans. That is the argument for dating a price rather than trusting one.

Hetzner's own pricing page renders its prices in the browser and quotes nothing at all to a plain HTTP client, so this figure is their published post-adjustment list price for CAX11 rather than a number read off the shop page. Specifications are from their cost-optimized page. Check it yourself before you buy.

$83.88 a year — $6.99 a month, monthly, plus an IPv4 charge

Germany, Finland, US, Singapore

docs.hetzner.com — read 12 September 2026

WingsHoster a friend's company Wings VPS 4G ₹10,685.52

A friend's company, and marked that way in the table for the same reason the DediRock row is marked: you are owed the reason a row might be liked, where you can see it. It came in through the same door as every other row — the price read off its own page, held to the same four gigabytes. Three things to know. The price is in rupees, so, like the netcup row, what leaves a dollar account depends on the day and on your card. Unmetered is on a shared 1 Gbps port, which means nobody counts your traffic and nobody promises its speed either. And it is a small Indian company — MSME-registered, the page says — so read its mention history on the forums below before you give it money, exactly as you would for anyone else in this table. The same company publishes HM360, billing and automation software for people running a hosting business, which is a strange and interesting thing to discover exists once you have a box of your own.

Priced in Indian rupees; the dollar cost moves with the exchange rate and any foreign-transaction fee on your card

₹10,685.52 a year — ₹890.46 a month, monthly, in Indian rupees

“Multiple locations”, not named on the plan; the company sells into India and Malaysia

wingshoster.com — read 14 September 2026

BuyVM Slice 4096 $180.00

Four times the price of the DediRock row for the same memory and one core instead of two, and people pay it on purpose. Unmetered means unmetered, block storage is available by the terabyte for a few dollars, and the company has been answering its own forum threads under its own name for over a decade. What you are buying above the specification is that nobody disappears.

$180.00 a year — $15.00 a month, monthly

Las Vegas, New York, Miami, Luxembourg

buyvm.net — read 12 September 2026

Contabo my other box Cloud VPS 8 (2026), US East $242.16

My other box, bought four days before this survey, and marked so you know why it is here. It is not the four-gigabyte machine — it is the plan I actually bought, with nearly six times the memory — and it is in this table for what its invoice teaches rather than for its specification. The plan is $14.28 a month on a twelve-month term with no setup fee. Choosing a location in the United States added $5.90 a month on top of that, which is 41%, and that line appears at checkout rather than on the plan. What actually left my account was $242.16 for the year, up front. The advertised number is real; it is just the price for somewhere you might not want your server to be. Pick the location first and read the total second, with every provider, not just this one.

Contabo answers anything that is not a person in a browser with a security check, so nothing here could read their page and just market-check cannot re-read this row. The price is from my own order confirmation of 10 September 2026, and memory, cores and disk were measured on the machine rather than copied from a listing. Nothing on the order states a traffic allowance, so the table does not guess one.

$242.16 a year — $20.18 a month, 12 months prepaid: $14.28 plan, plus $5.90 for a US location, no setup fee

US East, which is a paid option on this plan

contabo.com — read 10 September 2026

DigitalOcean Basic Droplet, 4 GB $288.00

Included so the rest of the table has a scale. This is the polished, well-documented, entirely respectable mainstream price for the same machine — six times the DediRock row and ten times the RackNerd one. You are paying for a console you will enjoy using, documentation that is genuinely excellent, and an invoice that grows every time you add a thing. It is a good product. It is not a cheap one.

$288.00 a year — $24.00 a month, monthly or hourly

15 regions

digitalocean.com — read 12 September 2026

Oracle Cloud Always Free — Ampere A1 $0.00

Genuinely free and genuinely more memory than any paid row here, and there are three catches you need before you build anything on it. Idle instances get reclaimed — Oracle's own documentation says a machine under 20% CPU and under 20% memory across seven days may be taken back. “Out of host capacity” when you try to create the instance is routine rather than exceptional. And it is one account per person, tied to a card for verification. Superb for a practice room; a strange place to put something you would be upset to lose.

$0.00 a year — $0.00 a month, free, indefinitely, with conditions

your account's home region

docs.oracle.com — read 12 September 2026

Where to look for yourself

This page is a snapshot taken by one person with one set of opinions. The market it is a snapshot of moves weekly — providers appear, run aggressive promotions, get bought, get overloaded, and occasionally stop answering tickets entirely. There is an entire corner of the internet that tracks this in real time, it has been doing it for fifteen years, and you should use it rather than my table.

To be completely clear about what this is not. This is not a deal site and it is never going to be one. It will not track promotions, it will not have a fresh table every Tuesday, and it does not want your traffic on that basis — the people linked above do that properly and have earned it. What this page is for is the part they assume you already know: what to buy, why that number and not the cheaper one, and what to do with the machine once the welcome email arrives. Take the sizing argument from here and take the current prices from them.
The ten minutes that will save you a bad year. Before you hand money to a company you have never heard of, search its name on LowEndTalk and read the last few months of mentions. You are not looking for perfection — every host has an angry thread. You are looking for a company that replies in public, under its own name, to complaints it did not have to answer. Cheap hosting is a reputation business and that forum is where the reputations are kept.

The floor: one core and one gigabyte is a demo, not a server

Every provider in that table sells something for a dollar or two a month, and for anything you intend to take seriously you should walk straight past all of them. The 1 GB, single-core box is not a smaller version of a real server. It is a different thing that fails in a specific, confusing way, and it fails at exactly the moment you finally have something worth losing.

Here is where a gigabyte actually goes. Ubuntu doing nothing at all is two to three hundred megabytes. ⁠nginx is twenty. A small ⁠Go service is fifteen, and you will run three of them before the month is out. ⁠Postgres on its defaults asks for a couple of hundred more before it has cached a single row. You are now at roughly half your memory and you have not built anything yet.

Then you run a build. A compiler is the most memory-hungry thing that will ever happen on this machine — and an agent working on your code runs builds and test suites continuously, which is the whole point of the box. On a gigabyte, that build does not run slowly. The kernel's out-of-memory killer wakes up, picks the process with the worst score, and destroys it. It usually picks the compiler or the database, so what you observe is not "I am out of memory". It is "my build randomly dies", or worse, "the site went away for thirty seconds and came back", which is a much harder sentence to search for.

The failure that actually ruins an evening is swap. A box under memory pressure with swap enabled does not stop — it thrashes, moving pages to a disk it is sharing with strangers. Load average goes to twenty, every request takes ninety seconds, and ssh stops answering. You cannot get in to fix it, because getting in needs memory. The only remaining move is the provider's reset button, which is a power cut, which is how a database gets corrupted. One core makes this likelier, because there is nothing left to schedule your shell on.

One core has a quieter cost too, which the second core in the plan above is there to buy: with a single core, everything queues behind everything. A deploy stalls the website. A certificate renewal stalls the deploy. You cannot tail a log while something compiles, which means you cannot watch the thing you are debugging while it happens.

None of which makes the tiny boxes useless. A 1 GB single-core machine is a perfectly good home for one static site, a chat bot, a webhook receiver, a ⁠WireGuard endpoint, a scheduled scraper, or a staging copy of something real. Buy three of them for three separate small jobs and you will be delighted. Just do not put the thing you care about, plus its database, plus an agent doing builds, on one of them — and notice that across that entire table, the gap between the 1 GB box and the 4 GB box is somewhere between fifteen and thirty dollars a year. It is the cheapest insurance in this guide and it is the one upgrade I will push you on twice.

The rule that actually sizes a box: your database should fit in memory

If you take one sizing heuristic away from this page, take this one. It is not precise and it is not the whole truth, and it will get you the right answer far more often than any amount of reasoning about request volume.

While your whole database is smaller than the machine's memory, every read your application makes is answered out of RAM after the first time. Cross that line and reads start going to the disk, and that transition is not a gentle ten or twenty per cent. A page from memory costs tens of nanoseconds. A random read from a shared network-attached disk costs hundreds of microseconds on a good day. That is four orders of magnitude, on the operation your application does most.

The reason it works as a rule of thumb is that it is not really about your database at all — it is about the operating system underneath it. Linux keeps recently read file pages in the page cache, using whatever memory nothing else currently wants. ⁠SQLite is a file, so it rides that cache for free; ⁠Postgres keeps its own shared_buffers on top of it. Either way, the practical question is not "how big is my database" but "is there room to keep it in memory after everything else has taken its share" — which is why the answer to a slow database is so often more memory rather than a cleverer query.

On a 4 GB boxReasonable resident size
Ubuntu, systemd, sshd, the usual daemons250–350 MB
nginx, a handful of workers20–40 MB
Three or four small Go services50–120 MB
Postgres, tuned for this box400–800 MB
Headroom for a build, a test run, an agent600 MB–1.5 GB
Left over to cache your data1–2 GB

So on the box this guide buys, the honest ground floor is roughly this: keep the database under about a gigabyte and you will never think about any of this again. Between one and two gigabytes, you are fine but you should know where the number is. Past that, on four gigabytes of total memory, you are no longer running a comfortable machine — you are running a machine with a performance problem that has not introduced itself yet.

What "it does not fit any more" looks like

It does not announce itself. Nothing errors, nothing crashes, no log line says the word memory. What you get instead is a page that is usually fine and occasionally terrible — the average stays respectable while the slowest one request in a hundred falls off a cliff, because that is the one that needed a page the cache had evicted. Reports arrive as "it felt slow earlier" and you cannot reproduce it, which is the most expensive category of bug there is.

These are the four numbers that answer it, and none of them need a monitoring stack:

# how big is the database, really
du -h /srv/myapp/app.db                        # SQLite: it is one file
sudo -u postgres psql -c "SELECT pg_size_pretty(pg_database_size('myapp'));"

# how much memory is left to cache it with
free -h                                        # read "available", not "free"

# is Postgres finding what it needs in memory? want 99%+
sudo -u postgres psql -c "SELECT sum(blks_hit)*100/sum(blks_hit+blks_read) AS hit_pct FROM pg_stat_database;"
"Available", not "free". A healthy Linux box reports almost no free memory, and beginners panic about it every single time. Memory holding cached file pages is doing useful work and will be handed over the instant a program asks for it. The number that matters in free -h is available. A box with 100 MB free and 2 GB available is perfectly happy; a box with 2 GB free is a box that is not caching your data yet.

The honest caveats, because it is a rule of thumb and not a law

It is really the working set, not the whole thing. What must fit is the part you actually touch — your indexes, and the rows people are reading this week. A table carrying four years of finished orders nobody ever queries can be far larger than memory and cost you nothing, as long as the indexes supporting the queries you do run still fit. The reason the cruder rule is the one worth teaching is that the working set is genuinely hard to estimate, it always grows in the direction you were not looking, and being wrong about it is invisible until it is urgent. Size for the whole database and you never have to be right.

Sequential and recent-only workloads are exempt. Append-only data you only ever query by recent window — logs, metrics, events, a queue — is fine at many times your memory, because the hot end is small and the cold end is never read. A hundred gigabytes of last year's telemetry on a 4 GB box is not a problem. A hundred gigabytes of customer records you filter across is a different story entirely.

The thing I got wrong: I put the files in the database. It is so convenient. One backup, one permission model, transactional deletes, no orphaned files on disk — and every argument for it is true right up until the images are ninety per cent of a database that no longer fits anywhere. This is the single most common way a modest application's database outgrows its server years before its records ever would. Keep the rows in the database and the bytes on the disk or in object storage, with a path in a column. The database section does this properly, and backups covers what changes about backing it up.

And when a database genuinely does outgrow the machine, the order of operations is almost always the same: move the blobs out, archive or partition what nobody reads, and only then buy memory. On the prices in the table above, doubling your memory costs between fifteen and forty dollars a year. There is a long tradition of spending a whole weekend tuning around a constraint that could have been bought off for the price of two coffees, and I have contributed to it personally.

⁠Choose Ubuntu, and stop thinking about it

The checkout page will offer you a dropdown with a dozen operating systems in it, and this is the one place where a beginner can lose a weekend to a decision that does not matter. Pick ⁠Ubuntu, the newest LTS. Every command in this guide is written for it, every error message you paste into a search box will have been seen by ten thousand people before you, and every answer you find will apply without translation.

Ubuntu LTS — pick this The most documented Linux on earth. Five years of security updates per release, and the version every command here is written against.
Debian What Ubuntu is built on. Superb, a little more conservative, packages a little older. A fine choice for your second server, not your first.
Rocky / Alma The Red Hat lineage. Different package manager, different config paths. Right if your job already uses it; otherwise every command here needs translating.
Alpine Tiny and excellent inside a container. On a whole server it hands you library problems you do not yet have the vocabulary to search for.

Ubuntu gets a lot of hate. Look at the numbers anyway

There is a genre of internet opinion about Ubuntu — the package format argument, the telemetry argument, the it-used-to-be-better argument — and some of it is fair. None of it matters to you this week. The thing that matters is the one thing the criticism accidentally proves: it is the distribution everybody else already assumed you were running, and that is not a popularity contest, it is a compatibility guarantee.

Linux servers, by distribution

Of the Linux websites that publicly say which distribution they are.

⁠Ubuntu 65%
⁠Debian 26%
CentOS 5%
Everything else 4%

Source: W3Techs, Usage statistics of Linux for websites, read 8 September 2026 — Ubuntu 15.1%, Debian 5.9%, CentOS 1.2% of all Linux sites. The percentages above are those figures re-based on the 23.1% that identify themselves. Read the next paragraph before quoting any of this at anybody.

Three quarters of Linux servers do not say what they are. W3Techs can only count the ones that announce it, and 76.9% do not — partly because hiding it is a mild security habit, which means the chart above is measuring the servers run by people who did not turn that off. It is a real signal, not a census, and anybody showing you this chart without this paragraph is selling something. The direction is not seriously in doubt; the exact ratio is.
Developers, by distribution

Percentage of all respondents, so these do not add to 100 — people use more than one.

⁠Ubuntu 28%
⁠Debian 11%
⁠Arch 10%

Source: Stack Overflow Developer Survey 2025, personal use, from more than 49,000 responses. Professional use is almost identical for Ubuntu at 27.7%, which is its own small finding: people do not switch distribution when they clock in.

And it is where the AI tooling actually lives

This is the part nobody mentions in the distribution arguments, and for a beginner in 2026 it is probably the strongest single reason. The entire machine-learning stack — drivers, containers, the notebooks you will paste from, the images the clouds hand you — is built and tested on Ubuntu first. Not by decree. Just by everybody making the same assumption for fifteen years until it became true.

You are not picking the best Linux. You are picking the one whose commands are already written down in the answer you are about to search for. Which, at eleven at night, three hours in, is the only property that has ever mattered.

The practical consequence is small and constant. A README says apt install and you can paste it. An error message has ten thousand results instead of eight, and the eight are from 2017. A driver has a .deb and an Ubuntu version number beside it rather than a build script and an apology. None of that is a technical virtue of the operating system — it is a property of everyone else, and it is yours for free because you picked the same box in a dropdown that they did.

“LTS” is the only part of the name worth understanding. It stands for long-term support: that release gets security updates for five years without you doing anything, so the machine does not need reinstalling because a version went out of date nine months in. The non-LTS releases are perfectly good and are supported for nine months — exactly long enough to forget about it and then be surprised. If the provider offers both, the one you want is the even year, 24.04 or 26.04, not the highest number on the list.

None of this is tribal. Debian is superb, Alpine is a marvel in the place it belongs, and anybody telling you the choice defines you is selling something. The argument for Ubuntu here is narrow and practical: it is the one where the answer to your problem already exists in a form you can paste. That is worth more to you this week than any technical difference between them, and you can form an opinion in a year from evidence you gathered yourself.

The rest of the checkout

  1. Operating system: Ubuntu, the newest LTS As above. If you are also offered a choice of “minimal” or “server”, either is fine — the difference is a handful of preinstalled packages you can add later in one command.
  2. Location: near your users, not near you The distance between the server and whoever visits your sites is what adds milliseconds. If that is mostly you and your friends, pick the continent you are on and stop thinking about it.
  3. Skip every upsell Managed support, control panels, backup add-ons, extra IP addresses — you need none of it. You are going to set this machine up yourself, which is the point, and backups get handled properly later with something considerably better than a checkbox on an order form.
  4. Wait for the email Within a few minutes you will get a message containing an IP address (four numbers with dots, like 203.0.113.40) and a root password. That is the address of your computer and the key to the front door.
That email is now the most sensitive thing in your inbox. Anybody holding those two pieces of information owns the machine outright. Do not forward it, do not paste it into a chat, do not drop it in a document that syncs somewhere. The very first thing you do — in the Lock it down section — is replace that password with something nobody can guess at all. Until you have, treat the box as disposable, because it is.

Buying the domain

The one purchase where almost everybody gets quietly fleeced — and the one section of this guide where a lot of readers arrive already holding something.

A signpost with three blank arrow boards pointing in different directions, a small figure at its foot.
A domain is a name pointing at a number. Everything hard about it is paperwork, and almost all of that paperwork is optional.

Your server has a number. A domain is the name that points at it. You want one for the obvious reason — myproject.com reads better than 203.0.113.40 and is possible to say out loud — and for one much less obvious reason: you cannot get an HTTPS certificate for a bare IP address. No domain means no padlock, and a browser in 2026 will make a site without a padlock look like a crime scene.

From here the guide forks, because you are one of three people. Pick your door. All three of them end up in the same place — a name that answers with your server's address — and none of them takes an afternoon.

⁠Door one — buy it from Cloudflare

Nearly every registrar runs the same trick, and it is a good trick: a very cheap first year, then a renewal several times higher, on the entirely correct assumption that transferring a domain is irritating enough that you will look at the invoice, sigh, and pay it. They are not wrong about you. They were not wrong about me.

⁠Cloudflare Registrar does not do this. It sells domains at exactly what the registry charges, with no markup at all, forever. Privacy protection — which keeps your name and home address out of the public WHOIS database — is included rather than sold as an extra. The DNS control panel you need later is the same account.

The trap, in one table

These are real Cloudflare prices as of September 2026. The left column is what a typical registrar advertises. The right column is what you pay every year after the first.

DomainFirst yearEvery year after
.site$4.99$27.705.5× jump
.online$4.99$27.705.5× jump
.space$4.99$25.205× jump
.tech$9.99$49.205× jump
Read the renewal price, never the sale price. A domain is something you keep for years, so the only number that matters is the one in year two, and it is almost never the number on the banner. This is true even at Cloudflare, which adds no markup of its own: some registries run the first-year promotion themselves and then charge full freight forever after.

Good domains for under ten dollars a year

Same price in year one and in year ten. All of these are perfectly respectable. Nobody has ever decided a piece of software was bad because of the letters after the dot, and anybody who would is not going to be a useful user.

DomainPer yearWorth knowing
.fyi$5.20The cheapest genuinely normal-looking option. Great for tools and documentation.
.uk$5.30Short and clean. Fine for a personal project regardless of where you live.
.us$6.50Requires a US connection — citizen, resident, or a business operating there.
.link$7.20Reads well for anything tool-shaped, and short enough to type on a phone.
.cc$8.00Widely used, unrestricted, and never mistaken for a scam.
.org$8.50 → $11.20The one exception worth making: first year cheap, then a modest, honest renewal.
.click / .page / .day$10.20Flat forever. .page in particular looks deliberate rather than budget.
.com$10.46Just over the line, and rising to about $11.15 on 1 November 2026. Still the one people assume.

For reference, the ones people usually ask about: .net is $11.86, .xyz renews at $11.20, .dev is $12.20, .app is $14.20, and .io — the one every startup wants — is $50.00 a year, every year.

Avoid the four-dollar bargain-bin extensions. You will see .trade, .stream, .bid, .date and friends at around $4.18. They are cheap because they are overwhelmingly used for spam, and an enormous number of corporate mail filters and network blocklists have simply written off the entire extension. Your email silently vanishes, your links get blocked, and you spend a fortnight debugging a problem that is not in your code. Spending six dollars instead of four is the best money in this whole guide.

Buying it

  1. Make a Cloudflare account Free. Go to ⁠dash.cloudflare.com and sign up. Turn on two-factor authentication immediately — this account will control where your domain points.
  2. Domain Registration → Register Domain Search for the name you want. The price shown is the price, in year one and in year five.
  3. Turn auto-renew on A lapsed domain can be bought by anybody who wants it, and getting one back is expensive when it is possible at all. This is one checkbox standing between you and a genuinely miserable week.
  4. Do not touch DNS yet You will point it at your server in the Pointing a domain section, once there is actually something for it to point at.

That is door one finished. Skip to buying the server if you have not yet, or straight on to getting in. The next two headings are for people who arrived holding a domain already.

Door two — you already own one somewhere else

Then you do not have to buy anything and you do not have to move anything. A domain is not welded to the company that sold it to you: who you pay is one question, and who answers when somebody looks the name up is a completely separate one. Almost every beginner conflates those, panics about "transferring", and puts the whole project off for a month.

You have two ways to be live today, and one of them is also step one of door three.

What you doTakesCosts
2a. Add two records where you already are about five minutes nothing
2b. Move DNS to Cloudflare, keep the registration ten minutes, then a wait nothing

2a is the shortest path to a working site. Every registrar on earth has a DNS panel, every one of them can hold an A record, and yours is fine. Log in, find the DNS or "Manage DNS" screen, and add the two records from the Pointing a domain section. Nothing else about your account changes. Do this if you want to see your own site on your own name in the next quarter of an hour, which is a perfectly good thing to want.

2b is what I would actually do, and it is worth being clear about why, because "use Cloudflare's DNS" sounds like brand loyalty and is not. Moving DNS means changing two nameserver entries at your registrar so that Cloudflare, rather than your registrar, is the thing the internet asks. You keep paying the same company the same renewal fee. What you get is a DNS panel that changes in seconds instead of hours, wildcard records that actually work, the free proxy and analytics if you ever want them — and, not incidentally, the exact state Cloudflare requires before it will accept a transfer. Door two done properly is door three already half finished.

Who you pay and who answers the phone are two different questions. Nearly every beginner treats them as one, and loses a month to it.

Where your registrar hides the two screens

You need exactly two things from whoever you bought the domain from: the DNS records screen (for 2a) and the nameservers screen (for 2b). Every registrar has both, every one of them calls them something slightly different, and all of them move the menus about once a year. Find your logo.

⁠Namecheap
Among the friendlier ones. Both screens are on the same page, two tabs apart.
  1. Domain ListManage next to your domain.
  2. DNS records: the Advanced DNS tab. "Host Records" is the table; @ is the bare domain.
  3. Nameservers: the Domain tab, the Nameservers dropdown → Custom DNS, then paste Cloudflare's two.
  4. Auth code: Domain tab → turn off the transfer lock, then Sharing & TransferTransfer Out for the code.
⁠GoDaddy
Everything is three clicks further in than you expect, and the upsell interstitials are relentless. It all works.
  1. My Products (or Domain Portfolio) → the domain → DNS.
  2. DNS records: the DNS Records list on that page. Delete ⁠GoDaddy's parked-page record before adding your own, or you will have two.
  3. Nameservers: same page, the Nameservers panel → Change NameserversI'll use my own.
  4. Auth code: Domain SettingsAdditional SettingsTransfer domain away. Unlock first; the code is emailed.
⁠Porkbun
The best panel of the lot, and honest renewal pricing. If you are here, you may not need door three at all — check what yours actually renews at first.
  1. Domain Management → the arrow beside your domain to expand it.
  2. DNS records: DNS Records right there in the expanded row.
  3. Nameservers: the NS / Authoritative Nameservers field in the same row.
  4. Auth code: unlock the domain and the Auth Code is shown in the panel rather than emailed. Refreshingly civilised.
⁠Squarespace
This is where every Google Domains name ended up when Google sold the business. If you have not logged in since, that is why the branding changed under you.
  1. Domains in the account dashboard → the domain.
  2. DNS records: DNSDNS SettingsCustom Records.
  3. Nameservers: the same DNS screen → NameserversUse custom nameservers.
  4. Auth code: the domain's Transfer or Domain Lock panel. Unlock, then request the code.
⁠NameSilo
Plain, fast, and cheap already. Like ⁠Porkbun, worth pricing your renewal before you bother transferring.
  1. Manage My Domains → the domain.
  2. DNS records: the Manage DNS icon in the domain's row.
  3. Nameservers: tick the domain, then the Change Nameservers action above the list.
  4. Auth code: unlock in the domain's Registry Lock setting; the EPP code is on the same page.
⁠Gandi
A serious registrar with a serious panel. Everything is where a technical person would put it, which is not always where a beginner would look.
  1. Domain list → the domain.
  2. DNS records: the DNS Records tab.
  3. Nameservers: the Nameservers tab → External.
  4. Auth code: the Transfer section → unlock, then reveal the authorisation code.
⁠IONOS
The panel is built for people buying hosting bundles, so the DNS screen is buried under the product rather than under the domain.
  1. Domains & SSL → the domain → the gear or menu.
  2. DNS records: DNSAdjust DNS settings.
  3. Nameservers: the same menu → NameserverUse own nameserver.
  4. Auth code: the domain's menu → Transfer domain / Request auth code.
Somebody else
Hover, Dynadot, Name.com, OVH, Cloudflare-adjacent resellers, the hosting company that sold you a website in 2019 — the shape is identical everywhere.
  1. Find the page that lists your domains and open the one you want.
  2. DNS records is behind a link saying DNS, Zone, Zone File, Advanced DNS or Manage DNS.
  3. Nameservers is behind a link saying Nameservers, NS or Delegation, and the option you want is the one called custom, external or third-party.
  4. The auth code is behind a link saying Transfer, Transfer Out, EPP, or Authorisation Code — and it is always greyed out until you turn the transfer lock off.

If you genuinely cannot find one of them, search the registrar's help centre for the words "EPP code" rather than for "transfer": the help article you want is the one written for people leaving, and it is usually more direct than the one written for people staying.

Not everybody should do door three. If you are already at Porkbun, ⁠NameSilo, ⁠Namecheap or somewhere else with an honest flat renewal, the saving is a couple of dollars a year and the transfer is not worth your Saturday. Look up what your domain actually renews at — the number on the renewals page, not the one you paid in year one — and if it starts with a small digit, stop at door two and go and build something. The point of this section is never to pay a 5× renewal by accident, not to move every domain on earth to one company.

⁠Door three — moving the registration to Cloudflare

A small figure carrying a suitcase from a closed shop front on the left towards an open archway on the right.
A transfer moves the billing relationship and nothing else. The name keeps working the entire time, provided you did door two first.

Worth saying plainly before you start: a domain transfer does not take your site down. The records are already at Cloudflare by this point — that is what door two was — and the transfer only changes which company bills you. The fear that it will break something is the single biggest reason people keep paying a renewal five times what the name is worth.

A transfer also is not free, exactly. You pay one year's registration, and that year is added to the time you already have rather than replacing it. So transferring a name with eight months left buys you twenty months. Nothing is lost.

Before Cloudflare will take it

Four conditions, all of them ICANN's or Cloudflare's rather than anything you can argue with. Check them first, because failing one at step six is much more annoying than reading them now.

ConditionWhy, and what to do
Registered 60+ days ago An ICANN rule against fraud, not a Cloudflare one. A brand-new domain simply has to wait. Same if it was transferred in the last 60 days.
No recent contact change Changing the registrant's name or email can start its own 60-day lock at some registrars. If you are about to do both, transfer first.
Already on Cloudflare DNS Cloudflare Registrar only holds names whose DNS it also serves. This is door two, and the zone has to read Active before the transfer option appears.
DNSSEC turned off If your registrar signed the zone, disable DNSSEC there and give it a day before changing nameservers. Skipping this is the classic way to take your own domain down mid-transfer.

The transfer itself

  1. Add the domain to Cloudflare and switch the nameservers Cloudflare reads your existing records and copies most of them for you. Check the copy against the old panel before you switch — it is good, not perfect, and MX records are the ones it most often needs help with.
  2. Wait for the zone to say Active Usually minutes, occasionally a day. Nothing below this line is available until it does, and refreshing does not help.
  3. Unlock the domain at the old registrar The setting is called the transfer lock or clientTransferProhibited. Turning it off is not dangerous; it is a seatbelt, and you are the one driving.
  4. Get the authorisation code Also called an EPP code, auth-info code or transfer code. Some registrars show it on screen, some email it to the address on the registration — which is another reason that address has to be one you can actually read.
  5. In Cloudflare: Domain Registration → Transfer Domains Pick the domain, paste the code, pay the one year, and confirm the contact details. Cloudflare's WHOIS privacy is on by default and included.
  6. Approve it, then leave it alone The old registrar emails you asking to confirm. Approving speeds it up; ignoring it still works, because the transfer completes on its own after five days unless somebody actively refuses it. Then turn auto-renew on in the new place, which is the step everybody forgets.
Cloudflare Registrar does not sell every extension. It supports a long list of common ones and a lot of country codes, but not all of them — and if yours is not on it, no amount of clicking will produce the transfer button. That is not a problem you have to solve: door two costs nothing, works with any registrar in the world, and gives you the same DNS panel. Check the list before you start unlocking things.
The thing I got wrong: I let a domain expire. Not a client's, thankfully — one of mine, with things linked to it. The card on file had rolled over to a new expiry date, the renewal failed quietly, the warning email went to an address I had stopped reading, and by the time I noticed, the name was in somebody else's shopping basket. There is no support ticket for that. Auto-renew, on a card that works, on an address you actually open. It is the cheapest insurance on this page — and a transfer is exactly the moment it gets switched off without you noticing, because auto-renew does not come across with the domain.
02

Make it safe

Before anything fun. Twenty minutes, once, and then never again.

The same creature standing in front of a door, holding a closed padlock.

Getting in

You already have everything you need. Yes, on Windows. No, you do not need to download anything.

A single door standing alone in empty space, slightly ajar, a small figure beside it holding a key up.
There is exactly one door and it is already installed. The next section is entirely about changing the lock.

A terminal is a window where you type commands instead of clicking things. That is the entire concept. It is intimidating for about four minutes, and then it is the fastest thing you have ever used, and then you start resenting software that will not let you type at it.

SSH is how you get a terminal on a computer that is somewhere else. It is encrypted, it is ancient, it has been attacked by everybody for thirty years and it is still standing, and it is already built into Windows 11 and ⁠macOS. You do not need PuTTY. You do not need WSL for this. You do not need to install one single thing. Anyone telling you otherwise learned this in 2009 and has not looked since. (WSL is worth having for an entirely different reason, later.)

Opening a terminal

Press Win and type terminal, then open Windows Terminal. It opens on PowerShell, which is exactly what you want. Check that SSH is present:

ssh -V

You should see something like OpenSSH_for_Windows_9.x. If the command is not found, open Settings → System → Optional features, click Add an optional feature, and add OpenSSH Client. One reboot and it is there for good.

Press Space, type terminal, hit return. SSH has been installed since forever. Confirm it:

ssh -V

Optional: a nicer terminal

Everything in this guide works in the terminal you already have, and I would rather you started there — one fewer download between you and the thing you actually came to do. But the built-in terminal is a fairly plain box, and once you are spending real hours in it you may want something with better manners. Warp is the one worth knowing about. It runs on Windows, macOS and Linux, it keeps each command and its output in a separate block you can scroll to and copy cleanly, and text editing in it behaves the way text editing behaves everywhere else in your life, which is a lower bar than the terminal has historically cleared.

Windows 11 ships with winget, a package manager, so this is one line in the terminal you already opened. There are x64 and ARM64 installers on the site if you would rather click:

winget install Warp.Warp

Then find it in the Start menu. It needs Windows 10 build 18362 or newer, which in practice means anything you are likely to be running.

With ⁠Homebrew, or from the download page if you do not have it:

brew install --cask warp
Skip the sign-in, and skip the AI tier. Warp will offer you an account on first launch and you can decline it — the terminal works signed out. It also sells its own AI assistant, with paid plans starting around $20 a month as of September 2026. You do not need it. You are already paying for ⁠Claude Code, which is doing the same job with a whole server underneath it rather than one shell. Use Warp as a terminal, on the free tier, and let the agent be the agent.
Every command in this guide is identical either way. A terminal is a window that shows you a shell; changing the window does not change the shell, and nothing below depends on which one you picked. If Warp is not to your taste, or you would rather not install anything, close this bit and carry on — Windows Terminal does all of it.

Your first connection

Take the IP address and root password from the provider's email. Replace 203.0.113.40 with your own address everywhere in this guide.

ssh root@203.0.113.40

The first time, you will be asked something like this:

The authenticity of host '203.0.113.40' can't be established.
ED25519 key fingerprint is SHA256:iRk3z...
Are you sure you want to continue connecting (yes/no/[fingerprint])?

Type yes and press enter. Your computer is saying I have never spoken to this machine before and I have no way of knowing it is really the one you meant. It writes down the server's fingerprint so that if it ever changes it can shout at you about it. You see this once per server and then never again, and on the day you do see it again you should stop and find out why.

Then paste the root password. Nothing appears as you type — no dots, no asterisks, not so much as a flicker. Your keyboard is fine, your paste worked, the terminal is simply declining to tell an onlooker how long your password is. Press enter.

Welcome to Ubuntu 26.04 LTS (GNU/Linux x86_64)
root@your-server:~#

That prompt is a computer in a data centre waiting for you to say something. Everything you type from here runs there, not here. Try a few harmless things just to feel it:

# who am I, and where am I?
whoami
pwd

# what is this machine made of?
free -h
df -h
nproc

# how long has it been switched on?
uptime
Nothing here can hurt your laptop. Your machine is drawing text and nothing else. Every command runs on the server. The worst case is that you destroy the server, at which point the provider reinstalls it from scratch in two clicks and you are back where you started twenty minutes later, slightly wiser. That is not a risk, it is the feature. Type exit to come home.

The bits nobody explains

Every guide skips these, because whoever wrote it learned them so long ago that they have stopped being knowledge and turned into instinct. They are the actual reason beginners bounce off a terminal, so here they are, in order of how much time they will save you.

The prompt is telling you four things.

sam@web-01:~$
│   │      │ └── $ means an ordinary user. A # means you are root:
│   │      │      be awake.
│   │      └───── where you are. ~ is your home folder.
│   └──────────── which machine. Check this before anything destructive.
└──────────────── who you are on it.
Pasting, which is where most people get stuck
In Windows Terminal, Ctrl+V pastes and so does a plain right-click. What you must not do is reach for Ctrl+C to copy while something is running — in a terminal that means stop what you are doing, and it has meant that since long before it meant copy. To copy, select the text and use Ctrl+Shift+C, or just select it and right-click.
Tab completes, remembers
Type the first few letters of a file or command and press Tab; the shell finishes it, and beeps or shows you the options if it is ambiguous. The up arrow walks back through everything you have typed. A surprising amount of being fast at this is simply not typing — people who look quick at a terminal are mostly pressing two keys.
~ is home, / is everything
/ is the top of the machine and every file lives somewhere beneath it. ~ is shorthand for your own folder, which is /home/sam for a user called sam. cd moves you, pwd tells you where you ended up, and cd with nothing after it always takes you home.
Getting out of something
Ctrl+C stops whatever is running and gives you the prompt back. It is the fire escape and you should use it freely; it is not destructive. exit — or Ctrl+D — closes the SSH session and returns you to your own machine. If a program has taken over the whole screen and ignores both, it is usually an editor, and it usually wants q.
Nothing scrolls away permanently
You can scroll up. The output of the last hour is still there. When something goes wrong, the answer is nearly always four lines above where you are looking, in the part you skipped because it had not gone wrong yet.

nano, and the editor you did not ask for

Every file this guide asks you to change is opened with nano, and every guide tells you the same two keys and stops: Ctrl+O saves, Ctrl+X quits. That is enough to paste a config in. It is not enough for the day an error message says line 212 of a file with four hundred lines in it. The two rows at the bottom of the screen are the rest of the manual — ^ means Ctrl and M- means Alt — and these are the five worth knowing before you need them:

Ctrl+W finds
Type a word, press enter, and the cursor lands on it. Alt+W jumps to the next one. In a long nginx config this is how you reach the line you came for.
Ctrl+/ goes to a line
Type the number an error message just gave you. Alt+G is the same thing, and nano +212 file opens the file already there.
Ctrl+K cuts a line, Ctrl+U puts it back
Press Ctrl+K several times in a row and the lines travel together, so cut, move, paste is how a block moves. It is also how a line is deleted: cut it and never paste.
Alt+U undoes
As many steps back as you like; Alt+E redoes. Nothing touches the disk until Ctrl+O, so a file you have made a mess of is also fixed by Ctrl+X and answering N.
The questions it asks
Ctrl+O shows the filename and waits: press enter to keep it. Ctrl+X on a changed file asks whether to save: Y then enter, or N to throw the changes away. Ctrl+C backs out of any question it asks.

Then there is the day a screen opens that is not nano. No rows at the bottom, and typing either does nothing or puts letters where you did not aim them. That is vi, and you are in it because some program opened its own editor rather than yours. It will not be the server: on the Ubuntu this box runs, every program falls back to nano, and vi is not installed at all. It will be your laptop. A Mac's git commit, typed without -m, opens vim, and Git for Windows offers vim as its editor during install, so the first time you forget the -m you are in it. Getting out is three keys: Esc, then :q!, then enter, which leaves without saving, and at that moment saving was not the plan anyway. Then tell every program which editor you meant, and it never happens again:

# Mac: ~/.zshrc.  Linux, or Ubuntu inside WSL: ~/.bashrc.  Then open a new terminal.
export EDITOR=nano

# PowerShell on Windows has no nano; point git at Notepad instead.
git config --global core.editor notepad

vi is not a mistake, for what it is worth. It is from 1976, when the screen was the expensive part, and every one of its habits is a habit for a machine that cannot afford to draw a menu. People who learned it then still fly in it. You are not obliged to, and nothing in this guide will put you in it on purpose.

Stop typing the address

You are going to type ssh sam@203.0.113.40 several hundred times, and you are going to mistype it. SSH reads a config file on your machine where you can give the whole thing a short name. Make it early; it is one of those small things that quietly removes friction for years.

# Windows:      C:\Users\you\.ssh\config
# macOS, Linux: ~/.ssh/config

Host box
    HostName 203.0.113.40
    User sam

From then on, the whole command is:

ssh box

Add a Host block per machine and the names are yours to choose. Once you have made a key in the next section, one more line in that block points at it, and you can stop thinking about any of this permanently.

When it will not connect

SSH has four failure messages that beginners read as one big "it is broken". They are not the same thing at all, and each one tells you precisely how far you got, which means each one has a different fix.

What it saysHow far you gotWhat it usually is
Connection timed out Nothing answered at all. Wrong IP address, the machine is not running yet, or a firewall is dropping you. New servers can take a couple of minutes to actually boot. Check the address against the provider's email, character by character.
Connection refused Something answered, and said no. You have the right machine but nothing is listening on that port. SSH is not running, or it is on a different port. This is the good failure — it means the address is right.
Permission denied (publickey,password) All the way to SSH. It does not like you. Wrong username, wrong password, or a key it has never seen. Note the username: root@ and sam@ are different people and only one of them exists yet.
Host key verification failed All the way, and it stopped on purpose. The machine's fingerprint is not the one your computer remembers. Almost always because the server was rebuilt. Almost — so find out why before you clear it, because the other explanation is that you are not talking to your server.
Add -v when you are stuck. ssh -v sam@203.0.113.40 narrates every step it takes, and the last line before it gives up is the one that matters. If you cannot make sense of it, that output is also the single best thing to paste at an agent — it is precise, it is short, and it describes a real machine rather than a feeling.

Lock it down

Do this before you build anything. It is not optional, it is not advanced, and it takes twenty minutes.

A barred gate with a large closed padlock on it, a small figure standing to one side.
Twenty minutes of work, done once. Everything after this point quietly assumes you did it.

Your server has a public address, and within minutes of it existing, automated scanners find it and start trying passwords. This is not paranoia and it is not personal — it is the background radiation of the internet, and it is aimed at everybody equally. On the machine serving this page there were 47 failed login attempts in a single day before it was locked down, and it had been online for a matter of hours.

Six steps, in order. Afterwards, guessing your way in stops being slow and starts being arithmetically impossible, which is a much better place to be.

One thing worth saying before you start, because it saves you years of anxiety later: most of what gets called a "vulnerability" on the internet is not the kind of thing that happens to you. A great many of them read like well, if the attacker were already inside your computer, and were also very small, he could easily bend your CPU pins. The correct response to that is not a patch. It is asking why there is a two-foot faerie rattling around inside the case in the first place. The six steps below are about the front door, which is the part that actually gets tried, every single day, by everybody.

1. Stop being root

root is the account that can do absolutely anything, with no confirmation and no undo. Living as root means every typo is a potential catastrophe and the logs cannot tell you who did what, because the answer is always "root". Make yourself an ordinary account that can become root when it has a reason to.

# Replace "sam" with whatever you want to be called.
adduser sam

# Put yourself in the sudo group, which grants the right to
# temporarily run a single command as root.
usermod -aG sudo sam

adduser asks for a password — make it long and put it in a password manager — and then asks for your full name, room number and phone number. Press enter through all of those. They are a relic of shared university machines and nobody has read one this century.

From now on, when a command needs root powers you put sudo in front of it and type your password. To get a full root shell for a while, use sudo -i.

2. Make a key, and use it instead of a password

This is the one that matters. An SSH key is a matched pair of files. The public half goes on the server and is not a secret in any sense — you could put it on a billboard. The private half stays on your laptop and never moves anywhere, ever. Logging in proves you hold the private half without ever sending it.

why a key beats a password Four messages. The interesting part is what is missing from all four.
  1. Your laptop names which key it would like to use. Not the key itself — which one.
  2. The server looks in authorized_keys, finds the public half, and sends back a number it has never sent before.
  3. Your laptop signs that number with the private half, which has not moved and will not move.
  4. The server checks the signature against the public half it already had, and opens the door.
  5. So nothing secret crossed the wire, and nothing on the server is worth stealing: authorized_keys can verify you and can never impersonate you.
A password is a secret you send. A key is a secret you prove you have without sending it.

A password can be guessed. A key cannot. The space to search through is so large that brute force stops being a strategy and becomes a genre of fiction.

Generate the key on your own computer, not on the server. A private key is created where it is going to live and it never travels. If somebody hands you a private key over chat or email, that key has been read by everything that carried it, and it is not private any more — it is just a file.

On your laptop, in a new terminal window — not the one connected to the server:

# Make the key. Press enter to accept the default location.
# When it asks for a passphrase, type one — it encrypts the key file
# itself, so a stolen laptop is not a stolen server.
ssh-keygen -t ed25519 -C "sam@laptop"

# Copy the PUBLIC half up to the server. scp is a file copy over ssh.
ssh sam@203.0.113.40 "mkdir -p ~/.ssh && chmod 700 ~/.ssh"
scp $env:USERPROFILE\.ssh\id_ed25519.pub sam@203.0.113.40:~/newkey.pub
ssh sam@203.0.113.40 "cat ~/newkey.pub >> ~/.ssh/authorized_keys && chmod 600 ~/.ssh/authorized_keys && rm ~/newkey.pub"

Windows has no ssh-copy-id, and piping the file straight into ssh from PowerShell can quietly corrupt it with Windows line endings and a byte-order mark. Copying the file with scp and appending it server-side, as above, avoids the whole problem.

# Make the key. Press enter to accept the default location, and
# choose a passphrase when asked.
ssh-keygen -t ed25519 -C "sam@laptop"

# macOS has a purpose-built tool for installing it.
ssh-copy-id sam@203.0.113.40

# Optional: let the macOS keychain remember the passphrase.
ssh-add --apple-use-keychain ~/.ssh/id_ed25519

Now test it, and do not skip this. Open a new terminal and log in:

ssh sam@203.0.113.40

If you are asked for your key passphrase, that is correct — that prompt comes from your own laptop. If you are asked for your server password, the key is not working, and you must fix that before continuing. Stop here until a passphrase-only login works.

This is also the moment to finish the Host box block you made back in Getting in. Add one line to it — IdentityFile ~/.ssh/id_ed25519 — and ssh box now means the right machine, the right user and the right key, forever.

3. Turn passwords off entirely

Now that keys work, close the door behind you. On the server:

sudo nano /etc/ssh/sshd_config.d/10-hardening.conf

Paste this in, then press Ctrl+O, enter, Ctrl+X:

# Keys only. Password guessing is now impossible rather than slow.
PasswordAuthentication no
KbdInteractiveAuthentication no
PubkeyAuthentication yes

# root can still be reached with a key, never with a password.
PermitRootLogin prohibit-password

# Three tries and twenty seconds, not six and two minutes.
MaxAuthTries 3
LoginGraceTime 20
X11Forwarding no

# Keep long connections alive through home-router NAT.
ClientAliveInterval 30
ClientAliveCountMax 6
Why the filename starts with 10-, and why it matters. SSH reads that directory in alphabetical order and keeps the first value it finds for each setting. Nearly every cloud image ships a 50-cloud-init.conf containing the single line PasswordAuthentication yes. Call your file 99-hardening.conf and it is silently ignored: your config looks immaculate, you feel great about it, and passwords still work. This exact trap caught the setup of the very server you are reading this on, which is why there is a verification command two paragraphs down. Never trust the file. Ask the daemon.

Check the config parses, apply it, and then verify what the server actually believes:

sudo sshd -t                # silence means the file is valid
sudo systemctl reload ssh

# The real test. This prints the settings in force, not the ones you wrote.
sudo sshd -T | grep -iE 'passwordauth|permitrootlogin|pubkeyauth'

You want to see exactly this:

permitrootlogin prohibit-password
pubkeyauthentication yes
passwordauthentication no
That command, on the machine serving this page. Real output, real timings — only the typing speed is staged.
Keep your current session open while you test. Reloading SSH never drops connections that already exist. So open a second terminal and log in with your key. If it works, the first one is safe to close. If it does not, you still have a live session to undo the change from, and that session is the only thing standing between you and a reinstall. Every provider also has a browser-based console that bypasses SSH entirely. Find that button today, while it is a curiosity rather than an emergency.

4. Close every port except the three you use

A firewall decides which doors on the machine are open to the internet. ⁠Ubuntu ships one called ufw that is genuinely simple.

sudo ufw allow OpenSSH   # 22  — your terminal. Allow this FIRST.
sudo ufw allow 80/tcp    # 80  — plain web, needed for certificates
sudo ufw allow 443/tcp   # 443 — encrypted web, the real one
sudo ufw enable

sudo ufw status verbose
Allow SSH before you enable the firewall. Enabling ufw with a default-deny policy and no SSH rule locks you out of your own server instantly and completely. It is the single most common way people lose a box on day one. The order above is correct — follow it exactly.

Notice what is not on that list. Nothing for a database, nothing for the application itself, nothing for anything you are about to build. That is deliberate, and it is the single most useful habit in this section: anything only your own programs talk to listens on 127.0.0.1, which is not reachable from outside the machine at all. The Databases section does this properly. A door that does not exist cannot be picked.

5. Ban the scanners automatically

fail2ban reads the logs and blocks addresses that keep failing. With keys-only login nobody was getting in regardless, so this is not really about security — it is about not having your logs buried under thousands of identical failures, so that the one interesting line is still visible when you go looking for it.

sudo apt install -y fail2ban
sudo nano /etc/fail2ban/jail.local
[DEFAULT]
bantime  = 1h
findtime = 10m
maxretry = 5

# Repeat offenders get progressively longer bans, up to a week.
bantime.increment = true
bantime.factor    = 2
bantime.maxtime   = 1w

# Ubuntu logs SSH to the systemd journal, not to a file.
backend   = systemd
banaction = ufw

[sshd]
enabled = true
mode    = aggressive
sudo systemctl enable --now fail2ban
sudo fail2ban-client status sshd

6. Actually turn security updates on

Ubuntu installs unattended-upgrades by default, and everybody assumes that means updates are happening. On a great many cloud images they are not. The service is running, it is enabled, it reports itself as healthy, and the periodic trigger that tells it to actually do something is set to zero. It has been enabled and doing nothing for months.

# Check what yours says. A "0" here means nothing has ever been installed.
cat /etc/apt/apt.conf.d/20auto-upgrades
sudo nano /etc/apt/apt.conf.d/20auto-upgrades
APT::Periodic::Update-Package-Lists "1";
APT::Periodic::Unattended-Upgrade "1";
APT::Periodic::AutocleanInterval "7";

Then catch up on everything that was missed while it was switched off:

sudo unattended-upgrade -v

# If a new kernel was installed, the machine needs a restart to use it.
# This file only exists when a reboot is genuinely pending.
cat /var/run/reboot-required 2>/dev/null && sudo reboot

Switching it on is step one of two. Step two is knowing what it does not cover, how to check it is really installing things, and when the software you are running stops being fixed at all — which is Staying up to date, in part six.

This same check found 78 uninstalled security updates on this box. Months of them, on a machine whose update service was enabled, running, and set to 0. Two lines in a text file, and almost nobody ever opens it. Assume nothing on a server is happening until you have seen the output that proves it.

apt, which you have now used without being told what it is

Steps five and six ran apt and edited its settings without stopping to introduce it, so, briefly: it is how software gets onto this machine. It is what Windows does with a download page, a setup wizard and a tray icon nagging about updates, done once, for everything. Ubuntu keeps an archive — the one this box reads lists about 78,000 packages — built by Ubuntu, signed by Ubuntu, and checked against that signature on the way in. apt install fetches from it. So does the nightly security run you just switched on. Every install in this guide comes from there unless it says otherwise, and that "unless" is what the notice below is about.

sudo apt update                  # refresh the list of what exists; installs nothing
sudo apt upgrade                 # install newer versions of what you already have
apt search --names-only sqlite   # is there a package called something like this?
apt show sqlite3                 # what is it, how big, what comes with it
sudo apt install sqlite3         # get it. -y answers the yes/no question for you
sudo apt remove sqlite3          # drop it. purge instead also removes its config
sudo apt autoremove              # and whatever was only there to support it
apt list --installed             # everything on here. This box: 1,114 packages

The pair everybody gets backwards is the first two. update updates nothing; it downloads the list of what is available, and that is all. It is upgrade that installs. The failure that teaches this is E: Unable to locate package for something that certainly exists, and it means the catalogue on your box is older than the package, which on a machine made yesterday is very likely. Run update first, then look again. The two go together so often that the one-liner in Linux on your laptop joins them with &&. And apt-get, which every guide older than 2014 uses instead, is the same tool with a face for scripts; apt is the face for people. Either works. Read apt-get as a sign of a guide's age, not of its wrongness.

Somebody else's archive is somebody else running as root on your machine. Sooner or later a guide will say add-apt-repository ppa:somebody/something to get a newer version of a thing. A PPA is one person's build, hosted on Launchpad, and installing from it means their packages run as root on your box, every time they update, for as long as the line stays in your sources. That is not a vulnerability. It is the arrangement, and it is the same arrangement you already have with Ubuntu; the difference is that Ubuntu has a security team, a release process and a name to lose, and ppa:somebody has somebody. Three honest reasons to step outside the archive: the software is not in it, the version in it is too old for you, or the vendor publishes their own repository, which is a PPA with a company standing behind it that you can hold responsible. This box has exactly one of those, ⁠GitHub's own, for gh: a file in /etc/apt/sources.list.d/ and GitHub's signing key in /etc/apt/keyrings/. Know what each line in that folder is for, because the release upgrade in two years' time switches every one of them off and asks you what to do, and "I do not remember" is the wrong answer to have ready. Ubuntu also ships snap, a second way to install things, with its own opinions. Nothing in this guide needs it.

The things that actually go wrong

Break-ins on small servers are tediously repetitive. Nobody is writing a novel exploit for your side project. Almost every lost box is one of these five, and you have just handled the first two. What to do on the day one of them happens anyway is in the backups section, because that is where the answer lives.

How people lose a box
  • A guessable password on an account exposed to the internet.
  • Root login left enabled with a password.
  • A database bound to 0.0.0.0 with a weak password — scanners find these in hours.
  • Secrets committed to a public ⁠GitHub repository. Bots scrape new commits within minutes.
  • A server never updated because nobody checked that automatic updates were running.
What you just did about it
  • Passwords no longer work over the network at all.
  • Root can only be reached with a key you hold.
  • Everything internal will listen on 127.0.0.1, invisible from outside.
  • .gitignore and environment variables — covered in GitHub.
  • Security updates install themselves, and you verified it rather than assumed it.
Once ⁠Claude Code is installed, it can audit all of this for you. "Check this server's SSH configuration, firewall rules, running services, listening ports and automatic updates. Tell me anything that is exposed or misconfigured, worst first." It is a genuinely good second pair of eyes and it reads config files with a patience you will never have at eleven at night.
The thing I got wrong: I used to think moving SSH to port 2222 was security. It is not. It is a noise filter, and a good one — the mass scanners knock on 22 and move on, so your logs get quiet and you feel like a genius. But the person who actually wants in scans all sixty-five thousand ports in under a minute and finds you anyway. Nothing in this section is obscurity. Keys have no guessable form, closed ports have nothing behind them, and installed updates have no hole left to exploit. Move the port too if you like the quiet. Just never mistake the quiet for a lock.

Your server as a drive

Optional, and genuinely lovely. Browse the server in Explorer or Finder.

A laptop at the left and a server cabinet at the right joined by one long looping cable, a small figure sitting on the loop.
One cable, in the sense that matters: the files are over there and the window is over here.

The argument of this guide is that you should stop living in a file explorer. That is not the same as never wanting to look at one. Sometimes you want to see a folder, drag a photo in, or open a file in an editor with fonts in it, and none of that requires giving up the terminal as the place where the work happens.

So mount the server as an ordinary drive. It rides on the SSH connection you already have, which means nothing new to secure, no new port, and no new password to lose.

SSHore is a small Windows 11 app that turns a saved SSH connection into a drive letter. Install it, add a profile with your server address, username and private key, and the whole filesystem appears in Explorer. It sits in the tray, remembers your profiles, mounts in one click, and stores credentials with Windows DPAPI rather than in a text file.

  1. Install WinFsp and SSHFS-Win These are the two components Windows needs to understand a remote drive. SSHore's installer points you at both.
  2. Add a profile Host 203.0.113.40, user sam, and your private key at C:\Users\you\.ssh\id_ed25519. Point the remote path at /srv — where all your projects will live — or at / for the whole machine.
  3. Mount it It becomes a drive letter. Open it in Explorer, drag files in and out, open them in any editor you like.

In Finder, press K and connect to sftp://203.0.113.40. It works with no installation at all, though it is read-mostly and a little slow.

For a proper read-write mount, install macFUSE and sshfs, then:

mkdir -p ~/server
sshfs sam@203.0.113.40:/srv ~/server -o volname=server,follow_symlinks
Mount it to look, not to build. A network drive is much slower than local disk, and running a compiler or a package installer across one is a special kind of misery — thousands of tiny files, each one a round trip. Let the agent work on the server, where the files already are. Use the mount for dragging things in and out and for opening the occasional file in something graphical.

⁠A second Linux, on your own laptop

Optional, free, and the best twenty minutes in this guide if you are on Windows. Not a replacement for the server — a practice room next door to it.

A large server cabinet on the right and a much smaller cabinet of the same shape on a desk at the left, a small figure standing between them.
The same operating system, twice: one rented and permanently on the internet, one on your desk that nobody can reach and nothing depends on.

Everything so far has been about one computer, the one you rent. This section is about putting the same operating system on the machine in front of you as well, so that the commands in this guide have somewhere to be wrong. It is genuinely optional. Skip it and nothing later breaks. But a beginner's biggest practical problem is that the only Linux they have is the one their website is running on, which makes every experiment feel like defusing something.

You will learn Linux roughly at the speed you are willing to break it. That speed is near zero on a machine that is serving your site.

What WSL actually is

WSL is the Windows Subsystem for Linux. Despite the name it is not an emulator or a compatibility shim: modern WSL runs a real Linux kernel in a very light virtual machine that Windows manages for you, with ⁠Ubuntu inside it — the same Ubuntu that is on your server. There is no dual boot, no partitioning, nothing to configure, and no decision at startup about which operating system you are in today. It is a program you open, and inside it is Linux.

It is built into Windows 11. You are not downloading a third-party tool, and this does not contradict what Getting in says — you still do not need WSL to reach your server, because Windows has had a perfectly good ssh for years. This is something else: not a way in, a second machine.

Why bother, when you already have a server

Somewhere to break things Destroy a WSL Ubuntu completely and it is back in two minutes. Try that on the box answering your domain.
Your laptop matches your server Same shell, same apt, same paths, same permissions. One set of habits instead of two.
The agent works locally too ⁠Claude Code runs in WSL on your own files, at local-disk speed, with no round trip to the data centre.
Fast, free iteration Compile and test a hundred times without touching production, then deploy the version that works.

The last one is the real argument. A server is for running things. A laptop is for trying things. When both of them are Ubuntu, the distance between the two is a git push rather than a translation exercise, and the "works on my machine" conversation you have heard people complain about for twenty years mostly stops happening.

WSL is not a substitute for the server, and this is not a small distinction. It has no address on the internet, no domain, no certificate, and it stops existing when your laptop sleeps. Nobody else can reach it — not your users, not a webhook, not you from your phone. Anything the world is supposed to be able to load lives on the VPS. WSL is where you find out whether it works before anybody is watching.

Installing it

Open PowerShell as Administrator — right-click the Start button, then Terminal (Admin) — and run one command.

wsl --install

That turns on the two Windows features it needs, fetches the kernel, installs Ubuntu, and tells you to reboot. Reboot. On the way back up, an Ubuntu window opens by itself and asks you to invent a username and a password.

That username and password are local, and they are not your server's. Pick the same username if you like consistency; do not reuse the server's password. This one only ever unlocks sudo on your own laptop, and it is fine for it to be something you can type quickly. The server's password is a different class of secret and is dealt with in Lock it down.

Two variations worth knowing, both run from the same PowerShell window:

# See which distributions are on offer before committing
wsl --list --online

# Ask for a specific Ubuntu rather than whatever is current
wsl --install -d Ubuntu-24.04

# If WSL was already installed years ago and is stale
wsl --update

Take the Ubuntu LTS that matches your server if you can. It is not a requirement and nothing breaks if the versions differ, but matching them removes an entire category of confusing afternoon: a package that exists in one and not the other, a default that changed between releases, a config file that moved.

The first five minutes inside it

You are now at an Ubuntu shell that behaves exactly like the one on your server, so the first thing you do is the same thing you did there.

sudo apt update && sudo apt upgrade -y

From here, every ⁠macOS tab in this guide is your tab. That is worth saying explicitly, because it looks wrong: the Windows tabs on this page are written for PowerShell, and PowerShell is not what you are in any more. Inside WSL you are on a Unix shell, so the macOS commands — the ones with ssh-keygen and ~/.ssh/config and forward slashes — are the ones that apply. apt replaces brew and everything else lines up.

Where your files actually are

This is the one part of WSL that confuses everybody, and getting it right on day one saves a genuinely miserable week later. There are two filesystems and they are not equally fast.

PathWhat it isSpeed
/home/you The Linux filesystem. Where your projects belong. fast
/mnt/c/… Your Windows C: drive, mounted so you can reach it. slow
\\wsl$\Ubuntu The Linux filesystem, seen from Windows Explorer. fine
Keep your code in /home/you, never in /mnt/c. Every read and write under /mnt/c crosses between two operating systems, and a build reads thousands of small files. Put a project on /mnt/c and a compile that takes four seconds takes ninety, an npm install takes a coffee break, and you conclude that WSL is slow. It is not; you asked it to do the one thing it is bad at. This is the single most common WSL mistake and it is invisible until you have made it.

To get at your Linux files from Windows — to drag one into an email, or open it in something graphical — type \\wsl$\Ubuntu into the Explorer address bar, or run explorer.exe . from inside WSL to open the folder you are standing in. Both are the correct way round: reach into Linux from Windows, rather than keeping the files on Windows and reaching out.

The terminal to use

Windows Terminal is already on Windows 11 and it grows an Ubuntu tab by itself the moment WSL is installed — the down-arrow next to the plus sign in the tab bar. That is the whole setup. You can pin it, split it, and have PowerShell in one pane and Ubuntu in another, which is a genuinely nice way to work: Windows things on the left, server-shaped things on the right.

If you preferred ⁠Warp from the Getting in section, it drives WSL too. Nothing here depends on which terminal you use, and switching later costs nothing.

Two things that will catch you out

You now have two SSH identities, in two places. Keys you made in PowerShell live in C:\Users\you\.ssh. Keys you make inside WSL live in /home/you/.ssh, which is a completely separate directory. WSL will not find the Windows ones, so an ssh mybox that worked yesterday in PowerShell says Permission denied today in Ubuntu, and it looks like the server broke. It did not. Either make a second key inside WSL and add it to the server the same way — which is cleaner, because a key should belong to one place — or copy the existing one across and fix its permissions:

# Only if you would rather have one key than two
mkdir -p ~/.ssh && cp /mnt/c/Users/you/.ssh/id_ed25519* ~/.ssh/
chmod 700 ~/.ssh
chmod 600 ~/.ssh/id_ed25519

# ssh refuses a private key that other users could read. This is why.
ls -l ~/.ssh

And systemd is on, but it is not your server's systemd. Recent WSL runs systemd, so the unit files in ⁠nginx, systemd, TLS can be written and started locally, which is a much better place to get a unit wrong than on the machine serving your site. What it cannot do is convince you the deployment works — different kernel, different network, no TLS, no real ports open to the world. Practise the shape here; prove it there.

What to practise in it

If you want a use for this beyond having it, here is the honest list of what is better learned locally than on the box, roughly in the order you will meet them.

Practise hereBecause
Users, sudo and permissionsLocking yourself out of WSL costs two minutes. Locking yourself out of the server costs a support ticket.
Writing a justfilePure text and fast feedback. Nothing about it needs a server at all.
⁠SQLiteA database is one file. Make ten, delete ten.
A systemd unitGetting a unit wrong locally teaches you the same lesson without any downtime.
Restoring a backupThe restore drill wants a machine you are happy to wreck. This is that machine.
On macOS you already have this and you can skip the whole section. macOS is a Unix, its terminal is a real shell, and the differences from Ubuntu are small and mostly about package managers — brew instead of apt. If you want a closer match to the server than macOS gives you, the answer is the same idea in a different wrapper: a small virtual machine running the same Ubuntu. Worth doing eventually, not worth doing today.
03

Bring in the agent

This is the part that changes what you are capable of building.

The creature working at a desk alongside a second, taller one of the same kind.

⁠Claude Code

Install it, point it at the root of the machine, and let it work.

Two creatures of the same kind sitting side by side at one desk in front of a single wide monitor.
Two of you now. The one that types is faster than you, and the one that decides is still you.

⁠Claude Code is Claude in a terminal, with hands. It reads and writes files, runs commands, installs packages, starts services, reads its own error output and has another go. It is not an autocomplete in a sidebar. It is closer to a colleague you hand a task to and then check on.

The honest version, because you will meet this on day two and it is better to hear it now: imagine hiring a mechanic for your garage who is genuinely pretty good, but who might also, without warning, start doing aerobics instead of fixing the car. He will confidently repair a gas leak by taping the windows shut, and if you question it he will defend the decision, fluently, with reasons. That is what you are working with. It is still an enormous amount of mechanic for the money — you just do not hand him the keys and go on holiday.

On the server, as your normal user:

curl -fsSL https://claude.ai/install.sh | bash

Then close the terminal and open a new one so your shell picks up the new location, and check it:

claude --version

The first time you run claude, it prints a URL. Open it in the browser on your laptop, sign in to the Claude account your subscription is on, paste the code back. That is the whole setup.

Start at the root

Here is the bit people find surprising, and it is the single most useful habit in this guide:

cd /
claude

/ is the top of the filesystem — every single thing on the machine is somewhere underneath it. Starting there means the agent can see the web server config, the systemd units, the logs, the database and every project at once. When you say "put this online at demo.mysite.com", nobody has to explain where ⁠nginx lives or which service to restart. It can go and look, the way you would.

On a normal development machine this would be an unhinged thing to do. Here it is the whole point. The box is disposable, the work is in git, and an agent with full context is the difference between "here is a code snippet, now go and configure your web server" and "it is live, here is the link".

Give it real permission. Claude Code asks before every command by default, which is exactly right on a laptop full of your own life and merely exhausting on a box that costs $3.74 a month. Press Shift+Tab to cycle the permission modes and let it go. The safety net was never the confirmation prompt — you stopped reading those on day two, everybody does. The safety net is that the machine is disposable, the work is in git, and you have backups.

Things worth knowing on day one

Just describe the outcome
"Make me a page at status.mysite.com that shows whether each of my services is running, dark theme, updates by itself." Not a list of steps — you are hiring for the result, not renting hands. It will ask when something genuinely matters, and it is usually right about what does not.
Esc stops it
Interrupts whatever it is doing without killing the session, so you can redirect mid-task. Use it the moment you see it heading somewhere you did not intend — earlier is much cheaper than later.
/clear wipes the conversation
Starts fresh, keeping the same folder and the same permissions. This matters more than it sounds and gets its own section below.
/model switches which Claude you are talking to
Different models are better at different halves of the job. Also below.
# at the start of a message writes it down
It saves the fact into the project's CLAUDE.md, so the next session already knows it. "# the database password lives in /srv/app/.env, never in git."
It can use sudo
Installing packages, editing nginx, restarting services — all of that wants root. Let it. Just actually read the line when the words rm -rf, DROP or --force go past. Those three are the ones with no undo.

A good first session

Something that exercises the whole pipeline and gives you a real URL at the end:

cd /
claude
Set me up a new project at /srv/hello with a small Go web server
that serves one nice-looking page. Give it a justfile with recipes
for dev, check and deploy, a systemd unit so it survives a reboot,
and an nginx config. Use port 8090.

Explain what each piece does as you go — I have not done this before.
Do not put it on the internet yet; I want to see it working locally
first with curl.

Watch what it does. Ask it why it did any of it — "why a systemd unit and not just running it?", "what is that nginx block for?" — and keep asking until the answers stop surprising you. That conversation is worth more than three tutorials, because it is about the actual machine in front of you rather than somebody's generic example, and because you can interrupt it.

The thing I got wrong: I mistook confidence for correctness. An agent is roughly as competent as the person driving it. It multiplies whatever you bring — but multiply zero by a billion and you still have zero, and for a while I was bringing zero to areas I did not know well, then nodding along at explanations that sounded superb. The fix is not distrust, it is verification: make it show you the thing working. Not a summary of it working. curl the endpoint, read the log, restart the service and watch it come back. It is very hard to hallucinate a green check you ran yourself.

Driving the agent

The habits that separate a good session from a frustrating one.

A small cart with a figure at the wheel, a taller figure of the same kind walking ahead pulling it.
The agent supplies the speed. Steering is a separate job, and it is the one thing here you cannot hand over.

Everything in this section is about one thing: context. An agent is only as good as what it can currently see and how much rubbish is piled on top of it. A fresh session with a clear brief beats a long session that has been wandering around for four hours, every single time, and it is not close.

Plan with one model, build with another

The two halves of building software want opposite things. Planning wants breadth, patience, and a tolerance for sitting in an ambiguous problem without grabbing at the first answer. Building wants speed, precision and something close to stubbornness. Asking one session to do both is how you end up with a confident implementation of the wrong idea. Use /model to switch — and revisit the pairing whenever the roster changes, which is the next section.

PhaseModelWhat you ask for
Plan Fable Architecture, the directory layout, the order of work, what could go wrong. Ends with a written plan in docs/ and no code at all.
Build Opus Execute the approved plan. Fast, thorough, good at holding a large codebase in its head and finishing things.
Review Either A fresh session, no memory of writing the code, asked to find what is wrong with it.

The rhythm looks like this:

# 1. Think it through — no code yet.
/model fable
Read the codebase. I want to add user accounts with passkeys.
Write docs/features/auth.md: the approach, the database changes,
the endpoints, and what could go wrong. Do not write any code.

# 2. Read the plan yourself. Argue with it. Then:
/clear
/model opus
Read docs/features/auth.md and implement it exactly.
Run just check before you tell me a piece is done.

The /clear in the middle is not decoration. The build session should start from the written plan, not from the rambling forty-minute conversation that produced it — including the three approaches you talked yourselves out of, which are still sitting there looking like options.

Clear early, and leave a note

A long conversation gets worse, not better. Everything you explored, abandoned and corrected along the way is still sitting in view, competing for attention with the thing that actually matters. Somewhere around 150,000 to 200,000 tokens — well before anything forces your hand — stop and reset.

Never just clear. That is throwing away the only copy. Hand over first, the way you would to a colleague going on leave:

We are about to run out of useful context. Write docs/HANDOFF.md
covering: what we set out to do, what is now done and working, what
is half-finished and exactly where, the decisions we made and why we
rejected the alternatives, and the precise next three steps.

Write it for someone competent who has never seen this project.

Then /clear, and open the next session with:

Read docs/HANDOFF.md and CLAUDE.md, then continue from step one.

This is not a workaround for a limitation. It is a better way to work even when you have plenty of room left, because writing the handoff forces every decision to be said out loud in plain words, and half the time that is when you notice the decision was wrong. You also get a project that documents itself as a side effect, which is the only kind of documentation that ever actually gets written.

I have written two hundred and thirty of these across three projects. I mention the number because "leave a handoff" sounds like the sort of advice people give and do not take, and I want you to know it is a practice rather than a suggestion. One warning from doing it badly first: put the date and time in the filenamedocs/handoffs/2026-09-06-1430-auth-done.md — not in a folder name. I started with a directory per day and it was useless. With the timestamp in the filename, a plain directory listing sorts itself into a readable log of the whole project, and finding "the session where we changed the schema" takes four seconds.

Audit what it remembers

CLAUDE.md is a file in each project that the agent reads at the start of every session. It fills up over time, and some of what collects in it goes stale — a convention you abandoned, a port you moved, a rule that was true for exactly one afternoon in March. Stale instructions are worse than no instructions, because they get followed confidently by something that has no way of knowing they expired.

Every couple of weeks:

Read CLAUDE.md and every .md file under docs/. For each claim,
check whether it is still true of the actual code. Give me a list:
still true, now wrong, or no longer relevant. Do not change
anything yet.

Good things to keep in CLAUDE.md:

  • Commands: "always just check before deploying".
  • Constraints: "this box has 4 GB of RAM; keep the service under 200 MB".
  • Decisions with a reason: "stdlib only — no web framework, deliberately".
  • Traps: "the formatter is duplicated in Go and JS on purpose; change both".
  • Ports already taken, so the next project does not collide.
The thing I got wrong: one of my CLAUDE.md files reached 24,000 words. For comparison, my median is about 650 and the shortest useful one I have is 29 words. Twenty-four thousand is not a briefing, it is a novella, and every session on that project paid for the whole thing in tokens before it got to a single useful instruction. The really funny part is what it was full of: notes about keeping context clean. Context rot, growing inside the file whose job was to prevent it. A CLAUDE.md points at docs/. It does not absorb it. If yours is over about 1,500 words, the excess is documentation wearing a briefing's coat, and it belongs in a file the agent can choose to open.

Let it look at other people's work first

Before you start anything unfamiliar, send it off to read how the problem is normally solved. This costs one prompt and routinely saves a day of quietly reinventing something in a shape nobody else uses — which you will only discover much later, when you try to explain your project to someone and watch their face.

Before you write anything: find two or three well-regarded Go
projects on GitHub that do something structurally similar to this.
Read how they lay out directories, name things, handle config and
structure tests. Summarise what you found, what you are copying,
and what you are deliberately doing differently.

Seven habits, in one place

  1. One task per session Not "build the app". "Add the login page." Finish, clear, next.
  2. Make it explain before it acts On anything with consequences, ask for the plan first. Arguing with a paragraph is far cheaper than arguing with a diff.
  3. Commit constantly Every working state gets a commit. Then a mistake costs minutes, and "undo that" is a thing you can actually say and mean.
  4. Verify from outside "It should work" is not a result. curl the URL, restart the service, reboot the box and watch it come back on its own.
  5. Interrupt early Esc the second it heads somewhere wrong. Sunk cost applies to agents too, and it applies to you watching one.
  6. Write things down as you learn them Every "oh, that is why" becomes a line in CLAUDE.md or a note in docs/, immediately, while you still care.
  7. Ask whether a script would do Before handing something to the agent, ask yourself "could I just write a script for this?" If the answer is even maybe, write the script. Determinism you can read beats cleverness you have to check, and the less you need to trust, the better your results get.

Keeping your models current

The one part of this whole setup that changes underneath you every few weeks. Prices move in both directions, names move, and the model you chose in spring is rarely the right answer by autumn — which is a chore if you ignore it and a small, automatable advantage if you do not.

Two tall cartons stand at the left with a third tilted over as it is lifted out, and the small creature at the right holds up a fresh carton of the same shape.
Same shape, newer date, and usually cheaper. The roster is a file you diff, not a page you remember to visit.

Two different habits live here, and people usually have one and not the other. The first is for any project of yours that calls a model — through ⁠Claude's API directly, or through a router that fronts hundreds of them. The second is for the tool you are holding: keeping the agent itself current, and knowing what you are spending.

If your app calls models, keep a roster

The catalogue is genuinely volatile. On a router like OpenRouter there are several hundred text models on any given day, and the spread is not subtle: top-tier models cost a hundred times what small fast ones do, and the same model hosted by different companies varies by a factor of ten or more, because they run it on different hardware at different precision. There is no "the price". There is a price per model per provider per day.

And it moves the pleasant way about as often as the unpleasant one. Newer versions frequently land cheaper than the thing they replace — the current top-end Claude model costs a third of what the previous generation's did — and one model on that list had an announced price rise cancelled and the introductory price made permanent. If you pinned a model a year ago and never looked again, the most likely outcome is not that you are being overcharged. It is that you are paying more than you need to for worse output than you could be getting.

# The whole catalogue, with prices, as JSON. No key needed to read it.
# Prices are dollars per token, so multiply by a million to get the
# number everybody actually quotes.
curl -s https://openrouter.ai/api/v1/models \
  | jq -r '.data[] | select(.id|startswith("anthropic/")) |
      "\(.id)  $\((.pricing.prompt|tonumber)*1000000|round)/M in  $\((.pricing.completion|tonumber)*1000000|round)/M out"'

That one command is the whole basis of the habit, because it means the roster is a file you can diff rather than a page you have to remember to visit. Save it weekly, compare against last week's, and have it shout when something you actually use has changed. Four rules make the difference between a roster that helps and one that quietly breaks your product:

  1. Pin exact versions in code; choose new ones deliberately. Aliases that always point at "the latest" are excellent in a terminal and a liability in production — the model changes without a deploy, and so does your output. Routers say this themselves: use the concrete name if you want reproducibility.
  2. Set a price ceiling in the request. Most routers let you refuse anything above a figure you name. Then a price rise fails loudly on a Tuesday instead of appearing silently on a bill six weeks later. Same reasoning as the hard spending cap on anything of yours that can spend money: the limit belongs in the code, not in your memory.
  3. Log which model actually answered. With fallbacks, cheapest-first routing or an alias, the model that served the request is often not the one you asked for. The answer comes back saying who wrote it. Keep that field; it is the only way to explain a change in quality afterwards.
  4. Keep twenty prompts with known-good answers. Run them before you switch anything. This is the entire difference between "I upgraded the model" and "I upgraded the model and my classifier quietly got worse". There is a well-known study of one hosted model whose accuracy on a simple task fell from 84% to 51% over three months with no version change at all — behaviour drifts even when nothing is announced, and the only defence is a handful of cases you check.
Watch for retirement, not just price. Models are withdrawn. Anthropic commits to at least 60 days' notice before retiring one and publishes the dates in advance; routers carry an expiry field per model and will email you when something you use is scheduled to go. A request to a retired model does not degrade gracefully — it fails. Put that same weekly script in charge of noticing, because a calendar reminder for "check if my model still exists" is a reminder you will cancel.
"Free" models have a meter too. Free tiers on these routers are usually a small daily allowance, not a free account. I built the little creatures dotted through this page on an image router, assuming a catalogue of free models meant free experimentation, and discovered the limit was three generations per day across the whole account — not three per model. The six characters took thirty-nine attempts. On the free tier that is a fortnight; on the cheapest paid model it was about two pence. Read the limit before you plan around it, and if a tool of yours can spend money, give it a hard ceiling it enforces by reading its own ledger rather than by trusting you.

Keep the agent itself current

Claude Code updates roughly constantly, and the difference between versions is not cosmetic — tool behaviour, context handling and model availability all change. If you installed it the normal way it updates itself in the background; if you installed it through a package manager, including on Windows, it does not, and it will happily sit six months behind while you wonder why a documented feature is missing.

claude doctor      # version, install method, update channel, last update result
claude update      # take one now rather than waiting for the background check

claude doctor is the one worth knowing, because it answers the question people actually have, which is "is this thing updating itself or not". On this box it reports the version, that auto-updates are enabled, which channel they come from, and when the last one succeeded. There are two channels: the default takes releases as they ship, and the other lags about a week and skips anything that turned out badly. If the agent is doing work you depend on, the slower one is the better default, and you can still pull an update by hand the moment you want one.

Inside a session, three commands are worth having in your fingers:

/usage
What you have spent and how much of your plan's window is left. Subscriptions have both a rolling few-hour allowance and a weekly one, and the difference between "I have hit a limit" and "I have hit which limit" decides whether you wait an hour or change how you work.
/context
What is actually in the conversation, drawn as a grid. This is how you find out that half your window is one enormous file somebody pasted, which is both the cost and the reason the answers got vaguer.
/model
Switch models, and set how hard it thinks. Worth revisiting whenever a new one lands, along with planning with one model and building with another — the right pairing changes with the roster.

The token habits are the same ones in Driving the agent, and they are worth restating here because they are also the cost habits: a long conversation is an expensive conversation and a worse one. Clearing early with a handoff note, one task per session, and keeping files small are not frugality measures that cost you quality. They buy quality and happen to be cheaper.

Write me a small script that saves the model catalogue from my provider's API
each week, compares it against last week's copy, and tells me: anything whose
price changed, anything I use that has a retirement date, and any new model
cheaper than the one I am using for each job. Keep the JSON files in the repo
so the history is the record. Run it from a systemd timer and email me only
when something changed.

That script is the whole discipline. It takes an evening to write with the agent sitting right there, it costs nothing to run, and it turns "keeping up with models" from a thing you feel vaguely guilty about into an email that arrives when something actually happened.

Laying out a project

The same skeleton every time, in every language. It is the reason nothing gets lost.

A tall shelving unit of identical boxes at the right, a small figure carrying one towards it.
Every project the same shape, so that finding something is remembering rather than searching.

Every project on this box lives in /srv/<name>/ and looks broadly the same, whether it is Go, ⁠Rust, ⁠Python or a pile of HTML someone emailed you. The sameness is the feature. You can come back after six months, or hand it to a friend, or point a brand new agent at it, and the shape alone tells you where things are before you have read a word.

/srv/myproject/
├── README.md          what this is, how to run it, in five minutes
├── CLAUDE.md          notes for the agent: conventions, traps, commands
├── TODO.md            the next few things, in order
├── ROADMAP.md         the next few months, deliberately vague
├── justfile           every command this project has
├── .gitignore         secrets and build output, never committed
├── docs/
│   ├── PLAN.md        the architecture, written before the code
│   ├── HANDOFF.md     where the last session stopped
│   ├── DECISIONS.md   what was chosen, and what was rejected, and why
│   └── features/
│       ├── auth.md    one file per feature, written before building it
│       └── search.md
├── deploy/            nginx config, systemd unit — the real copies
└── ...                the actual code, laid out however the language wants

Why docs/ earns its place

Written documents are the only thing that survives a /clear. The conversation is gone the moment you end it. The file is still there tomorrow, and next month, and for whoever picks the project up after you — including you, who will not remember any of this, and who will be quietly furious with the version of you that did not write it down.

docs/PLAN.md
Written before any code, by the planning model. What is being built, how it is structured, in what order. If you cannot write this, you do not yet know what you are building.
docs/features/<thing>.md
One per feature, written before that feature exists. The approach, the data it needs, the endpoints, the edge cases. This is what you hand a build session so it starts with a brief instead of a vibe. When a feature earns pictures, let the file become a folder — docs/features/auth/ with the write-up inside it and a screenshots/ next to it. Save those as .webp: same picture, a fraction of the bytes, and every browser has read them for years.
docs/DECISIONS.md
The single highest-value file in the repo. Every real choice with a date, a reason, and the options you rejected. It is the file that stops you and your agent relitigating the same argument every three weeks.
docs/HANDOFF.md
Written at the end of every long session. The bridge between one context window and the next. Once you have more than a couple, give them their own folder and put the date and time in the filename so the listing reads as a log.
TODO.md and ROADMAP.md
TODO is the next handful of concrete things. ROADMAP is the direction — deliberately vague, because a detailed six-month plan is fiction. Ask the agent to keep TODO current; it is good at it and it makes "what were we doing?" a one-second question.

The one file to copy into everything

If you take a single file from this section, take docs/quick-start.md. It is not in the tree above, because I only worked out that it was the important one fairly recently — and it is already in nearly as many of my projects as the justfile is. It is not a getting-started guide for users. It is an orientation file for whoever opens the repository next, human or machine, and it answers five questions in about a page: where things are, what the stack is, which ports it uses, the sixty-second tour, and the hard rules that must not drift.

That last heading is the one that earns its keep. Hard rules — do not drift. Three or four lines saying this project uses ⁠SQLite and not an ORM, the ⁠nginx config in deploy/ is the real one, never edit the copy under /etc. An agent reads that and stops rediscovering your preferences by trial and error, and so does a person.

Start every project by copying the last one's skeleton. "Set up /srv/newthing with the same structure as /srv/myproject — same justfile shape, same docs layout, same .gitignore approach — but for a Python service." Yes, you will copy in feature files you never write and delete them unread. That is not waste, that is what a skeleton is for. Your conventions compound: by the fourth project the agent already knows how you like things and has stopped asking.
The thing I got wrong: I did not do any of this until late 2025. Every project I started in 2024 is code and nothing else. No justfile, no notes for the agent, no docs/ — and I genuinely cannot tell you what several of them do without reading them from the top. They are not old. They are eighteen months old and they are already archaeology. Nothing in this section is twenty-year-old wisdom handed down from a mountain; it is a fairly recent fix for a problem I gave myself, which is precisely why I can tell you which parts paid off.

The boilerplate nobody remembers

Websites in particular carry a pile of small files that are individually trivial and collectively the whole difference between something that looks finished and something that looks like homework. Nobody remembers all of them. I do not remember all of them, and I have been forgetting them since 2004. That is fine. This is exactly the kind of tedious completeness a machine is better at than you are.

Audit this website against the standard set of files a public site
is expected to have: favicon in the sizes browsers actually request,
web manifest, robots.txt, sitemap.xml, Open Graph and Twitter card
tags, a real 404 page, security.txt, a health endpoint, correct
canonical URLs, and sane cache headers on static assets.

Tell me which are missing, then add the ones that make sense here.

There is a fuller version of this as a tickable launch checklist further down.

⁠just

One file that makes every project work the same way, whatever it is written in.

Every project has commands. How do I run it? How do I test it? How do I put it live? The answers are always different — go run ./cmd/server, npm run dev, cargo watch -x run, python -m uvicorn app:app --reload — and you will not remember a single one of them in three weeks. I promise you that you will not. I have gone back to my own projects and had to read the CI config to find out how to start them.

just is a small program that reads a file called justfile and gives every command a name. It is not tied to a language, a framework or an ecosystem, and it has no opinion about your project. It runs shell, and that is the entire pitch. That neutrality is why it is worth adopting on day one: it will still be the right tool when your next project is in a language you have not learned yet.

sudo apt install -y just
just --version

Create the justfile before you write any code. Not after. It is the first file in a new project, because it is where you and your agent agree on what this project's verbs are, and that agreement is worth more than any of the code you are about to write.

The shape

# Variables at the top. Change them here, not in five places below.
service := "myapp"
port    := "8090"

# Running plain `just` lists everything, which is why this is first.
default:
    @just --list --unsorted

# Run it locally while you work on it
dev:
    go run ./cmd/server -addr 127.0.0.1:{{port}}

# Everything CI would run. Do this before deploying.
check: fmt vet test

fmt:
    gofmt -l -w .

vet:
    go vet ./...

test:
    go test ./... -count=1

# Build, install, restart, and confirm it actually came back
deploy:
    go build -o bin/server ./cmd/server
    sudo systemctl restart {{service}}
    @sleep 1
    @curl -fsS http://127.0.0.1:{{port}}/healthz && echo " — up"

# Follow the log. Ctrl-C stops watching, not the service.
logs:
    journalctl -u {{service}} -f -o cat

deploy here builds over the running binary and restarts it, which is the right amount of machinery on day one. The night it matters you will want the version you can take back, and it is two changes away.

Now every project you ever build responds to the same four words:

just          # what can this project do?
just dev      # run it
just check    # is it broken?
just deploy   # put it live
This site's own justfile, recorded on the box it runs on. The menu is what just prints; the check is what has to pass before a deploy.

A ⁠Rust project's just dev runs cargo run. A ⁠Python one runs uvicorn. A static site runs a file server. You do not care, and neither does your muscle memory. Neither, and this is the part that matters, does the agent — put the recipes in CLAUDE.md and it uses them instead of inventing a slightly different command every session and leaving you three ways to start the same program.

The deploy recipe above is three lines and it is the most important shape in this whole file: restart, wait a beat, then ask the service whether it is actually alive. That curl at the end is not garnish. A deploy that does not verify is a deploy that reports success when the binary panicked on startup, and you will find out about it from a person instead of from a script. Mine has since grown up — it refuses to run on the main branch, runs just check, opens the pull request, merges it and then deploys — but it started as those three lines, and those three lines are already honest.

Never put a secret in a justfile. The justfile goes to git; the secret does not. Read secrets from environment variables and keep the real values in a file .gitignore excludes. just loads a .env automatically if you put set dotenv-load at the top, which gives you one clean seam: commands are shared with everybody, values are shared with nobody.
# At the top of the justfile — loads .env, which git never sees
set dotenv-load

# The recipe references the name; the value lives only in .env
migrate:
    psql "$DATABASE_URL" -f db/schema.sql

The vocabulary, and why you should never change it

Pick these names once and then use exactly these names in every project you ever build, no matter what language it is in or what it does. Not similar names. The same names. The whole return on just is that you can sit down in front of a project you have not opened in a year and already know that just dev starts it and just logs tells you why it stopped.

RecipeWhat it should do
defaultPrint the list. Always @just --list, never a build. Typing plain just should never surprise you.
devRun locally with fast reload. Never touches anything live.
buildProduce the artefact, and nothing else.
testJust the tests, so check can call it.
lintThe linter, same reason.
checkFormat, vet, test. The one thing the agent runs before it is allowed to say "done".
deployBuild, install, restart, and verify. A deploy that does not check is a hope with a progress bar.
shipThe grown-up deploy, once you have one: check, commit, push, then deploy. One word for the whole ritual.
logsFollow this service's output. The first thing you run when something is wrong.
statusIs it up. One line, no ceremony.
backupPut the data somewhere safe. A first-class recipe, not an afterthought. See Backups.
cleanDelete build output. Never anything else, ever.
smokeHit the real public URL from outside and check the status codes and headers.

The part I did not expect is that the names outlive whatever is reading them. When I started keeping a .claude/skills/ folder so an agent could be told how to deploy a project without me explaining it every time, I did not design what went in it. I wrote the same words down again — deploy, check, logs, status, handoff, migrate. A recipe is that list for a person at a shell. A skill file is that list for a model in a session. Ten of my projects now carry both, saying the same thing twice in two formats, and that turned out to be the point rather than the waste: I never had to teach the newer tool anything, because the words were the project's and not the tool's. That is the real reason not to rename logs — not tidiness, but that the name will still be there after the thing you picked it for is gone.

One more thing, and it is the habit that has paid me back the most: the justfile is where you write down the thing that bit you. Not the command — the reason. Why a permission has to be granted again after installing. Why there is a sleep before the health check. Why this one recipe must not run on the main branch. Those comments are long and they look like clutter, and every one of them is a note from a past version of you who lost an evening finding it out.

The thing I got wrong: for a long time I named these differently every time. run here, serve there, start in the one from last spring, a deploy that in one project meant "build" and in another meant "build and push to production". So every return visit began with reading a file to find out what the words meant this time. Twenty-nine projects now use the list above, unchanged, and the compounding is absurd for something that costs nothing: I stopped having to remember anything at all.

⁠GitHub

Your undo button, your backup, and the thing that builds your installers.

⁠Git records the history of a project: every change, when, and why. ⁠GitHub is where that history lives on the internet. Together they are the entire reason you can let an agent run loose without your stomach hurting, because every working state is recoverable and "undo the last hour" stops being a wish and becomes a command. Nothing else in this guide buys you as much nerve for as little effort.

Later, GitHub is also what builds your Windows, ⁠macOS and Linux installers for free. That is the GitHub Actions section.

If you have never used it

  1. Make an account ⁠github.com/signup. Free. Turn on two-factor authentication when it offers — it is required for pushing code anyway, so do it now rather than mid-task later.
  2. Install the command-line tool on the server gh handles authentication and repository creation without you ever touching a token by hand.
    sudo apt install -y gh git
  3. Log in
    gh auth login
    Choose GitHub.com, then HTTPS, then Login with a web browser. It shows an eight-character code; open the URL on your laptop, paste the code, approve. It stores the credential securely on the server and configures git to use it. You will never type a password to push again.
  4. Tell git who you are
    git config --global user.name "Sam"
    git config --global user.email "sam@example.com"

Putting a project up

Ask the agent to do it, or do it yourself — it is four commands:

cd /srv/myproject

git init -b main
git add -A
git commit -m "Initial commit"

# Creates the repo on GitHub and pushes, in one go.
# Swap --private for --public when you are ready to show it.
gh repo create myproject --private --source=. --push

The one that ruins people: committing a secret

Bots scrape newly pushed public commits within minutes, hunting for API keys, database passwords and cloud credentials. This is not a cautionary tale that happened to a friend of a friend. It is a constant, automated, industrial process, running right now, and it does not care how small your project is. A leaked cloud key can become a five-figure bill while you sleep.

Here is how it goes, and every step of it is reasonable:

  1. 1. You put the database password in config.js, just for now, to get it working.
  2. 2. It works. You commit everything, because everything is what changed.
  3. 3. A week later you make the repo public, because you are proud of it and you should be.
  4. 4. You notice config.js, delete the line, and push the fix within the hour.
  5. 5. Why is there a crypto miner on my server?

Step four is where people think they are safe, and it is the step that does nothing at all. Git kept the old version. It kept it because that is git's entire job.

Write .gitignore before your first commit, not after. Git remembers everything, deliberately and permanently. A secret committed and then deleted is still sitting in the history, perfectly readable, one click away, for as long as the repository exists. Removing it properly means rewriting history and force-pushing, and by then it has been scraped anyway. Thirty seconds of .gitignore before the first commit removes this entire category of afternoon from your life.

A sane starting .gitignore for any project:

# Secrets — never, in any form
.env
.env.*
!.env.example
*.pem
*.key
credentials.json
secrets/

# Build output — regenerate it, do not store it
bin/
dist/
build/
target/
node_modules/
__pycache__/
*.exe

# Local noise
.DS_Store
*.log
*.sqlite
.vscode/
.idea/

The !.env.example line is the trick worth stealing. You commit an example file listing every variable the project needs, with obviously fake values, so that anybody — a future you, a friend, an agent opening the repo cold — can see what has to be set without ever seeing a real value. The genuine .env never goes near git. Two files, one committed and one not, and the question "what does this thing need to run?" has an answer that does not involve reading the source.

# .env.example — committed, and completely safe
DATABASE_URL=postgres://user:password@127.0.0.1:5432/myapp
SESSION_SECRET=generate-a-long-random-string
STORAGE_BUCKET=my-bucket-name

Where secrets actually go

Secret used byWhere it lives
Your app, on the serverAn .env file readable only by the service user, or Environment= lines in the systemd unit. Never in the repo.
GitHub ActionsRepository → Settings → Secrets and variables → Actions. Referenced as ${{ secrets.NAME }} and masked in logs.
You, personallyA password manager. Not a text file, not a note, not a chat message to yourself.
Your agentNowhere special — it reads the same .env the app does. Tell it in CLAUDE.md that the file exists and must never be committed or printed.
If you do leak one, rotate it — do not just delete it. The only thing that works is invalidating the credential at its source: new database password, new API key, revoked token. Deleting the file, amending the commit or flipping the repo to private afterwards accomplishes nothing, because a copy of it exists somewhere you cannot reach. Rotate first, feel embarrassed second, clean up third. And do not skip the rotation because you are hoping nobody noticed; they are not people, they are scripts, and they noticed.

An environment variable, worked through once

That table says "environment variable" as though you had met one, and half this guide leans on the idea, so here is the one worked example. An environment variable is a named value that a program is handed by whatever started it. Not a file the program opens and not a line in its source: it arrives with the process, the way a letter arrives already inside an addressed envelope. The program asks for it by name and gets whatever the parent put there. Same binary, different envelope, different database. In a terminal, the envelope is the shell:

# For one command: the name, the value, then the command.
GREETING="hello from the shell" printenv GREETING

# For the rest of this terminal session. Every program you start from
# here on inherits it. Close the window and it is gone.
export GREETING="still hello"
printenv GREETING

# Everything your shell is currently handing out. PATH and HOME are in
# there, which is how every command you type gets found.
printenv

Every language reads one with a single call — process.env.GREETING, os.environ["GREETING"], os.Getenv("GREETING") — and that call is the whole integration. The program neither knows nor cares whether the value came from a terminal, a file or systemd, which is exactly what makes it the right seam for a secret: the code says the name, in the open, in git, and the value is somebody else's problem. The .env.example above is a list of names. The .env beside it is the values.

On the server, the thing that starts your app is not a terminal, it is systemd, so the envelope has to be systemd's. It offers two ways to fill it: Environment=NAME=value written into the unit, and EnvironmentFile= pointing at a file of them. Use the file for anything secret, for one plain reason: the unit in /etc/systemd/system is readable by every account on the machine — it was installed 644, like nearly all of /etc — and systemctl show prints its Environment= lines to anyone who asks. A file that only the service's user can open is neither of those things. It is the same .env the app reads in development, in the same format:

# /srv/myapp/.env — NAME=value, one per line. No spaces around the =,
# and no "export" in front of anything (see the notice below).
DATABASE_URL=postgres://myapp:a-long-random-password@127.0.0.1:5432/myapp
SESSION_SECRET=another-long-random-string
# The service's user owns it and nobody else can read it: the 600 from
# the permissions section, on the file that most deserves it.
sudo chown myapp:myapp /srv/myapp/.env
sudo chmod 600 /srv/myapp/.env

# Tell the unit about it without editing the unit. This opens an editor
# on a drop-in file, /etc/systemd/system/myapp.service.d/override.conf,
# and merges whatever you type there into the unit when you save.
sudo systemctl edit myapp
[Service]
EnvironmentFile=/srv/myapp/.env
sudo systemctl restart myapp

# What systemd was told: the unit, plus every drop-in, as one document.
systemctl cat myapp

# What the process actually received, asked of the process itself.
sudo cat /proc/$(systemctl show -p MainPID --value myapp)/environ | tr '\0' '\n'

The last command is the one to keep. The unit says what was asked for; the process's own environment says what happened, and when the two disagree the second one is the truth. Permissions made the same point about testing as the service rather than as yourself. This is the same point about reading the service rather than the file.

The line that starts with export. A .env written for the shell, the kind you load with source .env, has export at the start of every line, and systemd drops every one of those lines without a word. Checked on this box: a file with four variables, one of them written export NAME=value, handed the service three, with no error and no log line. The app starts, finds its connection string empty, and dies with a message about the database that says nothing about why. Write the file as bare NAME=value and both readers agree; just's set dotenv-load reads that shape too.

Committing like someone who wants to sleep

Commit every time something works, not once a day. Small commits mean small undos, and the size of your undo is the size of your worst afternoon. Ask the agent to do it — "commit that with a message describing what changed and why" — and it will write better commit messages than most people, largely because it has not yet learned to type "fixes" and go to lunch.

git log --oneline -10        # what happened recently
git diff                     # what have I changed but not committed
git restore <file>           # throw away my changes to one file
git revert <commit>          # undo a commit, safely, as a new commit
04

Decide what to build

What it really is, then what shape it takes, then what it is written in.

The creature kneeling over a dark blueprint, drawing four rectangles in a row.

Everything is CRUD

Before you pick anything, notice what you are actually building. It is almost certainly the same thing as everyone else.

Four large blank cards standing in a row across the frame, a small figure at the left of them.
Create, read, update, delete. Almost everything you will ever build is those four verbs and some nouns.

Here is the single most useful thing I know, and it took me an embarrassingly long time to see it. Nearly every piece of software that anybody pays for, uses daily, or gets excited about is four operations on some rows: create a row, read some rows, update a row, delete a row. Create, read, update, delete. CRUD. That is the job.

Not most software. Not business software. I have written games, marketing platforms, kiosks, scheduling systems and things I still could not describe at a party, and essentially all of it was getting some data into some tables and then making it pleasant enough to navigate that somebody who finds an ATM stressful would be comfortable using it. The world is not a holographic simulation. It is a collection of rows and columns, and a surprising amount of it is in a spreadsheet somewhere that Becky in accounting maintains by hand.

The exercise, on three ideas you might actually have

Pick whichever of these is closest to what is in your head. Then look along the row and notice that they are the same four tables wearing different clothes.

The tableA multiplayer gameA little online shopA booking system
Peopleplayers — name, avatar, ratingcustomers — email, addressclients — name, phone
Thingsitems, levels, charactersproducts — price, stockrooms, staff, services
Eventsmatches — who, when, resultorders — who bought what, whenbookings — who, what, when
Linesthe moves in a matchthe items in an orderthe slots in a booking

A game is a booking system where the room is a match and the client is a wizard. A shop is a booking system where the slot is a parcel. The clever, fiddly, genuinely hard parts of each — the matchmaking, the payment flow, the double-booking rules — are real work, but they sit on top of those four tables rather than replacing them. Get the tables right and the hard part has somewhere to stand. Get them wrong and no amount of clever fixes it.

Why this is worth knowing on day one

  • It makes the project describable. "Four tables, a form to add a booking, a page to list them, a way to cancel one" is a thing you can hand to an agent. "A booking app" is not.
  • It kills the paralysis. You are not attempting something nobody has done. You are attempting something everybody has done, in your own shape, which is a completely different feeling.
  • It tells you what to build first. Create and read. Get a row in and get it back out on a page. Update and delete are half an hour once the first two work.
  • It survives every fashion. The framework you use will be unpopular in four years. The four tables will be exactly the same.

So before you open the stack chooser, write your four tables down. On paper, in a text file, in a message to yourself. Then ask the agent to argue with them:

A good opening prompt. "Here is my idea in three sentences. Give me the four or five tables it needs, with columns, and tell me for each one what happens when a user creates, reads, updates and deletes a row. Do not write any code — I want to argue with the shape first." Ten minutes of this saves a fortnight, and it is the cheapest ten minutes in the whole guide.
The thing I got wrong: I thought the interesting part was somewhere else. For years I chased the clever bits — the algorithm, the animation, the architecture diagram with the boxes — and treated "just some tables" as the boring scaffolding you get past on the way to the real work. Then I noticed that every project which went badly went badly at the tables, and every project that went well had a boring, obvious, almost stupid data model that somebody could understand in a minute. The tables are the real work. The clever bits are the garnish, and you cannot garnish your way out of a wrong schema.

What are you actually building?

One chart, eight honest endings, and a strong opinion about seven of them.

Before anybody argues about languages, there is a question that decides almost everything and gets skipped: where does this thing have to run? Not what it does, not how big it might get — where it runs. A thing that runs in a browser and a thing that runs on an iPhone are different projects with different costs, different release cycles and different people who can say no to you, and the language is close to the last decision rather than the first.

So: answer the chart. Every ending tells you what to build, what it will cost you, and which languages and frameworks are genuinely available for that particular expedition — not the one I like best.

01
Do you want to make an app?

We are starting from the absolute beginning here.

02
Where does it have to run?

Pick the place it must work, not everywhere you would eventually like it. This one question decides most of the rest.

03
What does it need that a website cannot do?

Be honest rather than ambitious. The reality check two sections down is the long version of this question, and the answer is usually "nothing".

03
Does it need the computer itself?

Almost everything people assume needs a desktop program does not. This is the question that tells you which.

03
What kind of game?

The split is not 2D versus 3D exactly — it is whether the last ten per cent of performance and platform reach is the product.

04
Why does the store matter?

Because there is a good chance you can have the listing without writing a native app, and a small chance you genuinely cannot.

the only wrong answerThen you are finished! Close the browser.

Go outside. Have a nice time. Genuinely — there is no shame in not wanting to build software, and the world has plenty of people making things nobody asked for, myself very much included.

When you do want to make something, it will still be here, and it will be free, and it will probably have been updated since.

the answer most of the timeMake it a website. Put it on your own box.

No installer. No review. No signing certificate. No store account, no platform fee, no rule change. Everybody is on the newest version the second you deploy, it works on a phone and a laptop and a library computer, and the way you distribute it is that you send somebody a link. That is a remarkable set of properties and the industry spent fifteen years trying to talk you out of noticing.

What you give up: it mostly wants a connection, and a short list of device features are off the table — the reality check is that list, honestly labelled. For the overwhelming majority of ideas, you give up nothing you were going to use.

The page

HTML, CSS and JavaScript, because the browser runs nothing else. Start plain. Add ⁠htmx if you want interactivity without a build step, or ⁠React / ⁠Svelte / Vue once genuinely many things on screen have state at once. ⁠TypeScript if the project is big enough to forget your own function signatures, and ⁠Vite to build it.

The server

⁠Go is the default in this guide and the reasons are in Front end and back end. But this is the layer where language genuinely does not matter much: ⁠Python with FastAPI or Django, ⁠PHP with Laravel, ⁠Node with Express or Hono, C# with ASP.NET, Java with Spring, Ruby on Rails, Elixir with Phoenix. All of them serve HTTP correctly. Pick the one whose ecosystem has the thing you need.

The records

⁠SQLite in a file until you can name the reason for something else, then ⁠Postgres. See Where the data lives, and the memory rule in buying the server.

the opinionated answerBuild a website and let them install it

A progressive web app is a website with two extra files: a manifest that says what its name and icon are, and a service worker that can answer requests when the network cannot. Add those and a phone will offer to put it on the home screen, where it gets its own icon, its own splash screen, no address bar, and works with no signal. It is not a simulation of an app. For the kind of thing most people are building, it is an app, and it took an afternoon rather than a quarter.

And here is the honest part, because it is mostly Apple's fault and it is worth knowing before you commit. On an iPhone every browser is ⁠Safari underneath, so "it works in ⁠Chrome" tells you nothing about iOS. Push notifications do work — but only after the user has added your site to the home screen themselves, through a share-sheet ritual that no ordinary person discovers unaided, so anything load-bearing built on notifications will quietly not reach most of your iPhone users. Stored data can be evicted after a few weeks of the app not being opened, so offline is a convenience and never the only copy. Bluetooth, USB and serial do not exist in any iOS browser at all. Screen capture does not exist. Background work does not exist. Android, for what it is worth, is dramatically better at every single one of these.

Now the thing that is never on the other side of the ledger. Build native and you have signed up to a relationship where a company you have never met can reject your release, change a rule, deprecate the API you built on, take a cut of your revenue, demand a yearly fee to keep existing, and require you to do a round of work every autumn because the operating system moved. That happens on a schedule. The web's contract has not meaningfully changed in fifteen years: your HTML from 2010 still renders. The opinion of this guide is to take the web, eat the missing device APIs, and keep the ability to ship a fix in ninety seconds — because being able to fix things immediately is worth more, over the life of a project, than any single capability on that list.
What you actually need

A site.webmanifest, a service worker, and icons. No framework is required for any of it — this site ships a manifest and no build step. Workbox will generate a sensible service worker if you would rather not write the caching rules yourself.

Same as any website

The whole web stack applies unchanged, because it is a website. Nothing about your server, your database or your deploy changes.

The escape hatch, later

If you ever do need a store listing or one native capability, Capacitor wraps the site you already built and lets you add native plugins for the one thing. You are not choosing to never have an app; you are choosing not to start with one.

a listing, without a rewriteWrap the web app. Do not rewrite it.

You can have a real store listing without a native codebase. On Android, a Trusted Web Activity puts your site in the Play Store as an app, and it is your site — same URL, same deploy. On iOS, Capacitor produces a real app binary around your web app that you submit like any other.

Two warnings. Apple's review guidelines reject an app that is only a website in a wrapper with nothing added, so the wrapped version needs to be a genuine app experience rather than a bookmark — offline behaviour, a real icon, native share, push. And you have now bought a release cycle: review latency, a developer account, and a version of your app frozen on people's phones until they update.

On "discovery", since that is usually the reason. The stores were a discovery engine in 2012. Today they hold millions of apps, search is dominated by people with advertising budgets, and the median new app is found by nobody. A link that a person can tap in a message, an email, a QR code on a poster or a search result converts better than a store listing nobody will ever type the name of. Go in the store because a specific human is requiring it, not because you assume it is where users come from.
Android

Trusted Web Activity via Bubblewrap, or Capacitor. Play Store account is $25 once, ever.

iOS

Capacitor. ⁠Apple Developer Program is $99 a year, and you need a Mac to submit.

Unchanged

Your site, your server, your database, your deploy. That is the entire point of doing it this way.

read this before you buildSelling inside the app is the one place the store really has you

If digital goods are sold inside an iOS app, Apple requires their in-app purchase system and takes a cut — and the rules about even mentioning that a cheaper price exists on your website have been through several courtrooms and are still moving. Android's terms are similar in shape.

The structure that avoids the entire problem: sell on the web, where you keep your margin and use ⁠Stripe like any other business, and let the app be the thing people use after they have paid. Every large subscription business you can think of does exactly this, which should tell you how well it works. If your product genuinely is consumable purchases inside a game loop, then in-app purchase is the cost of that business model and you should price it in from day one rather than discovering it in month five.

Web payments

Stripe, Paddle or Lemon Squeezy. Paddle and Lemon Squeezy act as merchant of record, which means they handle sales tax in places you have never been.

In-app, if forced

StoreKit on iOS, Play Billing on Android. RevenueCat if you need both and would rather not maintain two receipt validators.

a real exceptionThis is one of the cases where native is the answer

There is a genuine list, it is short, and you are on it. Bluetooth peripherals. NFC and tap-to-pay. Health and fitness data. Background location and geofencing. A watch app, a widget, a lock-screen presence. CarPlay or Android Auto. VoIP that rings like a phone call. Serious augmented reality. On-device machine learning that needs the platform's own accelerator. Anything that must keep working while the app is closed. None of that is available to a browser on iOS, and most of it is not available on Android either.

So build it native, with your eyes open about what it costs: a developer programme, a review queue between you and your users, a yearly round of work when the platform moves, and either two codebases or one cross-platform codebase with its own tax.

Do it in this order, though. Ship the web version first, for everybody, including the people whose device cannot do the special thing. Then build the native app for the capability, talking to the same server and the same database you already have. Going the other way — native first, web extracted afterwards — is not a refactor, it is a rewrite with a kinder name.
iOS, properly

Swift with SwiftUI. Nothing else is a serious answer if the app is iOS-shaped, and it is a genuinely pleasant language.

Android, properly

Kotlin with Jetpack Compose. Same reasoning.

Both at once

Flutter (Dart) for a single codebase that draws its own UI; React Native (JS/TS) if your team is already a web team; Kotlin Multiplatform to share logic and write each UI natively; .NET MAUI if you live in C# already.

Cheapest route in

Capacitor around the web app you already have, with a native plugin for the one capability. Keeps one codebase and one deploy for ninety per cent of the work.

genuinely a programAn installed desktop application

Reading a folder in place, driving a label printer, sitting in the system tray, launching at startup, automating another application — the browser sandbox exists specifically to prevent all of that, on every platform, and no amount of cleverness gets around it. You need a program. Installers and signing covers what that costs, including the warning screen and the price of removing it.

The hybrid almost nobody suggests. One blocked capability rarely justifies making the whole thing an installed app. Build the website for everything that can be a website, then ship a small companion that does only the hardware part and talks to your server. Most users install nothing; the few who need the printer get a ten-megabyte download instead of an application you now maintain on three operating systems.
Your web skills carry over

⁠Electron (JS/TS, ~80–150 MB, best documented), Tauri (⁠Rust core, ~5–15 MB, uses the system browser engine), ⁠Wails (Go core, similar size). The interface is the HTML and CSS you already wrote.

Native toolkits

C# with WinUI or WPF for Windows-only. Swift for Mac-only. Qt (C++ or Python) when you need real native widgets on all three and have the patience.

Smallest possible

A single binary with no window at all, plus a tiny local web page for the interface. Go or Rust, two megabytes, ship it in a zip.

what the box is best atA service with no face

This is the easiest thing on the entire chart and the reason renting a server is such a good deal. No interface to design, no browser compatibility, no store, no installer, nobody's phone. A program that starts when the machine boots, does its job, writes to a log and gets restarted if it dies — which is a systemd unit, and the rest of this guide is already about exactly that.

Language

Go, for the reason that matters here: one static binary, no runtime to install, tiny memory, and concurrency that is boring to write. ⁠Rust if the work is genuinely heavy. ⁠Python if the library you need only exists there, which for anything touching data or machine learning it usually does.

Scheduling

A systemd timer for anything periodic — see Serving it. A goroutine and a ticker if it lives inside a service you are already running.

Queues

Start with a table in your database and a loop. Reach for ⁠Redis or a real queue when jobs must survive a crash and be retried, and not one day before.

the lowest-friction software there isA command-line tool

No window, no installer, no signing, no store, no warning screen. You hand somebody a file and it works. If the audience is you, your colleagues, or other developers, this is frequently the correct answer to something that was described to you as needing an interface — and it is by far the fastest thing to build, which means you find out whether the idea was any good this afternoon.

Go

The standout, for one reason: go build produces a single static binary with no dependencies, and you can cross-compile for Windows, Mac and Linux from this box in three commands. Distribution is "download this file".

Rust

Same single-binary story, with the best argument-parsing library going (clap) and a compiler that is stricter than your future self.

Python

Fine, and the right answer when the work is data-shaped. Ship it with uv or pipx so the person installing it does not have to understand virtual environments.

Bash

Under about a hundred lines, and only for glue. Past that it becomes a language you do not want to debug at midnight, and every experienced person has learned this the same way.

a solved problemA browser game

2D in a browser stopped being a compromise a long time ago. You get the best distribution mechanism in the history of games — a link — no store cut, no review, no install, and a player can be playing four seconds after hearing about it. For a first game this is not the lesser option, it is the sensible one.

JavaScript engines

Phaser for a full 2D framework, PixiJS if you want the renderer and none of the opinions, or plain canvas, which is genuinely fine and teaches you more.

Export to the web

Godot exports to HTML5 and is free and open source. LÖVE (Lua) exports via love.js. Bevy (⁠Rust) compiles to WebAssembly.

Where to put it

Your own box, obviously — it is a static folder and some JavaScript. Then also on itch.io, where people who like small games actually look.

use an engineA 3D or console game

Do not write this from scratch. An engine is twenty years of other people's solved problems — physics, shaders, asset pipelines, controller support, platform export — and the part you actually want to spend your life on is the game. This is also the one branch of the chart where the "just self-host it" instinct does not apply: consoles mean dev kits and non-disclosure agreements, and storefronts take their cut.

Godot

Free, open source, no royalties, no company that can change the licence on you. GDScript is easy to learn; C# is supported. The obvious starting point in 2026.

Unity

C#, the largest asset store and tutorial corpus by a wide margin, and a licensing history worth reading before you commit a multi-year project.

Unreal

C++ and Blueprints. What you use when the visuals are the product, with a royalty above a revenue threshold.

Bevy

⁠Rust, data-oriented, genuinely lovely, and still young enough that you will be reading source rather than documentation.

Notice what that chart never once asked you. It did not ask which language you know. It did not ask whether anything was object-oriented, functional or procedural, and it will not, because that is very nearly the least useful thing you can know about a language when you are choosing one. What decides a language's fate on a project is duller and much more important: where can it run, what can it reach, and does the library I need exist in it. JavaScript wins the browser because it is the only thing in the browser. Swift reaches HealthKit. Python has the data libraries. Go makes one file you can copy to a server. None of those facts are about syntax, and all of them will decide your project. There is a whole section further down on using several on purpose.

Pick your stack

Five questions. A recommendation, the reasoning, and a prompt to paste.

The chart above tells you what shape the thing is. This tells you what to type. They are two halves of the same question and the chart is the half that matters more — but once the shape is settled, the stack really does fall out of it, and the widget below will hand you an opening prompt for your agent.

"What should I build this in?" stalls more projects than any other question, and it is almost always asked in the wrong order. The language barely matters. What matters is the shape — who runs it, whether it needs the machine underneath it, and what has to still be there tomorrow. Answer those three and the stack falls out on its own, which is a relief, because arguing about languages is the most enjoyable way ever invented to not build something.

Answer five questions and get a specific recommendation, the reasoning behind each piece, and an opening prompt you can paste straight into ⁠Claude Code.

What it will probably tell you

The recommendations are opinionated on purpose, so here are the opinions in plain sight rather than hidden in a widget. They are defaults, not laws. If you already know you want something else, use it — you are allowed to have your own conventions, and you should end up with some. Being polyglot is nearly free now, which gets a whole section of its own.

  • Most things should be websites. No installer, no store review, no signing certificate, and everybody is on the newest version the moment you deploy.
  • ⁠Go for the server. One binary, no runtime to install, fast, genuinely good at doing many things at once, and a language that has barely changed in a decade — which means an agent's knowledge of it is not quietly three versions stale.
  • ⁠SQLite for records, until you can say why not. A file on the disk, backed up by copying the file. ⁠Postgres is excellent and it is the right answer the day you can name the reason; "it feels more professional" is not a reason.
  • ⁠Rust only where it earns it. Real systems programming, or compute heavy enough that the difference shows up on an invoice.
  • ⁠Electron when somebody genuinely needs an icon on their desktop — and only then.

Can it be a website?

Tick what your idea genuinely needs. The honest answer is usually yes — and where it is no, it is usually the iPhone.

The web can do vastly more than people assume, mostly because the people telling you what it cannot do stopped checking around 2016. It also genuinely cannot do a short list of things, and one platform is responsible for nearly every gap on that list. On iOS, every browser is ⁠Safari underneath⁠Chrome on an iPhone is Safari wearing a different hat. So "it works in Chrome" and "it works on phones" are different sentences, and you find out which one you meant at the worst possible moment, in front of somebody.

Reality check

Tick everything your idea genuinely needs — not what would be nice.

The hybrid nobody suggests, and usually should. One blocked feature almost never means giving up the web. Build the website for everything that can be a website, then ship a small installed companion for the one part that cannot — a helper that talks to the USB device and forwards to your server, say. Most of your users install nothing at all, and the handful who need the hardware get a ten-megabyte download rather than a whole application you now have to maintain on three operating systems.

Know what shape it is? The stack chooser turns that into a specific recommendation and an opening prompt.

Front end and back end

Two words people use constantly and almost never define.

A side view of a small car with its bonnet propped open, a small figure beside it with one arm raised.
The engine is under the bonnet and the rims are on the outside. People spend an astonishing amount of time arguing about rims.

Front end is what runs on the visitor's device. Back end is what runs on yours. That is the whole distinction, and every other rule falls out of it — the security ones especially, because you control exactly one of those two machines and it is not the one with your user sitting in front of it.

An analogy I have never managed to improve on. The HTML is the body of the car and the CSS is the paint and the rims. The JavaScript in the browser is the gas pedal — it makes things happen when you press them. Your program on the server is the engine under the bonnet, and the database is what is in the boot, which is where the things you actually care about are kept. People spend an astonishing amount of time arguing about rims.

front end
The browser, on their phone or laptop

HTML, CSS and JavaScript. Draws the page, reacts to taps. The user can read every line of it, so nothing secret ever goes here.

the wire
HTTPS across the internet

Encrypted, so nobody in between can read or alter it. This is what the certificate is for.

your box
⁠nginx, the front door

Terminates HTTPS, decides which of your programs a request belongs to, and hands it over. One nginx serves every site on the machine.

your box
Your program — the back end

Checks who is asking, decides what they may see, reads and writes the database, returns a page or some JSON. This is where the rules live, because this is the machine you control.

your box
Your database, on the box

The permanent record. A ⁠SQLite file next to the app, or ⁠Postgres listening on the loopback address only — either way, unreachable from the internet no matter what.

Everything in the front end is public. All of it. The user can read your JavaScript, watch every request go out, and change any of them before they are sent. An API key in front-end code is a published API key. A check that only happens in the browser is not a check, it is a polite request. Every rule that actually matters gets enforced again on the server, because the server is the only machine in the picture whose behaviour you control.

Why not JavaScript on the server

JavaScript has to be the front end; the browser runs nothing else. The usual next thought is that using it on the back end too gives you one language everywhere, and that was a genuinely strong argument back when a person had to hold both halves in their head at once. It is worth a great deal less when an agent is doing the typing for both.

What you are trading away for it:

  • A dependency tree you cannot see the bottom of. A modest ⁠Node server routinely pulls in hundreds of megabytes and thousands of packages, every one of them something that can break, change under you, or be taken over. The Go equivalent usually has one or two — a database driver, maybe a crypto package — and you can name both of them.
  • Churn. The correct way to do any given thing in the JavaScript ecosystem changes every couple of years, and the previous correct way becomes faintly embarrassing. That is bad for you and it is worse for an agent, whose knowledge of a fast-moving ecosystem is a photograph with a date on it. Go's syntax and idioms have barely moved in a decade, so what the model learned is still true.
  • Deployment weight. Node means shipping a runtime and a node_modules directory and keeping versions aligned. Go produces one file. Copy it to the server, run it. That is the deploy.
  • Memory. On a 4 GB box, a handful of Node services and their runtimes add up quickly. Small Go services measure in single-digit megabytes.
This is a default, not a decree. If you already know Node well, or the one library you need exists nowhere else, use it and do not feel bad. A monorepo will happily hold a Node service next to a Go one. This is only about which way to lean when you have no reason to lean either way, and for a beginner on a small box with an agent doing the typing, Go wins that comfortably. Nobody has ever been fired for finishing something.

Choosing, layer by layer

LayerStart withReach for something else when
PagePlain HTML, CSS, a little JavaScriptThe interface has real state in many places at once — then a framework starts paying for itself.
ServerGo standard library net/httpAlmost never for the web parts — routing, TLS, templates and JSON are all built in. Add a library when you can say out loud what it does for you.
RecordsSQLite, in a filePostgres when you have concurrent writers, real JSON querying, or a second service that needs the same data.
Big filesObject storage — R2 or GCSThe server's own disk, if it is small and you have backups.
Sessions / cacheThe database, or just memory⁠Redis, once you have genuinely measured that you need it. Not before, and not because a blog post frightened you.
Background jobsA goroutine, or a systemd timerA real queue, when jobs must survive a crash and be retried.
Systems work⁠RustYou are writing something where the last 10% of performance is the product.
The thing I got wrong: I spent years defending a language instead of finishing things. I could tell you exactly why the fashionable criticisms of my stack were unfair, and I was mostly right, and being right was worth nothing at all. Meanwhile the people who never joined the argument were shipping. The pattern is easy to spot once you have been it: someone wanders from language to language looking for the one that will finally make it easy, always about to start, and years later they have opinions instead of projects. Pick the boring default in this table, build the thing, and form your opinions afterwards from evidence you gathered yourself.

Installers and signing

When you really do need an icon on somebody's desktop, here is what it costs and what it does not.

⁠Electron is a browser with the address bar taken off. Your app is HTML, CSS and JavaScript — the same skills, and frequently the literal same code as your website — wrapped up so that it gets a window, an icon, a tray presence and access to the machine's files and devices. VS Code, Slack and ⁠Discord are all Electron, which should settle any argument about whether it is real software.

The trade is an honest one and you should make it with your eyes open: you get out of the browser sandbox, and you pay in size — an Electron app starts around 80–150 MB — and in the considerable friction of persuading a human being to install anything at all.

The alternatives, briefly

OptionSizeWhen it fits
Electron~80–150 MBThe default. Best documentation, best tooling, and an agent has seen an enormous amount of it.
Tauri~5–15 MBSame web front end, ⁠Rust core, uses the system's own browser engine. Much smaller — at the cost of rendering slightly differently on each OS.
Go + ⁠Wails~10–20 MBLike Tauri but Go instead of Rust. Neat if the rest of your stack is already Go.
A plain binary~2–15 MBCommand-line tools. No window, no installer, no signing drama. Ship it in a zip.

The warning screen, and what it costs to remove

Software you hand out is unsigned unless you pay somebody to vouch for you. Unsigned software runs perfectly well — the operating system simply shows a frightening box first, and a certain number of your users will read it and stop. Note what you are actually buying here: not security, and not quality. You are buying the absence of a scary sentence.

PlatformUnsigned looks likeSigning costs
Windows A blue "Windows protected your PC" box. More infoRun anyway gets past it. $120/yr via Azure Trusted Signing ($9.99/month, needs a verified business or sole trader), or $200–$600/yr for a traditional certificate, which since 2023 must live on a hardware token or HSM.
⁠macOS "Cannot be opened because the developer cannot be verified." Right-click → Open, or approve it in System Settings → Privacy & Security. $99/yr for the ⁠Apple Developer Program. This also gets you notarisation, which is what actually removes the warning.
Linux Nothing. Nobody asks. Free.
You almost certainly should not pay for this yet. Signing is optional and it stays optional for a long time. Inside a company, hand out unsigned builds and tell people to click through — an enormous amount of internal software has always worked exactly like that. For a personal project, put a screenshot of the warning and one sentence of instructions on the download page and get on with your life. Pay for a certificate on the day strangers are refusing to install your software and you can point at what it is costing you. Not the day before that.
Build the website first anyway. If the idea can be a website at all, build that version first and wrap it in Electron later — the code carries over almost entirely, and you will have shipped something to real people months earlier. Going the other way, starting native and extracting a web version out of it afterwards, is not a refactor. It is a rewrite with a nicer name. One order keeps every door open and the other quietly shuts them.
05

Put it on the internet

From a name, to a padlock, to installers for operating systems you do not own.

The creature on a hilltop raising an antenna, signal arcs coming off the tip.

Pointing a domain

DNS is a phone book. You are adding one entry to it.

A large open book with an arrow leaving its pages and travelling to a small distant tower.
One entry in the world’s phone book. That is the whole of it, and it is the part everybody is most afraid of.

When somebody types your domain, their computer asks the internet's directory service which address that name belongs to. The directory is DNS and the entries in it are records. You need two of them, and there are only three record types worth learning in the first place, which is a much smaller subject than its reputation suggests.

a lookup, end to end Five steps, and only the last one touches your machine. Point at any of them.
  1. Your browser has the name mysite.com and cannot use it. Every connection on the internet is made to a number.
  2. So it asks a resolver — your provider's, or one you chose, like 1.1.1.1. The resolver does the walking.
  3. If nobody has asked lately, the resolver goes to your nameservers: the only machine in this picture you control, answering with the record you typed.
  4. The answer comes back and the resolver keeps it for the length of the TTL, which is why changing a record is not instant for everybody.
  5. Now your browser opens a connection, straight to the number. DNS carried no part of your site. It answered one question.
None of that was your server working. It was the internet finding out where your server is.
A
Name → IPv4 address. mysite.com203.0.113.40. This is the one that matters.
AAAA
Same, for IPv6. Add it if your provider gave you an IPv6 address; skip it otherwise.
CNAME
Name → another name. www.mysite.commysite.com. Follow the second name to find the address.

In Cloudflare

  1. Open your domain, then DNS → Records
  2. Add the apex record Type A, Name @ (which means the bare domain itself), IPv4 address your server's IP, Proxy status DNS only, TTL Auto.
  3. Add www Type CNAME, Name www, Target your bare domain, Proxy status DNS only.
  4. Wait a few minutes ⁠Cloudflare updates almost immediately, but other networks hang on to old answers for a while. Ten minutes is typical, an hour is not alarming, and going and doing something else is more effective than refreshing.
Turn the orange cloud off, at least at first. Cloudflare's proxy — the orange cloud — is genuinely useful later on: it hides your server's real address and soaks up attacks. While you are still setting things up it is a machine for generating confusing symptoms. Certificate requests behave differently, your logs fill with Cloudflare's addresses instead of real visitors (fixable, later), and a broken site shows you a Cloudflare error page rather than your own, so you debug the wrong computer for an hour. Get it working grey. Then turn it orange deliberately, as a decision, and set SSL mode to Full (strict) so the connection stays encrypted the whole way to your box.

Checking it worked

From the server, or from your laptop:

dig +short mysite.com
dig +short www.mysite.com

# Ask a public resolver directly, bypassing anything cached locally
dig +short mysite.com @1.1.1.1

You want your server's address back. Nothing at all means the record has not spread yet, so wait. The wrong address means either a cached answer or that you edited a different domain than the one you think you edited — both are extremely common, neither is interesting, and both are fixed by waiting or by looking properly. Almost every "DNS is broken" is really one of these two.

The tools that answer “did it work?”

Everything above assumed dig, which is the right tool and is not on every machine. Here is the short list of things worth knowing about, split by what they actually answer — because “my site is broken” is four different questions wearing one coat, and the trick is to find out which one you have before you start changing things.

The four questions, in order

Run these top to bottom and stop at the first one that surprises you. Almost every “DNS is broken” is answered inside the first two.

# 1. Does the name resolve at all, and to what?
dig +short mysite.com
dig +short mysite.com @1.1.1.1     # bypass anything cached on your network

# 2. Which nameservers is the world being told to ask?
#    After a move to Cloudflare these should be the Cloudflare pair.
dig +short NS mysite.com

# 3. Is the server itself answering, and with what?
curl -I https://mysite.com

# 4. Who owns the name, and when does it expire?
whois mysite.com | head -20

On Windows without WSL, question one is nslookup mysite.com and question two is nslookup -type=NS mysite.com. Question three works as written: curl has shipped with Windows for years. Question four has no built-in equivalent, so use a WHOIS website, or the Ubuntu you installed on your laptop, where every command on this page works exactly as printed.

Question two is the one nobody thinks of, and it is the one that catches a half-finished move. If you changed nameservers at your registrar but the old ones are still coming back, then nothing you have typed into the new DNS panel has ever been read by anybody. You are editing a control panel that is not connected to anything, which feels exactly like a broken record and is not one. Give it an hour, check again, and only then go looking for a real problem.
The oldest trap in DNS is your own computer. Your laptop, your router and your ISP all keep old answers for a while, and they do not agree about how long. So the site works on your phone and not on your laptop, and you conclude the record is wrong. Add @1.1.1.1 to a dig and you have asked the world directly, skipping every cache between you and it. If that is right and your laptop is wrong, nothing is broken and the answer is a cup of tea.

Email, and the one thing not to self-host

Every other section here says run it yourself. This one says do not, and the reason has nothing to do with how hard it is.

Two postboxes at opposite ends of the frame with a single envelope floating between them.
Receiving is four clicks and free. Sending is somebody else’s reputation, which is precisely what you are paying for.

A man moves to a new town, prints himself some beautiful business cards, and starts knocking on doors offering to carry everyone's post. The cards are perfect. His van is clean. He is punctual, polite and completely willing. Nobody gives him a single letter, because the town has a short list of couriers it trusts and he is not on it, and there is no form to fill in to get on it. You get on that list by years of not being a nuisance — and, as it turns out, the previous tenant of his address was a spectacular nuisance.

That is email, exactly, with no part of the analogy left over. You do not own your server's IP address; you rent it, and it arrives carrying a history that is not yours and a reputation that is. Gmail, Outlook and every corporate mail filter on earth are employed to be suspicious of strangers, and a brand new mail server on a cheap VPS is the most suspicious thing they will see all day. So, flatly: do not run your own mail server.

Not because it is difficult to install — apt will hand you a working SMTP daemon in about ninety seconds, and that is the trap, because installing it is roughly five per cent of the job. The rest is spent proving to strangers that you are not a criminal, forever, using four separate mechanisms. Their names are worth knowing anyway, because the relay you end up using will ask you to set three of them up, and it is nicer to know what you are agreeing to:

SPF
A DNS record listing which servers are allowed to send mail claiming to be from your domain.
DKIM
A signature on every message, checked against a public key you publish in DNS, proving nothing was altered on the way.
DMARC
A DNS record telling receivers what to do when the first two fail — and, usefully, where to send you a daily report about it.
Reverse DNS
Your IP address resolving back to your mail server's name. Set by whoever owns the IP, which is your host, not you.

Get all four exactly right and you can still land in spam, because you share a subnet with somebody who did not. Meanwhile most VPS providers block outbound port 25 by default and would like to know why you want it opened, which stops the majority of people at step one — and that is your host doing you a favour it will never get any credit for. It goes like this, every time:

  1. You install a mail server. It takes an hour, which feels like proof that it was a real piece of engineering worth doing.
  2. You send a test message to yourself. It arrives. You are delighted, and you tell somebody.
  3. You send one to a friend on Gmail. It lands in spam, so you add SPF.
  4. You add DKIM. Then DMARC. Then you email your host about reverse DNS and wait a day for the reply.
  5. It works. For two weeks.
  6. Why is nobody getting the password reset emails?

So hand it to somebody whose entire business is being on that list of trusted couriers. There are two halves to this and they are separate problems, which is the part nobody says out loud — most people only actually need the first one.

Receiving — free, and about four clicks

What you want is you@yourdomain.com arriving in the mailbox you already read forty times a day, and no second inbox to remember to check. ⁠Cloudflare Email Routing does exactly that and costs nothing: turn it on, say which addresses to accept and where to forward them, confirm at the far end. It writes its own MX records while you watch. Pound for pound it is the best value available on a domain, and it is the one part of this section to go and do tonight.

Make one of them a catch-all and you get something quietly excellent: every service you ever sign up to can have an address of its own, so the day ancient-forum@yourdomain.com starts receiving crypto spam you know exactly who sold you, and you can switch off that one address without touching anything else.

Sending — a relay, for the two emails your app sends

The other half is transactional mail: the password reset, the receipt, the "somebody replied to you" notification. Your app hands the message to a relay over an API or an authenticated connection, and the relay's reputation carries it the rest of the way. That reputation is the entire product. You are not paying for a mail server — you can have one of those for nothing — you are paying for the benefit of the doubt.

RelayFree tierAfter that
⁠Resend 3,000 a month, 100 a day around $20/mo for 50,000
Amazon SES nothing worth planning around $0.10 per 1,000

Resend is API first and the quicker of the two to wire up — a few lines of code and a key in an environment variable. Amazon SES is the cheapest by a distance once you have real volume, and the fiddliest to get going: new accounts start in a sandbox that will only deliver to addresses you have verified yourself, until you ask for that to be lifted.

Either is fine, and both are free at the volume of a project nobody has heard of yet. Start with whichever one's documentation you can follow at eleven at night; you can change your mind later by editing one environment variable, which is the entire reason the key lives in one.

Whichever you pick will hand you two or three DNS records to add. You do not type these from memory — the relay prints the real values — but it is worth seeing the shape once, so that a working domain looks familiar rather than magic:

# Receiving: Cloudflare writes these for you when you turn Email Routing on
example.com.                    MX    10 route1.mx.cloudflare.net.
example.com.                    MX    20 route2.mx.cloudflare.net.
example.com.                    MX    30 route3.mx.cloudflare.net.

# Sending: your relay gives you the real values. The shape is always this.
example.com.                    TXT   "v=spf1 include:_spf.your-relay.com ~all"
relay._domainkey.example.com.   TXT   "v=DKIM1; k=rsa; p=MIGfMA0GCSq..."
_dmarc.example.com.             TXT   "v=DMARC1; p=none; rua=mailto:you@example.com"

Start DMARC at p=none, which means "tell me, do not act on it". Read the reports for a fortnight, then tighten it. Setting a policy straight to reject on day one is an excellent way to discover, silently and from the far side, which of your own services had been sending mail as you.

This is not a contradiction, and it is worth being clear about why. The rest of this guide argues that renting a whole computer for $44.88 a year beats paying a cloud vendor by the request. That argument holds here too, and it points the other way. Self-host the thing where the work is the product. Outsource the thing where somebody else's reputation is the product. A VPS sells you a computer, and a computer is a computer whoever's name is above the door. A mail relay sells you the benefit of the doubt, and there is no price at which you can buy that from yourself.
Then make sure somebody reads it. Put a real, forwarded address on your domain registration, because the day you need the registrar's password reset is a poor day to find out it goes nowhere. The launch checklist asks you to publish a security.txt naming an address a stranger can use to tell you about a bug — that address should be one that reaches you. An inbox nobody opens is worse than no inbox at all, because it looks like an invitation.

⁠nginx, systemd, TLS

The three pieces between a request and your code. Your agent configures all of them.

One tall doorway with a padlock on it, three smaller doorways standing in a row beside it.
One front door, several rooms behind it. nginx’s entire job is deciding which room a knock belongs to.

You do not need to be able to write any of these from memory, and you never will, because you touch them about four times a year. What you do need is to know what each one is for, so that when something breaks at eleven at night you know which of the three to go and look at first. That is the whole reason this section exists.

one request, all the way down The dashed box is the only security boundary in this whole guide.
  1. A visitor asks for mysite.com on port 443 — the one port open to the world.
  2. nginx decrypts it, reads which domain was asked for, and hands it to whichever of your programs owns that name. That is the whole trick behind unlimited sites on one address.
  3. It forwards to your app on 127.0.0.1:8090. Loopback means this machine only. nginx is on this machine; the internet is not.
  4. Your app reads your data — a file on disk, or a database also on loopback — and answers back the way it came.
  5. A scanner tries port 8090 within minutes of your box existing, and gets nothing, because nothing there has ever been listening to the internet.
One door in the wall, and everything behind it talking to itself.

⁠nginx — the front door

Only one program can hold port 443, and nginx is it. Every request to every site on the box arrives there, nginx reads which domain was asked for, and hands it to whichever of your programs owns that name. That is the entire trick behind one server hosting unlimited sites on one address, and once you have seen it you stop being impressed by hosting panels.

Which gives you the filing system for the whole machine, and it is worth adopting deliberately, because it is what makes the fifth project as easy as the first: every app gets a port, a systemd unit, an nginx server block and a health endpoint. The ports are yours to hand out — 8090, 8091, 8092, in a tidy band you keep a note of in each project's CLAUDE.md so nothing collides. Nothing binds a public interface, ever. nginx is the only thing the internet can see.

Each site gets a file in /etc/nginx/sites-available/, linked into sites-enabled/:

server {
    listen 443 ssl;
    http2  on;
    server_name mysite.com;

    ssl_certificate     /etc/letsencrypt/live/mysite.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/mysite.com/privkey.pem;

    location / {
        proxy_pass http://127.0.0.1:8090;
        proxy_set_header Host              $host;
        proxy_set_header X-Real-IP         $remote_addr;
        proxy_set_header X-Forwarded-For   $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}
sudo nginx -t                    # ALWAYS test before reloading
sudo systemctl reload nginx      # zero-downtime; a bad config cannot take the site down
nginx -t then reload, every single time. reload keeps the old config running if the new one is broken, so a typo cannot take your sites offline. restart does not have that property; it throws away what was working and then discovers the problem. Build the habit now, while it is free, rather than learning the difference at the worst available moment.
A name the other side chose for itself proves nothing. As a kid I got onto a game forum's staff with an IRC hostname that said Squaresoft, a story told up top, and it worked because nobody asks who set the label. The web has the same trick. That config sends your app the visitor's address in two headers: X-Real-IP, which nginx overwrites, and X-Forwarded-For, which nginx adds to, keeping whatever the visitor sent at the front of it. An app that reads the front of that list, or believes either header from anyone but nginx, lets every visitor be anybody — and every rate limit and ban by address with it. Believe X-Real-IP, and only on connections from 127.0.0.1.

Not getting flooded

Nobody is going to attack your site. The people who attack things work from a list, you are not on it, and on the day you are, a rate limit is not what saves you anyway. What actually arrives, and arrives at everybody, is one enthusiastic program. It goes like this:

  1. You post the link somewhere busy.
  2. Somebody's aggregator likes it and fetches it every two seconds, forever, because its author forgot the line that says to wait.
  3. That page runs a query.
  4. That query is the slow one, the one you were going to look at.
  5. Why is production down?

A rate limit is you deciding in advance that no one address gets more than a fair share, so that when the loop turns up the other visitors never find out it did. nginx does it in ten lines, the ten lines are the same on every site you will ever build, and the visitor it turns away is answered before your app hears a word about them. That is the entire point: 429 Too Many Requests costs nginx nothing to send, and your slow query costs whatever it costs.

# Above the server block. A zone is a shared counter, one per visitor
# address; 10 MB tracks about 160,000 of them. 20 a second is more
# than a person can click and less than a loop can loop.
limit_req_zone $binary_remote_addr zone=api:10m rate=20r/s;

server {
    ...
    location /api/ {
        # burst: how far ahead one visitor may run before hearing no.
        # nodelay: answer the burst at once instead of spacing it out.
        limit_req        zone=api burst=40 nodelay;
        limit_req_status 429;

        proxy_pass http://127.0.0.1:8090;
    }
}

Put it on the things that cost something — a login form, a search, an API, the page that runs the report — and leave the front page alone. A limit on a page that costs nothing to serve protects nothing and punishes the one visitor who opened six tabs. That block is the exact one on the box you are reading this on, where it guards the dashboard's API, and this is what it says when eighty requests arrive at once. Point it at your own site, not mine:

seq 80 | xargs -P 40 -I{} curl -s -o /dev/null -w '%{http_code}\n' \
  https://mysite.com/api/stats | sort | uniq -c
$ seq 80 | xargs -P 40 -I{} curl -s -o /dev/null -w '%{http_code}\n' \
    https://mecasasu.casa/api/stats | sort | uniq -c
     48 200
     32 429

Forty got in on the burst, eight more got in because the twenty-a-second allowance refilled during the half second it all took, and thirty-two were told no by a web server that never woke the app up. That limit has been on this site since the day it went up, and in ten days of logs it had fired exactly zero times before that test — which is the correct number. It is a seatbelt, not a steering wheel. Ten lines, once, and then you do not think about it until the morning it earns its keep.

Behind Cloudflare's orange cloud, every visitor is Cloudflare. The zone above counts by address, and once the proxy is on, the address on every connection is one of Cloudflare's own. One bucket for the whole world: the first busy hour, everybody gets a 429; the first fail2ban rule, you ban Cloudflare and nobody gets in at all; and every line of your logs says the same visitor came, which is the end of counting them. Cloudflare sends the real address in a header, CF-Connecting-IP, and the fix is to let nginx believe that header — from Cloudflare's addresses only, for the reason in the box above this one. Their list changes, so it is a script, and the script runs from a timer, monthly.
#!/usr/bin/env bash
# /usr/local/sbin/cloudflare-real-ip: rebuild the list of Cloudflare
# addresses nginx is allowed to believe about who is really calling.
set -euo pipefail
out=${1:-/etc/nginx/snippets/cloudflare-real-ip.conf}
tmp=$(mktemp)
{
  echo "# Generated $(date -u +%F) from cloudflare.com/ips-v4 and /ips-v6. Do not edit."
  for v in ips-v4 ips-v6; do
    curl -fsS --max-time 20 "https://www.cloudflare.com/$v" | awk 'NF { print "set_real_ip_from " $0 ";" }'
  done
  echo "real_ip_header CF-Connecting-IP;"
} > "$tmp"
grep -q '^set_real_ip_from ' "$tmp"    # an empty list would trust nobody, silently
install -m 644 "$tmp" "$out" && rm -f "$tmp"
nginx -t && systemctl reload nginx

Then one line in the site's server block, include snippets/cloudflare-real-ip.conf;, and $remote_addr means the visitor again everywhere at once: the zone, the access log, the X-Real-IP your app reads, and fail2ban. The script was run on this box before it was printed here, and it is not installed, because this domain's cloud is grey and there is nothing to see through. The awk is there for a reason you would only find by running it: the last line of Cloudflare's list has no newline on the end, and a plain sed glues the header directive onto the last address. nginx happened to accept that. It would not have to.

systemd — keeping it running

If you start your program by typing its name, it dies the moment you close the terminal, which is a surprisingly common way for a website to have a very short life. systemd is what turns it into a service: started at boot, restarted when it crashes, logging somewhere you can actually read. Nothing you build is genuinely online until it survives a reboot you did not plan.

# /etc/systemd/system/myapp.service
[Unit]
Description=My application
After=network-online.target

[Service]
User=myapp
WorkingDirectory=/srv/myapp
ExecStart=/srv/myapp/bin/server -addr 127.0.0.1:8090
Restart=always
RestartSec=2

# Sandboxing. Cheap to add, and it means a bug in your code
# cannot become a bug in the whole machine.
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
# strict makes the whole disk read-only to this service, even folders
# it owns. Name the one place it writes, which must already exist.
# If the app writes nothing, delete this line.
ReadWritePaths=/srv/myapp/data

# Memory ceiling, and the line that makes it mean anything: without
# MemorySwapMax the service is pushed into swap instead of stopped,
# and thrashes there under its "limit". See Stress testing.
MemoryMax=256M
MemorySwapMax=0

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now myapp    # start it, and start it at every boot
systemctl status myapp
journalctl -u myapp -f               # follow the log
Give every service its own user. sudo adduser --system --group --no-create-home myapp. If your program is ever compromised, whoever is in there gets the permissions of an account that owns nothing, can log in nowhere, and has no interesting files — rather than yours. One command, once, at the point where it costs nothing. The --group is not optional: without it the account lands in nogroup, and the myapp:myapp that the next subsection hands to chown names a group that does not exist. This page left it out until 16 September 2026, and found out by running its own command.

Permissions — who is allowed to touch what

Giving the service its own user buys you one failure you are all but guaranteed to meet, so meet it here: the program works perfectly when you run it, and fails the moment systemd runs it. Nothing about the program changed. The person did. You ran it as sam, the unit runs it as myapp, and every time anything touches any file, Linux asks the same two questions — who is asking, and what is that user allowed to do here?

Every file answers with three kinds of people and three permissions. The people are the owner, the file's group, and everyone else. The permissions are read, write and execute. ls -l prints all nine on every line, and it looks like line noise right up until somebody points at the columns once:

$ ls -l /srv/myapp/config.toml
-rw-r--r-- 1 sam sam 812 Sep 14 09:12 config.toml
│└┬┘└┬┘└┬┘   │   │
│ │  │  │    │   └─ group
│ │  │  │    └───── owner
│ │  │  └─ everyone: r
│ │  └──── group: r
│ └─────── owner: rw
└─ - file, d directory

As a number, read is 4, write is 2, execute is 1, and a permission like 644 is just the three sums side by side — owner, group, everyone. Which sounds like homework, and it is not, because there are really only three numbers you will ever type:

644rw-r--r--
For files. You can change it, anybody can read it. Almost every file on a web server wants to be exactly this.
755rwxr-xr-x
For directories, and for programs. On a program, execute means run it. On a directory it means walk through it, and that is the one that catches people: a file is a room at the end of a corridor, and an unlocked room is no use to you if one of the doors on the way there is locked. namei -l, below, shows you every door.
600 and 700
Owner only. You have typed both already — on your SSH keys, which ssh refuses to use if anyone else could read them, and on the Cloudflare token in Infinite subdomains. Secrets get these, and nothing else does.

When a log says Permission denied, test as the service, not as yourself. You are the one user on the machine it definitely works for, which makes you the worst possible person to reproduce it. sudo -u borrows the service's identity for a single command:

# Who does the service actually run as?
systemctl show -p User myapp

# Try the thing that failed, as that user
sudo -u myapp cat /srv/myapp/config.toml
sudo -u myapp test -w /srv/myapp/data && echo writable

# Every directory on the way down, with its owner and permissions
namei -l /srv/myapp/data

# The usual fix: the service owns what it writes
sudo chown -R myapp:myapp /srv/myapp/data

That last line is the correct fix nine times out of ten, and the tenth is when it gets aimed at the whole project instead of the folder the program writes to. Leave the program itself owned by you. A service that can rewrite its own binary is a service that, on its worst day, can make that bad day permanent.

And somewhere in your first month a forum answer will tell you to run chmod -R 777. It works. It makes the error go away in exactly the way that taking the front door off its hinges fixes a sticky lock, and the answer with the most upvotes is very rarely the one that has to live in the house afterwards.

The thing this page got wrong: for its first eight days, the unit above could not write its own database. It said ProtectSystem=strict, and Databases said to keep ⁠SQLite in /srv/myapp/data, owned by myapp. Each was good advice, and together they could never work, because strict makes the entire disk read-only to the service — including folders it owns. The ownership was perfect. sudo -u myapp could write there all day. The service could not, because the sandbox is a second lock that the owner and the permissions never see. The tell is supposed to be the wording, Read-only file system instead of Permission denied, except SQLite does not pass that on: it says unable to open database file and sends you off to stare at a path that is fine. It was found on 14 September 2026 by running that unit on this box, and the ReadWritePaths= line is the fix. The rule to keep: when it works as the user and fails as the service, it is not the user any more. It is the unit.

Timers — things that run while you sleep

Sooner or later something needs to happen every night with nobody awake: clearing out sessions that expired, rebuilding a report, deleting uploads nobody finished. The traditional tool for that is cron, and cron works, and has worked since before most of the people reading this were born. On a new server, use a systemd timer instead. The reasons sound small, and each one is the difference between a job that failed and a job you know failed:

  • It logs where everything else logs. journalctl -u, same as your app. Cron's idea of reporting a problem is to email the output to a mailbox on the server itself, and on a server with no mail system — which is every server in this guide — it writes one line saying it threw the output away.
  • It catches up. With Persistent=true, a job whose time came and went while the machine was off runs at the next boot instead of quietly skipping a night.
  • It answers "when does this next run?" with a command, rather than with you reading five-field cron syntax and doing time zones in your head.
  • It is the same kind of thing as your service. Same User=, same sandbox, same everything from Permissions just above. There is nothing new to learn, which is the best feature a tool can have.

This box does not have cron installed at all. Nothing has missed it.

A timer is two files. The .service is the job, and a job that runs and exits is Type=oneshot. The .timer says when, and has the same name so systemd can pair them up. The job here is one query against a sessions table, as an example; yours is whatever your app needs doing at four in the morning.

# /etc/systemd/system/myapp-prune.service
[Unit]
Description=Delete expired sessions

[Service]
Type=oneshot
User=myapp
ExecStart=/usr/bin/sqlite3 /srv/myapp/data/app.db "DELETE FROM sessions WHERE expires_at < datetime('now');"
ProtectSystem=strict
ReadWritePaths=/srv/myapp/data
# /etc/systemd/system/myapp-prune.timer
[Unit]
Description=Delete expired sessions, nightly

[Timer]
OnCalendar=*-*-* 04:10:00
RandomizedDelaySec=10m
Persistent=true

[Install]
WantedBy=timers.target
# Check the schedule means what you think, before you trust it
systemd-analyze calendar '*-*-* 04:10:00'

sudo systemctl daemon-reload
sudo systemctl enable --now myapp-prune.timer    # enable the timer, not the service
systemctl list-timers                            # every timer, and when each next fires
sudo systemctl start myapp-prune.service         # run the job once, right now
journalctl -u myapp-prune.service -n 20          # and see what it said
systemctl --failed                               # anything that has broken, timers included

OnCalendar also understands hourly, daily and things like Mon *-*-* 09:00. systemd-analyze calendar prints the next time it will actually fire, in the server's own time zone, which is very probably UTC and very probably not yours. That one command is the difference between a report that lands at nine in the morning and one that lands at four.

A scheduled job fails in private. Nobody is watching it, which was the point of scheduling it. So the story always goes the same way:
  1. You add a nightly job, and it works.
  2. It works every night for seven months.
  3. You stop thinking about it, which is correct.
  4. The disk fills, or a table gets renamed, and it starts failing every night.
  5. Why is the newest backup from March?
A timer at least writes the failure down. Something outside the box has to actually notice it, and Finding out before anybody tells you has the free way to do that.

TLS — the padlock

Let's Encrypt issues real certificates, free, in about fifteen seconds. certbot asks for one, proves you control the domain, writes the files, edits your nginx config, and sets up automatic renewal.

sudo apt install -y certbot python3-certbot-nginx

# Both names in one certificate. Point the DNS records first —
# certbot proves control by answering a request on that domain.
sudo certbot --nginx -d mysite.com -d www.mysite.com

Certificates last 90 days and renew themselves via a timer that is installed automatically. You can prove the machinery works without spending a real, rate-limited issuance:

sudo certbot renew --dry-run
sudo certbot certificates            # what exists and when it expires
The thing I got wrong: I used to edit these files under /etc. Which works, right up until you rebuild the box, or move the project, or simply cannot remember what you changed nine months ago or why. The nginx config and the systemd unit are not settings. They are source code for your project, every bit as much as the program is, and they belong in the repository — mine live in a deploy/ folder at the root of each project, and just deploy installs them from there. The copies under /etc are output. Treat them as disposable and you can rebuild a whole server from git in an afternoon; treat them as the original and you are one reinstall away from archaeology.

Infinite subdomains

One domain. Unlimited names. All free, all with certificates.

This is the part people consistently underestimate, and it is the single best-value thing on the whole box. mysite.com costs money. anything.mysite.com costs nothing, forever, and you can have as many as you can think of names for. Every project, every demo, every half-idea you want to show somebody gets a real address on the actual internet, which does something to how seriously you take it.

mysite.com          the main site
app.mysite.com      the actual product
api.mysite.com      the JSON API
docs.mysite.com     documentation
status.mysite.com   is everything up?
demo.mysite.com     the thing you show people
git.mysite.com      whatever you feel like

Adding one is three steps, and your agent will do all three:

  1. A DNS record Type A, name app, pointing at the same IP. Or use the wildcard below and skip this forever.
  2. An ⁠nginx server block Same as before, with server_name app.mysite.com and a different port.
  3. A certificate sudo certbot --nginx -d app.mysite.com

The wildcard shortcut

Once you are adding these regularly, stop doing it one at a time and do it once properly. A wildcard DNS record sends every name to your server, and a wildcard certificate covers the lot of them.

The catch is a genuinely interesting one. Proving you own *.mysite.com cannot be done over HTTP, because there is no single page to serve — the name you are claiming is infinite. So it has to be proved through DNS instead, which means certbot needs permission to write a temporary record into your domain, which means an API token. This is the first real secret in the guide, so it gets handled properly.

  1. Add the wildcard DNS record Type A, name *, your server's IP. Keep the specific records for anything you want to behave differently.
  2. Make a scoped ⁠Cloudflare token My Profile → API Tokens → Create Token → Edit zone DNS, and restrict it to this one zone. Scope it as tightly as the interface will let you: it should be able to edit DNS for one domain and do absolutely nothing else — not your account, not your billing, not your other domains. A leaked token that can do one small thing is an inconvenience. A leaked token that can do everything is a weekend.
  3. Store it so only root can read it
    sudo mkdir -p /root/.secrets
    echo "dns_cloudflare_api_token = YOUR_TOKEN_HERE" | sudo tee /root/.secrets/cloudflare.ini
    sudo chmod 600 /root/.secrets/cloudflare.ini
  4. Ask for the wildcard
    sudo apt install -y python3-certbot-dns-cloudflare
    
    sudo certbot certonly \
      --dns-cloudflare \
      --dns-cloudflare-credentials /root/.secrets/cloudflare.ini \
      -d mysite.com -d '*.mysite.com'

From then on a new subdomain is one nginx block. No DNS change, no certificate request, no waiting for anything to propagate. Point the new server block at the wildcard certificate files, reload, and the name is live. This is the moment the box stops feeling like a rented computer and starts feeling like somewhere you live.

Add a catch-all that refuses unknown names. With a wildcard record, every conceivable subdomain now resolves to your box, including the several thousand you have not configured. With no default server block, nginx answers those with whichever site happens to sort first alphabetically — so a typo silently shows somebody the wrong project, and you will not find out for months. A catch-all that simply hangs up makes unconfigured names behave predictably, which is all you want from them.
server {
    listen 443 ssl default_server;
    server_name _;
    ssl_certificate     /etc/letsencrypt/live/mysite.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/mysite.com/privkey.pem;
    return 444;   # close the connection without a response
}

Databases

Where everything that has to survive a restart actually lives.

A filing cabinet with one drawer pulled fully open, a small figure standing on the drawer holding a sheet up.
A filing cabinet you can ask questions. On this box it starts life as a single file you can copy.

A plain file works fine until two things write to it at once, or until you want to ask a question more complicated than "give me everything". A database handles concurrency, constraints and queries. On this box, start with ⁠SQLite.

I want to say this one out loud, because the internet will tell you otherwise and because I have given the opposite advice myself. Reach for Postgres when you can name the reason: several programs writing at once, real JSON querying, another service that needs the same data. Those are good reasons and the day you have one you should switch without agonising. "It feels more professional" is not a reason. It is a costume. One binary, one box, one writer and a few thousand requests is SQLite territory, and it is roomier territory than almost anybody expects.

OptionUse it when
SQLite defaultOne program, one machine. A single file on disk, no service to run, no password, no port. This is most projects, for much longer than you think.
⁠PostgreSQLConcurrent writers, real JSON querying, or a second service that needs the same data. Excellent, and it will not be the thing that limits you.
⁠RedisWhen you have measured a specific need for a fast in-memory cache. Not before. It eats RAM, which is the scarce resource here.
MongoDBRarely worth it at this size. Postgres stores and queries JSON perfectly well if that is what you wanted.

SQLite, properly

There is no service to install and start. The database is a file, so the only real decisions are where it lives and who owns it.

sudo apt install -y sqlite3

# The database is one file. Keep it beside the app, owned by the
# service user, and nowhere the web server can serve it from.
sudo mkdir -p /srv/myapp/data
sudo chown myapp:myapp /srv/myapp/data
# As the service user, so the file it creates belongs to the app
sudo -u myapp sqlite3 /srv/myapp/data/app.db
-- Write-ahead logging: readers stop blocking the writer. Set once,
-- stored in the file, and the single most useful line here.
PRAGMA journal_mode = WAL;

-- Foreign keys are OFF by default in SQLite, for very old reasons.
-- Your application must turn them on for every connection it opens.
PRAGMA foreign_keys = ON;

CREATE TABLE users (
  id         INTEGER PRIMARY KEY,
  email      TEXT NOT NULL UNIQUE,
  created_at TEXT NOT NULL DEFAULT (datetime('now'))
);
.quit

Then the path goes in .env, exactly as a connection string would:

DATABASE_PATH=/srv/myapp/data/app.db
The whole sequence, for real: journal mode, a table, two rows, and the file it all lives in. Twelve kilobytes.
"database is locked" is the trap, and it is a configuration mistake. SQLite allows one writer at a time. If a second write arrives while the first is in progress and no busy timeout is set, it does not wait politely — it fails immediately. Set a busy timeout of a few seconds on every connection your app opens, and turn on WAL as above, and this error essentially disappears. Nearly every "SQLite cannot handle real traffic" story on the internet is these two lines being missing.
The backup is a copy of the file — but not with cp. Copying a live SQLite database with cp can capture it mid-write and give you a file that looks fine and is not. Use VACUUM INTO, which reads one consistent snapshot while the app carries on writing:
sudo sqlite3 /srv/myapp/data/app.db "VACUUM INTO '/var/backups/myapp-$(date +%F).db'"
Almost everywhere else will tell you .backup, and on a quiet database it is fine. It copies in pieces and starts again from the top whenever the app writes in between, so on a busy one it never finishes. On this box, with an 80 MB file, it was still going at thirty seconds with twenty-five writes a second arriving. VACUUM INTO took a fifth of a second at a thousand.

That one line is your entire just backup recipe. See Backups that hold for where the copy then goes, which is the part that actually matters.

When to reach for Postgres

On the day you can name the reason — a second service wants the same data, you have genuinely concurrent writers, you need to query inside JSON documents — Postgres is already on this box's shopping list and it is a twenty-minute move. Do it then, on purpose, and not before.

sudo apt install -y postgresql

# Confirm it is listening ONLY on the loopback address.
# 127.0.0.1 = only this machine. 0.0.0.0 = the entire internet.
sudo ss -tlnp | grep 5432
A database on the open internet is found within hours. ⁠Ubuntu's Postgres listens on 127.0.0.1 by default, and there is essentially never a reason to change that. Your application runs on the same machine, so it connects locally. If you need to browse the data from your laptop, do not open the port — tunnel it over the SSH connection you already have:
ssh -L 5432:127.0.0.1:5432 sam@203.0.113.40
Now localhost:5432 on your laptop is the server's database, encrypted, with no new port open and no new password to manage.

A database per project

sudo -u postgres psql
CREATE DATABASE myapp;
CREATE USER myapp WITH PASSWORD 'a-long-random-string-from-your-password-manager';
GRANT ALL PRIVILEGES ON DATABASE myapp TO myapp;

-- Postgres 15 and later need this second grant as well. Without it
-- your app connects successfully and then cannot create a single
-- table, with an error that does not obviously point here.
\c myapp
GRANT ALL ON SCHEMA public TO myapp;
\q

Then the connection string goes in .env, which git never sees:

DATABASE_URL=postgres://myapp:the-password@127.0.0.1:5432/myapp
That schema grant catches nearly everyone. It was a real behaviour change in Postgres 15, and most tutorials online predate it. The symptom is "permission denied for schema public" at the moment your app first tries to create a table, long after the database appeared to be set up correctly.
The thing I got wrong: I recommended Postgres by reflex for years. Including, for a while, on this very page. It is a magnificent database and it was genuinely the right call on the projects where I learned it, which had several services and a team and real concurrent load. Then I looked at what I actually build now — fifteen projects on SQLite, one on Postgres, and a pile of ⁠MariaDB left over from the ⁠PHP years — and had to admit that my advice was a description of a job I stopped doing. Notice when you are recommending your past instead of your present. It is a very comfortable mistake and it is invisible from the inside.

⁠GitHub Actions

Free computers that build your software for operating systems you do not own.

A long conveyor belt with three boxes spaced along it, a small figure reaching for one at the right end.
You push. The boxes come out the other end, built for machines you do not own and have never touched.

⁠GitHub Actions runs commands on GitHub's machines whenever something happens in your repository. Public repositories get it free; private ones get a monthly allowance a personal project will not come close to spending.

Before the examples, the honest rule, because this is an area where people build elaborate machinery they do not need: CI earns its place when it does something your own computer cannot. Building a Windows installer when you do not own Windows — that is worth every minute. Cutting a release with four binaries attached to it — worth it. Running your tests on a robot when just check already runs them here, on this box, in four seconds, before you deploy — that is a pipeline built to look like a real company, and it will cost you an afternoon of YAML for a result you already had.

So, in order of how much they actually pay:

  1. Build releases On every version tag, produce binaries or installers for Windows, ⁠macOS and Linux and attach them to a page people can download from. This is the one. If you only ever set up one workflow, make it this one.
  2. Check every push Run just check on every commit. Genuinely useful the moment somebody else is committing too, and mild duplication of your own habits until then.
  3. Deploy on push Shown below because it is satisfying and because you will want it eventually. You do not need it while just deploy takes one second and you are the only person here.

The thing nobody tells you about cross-compiling

You do not need a Mac to build for macOS, or a Windows machine to build for Windows. ⁠Go compiles for every platform from any platform — one Linux runner produces all of them in about thirty seconds, for nothing. ⁠Rust and Zig will do the same with a bit more setup. This still feels like getting away with something, and it is the single best reason on this page to have Actions at all.

.github/workflows/release.yml — this builds four binaries and publishes them whenever you push a tag starting with v:

name: release

on:
  push:
    tags: ["v*"]

permissions:
  contents: write        # needed to create the release

jobs:
  build:
    runs-on: ubuntu-latest
    strategy:
      matrix:
        include:
          - { goos: linux,   goarch: amd64, ext: "" }
          - { goos: darwin,  goarch: arm64, ext: "" }    # Apple Silicon
          - { goos: darwin,  goarch: amd64, ext: "" }    # older Intel Macs
          - { goos: windows, goarch: amd64, ext: ".exe" }

    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-go@v5
        with:
          go-version: "1.26"

      - name: Build
        env:
          GOOS:        ${{ matrix.goos }}
          GOARCH:      ${{ matrix.goarch }}
          CGO_ENABLED: "0"
        run: |
          out="myapp-${{ matrix.goos }}-${{ matrix.goarch }}${{ matrix.ext }}"
          go build -trimpath -ldflags "-s -w" -o "$out" ./cmd/myapp

      - uses: softprops/action-gh-release@v2
        with:
          files: myapp-*
git tag v0.1.0
git push --tags

A minute later there is a release page with four downloads on it, built on hardware you do not own, for operating systems you may never have used. That is the entire distribution story for a command-line tool, and it is done.

Electron apps need real machines

Desktop installers are the exception: a .dmg has to be assembled on macOS, and a Windows installer on Windows. Actions gives you both, so the matrix changes from cross-compilation to running the same job on three different runners.

jobs:
  build:
    strategy:
      matrix:
        os: [ubuntu-latest, windows-latest, macos-latest]
    runs-on: ${{ matrix.os }}

    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: "22"
      - run: npm ci
      - run: npx electron-builder --publish always
        env:
          GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}

That produces a .exe installer, a .dmg, and an AppImage and .deb — all attached to the release automatically. GITHUB_TOKEN is provided by Actions itself; you do not create it.

Deploying to your own server on every push

Optional, genuinely satisfying, and not something you need yet. Push to main and the live site updates itself. Set this up when pushing a button twice a day has started to annoy you, or when a second person joins — not on day one because it looks professional.

  1. Make a deploy-only key On the server: ssh-keygen -t ed25519 -f ~/deploy_key -C "github-actions". Add deploy_key.pub to the authorized_keys of a dedicated deploy user that owns only this project.
  2. Put the private half in GitHub Repository → Settings → Secrets and variables → Actions → New repository secret. Call it DEPLOY_KEY. GitHub masks it in every log line. Then delete the private key from the server — GitHub is the only place it needs to exist.
  3. The workflow
    name: deploy
    on:
      push:
        branches: [main]
    
    jobs:
      ship:
        runs-on: ubuntu-latest
        steps:
          - name: Deploy over SSH
            env:
              KEY:  ${{ secrets.DEPLOY_KEY }}
              HOST: ${{ secrets.DEPLOY_HOST }}
            run: |
              mkdir -p ~/.ssh && chmod 700 ~/.ssh
              printf '%s\n' "$KEY" > ~/.ssh/id_ed25519
              chmod 600 ~/.ssh/id_ed25519
              ssh -o StrictHostKeyChecking=accept-new deploy@"$HOST" \
                'cd /srv/myapp && git pull --ff-only && just deploy'
Pin third-party actions, and use few of them. Every uses: line runs somebody else's code on a machine holding your secrets. Stick to actions/*, which is GitHub's own, plus a small number of well-known ones, and pin each to a version rather than a moving branch — a branch is a promise that whatever is there tomorrow is still fine. This is one of the very few places in this guide where being a bit paranoid costs you nothing at all, so spend it here.

Backups that hold

Not an advanced topic. The part that separates software from a demo.

Two boxes on a shelf at the left, one box alone on a mound far to the right, a small figure between them.
Three copies, two kinds of storage, one of them somewhere else. The box on the right is the entire point.

Everything in this guide rests on giving an autonomous agent broad permission on a real machine, and that is only a reasonable thing to do because of this section. With good backups the worst realistic outcome is losing an afternoon. Without them it is losing the project, and losing a project is a very quiet kind of disaster — nothing explodes, you just stop being able to continue.

Backups get filed under "advanced" by an astonishing number of tutorials, usually in a paragraph at the end that says you should probably look into it. They are not advanced. Load balancing is advanced. Failover is advanced. Backups are the price of admission, and they are four lines in a shell script.

the loop, and the arrow people skip Three copies is the famous part. The green one is the part that gets skipped.
  1. A dead disk, a wrong rm, a migration that ate a column, or an agent doing exactly what you asked. They all arrive at the same place.
  2. Your server holds the working copy — which is to say, the copy that just broke. One copy on one machine is not a backup.
  3. A nightly dump runs on a timer and writes a consistent copy whether or not anybody remembers it. That is the whole difference between a backup and an intention.
  4. It goes somewhere else: a different provider, on somebody else's disks. A copy that dies with the box has only ever protected you from your own mistakes.
  5. The restore is the only step that was ever the point, and it is the one almost nobody has run.
A backup you have never restored from is a folder of files you are hoping about.

And it is not really about the agent. Here is what actually destroys data:

  • A DROP TABLE against production because that terminal looked like the other terminal.
  • A migration that runs perfectly and quietly throws away a column.
  • The provider's disk failing, or your account being suspended over a billing mix-up nobody meant.
  • You, at midnight, entirely certain that directory was the old one.

An old mantra worth adopting: nothing is ever truly deleted. Prefer marking a row as gone to actually removing it. Prefer moving a directory to .old over rm -rf. If something is irrevocably destroyed, no good can come of that — there is no upside to being unable to get it back, ever, in any situation I have encountered in twenty years.

3–2–1

The rule that has survived decades: three copies of anything you care about, on two different kinds of storage, with one of them somewhere else entirely.

Copy one
Live

The database and files on your server. This is not a backup; it is the thing you are protecting.

Copy two
On disk

Nightly dumps in /var/backups. Instant to restore from — but useless if the machine itself is gone.

Copy three
Elsewhere

Object storage at a different company. This is the copy that survives fire, deletion and a suspended account.

Grandfather–Father–Son

Keeping every copy forever is expensive. Keeping only last night's is worse than useless, and this is the bit that catches people who think they are covered. Some damage takes weeks to notice — a silently broken migration, a table gradually emptied by a bug, a file corrupted in March. If your only backup is last night's, you have carefully and reliably preserved the broken state, once a day, with great discipline.

The answer is a rotation that gets sparser as it gets older. Frequent copies of the recent past, occasional copies of the distant past.

today older → DAILY kept 7 days WEEKLY kept 5 weeks MONTHLY kept 12 months
daily weekly monthly pruned away

Twenty-four files. Together they let you go back to any of the last seven days, any of the last five weeks, or any month in the past year — and they cost a few dollars a year to store, because a compressed database dump of a small project is measured in megabytes.

The point is the gap between damage and discovery. A bug that silently drops rows can go unnoticed for a month. Seven daily copies are no help whatsoever; one monthly copy from before the bug shipped saves you. Design your retention around how long it takes you to notice, not how long it takes to restore. Those are completely different numbers and only one of them is under your control.

Four things to back up

WhatWhere it isHow
DatabasesA ⁠SQLite file, or ⁠PostgresNightly. sqlite3 … "VACUUM INTO" for SQLite — never cp on a live file — or pg_dump -Fc for Postgres. Irreplaceable; this is the one that matters.
Code/srv/*Already on ⁠GitHub if you commit. That is a backup, and an off-site one.
Uploads/srv/*/dataAnything users gave you. Not in git, not regenerable.
Config & secrets/etc/nginx, /etc/systemd/system, every .envNot in git by design. Encrypt these before they leave the box.

The script

/usr/local/bin/backup.sh. It finds what you run on its own: every SQLite file in a data folder under /srv, and every Postgres database if there is a Postgres. Ask your agent to fit it to your actual setup — it will get the details right and it will not get bored halfway through.

The line worth reading twice is the --exclude. Until September 2026 this page printed a script that tarred /srv/*/data whole, live database and all, two sections after telling you never to copy a live database. Run against a busy one on this box, tar stopped with an error five times out of five, so the script never reached the step that sends anything off the machine. The database inside each tarball would not open, six out of six. It looked finished. It had four tidy numbered steps. What it actually did was write a broken file to the same disk it was meant to protect, and then stop.

#!/usr/bin/env bash
set -euo pipefail
shopt -s nullglob   # a pattern that matches nothing is skipped, not an error
umask 077           # backups hold every secret on the box: root reads them, nobody else

DEST=/var/backups/nightly
STAMP=$(date -u +%Y%m%d)
mkdir -p "$DEST"

# 1. Every SQLite database: a consistent snapshot, taken while the
#    app keeps running. Never tar or cp the live file.
for db in /srv/*/data/*.db; do
  name=${db#/srv/}          # myapp/data/app.db
  name=${name//\//-}        # myapp-data-app.db
  out="$DEST/${name%.db}-$STAMP.db"
  rm -f "$out"              # VACUUM INTO refuses to overwrite
  sqlite3 "$db" "VACUUM INTO '$out'"
done

# 2. Every Postgres database, if this box runs Postgres
if id postgres &>/dev/null; then
  dbs=$(sudo -u postgres psql -Atc \
    "SELECT datname FROM pg_database WHERE NOT datistemplate AND datname <> 'postgres'")
  for db in $dbs; do
    sudo -u postgres pg_dump -Fc "$db" > "$DEST/$db-$STAMP.dump"
  done
fi

# 3. Uploads and configuration, without the live database files.
#    Step 1 already has those, and a copy taken mid-write is corrupt.
tar czf "$DEST/files-$STAMP.tar.gz" \
    --exclude='*.db' --exclude='*.db-wal' --exclude='*.db-shm' \
    /srv/*/data /etc/nginx /etc/systemd/system /srv/*/.env

# 4. Encrypt tonight's files for the trip. Only the public key lives
#    on this box; the private key lives in your password manager.
AGE_RECIPIENT=age1yourpublickey...
OUT=$(mktemp -d)
trap 'rm -rf "$OUT"' EXIT
for f in "$DEST"/*-"$STAMP".*; do
  age -r "$AGE_RECIPIENT" -o "$OUT/${f##*/}.age" "$f"
done

# 5. Push them off this machine. rclone talks to R2, B2, S3, Google
#    Cloud Storage and about forty other things. copy, never sync:
#    sync makes the bucket match this folder, and step 6 empties it.
#    Sunday's set is also a weekly, the 1st's is also a monthly.
rclone copy "$OUT" remote:my-backups/daily
if [ "$(date -u +%u)" = 7 ]; then
  rclone copy "$OUT" remote:my-backups/weekly
fi
if [ "$(date -u +%d)" = 01 ]; then
  rclone copy "$OUT" remote:my-backups/monthly
fi

# 6. Prune: keep 7 days locally. The bucket prunes itself (below).
find "$DEST" -type f -mtime +7 -delete
sudo chmod +x /usr/local/bin/backup.sh

It needs two tools and somewhere to send things. remote: in step 5 is a name rclone looks up, and you make it once, with the keys from an R2 API token for the bucket or a B2 application key. It goes in with sudo because the backup runs as root, so root is the one who has to find it:

sudo apt install age rclone

# Cloudflare R2
sudo rclone config create remote s3 provider=Cloudflare \
  access_key_id=YOUR_KEY_ID secret_access_key=YOUR_SECRET \
  endpoint=https://YOUR_ACCOUNT_ID.r2.cloudflarestorage.com \
  acl=private no_check_bucket=true

# Or Backblaze B2, instead of the above
sudo rclone config create remote b2 account=YOUR_KEY_ID key=YOUR_APPLICATION_KEY

Run it every night with a systemd timer:

# /etc/systemd/system/backup.service
[Unit]
Description=Nightly backup

[Service]
Type=oneshot
ExecStart=/usr/local/bin/backup.sh
# /etc/systemd/system/backup.timer
[Unit]
Description=Run the nightly backup

[Timer]
OnCalendar=*-*-* 03:30:00
RandomizedDelaySec=15m
Persistent=true          # if the box was off at 03:30, run it at next boot

[Install]
WantedBy=timers.target
sudo systemctl enable --now backup.timer
systemctl list-timers backup.timer      # when does it next run?
sudo systemctl start backup.service     # run it right now, once

Let the storage do the forgetting

Grandfather–father–son is two jobs, and they belong in different places. Choosing which copies are the weeklies and the monthlies is two if lines in the script, above: Sunday's set goes into weekly/ as well, and the 1st's into monthly/. Deleting is the storage's job. Give each folder a lifecycle rule and the bucket throws away what has aged out on its own, so nothing on the server ever deletes anything off-site, which is the point of off-site.

Lifecycle rule onDelete afterWhat stays
daily/7 daysThe last seven nights
weekly/35 daysThe last five Sundays
monthly/365 daysThe 1st of each of the last twelve months

In R2 that is the bucket's Settings, then Object Lifecycle Rules, one rule per folder. In B2 it is the bucket's Lifecycle Settings, then Use custom lifecycle rules, with the folder as the File Path, daysFromUploadingToHiding set to the number above and daysFromHidingToDeleting set to 1. Both delete within a day or so of the date rather than on the minute, which for backups is as precise as it needs to be. The script above ran for four hundred simulated nights on this box with those three rules applied, and the bucket settled at seven, five and twelve: the chart, exactly.

What this section said until September 2026. That the script should end in rclone sync, and that you should turn on versioning and let the provider keep the history. sync makes the bucket match the local folder, and the local folder only ever holds a week, so every file the prune removed vanished from the bucket the following night. Forty simulated nights on this box left nine days off-site, under a chart promising a year. Versioning was meant to catch that, and R2 does not have it: Cloudflare lists the versioning calls as not implemented. B2 does, and keeps every version forever unless told otherwise, which is the same mistake in the other direction, and with a bill. The old script also uploaded the tarball unencrypted, every .env inside it, one box above a red notice telling you not to.

Both ⁠Cloudflare R2 and ⁠Backblaze B2 are cheap and have real free tiers — R2 notably charges nothing for egress, which is what makes restoring free rather than an unwelcome surprise. Fifty gigabytes of backups costs a few dollars a year.

Encrypt before it leaves the machine. Those tarballs contain your database passwords and your API keys. Handing that to another company in plain text undoes a good deal of the work in the Lock it down section. Step 4 is age, and it needs a key pair, made once:
age-keygen
That prints three lines and saves nothing. All three go in your password manager; the age1… public key goes in the script as AGE_RECIPIENT. The server only ever holds the half that locks. Keep the private key off it: a backup you cannot decrypt because the key burned with the machine is not a backup. Getting one back, into a scratch folder:
cd "$(mktemp -d)"
sudo rclone copy remote:my-backups/monthly/myapp-data-app-20260901.db.age .

# Paste the AGE-SECRET-KEY-1… line, press Enter, then Ctrl-D
age -d -i - -o app.db myapp-data-app-20260901.db.age
sqlite3 app.db "PRAGMA integrity_check;"

A backup you have never restored is a rumour

This is the step everybody skips, and it is the reason backups fail on the one day they are needed. A backup you have never restored is not a backup, it is a belief. The script that runs nightly and writes a zero-byte file will keep doing that, cheerfully, for a year. Once, by hand, so you know what it feels like before it matters:

# Restore last night's dump into a scratch database
sudo -u postgres createdb restore_test
# The dumps are root's alone, so root reads the file and postgres restores it
sudo cat /var/backups/nightly/myapp-20260906.dump | sudo -u postgres pg_restore -d restore_test

# Does it actually contain data?
sudo -u postgres psql restore_test -c '\dt'
sudo -u postgres psql restore_test -c 'SELECT count(*) FROM users;'

# Clean up
sudo -u postgres dropdb restore_test
The SQLite version of the drill above, recorded end to end: take the backup, restore it somewhere scratch, count the rows, check the file is sound.
Putting a SQLite backup back for real: move the leftovers out first. If the app crashed, app.db-wal and app.db-shm are still sitting beside the database, and SQLite will replay that -wal on top of whatever you copy in. Tried twice on this box: once the restored database would not open at all, and once it passed its integrity check while holding the broken data it was supposed to replace. The second one is worse. Nothing complained, and nothing ever would have. So the old files go into a folder, not over the top, and not in the bin either:
sudo systemctl stop myapp
cd /srv/myapp/data

# app.db and its -wal and -shm, if there are any, move aside together
old=damaged-$(date +%F-%H%M)
sudo mkdir "$old"
sudo mv app.db* "$old"/

# The backup goes in owned by the app, and WAL goes back on:
# VACUUM INTO writes its copy in the default journal mode
sudo install -o myapp -g myapp -m 600 /var/backups/nightly/myapp-data-app-20260914.db app.db
sudo -u myapp sqlite3 app.db "PRAGMA journal_mode = WAL;"

sudo systemctl start myapp

Then make the machine prove it, every night

Once a quarter is what people write down. The honest number is once, the week it was set up, and then never, and I know that because it is my number. It is not laziness. The backup has a timer, and a timer does not care whether anybody remembers it. The restore has you, and you have a Tuesday.

So give the restore a timer too. /usr/local/bin/restore-drill.sh, below, runs an hour after the backup and does what you just did by hand, mechanically, to every database the backup script found: it opens tonight's snapshot, checks it is sound, counts every table against the live database, checks that tonight's set actually left the box, and prints one line saying whether it all came back. It restores into scratch copies, and into a scratch database whose name says drill, and it only ever reads the live ones. If it dies half way through, the trap still prints the verdict and still writes the log line, because a drill that fails quietly is the exact thing it exists to catch.

#!/usr/bin/env bash
set -euo pipefail
shopt -s nullglob
umask 077

DEST=/var/backups/nightly
LOG=/var/backups/restore-drills.log
STAMP=$(date -u +%Y%m%d)     # tonight's set: this runs an hour after backup.sh
WORK=$(mktemp -d)
scratch_db=""
verdict=""
databases=0
tables=0

# The last line is the verdict and the exit code agrees with it. A trap prints
# it, so a drill that dies half way through still says so, in the log too.
finish() {
  rc=$?
  rm -rf "$WORK"
  [ -z "$scratch_db" ] || sudo -u postgres dropdb --if-exists "$scratch_db" \
    || echo "drop $scratch_db by hand" >&2
  [ -n "$verdict" ] || verdict="FAILED: stopped early, exit code $rc"
  echo "$(date -u +%FT%TZ) $verdict" >> "$LOG"
  echo "RESTORE DRILL $verdict"
  case $verdict in OK*) exit 0 ;; *) exit 1 ;; esac
}
trap finish EXIT
fail() { verdict="FAILED: $*"; exit 1; }

# A table with rows tonight must have rows in the restore. Equal is not
# required, because a day of activity sits between the two counts.
compare() {   # compare <database> <table> <live rows> <restored rows>
  printf '%-28s %-28s %8s live %8s restored\n' "$1" "$2" "$3" "$4"
  [ "$3" -eq 0 ] || [ "$4" -gt 0 ] || fail "$1: $2 has $3 rows live and none in the restore"
  tables=$((tables + 1))
}

# 1. Every SQLite database, found the way backup.sh finds them. Tonight's
#    snapshot has to exist, pass its integrity check, and hold the rows.
for db in /srv/*/data/*.db; do
  name=${db#/srv/}; name=${name//\//-}
  snap="$DEST/${name%.db}-$STAMP.db"
  [ -f "$snap" ] || fail "no snapshot of $db tonight; expected $snap"
  cp "$snap" "$WORK/drill.db"
  [ "$(sqlite3 "$WORK/drill.db" 'PRAGMA integrity_check;')" = ok ] || fail "$snap fails its integrity check"
  while IFS= read -r t; do
    compare "$name" "$t" "$(sqlite3 -readonly "$db" "SELECT count(*) FROM \"$t\"")" \
                         "$(sqlite3 "$WORK/drill.db" "SELECT count(*) FROM \"$t\"")"
  done < <(sqlite3 -readonly "$db" "SELECT name FROM sqlite_master WHERE type = 'table' AND name NOT LIKE 'sqlite_%'")
  databases=$((databases + 1))
done

# 2. Every Postgres database, if this box runs Postgres. Tonight's dump goes
#    into a scratch database whose name says drill, is counted against the
#    live one, and is dropped. The live database is only ever read.
if id postgres &>/dev/null; then
  dbs=$(sudo -u postgres psql -Atc \
    "SELECT datname FROM pg_database WHERE NOT datistemplate AND datname <> 'postgres'")
  for db in $dbs; do
    dump="$DEST/$db-$STAMP.dump"
    [ -f "$dump" ] || fail "no dump of $db tonight; expected $dump"
    scratch_db="${db}_drill_$$"
    sudo -u postgres createdb "$scratch_db"
    sudo -u postgres pg_restore --no-owner -d "$scratch_db" < "$dump" || fail "$dump did not restore cleanly"
    while IFS= read -r t; do
      compare "$db" "$t" "$(sudo -u postgres psql -Atc "SELECT count(*) FROM $t" "$db")" \
                        "$(sudo -u postgres psql -Atc "SELECT count(*) FROM $t" "$scratch_db")"
    done < <(sudo -u postgres psql -Atc \
      "SELECT quote_ident(schemaname) || '.' || quote_ident(relname) FROM pg_stat_user_tables" "$db")
    sudo -u postgres dropdb "$scratch_db"; scratch_db=""
    databases=$((databases + 1))
  done
fi

# 3. Tonight's set is off the box as well. The key that opens those copies is
#    not on this machine, on purpose, so this proves they arrived. Opening
#    one is the drill you do by hand, with the key, once a quarter.
sent=$(for f in "$DEST"/*-"$STAMP".*; do echo "${f##*/}.age"; done | sort)
missing=$(comm -23 <(echo "$sent") <(rclone lsf remote:my-backups/daily | sort))
[ -z "$missing" ] || fail "not off-site tonight: $(echo $missing)"

verdict="OK: databases restored $databases, tables counted $tables, files off-site $(echo "$sent" | wc -l)"
sudo chmod +x /usr/local/bin/restore-drill.sh
sudo /usr/local/bin/restore-drill.sh      # once, by hand, before you trust the timer

What it printed on this box, against a scratch app with three tables:

$ sudo /usr/local/bin/restore-drill.sh
myapp-data-app.db            users                               5 live        5 restored
myapp-data-app.db            sessions                            0 live        0 restored
myapp-data-app.db            posts                               4 live        4 restored
RESTORE DRILL OK: databases restored 1, tables counted 3, files off-site 2
$ tail -n 3 /var/backups/restore-drills.log
2026-09-16T03:06:04Z OK: databases restored 1, tables counted 3, files off-site 2

The rule it applies is loose on purpose: a table with rows in the live database has to have rows in the restore. Not the same number, because a day's activity sits between the snapshot and the count — the first run on this box, with a second process inserting the whole time, printed 10 live and 6 restored for the table being written to, and that is correct. Zero where there should be rows is the failure it is looking for, because it is the one that hides. A backup script that quietly started writing empty files would pass a size check and pass an integrity check, and it fails this. So does a night with no snapshot at all, a snapshot that will not open, a dump that will not load, and a file that never reached the bucket; each was tried here, and each got its own sentence and an exit code of 1.

Then the timer: the same two files as every other job on this page, with the one line from Finding out before anybody tells you that turns a failure into a message. systemd runs ExecStartPost only when the drill passed, so a failing drill never checks in, and the silence is the alarm. Tested here both ways: the passing run pinged, the failing run did not. Delete that line if you have not set a check up; the log still has the verdict.

# /etc/systemd/system/restore-drill.service
[Unit]
Description=Nightly restore drill

[Service]
Type=oneshot
ExecStart=/usr/local/bin/restore-drill.sh
ExecStartPost=/usr/bin/curl -fsS -m 10 --retry 3 -o /dev/null https://hc-ping.com/your-check-uuid
# /etc/systemd/system/restore-drill.timer
[Unit]
Description=Run the restore drill, an hour after the backup

[Timer]
OnCalendar=*-*-* 04:30:00
RandomizedDelaySec=10m
Persistent=true

[Install]
WantedBy=timers.target
sudo systemctl daemon-reload
sudo systemctl enable --now restore-drill.timer
sudo systemctl start restore-drill.service     # run it now, once
journalctl -u restore-drill -n 5 --no-pager     # and read the verdict

What it does not prove, and the page will not pretend otherwise: that the off-site copy opens. The private key is not on the box, by design, as the encryption notice above says, so the machine can only check that tonight's files arrived. Opening one is the drill with you at the keyboard — the four lines in that notice, once a quarter, with the key pasted from your password manager. The machine does the every-night part and you do the once-a-quarter part, and between the two of you nobody is hoping.

Put both in the justfile so they take ten seconds and there is no excuse:

backup:
    sudo /usr/local/bin/backup.sh

drill:
    sudo /usr/local/bin/restore-drill.sh
Hand the whole thing to the agent. "Set up nightly backups: snapshot every SQLite database with VACUUM INTO and pg_dump any Postgres, tar up the rest of the /srv data directories without the live database files, plus the ⁠nginx and systemd config, encrypt with age, rclone copy (not sync) to daily, weekly and monthly folders in my R2 bucket with a lifecycle rule on each, keep 7 days locally, and add a systemd timer. Then install the restore drill from this page on a timer an hour later, run it once by hand, and show me the verdict line." It is a well-defined task with a testable end state, which is exactly the kind of thing an agent does extremely well and you do badly on a Sunday.
The drill worth running once a year. Sit down and answer this honestly: suddenly every one of my servers is offline and inaccessible, including for backups — how long until I am operational again? If the answer is "a few hours", that is completely fine, and better than most companies. If the answer is "I would not be able to recover", you have just found the most important thing on your to-do list. It is a good question to ask an agent too — it will read your actual setup rather than the one you think you have.
The thing I got wrong: I truncated the wrong table on a live project. About twenty years in, twenty minutes before showing a client the thing. Fine, I thought — there is a backup right here. So I restored it. Except it was a .sql dump of the entire database, not of the one table I had just emptied, and I found this out afterwards rather than before. Two mistakes, and the first one was not the expensive one.

What changed after that was not "be more careful", because I had already been telling myself that for two decades and it does not work. What changed was mechanical: different terminal colours for live machines, one SSH session open at a time, and never a restore command typed until I have read what is actually in the file. Everyone truncates the wrong table eventually. The difference between a senior person and a beginner is not that it stops happening — it is that by then they have a plan B, and a C, and a D.

The day you think you have been hacked

The question people ask that day is how do I remove the malicious files. I have answered it on forums for years, and the honest answer is that it is the wrong question, because it assumes the thing you found is the thing that is there. You found the one that used all the CPU. Whoever put it there had a shell on your machine, as root or as near to it as they needed, for some length of time you do not know, before they got round to that. And every tool you would use to look for the rest — ls, ps, ss — runs on the machine you no longer trust, where a rootkit's first job is to teach ls to lie. You cannot audit a machine with the machine. Cleaning it up is asking the burglar whether anything is missing.

Before any of that, though: is it a break-in at all? Most of what feels like one is the weather. On this box, over ten days, 14,910 login attempts arrived for 2,071 usernames that do not exist, from 1,792 addresses — admin and ubuntu first, then crypto, wallet and blockchain, because what they are hoping to find is a coin wallet. Seventy-five logins were accepted, every one of them with a key and none with a password, and fail2ban banned two addresses and let the rest fail. None of that is a hack. It is Lock it down doing its job in public. A hack is one of these, and each is one command, run from a fresh terminal:

# A login you did not make. Ubuntu keeps these in the journal now;
# there is no wtmp, so the old `last` command prints nothing.
sudo journalctl -u ssh --since -30d -o cat | grep Accepted

# A key you did not add. You know how many lines there should be.
sudo cat /root/.ssh/authorized_keys /home/*/.ssh/authorized_keys

# A process you cannot account for, and who it is talking to
ps -eo user,pid,pcpu,etimes,cmd --sort=-pcpu | head -15
sudo ss -tnp state established

# A job you did not schedule
ls -la /etc/cron.d /var/spool/cron/crontabs 2>/dev/null
systemctl list-timers --all

# Files changed in the last two days where files should not change.
# Read it, do not count it: apt and adduser touch /etc for a living.
sudo find /etc /usr/local/bin -mtime -2 -type f

If any of those shows you something you did not do, the box stops being a patient and becomes a witness, and what happens next is the first ten minutes from the section on your own mistakes, with one change: nothing gets undone, ever, because nothing on it can be believed.

  1. Snapshot it at the provider, then power it off at the provider. The snapshot is the copy of the wreckage. The power button in the control panel is the one thing about the machine an intruder cannot argue with, which is why it is not shutdown typed into the box.
  2. Rotate every secret it ever held. Everything in .env, the ⁠Cloudflare token, the credential gh stored, the deploy key in Actions, and the key the backup script uses to reach the bucket — then go and look at the bucket, because that key can delete your backups and you want to know now whether it did. Your own SSH key only if its private half was ever on the box, which is the reason it never should have been.
  3. Rebuild on a fresh machine. Twenty minutes, the repository, the unit. This is the recovery from "everything is gone", and it is the same forty-five minutes, because you built it to be.
  4. Restore the data from a copy dated before the break-in, not from last night. That is the entire argument for the weeklies and monthlies above: the gap between the damage and your noticing it. And ask, once, whether the data was the door — an upload that was really a program, a row that was really a command — because restoring the door reinstalls it.
  5. Only then, ask how. On the snapshot, read from a different machine, never booted. It is one of the five ways from the hardening section, nearly always, and if you cannot find which, you will be doing this again in a month.

Which is why the drill above runs every night rather than once a quarter. The exercise I would give every team, and every person with one box, is this: suddenly all of your servers are offline and out of reach, backups included — how long until you are operational again? If the answer is "a few hours, probably", you are fine, and you will be fine on this day too. If the answer is "I could not recover from that", nothing in this section will help you when it comes, and the thing to fix is that answer, today, on a Tuesday, while the box is merely being scanned.

06

Habits that compound

The difference between having one project and having a hundred.

The creature watering a row of four plants, each taller than the last.

Launch checklist

The small things that make something look finished. Your ticks are saved.

A short half-built brick wall, a small figure at its unfinished end laying one more brick.
Nothing here is clever. It is one small thing, done again, until you have a hundred projects instead of one.

Not one of these is hard. Every single one is forgettable, and together they are the whole difference between something that feels shipped and something that feels like a draft somebody left open. Work down the list with your agent — most items are one sentence of instruction and thirty seconds of its time.

I am not being theoretical about the forgetting. I went and counted across my own repositories: favicons in about half of them, an apple-touch-icon in slightly fewer, a web manifest in a third, a robots.txt in a quarter, a sitemap in barely one in eight. That is a descending staircase and it has exactly one cause, which is that I got to the end of a project, felt finished, and stopped. This list exists so that "felt finished" and "is finished" line up more often for you than they have for me.

0 / 0
Steal the checklist from projects you admire. "Find three popular open-source Go web projects on ⁠GitHub. List every file in their repository root and their static directory that we do not have. Tell me which of those we should add and why." Conventions exist for reasons, usually reasons somebody learned the hard way, and this is the fastest way to absorb a hundred of them without having to earn any of them yourself.

Monorepo to the moon

Use the right language for each job. The tax for doing that has mostly gone.

The old reason to pick one language and stay there was entirely human. Every extra language was another set of idioms, another toolchain, another thing to remain fluent in while not using it. Nobody is excellent at six at once, so teams standardised, and they were right to.

That constraint has largely dissolved, and it is one of the genuinely new things about working this way. An agent is competent in all of them at the same time and does not get rusty. So the question stops being "what do I know?" and becomes "what is actually right for this piece?" — which is a much better question and one you were never able to afford before.

What that looks like in practice

/srv/myproject/
├── justfile              one entry point for the whole thing
├── docs/
├── web/                  Go — the site and the API
├── worker/               Rust — the CPU-heavy image pipeline
├── scripts/              Python — data wrangling and one-off jobs
├── desktop/              Electron — the installable companion app
└── deploy/               nginx, systemd, the workflows

One repository. One history. One just check that runs all four test suites. The pieces talk over HTTP or a queue, so any one of them can be rewritten or thrown away without the others noticing — which is the actual point of microservices, as distinct from the version people cargo-cult, where each service gets its own repo, its own pipeline, its own deployment ceremony and its own reasons to be broken on a Friday.

LanguageReach for it when
GoAnything that serves requests or runs as a service. The default, and it should be.
⁠RustGenuinely low-level work, or compute so heavy the runtime cost is the product. Not for a CRUD app, whatever the internet says.
⁠PythonData, scripting, machine learning, anything where the library you need only exists there.
JavaScriptThe browser, and ⁠Electron. It is the only option in one place and a fine one there.
C / C++Talking to hardware, or wrapping a library that has no bindings anywhere else.
SQLNot optional. Every stack has it, and it is worth understanding rather than hiding behind a layer that generates it badly.
The failure mode is fragmentation, not variety. Adding a language is a decision and it should have a reason written down in docs/DECISIONS.md, in a sentence, on the day you make it. Six languages because each piece genuinely needed one is engineering. Six languages because each was fashionable in the month that part got written is a maintenance problem you have posted to your future self, and he cannot refuse delivery.

Every language, and what it is actually for

Here is the thing I wish somebody had told me twenty years ago, because it would have saved me a decade of extremely enjoyable arguments. The useful facts about a programming language are almost never the ones people fight about. Whether it is object-oriented, functional or procedural; whether it has classes or traits or interfaces; whether the syntax is pretty — those are the things beginners are taught to compare, and they are close to irrelevant when you are choosing one for a job.

What decides it is duller and much more consequential. Where can this thing run? JavaScript wins the browser by walkover, because it is the only language the browser has. What can it reach? Swift can ask an iPhone for its health data and nothing else can. Does the library I need exist in it? If the model you want to run only has Python bindings, the language argument is already over. And what does shipping it cost? A Go binary is one file you copy; a Node service is a runtime and a thousand packages you have to keep aligned.

So: the map. Columns are where a language primarily runs, which is the property that matters most and is almost never how these charts are organised. The lines are real relationships — one compiles to another, one is written in another, or the two are so routinely paired that learning one drags in the other. Click anything to see what it is for, what it is genuinely good at, and one honest criticism, because every language on here has at least one and the ones I like best get theirs too.

A map of where things run, not a ranking. Tap a language for the honest version.

JavaScriptthe browser

For Making a page do something when somebody touches it. It is the only language a browser runs, which makes it the only language in the most widely deployed runtime ever built.

Strong Unavoidable, therefore universal. Enormous ecosystem, instant feedback loop, and every device on earth already has an interpreter for it.

Fair dig Designed in ten days in 1995 and it shows in the corners — the equality rules alone have cost the industry decades. The correct way to do any given thing also changes every two years, which is worse for an agent than for you: its knowledge of a fast-moving ecosystem is a photograph with a date on it.

TypeScriptthe browser

For JavaScript with the types written down. It compiles to JavaScript and runs everywhere JavaScript runs, which is everywhere.

Strong The editor can answer questions about your own code, and a whole class of "undefined is not a function" at eleven at night simply stops happening. On anything beyond a few hundred lines it pays for itself quickly.

Fair dig It is a type system bolted onto a language that does not have one, so the types are erased at runtime and they can lie to you. Also a build step you did not have before, which is a real cost on something small.

Goyour server

For Services, tools and anything that answers requests. Built at Google explicitly to be readable by people who did not write it.

Strong go build makes one static file with no dependencies — copy it to the server, run it, that is the deploy. Doing many things at once is boring to write rather than clever. The standard library covers HTTP, TLS, templates and JSON without a single third-party package. And the language has barely changed in a decade, which means what an agent learned about it is still true.

Fair dig Deliberately plain to the point of being repetitive; you will write if err != nil until you dream about it. Generics arrived late and are still a bit awkward. If you love expressive type systems you will find it joyless, and that is a fair thing to want.

Pythonyour server · data

For Data, scripting, automation, and every single thing to do with machine learning. If a model, a dataset or a scientific library exists, it has Python bindings first and possibly only.

Strong You can read it out loud. The library situation is not a close contest — it is the whole reason the field standardised on it. Superb for the script you write once and run for six years.

Fair dig Slow on raw computation, which mostly does not matter because the fast parts are C underneath. Deployment is its real weakness: versions, virtual environments and system packages produce the phrase "works on my machine" more reliably than any other language here.

PHPyour server

For Websites, which it has done since 1995 and still runs a large fraction of the web on, whatever anybody tells you at a conference.

Strong The deployment model is genuinely the simplest there is: a file in a folder is a page, and there is no build and no process to keep alive. Modern PHP with Laravel is a pleasant, fast, well-documented way to build a real application, and hosting for it costs nothing anywhere on earth.

Fair dig Twenty years of accumulated inconsistency in the standard library, and a reputation formed in an era it has genuinely left behind — which still costs you, because the tutorials you find are as likely to be from 2009 as from 2026. I learned on procedural PHP before it had real object orientation, so the affection here is not neutral.

Node.jsyour server

For Running JavaScript off the browser. One language for both halves of a web application, which used to be the strongest argument in this entire section.

Strong The largest package registry in existence, excellent at many simultaneous connections, and nobody on your team has to learn a second language. ⁠Deno and ⁠Bun are newer runtimes for the same language with better defaults and a fraction of the ceremony.

Fair dig A modest service pulls in hundreds of megabytes and thousands of packages, every one of which can break, change under you or be taken over — and on a 4 GB box, the runtimes add up. The "one language everywhere" argument was strong when a human held both halves in their head, and is worth considerably less when an agent is typing both.

C#your server · desktop · games

For Business applications, Windows desktop software, and — via Unity — an enormous share of all the games you have ever played.

Strong One of the best-designed large languages in use, with tooling that is genuinely excellent and a standard library that has thought of your problem already. Modern .NET is fast, cross-platform and open source, which surprises people whose impression of it is fifteen years old.

Fair dig Still carries a cultural gravity towards Microsoft's way of doing things, and the ecosystem assumes a certain size of organisation — a lot of documentation is written for a team with an architect. Heavier to deploy on a small Linux box than Go, though far lighter than it used to be.

Javayour server · Android

For Systems that have to run for fifteen years and be maintained by people who have not met each other. Banks, insurers, logistics, Android's foundations, and a colossal amount of the internet's plumbing.

Strong The JVM is one of the great engineering achievements of the field — the garbage collectors alone are extraordinary — and the ecosystem is deep, stable and extremely well documented. Nothing you write will be the first thing of its kind.

Fair dig Verbose in a way that stopped being fashionable, memory-hungry enough to matter on a small server, and its enterprise culture produced some of the most famously over-engineered code ever committed. The language itself has improved enormously; the reputation has not caught up.

Rubyyour server

For Web applications, mostly with Rails, which invented or popularised half the conventions every other web framework now copies.

Strong Optimised for the happiness of the person typing, and it shows — it is a delight to read and write. Rails will get one person from nothing to a working, tested, deployed application faster than almost anything else in this list, which is exactly the property that matters most for a first project.

Fair dig Slower than the compiled options and memory-hungry at scale. The magic that makes it so pleasant is also genuinely hard to debug when it goes wrong, because the method you are looking for may not appear anywhere in the source. Hiring and fashion have moved on, which affects how much help you find.

Elixiryour server

For Systems that must not go down, and anything with a very large number of simultaneous connections — chat, presence, live dashboards, telemetry.

Strong It runs on the Erlang virtual machine, which was built by a phone company for switches that were not allowed to fail, and it shows: processes are nearly free, failure is isolated and supervised rather than fatal, and Phoenix LiveView lets you build a live-updating interface with almost no front-end JavaScript at all.

Fair dig A small community, so the "somebody has already solved this" property is weaker than anywhere else on this map. Functional and concurrent-by-default is a genuine shift in how you think, which is a cost even though it is a worthwhile one. Not a first language for a first project.

Cthe metal

For Operating systems, drivers, embedded devices, and the inside of nearly everything else on this page. Your Python interpreter is C. So is your database, your web server and your kernel.

Strong Small, fast, portable to hardware that has never heard of anything else, and stable for fifty years. It is the lingua franca every other language speaks when it needs to talk to something: if a library exists anywhere, it has a C interface.

Fair dig It will let you do absolutely anything, including the thing that corrupts memory somewhere else entirely and crashes an hour later. A significant share of every serious security vulnerability of the last thirty years is a memory bug in C or C++. That is not a moral failing of the language; it is a property you must budget for.

C++the metal · games

For Anything where the last ten per cent of performance is the product: game engines, browsers, trading systems, computer vision, audio.

Strong Zero-cost abstraction is a real thing and C++ delivers it — you can write expressive code that compiles to exactly what you would have written by hand. Nothing else has quite this combination of control and altitude, and Unreal, Chrome and most of the software you use daily are built on it.

Fair dig Enormous. There are several C++ dialects inside C++ and no two codebases agree on which they use, so "knowing C++" means less than knowing most languages here. Compile times are a genuine daily cost, the error messages are legendary, and it inherits C's memory model along with its consequences.

Rustthe metal · the browser, via wasm

For Systems work where memory bugs are unacceptable, compute heavy enough to show up on an invoice, and command-line tools you want to be perfect.

Strong It proves at compile time that a whole category of bug cannot happen, without a garbage collector. Cargo is the best package and build tool of anything on this map, the compiler's error messages genuinely teach you, and it compiles to WebAssembly so the same code can run in a browser.

Fair dig The learning curve is real and the borrow checker will beat you for a fortnight. It is the wrong tool for a CRUD application no matter what the internet says — you will spend your attention on lifetimes instead of on the thing you are building. Slow to compile, and there is a strain of evangelism around it that has done the language no favours.

Zigthe metal

For The job C does, with fifty years of hindsight applied. Also, oddly, one of the best cross-compilers in existence — people use it to build C projects for other platforms without writing a line of Zig.

Strong Simple enough to hold in your head, explicit about allocation in a way that makes memory behaviour obvious, no hidden control flow, and it compiles C directly so you can adopt it one file at a time.

Fair dig Not finished. The language is still changing under its users, the ecosystem is small, and an agent's knowledge of it is likelier to be out of date here than anywhere else on this map. Wonderful, and not where you put something that has to work next year without attention.

Swifttheir iPhone · Mac

For Apple's platforms, properly. If your app needs HealthKit, a watch face, a widget, CarPlay or the camera pipeline at full depth, this is the language and there is no second answer.

Strong A genuinely modern, pleasant language with real safety guarantees, and SwiftUI builds interfaces quickly. It reaches every API Apple has, the day Apple ships it — which is the entire reason to be here.

Fair dig Practically speaking it is one company's platform: you need a Mac, a yearly developer programme, and a tolerance for a yearly round of work when the OS moves. It runs on Linux servers in theory and almost nobody does it. Excellent language, narrow door.

Kotlintheir Android · your server

For Android, officially and overwhelmingly. Also a very good server language, and via Multiplatform it can share business logic between an iPhone app and an Android one.

Strong Everything Java does, with about half the typing and null-safety built into the type system, on the same superb JVM and with the whole Java ecosystem available unchanged. Jetpack Compose is a real pleasure compared to what Android development used to be.

Fair dig Android itself is the tax, not the language: build times, emulators, and a fragmented device population you cannot test properly. On the server it is excellent but always slightly the second choice to Java in terms of how much has already been written down.

SQLyour database

For Asking questions of data. Not optional, not replaceable, and the one language on this map you will still be using in twenty years whatever else changes.

Strong Fifty years old, declarative, and astonishingly good at its job — you say what you want and something cleverer than you works out how to get it. It is also the highest-leverage thing on this page to learn properly: one correct query routinely replaces two hundred lines of application code that loops.

Fair dig Every database speaks a slightly different dialect, so "knowing SQL" does not perfectly transfer. It is easy to write something that works on a thousand rows and falls over at a million. And the tooling that hides it from you — the object-relational mapper — will generate queries you would be appalled by, which is exactly why you should be able to read them.

Bashthe shell

For Gluing programs to each other. Not writing applications. Every server you will ever touch already has it, and so does every deploy pipeline.

Strong Unbeatable at its actual job: take the output of one thing, feed it to another, do it a thousand times. Learning pipes, redirection and exit codes properly will pay you back every working day for the rest of your career, and most of this tutorial is that skill.

Fair dig Past about a hundred lines it becomes a language you do not want to debug at midnight — quoting rules that punish you for a space in a filename, no real data structures, and error handling you have to remember to ask for. Everybody learns this the same way, by writing a 400-line script and then rewriting it in Python.

That is eighteen and it is not everything. Dart carries Flutter; Lua is embedded in more games and devices than you would guess; R still owns whole corners of statistics; Haskell, OCaml and Clojure will each change how you think about code even if you never ship a line of them; Fortran and COBOL are both, right now, running things you depend on. None of them are missing because they are bad.

And the actual point of all of that. Every single card up there has a fair dig in it, including the ones I use every day and the one this website is written in. That is on purpose. A person who can tell you what is wrong with their favourite language is somebody you can trust about languages; a person who cannot is telling you about their identity rather than about the tool. Pick the one that reaches the thing you need, that has the library, and that you can ship on Friday — then get on with building, which is the part that was always the point.

Everything in this guide is one set of preferences that happens to work. Take the parts that suit you and bin the rest with my blessing — the thing that matters is not my conventions, it is that you have conventions at all, that they are written down, and that they compound instead of being reinvented every January. Games, marketing tools, dashboards, a side project that quietly turns into the real one: same skeleton, same box, same four words at the terminal.

What one box holds

Honestly: far more than you expect, right up until it suddenly does not.

Four gigabytes of memory is the binding constraint. Not disk — sixty gigabytes is an enormous amount of text. Not CPU — two modern cores are genuinely quick, and most web requests are the machine waiting for something rather than thinking. Memory is what you run out of, so it is worth carrying a rough idea of what things cost.

ThingRoughly
⁠Ubuntu itself250–400 MBUnavoidable.
⁠nginx~10 MBServes every site on the box for that.
⁠SQLite0 MBThere is no process. It runs inside your program, which is the entire point of it.
⁠PostgreSQL60–150 MBOne instance is plenty; give each project a database, not a server.
A Go service8–25 MBYou can run dozens. This is the whole argument for Go on a small box.
A ⁠Python service60–150 MBFine. A few of them, also fine.
A ⁠Node service60–200 MBPlus the runtime. They add up faster than you think.
A JVM app400 MB+One is a real commitment on this machine.
Elasticsearch1 GB+Do not. Postgres full-text search is genuinely good.
Comfortable
  • Dozens of small Go services, each on its own subdomain and port.
  • A SQLite file per project, or one Postgres holding every project's data.
  • Any number of static sites — they cost essentially nothing.
  • Real traffic. Thousands of visitors a day on a well-built site is not close to the limit, and the log says how many you are getting.
  • Compiling, testing, and running your agent all day.
Will hurt
  • Running a large language model locally. It will not fit; use an API.
  • Video transcoding. It works, but it will occupy both cores for a long time.
  • A dozen Node processes at once.
  • npm install on a large project while everything else is running — the memory spike is the problem, not the average.
  • Anything storing large media on the local disk instead of object storage.

When it feels slow

htop                          # what is using CPU and memory, live
free -h                       # how much memory is actually left
df -h                         # disk — check this before it bites
sudo journalctl -p err -n 50  # the last 50 errors from anything
systemctl --failed            # anything that died and did not come back

# The most useful single question you can ask the box:
dmesg -T | grep -i 'killed process'

If you want to know what this box does under real load rather than what it is doing right now, Stress testing is the same machine with somebody leaning on it.

That last one finds the out-of-memory killer, and it is the single most useful question you can ask this machine. When Linux runs out of memory it picks a process and kills it, and the service simply vanishes without logging an error of its own, because it did not get the chance. If something disappeared for no apparent reason and the logs end mid-sentence, that is nearly always what happened.

Who came, without spying on them

The first thing everybody does the day after launching is go and find out whether anybody came, and the way most people find out is to paste a script from a company that hands it out free. It is free because your readers are what it costs. Every one of them is followed to your page and then off it again, on your behalf, and in return you get a chart. You do not need the chart yet. You need one number, and you already have it: nginx has been writing a line per request into /var/log/nginx/access.log since the hour you installed it, and awk has been on every Unix since 1977. That log is an analytics database. It just does not have a dashboard, which is what the four lines below are for.

One line looks like this. This is a real one from this box, and I picked it because it is the funniest line in ten days of log: a stranger trying a hole from 2014 on a server built in 2026, by putting a shell command where the referrer goes. The address is printed whole because it is a rented machine in a hosting company's range, scanning the whole internet, and nobody lives there. A reader's address is a different thing, and this section never prints one.

156.239.229.239 - - [08/Sep/2026:22:44:10 +0000] "GET /cgi-bin/test.cgi HTTP/1.1" 404 19 "() { ignored; }; echo Content-Type: text/html; echo ; /bin/cat /etc/passwd" "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.5 Safari/605.1.61"

Address, time, request, status, bytes, then two quoted fields: where the visitor claims to have come from and what they claim to be. Claims, both times — the visitor typed those, and the line above is a visitor typing something else entirely. The address and the status are the only two facts on the line. Everything you want to know is a question about those columns, and awk counts by column:

log=/var/log/nginx/access.log

# How many people today: distinct addresses that loaded the front page.
sudo grep "$(date +%d/%b/%Y)" $log \
  | awk '$7 == "/" && $9 == 200 {print $1}' | sort -u | wc -l

# Which pages, this fortnight. Column 7 is the path, column 9 the status.
sudo awk '$9 == 200 && $7 !~ /^\/(static|api)\// {print $7}' $log \
  | sort | uniq -c | sort -rn | head

# Where they came from. The referrer is the fourth quoted field.
sudo awk -F'"' '$4 != "-" {print $4}' $log | grep -v mysite.com \
  | sort | uniq -c | sort -rn | head

# And what everybody who was not a person was looking for.
sudo awk '$9 == 404 {print $7}' $log | sort | uniq -c | sort -rn | head

The last one is the one to run first, because it tells you what the other three are counting. Here it is on this box, over the ten days of log it had when I ran it:

$ sudo awk '$9 == 404 {print $7}' $log | sort | uniq -c | sort -rn | head -8
    145 /auth/callback
    135 /account
    132 /signup
    131 /signin
    130 /register
    129 /admin
    108 /login
    108 /api/auth/signin
$ sudo awk '{n++} $9 == 404 {f++} END {printf "%d lines, %d not found (%.0f%%)\n", n, f, 100*f/n}' $log
23770 lines, 14826 not found (62%)

Sixty-two percent of everything that has ever arrived at this site was somebody trying a door that does not exist. A hundred and one requests for /.env, sixty-two for /.git/config, three hundred for a WordPress endpoint on a site with no WordPress, and a whole set for the login pages of a framework I have never installed. None of it is aimed at you. It is the weather, it arrives at every address on the internet within hours of the address existing, and the reason Lock it down came before anything else in this guide is sitting right there in column nine. The other lesson is who the biggest visitor was: this box. Three thousand four hundred lines, more than anyone, from the screenshot tool and the smoke test that check the page after every deploy. Your top visitor will be you too. The first thing to do with any count is take yourself out of it.

Then the number you were after. Yesterday, seventy-two distinct addresses loaded this front page, and twelve of them also fetched the stylesheet. A person's browser fetches both; a crawler takes the page and leaves; a person who has been here before has the stylesheet cached and takes the page only. So the honest answer is "somewhere between twelve and seventy-two", and it will stay a range no matter what you install, because the thing you want to count is whether a human was on the other end, and the line does not say. The paid charts do not know either. They round confidently.

The name a visitor gives is not evidence. More than two thousand lines here carry the name of a well-known crawler — a search engine, an AI company's agent, a shopping giant's — and 1,953 of them are 404s for /.env, /.s3cfg and /api/proc/self/environ, which no search engine has ever wanted. That is a scanner wearing a borrowed coat, because some sites let those names past the door. Count by address. And behind Cloudflare's orange cloud the address is Cloudflare's too, on every line, until the snippet in Not getting flooded is installed.

When you do want the chart, get it from the same log rather than from a script in your page. goaccess is one package, reads the file nginx already writes, and on this box turns ten days of it into a report in a quarter of a second:

sudo apt install goaccess
sudo goaccess /var/log/nginx/access.log --log-format=COMBINED          # live, in the terminal
sudo goaccess /var/log/nginx/access.log --log-format=COMBINED -o /srv/myapp/data/stats.html

The second form writes one self-contained HTML file, about a megabyte, that you can open from the mounted drive or serve from a location behind the same auth_basic password as anything else private. Nothing about a reader leaves the machine either way. There is no cookie, so there is nothing to put a banner up about, and the address is in the log for the fortnight Ubuntu's logrotate keeps by default and then it is gone.

When your app wants its own numbers

Sooner or later a project needs a count the log cannot give it — views per listing, plays per track, whatever the thing is that your users make — and the first version everyone writes, me included, is UPDATE things SET views = views + 1 in the page handler. It works for a year. Then the page that lists the things sorts by that column, the report that graphs it walks every row, and you have built the invite tree again with a smaller tree. Seven of my own Go projects count things, and after enough versions they all count the same way, in two halves that never meet in a request:

  • The app writes what happened, in batches. A page view goes into a queue in memory, and once a second one transaction writes the queue to a hits table — a thousand rows in the time one row would take, because the transaction is what costs. The page never waits on its own counter, and if the queue is ever full the view is dropped, which is the correct trade: a count that is a little low is a count; a page that is down is not.
  • A timer turns rows into totals. Once an hour, one query groups the raw rows into an hits_hourly table, and everything that has to be fast reads that. It is the same shape as the timer that deletes old sessions, and it is a job, not a page.
sqlite3 /srv/myapp/data/app.db
-- One row per page view, written by the app in batches. No name and no
-- address: a visitor is a salted hash the app draws fresh every day, so a
-- person is one visitor for a day and nobody the day after.
CREATE TABLE hits (
  at      INTEGER NOT NULL,   -- unix seconds
  path    TEXT    NOT NULL,
  visitor TEXT    NOT NULL
);
CREATE INDEX hits_at ON hits (at);

-- What the charts read. The raw rows go after a week; this is what you keep.
CREATE TABLE hits_hourly (
  hour     INTEGER NOT NULL,  -- unix seconds, on the hour
  path     TEXT    NOT NULL,
  views    INTEGER NOT NULL,
  visitors INTEGER NOT NULL,
  PRIMARY KEY (hour, path)
);
.quit
-- /srv/myapp/sql/rollup.sql, run by the timer below. Every finished hour
-- that still has raw rows is recomputed, so a row that arrived late is
-- counted next time, and a night the timer missed catches up by itself.
.timeout 5000
INSERT OR REPLACE INTO hits_hourly (hour, path, views, visitors)
SELECT (at / 3600) * 3600, path, COUNT(*), COUNT(DISTINCT visitor)
FROM hits
WHERE at < (unixepoch() / 3600) * 3600
GROUP BY 1, 2;

DELETE FROM hits WHERE at < unixepoch() - 7 * 86400;
# /etc/systemd/system/myapp-rollup.service
[Unit]
Description=Roll page views up into hourly counts

[Service]
Type=oneshot
User=myapp
ExecStart=/usr/bin/sqlite3 /srv/myapp/data/app.db ".read /srv/myapp/sql/rollup.sql"
ProtectSystem=strict
ReadWritePaths=/srv/myapp/data
# /etc/systemd/system/myapp-rollup.timer
[Unit]
Description=Roll page views up, hourly

[Timer]
OnCalendar=hourly
Persistent=true

[Install]
WantedBy=timers.target

.timeout 5000 is the line people leave out: the app is writing to the same file, and without it the one time the two collide the rollup fails instead of waiting five seconds. The visitor column is a hash of the address and the browser with a random salt the app throws away each day, so the table cannot be turned back into people, and the same person on Tuesday and Wednesday is two visitors, which is the right kind of wrong. That rollup ran on this box against a scratch copy before it was printed here, with a row inserted late and a row nine days old, and did what the comment says. The hourly table for a site with a hundred pages and a year of history is under a million rows, which SQLite will sum for a chart in the time it takes the chart to draw. The raw table never gets big, because it is not allowed to.

Outgrowing it is the good outcome. A bigger VPS is a checkbox and a reboot. And because everything here is a git repo plus a handful of config files that also live in that repo, moving to an entirely different machine at an entirely different company is an afternoon's work. Nothing you are building is locked to this box, which is worth more than the box.

Staying up to date, and what "end of life" really means

Nobody is coming for you personally. That is the point: the scanning is automatic, it is constant, and it does not need a reason to arrive at your address. The only thing that decides whether it finds anything is how long you left the door open.

A tall server cabinet at the right with one small panel missing from its front, and the small creature far to the left holding up the matching panel.
The scanning is automatic and it does not need a reason. The only thing it can find is the panel you left off.

The least glamorous habit in this guide is the one with the highest return, and it is this: run software that is still getting fixed, and take the fixes. That is most of security for a box like yours. Not a firewall you tuned for a weekend, not a clever nickname for the SSH port — just being current.

Why this got worse, with numbers

It is easy to think of security holes as rare events. They are a firehose, and it is widening. Public vulnerability reports went from about 14,600 in 2017 to 48,000 in 2025, and 2026 is on track to pass 65,000 — roughly 250 new ones every day. Nobody reads that. Nobody is supposed to. It is a machine problem now on both sides.

And the window between "a fix exists" and "somebody is using it against you" has collapsed. Google's incident responders measured the average time from a vulnerability becoming public to being exploited: 63 days in 2018, 32 in 2021, five days in 2023. Their 2026 report puts it at minus seven days — on average, exploitation now begins before the patch exists. ⁠Cloudflare watched attacks on one server product start 22 minutes after somebody published proof it was possible. Roughly 29% of the vulnerabilities that end up on the US government's "known exploited" list were already being exploited on or before the day they became public.

The traffic arriving at your box reflects that. More than half of all web traffic is automated, and about 40% of it is hostile bots — a number that has gone up every year for a decade. Verizon's 2026 breach report found that 31% of breaches now start with somebody exploiting a vulnerability, which for the first time in nineteen years of that report beats stolen passwords. None of that is aimed at you. It is aimed at every address there is, and yours is one.

And now both sides have agents. In 2025 an AI system at Google found a vulnerability that attackers already knew about and were using. A US defence-research competition had machines find 54 of 63 planted bugs in 54 million lines of code, and patch most of them, for about $152 a go. An autonomous tool reached the top of a bug-bounty leaderboard against human researchers. And Anthropic published a case where somebody turned an agent into roughly 80–90% of a real intrusion campaign. The honest counterweight: of a thousand-odd vulnerabilities found by AI so far, only about 1.3% have been seen exploited. The scanning is getting cheaper faster than the attacking is getting smarter — but "cheaper scanning" is precisely the thing that finds an unpatched box.

How end of life works

Every piece of software you run has a date after which nobody will fix it, whatever is found in it. Not a date it stops working — that is the trap. It keeps working perfectly, serving pages, right up until the day it is the reason you are reading your logs at two in the morning. The dates are published years in advance and almost nobody looks at them.

ThingFixed forThen
⁠Ubuntu LTS5 yearsUpgrade, or extend
Ubuntu Pro free10 years5 machines free
⁠PostgreSQL5 yearsMajor upgrade
⁠GoAbout a yearRebuild on a newer one
⁠Node.js30 monthsMove up a version
⁠PHP2 years, +2 securityMove up a version
⁠Python5 yearsMove up a version

A new Ubuntu LTS arrives every two years and you upgrade from one to the next, never skipping; PostgreSQL ships small fixes quarterly inside those five years, which are a restart rather than an afternoon; and Go counts its year in releases rather than months — a version is supported until two newer ones exist.

# What is the thing you are running, and when does it stop being fixed?
# endoflife.date tracks nearly everything. jq is `sudo apt install jq`.
curl -s https://endoflife.date/api/v1/products/ubuntu/releases/latest \
  | jq -r '.result | "\(.label) — security support ends \(.eolFrom)"'

# Swap ubuntu for postgresql, nodejs, php, python, go, nginx, debian...
Your version number is not upstream's version number, and that is fine. This box serves everything through ⁠nginx 1.28.3, and nginx's own website calls that branch legacy. It is not out of date: Ubuntu takes the security fixes from newer versions and applies them to the version they shipped, so the number stays still while the holes get closed. That is the deal you are buying with an LTS, and it is why "just install the newest of everything from the internet" is usually worse, not better. The deal has one edge, and it bit me — see below.

What to actually do, in descending order of value

  1. Turn on automatic security updates and then check they are running. Step 6 of Lock it down switches them on; installing it is not the same as confirming it. The commands below take ten seconds and answer "is this actually happening on my machine".
  2. Reboot when the kernel says so. An updated kernel is not the running kernel. Ubuntu leaves a file behind to tell you, and it is the single most-ignored file on every server on earth.
  3. Attach Ubuntu Pro if it is a personal box. It is free for five machines, and the thing it buys you is security support for the community-maintained half of the archive — which is where most interesting software lives, and which the normal five-year promise does not cover.
  4. Scan your own project's dependencies, which the operating system knows nothing about. Every ecosystem has a one-liner: govulncheck for Go, npm audit, pip-audit, and ⁠GitHub's Dependabot will open the pull requests for you if the code is there.
  5. Put the database's major version in your calendar, because that one needs planning rather than a restart. Minor versions are a restart. Majors are an afternoon.
apt list --upgradable                    # what is waiting
ls /var/run/reboot-required              # no such file = no reboot needed
systemctl list-timers 'apt-daily*'       # when the automatic run happens
journalctl -u unattended-upgrades --since '-7 days'

# The good one: exactly what tonight's automatic run would install.
sudo unattended-upgrade --dry-run -v
# Your own code's dependencies, which apt has never heard of.
# Go, from the project directory — it reports only what you actually call:
go run golang.org/x/vuln/cmd/govulncheck@latest ./...
The thing I got wrong: I wrote this section on a box whose compiler was eight security releases out of date. Ubuntu ships Go in the community half of the archive, which gets updated when somebody gets round to it, and nobody had — it was still the release from February. Automatic updates were on, working, and had nothing to say about it, because that promise does not cover that half of the archive. So the binary serving this page was compiled with a toolchain carrying twenty known holes that this code actually reaches, four of them in the very thing that renders every page you are reading. One command found it in thirty seconds. I had never run that command on my own project.

The fix took one line, because Go can fetch and verify its own compiler; the same shape of fix exists in every ecosystem. The lesson is not about Go. It is that "automatic updates are on" answers a narrower question than the one you think you asked, and the only way to find the gap is to point a scanner at your own project rather than at your operating system.
A live example, from this box, while writing this. At the moment of writing there is a security update for nginx sitting in the queue here, along with seven others. It is not installed yet, and nothing is wrong: the automatic run happens once a day, and the dry-run above confirms it will take it at the next one. That is the system working exactly as designed — and it is also why "I turned on automatic updates" and "my box is patched right now" are two different sentences. If a hole is loud enough to have a name and a news article, do not wait for the timer. Run it yourself.

The upgrade you are putting off

Somewhere there is a project of yours on a version of something that went out of support during a previous presidency. The reason it has not been upgraded is never technical — it is that the upgrade has no visible reward. Nothing looks better afterwards. The button is in the same place. You spend a weekend and the client cannot tell.

Do it anyway, and do it before it is urgent, because the cost only goes one way. Two versions behind is an afternoon. Six versions behind is a rewrite with a different name. And the day it becomes urgent is, by definition, the day you least want a large, uncertain, all-at-once change to your production system: you will be doing the upgrade and the incident at the same time, having lost the option to do them separately.

This is also the single best use of an agent I have found, and the reason is that the work is enormous, mechanical, boring, well-documented, and continuously testable. Point it at one version step at a time — not five — with the tests running after each, and a commit at every green point so you can walk back. It will do in an evening the thing you have been not-doing for two years.

Audit this box for anything out of support. Check the Ubuntu release and its
end-of-life date, every service I am running and its version against upstream,
whether unattended-upgrades is actually installing things, whether a reboot is
pending, and my project's own dependencies with the right scanner for its
language. Then give me one table: what is current, what is behind, what is past
end of life, and what you would do first. Do not change anything yet.
What this looks like when it goes wrong. The Equifax breach — 147 million people's personal records — was a vulnerability in a web framework with a fix available in March, in a system that was still unpatched when attackers walked in during May, and they were inside until July. There was no clever exploit and no genius attacker. Somebody did not apply an update that had been sitting there for two months. The bill was a settlement of at least $575 million, and it is the most expensive boring omission in the history of this industry.

It is down. What now?

Thirty sections on building the thing. This is the one about the night it stops working, which is the moment most people quietly give up and delete the box.

A dark switched-off tower with a small figure at its base holding up a lantern.
The order you check things in matters more than the commands do, because everybody checks what they touched last.

Your car will not start. You are certain it is the battery, because you left the lights on last night and the lights are the thing you touched. You buy a battery. It is not the battery. The mechanic finds a dead fuel pump in about four minutes, and not because he knows your car better than you do — he has never seen it. He is simply not the man who left the lights on, so he starts at the top of the same list he always starts at, and that list does not care what you did last night.

That is the whole trick, and it is the only thing in this section worth memorising. When your site goes down you will start debugging at the last thing you changed, because it is the freshest thing in your head and because some part of you already feels responsible for it. It is almost never there. Work the ladder instead, in order, top to bottom, one command per rung, and let the machine tell you where you are standing rather than guessing from the smell of smoke.

The questionWhy it is on this rung
1Is it me, or is it everyone?Load the site on your phone with wi-fi switched off. Costs five seconds and settles the biggest question there is.
2Is the machine even alive?ping, then ssh. If you cannot get in, nothing below this rung is knowable.
3Is your service running?systemctl status, then the last fifty lines of its log. Most outages end here.
4Is ⁠nginx up, and is its config valid?An invalid config does not take effect, so the site keeps serving the old one until something reloads it. Then it stops.
5Is the certificate still good?Certificates expire on a date, not on a trend. Nothing degrades first.
6Is it DNS?It is sometimes DNS. Ask what the world sees, not what your laptop remembers.
7Is the disk full?A full disk breaks things that look completely unrelated to disk, which is why it is worth reaching the bottom of the ladder.

Rung one is first because your own browser is a liar. It has a cache, your operating system has a DNS cache, and your network may have opinions of its own. A site that is down for everybody is an outage. A site that is down only for you is a Tuesday, and the two want completely different afternoons from you.

# 1. is it me, or is it everyone — ask from somewhere that is not this machine
curl -sI https://example.com | head -1

# 2. is the machine alive
ping -c 3 example.com
ssh you@example.com          # if it refuses you, run it again with -v

# 3. is the service running, and what did it say on the way out
systemctl status myapp --no-pager
sudo journalctl -u myapp -n 50 --no-pager

# 4. is nginx up, and does its config actually parse
sudo nginx -t
systemctl status nginx --no-pager

# 5. is the certificate still valid, and until when
sudo certbot certificates

# 6. what does the world think your domain points at
dig +short example.com

# 7. is the disk full
df -h

Rung three is where most nights end, and it is worth knowing the one failure that leaves no note. If the service is simply gone and its log stops mid-sentence with no error, that is usually the out-of-memory killer, which does not give a process the chance to complain on its way out. What one box holds has the command for that and the reasoning behind it.

The deploy that broke it, taken back in one command

Rung three has a favourite ending, and it is the one night the instinct is right: the service died in the minute after you shipped something. The log says it started, printed one line and stopped, and the freshest thing in your head really is the cause this time. What goes wrong is the next minute. The deploy from The shape builds over the only binary you have, so "put the old one back" means finding the commit, rebuilding it and hoping it was that one, at the exact moment you are least able to do any of those calmly.

Two changes make it a deploy you can take back. Every build goes into its own dated folder, and the service runs whatever a link called current points at. Deploying moves the link forward. Rolling back moves it to the folder before, and the folder you just left is still there, untouched, for when you have found the bug.

# In /etc/systemd/system/myapp.service, the one line that changes
ExecStart=/srv/myapp/releases/current/server -addr 127.0.0.1:8090
# Every build lands in its own dated folder; `current` points at the live one.
releases := "/srv/myapp/releases"
stamp    := `date -u +%Y%m%d-%H%M%S`

# Build into a new release, point current at it, restart, and confirm it came back
deploy:
    go build -o {{releases}}/{{stamp}}/server ./cmd/server
    ln -sfn {{stamp}} {{releases}}/current.next
    mv -T {{releases}}/current.next {{releases}}/current
    sudo systemctl restart {{service}}
    @sleep 1
    @curl -fsS http://127.0.0.1:{{port}}/healthz && echo " — up, on {{stamp}}"
    @ls -1d {{releases}}/2*/ | sort | head -n -5 | xargs -r rm -r

# Point current at the release before this one, restart, and confirm
rollback:
    #!/usr/bin/env bash
    set -euo pipefail
    cd {{releases}}
    now=$(readlink current)
    prev=$(find . -mindepth 2 -maxdepth 2 -name server -printf '%h\n' | sed 's|^\./||' | sort \
           | awk -v now="$now" '$0 == now { print p; exit } { p = $0 }')
    test -n "$prev" || { echo "nothing older than $now to go back to"; exit 1; }
    ln -sfn "$prev" current.next && mv -T current.next current
    sudo systemctl restart {{service}}
    sleep 1
    curl -fsS http://127.0.0.1:{{port}}/healthz && echo " — back on $prev ($now is still there)"

# Every release, newest first, with an arrow on the live one
releases:
    @ls -1t {{releases}} | grep -v current | sed "s|^$(readlink {{releases}}/current)$|& ← current|"
sudo systemctl daemon-reload
just deploy       # the first release; the old bin/server can go once this says up

Three details carry the weight. mv -T swaps the link in one step, so there is never a moment when current points nowhere; a plain ln -sfn onto an existing link is a delete and then a create, which would not matter for one binary behind a restart, but the same trick serves a folder of PHP that nginx is reading right now, where it does. rollback starts with #!, which makes it one script rather than a shell per line, so prev survives to the line that uses it. And it only considers folders that hold a binary, because the first version of it on this box rolled back, cheerfully, into an empty folder that a failed build had left behind, and the service went into a restart loop. deploy no longer makes the folder itself for the same reason: go build creates it, and only when the build succeeds. Five releases are kept, which is five bad decisions deep.

On this box, one deploy after another, then back:

$ just rollback
ok v7 — back on 20260916-030557 (20260916-030559 is still there)
$ just releases
20260916-030559
20260916-030557 ← current
20260916-030555
20260916-030552
20260916-030550
This is the one undo the next section permits, and it is permitted for one reason. The day you do something really stupid opens by telling you not to undo anything, and then tells the story of a firm that undid itself out of existence. A rehearsed rollback is different in every way that mattered there: it goes to a state you have run before, by a command you have run before, in about a second. What it does not do is undo data. If the deploy ran a migration, the code goes back and the schema does not, and the old build may not understand the tables it now finds. That is a restore, not a rollback, and it is the section above.

When several things break at once

One thing the ladder will not tell you, and it is worth having in your head before you need it: when a handful of things fail in the same second, they have almost certainly not failed independently. Stop working down the list and go looking for the one thing they share.

Half a dozen long-running jobs on this machine all stopped within a second of each other one morning, which is a very convincing impression of a server dying. The server was fine. Load was nothing, memory was nothing, and the logs held no errors because there were none to hold. What had happened was that my laptop's connection dropped, and every one of those jobs was a child of the same SSH session, so they went the way a branch goes when you cut through the trunk. The fault was not on the server at all, and no rung of the ladder was ever going to find it, because the ladder is bolted to the server.

The fix is one command you should learn before the wire drops rather than after. tmux keeps a session running on the server independently of the connection you started it from, so a dropped link detaches you instead of killing your work. Reconnect, reattach, and you are back in the same session with everything still running — the long build, the migration, the agent halfway through a task. screen does the same job and is often already installed. Use either. Use one of them for anything you would hate to lose.

sudo apt install -y tmux

tmux new -s work             # start a named session and work in it
# ctrl-b then d            detach and leave it running
tmux ls                      # what is still running
tmux attach -t work          # get back in, from anywhere
The thing I got wrong: I read the error instead of the ladder. A machine told me Permission denied (publickey) and I believed every word of it. I got as far as drafting instructions for getting into the hosting company's web console to add a key by hand — for a machine I was already logged into at that moment, from that terminal. The key was fine. The server was fine. ssh had looked in the five default filenames it always looks in, found nothing there, offered no key at all, and the server replied with the only sentence available to it when nobody offers it anything. One ssh -v shows you that in about a second, because it prints every key it tries before it gives up. An error names the step that quit. It does not name the step that was wrong, and those are very rarely the same step.

Finding out before anybody tells you

Everything above assumes you know it is down. The way most people actually find out is a message from somebody who noticed first, which is the most expensive monitoring system ever devised, and it only ever reports at the worst possible moment. Replace it once, for free, with a check that runs from somewhere else. It has to be somewhere else. A monitor running on the same box goes down with the box, and then reports absolutely nothing, perfectly.

There are two shapes of check, and a real project wants both.

Something that knocks: UptimeRobot
It requests a URL every few minutes and emails you when the answer stops being a 200. The free plan is 50 monitors at five-minute intervals, which it describes as for hobby and non-profit projects. Point one at your /healthz and one at your home page, and rung one of the ladder is answered before you have finished reading the message.
Something that waits to be told: Healthchecks.io
The opposite direction. Your server checks in, and it emails you when a check-in does not arrive. This is the only kind of monitor that can catch a timer that never ran, because a job that did not run produces no error for anything else to find. The free plan watches 20 jobs.

Wiring a timer to it is one line in the job's .service. systemd only runs ExecStartPost if the job itself succeeded, so a failure simply never checks in, and the silence is the alarm:

# Add under [Service] in myapp-prune.service. The URL is the one Healthchecks.io gives you.
ExecStartPost=/usr/bin/curl -fsS -m 10 --retry 3 -o /dev/null https://hc-ping.com/your-check-uuid

Then make the health endpoint worth knocking on. A handler that returns 200 OK proves the process is running and not one thing more. It will say ok, cheerfully, while the database file is unreadable and the disk is full, because nobody asked it about either. The question to answer is the one your visitors are asking — can it actually do the thing? — and for nearly every app that means does the database answer. The exception in the same breath: it should not check every outside service you depend on, or the day your payment provider has an outage your monitor wakes you up about somebody else's night.

Make GET /healthz tell the truth. Run one cheap, real query against the database with a two-second timeout. Return 200 with the version if it works, or 503 with a one-line reason if it does not. Do not call any third-party service from it, do not require a login, and put nothing secret in the response. Add a test for the 503 case.

In the interest of full disclosure, the /healthz on this very site is the lazy kind: it says ok if the process answers. It gets away with that for one reason only, which is that there is no database behind it. The day there is one, it becomes a liar.

Write the ladder down now, while nothing is broken. Put it in your project's README.md, where you will find it from your phone at some deeply unreasonable hour. You will not compose a calm diagnostic procedure while the site is down and someone is asking you when it will be back; nobody does. The version of you who is not panicking has to leave a note for the version who is. That is most of what operations actually is.

The day you do something really stupid

It is coming. Not as a possibility — as an appointment. And handled properly it is worth more to you than the six months before it, which is not a consolation prize, it is the actual mechanism by which people get good at this.

A tall tower of stacked boxes at the right leaning steeply with its top box tipping off, and the small creature well to the left pointing a small camera at it.
Photograph the wreckage before you touch it. Undoing is a second change, made by somebody upset, on a box whose state you no longer know.

The previous section is for when the box breaks. This one is for when you break it: the command that was right for the other terminal, the migration that ran against the wrong database, the cleanup script that was cleaner than intended. Everybody has one. The people who look unshakeable are not the people it has not happened to — they are the people it has already happened to, twice, and who therefore have somewhere to start.

The first ten minutes

The order matters more than the actions, and the order is not obvious, because every instinct you have in the first ten seconds is wrong. In particular the strongest one: undo it.

  1. Stop. Hands off the keyboard. If the thing is actively making it worse — a job still running, a script still looping, a service still writing — stop that, and nothing else. More damage has been done in the sixty seconds after a mistake than in the mistake.
  2. Do not un-do anything yet. Undoing is a second change, made by somebody upset, on a system whose state you no longer know. It is how a bad ten minutes becomes a bad week.
  3. Take a copy of the wreckage before you touch it. A snapshot, a tar, a dump — whatever is one command. Every recovery option below gets harder once you start overwriting things, and several become impossible.
  4. Write down what you actually did, in the order you did it, while you still remember. Ten minutes later you will be reconstructing it, and you will be reconstructing it wrong, in a direction that flatters you.
  5. Tell somebody. Now. Before you know how bad it is, before you have a plan, before the part of your brain that wants to quietly fix it first gets a vote. The next subsection is entirely about this, because it is the step everybody skips.
  6. Then work out which kind of problem you have, because there are only two and they have completely different playbooks.
The most expensive undo in history. In 2012 a trading firm deployed new code to seven of its eight servers. The eighth ran the old version, started buying and selling wildly, and the company was losing about ten million dollars a minute. In the panic, engineers rolled the new code back off the seven servers that were fine — which turned one broken server into eight. Forty-five minutes cost them more than $460 million and, within months, the company. The regulator's report reads like a tragedy in the classical sense: nobody involved was stupid, and every individual step made sense to the person taking it.

Own it immediately, upward and sideways

Here is the advice that sounds like a trap and is not: tell your colleagues and tell your boss, straight away, in plain words, before you have a fix. "I ran X against production, I think I have destroyed Y, I am looking at what we can recover and I will have an answer in twenty minutes."

It feels like handing somebody a weapon. It is the opposite, and for a reason that has nothing to do with virtue: the value of everything you own decays by the minute. Backups rotate off the end. Logs roll. Replicas resynchronise and overwrite the good copy with the bad one. Somebody else quietly builds on top of the broken data. Every minute you spend hiding it is a minute of options being deleted, and they are options you may need a second person's access to use anyway.

There is a second reason, which every experienced person knows and few say out loud: you will be forgiven for the error and you will not be forgiven for the gap. An hour between the mistake and the telling is a question you will be asked about for the rest of your time there, in a tone the mistake itself never earned. People are astonishingly relaxed about a colleague who broke something and said so instantly. They are never relaxed about finding out later.

The best-documented example in the industry is a company that did this in public. In 2017 a GitLab engineer, working late on a struggling database, ran a destructive command against the primary instead of the secondary and realised within seconds. About 300 GB went. Then they discovered that of five separate backup and replication mechanisms, not one was working — the dumps were failing silently because of a version mismatch, the snapshots were not enabled. What saved them was a six-hour-old manual snapshot somebody had taken by hand. They lost six hours of data for thousands of projects.

And then they did something remarkable: they kept a public document updated live through the recovery, streamed the whole thing on YouTube to a peak of five thousand viewers, and published a postmortem naming every failure of their own systems. Nobody remembers the engineer. Everybody who works in this field remembers that postmortem, and it did GitLab more good than a decade of marketing. The mistake was ordinary. The handling was not.

If you are on your own, the rule still applies — to yourself. There is nobody to confess to, and that is exactly why it is easy to skip straight past the honest accounting into frantic fixing. Write down what happened in the project's own notes, with the date and what you changed afterwards so it cannot happen the same way twice. Future you is a different person with your access and none of your memory, and is the colleague you are protecting.

Which of the two problems do you have?

Every disaster sorts into recoverable and not, and the single most valuable skill in the whole evening is telling them apart in the first few minutes, because they demand opposite behaviour. Recoverable means: stop, think, restore carefully, verify. Unrecoverable means: stop trying to fix it, start containing it, and start telling people — because when there is nothing to recover, everything that is left is about honesty and time.

What you didReach forCosts you
Lost code, bad rebasegit reflogMinutes
Deleted a repo or branch⁠GitHub, 90 daysMinutes
Deleted a file still open/proc/PID/fdBe quick
Wrecked the dataLast night's backupHours of writes
… and hours matterPoint-in-time recoverySet it up first
Destroyed the boxNew VPS, repo, backupUnder an hour
Published a secretRotate itAssume it is used
Leaked personal dataContain, then discloseLegal clocks start

Look at how much of that top half exists because somebody already made your mistake. The reflog is ninety days of everywhere your branches have been, and thirty more for commits nothing points at any more. GitHub's ninety-day window for a repository somebody deleted in a temper. The fact that deleting a file on Linux does not remove it while a program still holds it open — so a log, database or config you just destroyed is often still readable through the running process's own file handles. None of that is luck. Each one is a guardrail built by a person who had a very bad evening and then made sure it hurt less for the next person. That is the tradition you are joining, and it is why the first move is to stop rather than thrash: the recovery path frequently exists, and thrashing is what destroys it.

A leaked key is not a file, it is an event. Rewriting history does not help: the commit is in every clone, in forks, in caches, reachable by its hash. GitHub's own documentation says it plainly — revoke and rotate the secret first, because that is the step that actually ends the problem. And the timing is not theoretical: researchers watching public commits found new secrets a median of twenty seconds after they appeared, and a deliberately planted cloud key was being used by strangers within one minute of being pushed. If you committed a credential, it is spent. Replace it, then tidy up.
And the one that cannot be undone at all. You can restore data. You cannot un-disclose it. If personal information about actual people got out, the question stops being "can I fix this" and becomes "what am I obliged to do" — in Europe that is notifying a regulator within 72 hours, in the United States it is fifty state laws with their own definitions. There is one mercy worth knowing in advance: those rules generally treat properly encrypted data as not disclosed, which is a very concrete argument for the encrypted backups two sections ago being encrypted before they leave the box rather than after.

"Everything is gone" is a 45-minute problem

If you have followed this guide, the worst realistic case is smaller than it feels at two in the morning. Your code is on ⁠GitHub. Your data was backed up last night, off the box, encrypted. Your configuration is in the repo beside the code. So the recovery is: rent another box, run the same setup, pull the repo, restore the backup, point the domain at the new address. Most of that hour is waiting for things to download.

Two things will be outside your control, and both are worth knowing before you meet them:

The old box coming back to life
A machine you thought was dead, rebooting and resuming work — sending email, running scheduled jobs, writing to the database you just restored. Before you cut over, fence it: power it off at the provider, and revoke the keys and tokens it holds. The trading disaster above is exactly this shape — one server nobody accounted for, doing what it was last told.
DNS not moving as fast as you did
You change the address and some people still arrive at the old one, because a name has a cache lifetime and not every resolver on the internet honours it — a measured few per cent simply keep answers longer than they were told to. If you know a move is coming, lower that lifetime a day in advance. If your domain sits behind ⁠Cloudflare's proxy, this problem mostly disappears: the address the world has cached is Cloudflare's, and you are only changing where Cloudflare looks.

Afterwards: fix the tool, not the person

The instinct after a bad evening is to promise to be more careful. It does not work. I have been promising myself that for twenty years, in between truncating the wrong table, and carefulness is a resource that runs out at exactly the moment you need it — late, tired, under pressure, which is when every one of these happens.

What works is changing the shape of the thing so the mistake becomes harder to make. When Amazon took out a large part of the internet in 2017, it was because an engineer following an established procedure mistyped one input and removed more servers than intended. Their published response is a model: they did not promise better typing. They changed the tool so it removes capacity slowly, and refuses outright to take a system below its minimum. The same instinct is why the recipes in this guide's own justfile refuse to run on the main branch, why live machines get a different terminal colour, and why every destructive command in the backup section makes a copy first.

And when it is somebody else's mistake, this is the whole job. Google's reliability handbook is blunt about why blame is not merely unkind but actively expensive: where people are punished for errors, they stop reporting them, and you lose the information you needed to stop the next one. Etsy's version is sharper still — an engineer who feels safe is not only willing to be held accountable, they become enthusiastic about teaching everybody else how not to repeat it, because they are now the world expert in that particular error. If somebody breaks your thing and tells you within five minutes, the correct first sentence is "thank you for telling me straight away".

Which brings it back round to why this guide keeps insisting on a box that is not your laptop. A forty-five-dollar machine where a catastrophe costs you an afternoon is not just cheap hosting. It is the only place most people ever get to make a genuinely stupid mistake at full speed, find out what the recovery actually feels like, and discover that the floor is a lot closer than it looked. Everyone who seems calm during an outage is drawing on a memory of a night like that. Go and get yours early, on something that does not matter.

Caching, and the box that keeps up

Advanced, and optional right up until the day it is not. The one technique that lets a forty-five-dollar box do the work of a fleet — provided the rest of the project is shaped to let it.

A big cooking pot on a stool at the left, a shelf of identical sealed jars at the right, and the small creature carrying one jar between them.
Cook once, serve from the shelf. Every jar handed out is a meal the stove did not have to make.

The warning goes first, because it is the part people skip: do not build a cache on day one. A cache added before you have measured anything is a second copy of your data, with its own bugs, standing guard over a database that was never slow. Plenty of people go an entire career without ever needing to cache a query, and the scar at the end of the next section is me doing it anyway.

But for some projects the day does come, and when it comes it arrives all at once. Then caching is the difference between a box that keeps up and a box that falls over, and it costs a small fraction of what a second machine costs — which is why the checklist in More than one box puts it before adding one.

The whole idea, in one number

Most of the traffic to any site is the same handful of questions, asked over and over. The front page. The leaderboard. The total on the dashboard. If five thousand people load the same page in the same five seconds, the database does not need to answer five thousand times. It needs to answer once, and the other 4,999 get a copy.

one answer, five thousand readers The red line is the only hard part.
  1. Five thousand visitors ask the same question in the same second — the front page, the leaderboard, the report everybody opens on a Monday.
  2. The cache already has the answer, so 4,999 of them get it without the database ever hearing about it.
  3. The first request, or the first after the copy expired, is a miss, and the database does the slow work exactly once.
  4. The answer is stored with an expiry, so even a copy nobody remembers to throw away heals itself.
  5. And when somebody saves a change, the old copy is thrown away there and then, not five minutes later. That is the hard part, and very nearly the only part.
Caching is easy. Knowing when the copy has become a lie is the job.

Where to put it, cheapest first

There are five places a copy can live, and they are in this order for a reason. Each one costs more to run and more to get wrong than the one above it, so start at the top and stop at the first one that is enough.

The browser: files that never change
Give static files a name that changes when their contents do — app.3f9a1c.js, or a ?v= on the end — and tell the browser to keep them for a year. A returning visitor then downloads nothing but the page itself. It is an item on the launch checklist, and it is free.
The edge: a CDN holding whole pages
For visitors who are not logged in, an entire page can be served by ⁠Cloudflare without your server being asked at all. Two catches get everybody. Cloudflare does not cache HTML at all by default, so you have to add a cache rule that says to. And it will not cache any response that sets a cookie, so a session cookie sent with every page quietly switches the whole thing off and tells you nothing.
Precomputed: the expensive answer, built on a schedule
A report that takes twenty seconds to build takes no time at all when a timer builds it every ten minutes into a table and every visitor reads the table. This is the one that fixes the worst pages on a site, and it is the one that would have saved me the story below.
In memory: a map with an expiry, inside your app
A handful of lines in Go: a map, a lock and a timestamp. It is gone when the process restarts, which does not matter, because it was only ever a copy. This is where most of the speed on a single box actually comes from.
Shared: ⁠Redis or memcached
Only when more than one process, or more than one machine, needs to see the same copy. On one box running one Go service you do not need it — it is another service to run and watch, holding a copy of data you already have.

The first two are nothing more than a header on the response:

# Files whose name changes when their content does: keep for a year
Cache-Control: public, max-age=31536000, immutable

# A public page: browsers always check again, a CDN may hold it for five
# minutes, and may serve the old copy for a day while it fetches a new one
Cache-Control: public, max-age=0, s-maxage=300, stale-while-revalidate=86400

# Anything about one logged-in person: no shared cache, ever
Cache-Control: private, no-store

The in-memory kind is the one to hand to the agent, and the prompt is mostly about the two things that go wrong: a crowd arriving while the copy is being rebuilt, and a copy that outlives the data it was copied from.

Add an in-memory cache to GET /leaderboard. Keep the rendered result for 30 seconds. If many requests arrive while it is being rebuilt, build it once and have the rest wait for that result, rather than all of them querying the database at the same moment. Throw the cached copy away as soon as a score is saved. Never cache anything that belongs to one user. Add tests for the expiry, for the invalidation, and for twenty simultaneous requests causing exactly one query.
Some things are never served from a copy. Money. Stock. Prizes left. Seats left. Anything where "a few seconds out of date" means two people get the last one. A cached count of what remains is a promise you cannot keep, and the next part of this section is exactly that promise being broken in front of a room full of people.

The right backend for a crowd

Caching decides how much work each request costs. The other half is how many requests can be inside the building at the same time, and here the choice of language matters rather more than the people who say it does not would like. A classic ⁠PHP setup gives each request its own worker process, and when every worker is busy the next visitor queues. ⁠Node runs everything on one thread, so one slow request makes everyone wait for it. Go gives every connection its own goroutine at a few kilobytes apiece, so thousands of open connections cost memory you have, not processes you do not. That is the argument for Go in What one box holds, pointed at a crowd.

A live screen is the same idea in the other direction: one event, many watchers. The dashboard on this site works that way — one sampler, fanned out to every browser that is looking, so a hundred viewers cost barely more than one — and the rule that makes it safe is that a slow screen must never slow down the thing it is watching.

ScanQRTo.WIN is a QR-code contest system I built in Go: a code goes up on a screen, you scan it, and your phone tells you whether you won, while an admin panel shows every scan arrive live. A client paid to run it at a large conference, on a RackNerd VPS with three cores and 3.3 GB of memory, with ⁠SQLite as the database. The moment that matters here is the moment a code goes up on the big screen and a room full of phones hits the server at once, while the panel streams every one of them. At the live game on 11 October 2025 that was about three hundred people: 320 scans from 311 different addresses, 225 of them inside a single minute, nearly all over their own carriers' mobile data. It kept up, and the scan never waited on the panel.

The load tests beforehand had said it was ready, and they were right about speed. They were testing whether it was fast, not whether it was right, and a test script from one machine does none of the things a room full of people does:

  1. People share networks. Phones on the same WiFi share one IP address, so the duplicate-player check that used addresses started telling brand-new players they had already played — and showing them somebody else's win code. It was found and fixed four days before the crowd arrived. Never identify a person by their IP address.
  2. A crowd carries the same phones. 260 of those players were on an iPhone running Safari, and iPhones look so alike to a browser fingerprint that different people came out as the same device: 311 addresses, but only 301 fingerprints.
  3. A game capped at 50 winners gave out 56. Every scan asked "how many winners so far?", saw fewer than 50, and added one. When dozens of those land together, they all read the same number before any of them has written.

That last one is the most common concurrency bug there is, and the fix is a single statement: make the check and the change the same operation, so the database does both at once or neither.

-- Wrong: read the count, decide in your code, then write. Under a burst,
-- a whole crowd of requests reads "49" before any of them writes.
SELECT winner_count FROM games WHERE id = 1;

-- Right: one statement. The database checks and adds together.
UPDATE games
   SET winner_count = winner_count + 1
 WHERE id = 1 AND winner_count < max_winners;
-- One row changed: this scan won. No rows changed: the cap was reached.

Both versions were run on this box on 14 September 2026, against a cap of 50, with three hundred players arriving at once. The first gave away 79 prizes. The second gave away exactly 50. Nothing about it is faster, and nothing about it is cached — it is simply the difference between asking the database a question and asking it to keep a promise.

Eleven months later, somebody doubted the numbers, so we pushed it until it broke. On 14 September 2026, on the same three-core box with a Xeon from 2013, now sharing it with about fifteen other live services, every fake phone was released at the same instant. The October tests had only ever sent twelve requests at a time, so this one was far harsher:

  • Through the public site, 1,000 phones at once all got an answer, the median one in 11.5 seconds.
  • At 1,500, 180 of them failed, and every single failure was nginx, logging that 768 worker_connections were not enough. The app answered everything that reached it.
  • Straight to the app, 6,000 phones at once all got an answer, in 594 MB of memory. It processes sixty to eighty scans a second whatever you throw at it, so a bigger crowd becomes a longer wait, never an error. The test stopped there to protect the other services, not because anything broke.
  • The real crowd, for scale, peaked at 225 scans in a minute. That is under four a second.

So the first thing to fall over was not Go, not SQLite and not the memory. It was one number in a config file that Ubuntu ships as 768, which also sets the ceiling for every other site on the box — when it ran out, the neighbours returned errors for a couple of seconds too, which is what testing in production costs. A proxied visitor uses two connections, one from the phone and one to your app, so 768 is nearer four hundred people than eight hundred. It is worth checking before your own big day:

# Ubuntu ships 768. Every proxied visitor costs two: phone to nginx, nginx to app.
grep -n 'worker_connections\|worker_rlimit_nofile' /etc/nginx/nginx.conf
# /etc/nginx/nginx.conf. Every connection is an open file, so raise the file limit
# with it. Edit the events block that is already there; a second one will not load.
worker_rlimit_nofile 8192;

events {
    worker_connections 4096;
}
sudo nginx -t && sudo systemctl reload nginx

The report that dropped production twice

On 23 October 2007 police shut down OiNK, the invite-only music-piracy site, and several sites raced to replace it. I was the launch developer on one of them, waffles.fm, under the name Kcaj. It opened on 31 October, eight days after the raid, and made the file-sharing news the same week. I found the job in a random IRC channel while looking for work, and I did not know what the project was when I said yes. Somebody had already paid for the servers and the domain and a team was well into building it, so I went at it sideways: I took an open-source tracker-and-forum codebase, modified it for many hours without sleep, and launched that ahead of everybody.

It went from ten members to a hundred thousand in a couple of months, and it did not go smoothly. My code broke long before the hardware did — somewhere past ten thousand members, pages I had written started taking the server down, and bigger machines could not keep pace with bad queries. Getting it the rest of the way took caching, cleaned-up queries and help from people far above my pay grade.

Then I built the invite tree. Who invited whom, whom they invited, all the way down, with every branch's donations added up, so we could finally see who had been responsible for the most donations in the site's whole history. I ran it against my own account, and it dropped production. We got the site back online, I revised the query, ran it again, and dropped production again.

The thing I got wrong: I ran a job as if it were a page. That second outage was the end of my time on the project. It hardly mattered by then, because the proper caching was in, and a site whose expensive answers are built ahead of time very nearly runs itself — there was not much left to do day to day. But the lesson has been worth more to me than the project was. A report that walks a hundred thousand members and adds up every branch is not a page somebody loads. It is a job. Run it on a schedule, store the answer, and let people read that. Nobody needs to know who brought in the most donations to the second, and if one query can take the whole site down, one query will, at the least convenient moment available.

Stress testing, before the crowd does it for you

Advanced, optional, and a waste of an afternoon for most projects — right up until the day somebody links you somewhere busy. The useful part is not the number you get. It is finding out which thing breaks first, because it is never the thing you would have bet on.

Three small server boxes in a row at the left, a large round gauge on a tall post right of centre, and the small creature beside the post pulling a lever.
Push it until something gives, then find out which thing gave. It is never the one you would have bet on.

If you are not expecting thousands of people at once, you can skip this section entirely and lose nothing. Almost nobody's first project needs it. A cheap box handles more than people believe, the guide has already said so in What one box holds, and worrying about throughput before you have a single user is the same disease as buying a bigger server before you have a website on the small one.

But it costs about an hour to find out what your box actually does, and the answer is worth having before the day you need it rather than during. So here is this website, on the ⁠Go binary and the $44.88 box you have been reading about, pushed until things started falling over — and then the same box after an afternoon of fixing what the test found.

This box, pushed until it broke

The rules first, because a benchmark with no rules is marketing. A copy of this site's own binary was started as a second service with the same sandbox, memory ceiling and processor share as the live one, with a second ⁠nginx in front of it running this site's real configuration and real certificate. The live site kept serving visitors the whole time and was never the thing under test. The load generator ran on the same two cores as the server it was hitting, which means every figure below is a floor: a second machine pushing would have got more. It is all one script, just stress, and it writes what it measured into a file with a date on it.

Requests are 50 at a time, through ⁠nginx and TLS, as a browser arrives; "before" is the same test against the same code before the afternoon's fixes.

MeasuredNowBefore
The page7,109/s127/s
Slowest 1 in 10042 ms976 ms
A script or stylesheet19,208/s1,374/s
Live dashboards2,000froze past 3,000
6,000 asking at onceall 2,000 fed10% of frames
Copies of this app40, in 568 MB
1,000 connections2,411 errorsthe same

The last row is the one that has not been fixed, and it is the same wall the conference game hit in the section above: worker_connections ships at 768, a proxied visitor costs two of them, and the errors start at about four hundred people in the building at the same instant. It is one number in one file, it is in Caching, and the box that keeps up with the command to change it, and it is worth doing before your big day rather than during it.

What broke first, in the order it broke

Not one of these was the code anybody would have looked at. In order:

  1. Compression, on every single request. This page is 640 KB of HTML, and ⁠nginx was gzipping all of it, from scratch, for every visitor. One processor core, most of it, for about a hundred and twenty pages a second — while the app that generated the page sat at a tenth of its allowance, bored. The page is the same for everybody for two seconds at a time, so now it is built once, squeezed once, and handed out as-is. Same bytes, 56 times as many of them.
  2. And again, for the stylesheets and the scripts. Same mistake, different files, and this one decided how many first-time visitors a second the box could take, since a returning reader has them all in their browser already.
  3. Every open dashboard doing the same maths for itself. One reading of the machine, ten thousand watchers, ten thousand identical conversions to text, every two seconds. Now it is converted once and the same bytes go to everybody.
  4. The log eating itself. The app wrote a line per request. Under load that filled the journal's allowance — 10,000 messages every 30 seconds per service — in under half a second, and the system said so: "Suppressed 1257132 messages". Every line after that was thrown away, which would include the error you were about to go looking for. The web server's own log already had every request. The app logs failures and slow requests now, and nothing else.
  5. The memory limit that was not limiting anything. See below. This is the one worth your attention.
The thing I got wrong: my memory ceiling was a suggestion, and I had been telling you to use it. The service on this box is told MemoryMax=192M, which I believed meant "if this thing ever runs away, kill it before it takes the machine with it". So I opened twenty thousand live dashboards against a copy of it and went to look at how gracefully it died. It did not die. It sat there at 7 MB of memory and 545 MB of swap — comfortably "under" its 192 MB limit, by the only measure the limit counts — while the swap on the whole box filled up and everything else on it got slower. The limit caps memory. Swap is not memory, so swap is not capped, so a service under its ceiling can quietly page half a gigabyte onto the disk and call it obeying orders.

One line fixes it: MemorySwapMax=0. Proved rather than assumed — a little program that grabs memory as fast as it can, run under the exact unit this guide gives you: without that line it took 960 MB and finished normally, and with it the kernel stopped it dead at the limit, which is what the limit was always supposed to mean. The unit in part five has the line in it now. If you copied that block before today, add it.

With swap honestly switched off, the real shape appeared: about 40 KB of memory per live dashboard, 3,000 of them comfortable, and past that the service was throttled by the kernel into a coma — alive, answering its health check, delivering eight frames a second to six thousand people who each expected one every two seconds. Which is worse than dying, because nothing alerts on it. The fix was not a bigger box. It was to admit a limit and hold it: 2,000 live viewers, and everybody after that gets a polite refusal and keeps the numbers the page was built with. Turning somebody away in one millisecond is a feature. Serving six thousand people badly is not.

The bandwidth runs out first

Here is the part that surprised me most, and it has nothing to do with code. This box can now build and compress this page 7,109 times a second. At 207 KB each, that is about twelve gigabits a second of HTML — and the network port it would have to go out of measured 390 megabits going up and 730 coming down. The processor is roughly thirty times faster than the wire it is attached to.

So the honest limit is not "requests a second". It is bytes a month. Reading this page from top to bottom — every picture, every little video loop — costs about 4.7 MB across 71 requests, measured in a real browser; the landing screen alone is about 1.8 MB. The plan includes 6 TB a month. That is 1,276,595 full reads, or about forty thousand a day, every day, for $3.74 a month. Spread evenly it is a steady 18.5 megabits, less than a fifth of the pipe. Flat out, the port would spend the entire monthly allowance in about 34 hours.

Monthly trafficSpread evenly, that isFull reads of a page like this
1 TB3.1 Mbit/s~213,000
6 TB (this plan)18.5 Mbit/s~1,276,595
20 TB61.7 Mbit/s~4,255,000
A 1 Gbit port, flat out, all month1,000 Mbit/s324 TB, which no cheap plan includes

Two things follow from that table. The first is that media, not code, is what spends a hosting allowance: strip the pictures and the loops off this page and the same 6 TB would carry about thirty-two million reads instead of one and a quarter million. The second is the cheapest optimisation on this entire page — put the big files on somebody else's network. ⁠Cloudflare in front of the site serves the pictures from its own machines for free, and your allowance stops being the thing you think about. That is in the caching section, and it is five minutes of work.

Check how your provider counts. Some count only what leaves the box; some count both directions, which halves your real allowance if you accept uploads. ⁠Hetzner and ⁠DigitalOcean both say in their own documentation that incoming traffic is free. Cheaper hosts often count both, and say so in a sentence on a page nobody reads. Find that sentence before you build something that takes file uploads. And know what happens when you go over: some bill you, some throttle you to dial-up speed until the first of the month, which is a much worse surprise.

How many apps, and how many people

The question this guide is actually about is not "how fast is one app", it is "how much can I get away with on one box". So: 40 copies of this application were started at once, each in its own sandbox with its own memory ceiling, and then all forty were loaded simultaneously. All forty answered. They used 568 MB between them — about 14 MB each — and between them they served 17,831 requests a second.

That last number is the important one, and it is a warning rather than a boast. Forty copies did not do forty times the work. They did roughly the same total work, divided forty ways, because they share the same two cores. Memory decides how many things you can run; the processor decides how much they can all do together. With 2900 MB free on this box and a deliberate 30% left spare, the arithmetic says about 145 small Go services would fit. I would stop long before that, because the interesting number is not how many fit, it is how many are busy at the same moment — and on a box full of side projects, that is almost never more than one.

Turning that into people takes an assumption, so here is mine, stated plainly: a person reading a page makes one request and then goes quiet for a while. If an active user causes a request every ten seconds, then a hundred requests a second is a thousand people using the thing at once, and this box does eighty times that on a cached page. Real applications are far slower than a cached page — a database-heavy Ruby on Rails application, measured by VPSBenchmarks on a two-core box much like this one, topped out at 18 requests a second before latency ran away, which on the same assumption is still around a hundred and eighty people at once, and many thousands of visitors a day. The shape of the answer, for almost everybody reading this, is: more than you will get.

Other languages, roughly, and why it matters less than you think

Only Go was measured here, because only Go is what this site is written in, and a benchmark of a language I did not test is a rumour with a decimal point. What follows is other people's numbers, from the last round of the TechEmpower benchmarks — the longest-running public comparison of web frameworks, which shut down in March 2026, so this is the final scoreboard there will be. The test shown is their "Fortunes": read rows from a real database, sort them, render a template, escape the output. It is the closest thing they run to an actual web page.

Requests a second, next to Go's standard library

Each bar is that framework's Fortunes result as a multiple of Go's, on a logarithmic scale, because a straight one would make everything below Rust a smudge. Their hardware is a 56-core server, not a $44.88 box — the ratios travel, the absolute numbers do not.

Go, standard library 1.0×
Rust, axum 7.2×
C#, ASP.NET Core MVC 2.1×
Node, Fastify 1.7×
Java, Spring 1.6×
Elixir, Phoenix 1.1×
PHP, no framework 0.94×
Python, FastAPI 0.70×
Node, Express 0.50×
Ruby, Rails 0.27×
Python, Django 0.20×
PHP, Laravel 0.11×

Source: TechEmpower Framework Benchmarks, Round 23 (the final round), Fortunes test, read 16 September 2026 — 56-core Xeon Gold 6330, 64 GB, 40 Gb Ethernet. Go's standard library with a Postgres driver is the baseline at 1.0×. Two well-known entries are missing because they failed to run in that round, which is its own kind of data.

Read the bottom of that chart before the top. The gap between the fastest framework and the slowest is about seventy times — and the gap between raw ⁠PHP and PHP with a big framework on top of it is roughly nine times, in the same language. The language you pick matters less than what you pile on top of it, and both matter less than whether you do the same work twice. Fixing compression on this site, in one language, on one afternoon, was worth 56× — more than the entire distance between Go and Rust. Nobody's project was ever saved by a rewrite that could have been a cache.

Memory is the other half, and it is the half that decides how many things fit on a $44.88 box. Using the rough per-service sizes from What one box holds, and leaving 30% of this machine spare:

StackEach, roughlyThat fit here
Go, standard library14 MB measured~145
Node, Fastify130 MB~15
Java, Spring450 MB~4
Python, FastAPI105 MB~19

A dozen ⁠Python or ⁠Node services on one cheap box is perfectly normal. Four JVM applications is a decision you make on purpose. And all of them, whatever the language, are sharing two cores — so the second question is always how many are busy at the same time, which for side projects is almost always "one".

Testing yours, in about an hour

You need one tool and one rule. The tool is wrk, which is in Ubuntu's package list and does one thing well. The rule is that you never point it at the thing people are using if you can point it at a copy instead.

sudo apt install wrk

# 2 threads, 50 connections, 30 seconds, and report the slow tail honestly.
# Ask for compression, because every real browser does.
wrk -t2 -c50 -d30s --latency -H 'Accept-Encoding: gzip' https://yourdomain.com/

Read three things in what it prints, and ignore the rest. Requests/sec is how much. 99% in the latency block is how slow it got for the unluckiest one in a hundred, which is the number your visitors actually feel — an average hides exactly the problem you are looking for. And Socket errors or Non-2xx is whether it was still working at all, because a server that fails instantly looks wonderfully fast.

While it runs, watch the machine rather than the tool. Four terminals, or four commands one after another:

htop                                    # is it the processor, and which process
free -h                                 # is it memory, and is swap being used
sudo tail -f /var/log/nginx/error.log   # the web server's own complaints
journalctl -u myapp -f                  # your app's complaints

# And afterwards, the one nobody thinks to check: did the log itself
# give up and start dropping lines while you were not looking?
journalctl --since '-10 min' | grep -i suppressed
Four tests, in this order, and stop as soon as one is unhappy. First a smoke test — a handful of connections, just to prove the script and the URL are right. Then a load test at the busiest hour you actually expect. Then a stress test, at several times that, to see what the bad day looks like. Then, only if all three were boring, a breakpoint test: climb until something errors, and write down the number where it did. Those four names are borrowed from k6's guide to test types, which is the clearest short thing written on this.
Your load generator is probably lying to you about latency. wrk, hey and Apache's ab all wait for each answer before sending the next request. So when your server stalls for a second, the tool politely stops asking — and the stall never appears in the percentiles. It is called coordinated omission, and it is why a tool can report a lovely 99th percentile for a server that just froze. For "how much can it take" this does not matter and wrk is fine. For "how slow is it at exactly 500 requests a second", use a tool that keeps a fixed rate no matter how the server feels: oha with -q and --latency-correction, vegeta, or k6's constant-arrival-rate.

Where to run it from matters more than which tool you pick. From the box itself is easiest and gives you a floor, since the generator is stealing processor from the thing it is measuring. From your laptop at home you are mostly measuring your own internet connection. The honest version is a second cheap VPS, billed by the hour, in a different data centre — a dollar, deleted the same afternoon — and that is also the only setup that tells you anything about your network path.

And since this whole guide is about handing the typing to an agent, hand it this too. It is a good task for one: mechanical, repetitive, and it needs somebody patient enough to watch four terminals.

Stress-test my app on this box without touching the live one. Start a second
copy on another port as a transient systemd unit with the same properties as
its real unit file, then hit it with wrk at 10, 50, 200 and 1000 connections
for 30 seconds each. While each run goes, record processor, memory and swap
for that unit from its cgroup, and afterwards check the journal for dropped
messages and nginx's error log. Tell me the first thing that broke, what the
evidence was, and what the numbers were before and after any change you
suggest. Change nothing without asking me first.

Leave room, because the real day will not look like your test

Now the part that makes the whole exercise worth something. Whatever number you measured, do not plan to run anywhere near it. A test sends one kind of request, from one place, with a warm cache, with no logged-in users, no search queries, no reports being generated, no scrapers, and no everyone-retrying-at-once. A real peak is all of those at the same time.

The best cautionary tale in the business is Niantic's: before launching Pokémon GO they load-tested to five times their most optimistic traffic estimate, which sounds like paranoia. Real traffic arrived at roughly fifty times that estimate. Google's own reliability book, which tells that story, is also where the rule of thumb comes from: enough spare capacity to survive a planned outage and an unplanned one at the same time. Brendan Gregg puts the practical line at 70% — past that, a resource starts behaving badly rather than proportionally.

The working rule: your expected peak should be at most half of what you measured. Measured 1,000 a second? Plan for 500, and treat 700 as the moment to do something. That is not timidity, it is the gap between a test and a Tuesday — and the cheapest half of capacity you will ever buy, because you already own it.

And when it does happen to you — when the thing is genuinely too big for one box — the order is not a mystery, and it is mostly not "buy more servers":

  1. Cache the expensive thing. Nearly always the whole answer, and nearly always cheaper than anything else here. Caching, and the box that keeps up.
  2. Put the big files on somebody else's network. Pictures, video, downloads. A CDN in front of your box is free at the size you are at, and it is the difference between a bandwidth bill and no bandwidth bill.
  3. Buy a bigger box. A checkbox and a reboot, and it doubles or quadruples what you have. Unfashionable, immediate, and usually correct.
  4. Then, and only then, buy a second box and put something in front of it that shares the work. That is a different kind of project, with failover, replicas and a whole new class of way to go wrong — More than one box and the six-box setup are what that looks like, and both of them start by telling you not to.

Tuning: what moved the number, and what famously did not

Every change below was measured on its own, against the same test, one at a time. That is not diligence for its own sake: change three things and measure once, and you have learned nothing except that the three of them together did something. Here is the whole afternoon, including the parts that were a waste of time, because those are the ones nobody publishes. (The dashboard row was measured at ten thousand viewers while the service could still borrow swap; the row below it is what stopped it doing that.)

ChangeBeforeAfter
Build the page once per reading, compress it once127/s7,109/s
Compress each stylesheet and script once, not per request1,374/s19,208/s
Log failures and slow requests only35,459/s48,580/s
Convert each live reading once, not once per viewer66% of frames94% of frames
MemorySwapMax=0, and refuse viewers past a limitfroze silentlyheld, or refused
Keep-alive connections from the web server to the app7,264/s7,507/s — noise
GOMEMLIMIT, the famous Go memory knobNo effect. The pressure was socket buffers and swap, neither of which it governs.
Removing the processor cap on the service18,166/s37,898/s — and kept the cap anyway

That last row is a judgement rather than a measurement. Taking the processor limit off this service doubles what it can do, and it stays on, because this box also runs a database, a web server and somebody else's evening. A service that can take the whole machine will, on exactly the day you are not watching. The cap costs throughput nobody is asking for and buys a machine that stays usable while one thing misbehaves. That is a trade worth making on a shared box and a silly one on a dedicated one — which is the general shape of tuning advice: it depends on what else is in the building.

The two rows that did nothing are the ones I want you to take away. Upstream keep-alive is on every tuning list on the internet, and here it was worth 3%, within the noise, because the connection it saves is to 127.0.0.1. Go's memory limit is the first thing anybody reaches for, and it governed none of the memory that was actually filling up. Both are good advice in general. Neither was my problem, and I would have "fixed" my site with both of them and shipped exactly the same page at 127 requests a second. Measure first. The famous knob is famous because of somebody else's bottleneck.

Where to read properly about each of these. Go's own garbage collector guide explains GOMEMLIMIT and GOGC honestly, including leaving 5–10% of headroom under a container limit. net/http/pprof shows you where the time actually goes, and profile-guided optimisation turns that profile into a 2–14% faster binary for free. nginx's keepalive documentation is the one everybody quotes and the one that needs an upstream block to do anything at all. Microcaching is the web-server version of what fixed this page — their own test went from 5.5 to 2,185 requests a second by keeping each answer for one second. And Brendan Gregg's USE method is the one page to read before any of it: for every resource, check utilisation, saturation and errors, and change one thing at a time.

The tool that produced every number here is in this site's own repository, as one script and a paragraph explaining what it cannot tell you, because a number without its method is just a boast. Run it, keep the output, and put the date on it. Next year's you will want to know whether the box got slower or the project got heavier, and the only way to answer that is to have written it down.

When you actually need more than one box

Shown once, priced honestly — because you are almost certainly not going to need it. The section after this one is for the day you do.

Two identical server cabinets at opposite ends of the frame joined by a single straight cable.
The second box is a real answer to a real problem. It is just very rarely the first answer.

Somewhere in your reading you will be told that a single server is amateur hour and that real systems have load balancers, failover, health checks and a fleet. Most of the people saying this work somewhere with a compliance department, and all of it is true for them. It is worth understanding what they are actually buying, because it is not what most people think.

A second box buys you exactly two things: capacity beyond what one machine can do, and survival when one machine dies. That is it. Those are the only two. And the first one is much further away than anybody expects — an unmanaged VPS of the size in this guide will absorb millions of requests a day without complaining, provided you have not done anything silly, and "provided you have not done anything silly" is the part almost every scaling story is actually about.

The cheaper answer, first

Before a second machine, there is a bigger machine. Doubling the RAM on one box is a checkbox, a reboot and about twenty extra dollars a year, and it requires no change to anything you have built. There is no shame in this, whatever anyone tells you. People reach for horizontal scaling because it sounds like engineering and vertical scaling sounds like giving up, when in reality one of them is a checkbox and the other is a new distributed system with new failure modes you have never debugged.

Before you add a machine, the honest checklist:

  1. Have you measured? Not felt. Measured. Which resource is exhausted — memory, CPU, disk, or a database doing a full scan because a query has no index? Nine times in ten it is the last one, and it is one line of SQL.
  2. Have you added the index? The single most common "we need to scale" is a missing index. It is free and it takes a minute.
  3. Have you cached the expensive thing? The page that takes two seconds to build, built once a minute instead of once a request. Caching is the whole of that.
  4. Have you tried the bigger box? Twenty dollars a year, no architecture change, no new failure modes.

What it actually looks like

Here it is, once, so you have seen it. ⁠nginx will happily balance across several machines — this is the whole configuration, and it is genuinely this small:

# Two application servers behind one public name. nginx sends each
# request to whichever is healthy; "max_fails" and "fail_timeout"
# are what make it stop sending traffic to a box that has died.
upstream app {
    server 10.0.0.11:8090 max_fails=2 fail_timeout=10s;
    server 10.0.0.12:8090 max_fails=2 fail_timeout=10s;
}

server {
    listen 443 ssl;
    server_name mysite.com;

    location / {
        proxy_pass http://app;
        proxy_next_upstream error timeout http_502 http_503;
    }
}

Twenty lines. Looks easy. Now here is the bill, and the bill is not the money.

What you addCostsAnd then
A second app server+$44.88/yrEvery deploy now has to reach both, and they must never be running different versions for long.
Somewhere shared for the data+$44.88/yrTwo app servers cannot share a ⁠SQLite file. You now need a database on its own machine — which is itself a single point of failure, so you have not actually removed one, you have moved it.
Shared sessions and uploadsTimeAnything held in one server's memory or on one server's disk breaks the moment a user's second request lands on the other machine.
The load balancer itselfA third box, or a managed oneIf it is one box, it is a single point of failure. If it is two, you are now doing DNS failover or a floating IP, and you have built a small network.
Real failoverOngoing attentionHealth checks that are honest, a database that can promote a replica, and a rehearsal — because failover you have never tested is the same as no failover, only more expensive.

So the true price of "highly available" is not $90 a year. It is three or four machines, a distributed data story, a deployment process that updates several places atomically, and a class of bug where everything is fine except on one server, sometimes. That is a real and reasonable thing to build when the requirement is real. It is a miserable thing to build because a blog post implied you should.

Ask what an hour of downtime actually costs you. This is the whole question and it takes ten seconds to answer. For a personal project, a portfolio, a tool your friends use, a shop making a handful of orders a day: an hour offline costs you an apology. For that, one box with tested backups is not a compromise — it is the correct engineering answer, and adding machines would make your uptime worse, because complexity you do not understand fails more often than simplicity you do. When an hour of downtime costs real money, you will know, and you will have the money to do this properly.

The second box worth having

There is one that pays for itself immediately, and it is not a load balancer. Rent a second cheap VPS at a different company, in a different country, and make it the target for your backups and nothing else. No traffic, no load balancing, no shared state. It costs the same twenty-odd dollars, and it means the failure that actually ends projects — the provider losing a disk, or suspending an account over a billing mix-up on a card that expired — is an inconvenience instead of an ending.

That is redundancy. Two copies of the thing that cannot be recreated, in two places that cannot fail together. Load balancing is a performance decision and 99.9% of people reading this will never need one. Backups are not advanced and everybody needs them. Almost every tutorial on the internet has these two exactly the wrong way round.

The thing I got wrong: I built for scale I did not have. Cache layers for traffic that never arrived, queues for jobs that could have been a function call, a second server standing there costing money and adding a version of "it works on one of them" to my life. All of it defensible in a design document and none of it load-bearing. The tell is when you cannot point at the measurement that made you do it. Optimising for theoretical milliseconds in projects that do not exist yet is a hobby, and it is a hobby that looks exactly like work, which is what makes it dangerous.

The six-box setup, for when it is real

Load balancers, a floating address, a primary database and its replica, and the automation that fails over and then comes back. For mission-critical projects only. Read it so you know the shape; build none of it yet.

Three pairs of identical server boxes spaced across the frame, each pair joined by a short cable, and the small creature pulling a lever beside the middle pair.
Everything in pairs, and one lever. The pairs are the easy part. The lever is the whole job.
This is not for you yet, and it may never be. If you are new to running a server, skip this section and come back when a project of yours has an hour of downtime with a price on it. Six boxes means six machines to patch, a replication setup to watch, two kinds of failover to rehearse, and a whole class of bug that only happens on one of them, sometimes. For the great majority of projects it is simply not worth it — the section above has the bill — and built before it is needed, it makes a site less reliable, not more.

So why is it here at all? Because pieces of it are exactly what you reach for on the day a project suddenly takes off, and that is not a day you want to spend learning the shape of the thing. About one project in a hundred is built so that it can scale, and about one in a hundred ever needs to. They are very rarely the same project.

The shape

six boxes, every one of them replaceable The green line is failover. The red cross is where recovery starts.
  1. Visitors look up one name and get one address. They never learn there are six machines behind it, and they should not have to.
  2. That address is a floating IP: one your provider lets you move from one machine to another in seconds, usually with a single API call. Right now it belongs to balancer A.
  3. Balancer B is identical to A and does nothing except watch it. The heartbeat between them is B's entire job.
  4. Either balancer can send any request to either app server. They run the same code at the same version and keep nothing on their own disks, so it genuinely does not matter which one answers.
  5. Every write goes to the primary database. The replica copies it continuously, takes the heavy reads, and is the spare.
  6. When A goes quiet, B claims the floating IP. That is failover, and it is the easy half.
  7. When the primary dies, the replica is promoted — and then something has to rebuild the dead machine as the new replica and let it catch up. That is recovery. It is the half nobody rehearses, and until it is finished you are running on one database again with every dashboard green.
Nothing in it is special. Everything in it has a twin, and every twin has been tested.

What each pair is for

The floating address
Not every provider sells one, and that alone can decide where your balancers live. Without one, the fallback is DNS: a short TTL set days before you need it, and a record you change by hand or by script. It is slower, because some resolvers hold on to an old answer longer than they were told to, but it is how most small setups actually fail over, and it is a lot better than nothing.
Two balancers
⁠nginx with an upstream block, exactly the twenty lines in the section above, or HAProxy. The second one sits there doing nothing, which is not waste. It is the product.
Two app servers
The moment there are two, anything kept on one app server's disk or in its memory is a bug waiting for a user's second request to land on the other one. Sessions go in the database or ⁠Redis, uploads go to object storage, and every deploy reaches both before either serves the new version for long.
A primary and a replica
⁠Postgres ships replication built in. Writes go to the primary, the replica follows it, and reports and dashboards can read from the replica so the heavy readers stop fighting the writers — which is exactly where the invite tree in Caching should have run.

The health check has to ask the right question, and "is it up?" is not it. I learned this one the slow way, with a script that sent reads back to the primary whenever the replica was down. A replica can be completely up — answering every query, instantly — and still be serving old data, because replication quietly stopped. Down is not the only way to be wrong. Ask how far behind it is:

# On the primary: is every replica connected, and how far behind is each one?
sudo -u postgres psql -c "SELECT client_addr, state, replay_lag FROM pg_stat_replication;"

# On a replica: am I a replica, and how old is the newest change I have applied?
sudo -u postgres psql -c "SELECT pg_is_in_recovery(), now() - pg_last_xact_replay_timestamp();"

The second number grows on a quiet database even when everything is fine, because nothing new has been written for it to apply. Read it together with the first, never on its own.

Failover is half. Recovery is the other half.

Failover gets all the attention, because it is the dramatic part and it can be made automatic and fast. Recovery is what comes after, and it is where the real damage is done. The old primary is now out of date, and it must never come back as a primary on its own: two machines that both believe they are the primary will both accept writes, and the data splits in two in a way that no tool will ever merge back together. It has to be rebuilt as a replica of the new primary, allowed to catch up, checked, and only then are you back to two.

Automate both halves or neither. A tool such as Patroni handles promotion and rejoining for Postgres, and it is a serious piece of software that deserves a serious week of reading before it goes anywhere near your data. Then rehearse the whole cycle on a schedule, because failover you have never tested is not failover. It is a hope with a monthly bill.

The pieces you actually reach for in an emergency

The day a project takes off is not the day to build six boxes. It is the day to do the next cheapest thing, in this order, and to stop at the first one that is enough:

  1. Cache the expensive thing. Caching. Hours of work, and very often all you needed.
  2. Buy the bigger box. A checkbox and a reboot.
  3. Move the database to a machine of its own. The first real split, and the one that buys the most.
  4. Add a read replica. Point the reports and dashboards at it, so the heavy readers stop fighting the writers.
  5. Put the logged-out pages behind a CDN. Most visitors then never reach your server at all.
  6. Keep a warm standby at a different company. Set the DNS TTL short now and write the steps down now, so losing an entire provider means changing one A record.
  7. Only then, the whole six boxes. If an hour offline still costs more than everything above: the balancers, the floating address and the automation.
The thing I got wrong: my list of what lived where. In September 2026 a provider holding a stack of my projects went dark for about a day, with no explanation. The recovery went the way this guide says it should: a new box, the code from git, the data from off-site backups, the DNS repointed. And three domains stayed down after everything else was back, because they were not on my list of what ran on that machine — one of them sitting behind a ⁠Cloudflare error because its record still pointed at the dead box. None of the six-box machinery above would have helped with that. A list would have. What runs where is part of your recovery plan, and it is the part nobody writes down until the day they need it.

Glossary

Each of these is explained where it first turns up. That is no help at all on the Tuesday three weeks later when you meet it again in an error message, so here they all are in one place.

None of these words are difficult. Nearly every one of them is named after the thing it does, and stops being intimidating the moment somebody tells you what that is — which is roughly one sentence of work, and the reason a beginner is made to feel stupid for a fortnight is that nobody could be bothered to do it. So: one sentence each. If you know a term already, skip it; nothing here is a test.

The machine

VPS
Virtual private server. A computer you rent by the month in somebody else's building. Nobody else is using it — it is simply not in your house. Buying the server.
root
The account that is allowed to do anything, including all the things you did not mean. You log in as an ordinary user and borrow its powers deliberately.
sudo
"Run this one command as root." The password it asks for is yours, not root's, and typing it is the moment to read the command once more.
apt
How software gets onto an Ubuntu machine: an archive Ubuntu builds and signs, and one command that installs from it. update refreshes the list of what exists; upgrade is the one that installs. apt.
PPA
Somebody else's archive, added to yours. Their builds then run as root on your machine every time they update, which is the deal you already have with Ubuntu, minus Ubuntu. apt.
permissions
Who may read, write or run a file: its owner, its group, and everyone else, each answered separately. The usual reason something works for you and fails as a service. Permissions.
daemon
A program that runs quietly in the background from the moment the machine boots, with nobody watching it. ⁠nginx is one. Your app becomes one.
CVE
A published security hole, with an identifier so everybody is talking about the same one. Around 250 new ones a day, nearly all of them about software you do not run. Staying up to date.
end of life
The date after which nobody fixes a piece of software any more. It keeps working perfectly, which is the trap. Staying up to date.
p99
The slowest one request in a hundred. The number your visitors actually feel, and the one an average is best at hiding. Stress testing.
systemd
The thing that starts daemons, restarts them when they die, and writes down whatever they said on the way. If your app is running at all, systemd is why.
unit
One file telling systemd how to run one thing: which command, as which user, and what to do when it stops. Yours will live in /etc/systemd/system/.
timer
A unit that starts another unit on a schedule — systemd's replacement for cron, with logs. Timers.
port
A numbered door on the machine. 443 belongs to the web, 22 to SSH, and the 8090s are yours to hand out as you please.
loopback
127.0.0.1 — the address meaning "this machine and nothing else". A program listening there cannot be reached from the internet at all, which is exactly where your app should sit.
firewall
The list of which ports the outside world may knock on. ufw is the friendly face of it: shut everything, then open three. Lock it down.

Getting in

SSH
Encrypted remote control of the machine's command line. Ancient, hammered on by everybody for thirty years, and the only door you will use. Getting in.
key pair
Two matching files. The public half sits on the server and is not a secret; the private half never leaves your laptop, and is never emailed, pasted or committed.
passphrase
The password on the private key itself, so that a stolen laptop is not the same event as a stolen server.

Names, and the padlock

DNS
The internet's phone book: it turns a name people can remember into an address machines can use. Pointing a domain.
A record
One entry in that book: a name pointing at an IPv4 address. AAAA is the same thing for IPv6.
CNAME
A name pointing at another name rather than an address — www.mysite.com at mysite.com — so you only ever have one address to keep up to date.
TLS
The encryption behind the padlock. HTTPS is ordinary HTTP with TLS underneath it. You will also see it called SSL, which is the old name for the old version and has stuck around the way old names do.
certificate
A file proving the name is really yours, signed by somebody every browser already trusts. It is free, it expires, and renewing it is somebody else's cron job once you have set it up. nginx, systemd, TLS.

The web server

nginx
The program holding port 443. It reads which domain the request asked for and hands it to whichever of your apps owns that name, which is the whole trick behind one box serving unlimited sites.
load balancer
A front door that shares requests out across several copies of your app, and stops sending them to one that has died. Rarely needed on one box. The six-box setup.
reverse proxy
That job, given its proper name. A normal proxy fetches the internet on your behalf; a reverse proxy receives the internet on your app's behalf, so your app never has to face it directly.
rate limit
A ceiling on how often one address may ask, enforced by nginx before your app hears about it. Not for attackers, who work from a list you are not on; for the one script with a bug in its loop. Not getting flooded.
access log
The file nginx appends one line to for every request: who asked, for what, and what they got. Already your analytics; awk is the dashboard. Who came.

Your code, and its history

repository
A folder that git is watching, plus every version of it that has ever existed. Everyone says "repo".
commit
One saved point in that history, with a message saying why. Also the verb for making one. Cheap, so make them constantly.
branch
A line of history running alongside the main one, so that unfinished work does not have to be finished before anything else can start.
CI
Continuous integration: a robot on somebody else's computer that runs your tests every time you push, so "it worked on my machine" stops being a defence. ⁠GitHub Actions.
environment variable
A value handed to a program by whatever started it, rather than written inside it. This is where a secret goes: a file only the service can read, named from its systemd unit, and never in git. An environment variable, worked through once.
release
One build, in its own dated folder, beside the builds before it. The service runs whichever one a link called current points at.
rollback
Pointing current at the release before this one and restarting. One command, about a second, if the deploy was built to allow it. The deploy that broke it.

Data

CRUD
Create, Read, Update, Delete — the four things every application has ever done to anything. Everything is CRUD.
WAL
Write-ahead log. ⁠SQLite writes each change to a side file before folding it in, so readers are never blocked by a writer and a power cut leaves you something repairable rather than something shredded. Turn it on once and forget it. Databases.
cache
A copy of an answer kept so it does not have to be worked out again. Fast, and wrong the moment the original changes, which is why every copy needs an expiry. Caching.
rollup
A timer that turns a table of raw events into a table of hourly totals, so the pages that need the number read a sum that already exists. A job, not a page. When your app wants its own numbers.
replica
A second database that continuously copies a primary one. It can take the heavy reads and stand in if the primary dies, and it can be up and quietly out of date. The six-box setup.
restore drill
A timer that restores last night's backup into a scratch copy, counts what came back, and prints one line. The backup had a timer; this gives the restore one. Then make the machine prove it.
Jargon is mostly a fence with no field behind it. Read that list again and notice how much of it is one-sentence stuff. Every one of these words was invented by somebody who needed a short way to say a long thing, and not one of them was invented to keep you out — that part gets added later, by people who enjoy knowing something you do not. When you hit a word this page never taught you, and you will, the useful move is to ask what it is for rather than what it means. Your agent will tell you in a sentence and will not be smug about it, which is more than can be said for the internet.

Cheat sheet

The ones you will actually use. Bookmark this bit; nobody memorises these and nobody should.

Getting around

ssh sam@203.0.113.40      # connect
exit                      # leave
cd /srv/myproject         # go somewhere
ls -la                    # what is here, including hidden files
pwd                       # where am I
nano file.txt             # edit  (ctrl-O saves, ctrl-X quits)
cat file.txt              # print a file
tail -f file.log          # watch a file grow
grep -rn "thing" .        # find text in every file below here
du -sh *                  # what is taking up space

Packages

sudo apt update                   # refresh the catalogue; installs nothing
sudo apt upgrade                  # install what is newer
apt search --names-only thing     # is there a package for it
apt show thing                    # what is it, how big
sudo apt install -y thing         # get it
sudo apt remove thing             # drop it (purge: and its config too)
apt list --upgradable             # what is waiting

Services

systemctl status myapp            # is it running?
sudo systemctl restart myapp      # turn it off and on again
sudo systemctl enable myapp       # start it at every boot
journalctl -u myapp -f            # follow its log
journalctl -u myapp -n 200        # the last 200 lines
systemctl --failed                # what is broken

Web

sudo nginx -t                     # is the config valid?
sudo systemctl reload nginx       # apply it, safely
sudo certbot certificates         # what certs exist, expiring when
sudo certbot renew --dry-run      # prove renewal works
curl -I https://mysite.com        # see the response headers
dig +short mysite.com             # where does the name point
sudo tail -f /var/log/nginx/access.log  # who is here, live

Git

git status                        # what has changed
git add -A && git commit -m "..."  # save a checkpoint
git push                          # send it to GitHub
git log --oneline -10             # recent history
git diff                          # what have I changed
git restore <file>                # undo my changes to one file
git revert <commit>               # undo a commit, safely

The agent

cd / && claude                    # start it where it can see everything
/model                            # switch model: plan vs build
/clear                            # fresh context, same folder
# a fact worth keeping             # writes it into CLAUDE.md
Esc                               # stop it, without ending the session
Shift+Tab                         # cycle permission modes

Health

htop                              # live CPU and memory
free -h                           # memory left
df -h                             # disk left
sudo ss -tlnp                     # what is listening, and on which address
sudo ufw status verbose           # what the firewall allows
sudo fail2ban-client status sshd  # who has been banned
uptime                            # how long since the last reboot

Go and build something

That is the whole method. A server for $44.88 a year, a domain for the price of a coffee, and an agent with a terminal and permission to use it. Everything else on this page is detail hung off those three things.

The reason to write it down is that the hard part was never the code. The hard part was knowing that a server costs less than a streaming subscription, that a certificate is free and takes fifteen seconds, that subdomains are unlimited, and that the right move is to climb out of the graphical editor and let something capable work in a place where mistakes cost nothing. None of that is secret. It is just very rarely said plainly, in one place, by somebody with nothing to sell you.

I had an art teacher in middle school who was better at teaching than almost anyone I have met since. He asked us to draw a perfect circle. Everyone produced something wobbly and apologetic. Then he drew six fast, sloppy circles on top of each other and rubbed out everything that was not part of the very obvious, very perfect circle now sitting there in the middle of them. That is what building software is. Not one correct line after another — six rough passes and then removing what was not it. Your first version will be wobbly. It is supposed to be. It is one of the six.

Take all of this, change it, disagree loudly with the opinionated parts. The stacks here are one set of preferences that works well on a small box with an agent doing the typing, and they are not the only set. What compounds is not my conventions — it is having conventions, writing them down, and reusing them until your fifth project starts on the afternoon you think of it rather than the following spring.

And do not wait to be told you are allowed. There is no book, no course, no certificate and no job that hands out permission; the people waiting for one are missing the rest of the horse, not you. Go and build the smallest ugly version of the thing, tonight, on a machine that does not matter. The world is your delicious oyster. Good luck and Godspeed.

See it running

Real-time telemetry from the box serving this page. Same hardware, same price.

Built this way

Games, websites, marketing tools — all the same skeleton, all on boxes like this one.

Start now

Five questions and a prompt to paste. It takes a minute.

The server prices in the survey were read off each provider's own page on 14 September 2026 — 3 days ago — and every row links to the page it came from. Domain and signing prices were checked in September 2026. Being prices, all of them are subject to change, so verify before you buy. Nothing here is sponsored and no link on this site is an affiliate link.

Every chart on this page carries its source and the date it was read, underneath the chart, along with whatever is wrong with the data. There is no chart here whose numbers you cannot go and check yourself in about a minute, which is the only kind worth drawing.

The logos are from Simple Icons, drawn beside the name of whatever the sentence is already talking about. Each one is its owner's trademark, and none of them is here because anybody asked.

The six little creatures at the top of each part, and the wide drawing at the head of most sections, were generated by an image model on this box for about ten pence in total. They are nobody's mascot but ours — not the penguin, not the daemon, not the gopher — which is deliberate, because the real ones carry four different licences between them and one of them is a trademark you need written permission to use. No drawing here contains a word, a logo or a signature, and that is enforced in the prompt rather than hoped for. Everything about how they were made, including the ninety-odd drafts that did not survive and what each one cost, is in the repository.