cmd / r2

r2 is a backup tool and a Cloudflare R2 client. It snapshots SQLite databases and bundles git repositories, and it puts, gets, and lists objects. I wrote it in Go with the standard library only.

Why I wrote it

My VM holds the primary copy of my applications, and the provider sells no disk snapshots. A disk failure would take the databases and the repositories with it, so the copies have to live off the box.

The AWS SDK covers that in four API calls and a large dependency. I wrote the four calls instead.

Commands

r2 backup sqlite db.db  # gzipped snapshot, uploaded as db-<UTCts>.db.gz
r2 backup git repo/     # bundle of the branch, as repo-<UTCts>.bundle
r2 put file [key]       # upload a file
r2 get key [dest]       # download an object
r2 ls [prefix]          # list keys with sizes and times

The SQLite backup shells out to sqlite3 .backup, so the snapshot stays consistent while the application writes. I describe the database options in go / sqlite. The git backup shells out to git bundle, and git clone restores the bundle.

Signing requests from the spec

R2 speaks the S3 API, and an S3 request includes a Signature Version 4 signature. The SDK hides that work. r2 implements it from the AWS documentation in 123 lines.

The signer builds a canonical request from the method, path, query, and headers, then hashes it into a string to sign. It derives the signing key through four HMAC rounds: date, region, service, and aws4_request. It signs the string and sets the Authorization header. R2 names its region auto.

AWS publishes test vectors for SigV4, and the repository keeps them under testdata/. Each vector fixes the request and the clock, and the test compares the signature byte for byte. A signer that passes the vectors reaches R2 on the first request.

Design

One PUT per object, and no multipart upload. A single PUT uploads a file up to 5 GB, and my artifacts are megabytes.

No pruning logic. Seven lifecycle rules on the bucket expire objects by prefix, so retention lives next to the data. Hourly keys expire after 7 days, daily keys after 90, and weekly keys after 400. Four legacy prefixes drain after 1 day.

One bucket per VM. Keys carry a tier prefix and an application prefix: hourly/sports/db-<ts>.db.gz, daily/sports/repo-<ts>.bundle, and so on. The API token accepts requests from the VM addresses only.

Four environment variables configure the tool: R2_ENDPOINT, R2_BUCKET, R2_ACCESS_KEY_ID, and R2_SECRET_ACCESS_KEY. An optional R2_PREFIX prepends the application prefix to a generated key. The tool exits when one of the four is missing.

How sports uses it

The scoreboard backs up through a shell script that calls r2 put. A git post-commit hook pushes a bundle after each commit, so a commit reaches the bucket in seconds. The application snapshots the database on each clean ingest or saved summary, and those snapshots stay on the VM. An hourly cron job pushes the changed snapshots off the box. Each push writes the hourly prefix. The first push of the day also writes the daily prefix, and the first of the week also writes the weekly prefix.

A restore reads the most recent bundle and snapshot:

r2 ls hourly/sports/
r2 get hourly/sports/repo-<ts>.bundle
git clone repo-<ts>.bundle sports

The rehearsal

A backup is a claim until a restore proves it. A rehearsal on the VM runs each Wednesday: it downloads the newest snapshot and bundle, restores the database, and records the key and the byte count. The latest run restored an hourly snapshot of about 50 MB.

The rehearsal measures the recovery point I tolerate.

What it is not

r2 is not a sync tool and not a log streamer. Litestream ships each write-ahead log frame and restores to any second. r2 ships point-in-time artifacts and restores to the last trigger. I tolerate a recovery point of minutes, so I run the simpler tool.

← All articles