docs: update README and RUNBOOK for multi-deployment config

- README: remove hardcoded srv paths; add configs/ to components table;
  document BACKUP_CONF usage for multi-server setups; note DB_CONTAINER=
  makes docker optional
- RUNBOOK: generalize §1 setup to use config variables (source config,
  use $REPO / $BORG_PASSPHRASE_FILE / etc. instead of hardcoded paths);
  §4 day-2 ops now shows source-config pattern before direct borg commands;
  §5 recovery notes BACKUP_CONF for restore.sh; §6 drill uses config vars

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A3rQSEidP6Y61kaCtxjVV1
This commit is contained in:
2026-08-30 23:50:38 +01:00
co-authored by Claude Sonnet 4.6
parent 74390b78af
commit f3cf24a41c
2 changed files with 205 additions and 138 deletions
+28 -13
View File
@@ -1,23 +1,38 @@
# backup-agent # backup-agent
Backup and disaster-recovery tooling for a Linux server: [Borg](https://borgbackup.readthedocs.io/) Backup and disaster-recovery tooling for Linux servers: [Borg](https://borgbackup.readthedocs.io/)
backing up `/home/srv/files/content` — including a MariaDB database backing up a target directory tree — including a MariaDB database running in
running in Docker — with an offsite mirror on Scaleway S3 via `rclone`. Docker — with an offsite mirror on Scaleway S3 via `rclone`.
MariaDB is never stopped during backup: `dump_db.sh` takes a MariaDB is never stopped during backup: `dump_db.sh` takes a
transactionally-consistent logical dump (`mysqldump --single-transaction`) transactionally-consistent logical dump (`mysqldump --single-transaction`)
while the container keeps running, and the container's raw data directory while the container keeps running, and the container's raw data directory is
is excluded from the archive entirely (a `.nobackup` marker file), so only excluded from the archive entirely (a `.nobackup` marker file), so only the
the logical dump ever gets backed up. Zero DB downtime. logical dump ever gets backed up. Zero DB downtime.
## Components ## Components
| File | Purpose | | File | Purpose |
|---|---| |---|---|
| `configs/` | Per-deployment config files (one per server). Each sets `TARGET`, `REPO`, `DB_CONTAINER`, etc. |
| `borg-backup.sh` | Daily backup orchestrator (run from cron): dump → archive → prune → compact → integrity check → offsite sync. | | `borg-backup.sh` | Daily backup orchestrator (run from cron): dump → archive → prune → compact → integrity check → offsite sync. |
| `dump_db.sh` | Per-database `mysqldump`/`mariadb-dump`, atomic staging/swap. Invoked by `borg-backup.sh`. | | `dump_db.sh` | Per-database `mysqldump`/`mariadb-dump`, atomic staging/swap. Invoked by `borg-backup.sh`. |
| `restore.sh` | Recovery CLI: `full` (disaster recovery), `db <name>` (single database), `file <path>` (single file/dir), `--list-archives`. Every mode supports `--dry-run`. | | `restore.sh` | Recovery CLI: `full` (disaster recovery), `db <name>` (single database), `file <path>` (single file/dir), `--list-archives`. Every mode supports `--dry-run`. |
## Multi-server use
Both scripts are deployment-agnostic. Point them at a config with `BACKUP_CONF`:
```bash
# backup
BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf /opt/backup-agent/borg-backup.sh
# restore
BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf /opt/backup-agent/restore.sh --list-archives
```
See `configs/` for existing deployments and `RUNBOOK.md §2` for deployment steps.
## Quickstart ## Quickstart
```bash ```bash
@@ -37,16 +52,16 @@ Day to day:
## Encryption ## Encryption
The Borg repo at `/home/srv/files/backups/borg-2025` is **unencrypted** by The Borg repos on current deployments are **unencrypted** by deliberate operator
deliberate choice on this deployment`borg-backup.sh` will keep printing choice`borg-backup.sh` will keep printing a warning about it on every run,
a warning about it on every run, which is expected. See `RUNBOOK.md` if which is expected. See `RUNBOOK.md` if you want to switch to an encrypted repo.
you want to switch to an encrypted repo.
## Requirements ## Requirements
`borg`, `docker`, `rclone`, `flock`, a `mysql`/`mariadb` client — see `borg`, `docker`, `rclone`, `flock`, a `mysql`/`mariadb` client. When
`REQUIRED_CMDS` in `borg-backup.sh`. Targets Linux; `flock(1)` doesn't `DB_CONTAINER` is empty in the config, `docker` is not required. Targets
exist on macOS, so these scripts won't run as-is on a Mac. Linux; `flock(1)` doesn't exist on macOS, so these scripts won't run as-is
on a Mac.
## License ## License
+177 -125
View File
@@ -3,72 +3,111 @@
Covers `borg-backup.sh` (daily backup), `dump_db.sh` (MariaDB logical Covers `borg-backup.sh` (daily backup), `dump_db.sh` (MariaDB logical
dumps, invoked by the backup script), and `restore.sh` (recovery). dumps, invoked by the backup script), and `restore.sh` (recovery).
## 1. One-Time Setup All deployment-specific values (`TARGET`, `REPO`, `DB_CONTAINER`, etc.) live in
a config file under `configs/`. Both scripts source it automatically — see
**§2** for how to point them at the right one.
Run once, by hand, on the server: ---
1. Initialize the encrypted Borg repo: ## 1. One-Time Setup (per server)
```bash
mkdir -p /home/srv/files/backups Run once, by hand, on the server. Substitute values from your config file
borg init --encryption=repokey-blake2 /home/srv/files/backups/borg-2025 (`configs/<deployment>.conf`). The examples below use shell variables sourced
``` from it:
2. Create the passphrase file used by both backup and restore:
```bash ```bash
echo 'your-strong-passphrase' > /root/.borg-passphrase source /opt/backup-agent/configs/nexusvoice.conf # adapt path
chmod 600 /root/.borg-passphrase ```
```
3. Create the MariaDB `backup` user used by `dump_db.sh` (read-only, no ### 1.1 Initialize the Borg repo
stop/lock of the server required thanks to `--single-transaction`):
```sql ```bash
CREATE USER 'backup'@'%' IDENTIFIED BY 'a-strong-password'; mkdir -p "$(dirname "$REPO")"
GRANT SELECT, LOCK TABLES, SHOW VIEW, TRIGGER, PROCESS, RELOAD ON *.* TO 'backup'@'%'; borg init --encryption=repokey-blake2 "$REPO"
``` # If running unencrypted by deliberate choice:
```bash # borg init --encryption=none "$REPO"
echo 'a-strong-password' > /root/.mariadb-backup.pw ```
chmod 600 /root/.mariadb-backup.pw
``` ### 1.2 Passphrase file (backup and restore both read it)
4. Create the root password file used only by `restore.sh` (restore needs
CREATE/DROP privileges the `backup` user does not have): ```bash
```bash echo 'your-strong-passphrase' > "$BORG_PASSPHRASE_FILE"
echo 'the-mariadb-root-password' > /root/.mariadb-root.pw chmod 600 "$BORG_PASSPHRASE_FILE"
chmod 600 /root/.mariadb-root.pw ```
```
5. Exclude MariaDB's raw data directory from the archive. Find the Skip this step only if running without encryption and without a passphrase.
directory bind-mounted into the container as its datadir and drop a
marker file in it: ### 1.3 MariaDB `backup` user (used by `dump_db.sh`)
```bash
touch /home/srv/files/content/mariadb/data/.nobackup Read-only; `--single-transaction` makes the dump consistent without stopping
``` the container.
This is what allows backups to run with the container up: only the
logical dump under `mariadb/dump/` is ever archived or restored from. ```sql
6. Configure the `scaleway` rclone remote: CREATE USER 'backup'@'%' IDENTIFIED BY 'a-strong-password';
```bash GRANT SELECT, LOCK TABLES, SHOW VIEW, TRIGGER, PROCESS, RELOAD ON *.* TO 'backup'@'%';
rclone config ```
# create a remote named "scaleway", type S3, matching your Scaleway
# Object Storage credentials and region ```bash
``` echo 'a-strong-password' > /root/.mariadb-backup.pw
chmod 600 /root/.mariadb-backup.pw
```
### 1.4 MariaDB root password file (used only by `restore.sh`)
Restore needs CREATE/DROP privileges the `backup` user does not have.
```bash
echo 'the-mariadb-root-password' > "$ROOT_PASSWORD_FILE"
chmod 600 "$ROOT_PASSWORD_FILE"
```
### 1.5 Exclude the MariaDB raw data directory
Find the host directory bind-mounted into the container as its datadir and
drop a marker file in it. Borg skips any directory that contains `.nobackup`
(via `--exclude-if-present`), so live InnoDB files are never read:
```bash
touch "${TARGET}/mariadb/data/.nobackup"
```
This is what allows backups to run with the container up: only the logical
dump under `${DUMP_SUBDIR}/` is ever archived or restored.
### 1.6 Configure the Scaleway rclone remote
```bash
rclone config
# Create a remote named "scaleway", type S3, with your Scaleway
# Object Storage credentials and region.
```
---
## 2. Deploying the Scripts ## 2. Deploying the Scripts
### Single deployment (default) ### Single deployment
Copy `borg-backup.sh`, `dump_db.sh`, `restore.sh`, and your chosen config Copy `borg-backup.sh`, `dump_db.sh`, `restore.sh`, and the `configs/` directory
from `configs/` to `/opt/backup-agent/`. Symlink the config as `backup.conf` to `/opt/backup-agent/`. Symlink the relevant config as `backup.conf` beside
next to the scripts, or set `BACKUP_CONF` in the cron entry: the scripts, or pass `BACKUP_CONF` explicitly in the cron entry.
```bash ```bash
chmod +x /opt/backup-agent/borg-backup.sh /opt/backup-agent/restore.sh chmod +x /opt/backup-agent/borg-backup.sh \
chmod +x /opt/backup-agent/dump_db.sh /opt/backup-agent/restore.sh \
/opt/backup-agent/dump_db.sh
# Option A — symlink (scripts auto-discover backup.conf beside them): # Option A — symlink (scripts auto-discover backup.conf beside them):
ln -s /opt/backup-agent/configs/srv.conf /opt/backup-agent/backup.conf ln -s /opt/backup-agent/configs/srv.conf /opt/backup-agent/backup.conf
# Option B — explicit env var in the cron entry (see §3).
# Option B — explicit BACKUP_CONF in the cron entry (see §3).
``` ```
### Multiple deployments on different servers ### Multiple deployments on different servers
Each server gets its own config file. Deploy the same three scripts to Each server gets its own config. Deploy the same three scripts to
`/opt/backup-agent/` on each server. Point each server's cron entry at its `/opt/backup-agent/` on each server; point each cron entry at its config via
config via `BACKUP_CONF`: `BACKUP_CONF`:
``` ```
# /etc/cron.d/borg-backup (nexusvoice server) # /etc/cron.d/borg-backup (nexusvoice server)
@@ -76,10 +115,12 @@ BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf
30 2 * * * root /opt/backup-agent/borg-backup.sh >> /var/log/borg/cron.log 2>&1 30 2 * * * root /opt/backup-agent/borg-backup.sh >> /var/log/borg/cron.log 2>&1
``` ```
Before running, verify/update `configs/nexusvoice.conf`: Before the first run, verify `configs/nexusvoice.conf`:
- `TARGET` — confirm `/home/acid/nexusvoice` (or `/opt/nexusvoice` after the move) - `TARGET` — current path; update if files move (e.g. `/home/acid/nexusvoice``/opt/nexusvoice`)
- `DB_CONTAINER` — confirm the MariaDB container name on that server - `DB_CONTAINER` — confirm the MariaDB container name on that server
- `REPO` — must not overlap with the srv repo; uses separate Scaleway path - `REPO` — must not overlap with any other deployment's repo
---
## 3. Scheduling ## 3. Scheduling
@@ -97,128 +138,139 @@ ls -lt /var/log/borg/backup-*.log | head -1 # latest log file
tail -50 /var/log/borg/backup-*.log # inspect it tail -50 /var/log/borg/backup-*.log # inspect it
``` ```
---
## 4. Day-2 Operations ## 4. Day-2 Operations
List archives: Source your config first to get the right `BORG_REPO`, `BORG_PASSCOMMAND`, etc.:
```bash ```bash
source /opt/backup-agent/configs/nexusvoice.conf # adapt
export BORG_REPO="$REPO"
export BORG_PASSCOMMAND="cat $BORG_PASSPHRASE_FILE"
```
Then standard borg commands work without extra flags:
```bash
# List archives
./restore.sh --list-archives ./restore.sh --list-archives
# or directly:
borg list
# Repo size and health
borg info
# Rotate the passphrase
# Note: only re-encrypts the key, not the archive data; old passphrase
# still needed for archives created before the change until fully migrated.
borg key change-passphrase
# Break a stale lockfile (backup or restore aborted mid-run)
borg break-lock
``` ```
Check repo size and health: ---
```bash
BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg info /home/srv/files/backups/borg-2025
```
Rotate the passphrase (creates a new key, re-encrypts nothing — old
archives still need the old passphrase to read, so keep both until fully
migrated):
```bash
BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg key change-passphrase /home/srv/files/backups/borg-2025
```
Stale lockfile (backup or restore aborted mid-run and left the repo
locked):
```bash
BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg break-lock /home/srv/files/backups/borg-2025
```
## 5. Recovery Procedures ## 5. Recovery Procedures
All `restore.sh` commands accept `--dry-run` to preview exactly what would All `restore.sh` commands accept:
happen without touching anything, and `--archive NAME` to target a - `--dry-run` — preview exactly what would happen without touching anything
specific archive instead of the latest (see archive names via - `--archive NAME` — target a specific archive instead of the latest (get names via `--list-archives`)
`--list-archives`). Root DB credentials come from `MYSQL_ROOT_PASSWORD` in
the environment if set, otherwise from `/root/.mariadb-root.pw` — set Root DB credentials come from `MYSQL_ROOT_PASSWORD` in the environment if set,
whichever is more convenient for how you're invoking it. Restore logs go otherwise from the `ROOT_PASSWORD_FILE` in your config (default
to `/var/log/borg/restore-*.log` (the `RESTORE_LOGDIR` environment `/root/.mariadb-root.pw`). Restore logs go to `/var/log/borg/restore-*.log`;
variable overrides the directory, mainly useful for testing). `full` and `RESTORE_LOGDIR` overrides the directory (useful for testing).
`db` share `borg-backup.sh`'s lockfile, so a restore refuses to start
while the nightly backup is mid-run (and vice versa) rather than racing `restore.sh` shares `borg-backup.sh`'s lockfile (`LOCKFILE` in the config),
it. so a restore refuses to start while the nightly backup is mid-run and vice
versa.
For all `restore.sh` commands, select the deployment with `BACKUP_CONF`:
```bash
export BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf
```
Or run without it if `backup.conf` is already symlinked beside the script.
### 5.1 Full disaster recovery (new or wiped server) ### 5.1 Full disaster recovery (new or wiped server)
Use when the whole server/container is gone and you're rebuilding from Use when the whole server/container is gone and you're rebuilding from scratch.
scratch.
```bash ```bash
# 1. Reinstall borg, docker, and the mariadb container image/compose file # 1. Reinstall borg, docker, and the mariadb container image/compose file
# (not covered by restore.sh - this is infra provisioning). # (not covered by restore.sh this is infra provisioning).
# 2. Restore the passphrase file (from your password manager / secondary # 2. Restore the passphrase file (from your password manager / secondary
# backup - it is NOT stored in the repo it protects) to # backup it is NOT stored in the repo it protects) to the path set in
# /root/.borg-passphrase, and the root DB password to # BORG_PASSPHRASE_FILE, and the root DB password to ROOT_PASSWORD_FILE.
# /root/.mariadb-root.pw.
# 3. Preview: # 3. Preview:
./restore.sh full --dry-run ./restore.sh full --dry-run
# 4. Run for real. --force is only required if /home/srv/files/content # 4. Run for real. --force is only required if $TARGET already has data in it
# already has data in it (e.g. a stale mount); omit it on a genuinely # (e.g. a stale mount); omit it on a genuinely empty/fresh server:
# empty/fresh server:
./restore.sh full --force ./restore.sh full --force
``` ```
This extracts the full content tree from the archive, starts the This extracts the full content tree, starts the DB container and waits for it
`mariadb` container and waits for it to report healthy, then restores to report healthy, then restores every database dump (users/grants first).
every database dump (users/grants first) — the container must be running
before any of the dump restores, which is why it starts first.
**Verify afterward:** **Verify afterward:**
- `docker ps` shows `mariadb` running and healthy. - `docker ps` shows the DB container running and healthy.
- The application responds normally. - The application responds normally.
- Spot-check row counts on a couple of tables against what you'd expect. - Spot-check row counts on a couple of tables against what you'd expect.
### 5.2 Single database restore ### 5.2 Single database restore
Use when one database got corrupted or someone ran a bad migration/query Use when one database got corrupted or someone ran a bad migration/query
against it — this **drops and recreates** that database. this **drops and recreates** that database.
```bash ```bash
./restore.sh db shopdb --dry-run # preview ./restore.sh db shopdb --dry-run # preview
./restore.sh db shopdb # prompts: type "shopdb" to confirm ./restore.sh db shopdb # prompts: type "shopdb" to confirm
``` ```
Non-interactive (e.g. scripted from a monitoring alert): add `--yes` to Non-interactive (e.g. scripted from a monitoring alert): add `--yes` to
skip the typed confirmation. skip the typed confirmation.
**Verify afterward:** connect to the database and check the tables/row **Verify afterward:** connect to the database and check the tables/row counts
counts you expect. you expect.
### 5.3 Single file/directory restore ### 5.3 Single file/directory restore
Use for accidental deletion of a file, or to inspect an old version — this Use for accidental deletion or to inspect an old version — never touches the
never touches the running database or container. running database or container.
```bash ```bash
./restore.sh file path/relative/to/content/some-file.txt --dest /tmp/recovered ./restore.sh file path/relative/to/target/some-file.txt --dest /tmp/recovered
``` ```
The final location of the recovered item is printed at the end (it lands The final location of the recovered item is printed at the end. Borg
under `/tmp/recovered/home/srv/files/content/...` — Borg preserves the preserves the absolute path it was archived with, so the file lands under
absolute path it was archived with). `/tmp/recovered/<TARGET>/...`.
---
## 6. Restore Drill Cadence ## 6. Restore Drill Cadence
Quarterly, run a real `full` restore into a scratch directory (not Quarterly, run a real `full` extract into a scratch directory to confirm
`/home/srv/files/content`) to confirm backups are actually usable: backups are actually usable.
`borg extract` always extracts into the current directory (there's no `borg extract` always extracts into the current directory; the archive name
`--destination` flag — this is why `restore.sh` itself `cd`s into the must come from `borg list --short` (plain `borg list` prints a formatted
destination before extracting), and the archive name must come from line, not a bare name). Source your config to get the right values:
`borg list --short` (plain `borg list` prints a formatted line, not a bare
name), so:
```bash ```bash
source /opt/backup-agent/configs/nexusvoice.conf # adapt
export BORG_REPO="$REPO"
export BORG_PASSCOMMAND="cat $BORG_PASSPHRASE_FILE"
mkdir -p /tmp/restore-drill && cd /tmp/restore-drill mkdir -p /tmp/restore-drill && cd /tmp/restore-drill
export BORG_REPO=/home/srv/files/backups/borg-2025
export BORG_PASSCOMMAND="cat /root/.borg-passphrase"
LATEST=$(borg list --short | tail -1) LATEST=$(borg list --short | tail -1)
borg extract --lock-wait 600 "::$LATEST" borg extract --lock-wait 600 "::$LATEST"
``` ```
Confirm the dump files under `mariadb/dump/` are present, non-empty, and Confirm the dump files under `${DUMP_SUBDIR}/` are present, non-empty, and
importable (`mysql -u root -p < mariadb/dump/somedb.sql` against a importable (`mysql -u root -p < mariadb/dump/somedb.sql` against a throwaway
throwaway MariaDB container). Log the drill date and outcome somewhere MariaDB container). Log the drill date and outcome somewhere durable (e.g. a
durable (e.g. a team wiki page). team wiki page).