docs: update README and RUNBOOK for multi-deployment config

- README: remove hardcoded srv paths; add configs/ to components table;
  document BACKUP_CONF usage for multi-server setups; note DB_CONTAINER=
  makes docker optional
- RUNBOOK: generalize §1 setup to use config variables (source config,
  use $REPO / $BORG_PASSPHRASE_FILE / etc. instead of hardcoded paths);
  §4 day-2 ops now shows source-config pattern before direct borg commands;
  §5 recovery notes BACKUP_CONF for restore.sh; §6 drill uses config vars

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A3rQSEidP6Y61kaCtxjVV1
This commit is contained in:
2026-08-30 23:50:38 +01:00
co-authored by Claude Sonnet 4.6
parent 74390b78af
commit f3cf24a41c
2 changed files with 205 additions and 138 deletions
+28 -13
View File
@@ -1,23 +1,38 @@
# backup-agent
Backup and disaster-recovery tooling for a Linux server: [Borg](https://borgbackup.readthedocs.io/)
backing up `/home/srv/files/content` — including a MariaDB database
running in Docker — with an offsite mirror on Scaleway S3 via `rclone`.
Backup and disaster-recovery tooling for Linux servers: [Borg](https://borgbackup.readthedocs.io/)
backing up a target directory tree — including a MariaDB database running in
Docker — with an offsite mirror on Scaleway S3 via `rclone`.
MariaDB is never stopped during backup: `dump_db.sh` takes a
transactionally-consistent logical dump (`mysqldump --single-transaction`)
while the container keeps running, and the container's raw data directory
is excluded from the archive entirely (a `.nobackup` marker file), so only
the logical dump ever gets backed up. Zero DB downtime.
while the container keeps running, and the container's raw data directory is
excluded from the archive entirely (a `.nobackup` marker file), so only the
logical dump ever gets backed up. Zero DB downtime.
## Components
| File | Purpose |
|---|---|
| `configs/` | Per-deployment config files (one per server). Each sets `TARGET`, `REPO`, `DB_CONTAINER`, etc. |
| `borg-backup.sh` | Daily backup orchestrator (run from cron): dump → archive → prune → compact → integrity check → offsite sync. |
| `dump_db.sh` | Per-database `mysqldump`/`mariadb-dump`, atomic staging/swap. Invoked by `borg-backup.sh`. |
| `restore.sh` | Recovery CLI: `full` (disaster recovery), `db <name>` (single database), `file <path>` (single file/dir), `--list-archives`. Every mode supports `--dry-run`. |
## Multi-server use
Both scripts are deployment-agnostic. Point them at a config with `BACKUP_CONF`:
```bash
# backup
BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf /opt/backup-agent/borg-backup.sh
# restore
BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf /opt/backup-agent/restore.sh --list-archives
```
See `configs/` for existing deployments and `RUNBOOK.md §2` for deployment steps.
## Quickstart
```bash
@@ -37,16 +52,16 @@ Day to day:
## Encryption
The Borg repo at `/home/srv/files/backups/borg-2025` is **unencrypted** by
deliberate choice on this deployment`borg-backup.sh` will keep printing
a warning about it on every run, which is expected. See `RUNBOOK.md` if
you want to switch to an encrypted repo.
The Borg repos on current deployments are **unencrypted** by deliberate operator
choice`borg-backup.sh` will keep printing a warning about it on every run,
which is expected. See `RUNBOOK.md` if you want to switch to an encrypted repo.
## Requirements
`borg`, `docker`, `rclone`, `flock`, a `mysql`/`mariadb` client — see
`REQUIRED_CMDS` in `borg-backup.sh`. Targets Linux; `flock(1)` doesn't
exist on macOS, so these scripts won't run as-is on a Mac.
`borg`, `docker`, `rclone`, `flock`, a `mysql`/`mariadb` client. When
`DB_CONTAINER` is empty in the config, `docker` is not required. Targets
Linux; `flock(1)` doesn't exist on macOS, so these scripts won't run as-is
on a Mac.
## License
+177 -125
View File
@@ -3,72 +3,111 @@
Covers `borg-backup.sh` (daily backup), `dump_db.sh` (MariaDB logical
dumps, invoked by the backup script), and `restore.sh` (recovery).
## 1. One-Time Setup
All deployment-specific values (`TARGET`, `REPO`, `DB_CONTAINER`, etc.) live in
a config file under `configs/`. Both scripts source it automatically — see
**§2** for how to point them at the right one.
Run once, by hand, on the server:
---
1. Initialize the encrypted Borg repo:
```bash
mkdir -p /home/srv/files/backups
borg init --encryption=repokey-blake2 /home/srv/files/backups/borg-2025
```
2. Create the passphrase file used by both backup and restore:
```bash
echo 'your-strong-passphrase' > /root/.borg-passphrase
chmod 600 /root/.borg-passphrase
```
3. Create the MariaDB `backup` user used by `dump_db.sh` (read-only, no
stop/lock of the server required thanks to `--single-transaction`):
```sql
CREATE USER 'backup'@'%' IDENTIFIED BY 'a-strong-password';
GRANT SELECT, LOCK TABLES, SHOW VIEW, TRIGGER, PROCESS, RELOAD ON *.* TO 'backup'@'%';
```
```bash
echo 'a-strong-password' > /root/.mariadb-backup.pw
chmod 600 /root/.mariadb-backup.pw
```
4. Create the root password file used only by `restore.sh` (restore needs
CREATE/DROP privileges the `backup` user does not have):
```bash
echo 'the-mariadb-root-password' > /root/.mariadb-root.pw
chmod 600 /root/.mariadb-root.pw
```
5. Exclude MariaDB's raw data directory from the archive. Find the
directory bind-mounted into the container as its datadir and drop a
marker file in it:
```bash
touch /home/srv/files/content/mariadb/data/.nobackup
```
This is what allows backups to run with the container up: only the
logical dump under `mariadb/dump/` is ever archived or restored from.
6. Configure the `scaleway` rclone remote:
```bash
rclone config
# create a remote named "scaleway", type S3, matching your Scaleway
# Object Storage credentials and region
```
## 1. One-Time Setup (per server)
Run once, by hand, on the server. Substitute values from your config file
(`configs/<deployment>.conf`). The examples below use shell variables sourced
from it:
```bash
source /opt/backup-agent/configs/nexusvoice.conf # adapt path
```
### 1.1 Initialize the Borg repo
```bash
mkdir -p "$(dirname "$REPO")"
borg init --encryption=repokey-blake2 "$REPO"
# If running unencrypted by deliberate choice:
# borg init --encryption=none "$REPO"
```
### 1.2 Passphrase file (backup and restore both read it)
```bash
echo 'your-strong-passphrase' > "$BORG_PASSPHRASE_FILE"
chmod 600 "$BORG_PASSPHRASE_FILE"
```
Skip this step only if running without encryption and without a passphrase.
### 1.3 MariaDB `backup` user (used by `dump_db.sh`)
Read-only; `--single-transaction` makes the dump consistent without stopping
the container.
```sql
CREATE USER 'backup'@'%' IDENTIFIED BY 'a-strong-password';
GRANT SELECT, LOCK TABLES, SHOW VIEW, TRIGGER, PROCESS, RELOAD ON *.* TO 'backup'@'%';
```
```bash
echo 'a-strong-password' > /root/.mariadb-backup.pw
chmod 600 /root/.mariadb-backup.pw
```
### 1.4 MariaDB root password file (used only by `restore.sh`)
Restore needs CREATE/DROP privileges the `backup` user does not have.
```bash
echo 'the-mariadb-root-password' > "$ROOT_PASSWORD_FILE"
chmod 600 "$ROOT_PASSWORD_FILE"
```
### 1.5 Exclude the MariaDB raw data directory
Find the host directory bind-mounted into the container as its datadir and
drop a marker file in it. Borg skips any directory that contains `.nobackup`
(via `--exclude-if-present`), so live InnoDB files are never read:
```bash
touch "${TARGET}/mariadb/data/.nobackup"
```
This is what allows backups to run with the container up: only the logical
dump under `${DUMP_SUBDIR}/` is ever archived or restored.
### 1.6 Configure the Scaleway rclone remote
```bash
rclone config
# Create a remote named "scaleway", type S3, with your Scaleway
# Object Storage credentials and region.
```
---
## 2. Deploying the Scripts
### Single deployment (default)
### Single deployment
Copy `borg-backup.sh`, `dump_db.sh`, `restore.sh`, and your chosen config
from `configs/` to `/opt/backup-agent/`. Symlink the config as `backup.conf`
next to the scripts, or set `BACKUP_CONF` in the cron entry:
Copy `borg-backup.sh`, `dump_db.sh`, `restore.sh`, and the `configs/` directory
to `/opt/backup-agent/`. Symlink the relevant config as `backup.conf` beside
the scripts, or pass `BACKUP_CONF` explicitly in the cron entry.
```bash
chmod +x /opt/backup-agent/borg-backup.sh /opt/backup-agent/restore.sh
chmod +x /opt/backup-agent/dump_db.sh
chmod +x /opt/backup-agent/borg-backup.sh \
/opt/backup-agent/restore.sh \
/opt/backup-agent/dump_db.sh
# Option A — symlink (scripts auto-discover backup.conf beside them):
ln -s /opt/backup-agent/configs/srv.conf /opt/backup-agent/backup.conf
# Option B — explicit env var in the cron entry (see §3).
# Option B — explicit BACKUP_CONF in the cron entry (see §3).
```
### Multiple deployments on different servers
Each server gets its own config file. Deploy the same three scripts to
`/opt/backup-agent/` on each server. Point each server's cron entry at its
config via `BACKUP_CONF`:
Each server gets its own config. Deploy the same three scripts to
`/opt/backup-agent/` on each server; point each cron entry at its config via
`BACKUP_CONF`:
```
# /etc/cron.d/borg-backup (nexusvoice server)
@@ -76,10 +115,12 @@ BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf
30 2 * * * root /opt/backup-agent/borg-backup.sh >> /var/log/borg/cron.log 2>&1
```
Before running, verify/update `configs/nexusvoice.conf`:
- `TARGET` — confirm `/home/acid/nexusvoice` (or `/opt/nexusvoice` after the move)
Before the first run, verify `configs/nexusvoice.conf`:
- `TARGET` — current path; update if files move (e.g. `/home/acid/nexusvoice``/opt/nexusvoice`)
- `DB_CONTAINER` — confirm the MariaDB container name on that server
- `REPO` — must not overlap with the srv repo; uses separate Scaleway path
- `REPO` — must not overlap with any other deployment's repo
---
## 3. Scheduling
@@ -97,128 +138,139 @@ ls -lt /var/log/borg/backup-*.log | head -1 # latest log file
tail -50 /var/log/borg/backup-*.log # inspect it
```
---
## 4. Day-2 Operations
List archives:
Source your config first to get the right `BORG_REPO`, `BORG_PASSCOMMAND`, etc.:
```bash
source /opt/backup-agent/configs/nexusvoice.conf # adapt
export BORG_REPO="$REPO"
export BORG_PASSCOMMAND="cat $BORG_PASSPHRASE_FILE"
```
Then standard borg commands work without extra flags:
```bash
# List archives
./restore.sh --list-archives
# or directly:
borg list
# Repo size and health
borg info
# Rotate the passphrase
# Note: only re-encrypts the key, not the archive data; old passphrase
# still needed for archives created before the change until fully migrated.
borg key change-passphrase
# Break a stale lockfile (backup or restore aborted mid-run)
borg break-lock
```
Check repo size and health:
```bash
BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg info /home/srv/files/backups/borg-2025
```
Rotate the passphrase (creates a new key, re-encrypts nothing — old
archives still need the old passphrase to read, so keep both until fully
migrated):
```bash
BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg key change-passphrase /home/srv/files/backups/borg-2025
```
Stale lockfile (backup or restore aborted mid-run and left the repo
locked):
```bash
BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg break-lock /home/srv/files/backups/borg-2025
```
---
## 5. Recovery Procedures
All `restore.sh` commands accept `--dry-run` to preview exactly what would
happen without touching anything, and `--archive NAME` to target a
specific archive instead of the latest (see archive names via
`--list-archives`). Root DB credentials come from `MYSQL_ROOT_PASSWORD` in
the environment if set, otherwise from `/root/.mariadb-root.pw` — set
whichever is more convenient for how you're invoking it. Restore logs go
to `/var/log/borg/restore-*.log` (the `RESTORE_LOGDIR` environment
variable overrides the directory, mainly useful for testing). `full` and
`db` share `borg-backup.sh`'s lockfile, so a restore refuses to start
while the nightly backup is mid-run (and vice versa) rather than racing
it.
All `restore.sh` commands accept:
- `--dry-run` — preview exactly what would happen without touching anything
- `--archive NAME` — target a specific archive instead of the latest (get names via `--list-archives`)
Root DB credentials come from `MYSQL_ROOT_PASSWORD` in the environment if set,
otherwise from the `ROOT_PASSWORD_FILE` in your config (default
`/root/.mariadb-root.pw`). Restore logs go to `/var/log/borg/restore-*.log`;
`RESTORE_LOGDIR` overrides the directory (useful for testing).
`restore.sh` shares `borg-backup.sh`'s lockfile (`LOCKFILE` in the config),
so a restore refuses to start while the nightly backup is mid-run and vice
versa.
For all `restore.sh` commands, select the deployment with `BACKUP_CONF`:
```bash
export BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf
```
Or run without it if `backup.conf` is already symlinked beside the script.
### 5.1 Full disaster recovery (new or wiped server)
Use when the whole server/container is gone and you're rebuilding from
scratch.
Use when the whole server/container is gone and you're rebuilding from scratch.
```bash
# 1. Reinstall borg, docker, and the mariadb container image/compose file
# (not covered by restore.sh - this is infra provisioning).
# (not covered by restore.sh this is infra provisioning).
# 2. Restore the passphrase file (from your password manager / secondary
# backup - it is NOT stored in the repo it protects) to
# /root/.borg-passphrase, and the root DB password to
# /root/.mariadb-root.pw.
# backup it is NOT stored in the repo it protects) to the path set in
# BORG_PASSPHRASE_FILE, and the root DB password to ROOT_PASSWORD_FILE.
# 3. Preview:
./restore.sh full --dry-run
# 4. Run for real. --force is only required if /home/srv/files/content
# already has data in it (e.g. a stale mount); omit it on a genuinely
# empty/fresh server:
# 4. Run for real. --force is only required if $TARGET already has data in it
# (e.g. a stale mount); omit it on a genuinely empty/fresh server:
./restore.sh full --force
```
This extracts the full content tree from the archive, starts the
`mariadb` container and waits for it to report healthy, then restores
every database dump (users/grants first) — the container must be running
before any of the dump restores, which is why it starts first.
This extracts the full content tree, starts the DB container and waits for it
to report healthy, then restores every database dump (users/grants first).
**Verify afterward:**
- `docker ps` shows `mariadb` running and healthy.
- `docker ps` shows the DB container running and healthy.
- The application responds normally.
- Spot-check row counts on a couple of tables against what you'd expect.
### 5.2 Single database restore
Use when one database got corrupted or someone ran a bad migration/query
against it — this **drops and recreates** that database.
Use when one database got corrupted or someone ran a bad migration/query
this **drops and recreates** that database.
```bash
./restore.sh db shopdb --dry-run # preview
./restore.sh db shopdb # prompts: type "shopdb" to confirm
./restore.sh db shopdb # prompts: type "shopdb" to confirm
```
Non-interactive (e.g. scripted from a monitoring alert): add `--yes` to
skip the typed confirmation.
**Verify afterward:** connect to the database and check the tables/row
counts you expect.
**Verify afterward:** connect to the database and check the tables/row counts
you expect.
### 5.3 Single file/directory restore
Use for accidental deletion of a file, or to inspect an old version — this
never touches the running database or container.
Use for accidental deletion or to inspect an old version — never touches the
running database or container.
```bash
./restore.sh file path/relative/to/content/some-file.txt --dest /tmp/recovered
./restore.sh file path/relative/to/target/some-file.txt --dest /tmp/recovered
```
The final location of the recovered item is printed at the end (it lands
under `/tmp/recovered/home/srv/files/content/...` — Borg preserves the
absolute path it was archived with).
The final location of the recovered item is printed at the end. Borg
preserves the absolute path it was archived with, so the file lands under
`/tmp/recovered/<TARGET>/...`.
---
## 6. Restore Drill Cadence
Quarterly, run a real `full` restore into a scratch directory (not
`/home/srv/files/content`) to confirm backups are actually usable:
Quarterly, run a real `full` extract into a scratch directory to confirm
backups are actually usable.
`borg extract` always extracts into the current directory (there's no
`--destination` flag — this is why `restore.sh` itself `cd`s into the
destination before extracting), and the archive name must come from
`borg list --short` (plain `borg list` prints a formatted line, not a bare
name), so:
`borg extract` always extracts into the current directory; the archive name
must come from `borg list --short` (plain `borg list` prints a formatted
line, not a bare name). Source your config to get the right values:
```bash
source /opt/backup-agent/configs/nexusvoice.conf # adapt
export BORG_REPO="$REPO"
export BORG_PASSCOMMAND="cat $BORG_PASSPHRASE_FILE"
mkdir -p /tmp/restore-drill && cd /tmp/restore-drill
export BORG_REPO=/home/srv/files/backups/borg-2025
export BORG_PASSCOMMAND="cat /root/.borg-passphrase"
LATEST=$(borg list --short | tail -1)
borg extract --lock-wait 600 "::$LATEST"
```
Confirm the dump files under `mariadb/dump/` are present, non-empty, and
importable (`mysql -u root -p < mariadb/dump/somedb.sql` against a
throwaway MariaDB container). Log the drill date and outcome somewhere
durable (e.g. a team wiki page).
Confirm the dump files under `${DUMP_SUBDIR}/` are present, non-empty, and
importable (`mysql -u root -p < mariadb/dump/somedb.sql` against a throwaway
MariaDB container). Log the drill date and outcome somewhere durable (e.g. a
team wiki page).