diff --git a/README.md b/README.md index 1ecd442..5e5cb5b 100644 --- a/README.md +++ b/README.md @@ -1,23 +1,38 @@ # backup-agent -Backup and disaster-recovery tooling for a Linux server: [Borg](https://borgbackup.readthedocs.io/) -backing up `/home/srv/files/content` — including a MariaDB database -running in Docker — with an offsite mirror on Scaleway S3 via `rclone`. +Backup and disaster-recovery tooling for Linux servers: [Borg](https://borgbackup.readthedocs.io/) +backing up a target directory tree — including a MariaDB database running in +Docker — with an offsite mirror on Scaleway S3 via `rclone`. MariaDB is never stopped during backup: `dump_db.sh` takes a transactionally-consistent logical dump (`mysqldump --single-transaction`) -while the container keeps running, and the container's raw data directory -is excluded from the archive entirely (a `.nobackup` marker file), so only -the logical dump ever gets backed up. Zero DB downtime. +while the container keeps running, and the container's raw data directory is +excluded from the archive entirely (a `.nobackup` marker file), so only the +logical dump ever gets backed up. Zero DB downtime. ## Components | File | Purpose | |---|---| +| `configs/` | Per-deployment config files (one per server). Each sets `TARGET`, `REPO`, `DB_CONTAINER`, etc. | | `borg-backup.sh` | Daily backup orchestrator (run from cron): dump → archive → prune → compact → integrity check → offsite sync. | | `dump_db.sh` | Per-database `mysqldump`/`mariadb-dump`, atomic staging/swap. Invoked by `borg-backup.sh`. | | `restore.sh` | Recovery CLI: `full` (disaster recovery), `db ` (single database), `file ` (single file/dir), `--list-archives`. Every mode supports `--dry-run`. | +## Multi-server use + +Both scripts are deployment-agnostic. Point them at a config with `BACKUP_CONF`: + +```bash +# backup +BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf /opt/backup-agent/borg-backup.sh + +# restore +BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf /opt/backup-agent/restore.sh --list-archives +``` + +See `configs/` for existing deployments and `RUNBOOK.md §2` for deployment steps. + ## Quickstart ```bash @@ -37,16 +52,16 @@ Day to day: ## Encryption -The Borg repo at `/home/srv/files/backups/borg-2025` is **unencrypted** by -deliberate choice on this deployment — `borg-backup.sh` will keep printing -a warning about it on every run, which is expected. See `RUNBOOK.md` if -you want to switch to an encrypted repo. +The Borg repos on current deployments are **unencrypted** by deliberate operator +choice — `borg-backup.sh` will keep printing a warning about it on every run, +which is expected. See `RUNBOOK.md` if you want to switch to an encrypted repo. ## Requirements -`borg`, `docker`, `rclone`, `flock`, a `mysql`/`mariadb` client — see -`REQUIRED_CMDS` in `borg-backup.sh`. Targets Linux; `flock(1)` doesn't -exist on macOS, so these scripts won't run as-is on a Mac. +`borg`, `docker`, `rclone`, `flock`, a `mysql`/`mariadb` client. When +`DB_CONTAINER` is empty in the config, `docker` is not required. Targets +Linux; `flock(1)` doesn't exist on macOS, so these scripts won't run as-is +on a Mac. ## License diff --git a/RUNBOOK.md b/RUNBOOK.md index bccc92a..8787111 100644 --- a/RUNBOOK.md +++ b/RUNBOOK.md @@ -3,72 +3,111 @@ Covers `borg-backup.sh` (daily backup), `dump_db.sh` (MariaDB logical dumps, invoked by the backup script), and `restore.sh` (recovery). -## 1. One-Time Setup +All deployment-specific values (`TARGET`, `REPO`, `DB_CONTAINER`, etc.) live in +a config file under `configs/`. Both scripts source it automatically — see +**§2** for how to point them at the right one. -Run once, by hand, on the server: +--- -1. Initialize the encrypted Borg repo: - ```bash - mkdir -p /home/srv/files/backups - borg init --encryption=repokey-blake2 /home/srv/files/backups/borg-2025 - ``` -2. Create the passphrase file used by both backup and restore: - ```bash - echo 'your-strong-passphrase' > /root/.borg-passphrase - chmod 600 /root/.borg-passphrase - ``` -3. Create the MariaDB `backup` user used by `dump_db.sh` (read-only, no - stop/lock of the server required thanks to `--single-transaction`): - ```sql - CREATE USER 'backup'@'%' IDENTIFIED BY 'a-strong-password'; - GRANT SELECT, LOCK TABLES, SHOW VIEW, TRIGGER, PROCESS, RELOAD ON *.* TO 'backup'@'%'; - ``` - ```bash - echo 'a-strong-password' > /root/.mariadb-backup.pw - chmod 600 /root/.mariadb-backup.pw - ``` -4. Create the root password file used only by `restore.sh` (restore needs - CREATE/DROP privileges the `backup` user does not have): - ```bash - echo 'the-mariadb-root-password' > /root/.mariadb-root.pw - chmod 600 /root/.mariadb-root.pw - ``` -5. Exclude MariaDB's raw data directory from the archive. Find the - directory bind-mounted into the container as its datadir and drop a - marker file in it: - ```bash - touch /home/srv/files/content/mariadb/data/.nobackup - ``` - This is what allows backups to run with the container up: only the - logical dump under `mariadb/dump/` is ever archived or restored from. -6. Configure the `scaleway` rclone remote: - ```bash - rclone config - # create a remote named "scaleway", type S3, matching your Scaleway - # Object Storage credentials and region - ``` +## 1. One-Time Setup (per server) + +Run once, by hand, on the server. Substitute values from your config file +(`configs/.conf`). The examples below use shell variables sourced +from it: + +```bash +source /opt/backup-agent/configs/nexusvoice.conf # adapt path +``` + +### 1.1 Initialize the Borg repo + +```bash +mkdir -p "$(dirname "$REPO")" +borg init --encryption=repokey-blake2 "$REPO" +# If running unencrypted by deliberate choice: +# borg init --encryption=none "$REPO" +``` + +### 1.2 Passphrase file (backup and restore both read it) + +```bash +echo 'your-strong-passphrase' > "$BORG_PASSPHRASE_FILE" +chmod 600 "$BORG_PASSPHRASE_FILE" +``` + +Skip this step only if running without encryption and without a passphrase. + +### 1.3 MariaDB `backup` user (used by `dump_db.sh`) + +Read-only; `--single-transaction` makes the dump consistent without stopping +the container. + +```sql +CREATE USER 'backup'@'%' IDENTIFIED BY 'a-strong-password'; +GRANT SELECT, LOCK TABLES, SHOW VIEW, TRIGGER, PROCESS, RELOAD ON *.* TO 'backup'@'%'; +``` + +```bash +echo 'a-strong-password' > /root/.mariadb-backup.pw +chmod 600 /root/.mariadb-backup.pw +``` + +### 1.4 MariaDB root password file (used only by `restore.sh`) + +Restore needs CREATE/DROP privileges the `backup` user does not have. + +```bash +echo 'the-mariadb-root-password' > "$ROOT_PASSWORD_FILE" +chmod 600 "$ROOT_PASSWORD_FILE" +``` + +### 1.5 Exclude the MariaDB raw data directory + +Find the host directory bind-mounted into the container as its datadir and +drop a marker file in it. Borg skips any directory that contains `.nobackup` +(via `--exclude-if-present`), so live InnoDB files are never read: + +```bash +touch "${TARGET}/mariadb/data/.nobackup" +``` + +This is what allows backups to run with the container up: only the logical +dump under `${DUMP_SUBDIR}/` is ever archived or restored. + +### 1.6 Configure the Scaleway rclone remote + +```bash +rclone config +# Create a remote named "scaleway", type S3, with your Scaleway +# Object Storage credentials and region. +``` + +--- ## 2. Deploying the Scripts -### Single deployment (default) +### Single deployment -Copy `borg-backup.sh`, `dump_db.sh`, `restore.sh`, and your chosen config -from `configs/` to `/opt/backup-agent/`. Symlink the config as `backup.conf` -next to the scripts, or set `BACKUP_CONF` in the cron entry: +Copy `borg-backup.sh`, `dump_db.sh`, `restore.sh`, and the `configs/` directory +to `/opt/backup-agent/`. Symlink the relevant config as `backup.conf` beside +the scripts, or pass `BACKUP_CONF` explicitly in the cron entry. ```bash -chmod +x /opt/backup-agent/borg-backup.sh /opt/backup-agent/restore.sh -chmod +x /opt/backup-agent/dump_db.sh +chmod +x /opt/backup-agent/borg-backup.sh \ + /opt/backup-agent/restore.sh \ + /opt/backup-agent/dump_db.sh + # Option A — symlink (scripts auto-discover backup.conf beside them): ln -s /opt/backup-agent/configs/srv.conf /opt/backup-agent/backup.conf -# Option B — explicit env var in the cron entry (see §3). + +# Option B — explicit BACKUP_CONF in the cron entry (see §3). ``` ### Multiple deployments on different servers -Each server gets its own config file. Deploy the same three scripts to -`/opt/backup-agent/` on each server. Point each server's cron entry at its -config via `BACKUP_CONF`: +Each server gets its own config. Deploy the same three scripts to +`/opt/backup-agent/` on each server; point each cron entry at its config via +`BACKUP_CONF`: ``` # /etc/cron.d/borg-backup (nexusvoice server) @@ -76,10 +115,12 @@ BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf 30 2 * * * root /opt/backup-agent/borg-backup.sh >> /var/log/borg/cron.log 2>&1 ``` -Before running, verify/update `configs/nexusvoice.conf`: -- `TARGET` — confirm `/home/acid/nexusvoice` (or `/opt/nexusvoice` after the move) +Before the first run, verify `configs/nexusvoice.conf`: +- `TARGET` — current path; update if files move (e.g. `/home/acid/nexusvoice` → `/opt/nexusvoice`) - `DB_CONTAINER` — confirm the MariaDB container name on that server -- `REPO` — must not overlap with the srv repo; uses separate Scaleway path +- `REPO` — must not overlap with any other deployment's repo + +--- ## 3. Scheduling @@ -97,128 +138,139 @@ ls -lt /var/log/borg/backup-*.log | head -1 # latest log file tail -50 /var/log/borg/backup-*.log # inspect it ``` +--- + ## 4. Day-2 Operations -List archives: +Source your config first to get the right `BORG_REPO`, `BORG_PASSCOMMAND`, etc.: ```bash +source /opt/backup-agent/configs/nexusvoice.conf # adapt +export BORG_REPO="$REPO" +export BORG_PASSCOMMAND="cat $BORG_PASSPHRASE_FILE" +``` + +Then standard borg commands work without extra flags: + +```bash +# List archives ./restore.sh --list-archives +# or directly: +borg list + +# Repo size and health +borg info + +# Rotate the passphrase +# Note: only re-encrypts the key, not the archive data; old passphrase +# still needed for archives created before the change until fully migrated. +borg key change-passphrase + +# Break a stale lockfile (backup or restore aborted mid-run) +borg break-lock ``` -Check repo size and health: - -```bash -BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg info /home/srv/files/backups/borg-2025 -``` - -Rotate the passphrase (creates a new key, re-encrypts nothing — old -archives still need the old passphrase to read, so keep both until fully -migrated): - -```bash -BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg key change-passphrase /home/srv/files/backups/borg-2025 -``` - -Stale lockfile (backup or restore aborted mid-run and left the repo -locked): - -```bash -BORG_PASSCOMMAND="cat /root/.borg-passphrase" borg break-lock /home/srv/files/backups/borg-2025 -``` +--- ## 5. Recovery Procedures -All `restore.sh` commands accept `--dry-run` to preview exactly what would -happen without touching anything, and `--archive NAME` to target a -specific archive instead of the latest (see archive names via -`--list-archives`). Root DB credentials come from `MYSQL_ROOT_PASSWORD` in -the environment if set, otherwise from `/root/.mariadb-root.pw` — set -whichever is more convenient for how you're invoking it. Restore logs go -to `/var/log/borg/restore-*.log` (the `RESTORE_LOGDIR` environment -variable overrides the directory, mainly useful for testing). `full` and -`db` share `borg-backup.sh`'s lockfile, so a restore refuses to start -while the nightly backup is mid-run (and vice versa) rather than racing -it. +All `restore.sh` commands accept: +- `--dry-run` — preview exactly what would happen without touching anything +- `--archive NAME` — target a specific archive instead of the latest (get names via `--list-archives`) + +Root DB credentials come from `MYSQL_ROOT_PASSWORD` in the environment if set, +otherwise from the `ROOT_PASSWORD_FILE` in your config (default +`/root/.mariadb-root.pw`). Restore logs go to `/var/log/borg/restore-*.log`; +`RESTORE_LOGDIR` overrides the directory (useful for testing). + +`restore.sh` shares `borg-backup.sh`'s lockfile (`LOCKFILE` in the config), +so a restore refuses to start while the nightly backup is mid-run and vice +versa. + +For all `restore.sh` commands, select the deployment with `BACKUP_CONF`: + +```bash +export BACKUP_CONF=/opt/backup-agent/configs/nexusvoice.conf +``` + +Or run without it if `backup.conf` is already symlinked beside the script. ### 5.1 Full disaster recovery (new or wiped server) -Use when the whole server/container is gone and you're rebuilding from -scratch. +Use when the whole server/container is gone and you're rebuilding from scratch. ```bash # 1. Reinstall borg, docker, and the mariadb container image/compose file -# (not covered by restore.sh - this is infra provisioning). +# (not covered by restore.sh — this is infra provisioning). # 2. Restore the passphrase file (from your password manager / secondary -# backup - it is NOT stored in the repo it protects) to -# /root/.borg-passphrase, and the root DB password to -# /root/.mariadb-root.pw. +# backup — it is NOT stored in the repo it protects) to the path set in +# BORG_PASSPHRASE_FILE, and the root DB password to ROOT_PASSWORD_FILE. # 3. Preview: ./restore.sh full --dry-run -# 4. Run for real. --force is only required if /home/srv/files/content -# already has data in it (e.g. a stale mount); omit it on a genuinely -# empty/fresh server: +# 4. Run for real. --force is only required if $TARGET already has data in it +# (e.g. a stale mount); omit it on a genuinely empty/fresh server: ./restore.sh full --force ``` -This extracts the full content tree from the archive, starts the -`mariadb` container and waits for it to report healthy, then restores -every database dump (users/grants first) — the container must be running -before any of the dump restores, which is why it starts first. +This extracts the full content tree, starts the DB container and waits for it +to report healthy, then restores every database dump (users/grants first). **Verify afterward:** -- `docker ps` shows `mariadb` running and healthy. +- `docker ps` shows the DB container running and healthy. - The application responds normally. - Spot-check row counts on a couple of tables against what you'd expect. ### 5.2 Single database restore -Use when one database got corrupted or someone ran a bad migration/query -against it — this **drops and recreates** that database. +Use when one database got corrupted or someone ran a bad migration/query — +this **drops and recreates** that database. ```bash ./restore.sh db shopdb --dry-run # preview -./restore.sh db shopdb # prompts: type "shopdb" to confirm +./restore.sh db shopdb # prompts: type "shopdb" to confirm ``` Non-interactive (e.g. scripted from a monitoring alert): add `--yes` to skip the typed confirmation. -**Verify afterward:** connect to the database and check the tables/row -counts you expect. +**Verify afterward:** connect to the database and check the tables/row counts +you expect. ### 5.3 Single file/directory restore -Use for accidental deletion of a file, or to inspect an old version — this -never touches the running database or container. +Use for accidental deletion or to inspect an old version — never touches the +running database or container. ```bash -./restore.sh file path/relative/to/content/some-file.txt --dest /tmp/recovered +./restore.sh file path/relative/to/target/some-file.txt --dest /tmp/recovered ``` -The final location of the recovered item is printed at the end (it lands -under `/tmp/recovered/home/srv/files/content/...` — Borg preserves the -absolute path it was archived with). +The final location of the recovered item is printed at the end. Borg +preserves the absolute path it was archived with, so the file lands under +`/tmp/recovered//...`. + +--- ## 6. Restore Drill Cadence -Quarterly, run a real `full` restore into a scratch directory (not -`/home/srv/files/content`) to confirm backups are actually usable: +Quarterly, run a real `full` extract into a scratch directory to confirm +backups are actually usable. -`borg extract` always extracts into the current directory (there's no -`--destination` flag — this is why `restore.sh` itself `cd`s into the -destination before extracting), and the archive name must come from -`borg list --short` (plain `borg list` prints a formatted line, not a bare -name), so: +`borg extract` always extracts into the current directory; the archive name +must come from `borg list --short` (plain `borg list` prints a formatted +line, not a bare name). Source your config to get the right values: ```bash +source /opt/backup-agent/configs/nexusvoice.conf # adapt +export BORG_REPO="$REPO" +export BORG_PASSCOMMAND="cat $BORG_PASSPHRASE_FILE" + mkdir -p /tmp/restore-drill && cd /tmp/restore-drill -export BORG_REPO=/home/srv/files/backups/borg-2025 -export BORG_PASSCOMMAND="cat /root/.borg-passphrase" LATEST=$(borg list --short | tail -1) borg extract --lock-wait 600 "::$LATEST" ``` -Confirm the dump files under `mariadb/dump/` are present, non-empty, and -importable (`mysql -u root -p < mariadb/dump/somedb.sql` against a -throwaway MariaDB container). Log the drill date and outcome somewhere -durable (e.g. a team wiki page). +Confirm the dump files under `${DUMP_SUBDIR}/` are present, non-empty, and +importable (`mysql -u root -p < mariadb/dump/somedb.sql` against a throwaway +MariaDB container). Log the drill date and outcome somewhere durable (e.g. a +team wiki page).