Backup & disaster-recovery guide
How to back the platform up, and how to get it back after anything from a fat- fingered config change to a total host loss. Pairs with the administrator guide (§5 covers the everyday backup/ restore commands); this document is the complete strategy and the recovery runbooks.
The honeypot mindset shapes recovery too. The persona containers are disposable by design — losing one is a non-event, you rebuild it from the image. What actually needs protecting is the evidence (the events database and audit log) and the configuration that enforces containment (firewall, egress rules, VPN, compose project). Those are what the plan below protects.
1. What to protect (and what not to)
| Asset | Where | Back up? | Why |
|---|---|---|---|
| Events + audit log | PostgreSQL (pgdata volume) |
Yes — critical | The captured attack evidence. Irreplaceable. |
| Compose project | /opt/honeypot-platform (COMPOSE_DIR) |
Yes | Personas, service configs, migrations, .env. |
.env secrets |
inside the compose project | Yes, encrypted | Loss = new secrets + re-issue tokens. |
| Host firewall / egress | /etc/nftables.conf |
Yes | The containment boundary. |
| WireGuard config + keys | /etc/wireguard |
Yes, encrypted | Admin access. |
| SSH config | /etc/ssh/sshd_config |
Yes | Reproduces host hardening. |
| Persona containers / images | Docker | No | Stateless — rebuilt from Dockerfile + compose. |
| Cowrie/OpenCanary log volumes | Docker volumes | No | Transient; already normalized into PostgreSQL. |
| Tailscale node identity | /var/lib/tailscale |
No (re-auth) | Re-running setup-tailscale.sh re-authenticates; the ACL lives in the Tailscale console and as infrastructure/tailscale/acl-example.hujson in the repo. |
The single most important artifact is the PostgreSQL dump. Everything else is reproducible from the git repo plus your secrets; the captured events are not.
2. The backup procedure
infrastructure/backup/backup.sh captures the host configs, the compose project,
and a compressed PostgreSQL dump into one timestamped tarball, then prunes old
archives.
sudo BACKUP_ROOT=/var/backups/honeypot \
RETENTION=14 \
COMPOSE_DIR=/opt/honeypot-platform \
DB_CONTAINER=honeypot-dev-postgres-1 \
DB_NAME=honeypot \
DB_USER=honeypot \
infrastructure/backup/backup.sh
Produces /var/backups/honeypot/<YYYYMMDD-HHMMSS>.tar.gz containing
/etc/wireguard, /etc/nftables.conf, sshd_config, the compose project, and
honeypot.dump (a pg_dump -Fc custom-format dump). Archives beyond RETENTION
are deleted.
Set
DB_CONTAINER/DB_NAMEor the database is not dumped — the script skips the DB when they're empty. Confirm the container name withdocker ps --format '{{.Names}}'(compose projecthoneypot-dev→honeypot-dev-postgres-1).
2.1 Encrypt before it leaves the host
The tarball contains .env and VPN keys. Never store or move it unencrypted.
Encrypt to an age recipient (public-key — the private key stays off the host):
STAMP=$(ls -1t /var/backups/honeypot/*.tar.gz | head -1)
age -r age1yourpublickey... -o "$STAMP.age" "$STAMP" && shred -u "$STAMP"
Store *.age files; keep the age identity (private key) somewhere separate from
the backups — a password manager or an offline device. Without it the backups are
useless to an attacker and to you, so protect it deliberately.
2.2 Schedule it
Mirror the disk-monitor timer that harden.sh already installs. Create
/etc/systemd/system/honeypot-backup.service:
[Unit]
Description=Honeypot platform backup
[Service]
Type=oneshot
Environment=BACKUP_ROOT=/var/backups/honeypot RETENTION=14
Environment=DB_CONTAINER=honeypot-dev-postgres-1 DB_NAME=honeypot
ExecStart=/opt/honeypot-platform/infrastructure/backup/backup.sh
and /etc/systemd/system/honeypot-backup.timer:
[Unit]
Description=Nightly honeypot backup
[Timer]
OnCalendar=*-*-* 02:30:00
Persistent=true
[Install]
WantedBy=timers.target
sudo systemctl daemon-reload && sudo systemctl enable --now honeypot-backup.timer
systemctl list-timers honeypot-backup.timer # confirm next run
Add the encrypt-and-ship step (§2.1 + a copy to offsite storage) either inside a
wrapper ExecStart script or as a second ExecStartPost= line.
2.3 Follow 3-2-1
- 3 copies of the data,
- on 2 different media,
- 1 of them offsite and offline.
Concretely: the on-host /var/backups/honeypot, an encrypted copy pulled to a
management workstation, and an encrypted copy in object storage or on a rotated
external disk. The offsite copy is your defence against the host being destroyed,
ransomwared, or seized. Because backups contain secrets, the offsite copy being
age-encrypted is non-negotiable.
3. The restore procedure
sudo COMPOSE_DIR=/opt/honeypot-platform \
DB_CONTAINER=honeypot-dev-postgres-1 \
DB_NAME=honeypot DB_USER=honeypot \
infrastructure/backup/restore.sh /var/backups/honeypot/<stamp>.tar.gz
(If the archive is age-encrypted, decrypt first:
age -d -i identity.txt <stamp>.tar.gz.age > <stamp>.tar.gz.)
Restore extracts the archive and:
- Restores
/etcconfigs (saving current ones as*.bak). - Restores the compose project to
COMPOSE_DIR. - If a
*.dumpis present andDB_CONTAINER/DB_NAMEare set, runspg_restore --clean --if-existsinto the database. - Reloads nftables and restarts WireGuard.
After it finishes: make -C server up (if the stack isn't running),
make -C server migrate (bring schema to head — harmless if already current),
then make -C server smoke to prove the pipeline works end to end.
4. Recovery targets
| Target | Rationale | |
|---|---|---|
| RPO (max data loss) | ≤ 24 h | Nightly backup. Tighten to hourly for the DB if the evidence stream is high-value. |
| RTO — persona | minutes | down && up; stateless rebuild. |
| RTO — DB restore | ≤ 30 min | One pg_restore on existing hardware. |
| RTO — full host rebuild | 1–3 h | Provision host + harden + restore + verify. |
If you need a tighter RPO on the evidence, add an hourly DB-only dump alongside
the nightly full backup (pg_dump -Fc on its own is cheap and fast).
5. Disaster runbooks
5.1 A persona is compromised (or misbehaving)
Impact: none to the platform — this is the designed outcome. Nothing an attacker did inside a persona persists.
make -C server down && make -C server up # every persona back to clean state
If you suspect the host boundary was tested, also re-verify containment (§6) before returning to normal. No backup needed — the containers are rebuilt from images.
5.2 Accidental config change / bad deploy
Symptom: a firewall edit locked something out, or a deploy broke ingestion.
- Config in git:
git checkout <good-commit> -- <path>, rebuild,make up. - Config not in git (e.g. hand-edited
/etc/nftables.conf): restore just that file from the latest backup, or re-run the infra scripts (sudo nft -f infrastructure/host/nftables.conf). - A migration went wrong:
make -C server migrate-downto step back, fix the revision, re-apply.
5.3 Database corruption or data loss
# Restore only the DB from the most recent good archive:
tar -xzf <stamp>.tar.gz -C /tmp/restore
docker exec -i honeypot-dev-postgres-1 pg_restore -U honeypot -d honeypot \
--clean --if-exists < /tmp/restore/<stamp>/honeypot.dump
make -C server smoke
You lose only events since the last backup (your RPO). The audit log restores with the rest of the database.
5.4 Total host loss (hardware death, ransomware, seizure)
Full rebuild on new hardware:
- Provision a clean Debian/Ubuntu host in the isolated VLAN/DMZ.
- Harden:
sudo infrastructure/host/harden.sh. - Recover the repo:
git cloneyour private repo to/opt/honeypot-platform. - Restore from the offsite encrypted backup (decrypt, then §3). This brings
back
.env, host configs, and the database. - Firewall + egress:
sudo nft -f infrastructure/host/nftables.confandsudo infrastructure/host/docker-egress.sh. - Admin VPN: re-run
setup-tailscale.sh(re-authenticate the node) or restore WireGuard; re-apply the Tailscale ACL. - Bring it up:
make -C server up && make -C server migrate && make -C server smoke. - Verify containment (§6) before exposing the trap ports.
If you have no backup but do have the git repo and your secrets, you can still rebuild a working platform from steps 1–3, 5–8 — you only lose the historical events. That's the worst case, and it's why the repo + secrets + offsite DB dump together are the recovery kit.
5.5 Ransomware reaching the backups
This is why §2.1 and §2.3 exist. If the on-host and management-workstation copies
are encrypted or destroyed, recover from the offline offsite copy. Backups
that are always-online and unencrypted are not a DR plan — they're a second
victim. Keep at least one copy offline and age-encrypted.
6. Verify containment after any recovery
A restore is not "done" until you've re-confirmed the honeypot can't reach your real network — the whole point of the platform. From the host:
# Honeypot subnet must NOT reach the management plane or the admin VPN:
sudo nft list ruleset | grep -A2 honeypot # drop rules present
# A persona container must have no outbound internet:
docker compose -f server/docker/docker-compose.dev.yml exec cowrie \
sh -c 'curl -m5 https://example.com' ; echo "exit=$?" # must fail/time out
Both checks must show containment holding. Only then re-expose the trap ports.
7. Test-restore drill (do this — an untested backup is a guess)
Quarterly, prove the backups actually restore:
- Copy the latest encrypted archive to a throwaway host (or VM).
- Decrypt, run
restore.sh,make up,make migrate,make smoke. - Log in via the API and confirm event counts match what you expect
(
GET /stats). - Tear the throwaway host down.
Record the date and result. A backup you have never restored is a hope; a backup you restored last quarter is a plan.
8. Checklist
-
backup.shscheduled nightly via systemd timer, withDB_CONTAINER/DB_NAMEset. - Backups
age-encrypted; the identity stored separately from the backups. - 3-2-1: on-host + management copy + offline offsite copy.
- Retention set (e.g. 14) and old archives pruning.
- Git repo is private and current (it's half your recovery kit).
- Test-restore drill run and dated within the last quarter.
- Containment re-verified after the last restore.