Backups and recovery
What OBSESC stores where, how the built-in volume snapshots work, and how to restore a node's data after a failure.
OBSESC keeps your events in two places: your S3 bucket, which is the system of record, and fast local tiers on EBS volumes attached to the node. This page covers how each is protected and how to get a node back after a failure.
What lives where
| Data | Where | Protection | If you lose it |
|---|---|---|---|
| Raw events: every event, as Parquet with Iceberg metadata | Your S3 bucket | S3 durability. The stack never deletes the bucket. Optional S3 Object Lock | This is your source of truth. Protect it like any other system of record |
| Write-ahead log: acknowledged events not yet committed to S3 | WAL volume, /var/lib/obsesc/wal | Daily EBS snapshot | Can hold events that aren’t in S3 yet. Reattach the volume so they’re recovered |
| Navigation tier: what makes exploring history fast | Summary volume, /var/lib/obsesc/summary | Daily EBS snapshot | Rebuildable from your raw events (see Rebuilding the navigation tier) |
| Anomaly records | Anomaly volume, /var/lib/obsesc/anomaly | Daily EBS snapshot | Restore from a snapshot |
| Node configuration | Root volume | Regenerated from stack parameters at first boot | Hand edits are lost when the instance is replaced |
Once events are committed to S3, the node no longer needs them in its write-ahead log.
Built-in volume snapshots
With EnableDataVolumeSnapshots=true (the default), the stack creates an Amazon Data Lifecycle Manager policy. It snapshots every EBS volume tagged Project=<stack-name> once a day at 03:00 UTC and keeps the latest seven. Each volume’s tags, including obsesc-purpose (wal, summary or anomaly), are copied onto its snapshots.
KMS key policy. The data volumes are encrypted with the key you passed as KmsKeyArn. For the snapshot policy to work, that key’s key policy must allow the stack’s snapshot lifecycle role to use it. An IAM grant alone isn’t enough. See Encryption.
Check that snapshots are being taken:
STACK=my-obsesc
aws dlm get-lifecycle-policies --query 'Policies[].{id:PolicyId,state:State,desc:Description}' --output table
aws ec2 describe-snapshots --owner-ids self \
--filters "Name=tag:Project,Values=$STACK" "Name=tag:obsesc-purpose,Values=summary" \
--query 'sort_by(Snapshots,&StartTime)[-1].{id:SnapshotId,when:StartTime,state:State}' --output table
No rows means either the first 03:00 run hasn’t happened yet, or the snapshot role can’t use your KMS key.
Run a restore drill
An untested snapshot isn’t a backup. After the first snapshot lands, and then every quarter:
- Create a volume from the newest summary snapshot in the node’s Availability Zone.
- Attach it to a scratch instance (or to the node, at a spare device name such as
/dev/xvdr). - Mount it read-only and confirm it mounts cleanly and isn’t empty.
- Detach and delete the volume.
The commands are the same as steps 2, 3 and 6 of the copy-in restore below. Record the snapshot ID, its age and whether it mounted.
Restoring the navigation tier
Restoring a recent snapshot is faster than a full rebuild. The node then only catches up on data that arrived after the snapshot was taken.
Find the snapshot
aws ec2 describe-snapshots --owner-ids self \
--filters "Name=tag:Project,Values=$STACK" "Name=tag:obsesc-purpose,Values=summary" \
--query 'sort_by(Snapshots,&StartTime)[-1].{id:SnapshotId,when:StartTime,gib:VolumeSize}' --output table
aws ec2 describe-instances --instance-ids <instance-id> \
--query 'Reservations[].Instances[].Placement.AvailabilityZone' --output text
Copy-in restore
Use this when the stack’s summary volume still exists but its contents are bad. It keeps the stack template in step with what’s actually attached.
-
Stop your producers, then stop the services:
sudo systemctl stop obsesc.target -
Create a volume from the snapshot and attach it at a spare device name. Don’t use
xvdf,xvdgorxvdh; those belong to the data volumes.VOL=$(aws ec2 create-volume --snapshot-id <snap-id> --availability-zone <az> \ --volume-type gp3 --iops 6000 --throughput 500 \ --tag-specifications "ResourceType=volume,Tags=[{Key=Project,Value=$STACK},{Key=obsesc-purpose,Value=summary-restore}]" \ --query VolumeId --output text) aws ec2 wait volume-available --volume-ids "$VOL" aws ec2 attach-volume --volume-id "$VOL" --instance-id <instance-id> --device /dev/xvdr -
Mount it read-only and check it:
DEV=$(for n in /dev/nvme*n1; do ebsnvme-id -b "$n" 2>/dev/null | grep -q xvdr && echo "$n"; done) sudo mkdir -p /mnt/summary-restore sudo mount -o ro "$DEV" /mnt/summary-restore ls /mnt/summary-restore | head -
Replace the volume’s contents:
sudo find /var/lib/obsesc/summary -mindepth 1 -maxdepth 1 -exec rm -rf {} + sudo rsync -aHAX /mnt/summary-restore/ /var/lib/obsesc/summary/ sudo chown -R obsesc:obsesc /var/lib/obsesc/summary -
Start the node and check it’s ready:
sudo systemctl start obsesc.target curl -s -w '\n%{http_code}\n' http://127.0.0.1:18080/readyThe node catches up on anything that arrived after the snapshot in the background. Re-enable your producers once
/readyreturns200. -
Clean up the scratch volume:
sudo umount /mnt/summary-restore aws ec2 detach-volume --volume-id "$VOL" && aws ec2 wait volume-available --volume-ids "$VOL" aws ec2 delete-volume --volume-id "$VOL"
Don’t try to force a full rebuild by deleting some of the files on a restored volume. That can leave the navigation tier inaccurate. A full rebuild has to start from an empty volume (see Rebuilding the navigation tier).
Swap restore
Use this when the stack’s summary volume is gone. Create the volume from the snapshot as in step 2, but attach it at /dev/xvdg. Then point /etc/fstab at it and mount it at /var/lib/obsesc/summary, following the mount steps in “Rebuild onto a new volume” in Rebuilding the navigation tier, but without mkfs.xfs. Tag the volume obsesc-purpose=summary so the daily snapshots continue.
The stack template still names the original volume. Before any stack update that replaces the instance, such as an upgrade, import the restored volume into the stack as SummaryVolume. The alternative is a copy-in restore onto a fresh volume the stack manages.
Recovering write-ahead log data from a failed node
If a node stopped unexpectedly, its WAL volume may hold acknowledged events that aren’t in S3 yet. Don’t delete the WAL volume. When the node starts again with that volume attached, it recovers the events held there and commits them to S3. If the instance can’t be brought back, contact us before you change anything on the volume.
Protecting your raw events
- The stack never deletes your bucket. If the stack created it (
CreateRawBucket=true), the bucket has aRetaindeletion policy. If you brought your own bucket, the stack doesn’t manage it at all. - S3 Object Lock is available on a bucket the stack creates (
EnableObjectLock=true, withObjectLockModeandObjectLockRetainDays). It stops objects being deleted or overwritten before the retention period ends. It can only be turned on when the bucket is created. - Your raw events are standard Parquet with Iceberg metadata. Even with no OBSESC node running, you can read them with Athena, Trino, Spark or DuckDB.