Rebuilding the navigation tier
Rebuild a node's navigation tier from the raw events in your S3 bucket after a lost, corrupt or replaced summary volume.
The navigation tier lives on the node’s summary EBS volume and makes exploring history fast. Your original events stay in your S3 bucket in open formats, and the navigation tier can be rebuilt from them.
What to expect
- An empty summary volume is rebuilt from every raw event still in your bucket.
- A summary volume restored from a snapshot only needs to catch up on data that arrived after the snapshot was taken.
AWS charges you as normal for the S3 requests, and for retrievals from Standard-IA or Glacier Instant Retrieval if your bucket’s lifecycle rules have moved objects there.
Limits
- Only events still in your bucket can be rebuilt. Events that have passed your retention period and been deleted are gone for good.
- Time ranges that aren’t rebuilt yet return no rows. The console reads the navigation tier, and returns an empty result for a period that hasn’t been rebuilt, not an error. Use SQL over the
raw_eventstable (see SQL reference) to read those events in the meantime. - Rebuild time scales with the amount of raw data being re-read.
Choose your starting point
| Situation | Do this |
|---|---|
| You have a recent snapshot of the summary volume | Restore it (see Backups and recovery). The node only catches up from the time of the snapshot, which is the fastest option |
| The summary volume is intact but its contents are wrong or corrupt | Full rebuild on the existing volume (below) |
| The summary volume is gone | Rebuild onto a new volume (below) |
Full rebuild on the existing volume
Don’t delete only part of the volume’s contents. Rebuilding on top of what’s left can make the navigation tier inaccurate. A full rebuild has to start from an empty volume.
-
Stop writes to the node. Pause your producers or accept a short gap in ingest; shippers retry. Then stop the services:
sudo systemctl stop obsesc.target -
Snapshot the summary volume, then empty it. The snapshot is your way back if you change your mind (see Backups and recovery). Once it’s complete, clear the volume:
sudo find /var/lib/obsesc/summary -mindepth 1 -maxdepth 1 -exec rm -rf {} + -
Start the node and check it’s ready:
sudo systemctl start obsesc.target curl -s -w '\n%{http_code}\n' http://127.0.0.1:18080/ready -
Resume traffic. Restart your producers once
/readyis200. You can accept new data while the rebuild runs, and the rebuild keeps going in the background. Older history reappears in the console as it’s rebuilt. Once the console shows the history you expect, you can delete the snapshot from step 2. If you need to confirm the rebuild has finished, contact us.
Rebuild onto a new volume
Use this when the summary volume has been deleted or can’t be used. The node finds each data volume by its device name. Anything attached at /dev/xvdg is the summary volume.
-
Stop the services:
sudo systemctl stop obsesc.target. -
Create an encrypted gp3 volume in the node’s Availability Zone. Tag it so the daily snapshot policy picks it up:
STACK=my-obsesc VOL=$(aws ec2 create-volume --availability-zone <az> --size 500 \ --volume-type gp3 --iops 6000 --throughput 500 --encrypted --kms-key-id <your-kms-key-arn> \ --tag-specifications "ResourceType=volume,Tags=[{Key=Project,Value=$STACK},{Key=obsesc-purpose,Value=summary}]" \ --query VolumeId --output text) aws ec2 wait volume-available --volume-ids "$VOL" aws ec2 attach-volume --volume-id "$VOL" --instance-id <instance-id> --device /dev/xvdg -
Format it and mount it at the summary path, replacing the old
/etc/fstabentry:DEV=$(for n in /dev/nvme*n1; do ebsnvme-id -b "$n" 2>/dev/null | grep -q xvdg && echo "$n"; done) sudo mkfs.xfs -f "$DEV" sudo sed -i '\#/var/lib/obsesc/summary#d' /etc/fstab echo "UUID=$(sudo blkid -s UUID -o value "$DEV") /var/lib/obsesc/summary xfs defaults,nofail 0 2" | sudo tee -a /etc/fstab sudo mount /var/lib/obsesc/summary sudo chown obsesc:obsesc /var/lib/obsesc/summary && sudo chmod 0750 /var/lib/obsesc/summary -
Start the node and follow steps 3 and 4 of the full rebuild above.
The stack template still refers to the original summary volume. Before your next stack update that replaces the instance, such as an upgrade, bring the new volume under the stack. The simplest way is to import it as the SummaryVolume resource. Otherwise the update will try to attach the old volume at the same device name.
Multi-node deployments
The standard stack runs a single node. For a multi-node deployment, contact us before you rebuild.