Docs: Rebuilding the navigation tier
Documentation / Operations

Rebuilding the navigation tier

Rebuild a node's navigation tier from the raw events in your S3 bucket after a lost, corrupt or replaced summary volume.

The navigation tier lives on the node’s summary EBS volume and makes exploring history fast. Your original events stay in your S3 bucket in open formats, and the navigation tier can be rebuilt from them.

What to expect

  • An empty summary volume is rebuilt from every raw event still in your bucket.
  • A summary volume restored from a snapshot only needs to catch up on data that arrived after the snapshot was taken.

AWS charges you as normal for the S3 requests, and for retrievals from Standard-IA or Glacier Instant Retrieval if your bucket’s lifecycle rules have moved objects there.

Limits

  • Only events still in your bucket can be rebuilt. Events that have passed your retention period and been deleted are gone for good.
  • Time ranges that aren’t rebuilt yet return no rows. The console reads the navigation tier, and returns an empty result for a period that hasn’t been rebuilt, not an error. Use SQL over the raw_events table (see SQL reference) to read those events in the meantime.
  • Rebuild time scales with the amount of raw data being re-read.

Choose your starting point

SituationDo this
You have a recent snapshot of the summary volumeRestore it (see Backups and recovery). The node only catches up from the time of the snapshot, which is the fastest option
The summary volume is intact but its contents are wrong or corruptFull rebuild on the existing volume (below)
The summary volume is goneRebuild onto a new volume (below)

Full rebuild on the existing volume

Don’t delete only part of the volume’s contents. Rebuilding on top of what’s left can make the navigation tier inaccurate. A full rebuild has to start from an empty volume.

  1. Stop writes to the node. Pause your producers or accept a short gap in ingest; shippers retry. Then stop the services:

    sudo systemctl stop obsesc.target
  2. Snapshot the summary volume, then empty it. The snapshot is your way back if you change your mind (see Backups and recovery). Once it’s complete, clear the volume:

    sudo find /var/lib/obsesc/summary -mindepth 1 -maxdepth 1 -exec rm -rf {} +
  3. Start the node and check it’s ready:

    sudo systemctl start obsesc.target
    curl -s -w '\n%{http_code}\n' http://127.0.0.1:18080/ready
  4. Resume traffic. Restart your producers once /ready is 200. You can accept new data while the rebuild runs, and the rebuild keeps going in the background. Older history reappears in the console as it’s rebuilt. Once the console shows the history you expect, you can delete the snapshot from step 2. If you need to confirm the rebuild has finished, contact us.

Rebuild onto a new volume

Use this when the summary volume has been deleted or can’t be used. The node finds each data volume by its device name. Anything attached at /dev/xvdg is the summary volume.

  1. Stop the services: sudo systemctl stop obsesc.target.

  2. Create an encrypted gp3 volume in the node’s Availability Zone. Tag it so the daily snapshot policy picks it up:

    STACK=my-obsesc
    VOL=$(aws ec2 create-volume --availability-zone <az> --size 500 \
      --volume-type gp3 --iops 6000 --throughput 500 --encrypted --kms-key-id <your-kms-key-arn> \
      --tag-specifications "ResourceType=volume,Tags=[{Key=Project,Value=$STACK},{Key=obsesc-purpose,Value=summary}]" \
      --query VolumeId --output text)
    aws ec2 wait volume-available --volume-ids "$VOL"
    aws ec2 attach-volume --volume-id "$VOL" --instance-id <instance-id> --device /dev/xvdg
  3. Format it and mount it at the summary path, replacing the old /etc/fstab entry:

    DEV=$(for n in /dev/nvme*n1; do ebsnvme-id -b "$n" 2>/dev/null | grep -q xvdg && echo "$n"; done)
    sudo mkfs.xfs -f "$DEV"
    sudo sed -i '\#/var/lib/obsesc/summary#d' /etc/fstab
    echo "UUID=$(sudo blkid -s UUID -o value "$DEV") /var/lib/obsesc/summary xfs defaults,nofail 0 2" | sudo tee -a /etc/fstab
    sudo mount /var/lib/obsesc/summary
    sudo chown obsesc:obsesc /var/lib/obsesc/summary && sudo chmod 0750 /var/lib/obsesc/summary
  4. Start the node and follow steps 3 and 4 of the full rebuild above.

The stack template still refers to the original summary volume. Before your next stack update that replaces the instance, such as an upgrade, bring the new volume under the stack. The simplest way is to import it as the SummaryVolume resource. Otherwise the update will try to attach the old volume at the same device name.

Multi-node deployments

The standard stack runs a single node. For a multi-node deployment, contact us before you rebuild.