Docs: Upgrades
Documentation / Operations

Upgrades

Move an OBSESC stack to a new AMI release without losing acknowledged data, verify the result, and roll back if needed.

An upgrade replaces the node’s instance with one booted from the new OBSESC AMI. Your data volumes and S3 bucket carry over unchanged. The steps below keep every acknowledged event safe along the way: check compatibility, stop ingest, snapshot, roll, verify, then resume.

Before you start

You need:

  • the new release’s AMI ID and version string
  • shell access to the node (SSM Session Manager works out of the box)
  • AWS CLI credentials that can update the CloudFormation stack and create EBS snapshots

Stack parameters removed in a release must be dropped from your update, because CloudFormation rejects parameters the template no longer declares. The release notes list any such changes.

1. Check compatibility

Read the release notes for the version you’re moving to. If you’re unsure whether you can move straight to it, or if your node currently reports a write-ahead log problem, contact us before you go further.

2. Stop ingest

Stop your producers, then give the node a few minutes to commit everything it has accepted to S3. Then stop the services:

sudo systemctl stop obsesc.target

Stopping with data still pending is safe, because the node recovers it on the next start. But that data is then read by the new release, which is the case you want to avoid. If you can’t stop producers cleanly, contact us before you roll.

3. Snapshot

The daily snapshots can be up to 24 hours old. Take fresh ones so your rollback point is the state you just stopped at:

STACK=my-obsesc
INSTANCE=<instance-id>
NEW=<new-version>
for purpose in wal summary anomaly; do
  vol=$(aws ec2 describe-volumes \
    --filters "Name=tag:Project,Values=$STACK" "Name=tag:obsesc-purpose,Values=$purpose" \
              "Name=attachment.instance-id,Values=$INSTANCE" \
    --query 'Volumes[0].VolumeId' --output text)
  aws ec2 create-snapshot --volume-id "$vol" \
    --description "$STACK pre-upgrade $purpose" \
    --tag-specifications "ResourceType=snapshot,Tags=[{Key=Project,Value=$STACK},{Key=obsesc-purpose,Value=$purpose},{Key=obsesc-pre-upgrade,Value=$NEW}]"
done
aws ec2 wait snapshot-completed --filters "Name=tag:obsesc-pre-upgrade,Values=$NEW"

4. Roll to the new AMI

Update the stack with the new AmiId. Keep every other parameter as it is:

PARAMS=$(aws cloudformation describe-stacks --stack-name "$STACK" \
  --query 'Stacks[0].Parameters[?ParameterKey!=`AmiId`].ParameterKey' --output text \
  | tr '\t' '\n' | sed 's/.*/ParameterKey=&,UsePreviousValue=true/' | tr '\n' ' ')
aws cloudformation update-stack --stack-name "$STACK" --use-previous-template \
  --parameters ParameterKey=AmiId,ParameterValue=<new-ami-id> $PARAMS \
  --capabilities CAPABILITY_NAMED_IAM
aws cloudformation wait stack-update-complete --stack-name "$STACK"

If the new release ships a new template, pass it instead of --use-previous-template.

What happens during the update:

  • The instance is replaced. CloudFormation creates a new instance, moves the three data volumes across, and terminates the old one. The volumes are stack resources, so their data is kept.
  • First-boot setup runs again on the new instance and rewrites the node’s configuration file from the stack parameters. Any change you made to that file by hand is lost.
  • The node’s private IP changes. Point shippers at the stack outputs (IngestEndpointOtlp, IngestEndpointEs) or a load balancer, not at an IP address you copied down. Open the console from the UiUrl output.
  • If you restored a summary volume out of band with a swap restore, fold it back into the stack first. See Backups and recovery.

If the update ends in UPDATE_ROLLBACK_COMPLETE, the old instance is back in service. If it ends in UPDATE_ROLLBACK_FAILED, use aws cloudformation continue-update-rollback. Don’t delete and recreate the stack. The raw bucket survives a delete, but the data volumes don’t.

5. Verify before resuming traffic

On the new instance:

curl -s -w '\n%{http_code}\n' http://127.0.0.1:18080/ready
journalctl -u obsesc-node --since "15 min ago" --no-pager | tail -100

A good upgrade looks like this:

  • /ready returns 200 with the body ready
  • the journal shows no errors
  • the console opens from UiUrl and a query runs end to end

If something’s wrong:

SignalMeaningAction
/ready 503, reporting that the write-ahead log can’t be readThe new release can’t read data the old one wrote. Nothing is lost: the data is held on the write-ahead log volumeRoll back (below)
/ready 503, reporting that the write-ahead log is damagedDamaged data on the volume, unrelated to the upgradeNot a reason to roll back. See Troubleshooting
/ready 503, reporting that some data can’t be written to S3Retrying won’t help. The events are still on the write-ahead log volumeSave journalctl -u obsesc-node and contact us

6. Resume

Re-enable your producers, and check that new events appear in the console.

Rolling back

The usual case: run step 4 again with the previous AmiId. The old release reads the data it wrote itself, and /ready returns to 200 once the node has recovered.

If the old release refuses to start because the newer release wrote data it can’t read, that’s deliberate: it refuses rather than silently dropping data. Start the newer release long enough for it to commit everything to S3, stop it, then roll back. If the newer release can’t run at all, contact us before you change anything on the write-ahead log volume.