Troubleshooting
Diagnose and fix common OBSESC problems: failed deployments, rejected ingest, missing recent data, S3 write failures and write-ahead log alerts.
Start with the node’s own report, then find the symptom below. Each built-in alert also says what to do; it’s the quickest guide when you’ve been paged.
First steps
On the node (connect with SSM Session Manager):
curl -s -w '\n%{http_code}\n' http://127.0.0.1:18080/ready # read the body, not just the code
systemctl status obsesc-node --no-pager
journalctl -u obsesc-node --since "30 min ago" --no-pager | tail -100
The /ready body names the problem when there is one. The console’s Settings area also shows the health of the node serving it.
The stack fails to create
The stack waits up to 15 minutes for the node to report that it’s up. If first-boot setup fails, the stack rolls back with a reason:
| Reason in the stack events | Cause |
|---|---|
userdata-failed-see-/var/log/cloud-init-output.log | A first-boot step failed. Read that log on the instance before the rollback terminates it, or re-create the stack with rollback disabled (--disable-rollback) so the instance stays up |
obsesc-node-not-ready-within-300s-see-journalctl | The node started but /ready didn’t return 200 within 5 minutes. Check journalctl -u obsesc-node |
Common causes:
- A data volume wasn’t attached in time.
- An
AccessDeniedreading the TLS or authentication parameters you passed: the node role can’t read them, or can’t use the key they’re encrypted with. - An S3 or KMS
AccessDeniedinjournalctl -u obsesc-node: the node role can’t reach the bucket or the key. For a bucket in another account, apply theRawBucketPolicySnippetstack output as that bucket’s policy. See IAM and permissions.
The node restarts over and over after a config change
The node’s configuration file is strict: an unknown or misspelt key stops the node from starting, and the service keeps restarting. The journal names the key it couldn’t load:
journalctl -u obsesc-node -n 50 --no-pager
Undo the edit and the node starts again. Hand edits are also lost whenever the instance is replaced, because first-boot setup rewrites the file from your stack parameters. Prefer stack parameters, and contact us before you change anything else.
Shippers get HTTP 503 or gRPC UNAVAILABLE
A 503 from ingest means the node is applying backpressure. It’s retryable, and shippers with retries enabled will buffer and resend. The /ready body and OBSESC’s built-in alerts tell you which of these it is:
| Cause | Fix |
|---|---|
| Data isn’t reaching S3 fast enough | Usually a permissions or bucket problem, not load. See “The node can’t write to S3” below |
| The node is at capacity | See Scaling |
| The write-ahead log volume is full | Grow it (see Scaling) and fix whatever is stopping data reaching S3 |
If the node is refusing every batch, /ready returns 503. Save the journal before you restart, because a restart hides the cause.
The console shows no data for recent time ranges
The console reads the navigation tier. A period that hasn’t reached the navigation tier yet returns an empty result, not an error.
- If OBSESC’s built-in alerts say navigation-tier data is waiting to be written, don’t restart the node until that clears. It usually points to a disk or storage problem on the node. Fix it and the node resumes on its own.
- If the built-in alerts say the navigation tier is falling behind, it’ll catch up, or see Scaling.
- If the built-in alerts say the navigation tier has stopped producing anything, save the journal, then restart the node. If it keeps happening, contact us.
In the meantime, the events are safe in your S3 bucket. Read them with SQL over the raw_events table (see SQL reference).
Navigation-tier integrity alerts
When a built-in alert says navigation-tier data may be inaccurate for a period, this usually follows a hard stop, such as the instance being killed or running out of memory. The raw events in S3 are unaffected.
- Save the journal before it rotates:
journalctl -u obsesc-node --since "24 hours ago" --no-pager > obsesc-node.log. - Use SQL over
raw_eventsas the authoritative answer for the affected period. If it matters, rebuild the navigation tier. - Find out why the node stopped hard:
journalctl -k | grep -i 'killed process'for out-of-memory kills, and the EC2 console for terminations. - Silence the alert for this instance until its next restart. Don’t silence the rule itself.
If you’re unsure, contact us with the saved journal.
The node can’t write to S3
Signs: a built-in alert that data isn’t reaching your bucket, or ingest backpressure that isn’t explained by load. The node keeps retrying failed S3 writes, so it recovers on its own once the cause is fixed.
- Check the node role’s S3 and KMS permissions, the bucket policy, and, for a bucket in another account, that account’s KMS key policy.
- If you use S3 Object Lock or a bucket in another account, re-check those settings after any stack update.
- If
/readyreports that some data can’t be written and retrying won’t help, it isn’t a permissions problem and it won’t clear on its own. The affected events are still held on the node’s write-ahead log volume, so don’t delete that volume. Save the journal and contact us.
Write-ahead log problems
If /ready reports that part of the write-ahead log is damaged or can’t be read, the node has taken itself out of rotation.
- Just after an upgrade, the new release may not be able to read what the old one wrote. Nothing is lost. Roll back (see Upgrades). A plain restart hits the same problem again.
- Otherwise, it’s usually disk damage. Events after the damaged point may not be in S3.
Either way, don’t delete anything on the write-ahead log volume. Save the journal and contact us. Resolve this before you upgrade.
Alerts never arrive
Check the stack’s AlertDelivery output. NOT DELIVERING means no Alertmanager is configured. See Monitoring. If an Alertmanager is configured and alerts still don’t arrive, check that the node can reach it on that address, then contact us.
No volume snapshots
The snapshot role can’t use your KMS key. Grant the stack’s snapshot lifecycle role access in the key’s key policy. See Backups and recovery.
Multi-node deployments
The standard stack runs a single node. Contact us about multi-node deployments.