Monitoring
How to watch an OBSESC deployment: the readiness check, the console, the built-in alerts and how to deliver them to your on-call.
Every OBSESC node reports on its own health. It answers a readiness check, shows its status in the console, and ships with built-in alerts that deliver to your Alertmanager. This page covers what to watch, and how to get alerts to the people on call.
What runs on each node
| Port | Purpose |
|---|---|
18080 | The console, the query API and the /ready check |
4317 | OTLP/gRPC ingest |
4318 | OTLP/HTTP ingest |
9200 | Elasticsearch _bulk ingest |
The node security group opens these ports to AllowedIngestCidr. See Network configuration.
The console
Open the UiUrl stack output. Without the UI load balancer, that’s the node’s query port (http://<node-private-ip>:18080/, or https:// if you set TlsCertSsmParameter). With EnableUiLoadBalancer=true, it’s the internal load balancer’s address. With the load balancer and EnableUiPerUserRbac=true, people sign in through your identity provider and each person works with their own role. See Security model.
The console’s Settings area shows the health of the node serving it.
Health check
Point load balancer and shipper health checks at GET :18080/ready. It doesn’t need a token.
200means the node is ready. The body can still carry a warning, for example that the node is applying backpressure, so read the body as well as the status code.503means the node has taken itself out of rotation. The body gives a short reason.
Quick check from the instance:
curl -s -w '\n%{http_code}\n' http://127.0.0.1:18080/ready
If you set TlsCertSsmParameter at deploy time, the node serves HTTPS on 18080. Use https:// in this command.
Built-in alerts
OBSESC ships alert rules that evaluate on every node and deliver to your Alertmanager. Each alert has a severity and says what to do:
- critical: act now. Data is at risk, or the node has stopped working even though it looks healthy from outside.
- warning: act today. A limit is getting close, or something needs your attention.
OBSESC’s built-in alerts will tell you when:
- the node is refusing ingest, or has been applying backpressure for a sustained period
- accepted data isn’t reaching your S3 bucket
- the write-ahead log volume is getting full, or part of it is damaged or unreadable
- the navigation tier is falling behind ingest, or may be inaccurate for a period
- alert delivery itself is failing
One alert fires all the time, on purpose, to prove the alert pipeline works. Another fires for as long as no Alertmanager is configured.
Troubleshooting walks through the most common situations.
Deliver alerts to your on-call
Out of the box, alerts evaluate but nothing pages anyone. The stack’s AlertDelivery output says NOT DELIVERING until you configure an Alertmanager.
To turn on delivery, update the stack with the AlertmanagerEndpoint parameter set to the host:port of an Alertmanager the node can reach:
STACK=my-obsesc
# Keep every other parameter at its current value.
PARAMS=$(aws cloudformation describe-stacks --stack-name "$STACK" \
--query 'Stacks[0].Parameters[?ParameterKey!=`AlertmanagerEndpoint`].ParameterKey' --output text \
| tr '\t' '\n' | sed 's/.*/ParameterKey=&,UsePreviousValue=true/' | tr '\n' ' ')
aws cloudformation update-stack --stack-name "$STACK" --use-previous-template \
--parameters ParameterKey=AlertmanagerEndpoint,ParameterValue=alertmanager.your-domain.internal:9093 $PARAMS \
--capabilities CAPABILITY_NAMED_IAM
The node applies this setting during first-boot setup, so a stack update that changes it replaces the node instance. Read Upgrades before you run a stack update on a live node. If you need to change alert delivery without replacing the instance, contact us.
Route the always-firing alert to a dead man’s switch (a heartbeat receiver that pages when the heartbeat stops), not to a person. If it goes quiet, your alerting is broken. If you’ve deliberately chosen not to page anyone, silence the “no Alertmanager” alert so the decision is on record.
Logs
The node logs to the systemd journal. Connect with SSM Session Manager (the AMI runs the SSM agent, and the node role allows it), then:
journalctl -u obsesc-node --since "1 hour ago" --no-pager
journalctl -u obsesc-node -f # follow
systemctl status obsesc-node --no-pager
If you need help reading what you find, contact us with the saved journal.