Docs: Configuration examples
Documentation / Ingestion

Configuration examples

Node-side ingest settings and ready-to-adapt configurations for AWS sources, Loki, GELF, Kafka, journald, Windows Event Log, rsyslog and Cribl.

This page collects the node-side ingest settings, plus configurations for sources that don’t have their own page. Every shipper example here is dual-write: your existing destination stays in place and OBSESC is added next to it.

Where node settings live

On the AMI, the node reads /etc/obsesc/config.yaml. Ingest settings sit under ingest:, and authentication sits under security:. After you edit the file, restart the node so it picks up the change:

sudo systemctl restart obsesc-node

Tokens and keys are read only at startup, so a restart is also needed after you rotate them. For every setting, see the Configuration reference.

The CloudFormation stack writes this file when it creates the instance. If a stack update replaces the instance (for example, to move to a new AMI), the file is written again from the template, so keep a copy of any changes you make by hand and reapply them afterwards. Contact us if you’d like help carrying settings across an upgrade.

Authentication

You can set ingest credentials directly in config.yaml:

security:
  ingest_tokens:          # OTLP, Elasticsearch bulk, Vector native, Loki
    - "<long-random-token>"
  hec_tokens:             # Splunk HEC. An empty list rejects everything
    - "<long-random-hec-token>"
  fluent_shared_key: "<long-random-key>"   # Fluent Forward handshake

The better option is to keep them out of the file. Store them in an AWS Secrets Manager secret and point the node at it. When you deploy with the CloudFormation stack’s AuthSecretArn parameter, this is wired up for you:

security:
  secrets_manager_secret_arn: "arn:aws:secretsmanager:ap-southeast-2:123456789012:secret:obsesc-auth"

The secret’s value is a JSON object. Any field present in it overrides the matching YAML field. When the stack sets AuthSecretArn, include a query_api_token as well. Unless console sign-in through the load balancer is configured, the node won’t start without one.

{
  "query_api_token": "<token>",
  "ingest_tokens": ["<token>"],
  "hec_tokens": ["<hec-token>"],
  "fluent_shared_key": "<key>",
  "firehose_access_key": "<key>",
  "kafka_sasl_password": "<password>"
}

The node’s instance role needs secretsmanager:GetSecretValue on that secret. Setting up the secret through the stack grants this for you. For TLS on the listeners, set the stack’s TLS certificate parameters. See Encryption.

Service naming

ingest:
  default_service: unknown                  # where events with no service land
  service_from: kubernetes.namespace_name   # optional: derive a missing service from this attribute

The node treats an event as having no service when its service is empty, unknown, default, or the OpenTelemetry SDK’s unknown_service[:<process>]. service_from is checked first for those events, and default_service applies if that attribute is missing too. A service the shipper does set always wins.

Dimensions

The node fills four built-in dimensions (host, env, namespace and tenant) from each shipper’s native attribute names. It checks keys such as host.name, hostname, kubernetes.host, deployment.environment, k8s.namespace.name, kubernetes.namespace_name and tenant_id. You can override the candidate keys globally or for a single protocol, or assign a tenant per ingest token:

ingest:
  dimensions:
    enabled: true
    env: ["deployment.environment", "stage"]
    per_source:
      fluent-forward:
        host: ["kubernetes.host"]
    tenant_by_token:
      - token: "<ingest-token-for-team-a>"
        tenant: team-a

per_source accepts otlp-http, otlp-grpc, es-bulk, splunk-hec, fluent-forward and vector-native. A token’s tenant replaces any tenant the shipper sent. It applies to OTLP, Elasticsearch bulk, HEC and Vector.

Multiline events

Assemble multiline events in the shipper. The node treats each record as one event, so a 30-line stack trace sent as 30 records becomes 30 events. Every shipper example in this section turns on multiline support: multiline.parser in Fluent Bit, parsers: multiline in Filebeat, [sources.*.multiline] in Vector, and the recombine operator for an OTel Collector filelog receiver.

Two paths have no shipper stage to do this: a forwarder posting to HEC /services/collector/raw, and a Fluent Forward source you can’t put a parser in front of. For those, the node can do the assembly itself:

ingest:
  multiline:
    enabled: true                        # default false
    start_pattern: '^\d{4}-\d{2}-\d{2}'  # optional; otherwise common stack-trace formats are recognised

The trade-off: the node acknowledges each line as it arrives, but briefly holds an unfinished multiline event while it waits for more lines. If the node stops during that time, that event can be lost. Don’t turn it on for data a shipper has already assembled. Contact us if you need to tune node-side multiline assembly.

Timestamp and size guards

ingest:
  max_future_skew_secs: 604800   # 7 days; 0 disables
  max_event_age_secs: 0          # 0 = no lower bound (the default)
  timestamp_action: clamp        # or `reject`
  max_event_bytes: 8388608       # per event; over it rejects the batch
  max_attrs_per_event: 4096

An event timestamped more than 7 days in the future is clamped to that boundary by default. With reject, the whole batch is refused instead. Size and attribute limits always reject the whole batch, never truncate it. See Limits.

Amazon Data Firehose (CloudWatch Logs)

This path needs no shipper. A CloudWatch Logs subscription filter streams a log group to a Firehose delivery stream, and the stream’s HTTP endpoint destination is the node.

ingest:
  firehose:
    enabled: true
    port: 8555
    access_key: "<long-random-key>"   # required; the listener fails closed without it
    # service: lambda                 # optional: one service for the whole stream

Firehose delivers only to HTTPS, so the node needs TLS configured, or an HTTPS load balancer in front of port 8555. Add an inbound security group rule for the port as well. Recommended delivery-stream settings: GZIP content encoding, 1 MiB / 60 s buffering, S3 backup for failed data only, and a service common attribute to name the stream. By default, each CloudWatch log group (for example /aws/lambda/checkout) becomes the service. If a log message is itself a JSON object with a service, service_name, app or job key, that key names the service instead. Delivery is at-least-once.

Loki push (Promtail, Grafana Alloy)

ingest:
  loki:
    enabled: true
    port: 3100

Promtail with two clients: your existing Loki, plus OBSESC.

clients:
  - url: http://loki.your-domain.internal:3100/loki/api/v1/push      # existing, unchanged
  - url: http://obsesc.your-domain.internal:3100/loki/api/v1/push    # added
    bearer_token: "<token from security.ingest_tokens>"
    backoff_config:
      min_period: 500ms
      max_period: 30s
      max_retries: 20

scrape_configs:
  - job_name: app
    static_configs:
      - targets: [localhost]
        labels:
          job: app
          app: checkout
          __path__: /var/log/app/*.log
    pipeline_stages:
      - multiline:
          firstline: '^\d{4}-\d{2}-\d{2}'
          max_wait_time: 1s

The service comes from the first label present out of service, service_name, app and job. Other labels and structured metadata become attributes, and X-Scope-OrgID becomes the tenant attribute. The node returns 204 once the batch is durable.

GELF (Docker, log4j2, logback)

ingest:
  gelf:
    enabled: true
    port: 12201        # UDP and TCP
    udp: true
    tcp: true

Docker daemon (/etc/docker/daemon.json):

{
  "log-driver": "gelf",
  "log-opts": {
    "gelf-address": "tcp://obsesc.your-domain.internal:12201",
    "tag": "{{.Name}}"
  }
}

GELF has no authentication and no acknowledgement, so restrict network access to the port. Prefer TCP: UDP loses any datagram that’s dropped. The service comes from _service, _app, _application, _container_name or _image_name, in that order.

Kinesis Data Streams

ingest:
  kinesis:
    enabled: true
    stream_name: app-logs
    start: latest                # or earliest; only used the first time the stream is read

The node’s instance role needs kinesis:ListShards, kinesis:DescribeStreamSummary, kinesis:GetShardIterator and kinesis:GetRecords on the stream, plus kms:Decrypt if the stream is encrypted. CloudWatch Logs subscription payloads are unpacked into one event per log event. Disable KPL record aggregation on your producers, because aggregated records aren’t split.

S3 objects through SQS

ingest:
  sqs:
    enabled: true
    queue_url: https://sqs.ap-southeast-2.amazonaws.com/123456789012/obsesc-s3-events

Send s3:ObjectCreated:* notifications to the queue, and give the queue a dead-letter queue. The node’s role needs sqs:ReceiveMessage, sqs:DeleteMessage and sqs:GetQueueAttributes on the queue, and s3:GetObject on the bucket prefix. Objects are read as NDJSON or text lines, and gzip is detected automatically. A message is deleted only after every event in it is durable, so set the queue’s visibility timeout above the time your largest object takes to ingest.

Kafka

ingest:
  kafka:
    enabled: true
    brokers: ["b-1.my-cluster.kafka.ap-southeast-2.amazonaws.com:9096"]
    topic: app-logs
    start: latest
    tls: true
    sasl_mechanism: scram-sha-512  # plain | scram-sha-256 | scram-sha-512
    sasl_username: obsesc
    # sasl_password: set as kafka_sasl_password in the Secrets Manager secret

The node reads the topic as an independent consumer. It uses no consumer group, so your existing consumers aren’t affected. On Amazon MSK, use SASL/SCRAM with TLS. MSK IAM authentication and client-certificate mTLS aren’t supported. Producers must use gzip, snappy or zstd compression, because lz4 isn’t supported.

journald

Use Vector’s journald source with the native Vector sink:

[sources.journal]
type = "journald"

[transforms.journal_shape]
type = "remap"
inputs = ["journal"]
source = '''
unit = string(._SYSTEMD_UNIT) ?? "journal"
.service = replace(unit, r'\.service$', "")
.host = ._HOSTNAME
.level = .PRIORITY
'''

[sinks.obsesc]
type = "vector"
inputs = ["journal_shape"]
address = "obsesc.your-domain.internal:9000"

Windows Event Log

Vector on Windows, using the windows_event_log source with the same vector sink. Alternatively, Winlogbeat with output.elasticsearch, as shown on the Filebeat page. Ship the rendered message rather than the raw XML.

rsyslog

OBSESC has no syslog listener. rsyslog sends through omelasticsearch (the rsyslog-elasticsearch package) to the bulk port. Save this as /etc/rsyslog.d/50-obsesc.conf:

module(load="omelasticsearch")

template(name="obsesc-json" type="list" option.jsonf="on") {
    property(outname="@timestamp" name="timereported" dateFormat="rfc3339" format="jsonf")
    property(outname="service" name="programname" format="jsonf")
    property(outname="host" name="hostname" format="jsonf")
    property(outname="severity" name="syslogseverity-text" format="jsonf")
    property(outname="message" name="msg" format="jsonf")
}

*.* action(type="omelasticsearch"
           server="obsesc.your-domain.internal"
           serverport="9200"
           searchIndex="logs"
           bulkmode="on"
           template="obsesc-json"
           action.resumeRetryCount="-1")

omelasticsearch supports basic authentication only, which the node doesn’t accept. If security.ingest_tokens is enforced, either restrict port 9200 to your syslog hosts or relay through a shipper that can send a token.

Cribl Stream

Add an Elasticsearch destination with Bulk API URL http://obsesc.your-domain.internal:9200/_bulk. If ingest tokens are enforced, add the extra header Authorization: ApiKey <token>. Alternatively, use a Splunk HEC destination pointing at /services/collector/event. Keep your existing route and add a second route or a Clone function. Make sure a service field is set, adding it with an Eval function if needed.