Docs: Fidelity tier
Documentation / Concepts

Fidelity tier

How OBSESC stores every event as open Parquet files with Iceberg metadata in your own S3 bucket.

The fidelity tier is the territory: every event OBSESC accepts, stored complete in your own S3 bucket as open columnar files. It is the source of truth for everything else OBSESC does, and you can read it without OBSESC.

What “complete” means

Each event keeps its timestamp, service, original body bytes as your shipper sent them, and every attribute key and value. Nothing is sampled. OBSESC adds a few derived columns alongside the originals. For example, when an event carries a host attribute, a host column is added. Native attributes are never rewritten. The one exception is redaction you configure at ingest, which runs before anything is stored. See Data flow.

Where it lives

With the CloudFormation template you can use an existing bucket (the default, CreateRawBucket=false, with RawBucketName) or let the stack create one (CreateRawBucket=true). The node writes under these prefixes:

PrefixContents
raw/Parquet data files
catalog/OBSESC’s table catalog and the Iceberg metadata under catalog/iceberg/metadata/. Do not modify
cluster/Reserved for multi-node deployments
query/audit/Query audit trail
holds/Legal-hold registry

The node’s IAM policy only allows object and list actions on these prefixes, so it cannot see anything else you keep in the bucket. Use one OBSESC deployment per bucket.

Object layout

Data files are partitioned by the event time of their earliest event, not by the time they were written:

s3://my-obsesc-bucket/raw/<yyyy>/<mm>/<dd>/<hh>/<node_id>-<uuid>.parquet

If a single write contains events from more than one UTC day, for example a shipper replaying a buffer after an outage, OBSESC splits it into one object per day. Last Tuesday’s events land under last Tuesday’s prefix, so prefix-scoped S3 lifecycle rules and anyone browsing the bucket see data where its events belong.

File format

  • Apache Parquet, compressed with zstd.
  • Rows sorted by service and time within each file.
  • Standard Parquet statistics, so external engines can skip data that doesn’t match.

Columns

ColumnTypeNotes
timestamp_nsint64Event time, nanoseconds since the Unix epoch
sourceuint8Ingest protocol, e.g. 0 OTLP/HTTP, 2 Elasticsearch bulk, 3 Splunk HEC
servicestringLogical service name
bodybinaryThe original message bytes
idempotency_keyfixed binary (16)Stable identifier used for deduplication
attr_keyslist<string>Attribute keys, parallel with attr_values
attr_valueslist<binary>Attribute values
host, env, namespace, tenantstring, nullableCanonical dimensions, derived from the matching attribute. Null when the event has no such attribute
event_timetimestamp (µs, UTC)Same instant as timestamp_ns. The Iceberg partition source
attributesmap<string, binary>Optional, off by default. Lets external engines read attributes by key (attributes['trace_id']). Contact us to turn it on

The schema only ever gains nullable columns. Older files read the newer columns as null, so files of different ages sit in one table.

Iceberg metadata

OBSESC keeps Apache Iceberg v2 metadata for the files in your bucket up to date:

s3://my-obsesc-bucket/catalog/iceberg/metadata/v<N>.metadata.json
s3://my-obsesc-bucket/catalog/iceberg/metadata/version-hint.text

The exported table:

  • is partitioned by day(event_time)
  • supports ordinary scans and time travel across snapshots
  • resolves columns by name, so older files still read correctly
  • does not support Iceberg incremental or changelog scans. Compaction and retention appear as new snapshots listing fewer or different files.

Reading it with your own tools

Point any Iceberg-capable engine at the table’s current metadata file (version-hint.text names the latest v<N>). For example, with DuckDB:

INSTALL iceberg; LOAD iceberg;
SELECT service, count(*)
FROM iceberg_scan('s3://my-obsesc-bucket/catalog/iceberg/metadata/v<N>.metadata.json')
WHERE event_time >= TIMESTAMP '2026-01-01' AND event_time < TIMESTAMP '2026-01-02'
GROUP BY service;

Filter on event_time to benefit from partition pruning. service is not a partition column, because each file holds many services.

Storage classes

The fidelity tier is designed for three S3 storage classes only: Standard, Standard-IA and Glacier Instant Retrieval. All three return objects in milliseconds, which keeps drill-down immediate. Glacier Flexible Retrieval and Deep Archive are not supported.

With CreateRawBucket=true the stack creates the bucket with lifecycle transitions (StandardIaAfterDays, default 30; GlacierIrAfterDays, default 180) and tells the node about them. With your own bucket, you own the lifecycle rules and declare them to OBSESC. See Storage and formats.

Encryption and immutability

  • The stack’s KmsKeyArn parameter makes every raw and catalog write use SSE-KMS with your key.
  • With CreateRawBucket=true and EnableObjectLock=true, the bucket is created with S3 Object Lock. Every raw object is then locked for ObjectLockRetainDays in ObjectLockMode (governance or compliance). Object Lock can only be enabled when a bucket is created.

See Encryption for details.

Compaction

Optional compaction, off by default, tidies the raw files in your bucket into fewer, larger files. Compacted files appear in your bucket as new files and in the Iceberg table as a new snapshot. Contact us if you want it turned on.