Fidelity tier
How OBSESC stores every event as open Parquet files with Iceberg metadata in your own S3 bucket.
The fidelity tier is the territory: every event OBSESC accepts, stored complete in your own S3 bucket as open columnar files. It is the source of truth for everything else OBSESC does, and you can read it without OBSESC.
What “complete” means
Each event keeps its timestamp, service, original body bytes as your shipper sent them, and every attribute key and value. Nothing is sampled. OBSESC adds a few derived columns alongside the originals. For example, when an event carries a host attribute, a host column is added. Native attributes are never rewritten. The one exception is redaction you configure at ingest, which runs before anything is stored. See Data flow.
Where it lives
With the CloudFormation template you can use an existing bucket (the default, CreateRawBucket=false, with RawBucketName) or let the stack create one (CreateRawBucket=true). The node writes under these prefixes:
| Prefix | Contents |
|---|---|
raw/ | Parquet data files |
catalog/ | OBSESC’s table catalog and the Iceberg metadata under catalog/iceberg/metadata/. Do not modify |
cluster/ | Reserved for multi-node deployments |
query/audit/ | Query audit trail |
holds/ | Legal-hold registry |
The node’s IAM policy only allows object and list actions on these prefixes, so it cannot see anything else you keep in the bucket. Use one OBSESC deployment per bucket.
Object layout
Data files are partitioned by the event time of their earliest event, not by the time they were written:
s3://my-obsesc-bucket/raw/<yyyy>/<mm>/<dd>/<hh>/<node_id>-<uuid>.parquet
If a single write contains events from more than one UTC day, for example a shipper replaying a buffer after an outage, OBSESC splits it into one object per day. Last Tuesday’s events land under last Tuesday’s prefix, so prefix-scoped S3 lifecycle rules and anyone browsing the bucket see data where its events belong.
File format
- Apache Parquet, compressed with zstd.
- Rows sorted by
serviceand time within each file. - Standard Parquet statistics, so external engines can skip data that doesn’t match.
Columns
| Column | Type | Notes |
|---|---|---|
timestamp_ns | int64 | Event time, nanoseconds since the Unix epoch |
source | uint8 | Ingest protocol, e.g. 0 OTLP/HTTP, 2 Elasticsearch bulk, 3 Splunk HEC |
service | string | Logical service name |
body | binary | The original message bytes |
idempotency_key | fixed binary (16) | Stable identifier used for deduplication |
attr_keys | list<string> | Attribute keys, parallel with attr_values |
attr_values | list<binary> | Attribute values |
host, env, namespace, tenant | string, nullable | Canonical dimensions, derived from the matching attribute. Null when the event has no such attribute |
event_time | timestamp (µs, UTC) | Same instant as timestamp_ns. The Iceberg partition source |
attributes | map<string, binary> | Optional, off by default. Lets external engines read attributes by key (attributes['trace_id']). Contact us to turn it on |
The schema only ever gains nullable columns. Older files read the newer columns as null, so files of different ages sit in one table.
Iceberg metadata
OBSESC keeps Apache Iceberg v2 metadata for the files in your bucket up to date:
s3://my-obsesc-bucket/catalog/iceberg/metadata/v<N>.metadata.json
s3://my-obsesc-bucket/catalog/iceberg/metadata/version-hint.text
The exported table:
- is partitioned by
day(event_time) - supports ordinary scans and time travel across snapshots
- resolves columns by name, so older files still read correctly
- does not support Iceberg incremental or changelog scans. Compaction and retention appear as new snapshots listing fewer or different files.
Reading it with your own tools
Point any Iceberg-capable engine at the table’s current metadata file (version-hint.text names the latest v<N>). For example, with DuckDB:
INSTALL iceberg; LOAD iceberg;
SELECT service, count(*)
FROM iceberg_scan('s3://my-obsesc-bucket/catalog/iceberg/metadata/v<N>.metadata.json')
WHERE event_time >= TIMESTAMP '2026-01-01' AND event_time < TIMESTAMP '2026-01-02'
GROUP BY service;
Filter on event_time to benefit from partition pruning. service is not a partition column, because each file holds many services.
Storage classes
The fidelity tier is designed for three S3 storage classes only: Standard, Standard-IA and Glacier Instant Retrieval. All three return objects in milliseconds, which keeps drill-down immediate. Glacier Flexible Retrieval and Deep Archive are not supported.
With CreateRawBucket=true the stack creates the bucket with lifecycle transitions (StandardIaAfterDays, default 30; GlacierIrAfterDays, default 180) and tells the node about them. With your own bucket, you own the lifecycle rules and declare them to OBSESC. See Storage and formats.
Encryption and immutability
- The stack’s
KmsKeyArnparameter makes every raw and catalog write use SSE-KMS with your key. - With
CreateRawBucket=trueandEnableObjectLock=true, the bucket is created with S3 Object Lock. Every raw object is then locked forObjectLockRetainDaysinObjectLockMode(governanceorcompliance). Object Lock can only be enabled when a bucket is created.
See Encryption for details.
Compaction
Optional compaction, off by default, tidies the raw files in your bucket into fewer, larger files. Compacted files appear in your bucket as new files and in the Iceberg table as a new snapshot. Contact us if you want it turned on.