Docs: Storage and formats
Documentation / Concepts

Storage and formats

Every place OBSESC keeps data in your account, the format of each, and how retention and lifecycle apply.

OBSESC keeps data in two kinds of storage inside your AWS account: EBS volumes attached to the node, and one S3 bucket. This page lists what lives where, the format of each, and how retention, lifecycle and encryption apply.

At a glance

StoreWhereFormatSource of truth?Readable without OBSESC
Write-ahead logEBS, /var/lib/obsesc/walOBSESC internalOnly until stored in S3No
Navigation tierEBS, /var/lib/obsesc/summaryOBSESC proprietaryNo: rebuildable from S3No
Records of unusual behaviourEBS, /var/lib/obsesc/anomalyOBSESC proprietaryNoNo
Fidelity tier (events)S3, raw/Apache Parquet, zstdYesYes
Table catalogS3, catalog/OBSESC catalog + Apache Iceberg v2 metadataYes (the catalog)Yes, via the Iceberg metadata
ReservedS3, cluster/OBSESC internalFor multi-node deploymentsn/a
Query audit trailS3, query/audit/OBSESC internalFor auditn/a
Legal holdsS3, holds/OBSESC internalFor holdsn/a

Only one thing is irreplaceable: the Parquet files and the catalog in your bucket. Everything on EBS is either short-lived (the WAL) or can be rebuilt from S3 (the navigation tier).

EBS volumes

The CloudFormation stack creates three gp3 volumes, encrypted with your KMS key:

VolumeMountParameterDefault size
WAL/var/lib/obsesc/walWalVolumeSizeGiB100 GiB
Summary/var/lib/obsesc/summarySummaryVolumeSizeGiB500 GiB
Anomaly/var/lib/obsesc/anomalyAnomalyVolumeSizeGiB100 GiB

The node does not use instance-local NVMe storage, so volumes can be detached and moved to a replacement instance.

  • WAL. Holds accepted events until they are stored in S3, then frees the space. It is not a history store, and retention does not apply to it. Its size bounds how long the node can keep accepting data while S3 is unreachable. If too much data is waiting to reach S3, ingest returns 503 until it catches up.
  • Summary and anomaly. Much smaller than your raw data. Size them for your retention window. See Sizing.

EnableDataVolumeSnapshots=true (the default) snapshots the stack’s data volumes daily and keeps seven snapshots. The key policy of your KMS key must allow the stack’s snapshot role to use it. See Backups and recovery.

S3 bucket

One bucket per deployment. With the CloudFormation template, name an existing bucket with RawBucketName (the default, CreateRawBucket=false) or let the stack create one (CreateRawBucket=true). If the bucket belongs to another account, set RawBucketAccountId. The node names the expected bucket owner on every request to your raw data, so S3 refuses the call if the bucket belongs to a different account.

For a self-managed configuration, the equivalent settings are:

storage:
  s3_bucket: "my-obsesc-bucket"
  s3_prefix: "raw/"
  s3_sse_kms_key_arn: "arn:aws:kms:us-east-1:123456789012:key/00000000-0000-0000-0000-000000000000"
  # Only when the bucket is in another region or account:
  # s3_region: "eu-west-1"
  # assume_role_arn: "arn:aws:iam::123456789012:role/obsesc-raw-access"
  # expected_bucket_owner: "123456789012"

storage.s3_prefix namespaces the raw data files only. The catalog, audit and hold prefixes always sit at the bucket root. See Fidelity tier for the object layout and Parquet schema.

Formats

Apache Parquet (events)

Open, columnar and self-describing. Compressed with zstd. Any Parquet reader can open the files directly.

Apache Iceberg v2 (table metadata)

OBSESC writes Iceberg metadata under catalog/iceberg/metadata/. The table is partitioned by day(event_time). Engines such as Athena, Trino, Spark and DuckDB can query it in place. If an Iceberg metadata update fails, ingest carries on, and the next successful update brings external engines fully up to date.

OBSESC’s own proprietary format. Only OBSESC reads it, and it is not an interchange format. Because the navigation tier can be rebuilt from the Parquet files, you never need to export it.

Retention

Retention is off by default: OBSESC keeps everything until you set a window.

storage:
  retention_window_days: 1095          # unset = keep forever
  retention_check_interval_secs: 3600  # how often retention runs (default hourly)
  retention_grace_period_secs: 300     # delay before expired objects are deleted from S3

retention:
  min_window_days: 30                  # compliance floor; no window (global or policy) may be shorter, or the config is rejected
  policies:
    - match: { service_glob: "debug-*" }
      retention_window_days: 30
    - match: { service_glob: "payments" }
      retention_window_days: 2555

How it applies:

  • Age is measured by event time, not write time.
  • Per-service policies apply to the navigation tier and records of unusual behaviour. The most specific match wins: an exact name beats a glob, a glob with more literal characters beats a looser one, and on a tie the first declared policy wins. A service no policy matches uses storage.retention_window_days.
  • Raw Parquet files hold many services each, so a raw file is kept until the longest window that applies to any of its services has passed. To remove one service’s events sooner you need erasure, not retention. Erasure must be enabled before you need it, so contact us to enable it. See Compliance.
  • Expired raw files drop out of queries first. The S3 objects are deleted after retention_grace_period_secs, so in-flight queries can finish.
  • Legal holds override every retention policy.

Retention and S3 lifecycle

Standard-IA and Glacier Instant Retrieval bill a minimum storage duration (30 and 90 days) that starts at the transition. If retention deletes an object partway through that period, AWS bills you for the rest. With retention R, transition age T and minimum M:

  • R < T: safe. Data is deleted before it transitions.
  • T ≤ R < T + M: you pay early-deletion charges on every deleted object.
  • R ≥ T + M: safe.

OBSESC cannot read your bucket’s lifecycle rules. With CreateRawBucket=true the stack declares them for you. With your own bucket, mirror them into config:

storage:
  lifecycle:
    standard_ia_after_days: 30
    glacier_ir_after_days: 180

The node then warns at startup when your retention falls in the paying range. Contact us if you want a configuration checked before you deploy it.

Glacier Flexible Retrieval and Deep Archive are not supported for the raw tier. Objects there cannot be read back in milliseconds, which drill-down and search depend on.

Encryption

  • S3. The stack’s KmsKeyArn parameter (your customer-managed key) makes every raw and catalog write request SSE-KMS under that key. In a self-managed configuration without storage.s3_sse_kms_key_arn, the bucket’s default encryption applies.
  • EBS. The stack encrypts every data volume with the same customer-managed key.

See Encryption and IAM and permissions.

Sizing

The raw tier grows with ingest volume and retention. Size the summary and anomaly volumes for your retention window and number of services. The WAL only needs to cover your longest expected S3 interruption at peak ingest. See System requirements for sizing guidance.