Read the S3 Inventory report to reconcile the scrub (AWS-native item 4 activation)

Contribution Date
Contribution Project
Contribution Details
Read the S3 Inventory report to reconcile the scrub (AWS-native item 4 activation) The mechanism landed in #48 — a Capabilities flag, an InventoryEntry, and a scrub that reads a store's inventory once instead of a HEAD per placement, with a per-object fallback for stores that publish none. Nothing produced an inventory, so the path never fired. This wires it to real S3. S3Store::inventory discovers the latest /manifest.json under the configured prefix (those names sort chronologically, so the lexical maximum is the newest run), reads the manifest for its column order and data-file list, then gunzips and CSV-parses each data file into InventoryEntry rows. The parsing is a hand-rolled RFC-4180 line reader rather than a new csv crate — S3 Inventory quotes every field and the only escape is a doubled quote, so the format does not warrant the dependency tree. flate2 (already vetted transitively) does the gunzip. The reader honours the manifest's declared fileSchema rather than assuming a column order, refuses a report without a Size column with a message naming the fix, and URL-decodes keys. with_inventory_prefix flips the object_inventory capability and the prefix together, so a store can never claim inventory it has no source for. It is wired only on the AWS branch of each binary's store builder (inventory is AWS-only; a non-AWS endpoint that never fills the prefix is flagged by config.advisories rather than silently verifying nothing). S3 entries carry no comparable checksum: the inventory's ETag is an MD5, not the blake3 dam records, so comparing them would report every object corrupt. The scrub keeps its size + first-byte probe on S3 — the same floor the per-object S3 path already had, since S3 claims no server checksum. Terraform configures a daily CSV inventory on the bucket (Size + StorageClass, current versions only), with the bucket-policy grant S3 Inventory delivery requires even same-bucket same-account, and threads the prefix through DAMRS_STORAGE__INVENTORY_PREFIX. Inventory files use SSE-S3 to avoid a CMK key-policy grant for the delivery. The parsing is unit-tested (gzip, CSV quoting, URL-encoded keys, column reorder, missing-Size refusal, unknown storage class). Real-S3 delivery is a ~24h cycle, so the live path is an ops-verify follow-up, not a CI assertion.
Contribution Author
Bassam Ismail
Files count
0
Patches count
1