Issue a bulk restore as one S3 Batch job (AWS-native item 3)

Contribution Date
Contribution Project
Contribution Details
Issue a bulk restore as one S3 Batch job (AWS-native item 3) The restore poll issued a RestoreObject per placement, which at collection scale is hundreds of round trips. A store that can batch now folds a backlog into one server-side S3 Batch job with the backend's own retries, throttling and completion report. The seam is a Capabilities::batch_operations flag and a BlobStore::bulk_restore(keys, tier, keep_for) -> BatchHandle, defaulting to Unsupported so a driver with no batch plane keeps the per-object fallback. FakeS3Store implements it (sharing begin_restore with the single restore, so the two cannot drift) and records its calls, which is the seam the pipeline test drives; the shared conformance suite gains a bulk-restore case. S3Store::bulk_restore uploads a "bucket,key" CSV manifest, then creates an S3 Batch S3InitiateRestoreObject job via aws-sdk-s3control, parsing the account from the batch role ARN so no separate account id is threaded through. with_batch_role flips the capability and the role together, the same lockstep as with_inventory_prefix. Only Standard and Bulk batch; Expedited is refused rather than silently downgraded. The poll claims a larger slice when the store can batch and, at or above a threshold, groups the claim by (tier, keep-warm) and issues one job per group; Expedited, an unknown tier, or a sub-threshold backlog fall back to the per-object loop. The job runs asynchronously and its objects are still observed becoming readable through the ordinary per-object head reconciliation, so the handle is for correlation, not the availability signal. Terraform adds the S3 Batch IAM role (assumed by batchoperations.s3, granted restore + report + CMK use on this bucket) and the app's CreateJob/PassRole grants, wired through DAMRS_STORAGE__BATCH_ROLE_ARN. SeaweedFS claims none of this and stays on the loop. Verified: workspace clippy -D warnings, fmt, terraform validate, the fake conformance suite (now exercising bulk restore), and pipeline tests proving a backlog past the threshold becomes exactly one job covering every object while a handful stays per-object. Real-S3 completion is a ~24h ops-verify follow-up.
Contribution Author
Bassam Ismail
Files count
0
Patches count
1