S3 Credential Provider SPI: Design Notes#

This page captures why the org.apache.comet.cloud.s3.CometS3CredentialProvider SPI is shaped the way it is. The user-facing contract and operator setup live in the user guide page on S3 credential providers; this page is for maintainers and reviewers who want the design rationale.

The gap the SPI fills#

Comet’s native scan paths (object_store for raw Parquet, opendal via iceberg-rust for Iceberg) bypass Spark’s Hadoop S3A code path. That means credentials cannot flow through any of the contracts that vendors typically wire into for S3A:

  • org.apache.spark.deploy.security.cloud.CloudCredentialsProvider yields a single JWT per service name. No path argument, no AWS credential.

  • Hadoop S3A custom signers hide path-aware logic inside Signer.sign(request, credentials). The credential never leaves the signing pipeline, and the underlying secret is an HMAC key that is not present in the signed output, so running the signer against a synthesized request cannot recover it.

  • AWSCredentialsProvider.getCredentials() (AWS SDK v1) and AwsCredentialsProvider.resolveCredentials() (v2) are parameterless. They cannot vend per-path credentials.

  • Reflecting into vendor singletons would encode per-vendor class names and lifecycles in Comet and would silently break on vendor upgrades.

A Comet-specific SPI is the narrowest fit: a single Java method that takes a CometS3CredentialContext (today wrapping bucket, path, and access mode; new fields can be added without breaking vendors compiled against earlier versions) and returns CometS3Credentials.

Why config-driven activation, not META-INF/services#

An earlier iteration used ServiceLoader discovery. That was rejected because:

  • Peer SPIs in the same space (Hadoop AWSCredentialsProvider, AWS SDK v2 AwsCredentialsProvider, Iceberg AwsClientFactory, S3A custom signers) are all class-name-in-config. Vendors are already familiar with that model.

  • ServiceLoader makes activation implicit on classpath presence. A vendor JAR drifting onto a cluster could silently change S3 auth behavior. The config key makes activation explicit.

  • The activation key (fs.s3a.comet.credential.provider.class, with per-bucket override) follows the same shape as fs.s3a.bucket.<name>.aws.credentials.provider, so operators do not learn a new pattern.

Activation is modeled on parquet.crypto.factory.class (Parquet Modular Encryption KMS, see Comet #2447): the user names a single vendor class and the vendor dispatches across multiple credential backends inside that class if they need to. This mirrors how Iceberg’s DecryptionPropertiesFactory already behaves for Parquet keys.

Why (FQCN, dispatchKey, catalogProperties) keying#

Comet caches one provider instance per (FQCN, dispatchKey, catalogProperties) triple. The dispatch key is the Spark V2 catalog name on the Iceberg path and the bucket on the Parquet path.

  • Two catalogs that share one provider class (typical in multi-tenant deployments) need isolated initialize maps because their catalogProperties differ. Without dispatchKey, the second initialize would either overwrite the first or be silently skipped.

  • The bucket as dispatchKey for Parquet gives vendors per-bucket isolation when the same provider is named under several fs.s3a.bucket.<name>.comet.credential.provider.class keys.

  • catalogProperties enters the key to handle multi-tenant JVMs (Spark Connect, Thrift Server, SparkSession.newSession()) where two sessions can configure the same provider class against the same dispatchKey but with different REST endpoints, OAuth tokens, or vendor keys. Without it the second session would silently use the first session’s credentials.

  • Keying solely by FQCN would force vendors to encode multi-tenant routing in static state. The triple-key keeps each call site independent.

ensureInitialized returns a long handle that the native bridge stashes and replays on every per-request call. Routing per-request lookups by handle avoids re-sending the property bag across JNI on the hot path and unambiguously selects the right provider when the same (FQCN, dispatchKey) pair maps to multiple instances.

Why fresh construction in initialize, not probing a JVM-wide static#

A provider implementation might be tempted to probe an existing static populated elsewhere (e.g. by a Hadoop S3A signer’s registerStore callback) and reuse the credential cache that the Hadoop path uses. That fails on Comet-only executors:

  • The driver JVM hits S3AFileSystem.initialize during analysis (raw s3a:// paths) or during Hadoop catalog manifest reads (Iceberg with Hadoop catalog), so the static is populated there.

  • The driver may not hit S3AFileSystem at all under Iceberg with REST catalog plus S3FileIO, because S3FileIO calls AWS SDK directly without going through the Hadoop layer. The static stays null.

  • Executors with Comet-only reads never instantiate S3AFileSystem. The data path is object_store (raw Parquet) or opendal via iceberg-rust (Iceberg native scan). Neither touches Hadoop S3A. The static stays null on every executor.

Constructing a fresh provider from catalogProperties plus SparkEnv is the only strategy that works across all four cases. The trade-off is that on the driver (and any JVM where Hadoop S3A is also active), two credential caches now exist for the same identity: one inside the Hadoop signer’s provider, one inside the SPI implementation’s. The vendor pays for this with a small number of extra AS round-trips on cold starts and TTL boundaries. A future optional optimization could probe the static first and reuse if non-null, falling back to fresh construction otherwise.

Credential reuse, bounded by the reported expiry#

Comet’s bridge does not schedule refresh or broadcast catalog state. That stays the vendor’s responsibility:

  • Iceberg vendors get software.amazon.awssdk.utils.cache.CachedSupplier for free inside org.apache.iceberg.aws.s3.VendedCredentialsProvider.

  • Custom-STS vendors write whatever cache fits their refresh model.

  • Driver-only state is distributed via initialize’s catalogProperties (Iceberg path) or read from Hadoop conf via SparkEnv (Parquet path). Both are plan-time snapshots: Comet does not re-execute the catalog or push fresh values to running scans. Vendors that need a refreshing bearer compose with Spark’s HadoopDelegationTokenProvider, which mints and renews on the driver and propagates to executors via UserGroupInformation. The two SPIs are orthogonal: Spark covers bearer lifecycle, this SPI covers path-aware AWS credential minting.

What the bridge does do is reuse a credential for as long as the vendor’s expiry allows. Each bridge keeps its last credential with a known expirationEpochMillis and serves it until five minutes before that expiry (REFRESH_BEFORE_EXPIRY in credential_bridge.rs), then asks the provider again. A bridge makes at most one provider call at a time, and a request that waited for one shares its outcome, a credential or an error, even when the credential cannot be kept. A bridge stands for one location on a location-scoped store, so requests for different locations never wait on each other. A credential whose expiry is unknown (0), or that does not expire (Long.MAX_VALUE), is not kept, and the provider is asked again for every request that does not overlap a call in flight.

The bridge used to forward every call, which left each path with a different and partly wrong answer to expiry:

  • object_store asks for a credential on every request and has no notion of expiry, so the Parquet path called the provider once per HTTP request and used a credential however close it was to expiring. object_store also signs a request once and sends that signature again on every retry, for up to 3 minutes by default, so a retry after backoff could arrive after the session token expired, and a 403 is not retried.

  • iceberg-rust builds a new opendal operator, and with it a new reqsign signer, for every storage call, so reqsign’s own cache lasts one operation, and a scan called the provider at least once per data and delete file, even with one-hour tokens.

The reuse window is not a tuning knob. It follows the expiry the vendor reports, and a vendor that wants every request to reach it reports 0. Five minutes is the refresh-ahead of Comet’s other credential caches, the native Parquet credential chain and the IRSA web-identity provider, and it covers object_store’s retries. A vendor whose credentials can be revoked before the expiry it reports should report an earlier one.

Location-scoped credentials on the Parquet path#

object_store::CredentialProvider::get_credential receives no request path, so one AmazonS3 store presents one credential, and the process-wide object_store cache holds one store per (scheme://bucket, config_hash, hdfs_backend). A base provider therefore gets one credential per bucket, requested with the path of the first file Comet reads. A bucket whose policies differ by location (one for warehouse/sales, another for warehouse/finance) needs a store per location and something that picks among them for each request.

CometS3LocationScopedCredentialProvider supplies that. getPolicyLocations(bucket) returns every location in the bucket that has its own policy, and create_store returns a LocationScopedObjectStore for the bucket instead of a plain store:

  • It is cached and registered like any other S3 store, one per key, so later scans on the bucket share it and the cache and registry keep one store per identity. As for any store, scans that miss the cache at the same moment each build one, and the last one cached wins.

  • Each request is served by the store of the longest location covering its path, matched one segment at a time after percent-decoding, with the bucket root as an implicit location. Routing is per request, so a partition whose files span several locations reads each file with its own location’s credential.

  • A location’s store is an AmazonS3 whose bridge is bound to the location itself, so the vendor sees one stable path per location. The bridge is derived from the bucket’s bridge, sharing its provider registration, so creating it calls no ensureInitialized and loads no classes on the Tokio worker that usually creates it. It is built on first use, usually inside an async read, from an S3StoreTemplate that resolved the region when the bucket’s store was created, so building never blocks on the Tokio runtime.

  • The locations are a snapshot. A 403 is either a real denial or a location added or removed since the snapshot, and so is a failure to get a location’s credential, because a provider with no policy for a path throws rather than vending a credential S3 would reject. On either, the store fetches the locations again and retries the request once if its path now routes elsewhere. The bridge gives its failures a CredentialProviderError source, which object_store passes through to the read unchanged, so the store can tell them from other errors. Every attempt, successful or not, starts a new snapshot generation, so requests routed from the same generation share one attempt and a failed attempt fails them all instead of each calling the provider in turn. The retry budget is per request, with no state that outlives it.

The locations come from asking for the bucket’s whole list rather than which prefixes one session covers. Asking per session leaves Comet to discover the other scopes from 403s, which needs mutable per-store state and cannot route a partition that spans scopes. With the whole list up front, routing is a function of the path.

The store itself caches no credentials. Each location’s bridge reuses its own credential until shortly before its expiry, as Credential reuse, bounded by the reported expiry describes. What it keeps is the provider’s location list and a store for each location that has been read, so it grows with the vendor’s policy list, not with the paths read. The list is only fetched again after a 403 or a credential failure, so a change that causes neither is not seen until the executor builds a new store.

The dispatcher returns null for a provider that does not implement the interface without calling it, and Comet builds the same plain store as before. The Iceberg path routes by the same locations; see the next section. Operations other than reads route by path without the retry, since Comet only reads through these stores.

Location-scoped credentials on the Iceberg path#

iceberg-rust’s OpenDAL S3 storage attaches one credential loader to the operator it builds for every file, and Comet builds that loader from one reference path: the table’s metadata location for a scan, its data location for a write. A base provider therefore gets one credential per table, requested with that path. That is wrong for a table whose files span locations with different policies: data outside the table location (write.data.path, the object-storage layout, files added by add_files or migrate), a narrower policy nested under the table location, or files in another bucket.

For a location-scoped provider, storage_factory_for returns a LocationScopedS3StorageFactory (execution/operators/iceberg_location_scoped.rs) instead of OpenDalStorageFactory::S3, following the pattern of BlobHostPromotingS3StorageFactory. Its storage implements iceberg-rust’s Storage, whose every method receives the path it operates on:

  • Each call is routed by its path to the longest location covering it in that path’s bucket, using the same LocationIndex as the Parquet store (cloud/s3/policy_locations.rs), and delegated to an OpenDAL S3 storage whose loader is a bridge bound to that location. Bucket snapshots and location storages are built on first use and belong to a SharedLocations that every FileIO of the provider registration shares; see Executor FileIO cache on the Iceberg path.

  • The key a call is routed by is the key the OpenDAL storage asks S3 for: the rest of its path after {scheme}://{bucket}/, normalized as opendal normalizes every path, which trims whitespace and drops leading and empty segments. The Parquet store percent-decodes a request’s URI to get its key, but an Iceberg path is not a URI: a partition value’s escape, such as the %3A in ts=2024-01-01T00%3A00, is part of the key. Routing stops at the first segment no location can have, such as .., since no location covers the key beyond it.

  • The reference bucket’s locations are fetched when the registration’s first FileIO is built, on the thread that loads it, as create_store does on the Parquet path. A later FileIO whose reference bucket is new to the registration fetches that bucket’s on its own loading thread. Another bucket’s are fetched when a call first names it, inside an async storage call, so the JVM call runs in tokio::task::block_in_place. Bridges for other locations and buckets are derived from the reference bridge and reuse its dispatcher handle: on this path the dispatch key is the catalog name, so one registration serves every bucket a table touches.

  • new_input and new_output return files bound to this storage rather than to a location’s, so every later read of a file is routed and can retry. reader returns a FileRead that routes each range read at the current snapshot, as a fresh request is routed, and reopens the file when that snapshot routes it elsewhere. So a refresh by another request moves an open reader too, and a failed range read shares only a refresh attempted after it was routed, which keeps a burst of failures on one refresh.

  • The refresh semantics are the Parquet store’s, and PolicyLocations shares the per-generation bookkeeping between the two: a request that gets a 403, or that fails because the provider could not produce its location’s credential, fetches the bucket’s locations again, once per snapshot generation, and retries once if its path now routes elsewhere. The second trigger looks different here. reqsign’s credential chain logs a provider exception and reports no credential (see the error message fidelity caveat), so the signer fails with CredentialInvalid and the request is never sent. The storage looks for opendal’s PermissionDenied or reqsign’s CredentialInvalid in the source chain of the iceberg error. Unlike the Parquet store, which Comet only reads through, this storage also serves native Iceberg writes, so writes and deletes refresh and retry the same way. A streaming writer is the exception: bytes it already handed to a location’s writer cannot be sent again without buffering the whole file, so a failed writer fetches the locations but returns the error, and Spark’s retry of the task writes the file through the new route. delete_stream deletes a failed batch again, through the new routes, if any of its paths now routes elsewhere.

  • For an s3-compliant alias scheme, BlobHostPromotingS3Storage wraps this storage, so routing sees the promoted host as the bucket. A hostless reference path, such as blob:///bucket/..., is not covered: build_s3_access finds no bucket in it and returns S3Access::Loader(None) before it looks up the provider, so the table uses opendal’s default chain, as it does with a base provider.

  • A provider whose getPolicyLocations throws when the factory is built fails the read or write, as on the Parquet path, rather than falling back to one credential for every file.

The IRSA web-identity provider keeps its own cache#

The bridge’s reuse above is bounded by the expiry a vendor reports. The EKS/IRSA web-identity provider in native/core/src/cloud/s3/web_identity.rs is Comet’s own and keeps a fuller cache. It is not a vendor path – there is no JVM SPI involved – so the reasoning above does not apply.

It exists because the implicit IRSA credential path on the Iceberg scan mishandles STS throttling. AssumeRoleWithWebIdentity is resolved per reader thread; a concurrent startup burst throttles STS; opendal’s default reqsign chain does not retry the throttle and falls through to the EKS node instance role, which lacks bucket access, turning a transient throttle into a hard 403 (this is the reported failure). The raw-Parquet path is intentionally left alone: its AWS SDK default chain already retries and stops on a provider error rather than downgrading, and — because NativeConfig.extractObjectStoreOptions forwards Hadoop’s core-default.xml default for fs.s3a.aws.credentials.provider — a Parquet take-over gated on “no provider configured” could never engage anyway. So the take-over is Iceberg-only, and covers both reads and writes (both build their FileIO through load_file_io): when both AWS_WEB_IDENTITY_TOKEN_FILE and AWS_ROLE_ARN are set, a region is set, and the catalog configures no explicit credentials, iceberg_common.rs::build_s3_access installs WebIdentityCredentialProvider (via a CustomAwsCredentialLoader) instead of returning S3Access::Loader(None). It:

  • builds its STS client from the AWS SDK’s fully-resolved SdkConfig (aws_config::defaults(...).load()) with a raised RetryConfig, and calls AssumeRoleWithWebIdentity on it; because the client comes from the resolved config rather than a hand-assembled one, it honors region, use_fips, use_dual_stack and any profile/custom STS endpoint with no per-setting copying to keep in sync,

  • reports the credential’s real STS expiry (not an early deadline). reqsign’s signer refuses to sign within ~10s of the reported expiry and reloads within 120s of it, so an early deadline would open a signing dead zone before every refresh; instead our own cache margin (min_ttl, floored to 120s) refreshes at or before the signer’s reload point, and a still-valid credential keeps being served if a refresh is throttled,

  • only ever calls AssumeRoleWithWebIdentity – no credential chain, no IMDS/instance-role fallback – so a throttle that outlasts the retries errors instead of downgrading, and

  • caches one credential per process keyed by identity (role_arn, token_file, region) plus the resolved settings (so a catalog’s own knobs are honored regardless of init order) and single-flights refreshes, so a startup burst makes one STS call per executor rather than one per reader thread; a failed refresh is briefly remembered so a throttled burst also costs one call rather than one per reader.

Unlike the bridge, this provider owns the whole credential lifecycle (there is no vendor to delegate to), which is why caching lives here. Config keys are catalog properties under the s3. prefix (s3.comet.credential.webIdentity.*), the same route as s3.comet.credential.provider.class; see the user guide “EKS / IRSA” section. It stands aside for any credential source opendal’s default chain ranks ahead of web-identity – a bridge class, catalog static keys / client.assume-role.arn, static AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY, or a configured profile (AWS_PROFILE, or a shared credentials / config file) – and when no region is set (a web-identity STS client with no region fails opaquely, while the default chain has a no-region global fallback). Deferring on a present profile/config file also covers a profile-configured custom STS endpoint_url. So it only changes the otherwise-default behavior.

STS endpoint selection (configure_sts_endpoint) overrides the SDK’s own endpoint resolution only in the one narrow case where doing so is safe and matches reqsign, and defers to the SDK for everything else. It overrides to the global sts.amazonaws.com (signed us-east-1) only when the region is in the commercial partition, FIPS is off, dual-stack is off, no custom/profile STS endpoint_url is set, and AWS_STS_REGIONAL_ENDPOINTS is legacy or unset. That is the case the old reqsign chain hit, so an executor whose egress only reaches the global endpoint keeps working after the take-over. Every other combination keeps the SDK’s regional resolution: FIPS (AWS_USE_FIPS_ENDPOINT / profile) uses the regional FIPS endpoint because there is no global FIPS STS endpoint (an accompanying AWS_STS_REGIONAL_ENDPOINTS=legacy is ignored with a one-time warning); dual-stack (AWS_USE_DUALSTACK_ENDPOINT) uses the dual-stack regional endpoint, because the SDK rejects a custom endpoint combined with dual-stack before it sends anything; a custom/VPC/profile endpoint_url is honored rather than silently replaced; AWS_STS_REGIONAL_ENDPOINTS=regional stays regional; and the China, GovCloud, ISO and EUSC partitions (and any region whose prefix Comet does not recognize) keep their own regional hosts (none of them has a sts.amazonaws.com global endpoint). The commercial-partition check is an allowlist of known commercial region prefixes, so an unrecognized prefix from a future partition defers to the SDK rather than being sent to the global endpoint.

Path-specific behavior#

object_store::CredentialProvider and reqsign_core::ProvideCredential differ in what they consume:

Concern

Parquet (object_store)

Iceberg (opendal via reqsign-core)

Trait method

get_credential() -> AwsCredential

provide_credential(...) -> Option<IcebergAwsCredential>

Returns expiry?

No (only key/secret/token)

Yes (expires_in: Option<Timestamp>)

Comet-side TTL wrapper?

The bridge’s own reuse, until 5 minutes before the expiry

The bridge’s own reuse, until 5 minutes before the expiry

When SPI is called

When the bridge’s credential is due, or has no expiry

When the bridge’s credential is due, or has no expiry

Vendor returns 0 expiry

Not reused; the SPI is called per request

Not reused; the bridge tells reqsign it expires in 5 minutes

The 5-minute fallback on the Iceberg path bounds how long reqsign reuses a credential with no reported expiry within one long-lived reader or writer. It can outlast a short-lived token, which is why a vendor that knows the expiry should report it. It is intentionally not a configuration knob. On both paths a value of Long.MAX_VALUE means no expiry, and a value before 2000, almost always seconds sent as milliseconds, is treated as unknown with a warning.

Property-bag handling on the Iceberg path#

The full unfiltered FileIO property bag crosses JNI as catalog_properties. The storage-prefix filter (s3./gcs./adls./client.) is applied native-side in iceberg_common.rs::build_file_io immediately before FileIOBuilder.with_prop. This means the bridge sees credentials.uri, OAuth tokens, and any vendor-custom keys with no parallel field on the operator and no driver-side broadcast. Vendors set their own keys on the catalog config and read them back inside initialize(Map).

IcebergScanExec derives a redacting Debug, and the FileIO cache key’s Debug omits the property bag, so plan dumps and tracing do not leak it.

Executor FileIO cache on the Iceberg path#

load_file_io in iceberg_common.rs serves clones from a per-executor cache instead of building a FileIO per task. An entry is keyed by access mode, catalog name, the full reference path and the whole catalog property bag, and it holds the FileIO together with the bridge it was built with, so a bridge and its dispatcher handle live as long as the entry. The cache holds 64 entries, evicts the least recently used one, and is drained with the Tokio runtime. The reference path has to stay in the key: the bridge is constructed with the bucket and url.path() of that location and the JVM provider is called with exactly that pair, so dropping the path from the key would hand one table’s bridge to another table in the same bucket. A location-scoped provider’s bucket snapshots and location storages are not kept per FileIO, since the metadata location in the key changes with every commit and each new one would ask the provider again. build_s3_access keeps them per provider registration and access mode instead, keyed like the cache but without the reference path, so on an executor the commits of a table, and the tables whose catalog properties match, share one set. The key is the whole property bag, and CometScanRule adds each table’s FileIO properties to it, so a REST catalog that vends per-table properties, such as per-table credentials, gets one set per table. Sharing across different bags would cross provider instances, since ensureInitialized registers one per bag. Builds for one registration run one at a time, so concurrent tasks on a new snapshot wait for the first rather than each asking the provider. The registry holds weak references: the state lives as long as a cached FileIO uses it, and a registration whose FileIOs are all evicted asks again next time. Two builds are never cached: memory:///, whose namespace the write path uses per task, and a read whose configured provider failed to initialise and fell back to the default chain, so the next task retries the provider instead of inheriting the fallback.

The cache holds the FileIO and its bridge. The bridge reuses its credential until shortly before its expiry, as Credential reuse, bounded by the reported expiry describes, so provide_credential reaches the vendor about once per credential rather than once per storage call.

Property-bag handling on the Parquet path#

The Parquet path forwards the full fs.s3a.* config subset as catalog_properties, so an SPI provider sees the same fs.s3a.* config Spark would. forward_catalog_properties in native/core/src/parquet/objectstore/s3.rs keeps every fs.s3a.* key, including the static credentials (*.access.key, *.secret.key, *.session.token). Forwarding them is deliberate: AWSCredentialProviderList skips a provider that throws NoAwsCredentialsException and moves to the next entry, so stripping the static keys would let a chain such as SimpleAWSCredentialsProvider,customProvider silently resolve through a different provider than Spark, reading data as a different principal. These keys already cross JNI for the non-adapter native path (build_credential_provider reads them), so this matches existing behavior rather than widening exposure. Because NativeConfig.extractObjectStoreOptions only collects fs.s3a.* (plus the fs.comet.* scheme keys), config a provider reads that is not under fs.s3a.* – notably hadoop.security.credential.provider.path – never crosses JNI. AdapterSupport.toConfiguration therefore seeds the Hadoop Configuration from the executor’s own Spark-derived Hadoop conf (spark.hadoop.*) and overlays the forwarded fs.s3a.* on top, so those keys are present.

Built-in adapters#

Comet ships two reference SPI implementations under org.apache.comet.cloud.s3, so standard provider classes that the native Rust list does not match work with a config change instead of bespoke code:

  • HadoopS3ACredentialProviderAdapter delegates to Hadoop S3A’s own provider construction (S3AUtils.createAWSCredentialProviderSet on Hadoop 3.3.x, CredentialProviderListFactory.createAWSCredentialProviderList on 3.4+).

  • AwsSdkCredentialProviderAdapter wraps a raw AWS SDK provider named in fs.s3a.comet.credential.adapter.class.

Each has a spark-3.x (SDK v1) and a spark-4.x (SDK v2) body under the same FQCN, selected by the shims.majorVerSrc source set, so each Comet build compiles against exactly the one AWS SDK its Hadoop line ships. The SDK and hadoop-aws are provided scope only (see the hadoop-aws.version property in the root pom.xml), so Comet does not bundle a second copy.

Returns or throws, not a fall-through value#

The SPI returns a CometS3Credentials or throws. There is no sentinel “I do not know” return. Vendors that are only authoritative for some paths resolve the default AWS chain themselves for the rest and return the result. This matches the contract on every other AWS credential SPI in the JVM ecosystem (AWS SDK v1/v2, Hadoop S3A, Iceberg VendedCredentialsProvider).

Lifecycle: AutoCloseable plus a JVM shutdown hook#

CometS3CredentialProvider extends AutoCloseable with a default no-op close(). The dispatcher installs a JVM shutdown hook that iterates every cached instance and calls close(), swallowing per-provider exceptions so a slow or buggy vendor cannot block other providers from cleaning up. Stateless providers ignore this entirely; vendors that hold long-lived resources (HTTP clients, scheduled-refresh executors, STS connection pools) override close() to release them. Shutdown hooks are best-effort, so a SIGKILL or abrupt JVM termination skips them; vendors must not rely on close() for correctness, only for resource hygiene.

Iceberg path: error message fidelity caveat#

When the bridge is wired into iceberg-rust, the outer reqsign-core::ProvideCredentialChain currently swallows thrown exceptions into “no credential” before the request reaches opendal. The credential is still not issued and the request still fails, but the message is degraded to reqsign’s generic failed to load signing credential, and the request is never sent. No Comet change fixes this; it is resolved upstream when iceberg-rust stops wrapping custom loaders in its outer chain or moves its S3 backend to object_store.