S3 Credential Providers#
Comet’s native S3 readers and native Iceberg writer normally fetch credentials from the standard AWS credential chain (static keys, instance profiles, environment variables, etc.). Some clusters use a vendor-managed mechanism instead, where credentials are issued per request based on a JWT or per S3 path. For those clusters, Comet supports loading a vendor-supplied bridge class that routes every native credential request through the vendor’s Java code.
Do I need this?#
You don’t, if any of the following describe your cluster:
You use static AWS credentials (
fs.s3a.access.key/fs.s3a.secret.key).You use EC2 instance profiles, EKS pod identities, ECS task roles, or environment variables.
Your S3 access works in Spark today via the default AWS credential chain.
You probably do, if any of these are true:
You have a Hadoop S3A custom signer configured (
fs.s3a.custom.signers=...).You have a Spark
CloudCredentialsProviderthat issues a JWT for a vendor STS service.You have a custom Iceberg
client.factorythat injects a configured S3 client.Spark queries against your S3 paths work, but the same queries with Comet enabled fail with 403.
Built-in adapters#
If a native Parquet scan fails with Unsupported credential provider: <class> (for example com.amazonaws.auth.DefaultAWSCredentialsProviderChain), the class you named in fs.s3a.aws.credentials.provider is one that plain Spark/Hadoop accepts but Comet’s native reader does not reimplement. Comet ships two built-in CometS3CredentialProvider adapters that fix this with a one-line config change; you leave your existing fs.s3a.aws.credentials.provider untouched.
These adapters cover the Parquet native scan path only. Enabling one is opt-in: naming it is what activates it. Note the native side forwards the fs.s3a.* config to initialize() for any provider class named on the Parquet path, not just these two adapters: a vendor CometS3CredentialProvider that received an empty map in Comet 1.0 now receives the fs.s3a.* subset (including static keys), and one cached instance per distinct fs.s3a.* config rather than one per bucket. This is additive, but a provider that logs the map or treats an empty map as “the Parquet path” should be aware of it.
The adapters and the AWS SDK are loaded through the class loader that loaded Comet, so hadoop-aws and the matching AWS SDK must be visible from there — put them on the same classpath as Comet (spark.executor.extraClassPath / spark.driver.extraClassPath, or $SPARK_HOME/jars), not only via --packages. If they are only on the user-jar loader, credential resolution fails at planning with NoClassDefFoundError before the adapter can report anything useful.
HadoopS3ACredentialProviderAdapter (recommended)#
Delegates to Hadoop S3A’s own provider construction, so it accepts everything the fs.s3a.aws.credentials.provider chain accepts (the default chain, web-identity, assumed-role, per-bucket config). This is the general answer for the failure above. It does not cover fs.s3a.custom.signers — the native reader signs SigV4 itself with whatever the chain returns and never invokes a custom signer, so a signer’s identity would not be applied; a custom-signer setup needs a vendor bridge (see “Do I need this?”). And it refuses fs.s3a.delegation.token.binding: Spark would use the delegation-token provider and bypass the chain, so rather than resolve a different identity the adapter fails with a message naming that key.
spark.hadoop.fs.s3a.comet.credential.provider.class=org.apache.comet.cloud.s3.HadoopS3ACredentialProviderAdapter
# leave your existing config as-is, for example:
spark.hadoop.fs.s3a.aws.credentials.provider=com.amazonaws.auth.DefaultAWSCredentialsProviderChain
It needs no extra config: it reads the standard fs.s3a.aws.credentials.provider (and the per-bucket fs.s3a.bucket.<bucket>.aws.credentials.provider) itself, and Comet forwards the full fs.s3a.* config to it, so a provider chain (including static keys and assumed-role) resolves the same way it would under Spark.
Anonymous access is the one exception. A bucket whose chain resolves to AnonymousAWSCredentialsProvider (common for public datasets) is not supported: the SPI has no way to express “anonymous”, so the adapter fails with an error naming the bucket rather than reading it unsigned. To read such a bucket, opt it out of the adapter with an empty per-bucket class so the native reader accesses it directly:
spark.hadoop.fs.s3a.bucket.<public-bucket>.comet.credential.provider.class=
AwsSdkCredentialProviderAdapter#
Wraps a single raw AWS SDK credential-provider class that is not registered through S3A. Name the delegate in a separate key:
spark.hadoop.fs.s3a.comet.credential.provider.class=org.apache.comet.cloud.s3.AwsSdkCredentialProviderAdapter
spark.hadoop.fs.s3a.comet.credential.adapter.class=<FQCN of your credential provider>
# per-bucket variant:
spark.hadoop.fs.s3a.bucket.<bucket>.comet.credential.adapter.class=<FQCN>
Which one, and which Spark version#
Use HadoopS3ACredentialProviderAdapter unless you have a plain SDK provider not wired through S3A. Both class names are the same on every Comet build; each build automatically uses the AWS SDK its Hadoop line ships (v1 on the Spark 3.4/3.5 builds, v2 on 4.0+), so you configure one name and get the right implementation.
Enabling a bridge#
A bridge is activated by naming the vendor’s class in a Spark config. Putting a JAR on the classpath alone has no effect; the config key must be set.
For raw Parquet (the object_store path), set the Hadoop S3A config:
spark.hadoop.fs.s3a.comet.credential.provider.class=com.vendor.MyCometCredentialProvider
A per-bucket override is supported and follows the same shape as fs.s3a.bucket.<name>.aws.credentials.provider:
spark.hadoop.fs.s3a.bucket.<bucket>.comet.credential.provider.class=com.vendor.MyCometCredentialProvider
For Iceberg (the opendal path), set the per-catalog property under the s3. namespace Iceberg already uses for its S3 settings:
spark.sql.catalog.<catalog>.s3.comet.credential.provider.class=com.vendor.MyCometCredentialProvider
On the Iceberg path Comet builds a table’s bridge from the bucket in the table’s metadata location, or its data location for a write. A location with no host, such as the blob:///bucket/... form of an s3-compliant alias scheme, names no bucket there, so Comet does not use the configured provider for that table at all: opendal’s default credential chain signs every request, and a location-scoped provider’s locations are not used.
Add the vendor JAR to your Spark executor classpath:
spark-submit --jars vendor-comet-bridge.jar ...
OSS Comet ships no vendor-specific bridges. Get one from the same vendor that supplies your Hadoop S3A signer or Iceberg client factory. If they do not yet provide one, send them to the “Writing a bridge” section below.
Wire encryption#
Catalog configuration you set under spark.sql.catalog.<catalog>.* (which may include endpoint URLs, OAuth tokens, and other vendor keys your provider reads) is sent from the driver to executors over Spark’s Netty RPC channel so the provider can run there. Live AWS credentials fetched by the provider stay on the executor that fetched them.
Spark’s RPC channel is plaintext by default and offers two opt-in encryption mechanisms, both off by default:
spark.network.crypto.enabledfor AES-based RPC encryption keyed off the auth shared secret.spark.ssl.rpc.enabledfor TLS on the same channel (Spark 3.4+).
The same defaults apply to Hadoop delegation tokens, CloudCredentialsProvider JWTs, and any other secrets Spark already ships driver-to-executor, so this is a deployment-wide call rather than something specific to this SPI. See Spark’s security guide for the full set of knobs.
Verification#
With the config set and the JAR on the classpath, executor logs show on first S3 access:
Info level:
Instantiated CometS3CredentialProvider <fully.qualified.VendorClassName>Debug level:
Fetching credentials via <class> (dispatchKey=<key>) for bucket=... path=... mode=...
Without the config set, no credential-related log lines appear at startup; native readers use the default AWS credential chain.
Troubleshooting#
Generic S3 error: Unsupported credential provider: <class> (native Parquet scan). The class in fs.s3a.aws.credentials.provider is one Hadoop S3A accepts but Comet’s native reader does not reimplement. Name HadoopS3ACredentialProviderAdapter as the Comet provider class (see Built-in adapters) and leave your existing config alone.
CometS3CredentialProvider class not found: <name>. The class named in the config is not on the executor classpath. Re-check --jars / spark.jars. On YARN or Kubernetes, confirm the JAR actually reached the executor and not only the driver.
<class> does not implement org.apache.comet.cloud.s3.CometS3CredentialProvider. The configured class exists but does not implement the SPI. Double-check the FQCN against the vendor’s documentation.
<class> must declare a public no-arg constructor. Vendor classes are instantiated reflectively with Class.forName(name).getDeclaredConstructor().newInstance(). A non-default constructor is not supported; ask the vendor to expose a no-arg form that reads any state it needs from initialize(Map) or environment.
CometS3CredentialProvider <class> (dispatchKey=...) was not initialized. initialize(Map) was not called before a credential request. Comet should always invoke ensureInitialized synchronously when it builds the bridge at plan time, so this indicates the bridge skipped the init call (a Comet bug) or the vendor’s initialize threw and the bridge fell through to the default chain.
403 AccessDenied with the bridge configured. The provider returned credentials but they were rejected by S3. Most often a region mismatch (see Iceberg section below) or expired session token; enable debug logging on the vendor’s class to confirm what it returned.
Credentials silently going stale during long-running jobs. When a vendor returns expirationEpochMillis=0, the bridge substitutes a 5-minute expiry before handing the credential to opendal, so opendal’s cache cannot hold a stale credential indefinitely. Returning a real expiry is preferred; the 5-minute fallback is a safety net, not a knob.
EKS / IRSA: STS throttling protection (native Iceberg reads and writes)#
This is automatic; there is nothing to configure to get the protection, and it does not involve a bridge class. It applies to native Iceberg reads and writes (both build their FileIO through the same path). The raw-Parquet path is unaffected: it uses the AWS SDK default chain, which already retries and stops on a provider error rather than downgrading.
On EKS with IAM Roles for Service Accounts (IRSA), executors assume the application role by calling STS AssumeRoleWithWebIdentity. Under a large concurrent startup burst (many executors times many cores, all fetching credentials at once), STS can throttle that call. On the Iceberg path, opendal’s default credential chain does not retry the throttle and falls through to the EKS node instance role, which usually lacks bucket access, so every native read then fails with a hard 403 AccessDenied even though the throttle was transient.
When Comet detects IRSA (both AWS_WEB_IDENTITY_TOKEN_FILE and AWS_ROLE_ARN are set), a region is set, and the catalog configures no explicit credentials, the native Iceberg reader resolves web-identity credentials itself instead of using opendal’s default chain. It builds the STS client from the AWS SDK’s fully-resolved config, so region, AWS_USE_FIPS_ENDPOINT, AWS_USE_DUALSTACK_ENDPOINT, and any profile/custom STS endpoint are honored. The provider:
retries the throttled
AssumeRoleWithWebIdentitycall with the AWS SDK’s exponential backoff and jitter (a throttle that outlasts the retries becomes an error that Spark’s task retry then picks up),never falls back to the node instance role, and
caches one assumed-role credential per executor process, shared across all reader threads and scans, so a startup burst makes one STS call per executor rather than one per thread. If a refresh is throttled while the current credential is still valid, it keeps serving that credential.
STS endpoint selection uses the global sts.amazonaws.com endpoint only in the plain commercial case, and leaves everything else to the AWS SDK. The global endpoint is used when your region is in the standard commercial partition, FIPS and dual-stack are both off, no custom STS endpoint is set, and AWS_STS_REGIONAL_ENDPOINTS is legacy or unset. That matches the previous behavior, so a network that only reaches the global endpoint keeps working. In every other case the SDK’s own regional endpoint is used: FIPS (AWS_USE_FIPS_ENDPOINT) always uses the regional FIPS endpoint since there is no global FIPS STS endpoint (an accompanying AWS_STS_REGIONAL_ENDPOINTS=legacy is ignored and a warning is logged), dual-stack (AWS_USE_DUALSTACK_ENDPOINT) uses the dual-stack regional endpoint, a custom STS endpoint (AWS_ENDPOINT_URL_STS or a profile) is honored, AWS_STS_REGIONAL_ENDPOINTS=regional stays regional, and the China, GovCloud, ISO and EUSC regions (and any region Comet does not recognize as commercial) use their own regional endpoints.
It stands aside whenever a higher-precedence credential source is configured – a Comet bridge class or catalog static keys / client.assume-role.arn, static credentials in the environment (AWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY), or a configured profile (AWS_PROFILE, or a shared credentials / config file such as ~/.aws/credentials or ~/.aws/config) – and when no region is set (AWS_REGION / AWS_DEFAULT_REGION). These all rank ahead of web-identity in opendal’s default chain or need its no-region fallback, so the take-over only changes the otherwise-default behavior and never switches away from an identity – or a profile-configured STS endpoint – you set explicitly. On an EKS/IRSA pod none of these apply, so the take-over still engages there.
Tuning is rarely needed. Set these as catalog properties under the s3. prefix (the same namespace as the credential-provider SPI key):
Property (per Iceberg catalog) |
Default |
Meaning |
|---|---|---|
|
|
Set |
|
|
STS attempts before the assume-role call is treated as failed. |
|
|
Refresh this many seconds before expiry (floored to 120s). |
|
|
Random slack added on top of |
For example, to opt out of the take-over or raise the retry count for one catalog:
spark.sql.catalog.<catalog>.s3.comet.credential.webIdentity.enabled=false
spark.sql.catalog.<catalog>.s3.comet.credential.webIdentity.maxAttempts=8
Use the s3. prefix. Comet forwards the unfiltered catalog property bag, so a bare, unprefixed key does technically reach the native reader, but the settings are looked up under s3. (matching the credential-provider SPI key), so only the prefixed spelling takes effect.
Iceberg: explicit S3 region required#
iceberg-storage-opendal does not auto-detect a bucket’s region, with or without the bridge. When neither the catalog (s3.region or client.region) nor the executor environment (AWS_REGION / AWS_DEFAULT_REGION) supplies a region, Comet uses us-east-1. That suits most non-AWS S3-compatible services but fails for AWS buckets in other regions, so set the region (and the endpoint for non-AWS) explicitly on the Spark catalog:
spark.sql.catalog.<catalog>.s3.region = us-east-1
spark.sql.catalog.<catalog>.s3.endpoint = https://... (non-AWS only)
spark.sql.catalog.<catalog>.s3.path-style-access = true (path-style endpoints only)
Writing a bridge#
Comet’s native scan paths (object_store for raw Parquet, opendal via iceberg-rust for Iceberg) bypass Hadoop S3A entirely. The standard AWSCredentialsProvider.getCredentials() has no path argument, so vendors that issue per-path STS credentials cannot expose them through it. The CometS3CredentialProvider SPI fills that gap.
Implement org.apache.comet.cloud.s3.CometS3CredentialProvider:
package org.apache.comet.cloud.s3;
public interface CometS3CredentialProvider extends AutoCloseable {
/** Called once per (FQCN, dispatchKey, catalogProperties) before any per-request call. Optional. */
default void initialize(java.util.Map<String, String> catalogProperties) {}
CometS3Credentials getCredentialsForPath(CometS3CredentialContext context) throws Exception;
/** Invoked from the dispatcher's JVM shutdown hook. Default is a no-op. */
@Override default void close() throws Exception {}
}
The class must have a public no-arg constructor. getCredentialsForPath may be invoked concurrently from many native tokio worker threads; the implementation must be thread-safe. The supplied CometS3CredentialContext exposes getBucket(), getPath(), and getMode(); future Comet releases may add accessors here without changing the method signature, so vendors compiled against today’s API stay binary-compatible.
Lifecycle#
Comet keys provider instances by (FQCN, dispatchKey, catalogProperties). The dispatch key is the Spark V2 catalog name on the Iceberg path and the S3 bucket name on the Parquet path. The first time a given key is seen on an executor, Comet reflects the class, calls initialize(Map) exactly once, and caches the instance for the JVM lifetime. Two catalogs sharing one provider FQCN therefore get isolated instances with their own initialize maps. Including catalogProperties in the key matters in multi-tenant JVMs (Spark Connect, Thrift Server, SparkSession.newSession()) where two sessions can otherwise collide on the same (FQCN, dispatchKey) and have the second session silently use the first session’s credentials.
initialize should be cheap and non-blocking. Defer real credential fetches (REST round-trips, STS calls) to the first getCredentialsForPath invocation. On the Iceberg path the supplied catalogProperties carries the unfiltered FileIO bag, including REST-vended fields like credentials.uri, OAuth tokens, and any vendor-custom keys you set on the catalog config. On the Parquet path it carries the fs.s3a.* config subset (including static keys such as fs.s3a.access.key); in Comet 1.0 this map was empty on the Parquet path. The map may contain secrets, so do not log it.
close() is invoked from a JVM shutdown hook installed by the dispatcher. The default no-op is fine for stateless providers. Override it to release HTTP clients, scheduled-refresh executors, or STS connection pools. Shutdown hooks are best-effort: a SIGKILL or abrupt JVM termination skips them, so do not depend on close() for correctness.
Caching, refresh, and distribution#
Comet does not broadcast catalog state or schedule refresh, and it keeps a credential only as long as the expiry you report allows (see below). Vendors decide:
Whether to cache credentials and for how long. Iceberg vendors get
software.amazon.awssdk.utils.cache.CachedSupplierfor free insideVendedCredentialsProvider; vendors with custom STS write whatever cache fits.When to refresh: proactive timer, on-demand at expiry, on
403retry, etc.How to distribute and refresh state.
catalogPropertiesis a plan-time snapshot. Comet does not re-execute the catalog or push fresh values to running scans, so any state that must refresh during a scan is the vendor’s responsibility. Options: call back to a vendor service from the executor on eachgetCredentialsForPath, run your own Spark broadcast inside the class, or compose with Spark’sHadoopDelegationTokenProviderwhen the underlying bearer needs scheduled driver-side renewal.
When using HadoopDelegationTokenProvider, Spark mints and renews the token on the driver and propagates it to executors via UserGroupInformation. The provider reads the current value inside getCredentialsForPath and exchanges it for AWS credentials:
@Override
public CometS3Credentials getCredentialsForPath(CometS3CredentialContext ctx) throws Exception {
Token<?> token = UserGroupInformation.getCurrentUser()
.getCredentials()
.getToken(new Text("vendor-service"));
return vendorSts.exchangeForS3Credentials(token, ctx.getBucket(), ctx.getPath(), ctx.getMode());
}
Spark delegation token propagation is supported on YARN and Kubernetes only. Standalone deployments need a different refresh path, typically a vendor-side service callback authenticated by long-lived state in catalogProperties or Hadoop conf.
Publish a real expirationEpochMillis when you have one. On both paths Comet reuses a credential until five minutes before that expiry and then asks you again. Comet makes at most one call at a time for each location (each bucket, for a provider without locations), and requests that arrive while it is in flight share its answer, even one Comet cannot keep. On the Iceberg path, the reuse and the sharing happen within a task, so each task asks you at least once, except that a location-scoped provider’s credentials are shared by the tasks that run at the same time on an executor. 0 means unknown: Comet does not keep the credential and asks you again for every request that does not overlap a call already in flight, and on the Iceberg path it assumes the credential lasts five minutes. Long.MAX_VALUE means the credential does not expire, and Comet does not keep it either. A value before 2000, almost always seconds sent as milliseconds, is treated as unknown, with a warning. If a credential can be revoked before the expiry you report, report an earlier one, or 0.
On the Parquet/object_store path, object_store signs a request once and sends the same signature on every retry, for up to 3 minutes by default, and it does not retry a 403. The five minutes Comet leaves before an expiry cover those retries. A credential you return with less time left than that is used for the request that asked for it and not kept.
The built-in adapters report an expiry when the AWS SDK exposes one, which it does on the Spark 4.x builds (SDK v2). On the Spark 3.4 and 3.5 builds (SDK v1) they cannot, and report 0.
Returned fields#
Field |
Notes |
|---|---|
|
Required. |
|
Required. |
|
|
|
When the credential stops working. Comet reuses it until 5 minutes before. |
Provide a real expirationEpochMillis whenever you have one. Without it, Comet asks for a credential on every request on the Parquet path, and on every storage call on the Iceberg path, apart from requests that overlap a call already in flight.
Returns or throws#
The SPI follows the same shape as the other JVM AWS-credential SPIs (AWS SDK v1/v2, Hadoop S3A, Iceberg VendedCredentialsProvider): return credentials or throw. There is no “fall-through” return value.
If your provider is authoritative only for some paths, resolve the default AWS chain yourself for the rest:
private final DefaultCredentialsProvider defaultChain = DefaultCredentialsProvider.create();
@Override
public CometS3Credentials getCredentialsForPath(CometS3CredentialContext ctx) throws Exception {
if (handlesPath(ctx.getBucket(), ctx.getPath())) {
return mintFromMyVendorService(ctx.getBucket(), ctx.getPath(), ctx.getMode());
}
AwsCredentials c = defaultChain.resolveCredentials();
String token = null;
long expiresAt = 0L;
if (c instanceof AwsSessionCredentials) {
AwsSessionCredentials session = (AwsSessionCredentials) c;
token = session.sessionToken();
expiresAt = session.expirationTime().map(Instant::toEpochMilli).orElse(0L);
}
return new CometS3Credentials(c.accessKeyId(), c.secretAccessKey(), token, expiresAt);
}
Credentials per location#
On the Parquet path, a CometS3CredentialProvider gets one credential per bucket: Comet requests it with the path of the first file it reads from the bucket and uses it for every file there. On the Iceberg path it gets one per table: Comet requests it with the path of the table’s metadata file (its data location, for a write) and uses it for every file the table reads or writes. If your policies differ by location within a bucket, for example one policy for warehouse/sales and another for warehouse/finance, implement CometS3LocationScopedCredentialProvider and tell Comet where those locations are:
public final class MyLocationProvider implements CometS3LocationScopedCredentialProvider {
@Override
public List<String> getPolicyLocations(String bucket) throws Exception {
return policyService.locationsWithPolicies(bucket); // ["warehouse/sales", "warehouse/finance"]
}
@Override
public CometS3Credentials getCredentialsForPath(CometS3CredentialContext ctx) throws Exception {
// ctx.getPath() is a returned location with a leading slash, or "/" for the bucket root.
return sessionForLocation(ctx.getBucket(), ctx.getPath(), ctx.getMode());
}
}
Comet serves each request with the credential of the longest location that covers its path. A location covers a path when the path is the location itself or lies below it, compared one /-separated segment at a time, so warehouse/sales covers warehouse/sales/part-0.parquet but not warehouse/sales_eu/part-0.parquet. The bucket root covers every path that no returned location covers, and an empty list serves the whole bucket with the root’s credential. Write locations the way CometS3CredentialContext.getPath() writes paths: percent-encoded, without the scheme or bucket name. A literal % must be written as %25; other characters may be left unencoded, and a leading or trailing / is optional. When several locations decode to the same path, Comet keeps the first. A literal % is common on the Iceberg path, where Comet compares a location with each file’s key as Iceberg wrote it. Iceberg escapes partition values, so a location for the partition directory ts=2024-01-01T00%3A00 is written ts=2024-01-01T00%253A00.
Comet requests a location’s credential by calling getCredentialsForPath with the location as the path, as you returned it but with a leading slash. Every request under a location shares that credential, so it must authorize every path the location is the longest match for, and your cache can key on the location. Locations apply to Comet’s native Parquet reads and to its native Iceberg reads and writes. On the Iceberg path each data and delete file is routed by its own path, in its own bucket, so a table whose files span several locations or buckets gets each file’s credential right. A table whose metadata location has no host is the exception; see Enabling a bridge.
When Comet asks. Comet calls getPolicyLocations when it creates the store for a bucket on an executor and keeps the answer for later reads of that bucket with the same S3 configuration. On the Iceberg path it asks for the bucket of a table’s metadata or data location, and again when a table or file in another bucket is first used. Tables share the answers when their catalog properties and access mode match, but Comet keeps them only while a task that uses them is running on the executor, so a task asks again unless another task that uses them is already running there. Comet adds each table’s FileIO properties to its catalog properties, so a REST catalog that vends per-table properties, such as per-table credentials, gets its own getPolicyLocations calls for each table. Reads that start at the same moment may each create a store and call it. If a read then fails with 403, or because getCredentialsForPath threw for the location Comet sent it to, Comet asks again, once for all the reads that failed on the same answer, and retries each read once if its path now falls under a different location. On the Iceberg path writes and deletes do the same, except that a file being streamed is not sent again: its task fails, and Spark’s retry of the task writes it with the new answer. So a location added while a job runs is picked up even when you vend no credential for the bucket root. A location you drop stops being used once Comet next asks for its credential and you refuse it. Comet reuses a location’s credential until five minutes before the expiry you reported for it, so a dropped location stays in use until then, or until the next request if you reported 0. A location added or removed without a read failing on it is not seen until the executor creates a new store. Make getPolicyLocations thread-safe and independent of where it runs; it may be called on the driver or on executors.
Failures. If getPolicyLocations throws or returns null, or returns a location that is null or invalid, the read fails. A location is invalid if, once decoded, it is not valid UTF-8 or has a segment that is empty, ., .., or contains a control character, so a URI such as s3://bucket/a is invalid too. Comet does not fall back to a broader credential.
Backward compatibility. Providers that implement only CometS3CredentialProvider are unaffected. Comet calls nothing new on them and keeps one credential per bucket on the Parquet path and one per table on the Iceberg path, as before.
Composing multiple credential backends#
A single configured provider class is the dispatcher. If a vendor needs to route across several credential backends (per bucket, per path prefix, per tenant), the dispatch lives inside the vendor’s class:
public final class MyCometCredentialProvider implements CometS3CredentialProvider {
private final ProdVendor prod = ...;
private final StsVendor sts = ...;
private final DefaultVendor fallback = ...;
@Override
public CometS3Credentials getCredentialsForPath(CometS3CredentialContext ctx) throws Exception {
if (ctx.getBucket().startsWith("data-prod-")) return prod.fetch(ctx);
if (ctx.getBucket().equals("partner-shared")) return sts.assumeRole(ctx);
return fallback.fetch(ctx.getBucket(), ctx.getPath());
}
}
Per-bucket Hadoop overrides (fs.s3a.bucket.<name>.comet.credential.provider.class) are also available if you prefer to ship multiple vendor classes and pick by bucket in config rather than in code.
For Iceberg deployments where two catalogs share one provider class but need isolated state, configure the same FQCN on both catalogs and read your discriminator from initialize’s catalogProperties. Each catalog gets its own provider instance because Comet keys by (FQCN, catalogName, catalogProperties):
public final class MyMultiTenantProvider implements CometS3CredentialProvider {
private volatile String tenantId;
@Override
public void initialize(Map<String, String> catalogProperties) {
this.tenantId = catalogProperties.get("vendor.tenant-id");
}
@Override
public CometS3Credentials getCredentialsForPath(CometS3CredentialContext ctx) {
return mintForTenant(tenantId, ctx.getBucket(), ctx.getPath(), ctx.getMode());
}
}
Reference implementation: Iceberg REST vended credentials#
For Iceberg REST catalogs that vend AWS credentials (LoadTableResponse.credentials), the canonical implementation wraps Iceberg’s existing VendedCredentialsProvider:
public final class IcebergRESTVendedS3Provider implements CometS3CredentialProvider {
private volatile VendedCredentialsProvider provider;
@Override
public void initialize(Map<String, String> catalogProperties) {
this.provider = VendedCredentialsProvider.create(catalogProperties);
}
@Override
public CometS3Credentials getCredentialsForPath(CometS3CredentialContext ctx) {
AwsCredentials c = provider.resolveCredentials();
String token = null;
long expiresAt = 0L;
if (c instanceof AwsSessionCredentials) {
AwsSessionCredentials session = (AwsSessionCredentials) c;
token = session.sessionToken();
// Report the vended expiry, so Comet reuses the credential until shortly before it.
expiresAt = session.expirationTime().map(Instant::toEpochMilli).orElse(0L);
}
return new CometS3Credentials(c.accessKeyId(), c.secretAccessKey(), token, expiresAt);
}
}
VendedCredentialsProvider reads credentials.uri, the catalog endpoint, and OAuth tokens from the supplied map (Comet forwards the unfiltered FileIO bag to initialize), and refreshes through its own CachedSupplier. Refresh-near-expiry and the REST round-trip live in Iceberg. Comet only reuses the credential this returns until five minutes before the expiry it reports, so report the expiry: returning 0 turns that reuse off. Comet ships a copy of this class under spark/src/test as a reference; copy it into your runtime jar alongside iceberg-aws and AWS SDK v2.
Access mode#
Value |
Used for |
|---|---|
|
All native scan paths (raw Parquet, Iceberg). |
|
The native Iceberg writer (see Iceberg Writes). If the configured provider fails to initialize, the write fails rather than falling back to the default credential chain. |
A WRITE credential is not implicitly read-capable. Vendors that need read-during-write workflows include the required read permissions in the IAM policy attached to their WRITE credentials.
Build setup#
Vendor implementations need the Comet SPI classes at compile time only. Use provided-scope:
<dependency>
<groupId>org.apache.datafusion</groupId>
<artifactId>comet-spark-spark${spark.version.short}_${scala.binary.version}</artifactId>
<version>${comet.version}</version>
<scope>provided</scope>
</dependency>
Iceberg path: error message fidelity#
When the bridge is wired into iceberg-rust, the outer reqsign-core::ProvideCredentialChain currently swallows thrown exceptions into “no credential” before the request reaches opendal. The credential is still not issued and the request still fails, only the message is degraded to reqsign’s generic failed to load signing credential. No Comet change fixes this; it is resolved upstream if iceberg-rust stops wrapping custom loaders in its outer chain or moves its S3 backend to object_store.