Spark Version Compatibility#

This page documents known issues and limitations specific to each supported Apache Spark version.

For general compatibility information that applies across all Spark versions, see the other pages in this compatibility guide. For the rule that governs how long each Spark minor is supported, see the Apache Spark version support section of the versioning policy.

Spark 3.4#

Spark 3.4.3 is supported with Java 17 and Scala 2.12/2.13.

Warning

Spark 3.4 support is deprecated as of the 1.0.0 release and will be removed in a future release. Comet continues to build and publish Spark 3.4 binaries in the meantime, but Apache Spark’s own SQL test suite no longer runs against Spark 3.4 automatically: it runs only when a contributor opts a pull request into it. Regressions specific to Spark 3.4 are therefore more likely to reach a release than on the other supported versions. We recommend moving to Spark 3.5 or later.

Known Limitations#

  • Extra SparkException layer in Parquet schema mismatch errors: when a Parquet read is rejected because a file’s type cannot be converted to the requested type, the error’s cause chain has one more SparkException layer than Spark’s own reader produces. See Parquet Compatibility for details.

Spark 3.5#

Spark 3.5.9 is supported with Java 17 and Scala 2.12/2.13.

Known Limitations#

  • Extra SparkException layer in Parquet schema mismatch errors: when a Parquet read is rejected because a file’s type cannot be converted to the requested type, the error’s cause chain has one more SparkException layer than Spark’s own reader produces. See Parquet Compatibility for details.

Spark 4.0#

Spark 4.0.4 is supported with Java 17/21 and Scala 2.13.

Known Limitations#

  • Collation support: Spark 4.0 introduced collation support. Non-default collated strings are not yet supported by Comet and will fall back to Spark.

Spark 4.1#

Spark 4.1.3 is supported with Java 17/21 and Scala 2.13.

Known Limitations#

  • NullType columns in Parquet files (#4199): Spark encodes a NullType column as a Parquet BOOLEAN physical type annotated with LogicalType::Unknown. The Rust parquet crate that Comet depends on accepts Unknown only when paired with INT32 and rejects any other physical type with Parquet error: Cannot annotate Unknown from BOOLEAN for field '<name>'. Any attempt to read a Parquet file that contains a NullType column fails at decode time before Comet’s scan runs. Workaround: project the column away, cast it to a concrete type before persisting, or read the file with Comet disabled for that query.

Spark 4.2#

Spark 4.2.0 is supported with Java 17 and Scala 2.13.

Known Limitations#

  • FILTER on window aggregates: Spark 4.2 accepts a FILTER (WHERE ...) clause on an aggregate function used as a window function. Comet does not support the clause there, so the window operator falls back to Spark.

  • Merged subqueries in UNION branches (#4949): Spark 4.2 turns a single-row aggregate in some queries, for example TPC-DS q77a, into a scalar subquery that returns a struct, and reads its fields in a projection over a OneRowRelation. Comet does not support a struct-typed scalar subquery result yet (#5834), so that projection, and the union and the aggregates above it, fall back to Spark.

  • ANSI arithmetic overflow (#4967): Spark 4.2 changed some ANSI-mode arithmetic overflow behavior, and Comet does not match these changes yet. See ANSI-mode error classes and messages.