Spark Version Compatibility#
This page documents known issues and limitations specific to each supported Apache Spark version.
For general compatibility information that applies across all Spark versions, see the other pages in this compatibility guide. For the rule that governs how long each Spark minor is supported, see the Apache Spark version support section of the versioning policy.
Spark 3.4#
Spark 3.4.3 is supported with Java 17 and Scala 2.12/2.13.
Warning
Spark 3.4 support is deprecated as of the 1.0.0 release and will be removed in a future release. Comet continues to build and publish Spark 3.4 binaries in the meantime, but Apache Spark’s own SQL test suite no longer runs against Spark 3.4 automatically: it runs only when a contributor opts a pull request into it. Regressions specific to Spark 3.4 are therefore more likely to reach a release than on the other supported versions. We recommend moving to Spark 3.5 or later.
Known Limitations#
Extra
SparkExceptionlayer in Parquet schema mismatch errors: when a Parquet read is rejected because a file’s type cannot be converted to the requested type, the error’s cause chain has one moreSparkExceptionlayer than Spark’s own reader produces. See Parquet Compatibility for details.
Spark 3.5#
Spark 3.5.9 is supported with Java 17 and Scala 2.12/2.13.
Known Limitations#
Extra
SparkExceptionlayer in Parquet schema mismatch errors: when a Parquet read is rejected because a file’s type cannot be converted to the requested type, the error’s cause chain has one moreSparkExceptionlayer than Spark’s own reader produces. See Parquet Compatibility for details.
Spark 4.0#
Spark 4.0.4 is supported with Java 17/21 and Scala 2.13.
Known Limitations#
Collation support: Spark 4.0 introduced collation support. Non-default collated strings are not yet supported by Comet and will fall back to Spark.
Spark 4.1#
Spark 4.1.3 is supported with Java 17/21 and Scala 2.13.
Known Limitations#
NullTypecolumns in Parquet files (#4199): Spark encodes aNullTypecolumn as a ParquetBOOLEANphysical type annotated withLogicalType::Unknown. The Rustparquetcrate that Comet depends on acceptsUnknownonly when paired withINT32and rejects any other physical type withParquet error: Cannot annotate Unknown from BOOLEAN for field '<name>'. Any attempt to read a Parquet file that contains aNullTypecolumn fails at decode time before Comet’s scan runs. Workaround: project the column away, cast it to a concrete type before persisting, or read the file with Comet disabled for that query.
Spark 4.2#
Spark 4.2.0 is supported with Java 17 and Scala 2.13.
Known Limitations#
FILTERon window aggregates: Spark 4.2 accepts aFILTER (WHERE ...)clause on an aggregate function used as a window function. Comet does not support the clause there, so the window operator falls back to Spark.Merged subqueries in
UNIONbranches (#4949): Spark 4.2 turns a single-row aggregate in some queries, for example TPC-DS q77a, into a scalar subquery that returns a struct, and reads its fields in a projection over aOneRowRelation. Comet does not support a struct-typed scalar subquery result yet (#5834), so that projection, and the union and the aggregates above it, fall back to Spark.ANSI arithmetic overflow (#4967): Spark 4.2 changed some ANSI-mode arithmetic overflow behavior, and Comet does not match these changes yet. See ANSI-mode error classes and messages.