Spark Data Type Support#
This page is the complete reference for how Apache Comet handles each Spark data type. Comet’s native execution path is built on Apache Arrow, so the set of types Comet can express natively is constrained by Arrow’s type system. When a query references a type Comet does not support, the relevant operator falls back to Spark; results are unaffected.
For per-scan and per-operator type caveats (for example, Parquet read-time conversions or hash-aggregate group-key restrictions), see the Compatibility Guide.
Status legend#
Status |
Meaning |
|---|---|
✅ Supported |
Native support; enabled by default. |
⚠️ Supported (caveats) |
Works, but with limits: certain values, contexts, or configurations fall back to Spark. |
🔜 Planned |
Intended; tracked by an open issue or pull request. |
Not currently planned#
The following types fall back to Spark and are not on the current roadmap. They are omitted from the tables below and may be reconsidered based on demand:
UserDefinedType: user-defined types are application-specific and outside the scope of native acceleration; queries referencing UDTs fall back to Spark.
Numeric#
Type |
Status |
Notes |
|---|---|---|
|
✅ |
|
|
✅ |
|
|
✅ |
|
|
✅ |
|
|
✅ |
NaN and signed-zero handling can diverge from Spark in comparisons and aggregations. See Floating-point Compatibility. |
|
✅ |
NaN and signed-zero handling can diverge from Spark in comparisons and aggregations. See Floating-point Compatibility. |
|
✅ |
String and binary#
Type |
Status |
Notes |
|---|---|---|
|
✅ |
Default UTF-8 binary collation is supported. Non-default collations (Spark 4.0+) fall back (#2190). |
|
✅ |
|
|
✅ |
Spark normalizes |
|
✅ |
Spark normalizes |
Boolean#
Type |
Status |
Notes |
|---|---|---|
|
✅ |
Datetime#
Type |
Status |
Notes |
|---|---|---|
|
✅ |
|
|
✅ |
|
|
✅ |
|
|
⚠️ |
Spark 4.1+. Native serialization is in place; some operators (sort, min/max) are still being wired up (#4288). |
Interval#
All three interval types are mapped to Arrow and flow through serde, native shuffle, and the
codegen dispatcher, so interval columns and the interval-producing expressions run natively.
Several operators still gate on the type and fall back: Parquet scans of ANSI interval columns,
single-column sorts, hash aggregates (min / max / sum / avg), GROUP BY, window
functions. Remaining work is tracked by
#5061.
Type |
Status |
Notes |
|---|---|---|
|
⚠️ |
Parquet scan, single-column sort, aggregate, |
|
⚠️ |
Parquet scan, single-column sort, aggregate, |
|
⚠️ |
Parquet scan, single-column sort, aggregate, |
Complex#
Type |
Status |
Notes |
|---|---|---|
|
✅ |
Empty structs (no fields) and structs with duplicate field names fall back. |
|
✅ |
|
|
✅ |
Hash aggregate group keys cannot contain a |
Variant#
Type |
Status |
Notes |
|---|---|---|
|
⚠️ |
Spark 4.0+. Native Parquet scans support direct projection of top-level Variant columns. Non-null existence defaults fall back. |
Direct projection requires explicit configuration on every supported Spark version:
spark.sql.variant.allowReadingShredded=true (defaults to false in Spark 4.0) and
spark.sql.variant.pushVariantIntoScan=false (defaults to true in Spark 4.1+), with the default
Parquet timestamp inference settings. Support for Spark’s whole-value pushdown rewrite is tracked
by #5519. Nested Variant columns, pushed-down
Variant field extraction, expressions, writes, shuffle and spill, Python operators, encrypted
files, and Iceberg scans that read a Variant column fall back to Spark. Iceberg scans of tables
whose Variant columns the query does not read run natively. Spark also handles columnar-to-row
conversion of the native scan output and strict reads with allowReadingShredded=false. Broader
support is tracked by #4295 and
#3983.
Shredded reconstruction can be slower than Spark’s reader; see the focused scan and allocation measurements in PR #5868.
Other#
Type |
Status |
Notes |
|---|---|---|
|
✅ |
See also#
Comet Compatibility Guide - known incompatibilities and edge cases.
Parquet Scan Compatibility - per-type behavior at scan time.
Supported Spark Operators - the equivalent reference for operators.
Supported Spark Expressions - the equivalent reference for expressions.