Spark Data Type Support#

This page is the complete reference for how Apache Comet handles each Spark data type. Comet’s native execution path is built on Apache Arrow, so the set of types Comet can express natively is constrained by Arrow’s type system. When a query references a type Comet does not support, the relevant operator falls back to Spark; results are unaffected.

For per-scan and per-operator type caveats (for example, Parquet read-time conversions or hash-aggregate group-key restrictions), see the Compatibility Guide.

Status legend#

Status

Meaning

✅ Supported

Native support; enabled by default.

⚠️ Supported (caveats)

Works, but with limits: certain values, contexts, or configurations fall back to Spark.

🔜 Planned

Intended; tracked by an open issue or pull request.

Not currently planned#

The following types fall back to Spark and are not on the current roadmap. They are omitted from the tables below and may be reconsidered based on demand:

  • UserDefinedType: user-defined types are application-specific and outside the scope of native acceleration; queries referencing UDTs fall back to Spark.

Numeric#

Type

Status

Notes

ByteType

✅

ShortType

✅

IntegerType

✅

LongType

✅

FloatType

✅

NaN and signed-zero handling can diverge from Spark in comparisons and aggregations. See Floating-point Compatibility.

DoubleType

✅

NaN and signed-zero handling can diverge from Spark in comparisons and aggregations. See Floating-point Compatibility.

DecimalType

✅

String and binary#

Type

Status

Notes

StringType

✅

Default UTF-8 binary collation is supported. Non-default collations (Spark 4.0+) fall back (#2190).

BinaryType

✅

CharType

✅

Spark normalizes CHAR(n) to StringType for evaluation; same caveats apply.

VarcharType

✅

Spark normalizes VARCHAR(n) to StringType for evaluation; same caveats apply.

Boolean#

Type

Status

Notes

BooleanType

✅

Datetime#

Type

Status

Notes

DateType

✅

TimestampType

✅

TimestampNTZType

✅

TimeType

⚠️

Spark 4.1+. Native serialization is in place; some operators (sort, min/max) are still being wired up (#4288).

Interval#

All three interval types are mapped to Arrow and flow through serde, native shuffle, and the codegen dispatcher, so interval columns and the interval-producing expressions run natively. Several operators still gate on the type and fall back: Parquet scans of ANSI interval columns, single-column sorts, hash aggregates (min / max / sum / avg), GROUP BY, window functions. Remaining work is tracked by #5061.

Type

Status

Notes

YearMonthIntervalType

⚠️

Parquet scan, single-column sort, aggregate, GROUP BY, and window operators fall back.

DayTimeIntervalType

⚠️

Parquet scan, single-column sort, aggregate, GROUP BY, and window operators fall back.

CalendarIntervalType

⚠️

Parquet scan, single-column sort, aggregate, GROUP BY, and window operators fall back.

Complex#

Type

Status

Notes

StructType

✅

Empty structs (no fields) and structs with duplicate field names fall back.

ArrayType

✅

MapType

✅

Hash aggregate group keys cannot contain a MapType (transitively): Arrow’s row format used by DataFusion’s grouped hash aggregate does not support Map, so such groupings fall back.

Variant#

Type

Status

Notes

VariantType

⚠️

Spark 4.0+. Native Parquet scans support direct projection of top-level Variant columns. Non-null existence defaults fall back.

Direct projection requires explicit configuration on every supported Spark version: spark.sql.variant.allowReadingShredded=true (defaults to false in Spark 4.0) and spark.sql.variant.pushVariantIntoScan=false (defaults to true in Spark 4.1+), with the default Parquet timestamp inference settings. Support for Spark’s whole-value pushdown rewrite is tracked by #5519. Nested Variant columns, pushed-down Variant field extraction, expressions, writes, shuffle and spill, Python operators, encrypted files, and Iceberg scans that read a Variant column fall back to Spark. Iceberg scans of tables whose Variant columns the query does not read run natively. Spark also handles columnar-to-row conversion of the native scan output and strict reads with allowReadingShredded=false. Broader support is tracked by #4295 and #3983.

Shredded reconstruction can be slower than Spark’s reader; see the focused scan and allocation measurements in PR #5868.

Other#

Type

Status

Notes

NullType

✅

See also#