datetime_funcs Expression Audits#
Audit notes for expressions in this category that have been audited. Absence of an entry means the expression has not been audited yet, not that it is unsupported. See the user guide Spark Expression Support for current support status.
curdate#
Alias of
current_date; constant-folded to a literal by Spark’sComputeCurrentTimerule before Comet sees the plan.
current_date#
Constant-folded to a literal by Spark’s
ComputeCurrentTimerule before Comet sees the plan.
current_timestamp#
Constant-folded to a literal by Spark’s
ComputeCurrentTimerule before Comet sees the plan.
dayname#
Spark 4.0+. Implemented natively: maps a
DateTypevalue to a fixed US-English abbreviated day name (DayOfWeek.getDisplayName(TextStyle.SHORT, Locale.US)), with no session-locale or timezone dependence.
dayofweek#
Performance (tuned 2026-09-08, PR #5771): computed directly from the epoch day (
((days + 4).rem_euclid(7) + 1)) by the nativespark_dayofweekkernel, replacingdatepart('dow', ..)– which reconstructs aNaiveDateTimeper row and recomputes the null mask viaunary_opt– plus a separate+ 1arithmetic node. About 9x faster on flat input with no or sparse nulls, 3-4x at 87.5% nulls, and roughly 1.1x on an all-null batch or a cardinality-8 dictionary, where the replaced path already did little work (about 2x at cardinality 1024). Note the old path returned NULL for epoch days outside chrono’s range; the kernel is correct across the fulli32domain. Benchmark:benches/dayofweek_weekday.rs.
from_utc_timestamp#
Spark 3.4.3 (audited 2026-05-12): identical to 3.5.8.
Spark 3.5.8 (audited 2026-05-12): baseline.
Spark 4.0.1 (audited 2026-05-12):
inputTypeswidened toStringTypeWithCollation; behaviour unchanged for ASCII timezone strings.Marked
Incompatible: Comet’s native timezone parser only accepts IANA zone IDs (e.g.America/Los_Angeles) and fixed+HH:MMoffsets, while Spark also accepts legacy forms (GMT+1,UTC+1, three-letter abbreviations likePST). By default it runs through the codegen dispatcher (Spark-correct) and uses the native path only when incompatible expressions are explicitly allowed, where legacy zone forms throw a native parse error at execution.
hour#
Performance (tuned 2026-09-08, PR #5771):
hour/minute/secondtake an integer fast path when no timezone offset applies –TimestampNTZ, or a timezone-aware timestamp in a UTC session – computing the field from the stored microseconds with Euclidean division instead of building achronodatetime per row viadate_part. Dictionaries, non-microsecond units and offset timezones keep the general path. Every fast-path shape with a baseline of at least 2 µs takes 82-94% less time; offset session zones and dictionary input keep the general path and stay within noise on a repeat. Benchmark:benches/extract_date_part.rs.
make_timestamp_ltz#
The 6-argument form rewrites to
MakeTimestampand runs via the codegen dispatcher. The 2-argument(date, time)form requires the Spark 4.1 TIME type and falls back.
make_timestamp_ntz#
The 6-argument form rewrites to
MakeTimestampand runs via the codegen dispatcher. The 2-argument(date, time)form requires the Spark 4.1 TIME type and falls back.
minute#
Performance (tuned 2026-09-08, PR #5771):
hour/minute/secondtake an integer fast path when no timezone offset applies –TimestampNTZ, or a timezone-aware timestamp in a UTC session – computing the field from the stored microseconds with Euclidean division instead of building achronodatetime per row viadate_part. Dictionaries, non-microsecond units and offset timezones keep the general path. Every fast-path shape with a baseline of at least 2 µs takes 82-94% less time; offset session zones and dictionary input keep the general path and stay within noise on a repeat. Benchmark:benches/extract_date_part.rs.
monthname#
Spark 4.0+. Implemented natively: maps a
DateTypevalue to a fixed US-English abbreviated month name (Month.getDisplayName(TextStyle.SHORT, Locale.US)), with no session-locale or timezone dependence.
now#
Alias of
current_timestamp; constant-folded to a literal by Spark’sComputeCurrentTimerule before Comet sees the plan.
second#
Performance (tuned 2026-09-08, PR #5771):
hour/minute/secondtake an integer fast path when no timezone offset applies –TimestampNTZ, or a timezone-aware timestamp in a UTC session – computing the field from the stored microseconds with Euclidean division instead of building achronodatetime per row viadate_part. Dictionaries, non-microsecond units and offset timezones keep the general path. Every fast-path shape with a baseline of at least 2 µs takes 82-94% less time; offset session zones and dictionary input keep the general path and stay within noise on a repeat. Benchmark:benches/extract_date_part.rs.
to_date#
Rewrites to
Cast(no format, native) orCast(GetTimestamp(...))(with format, via the codegen dispatcher) before Comet sees the plan.
to_timestamp#
Rewrites to
Cast(no format, native) orGetTimestamp(with format, via the codegen dispatcher) before Comet sees the plan.
to_timestamp_ltz#
Rewrites to
to_timestampwithTimestampType; same support asto_timestamp.
to_timestamp_ntz#
Rewrites to
to_timestampwithTimestampNTZType; same support asto_timestamp.
to_utc_timestamp#
Spark 3.4.3 (audited 2026-05-12): identical to 3.5.8.
Spark 3.5.8 (audited 2026-05-12): baseline.
Spark 4.0.1 (audited 2026-05-12):
inputTypeswidened toStringTypeWithCollation; behaviour unchanged for ASCII timezone strings.Marked
Incompatible: Comet’s native timezone parser only accepts IANA zone IDs (e.g.America/Los_Angeles) and fixed+HH:MMoffsets, while Spark also accepts legacy forms (GMT+1,UTC+1, three-letter abbreviations likePST). By default it runs through the codegen dispatcher (Spark-correct) and uses the native path only when incompatible expressions are explicitly allowed, where legacy zone forms throw a native parse error at execution.
try_make_timestamp#
Rewrites to
MakeTimestamp(failOnError = false)and runs through the codegen dispatcher (CometMakeTimestamp), so invalid inputs return NULL to match Spark.
try_to_date#
Spark 4.1+. Rewrites to
Cast(..., EvalMode.LEGACY)(no format, native) orCast(GetTimestamp(..., failOnError = false))(with format, via the codegen dispatcher) before Comet sees the plan. In non-ANSI mode the rewritten tree is identical toto_date; invalid inputs return NULL to match Spark.
try_to_timestamp#
Rewrites to
Cast(..., EvalMode.LEGACY)(no format, native) orGetTimestamp(..., failOnError = false)(with format, via the codegen dispatcher) before Comet sees the plan. In non-ANSI mode the rewritten tree is identical toto_timestamp; invalid inputs return NULL to match Spark.
unix_timestamp#
Spark 3.4.3 (audited 2026-09-13): baseline. String inputs accept literal or column formats. Date, timestamp, and timestamp without time zone inputs ignore the format argument.
Spark 3.5.8 (audited 2026-09-13): parsing failures use structured timestamp parsing errors.
Spark 4.0.1 (audited 2026-09-13):
inputTypeswidened toStringTypeWithCollationfor the input and format arguments.Spark 4.1.1 (audited 2026-09-13): same input types and parsing behavior as Spark 4.0.1.
String inputs use Spark’s generated parser through codegen dispatch, including collated strings and formats. Literal and column formats preserve null handling, ANSI errors, parser policy, and session time zone.
Date, timestamp, and timestamp without time zone inputs retain native execution and ignore the format, including its collation. String input stays unsupported by the native serializer so
allowIncompatible=truecannot send it to the native kernel.Native timestamp conversion truncates fractional seconds toward zero, matching Spark’s
ToTimestamp. This fixes the previous use of floor division for negative fractional timestamps: at UTC,1969-12-31 23:59:58.5produces-1, not-2. Casting a timestamp toBIGINTdeliberately uses floor division in Spark and Comet, so that cast still produces-2.
weekday#
Performance (tuned 2026-09-08, PR #5771): computed directly from the epoch day (
(days + 3).rem_euclid(7)) by the nativespark_weekdaykernel, replacingdatepart('isodow', ..)plus a- 1arithmetic node. About 9x faster on flat input with no or sparse nulls, 3-4x at 87.5% nulls, and roughly 1.1x on an all-null batch or a cardinality-8 dictionary, where the replaced path already did little work (about 2x at cardinality 1024).weekdaynumbers Monday = 0 through Sunday = 6, a different convention fromdayofweek. Benchmark:benches/dayofweek_weekday.rs.