Floating-Point Semantics#
This page describes how Comet matches Spark’s handling of -0.0 and NaN in FLOAT and DOUBLE
values. It is aimed at contributors working on an expression or operator that compares, orders,
hashes, deduplicates or groups floating-point values, including values nested in arrays, structs
and maps. For user-facing differences from Spark, see
Floating-point Number Comparison. The work
to apply these rules systematically is tracked in
#6385.
The short version: Spark has no single rule for floating-point equality. Each function inherits one
from the Java API that its implementation calls. Arrow and DataFusion follow IEEE 754 total order,
which matches none of them. A native implementation must follow the rule of the Spark function it
replaces, using the shared helpers in native/spark-expr/src/float_semantics/, and its tests must
include a NaN with the sign bit set.
Spark’s rules#
Rule |
|
NaN |
Used by |
|---|---|---|---|
SQL ordering ( |
Equal |
All NaNs are equal, and NaN sorts above all values |
Comparisons, |
|
|
Every NaN becomes the canonical NaN |
Grouping keys, join keys and window partition keys |
|
Same hash |
Hashed through |
|
|
Distinct |
All NaNs are equal |
Java hash sets and maps of boxed values, and Spark’s |
Scala |
Equal |
A NaN equals nothing, not even another NaN |
Scala sets and maps of |
|
|
As in SQL ordering |
An ascending |
Some functions changed rules in a Spark release, so a native path has to follow the Spark version it runs against:
collect_setkeys its buffer with Scala’s==before Spark 4.2, so-0.0and0.0are one value and every NaN is a value of its own. Spark 4.2 normalizes NaN and-0.0first and keys the buffer by the normalized bits, so all NaNs are one value too (SPARK-57298).Spark’s
OpenHashSet, and theOpenHashMapbuilt on it, followDouble.equalsonly from Spark 3.5.2 and 4.0.0 (SPARK-45599). In every 3.4 release and in 3.5.0 and 3.5.1 they match a key with==but hash its bits (doubleToLongBits), so two NaNs never match, and-0.0and0.0match only when probing from one reaches the other.modeandpercentilecount values in anOpenHashMap.array_distinct,array_union,array_intersectandarray_exceptlook up the elements of a flat array in anOpenHashSet, but treat all NaNs as one value.modefollowsDouble.equalsfrom Spark 3.5.2 and 4.0.0 until Spark 4.2, which folds-0.0into0.0first (SPARK-57329).array_distinctandarray_unionkeep signed zeros apart in a flat array before Spark 4.2.0 (before 3.5.2, unless probing merges them). Spark 4.2.0 normalizes their arguments in the plan (SPARK-54918). From 4.0.5, 4.1.4 and 4.2.1 they normalize while they evaluate instead (SPARK-59602).Map construction (
ArrayBasedMapBuilder) finds duplicate keys byDouble.equalsin Spark 3.4 and 3.5. From Spark 4.0 it normalizes each key first, unlessspark.sql.legacy.disableMapKeyNormalizationis set (#6549).
How Arrow and DataFusion differ#
Arrow compares floats by IEEE 754 total order. -0.0 sorts below 0.0, NaNs compare by their bit
patterns, and a NaN with the sign bit set sorts below -Infinity. Arrow’s sort, row format and
hash kernels all work from the same bits. DataFusion 55 folds -0.0 into 0.0 in some kernels but
does not canonicalize NaN, and that has changed between DataFusion releases. Don’t rely on it:
normalize the values, or compare them with the helpers below.
NaNs with the sign bit set are the normal case on x86-64, not a corner case. Every NaN that
arithmetic produces at run time, such as sqrt(-1) or Infinity - Infinity, is
0xfff8000000000000 in both Rust and the JVM on x86-64, while aarch64 produces
0x7ff8000000000000. Spark hides the difference through doubleToLongBits. A native path that
compares raw values puts those NaNs below every other value, so a query that passes on an Apple
Silicon laptop can fail on a Linux x86 cluster.
Where Comet applies the rules#
The float_semantics module#
native/spark-expr/src/float_semantics/ holds the per-value rules and the kernels built on them.
Its module documentation is the reference. In short:
Helper |
Rule |
Use it to |
|---|---|---|
|
SQL ordering |
Compare values in a kernel that returns the original bits, as |
|
SQL ordering, at any depth |
Compare arrays and structs in place |
|
|
Normalize values before an Arrow kernel sorts, row-encodes, hashes or compares them |
|
|
Wrap a key or operand expression. |
|
|
Key a hash set or map the way boxed Java values do |
|
|
Sort the way |
|
|
Get the value to hash in place of a float |
Once -0.0 is folded and every NaN is canonical, Arrow’s total order agrees with Spark’s SQL
ordering. That is why normalizing the inputs of an Arrow kernel works, as long as only the keys are
normalized and the output keeps the original values.
Keys, sorting and comparisons#
Spark’s optimizer wraps grouping, join and window partition keys in
NormalizeNaNAndZero(NormalizeFloatingNumbers). Comet serializes it like any other expression.create_normalized_key_exprinnative/core/src/execution/planner.rsnormalizes sort keys, window order and partition keys, and the partition keys ofWindowGroupLimit, so that Arrow’s sort orders them the Spark way.spark_comparisoninnative/spark-expr/src/array_funcs/nested_comparison.rsbuilds every native=,<>,<=>,<,<=,>,>=andIS DISTINCT FROM, whichever operator evaluates it. It normalizes float operands, and folds a literal while the plan is built so that it stays a literal. Nested=and<>compare in place withspark_equality.A scan’s pushed-down data filters leave a float column compared with a literal other than NaN unwrapped (
FloatOperands::Raw) so that Parquet pruning still recognizes it, but only while the reader prunes with them without filtering rows. A bloom filter probe hashes the literal’s bits, so=against either zero becomes= -0.0 OR = 0.0. Withspark.comet.parquet.rowFilterPushdown.enabled=truethey normalize both sides, because the reader drops the rows a filter rejects, and a raw column would reject a stored NaN that Spark matches (#6702). Spark’s own reader keeps the two zeros apart in its dictionary and bloom filters, so it can skip a row group that Comet reads. A test of this pruning writes out the expected rows instead of comparing them with Spark’s.INnormalizes its operands in the serde (normalizeInOperandinpredicates.scala).hash,xxhash64, the native shuffle’s hash partitioner andapprox_count_distincthash floats throughhash_input.
Spark version differences#
Rules that depend on the Spark version are decided in Scala. The serde either chooses the native
path or passes a flag to native code, as Mode’s normalize_neg_zero does. Where a patch release
changed the rule, check the runtime version down to the patch with
Utils.majorMinorPatchVersion(SPARK_VERSION), as ArraySetSupport in arrays.scala does.
Until an expression has a native implementation that follows its rule, report it Incompatible
for floating-point input so that Spark evaluates it.
spark.comet.exec.strictFloatingPoint=true makes Comet fall back for floating-point operations
that can still differ from Spark. To gate such an operation, use
SupportLevel.strictFloatingPointReason(dataType, what). It returns a reason only in strict mode,
and only for a type that contains a FLOAT or DOUBLE.
Choosing the rule for an expression#
Read Spark’s implementation, both
evalanddoGenCode, in every Spark version that Comet supports, at the latest patch release of each. What it calls decides the rule.ordering.compare,ordering.equiv,ctx.genCompandctx.genEqualmean SQL ordering. Ajava.util.HashMapor another Java collection of boxed values meansDouble.equals, and so does anOpenHashSetorOpenHashMapfrom Spark 3.5.2 and 4.0.0. A Scala collection ofAny, such asmutable.HashSet[Any], means Scala’s==.java.util.Arrays.sorton a primitive array meansDouble.compare.Check whether Spark’s optimizer normalizes the input in the plan. If it does, Comet receives normalized values on that version only.
Apply the rule at every depth. A float inside an array, struct or map follows the same rule, and some DataFusion kernels fold
-0.0only in a flat array.Record what you found in the expression’s audit and, for a difference users can see, in the compatibility guide.
Guidelines#
Don’t compare floating-point values with Arrow’s comparison, sort or hash kernels, or with
total_cmporpartial_cmp, where Spark uses SQL ordering, unless you normalize them first.Normalize keys and operands, not results. Spark returns the original bits:
greatest(-0.0D, 0.0D)returns-0.0,maxreturns whichever zero it reads first, andarray_minreturns the first of equal elements.Don’t write a local copy of a rule. If a helper is missing, add it to
float_semantics.In a hot loop, use the
float_ltandfloat_gtpredicates. Comparingcompare_floats(a, b)with an ordering known only at run time was several times slower in a scan for a minimum.When equality on nested values can stop early, keep that exit.
spark_equalityrejects lists of different lengths without comparing their elements, which an equality built on an ordering comparator loses.
Testing#
Use the edge values
0.0,-0.0, a canonical NaN, a NaN with the sign bit set,1.0,-1.0,Infinity,-InfinityandNULL, for bothFLOATandDOUBLE.Make the sign-bit NaN at query time. Spark’s Parquet writer canonicalizes NaN, so a stored NaN reads back canonical. Negate it in the query (
-d), which flips the sign bit on every platform. Arithmetic such assqrt(-1)yields a sign-bit NaN only on x86-64, so a test that relies on it passes on Apple Silicon without exercising anything. The writer keeps the sign of zero, so-0.0can be stored directly.Read the values from a Parquet table, and also compare them with literals on both sides. The Comet SQL Tests turn constant folding off, so
-0.0Dstays a literal with the sign bit set, whiledouble('NaN')is a cast.Cover every place the rule applies. For a comparison, that is
Project,Filter, aggregate arguments,FILTERclauses, join conditions, sort keys, generators and windows.Nest the values in arrays and structs, two levels deep, and make the nested NaN a sign-bit one.
When arithmetic produces the NaN, run with ANSI mode both on and off (
-- ConfigMatrix: spark.sql.ansi.enabled=false,true), because division and remainder take different native paths.Where the rule changed in a Spark release, gate the fixtures with
MinSparkVersionorMaxSparkVersion, and add therun-all-spark-profileslabel to the pull request. CI pins one patch release per Spark line, so when a patch release changed the rule, also test the routing in Scala.Add the expression to
CometFloatSemanticsSuite, which crosses the edge values with operator contexts and with every expression and aggregate that accepts floating-point input. New expressions go in itsexpressionsoraggregateslist. If one does not match Spark yet, add aKnownGapthat links the issue. Known gaps are strict: a case that starts matching Spark fails, so the fix also removes the entry.
Fixtures to start from include expressions/conditional/float_comparisons.sql,
expressions/math/nan_divisor.sql, expressions/math/greatest_least_floating_point.sql,
expressions/aggregate/min_max_floating_point.sql, expressions/array/sort_array_floating_point.sql
and windows/nested_float_order_keys.sql, all under spark/src/test/resources/sql-tests/.