DataFusion Comet 1.1.0 Changelog#

This release consists of 416 commits from 41 contributors. See credits at the end of this changelog for more information.

Fixed bugs:

  • fix: skip null slots when checking overflow in unary negation #5162 (Smallfu666)

  • fix: surface next_day and make_date ANSI errors as Spark exceptions #5167 (peterxcli)

  • fix: normalize nested field nullability in ShuffleScanExec and ExpandExec #5138 (andygrove)

  • fix: propagate the Spark task ClassLoader to JVM UDF calls #5282 (andygrove)

  • fix: avoid duplicate CheckOverflow evaluation for decimal division #5225 (peterxcli)

  • fix: match Spark whitespace trimming in to_time and try_to_time #5364 (sunchao)

  • fix: preserve Catalyst nullability and field IDs in native Parquet writes #5369 (sunchao)

  • fix: guard against silent fail_on_error loss in scalar wiring (#5074) #5359 (sam-1112)

  • fix: preserve Spark semantics for dictionary-encoded Parquet inputs and reject dictionary targets #5234 (peterxcli)

  • fix: canonicalize NaN in flat arrays_overlap float keys #5376 (sunchao)

  • fix: report native shuffle write metrics accurately #5370 (sunchao)

  • fix: Native shuffle fails with a 2GB task serialization OOM on jobs with many partitions #5392 (parthchandra)

  • fix: format NativeMemoryConsumer id in toString #5398 (ywskycn)

  • fix: ignore reader-side parquet.hadoop.vectored.io.enabled in Iceberg native-write detection #5410 (snmvaughan)

  • fix: support map casts with NullType elements #5045 (peterxcli)

  • fix: distinguish reflection failure from absent accessor in IcebergReflection #5412 (unikdahal)

  • fix: use current shuffle config in aggregate test #5439 (peterxcli)

  • fix: support wide years in native make_date #5443 (peterxcli)

  • fix: report native child spill metrics in shuffle tasks #5445 (sunchao)

  • fix: expose native memory usage to Spark #5408 (ywskycn)

  • fix: release native shuffle reservation after spill failure #5461 (peterxcli)

  • fix: track peak memory before JVM shuffle spill #5463 (peterxcli)

  • fix: RAII for memory pool registration #5464 (peterxcli)

  • fix: tighten RSS writer visibility and JNI array limits #5475 (pingzh)

  • fix: report native operator spill metrics in Spark task metrics for non-shuffle stages #5497 (peterxcli)

  • fix: report bounded shuffle allocator memory usage #5516 (ywskycn)

  • fix: use Spark type names in ANSI abs overflow errors #5357 (Smallfu666)

  • fix: make collect_list/collect_set argument coercion a normalization barrier #5159 (andygrove)

  • fix: make task-shared memory pool as ref-counted RAII guard #5494 (peterxcli)

  • fix: accept UTC timezone aliases in Python Arrow input #5556 (sunchao)

  • fix: normalize scalar float sort and window rank keys #5469 (sunchao)

  • fix: make CometExplodeExec respect batch size #5362 (andygrove)

  • fix: align Spark 4.2 Python worker configuration #5561 (sunchao)

  • fix: prevent memory leak after failed Arrow vector import #5539 (1fanwang)

  • fix: preserve Arrow Field metadata across C Data exports #5552 (peterxcli)

  • fix: rebase map offsets in mapsort so sliced maps do not overrun entries #5630 (viirya)

  • fix: scope Celeborn bootstrap hooks to Comet clients #5627 (pingzh)

  • fix: skip codegen dispatcher null short-circuit when a foldable subtree can raise #5623 (andygrove)

  • fix: accept dictionary encodings in remote shuffle #5650 (pingzh)

  • fix: make CometDiskBlockWriter spill registry per-task instead of executor-global #5493 (peterxcli)

  • fix: Bump iceberg-rust so native Iceberg writes URL-escape partition paths #5651 (andygrove)

  • fix(celeborn): reject unsafe native push completion tracking #5665 (pingzh)

  • fix: match Spark’s null short-circuiting in array_join and enable it natively #5558 (Visorgood)

  • fix: make native shuffle spill metrics independent of input batching #5628 (sunchao)

  • fix: expand object store option references, uniquify constant metadata names, drop dead parquet JNI #5653 (dwsmith1983)

  • fix: match Spark’s ANSI bound check for float/double to integral casts #5683 (peterxcli)

  • fix: return NULL from rpad/lpad when the length column is NULL instead of panicking #5680 (peterxcli)

  • fix: fall back for concat_ws with array arguments instead of failing natively #5679 (peterxcli)

  • fix: prevent silent overflow when reading Parquet TIMESTAMP_MILLIS values #5177 (peterxcli)

  • fix: recover native Celeborn shuffle from oversized rows #5668 (pingzh)

  • fix: report native shuffle read metrics #5554 (peterxcli)

  • fix: correctly rounded decimal to double/float cast matching BigDecimal.doubleValue/floatValue #5684 (peterxcli)

  • fix: restore columnar transitions under the native Iceberg write #5696 (andygrove)

  • fix: make columnar-to-row benchmarks exercise Comet #5718 (rich7420)

  • fix: distinguish “nothing spilled” from a spill backend with no local path #5726 (andygrove)

  • fix: propagate Arrow array copy errors #5747 (rich7420)

  • fix: keep the dictionary hash fast path off nested and reseeded buffers #5757 (viirya)

  • fix: preserve ANSI errors for rejected TIMESTAMP_NTZ casts #5752 (peterxcli)

  • fix: apply the parent struct’s null mask before hashing its fields #5754 (viirya)

  • fix: read Iceberg tables partitioned by an unknown transform #5759 (andygrove)

  • fix: native Iceberg write panics on an evolved partition spec and on a timestamptz partition path #5729 (andygrove)

  • fix: let AQE optimize queries over Comet caches #5733 (peterxcli)

  • fix: dispatch Iceberg system functions wrapped as ApplyFunctionExpression #5773 (andygrove)

  • fix: attach tokio runtime threads to the JVM as daemon threads #5748 (zhangfengcdt)

  • fix: check nested TIMESTAMP_MILLIS overflow in unfiltered scans #5740 (peterxcli)

  • fix: match iceberg-java’s exception for unclustered input to a clustered Iceberg write #5779 (andygrove)

  • fix: enable FIRST/LAST partial merge #5041 (peterxcli)

  • fix: preserve aggregate result identity during exchange reuse #5470 (sunchao)

  • fix: ignore structural tags when lifting expression coverage #5471 (sunchao)

  • fix: align string to timestamp parsing with Spark’s segment rules #5682 (peterxcli)

  • fix: roll native Iceberg data files on iceberg-java’s 1000-row grid #5780 (andygrove)

  • fix: decide libhdfs routing from the scheme as written #5825 (comphead)

  • fix: list a fanout Iceberg write’s data files in a stable order #5810 (andygrove)

  • fix: isolate object-store registration by backend and configuration #5503 (sunchao)

  • fix: Delete completed tasks’ data files when an Iceberg write job fails #5663 (andygrove)

  • fix: apply Spark’s Parquet conversion rules to nested struct/list/map fields #5681 (peterxcli)

  • fix: preserve nulls for Boolean/Byte/Short/Integer columns in FuzzDataGenerator #5855 (Smallfu666)

  • fix: Accept explicit positive years in timestamp casts #5858 (peterxcli)

  • fix: decline structs with duplicate field names before they reach Java Arrow #5866 (dwsmith1983)

  • fix: explain ObjectHashAggregate fallback when Comet shuffle is disabled #5746 (0lai0)

  • fix: preserve join and generator semantics in plan identity #5828 (ErikBPF)

  • fix: propagate Parquet field-name folding failures #5845 (sunchao)

  • fix: decode dictionary input for PyArrow UDFs #5560 (sunchao)

  • fix: size JVM shuffle pointer array growth from the array, not the data pages #5907 (andygrove)

  • fix: render float and double Iceberg partition values like iceberg-java #5840 (andygrove)

  • fix: align time parsing and native second extraction with Spark #5738 (peterxcli)

  • fix: normalize floating-point values in native collect_set #5166 (peterxcli)

  • fix: keep Iceberg complex null checks on native scans #5732 (ErikBPF)

  • fix: decode invalid UTF-8 at the JVM to native FFI import boundary #5310 (manuzhang, andygrove)

  • fix: fall back when a struct repeats a Parquet field id #6004 (comphead)

  • fix: normalize signed zero in nested float array comparisons #5235 (divyankshah)

  • fix: fall back to Spark for native Iceberg writes to gs:// through HadoopFileIO #5935 (zhangfengcdt)

  • fix: gate the regr_r2 degenerate-case swap on the Spark patch release #6042 (dwsmith1983)

  • fix: include ABFS container in object store cache key #5053 (peterxcli)

  • fix: preserve Spark row index read errors #6046 (liupoyi-1031)

  • fix: remove the ineffective spark.executor.memoryOverhead adjustment from the driver plugin #6054 (andygrove)

  • fix: count each memory pool once in analyze_trace #5991 (andygrove)

  • fix: revert unsafe partial aggregates after final fallback #5421 (sunchao)

  • fix: restore the site’s mermaid diagrams and make a dropped one fail CI #6064 (andygrove)

  • fix: write cached batches to the schema width, not the batch width #6090 (andygrove)

  • fix(iceberg): guard native Iceberg scan driver-metric double-post, add metrics docs and tests #6085 (parthchandra)

  • fix: let decimal SUM recover from an intermediate overflow like Spark #6041 (dwsmith1983)

  • fix: Nested floating-point IN membership does not match Spark for signed zero #6073 (mizulun)

  • fix: let Comet memory pools overcommit on grow instead of panicking #6128 (andygrove)

  • fix: correct two nightly test failures on Spark 3.4 and 4.2 #6156 (andygrove)

  • fix: support null calendar interval literals #5133 (peterxcli)

  • fix: wrap Iceberg split-write failures the way Spark does when abort fails #6153 (andygrove)

  • fix: support Utf8/LargeUtf8/Utf8View in native RLike without panicking #5215 (sam-1112)

  • fix: normalize noncanonical NaN literals in comparisons #5472 (sunchao)

  • fix: preserve map field metadata and honor target sorted flag in cast_map_to_map #5227 (Smallfu666)

  • fix: reject duplicate Parquet field names before decoding #5786 (ErikBPF)

  • fix: preserve Parquet conversion errors during join filtering #6067 (pingzh)

  • fix: plan Iceberg writes with Spark’s own operator when Comet is disabled #6151 (andygrove)

  • fix: read shuffle write buffer, spill limit and off-heap sizes in bytes #6191 (andygrove)

  • fix: bound shuffle schema cache retention and preserve eviction order #6098 (sunchao)

  • fix: do not run Comet in on-heap mode without spark.comet.exec.onHeap.enabled #6195 (andygrove)

  • fix: remove misleading native opt-in for dispatch-only datetime expressions #6182 (LinSimon-901101)

  • fix: support CalendarIntervalType hashing #5135 (peterxcli)

  • fix: report Input column when native Iceberg scan is enabled or native shuffle is enabled #5265 (hsiang-c)

  • fix: skip the executor memory overhead warning when a factor is set or in local mode #6198 (andygrove)

  • fix: keep operators above a cached relation native after the AQE re-plan #6208 (andygrove)

  • fix: preserve current AQE logical-stage links on Comet operators #5483 (sunchao)

  • fix: check each fair_unified reservation against its own share #6205 (andygrove)

  • fix(iceberg): don’t push transform residuals as their source column, fail on residual errors #6154 (andygrove)

  • fix: match Spark’s duplicate field and field id semantics in parquet field lookup #5654 (dwsmith1983)

  • fix: park the native scan loop instead of busy-polling while waiting on native I/O #6219 (andygrove, mixermt)

  • fix: refresh S3 policy locations when a location’s credential fails #6223 (andygrove, snmvaughan)

  • fix: inject Comet session extension rules only once per session #6230 (andygrove)

  • fix: [branch-1.1] reject a file without field ids at any depth whether or not id matching is on (#6116) #6266 (andygrove, dwsmith1983)

  • fix: [branch-1.1] decline native Iceberg writes with a custom location provider (#6216) #6305 (andygrove, liupoyi-1031)

  • fix: [branch-1.1] Native S3 scan on EKS/IRSA turns a transient STS throttle into a hard 403 storm (#6025) #6323 (andygrove, parthchandra)

  • fix: [branch-1.1] log partial memory grants at DEBUG and drop the memory usage dump (#6269) #6346 (andygrove)

  • fix: [branch-1.1] zero sliced boolean offsets at every level before exporting to the JVM (#6339) #6449 (andygrove)

  • fix: [branch-1.1] build with Rust 1.99, which deprecates the legacy f64 constants and fetch_update (#6507) #6526 (andygrove)

  • fix: [branch-1.1] make adaptive aggregation skipping opt-in (#6474) #6488 (andygrove)

  • fix: [branch-1.1] fall back to Spark for _metadata.file_block_start and file_block_length (#6510) #6511 (andygrove)

  • fix: [branch-1.1] fall back for rank limits over nested float keys (#6468) #6487 (andygrove)

  • fix: [branch-1.1] fall back for incompatible regression aggregates (#6451) #6489 (andygrove)

  • fix: [branch-1.1] coerce native IF branches to a common type (#6458) #6491 (andygrove)

  • fix: [branch-1.1] match pre-epoch Iceberg temporal rounding (#6456) #6486 (andygrove)

  • fix: [branch-1.1] defer Parquet conversion errors until a row group is decoded, as Spark does (#6515) #6540 (andygrove)

  • fix: [branch-1.1] dispatch StaticInvoke and Invoke only into Spark’s own classes (#6542) #6554 (andygrove)

  • fix: [branch-1.1] gate array distinct and union signed-zero semantics by Spark version (#5750) #6561 (andygrove, peterxcli)

  • fix: [branch-1.1] fall back from the native Iceberg scan when a nested field was added or renamed (#6543) #6560 (andygrove)

Performance related:

  • perf: optimize spark_base64 in spark-expr #4885 (andygrove)

  • perf: optimize spark_floor (up to 4x faster) #4911 (andygrove)

  • perf: cache Iceberg reflection lookups on the planning path #5222 (andygrove)

  • perf: vectorize integer-to-decimal cast #4939 (andygrove)

  • perf: compute spark_size list lengths with Arrow length kernel #5233 (0lai0)

  • perf: make CometShuffleBenchmark completable and fast #5388 (andygrove)

  • perf: vectorize Map in spark_size via offset buffer #5395 (0lai0)

  • perf: improve ArrowWriter performance for fixed-length vectors #5046 (peterxcli)

  • perf: reuse Arrow IPC compression context across shuffle blocks #5038 (peterxcli)

  • perf: use Arrow cast for decimal rescale check #5440 (peterxcli)

  • perf: bulk copy fixed-width columns in ArrowWriter #5442 (peterxcli)

  • perf: serialize Python input directly from Comet Arrow vectors #5368 (sunchao)

  • perf: reuse per-partition scratch in the shuffle write path #5568 (dwsmith1983)

  • perf: validate shuffle IPC context reuse savings #5727 (peterxcli)

  • perf: slice the child instead of gathering it when unnesting #5667 (andygrove)

  • perf: evaluate posexplode array expressions once per batch #5737 (rich7420)

  • perf: cache expected schemas for remote shuffle decoding #5722 (pingzh)

  • perf: use Arrow cast for date to timestamp NTZ #5735 (peterxcli)

  • perf: reduce allocations when collecting cache statistics #5734 (peterxcli)

  • perf: avoid repeated decimal promotion in expression serialization #5736 (peterxcli)

  • perf: give collect_list and collect_set a native GroupsAccumulator #5803 (andygrove)

  • perf: vectorize the native map lookup behind element_at and GetMapValue #5806 (andygrove)

  • perf: compile user regex patterns once per planned expression #5612 (dwsmith1983)

  • perf: batch the nested-element list hash for flat struct elements #5778 (viirya)

  • perf: optimize map_sort singleton normalization (18x faster) #5887 (viirya)

  • perf: optimize map_sort for multi-entry string maps (up to 3x faster) #5901 (viirya)

  • perf: skip calendar reconstruction in hour/minute/second and dayofweek/weekday #5771 (peterxcli)

  • perf: reuse zstd compression contexts across shuffle blocks #5565 (dwsmith1983)

  • perf: cache parsed plan data across a stage tasks #5615 (dwsmith1983)

  • perf: spill every shuffle partition of a task into one file #5916 (peterxcli)

  • perf: decode shuffle blocks against a cached schema instead of re-parsing per block #5809 (peterxcli)

  • perf: project cached batches by buffer selection, prune on collated strings #5543 (andygrove)

  • perf: make one thread-local access per tracked allocation #6166 (andygrove)

  • perf: reuse quantile summary buffers during merge #4932 (peterxcli)

  • perf: optimize list_extract without defaults using Arrow take #5174 (peterxcli)

  • perf: track decimal overflow without rescanning results #5044 (peterxcli)

Implemented enhancements:

  • feat: support explode_outer #5192 (comphead)

  • feat: support timestampadd and timestampdiff via codegen dispatch #5030 (andygrove)

  • feat: support _metadata constant columns in native Parquet scan #5237 (mbutrovich)

  • feat: add make_interval support (codegen dispatch + native) #5039 (peterxcli)

  • feat: Optionally split the Iceberg V2 write operator into distinct writer and committer operations #4658 (jordepic)

  • feat: detect Iceberg V2 writes and emit fall-back reasons #5298 (jordepic)

  • feat: remove native cast from boolean to decimal #5185 (andygrove)

  • feat: add micro benchmark runner and EC2 guide #5374 (andygrove)

  • feat: build gate + inert wiring for contrib Delta scans [Delta contrib split, part 2] #4952 (schenksj)

  • feat: Support native scans with unprojected Spark 4 VARIANT columns #5377 (sunchao)

  • feat: expose native aggregate spill and memory metrics #5423 (sunchao)

  • feat: support WindowGroupLimitExec #4870 (comphead)

  • feat: add RSS partition writer and task-owned JNI callback (1/n) #5473 (pingzh)

  • feat: native spark_unbase64 kernel #5451 (0lai0)

  • feat: add partition writer destinations to native shuffle plans (2/n) #5476 (pingzh)

  • feat: add destination-aware native shuffle execution (3/n) #5481 (pingzh)

  • feat: bind task-owned RSS callbacks to native shuffle plans (4/n) #5491 (pingzh)

  • feat: add Celeborn shuffle manager and partition pusher (5/n) #5501 (pingzh)

  • feat: complete Celeborn map-side shuffle push lifecycle (6/n) #5513 (pingzh)

  • feat: add experimental native support for in-memory cache, disabled by default #5051 (andygrove)

  • feat: add raw Celeborn native shuffle reader (7/n) #5531 (pingzh)

  • feat: support Map for CreateArray literal #5452 (comphead)

  • feat: route round on float/double through the codegen dispatcher #5600 (andygrove)

  • feat: enable native-only Celeborn shuffle planning (8/n) #5537 (pingzh)

  • feat: support unicode case sensitive field names for reading parquet #5602 (comphead)

  • feat: implement native Iceberg V2 writer via iceberg-rust #5361 (jordepic)

  • feat: Support Iceberg system functions (bucket, truncate, years/months/days/hours) natively #5638 (andygrove)

  • feat: implement regr_slope, regr_intercept, regr_r2, regr_sxx, regr_syy, regr_sxy aggregates #4775 (andygrove)

  • feat: carry VariantType identity through schema serialization #5631 (peterxcli)

  • feat: route unrecognized StaticInvoke and Invoke through the codegen dispatcher #5692 (andygrove)

  • feat: support S3 compliant filesystems #5314 (comphead)

  • feat: add native spark_sequence kernel for integral element types #5614 (0lai0)

  • feat: support nested types as native shuffle hash partitioning keys #5567 (viirya)

  • feat: route next_day and levenshtein collated input through the codegen dispatcher #5720 (0lai0)

  • chore: add_benches hash function aggregators #5730 (coderfender)

  • test: cover Spark-to-Arrow batch conversion edge cases #5713 (peterxcli)

  • feat: normalize marked Variant arrays at the native Parquet boundary #5715 (peterxcli)

  • feat: support native concat_ws with string arrays #5725 (peterxcli)

  • feat: native dynamic filter pushdown for hash join into Parquet scans #5699 (pingzh)

  • feat: enable codegen dispatch for lpad and rpad #5764 (Satyr09)

  • feat: address remaining issues for CreateArray #5766 (comphead)

  • ci: stop non-gating labels from skipping Preflight #5784 (comphead)

  • feat: route abs on interval types through the codegen dispatcher #5622 (kazantsev-maksim)

  • ci: shard Iceberg Spark tests across four runners #5459 (sunchao)

  • test: expand ANSI coverage for round, conv and elt #5799 (rich7420)

  • test: strengthen ANSI exception assertions #5800 (rich7420)

  • chore: Improve network retry configuration for maven and artifact upload #5782 (comphead)

  • ci: gate the Delta contrib build on symbols, not on libcomet size #5827 (andygrove)

  • test: add helpers to assert whether an expression ran natively or via codegen dispatch #5610 (andygrove)

  • feat: support Spark encode expression via codegen dispatch #5037 (andygrove)

  • feat: support native aggregate function mode #4782 (andygrove)

  • test: name the whole dispatched subtree in the decimal promotion assertion #5849 (andygrove)

  • feat: route translate and to_csv through codegen dispatch by default #5032 (andygrove)

  • feat: expose native Parquet scan I/O and read-amplification metrics #5453 (sunchao)

  • ci: move the job routing policy out of ci.yml expressions and into compute-changes.py #5850 (andygrove)

  • feat: support max_by and min_by aggregate expressions #4817 (andygrove)

  • ci: add a Required Checks aggregator job so main can have a required status check #5842 (andygrove)

  • feat: adapt Parquet storage for Variant projection #5794 (peterxcli)

  • ci: move the Spark 3.4/3.5/4.0, Iceberg, macOS and benchmark suites behind a merge queue #5843 (andygrove)

  • feat(iceberg): Iceberg table format V3: apply deletion vector on reads #5853 (mbutrovich)

  • test: restore ANSI array access error coverage #5798 (rich7420)

  • test: cover native memory accounting boundaries #5856 (rich7420)

  • test: cover date maps in Parquet temporal fuzz tests #5877 (rich7420)

  • feat: Enable adaptive partial aggregation for eligible native shuffle plans #5723 (peterxcli)

  • ci: move the Spark 4.1 sql_hive shards behind the merge queue #5871 (andygrove)

  • ci: retry the Maven wrapper bootstrap in every job that calls ./mvnw directly #5852 (andygrove)

  • test: cover slice over expression-produced non-null element arrays (#… #5839 (sam-1112)

  • test: run the libhdfs suite manually instead of in CI #5892 (andygrove)

  • feat: Add Lance contrib build gate #5728 (wirybeaver)

  • chore: drop support for JDK 11 #5897 (manuzhang)

  • test: cover regex expression routing configurations #5917 (rich7420)

  • test: cover round expression routing configurations #5878 (rich7420)

  • test: cover to_csv routing configurations #5921 (rich7420)

  • test: cover interval expression routing configurations #5920 (rich7420)

  • test: cover array and map expression routing configurations #5918 (rich7420)

  • test: make the cancelled Iceberg abort test yield deterministically #5919 (andygrove)

  • chore: deprecate Spark 3.4 rather than removing it in 1.1.0 #5885 (andygrove)

  • test: restore the TPC-H suite’s 2 GiB off-heap budget #5904 (ErikBPF)

  • test: cover string expression routing and native opt-in #5915 (rich7420)

  • chore: run only the cache-writing jobs on push to main #5930 (andygrove)

  • ci: publish nightly SNAPSHOT jars to repository.apache.org #5902 (andygrove)

  • ci: fold the Delta build gate and PyArrow UDF suite into the merge queue tiers #5926 (andygrove)

  • refactor: use Arrow decimal precision validation #5160 (Smallfu666)

  • ci: shrink the pull request tier to the Linux build on the default Spark profile #5939 (andygrove)

  • ci: run the non-default Spark and Iceberg suites nightly instead of in the merge queue #5963 (andygrove)

  • test: cover lpad and rpad routing configurations #5895 (rich7420)

  • test: preserve Spark SQL baselines for ordering-sensitive fixtures #5914 (rich7420)

  • test: cover json_array_length routing configurations #5940 (peterxcli)

  • test: cover SQL fixture metadata and statement parsing #5941 (peterxcli)

  • test: cover crc32 on binary inputs #5942 (peterxcli)

  • test: cover datetime and timezone routing configurations #5950 (rich7420)

  • test: cover collation and predicate routing configurations #5951 (rich7420)

  • test: cover array_intersect routing configurations #5952 (rich7420)

  • test: guard native plan equality against omitted parameters #5953 (rich7420)

  • test: expand replace compatibility regression coverage #5409 (sam-1112)

  • feat: add native allocation accounting for memory observability #5934 (andygrove)

  • feat: run rlike natively by default for Java-equivalent literal patterns #5415 (sam-1112)

  • ci: write large actions/cache entries only on push to main #5973 (andygrove)

  • refactor: extract shared runtime filter components #5937 (pingzh)

  • test: skip the Spark RocksDB state-store suites when Comet is enabled #5987 (comphead)

  • chore: bench nondeterministic and json kernels #5989 (coderfender)

  • feat: add BinaryType support for SortMergeJoin #5928 (xhumanoid)

  • refactor: replace remaining hand-rolled loops with Arrow kernels #5367 (0lai0)

  • feat: support Spark 4 EmptyRelationExec as a native input #5821 (jeffw13)

  • chore: Bench agg (welford) stats #5988 (coderfender)

  • ci: rename the umbrella workflow from CI to Comet CI #6003 (andygrove)

  • ci: bootstrap Maven from the setup actions so every mvnw job is covered #5881 (andygrove)

  • ci: add dev/local-ci.sh to run the Spark SQL and Iceberg suites locally #5974 (comphead)

  • feat: narrow strict floating-point admission for scalar sort keys #5981 (0lai0)

  • feat: route compatible xxhash64 args through SparkXxhash64 #5960 (sam-1112)

  • feat: support string arrays in Spark-to-Comet conversion #5954 (rich7420)

  • feat: run length on binary input natively #5874 (dwsmith1983)

  • chore: drop the redundant width_bucket shim registrations and guard serde uniqueness #5873 (dwsmith1983)

  • feat: hook native Parquet writes into Spark’s WriteFilesExec seam on Spark 4.0+ #5763 (andygrove)

  • feat(iceberg): report native Iceberg scan planning metrics and scan time in the Spark UI #6027 (parthchandra)

  • test: guard page skipping in the native Iceberg scan #6040 (dwsmith1983)

  • test: stop claiming Spark agrees on signed-zero array literals #6055 (cestercian)

  • feat: trace Arrow memory held on the JVM side #6048 (andygrove)

  • feat: unix_timestamp codegen dispatch for strings, and fix pre-epoch fractional rounding #5789 (Satyr09)

  • chore: remove JaCoCo from the build #6084 (andygrove)

  • refactor: compose CometScanRule and CometExecRule into a single CometRule #6082 (andygrove)

  • chore: add a native Iceberg write benchmark #6038 (0lai0)

  • ci: retry the Spark test pre-compile step on resolution failures #6079 (andygrove)

  • ci: focus Miri on unsafe row and hash tests #6072 (rich7420)

  • test: run the Iceberg Spark tests with the native Iceberg writer enabled #5677 (andygrove)

  • feat: support listagg / string_agg aggregate (Spark 4.0+) #4816 (andygrove)

  • ci: run label-triggered CI as a separate workflow #6161 (andygrove)

  • feat: always count native allocations and log executor native memory usage #6162 (andygrove)

  • feat: use DataFusion unnest_outer instead of Comet ListEmptyToNullExpr #6132 (comphead)

  • chore: recommend –force-with-lease when updating PR branches #6171 (mizulun)

  • feat: positional round robin shuffle keyed on a row ordinal #6095 (andygrove)

  • chore: deprecate spark.comet.exec.memoryPool.fraction #6163 (andygrove)

  • feat: add spark.comet.explain.planOnly.enabled #5394 (andygrove)

  • feat: support direct Variant projection in native Parquet scans #5868 (peterxcli)

  • test: cover JSON and cast expression routing #6068 (rich7420)

  • feat: remove memory accounting from Comet’s on-heap mode #6066 (andygrove)

  • test: accept CometHashAggregateExec in the Spark 4.0 CollationSuite hash agg check #6220 (andygrove)

  • feat: per-location credentials for the S3 credential SPI #6031 (snmvaughan)

  • refactor: centralize data type support predicates #5025 (peterxcli)

  • chore: [branch-1.1] change version from 1.1.0-SNAPSHOT to 1.1.0 #6236 (andygrove)

  • feat: [branch-1.1] built-in S3 credential provider adapters for the native Parquet scan (#6023) #6318 (andygrove, parthchandra)

  • chore: [branch-1.1] Add 1.1.0 changelog #6282 (andygrove)

  • test: [branch-1.1] fix flaky mixed field id directory test on Spark 3.4 and 3.5 (#6312) #6527 (andygrove)

  • test: [branch-1.1] cover sliced boolean arrays in explode (#6473) #6492 (andygrove)

  • test: [branch-1.1] backport the Iceberg write report and the mid-write retry test (#6155, #6111) #6307 (andygrove, sam-1112)

  • feat: [branch-1.1] reuse S3 credentials until shortly before their reported expiry (#6509) #6545 (andygrove, snmvaughan)

Documentation updates:

  • docs: add note about run-iceberg-tests label in CI #5247 (mbutrovich)

  • docs: add 1.0.0 TPC-DS benchmark results, remove TPC-H #5284 (mbutrovich)

  • docs: add 1.0.0 changelog to main #5263 (andygrove)

  • docs: correct Spark 4.2 version and CI test status in installation guide #5315 (andygrove)

  • docs: add suggest-native-expression skill for assessing native expression candidates #5348 (andygrove)

  • docs: document C2R cost for wide/nested schemas in tuning guide #5458 (DebadityaHait)

  • docs: update stale interval multiplication note in expressions.md #5518 (peterxcli)

  • docs: add contributing guide link #5521 (Dharan-K)

  • doc: fix benchmark examples for MacOS #5522 (xhumanoid)

  • docs: add comet meeting link #5621 (coderfender)

  • docs: add a contributor guide page for CI and the merge queue #5863 (andygrove)

  • docs: add contributor guide page on memory management #5933 (andygrove)

  • docs: how to check the scheduled CI runs are actually running #6000 (andygrove)

  • docs: explain allocator hazards and diagram where memory is allocated #6014 (andygrove)

  • docs: pre-render mermaid diagrams to SVG at build time #6021 (andygrove)

  • docs: split the PR review skill by area and correct the shuffle contributor docs #6018 (andygrove)

  • docs: correct the stale range-partitioning strict floating-point rule #6049 (andygrove)

  • docs: recommend setting spark.executor.memoryOverhead alongside off-heap memory #6051 (andygrove)

  • docs: exempt testing-category and internal configs from the versioning policy #6089 (andygrove)

  • docs: make doc changes weekly comet sync #6160 (coderfender)

  • docs: correct the plugin and shuffle sections of the plugin overview #6197 (andygrove)

  • docs: fix stale and missing native Iceberg write details #6150 (andygrove)

  • docs: add an Iceberg writes contributor guide and review skill #6149 (andygrove)

  • docs: correct which operators run separate native plans in a task #6194 (andygrove)

  • docs: update the user guide for the 1.1.0 release #6168 (andygrove)

  • docs: [branch-1.1] backport the 1.1.0 upgrade notes and tuning guide updates (#6237, #6244, #6248) #6265 (andygrove)

  • ci: [branch-1.1] run every tier on release-branch pull requests, and the full suite before an RC (#6218) #6284 (andygrove)

  • docs: [branch-1.1] generate release docs for 1.1.0 #6344 (andygrove)

  • docs: [branch-1.1] credit co-authors in the 1.1.0 changelog #6359 (andygrove)

Other:

  • chore: start 1.1.0 development #5242 (andygrove)

  • refactor: replace hand-coded rollup of expression fallback reasons onto operators #5236 (andygrove)

  • test: rename CometCastSuite to CometNativeCastSuite #5268 (andygrove)

  • refactor: use arity helper for Int to Decimal128 reinterpretation #5193 (0lai0)

  • test: add guard for Iceberg version on Variant fallback test #5278 (mbutrovich)

  • chore: update documentation links for 1.0.0 release #5290 (andygrove)

  • chore(deps): bump reqsign-core from 3.2.0 to 3.2.1 in /native in the all-other-cargo-deps group #5287 (dependabot[bot])

  • chore(deps): bump object_store_opendal from 0.57.0 to 0.58.0 in /native #5289 (dependabot[bot])

  • chore(deps): bump the codeql-actions group with 2 updates #5286 (dependabot[bot])

  • bug: Revert “chore(deps): bump object_store_opendal from 0.57.0 to 0.58.0 in /native” #5332 (coderfender)

  • test: cover the narrowing direction of cast_and_stamp_schema #5285 (andygrove)

  • chore(deps): bump opendal from 0.57.0 to 0.58.1 in /native #5324 (manuzhang)

  • chore: respect Cargo parallelism settings for native release builds #5344 (pingzh)

  • test: make expression benchmark harness fair and reproducible #5371 (andygrove)

  • chore(deps): bump the all-other-cargo-deps group in /native with 5 updates #5360 (dependabot[bot])

  • test: fix vacuous signed-zero coverage in SQL file tests #5393 (sam-1112)

  • chore: fix clippy warnings for Rust 1.98 #5400 (ywskycn)

  • chore(deps): bump the codeql-actions group with 2 updates #5405 (dependabot[bot])

  • test: strengthen signed-zero SQL assertions #5404 (sunchao)

  • chore(deps): bump actions/checkout from 6 to 7 #5406 (dependabot[bot])

  • test: add expression fallback-invariance suite #5329 (4ktLuffy)

  • refactor: delegate ANSI integer arithmetic to arrow checked kernels #5280 (kazantsev-maksim)

  • test: docs and test-coverage hardening for native collect_list / collect_set #5055 (andygrove)

  • test: deduplicate Throwable cause-chain traversal #5441 (peterxcli)

  • refactor: use Arrow timezone type #5129 (Hashim1999164)

  • chore: add md formatting to make format #5460 (comphead)

  • test: fail unexpected vacuous fallback-invariance checks #5417 (sunchao)

  • ci: cache Maven distributions and retry bootstrap downloads #5422 (sunchao)

  • chore: add math benches #5479 (coderfender)

  • chore: Bench string functions #5492 (coderfender)

  • chore(deps): bump actions/setup-java from 5 to 6 #5524 (dependabot[bot])

  • chore(deps): bump the codeql-actions group with 2 updates #5523 (dependabot[bot])

  • chore: bench additional math function #5520 (coderfender)

  • chore: remove redundant condition and stray println in shuffle code #5562 (viirya)

  • test: cover struct data columns in native shuffle #5564 (viirya)

  • chore: rename .claude directory to vendor-neutral .ai #5594 (andygrove)

  • chore: bench additional scalar functions #5598 (coderfender)

  • test: verify Celeborn reflection compatibility #5604 (pingzh)

  • test: enable native path in lower/upper_enabled sql fixtures #5619 (cestercian)

  • chore: add benches datetime #5620 (coderfender)

  • chore: move dead and defensive serde guards out of convert #5595 (andygrove)

  • test: exercise the Iceberg write split-operator plan in Iceberg’s own suites #5640 (andygrove)

  • chore(deps): bump the codeql-actions group with 2 updates #5669 (dependabot[bot])

  • test: add explode operator microbenchmark #5381 (andygrove)

  • deps: bump DataFusion 55.0 and Arrow/Parquet 59.2 #5262 (mbutrovich)

  • chore: add benches array functions #5700 (coderfender)

  • test: restore Comet coverage for recursive HAVING and ORDER BY #5755 (rich7420)

  • test: restore Parquet V2 writer and delta encoding coverage #5760 (rich7420)

  • chore: Add benches for datetime funcs #5767 (coderfender)

  • bench: add a benchmark for the Spark hash kernels #5765 (viirya)

  • test: restore Spark 4.1 Variant shredding suites #5745 (rich7420)

  • refactor: share one helper for pushing a struct’s null mask into its children #5769 (viirya)

  • ci: label pull requests by changed paths and title prefix #5762 (dwsmith1983)

  • test: cover ambiguous exact nested Parquet field matches #5751 (peterxcli)

  • bench: measure nested types as native shuffle hash partitioning keys #5788 (viirya)

  • deps: bump to datafusion 55.1.0 #5865 (comphead)

  • bench: isolate map normalization and nested key hashing #5822 (viirya)

  • bench: add a shuffle read benchmark covering the per-block schema parse #5805 (peterxcli)

  • chore(deps): bump actions/setup-java from 4 to 6 #6012 (dependabot[bot])

  • chore(deps): bump the codeql-actions group with 2 updates #6011 (dependabot[bot])

  • deps: bump the iceberg-rust pin to bb1e4a4 and document why it is pinned #6094 (andygrove)

Credits#

Thank you to everyone who contributed to this release. Here is a breakdown of commits (PRs merged) per contributor. A PR with commits from more than one person counts for each of them.

   147	Andy Grove
    62	Peter Lee
    26	Chao Sun
    26	KUAN-HAO HUANG
    19	Ping Zhang
    15	Oleks V
    14	dustin
    13	Bhargava Vadlamani
    13	Liang-Chi Hsieh
    11	dependabot[bot]
    10	ChenChen Lai
     8	sam-1112
     6	Matt Butrovich
     5	Han-Yin Chang
     5	Parth Chandra
     4	Erik Bogado
     4	Steve Vaughan
     4	Wei Yan
     3	Jordan Epstein
     3	Manu Zhang
     2	Alexey
     2	Cestercian
     2	Daipayan Mukherjee
     2	Feng Zhang
     2	Kazantsev Maksim
     2	YuLun Mao
     2	liupoyi-1031
     1	Dharan-K
     1	Hashim Khan
     1	Henos D
     1	Jeff Wang
     1	LinSimon-901101
     1	Michael Taranov
     1	Scott Schenkein
     1	Stefan Wang
     1	Unik Dahal
     1	Viacheslav Inozemtsev
     1	Xuanyi Li
     1	debaditya
     1	divyank sameer shah
     1	hsiang-c

Thank you also to everyone who contributed in other ways such as filing issues, reviewing PRs, and providing feedback on this release.