DataFusion Comet 1.1.0 Changelog#
This release consists of 416 commits from 41 contributors. See credits at the end of this changelog for more information.
Fixed bugs:
fix: skip null slots when checking overflow in unary negation #5162 (Smallfu666)
fix: surface next_day and make_date ANSI errors as Spark exceptions #5167 (peterxcli)
fix: normalize nested field nullability in ShuffleScanExec and ExpandExec #5138 (andygrove)
fix: propagate the Spark task ClassLoader to JVM UDF calls #5282 (andygrove)
fix: avoid duplicate CheckOverflow evaluation for decimal division #5225 (peterxcli)
fix: match Spark whitespace trimming in to_time and try_to_time #5364 (sunchao)
fix: preserve Catalyst nullability and field IDs in native Parquet writes #5369 (sunchao)
fix: guard against silent fail_on_error loss in scalar wiring (#5074) #5359 (sam-1112)
fix: preserve Spark semantics for dictionary-encoded Parquet inputs and reject dictionary targets #5234 (peterxcli)
fix: canonicalize NaN in flat arrays_overlap float keys #5376 (sunchao)
fix: report native shuffle write metrics accurately #5370 (sunchao)
fix: Native shuffle fails with a 2GB task serialization OOM on jobs with many partitions #5392 (parthchandra)
fix: format NativeMemoryConsumer id in toString #5398 (ywskycn)
fix: ignore reader-side parquet.hadoop.vectored.io.enabled in Iceberg native-write detection #5410 (snmvaughan)
fix: support map casts with NullType elements #5045 (peterxcli)
fix: distinguish reflection failure from absent accessor in IcebergReflection #5412 (unikdahal)
fix: use current shuffle config in aggregate test #5439 (peterxcli)
fix: support wide years in native make_date #5443 (peterxcli)
fix: report native child spill metrics in shuffle tasks #5445 (sunchao)
fix: expose native memory usage to Spark #5408 (ywskycn)
fix: release native shuffle reservation after spill failure #5461 (peterxcli)
fix: track peak memory before JVM shuffle spill #5463 (peterxcli)
fix: RAII for memory pool registration #5464 (peterxcli)
fix: tighten RSS writer visibility and JNI array limits #5475 (pingzh)
fix: report native operator spill metrics in Spark task metrics for non-shuffle stages #5497 (peterxcli)
fix: report bounded shuffle allocator memory usage #5516 (ywskycn)
fix: use Spark type names in ANSI abs overflow errors #5357 (Smallfu666)
fix: make collect_list/collect_set argument coercion a normalization barrier #5159 (andygrove)
fix: make task-shared memory pool as ref-counted RAII guard #5494 (peterxcli)
fix: accept UTC timezone aliases in Python Arrow input #5556 (sunchao)
fix: normalize scalar float sort and window rank keys #5469 (sunchao)
fix: make CometExplodeExec respect batch size #5362 (andygrove)
fix: align Spark 4.2 Python worker configuration #5561 (sunchao)
fix: prevent memory leak after failed Arrow vector import #5539 (1fanwang)
fix: preserve Arrow Field metadata across C Data exports #5552 (peterxcli)
fix: rebase map offsets in mapsort so sliced maps do not overrun entries #5630 (viirya)
fix: scope Celeborn bootstrap hooks to Comet clients #5627 (pingzh)
fix: skip codegen dispatcher null short-circuit when a foldable subtree can raise #5623 (andygrove)
fix: accept dictionary encodings in remote shuffle #5650 (pingzh)
fix: make CometDiskBlockWriter spill registry per-task instead of executor-global #5493 (peterxcli)
fix: Bump iceberg-rust so native Iceberg writes URL-escape partition paths #5651 (andygrove)
fix(celeborn): reject unsafe native push completion tracking #5665 (pingzh)
fix: match Spark’s null short-circuiting in array_join and enable it natively #5558 (Visorgood)
fix: make native shuffle spill metrics independent of input batching #5628 (sunchao)
fix: expand object store option references, uniquify constant metadata names, drop dead parquet JNI #5653 (dwsmith1983)
fix: match Spark’s ANSI bound check for float/double to integral casts #5683 (peterxcli)
fix: return NULL from rpad/lpad when the length column is NULL instead of panicking #5680 (peterxcli)
fix: fall back for concat_ws with array arguments instead of failing natively #5679 (peterxcli)
fix: prevent silent overflow when reading Parquet TIMESTAMP_MILLIS values #5177 (peterxcli)
fix: recover native Celeborn shuffle from oversized rows #5668 (pingzh)
fix: report native shuffle read metrics #5554 (peterxcli)
fix: correctly rounded decimal to double/float cast matching BigDecimal.doubleValue/floatValue #5684 (peterxcli)
fix: restore columnar transitions under the native Iceberg write #5696 (andygrove)
fix: make columnar-to-row benchmarks exercise Comet #5718 (rich7420)
fix: distinguish “nothing spilled” from a spill backend with no local path #5726 (andygrove)
fix: propagate Arrow array copy errors #5747 (rich7420)
fix: keep the dictionary hash fast path off nested and reseeded buffers #5757 (viirya)
fix: preserve ANSI errors for rejected TIMESTAMP_NTZ casts #5752 (peterxcli)
fix: apply the parent struct’s null mask before hashing its fields #5754 (viirya)
fix: read Iceberg tables partitioned by an unknown transform #5759 (andygrove)
fix: native Iceberg write panics on an evolved partition spec and on a timestamptz partition path #5729 (andygrove)
fix: let AQE optimize queries over Comet caches #5733 (peterxcli)
fix: dispatch Iceberg system functions wrapped as ApplyFunctionExpression #5773 (andygrove)
fix: attach tokio runtime threads to the JVM as daemon threads #5748 (zhangfengcdt)
fix: check nested TIMESTAMP_MILLIS overflow in unfiltered scans #5740 (peterxcli)
fix: match iceberg-java’s exception for unclustered input to a clustered Iceberg write #5779 (andygrove)
fix: enable FIRST/LAST partial merge #5041 (peterxcli)
fix: preserve aggregate result identity during exchange reuse #5470 (sunchao)
fix: ignore structural tags when lifting expression coverage #5471 (sunchao)
fix: align string to timestamp parsing with Spark’s segment rules #5682 (peterxcli)
fix: roll native Iceberg data files on iceberg-java’s 1000-row grid #5780 (andygrove)
fix: decide libhdfs routing from the scheme as written #5825 (comphead)
fix: list a fanout Iceberg write’s data files in a stable order #5810 (andygrove)
fix: isolate object-store registration by backend and configuration #5503 (sunchao)
fix: Delete completed tasks’ data files when an Iceberg write job fails #5663 (andygrove)
fix: apply Spark’s Parquet conversion rules to nested struct/list/map fields #5681 (peterxcli)
fix: preserve nulls for Boolean/Byte/Short/Integer columns in FuzzDataGenerator #5855 (Smallfu666)
fix: Accept explicit positive years in timestamp casts #5858 (peterxcli)
fix: decline structs with duplicate field names before they reach Java Arrow #5866 (dwsmith1983)
fix: explain ObjectHashAggregate fallback when Comet shuffle is disabled #5746 (0lai0)
fix: preserve join and generator semantics in plan identity #5828 (ErikBPF)
fix: propagate Parquet field-name folding failures #5845 (sunchao)
fix: decode dictionary input for PyArrow UDFs #5560 (sunchao)
fix: size JVM shuffle pointer array growth from the array, not the data pages #5907 (andygrove)
fix: render float and double Iceberg partition values like iceberg-java #5840 (andygrove)
fix: align time parsing and native second extraction with Spark #5738 (peterxcli)
fix: normalize floating-point values in native collect_set #5166 (peterxcli)
fix: keep Iceberg complex null checks on native scans #5732 (ErikBPF)
fix: decode invalid UTF-8 at the JVM to native FFI import boundary #5310 (manuzhang, andygrove)
fix: fall back when a struct repeats a Parquet field id #6004 (comphead)
fix: normalize signed zero in nested float array comparisons #5235 (divyankshah)
fix: fall back to Spark for native Iceberg writes to gs:// through HadoopFileIO #5935 (zhangfengcdt)
fix: gate the regr_r2 degenerate-case swap on the Spark patch release #6042 (dwsmith1983)
fix: include ABFS container in object store cache key #5053 (peterxcli)
fix: preserve Spark row index read errors #6046 (liupoyi-1031)
fix: remove the ineffective spark.executor.memoryOverhead adjustment from the driver plugin #6054 (andygrove)
fix: count each memory pool once in analyze_trace #5991 (andygrove)
fix: revert unsafe partial aggregates after final fallback #5421 (sunchao)
fix: restore the site’s mermaid diagrams and make a dropped one fail CI #6064 (andygrove)
fix: write cached batches to the schema width, not the batch width #6090 (andygrove)
fix(iceberg): guard native Iceberg scan driver-metric double-post, add metrics docs and tests #6085 (parthchandra)
fix: let decimal SUM recover from an intermediate overflow like Spark #6041 (dwsmith1983)
fix: Nested floating-point IN membership does not match Spark for signed zero #6073 (mizulun)
fix: let Comet memory pools overcommit on grow instead of panicking #6128 (andygrove)
fix: correct two nightly test failures on Spark 3.4 and 4.2 #6156 (andygrove)
fix: support null calendar interval literals #5133 (peterxcli)
fix: wrap Iceberg split-write failures the way Spark does when abort fails #6153 (andygrove)
fix: support Utf8/LargeUtf8/Utf8View in native RLike without panicking #5215 (sam-1112)
fix: normalize noncanonical NaN literals in comparisons #5472 (sunchao)
fix: preserve map field metadata and honor target sorted flag in cast_map_to_map #5227 (Smallfu666)
fix: reject duplicate Parquet field names before decoding #5786 (ErikBPF)
fix: preserve Parquet conversion errors during join filtering #6067 (pingzh)
fix: plan Iceberg writes with Spark’s own operator when Comet is disabled #6151 (andygrove)
fix: read shuffle write buffer, spill limit and off-heap sizes in bytes #6191 (andygrove)
fix: bound shuffle schema cache retention and preserve eviction order #6098 (sunchao)
fix: do not run Comet in on-heap mode without
spark.comet.exec.onHeap.enabled#6195 (andygrove)fix: remove misleading native opt-in for dispatch-only datetime expressions #6182 (LinSimon-901101)
fix: support CalendarIntervalType hashing #5135 (peterxcli)
fix: report
Inputcolumn when native Iceberg scan is enabled or native shuffle is enabled #5265 (hsiang-c)fix: skip the executor memory overhead warning when a factor is set or in local mode #6198 (andygrove)
fix: keep operators above a cached relation native after the AQE re-plan #6208 (andygrove)
fix: preserve current AQE logical-stage links on Comet operators #5483 (sunchao)
fix: check each fair_unified reservation against its own share #6205 (andygrove)
fix(iceberg): don’t push transform residuals as their source column, fail on residual errors #6154 (andygrove)
fix: match Spark’s duplicate field and field id semantics in parquet field lookup #5654 (dwsmith1983)
fix: park the native scan loop instead of busy-polling while waiting on native I/O #6219 (andygrove, mixermt)
fix: refresh S3 policy locations when a location’s credential fails #6223 (andygrove, snmvaughan)
fix: inject Comet session extension rules only once per session #6230 (andygrove)
fix: [branch-1.1] reject a file without field ids at any depth whether or not id matching is on (#6116) #6266 (andygrove, dwsmith1983)
fix: [branch-1.1] decline native Iceberg writes with a custom location provider (#6216) #6305 (andygrove, liupoyi-1031)
fix: [branch-1.1] Native S3 scan on EKS/IRSA turns a transient STS throttle into a hard 403 storm (#6025) #6323 (andygrove, parthchandra)
fix: [branch-1.1] log partial memory grants at DEBUG and drop the memory usage dump (#6269) #6346 (andygrove)
fix: [branch-1.1] zero sliced boolean offsets at every level before exporting to the JVM (#6339) #6449 (andygrove)
fix: [branch-1.1] build with Rust 1.99, which deprecates the legacy f64 constants and fetch_update (#6507) #6526 (andygrove)
fix: [branch-1.1] make adaptive aggregation skipping opt-in (#6474) #6488 (andygrove)
fix: [branch-1.1] fall back to Spark for _metadata.file_block_start and file_block_length (#6510) #6511 (andygrove)
fix: [branch-1.1] fall back for rank limits over nested float keys (#6468) #6487 (andygrove)
fix: [branch-1.1] fall back for incompatible regression aggregates (#6451) #6489 (andygrove)
fix: [branch-1.1] coerce native IF branches to a common type (#6458) #6491 (andygrove)
fix: [branch-1.1] match pre-epoch Iceberg temporal rounding (#6456) #6486 (andygrove)
fix: [branch-1.1] defer Parquet conversion errors until a row group is decoded, as Spark does (#6515) #6540 (andygrove)
fix: [branch-1.1] dispatch StaticInvoke and Invoke only into Spark’s own classes (#6542) #6554 (andygrove)
fix: [branch-1.1] gate array distinct and union signed-zero semantics by Spark version (#5750) #6561 (andygrove, peterxcli)
fix: [branch-1.1] fall back from the native Iceberg scan when a nested field was added or renamed (#6543) #6560 (andygrove)
Performance related:
perf: optimize spark_base64 in spark-expr #4885 (andygrove)
perf: optimize
spark_floor(up to 4x faster) #4911 (andygrove)perf: cache Iceberg reflection lookups on the planning path #5222 (andygrove)
perf: vectorize integer-to-decimal cast #4939 (andygrove)
perf: compute spark_size list lengths with Arrow length kernel #5233 (0lai0)
perf: make CometShuffleBenchmark completable and fast #5388 (andygrove)
perf: vectorize Map in spark_size via offset buffer #5395 (0lai0)
perf: improve ArrowWriter performance for fixed-length vectors #5046 (peterxcli)
perf: reuse Arrow IPC compression context across shuffle blocks #5038 (peterxcli)
perf: use Arrow cast for decimal rescale check #5440 (peterxcli)
perf: bulk copy fixed-width columns in ArrowWriter #5442 (peterxcli)
perf: serialize Python input directly from Comet Arrow vectors #5368 (sunchao)
perf: reuse per-partition scratch in the shuffle write path #5568 (dwsmith1983)
perf: validate shuffle IPC context reuse savings #5727 (peterxcli)
perf: slice the child instead of gathering it when unnesting #5667 (andygrove)
perf: evaluate posexplode array expressions once per batch #5737 (rich7420)
perf: cache expected schemas for remote shuffle decoding #5722 (pingzh)
perf: use Arrow cast for date to timestamp NTZ #5735 (peterxcli)
perf: reduce allocations when collecting cache statistics #5734 (peterxcli)
perf: avoid repeated decimal promotion in expression serialization #5736 (peterxcli)
perf: give collect_list and collect_set a native GroupsAccumulator #5803 (andygrove)
perf: vectorize the native map lookup behind element_at and GetMapValue #5806 (andygrove)
perf: compile user regex patterns once per planned expression #5612 (dwsmith1983)
perf: batch the nested-element list hash for flat struct elements #5778 (viirya)
perf: optimize
map_sortsingleton normalization (18x faster) #5887 (viirya)perf: optimize
map_sortfor multi-entry string maps (up to 3x faster) #5901 (viirya)perf: skip calendar reconstruction in hour/minute/second and dayofweek/weekday #5771 (peterxcli)
perf: reuse zstd compression contexts across shuffle blocks #5565 (dwsmith1983)
perf: cache parsed plan data across a stage tasks #5615 (dwsmith1983)
perf: spill every shuffle partition of a task into one file #5916 (peterxcli)
perf: decode shuffle blocks against a cached schema instead of re-parsing per block #5809 (peterxcli)
perf: project cached batches by buffer selection, prune on collated strings #5543 (andygrove)
perf: make one thread-local access per tracked allocation #6166 (andygrove)
perf: reuse quantile summary buffers during merge #4932 (peterxcli)
perf: optimize list_extract without defaults using Arrow take #5174 (peterxcli)
perf: track decimal overflow without rescanning results #5044 (peterxcli)
Implemented enhancements:
feat: support
explode_outer#5192 (comphead)feat: support timestampadd and timestampdiff via codegen dispatch #5030 (andygrove)
feat: support _metadata constant columns in native Parquet scan #5237 (mbutrovich)
feat: add
make_intervalsupport (codegen dispatch + native) #5039 (peterxcli)feat: Optionally split the Iceberg V2 write operator into distinct writer and committer operations #4658 (jordepic)
feat: detect Iceberg V2 writes and emit fall-back reasons #5298 (jordepic)
feat: remove native cast from boolean to decimal #5185 (andygrove)
feat: add micro benchmark runner and EC2 guide #5374 (andygrove)
feat: build gate + inert wiring for contrib Delta scans [Delta contrib split, part 2] #4952 (schenksj)
feat: Support native scans with unprojected Spark 4 VARIANT columns #5377 (sunchao)
feat: expose native aggregate spill and memory metrics #5423 (sunchao)
feat: support
WindowGroupLimitExec#4870 (comphead)feat: add RSS partition writer and task-owned JNI callback (1/n) #5473 (pingzh)
feat: native
spark_unbase64kernel #5451 (0lai0)feat: add partition writer destinations to native shuffle plans (2/n) #5476 (pingzh)
feat: add destination-aware native shuffle execution (3/n) #5481 (pingzh)
feat: bind task-owned RSS callbacks to native shuffle plans (4/n) #5491 (pingzh)
feat: add Celeborn shuffle manager and partition pusher (5/n) #5501 (pingzh)
feat: complete Celeborn map-side shuffle push lifecycle (6/n) #5513 (pingzh)
feat: add experimental native support for in-memory cache, disabled by default #5051 (andygrove)
feat: add raw Celeborn native shuffle reader (7/n) #5531 (pingzh)
feat: support Map for
CreateArrayliteral #5452 (comphead)feat: route
roundon float/double through the codegen dispatcher #5600 (andygrove)feat: enable native-only Celeborn shuffle planning (8/n) #5537 (pingzh)
feat: support unicode case sensitive field names for reading parquet #5602 (comphead)
feat: implement native Iceberg V2 writer via iceberg-rust #5361 (jordepic)
feat: Support Iceberg system functions (bucket, truncate, years/months/days/hours) natively #5638 (andygrove)
feat: implement regr_slope, regr_intercept, regr_r2, regr_sxx, regr_syy, regr_sxy aggregates #4775 (andygrove)
feat: carry VariantType identity through schema serialization #5631 (peterxcli)
feat: route unrecognized StaticInvoke and Invoke through the codegen dispatcher #5692 (andygrove)
feat: support S3 compliant filesystems #5314 (comphead)
feat: add native spark_sequence kernel for integral element types #5614 (0lai0)
feat: support nested types as native shuffle hash partitioning keys #5567 (viirya)
feat: route next_day and levenshtein collated input through the codegen dispatcher #5720 (0lai0)
chore: add_benches hash function aggregators #5730 (coderfender)
test: cover Spark-to-Arrow batch conversion edge cases #5713 (peterxcli)
feat: normalize marked Variant arrays at the native Parquet boundary #5715 (peterxcli)
feat: support native concat_ws with string arrays #5725 (peterxcli)
feat: native dynamic filter pushdown for hash join into Parquet scans #5699 (pingzh)
feat: enable codegen dispatch for
lpadandrpad#5764 (Satyr09)feat: address remaining issues for
CreateArray#5766 (comphead)ci: stop non-gating labels from skipping Preflight #5784 (comphead)
feat: route
abson interval types through the codegen dispatcher #5622 (kazantsev-maksim)ci: shard Iceberg Spark tests across four runners #5459 (sunchao)
test: expand ANSI coverage for round, conv and elt #5799 (rich7420)
test: strengthen ANSI exception assertions #5800 (rich7420)
chore: Improve network retry configuration for maven and artifact upload #5782 (comphead)
ci: gate the Delta contrib build on symbols, not on libcomet size #5827 (andygrove)
test: add helpers to assert whether an expression ran natively or via codegen dispatch #5610 (andygrove)
feat: support Spark encode expression via codegen dispatch #5037 (andygrove)
feat: support native aggregate function
mode#4782 (andygrove)test: name the whole dispatched subtree in the decimal promotion assertion #5849 (andygrove)
feat: route translate and to_csv through codegen dispatch by default #5032 (andygrove)
feat: expose native Parquet scan I/O and read-amplification metrics #5453 (sunchao)
ci: move the job routing policy out of ci.yml expressions and into compute-changes.py #5850 (andygrove)
feat: support max_by and min_by aggregate expressions #4817 (andygrove)
ci: add a Required Checks aggregator job so main can have a required status check #5842 (andygrove)
feat: adapt Parquet storage for Variant projection #5794 (peterxcli)
ci: move the Spark 3.4/3.5/4.0, Iceberg, macOS and benchmark suites behind a merge queue #5843 (andygrove)
feat(iceberg): Iceberg table format V3: apply deletion vector on reads #5853 (mbutrovich)
test: restore ANSI array access error coverage #5798 (rich7420)
test: cover native memory accounting boundaries #5856 (rich7420)
test: cover date maps in Parquet temporal fuzz tests #5877 (rich7420)
feat: Enable adaptive partial aggregation for eligible native shuffle plans #5723 (peterxcli)
ci: move the Spark 4.1 sql_hive shards behind the merge queue #5871 (andygrove)
ci: retry the Maven wrapper bootstrap in every job that calls ./mvnw directly #5852 (andygrove)
test: cover slice over expression-produced non-null element arrays (#… #5839 (sam-1112)
test: run the libhdfs suite manually instead of in CI #5892 (andygrove)
feat: Add Lance contrib build gate #5728 (wirybeaver)
chore: drop support for JDK 11 #5897 (manuzhang)
test: cover regex expression routing configurations #5917 (rich7420)
test: cover round expression routing configurations #5878 (rich7420)
test: cover to_csv routing configurations #5921 (rich7420)
test: cover interval expression routing configurations #5920 (rich7420)
test: cover array and map expression routing configurations #5918 (rich7420)
test: make the cancelled Iceberg abort test yield deterministically #5919 (andygrove)
chore: deprecate Spark 3.4 rather than removing it in 1.1.0 #5885 (andygrove)
test: restore the TPC-H suite’s 2 GiB off-heap budget #5904 (ErikBPF)
test: cover string expression routing and native opt-in #5915 (rich7420)
chore: run only the cache-writing jobs on push to main #5930 (andygrove)
ci: publish nightly SNAPSHOT jars to repository.apache.org #5902 (andygrove)
ci: fold the Delta build gate and PyArrow UDF suite into the merge queue tiers #5926 (andygrove)
refactor: use Arrow decimal precision validation #5160 (Smallfu666)
ci: shrink the pull request tier to the Linux build on the default Spark profile #5939 (andygrove)
ci: run the non-default Spark and Iceberg suites nightly instead of in the merge queue #5963 (andygrove)
test: cover lpad and rpad routing configurations #5895 (rich7420)
test: preserve Spark SQL baselines for ordering-sensitive fixtures #5914 (rich7420)
test: cover json_array_length routing configurations #5940 (peterxcli)
test: cover SQL fixture metadata and statement parsing #5941 (peterxcli)
test: cover crc32 on binary inputs #5942 (peterxcli)
test: cover datetime and timezone routing configurations #5950 (rich7420)
test: cover collation and predicate routing configurations #5951 (rich7420)
test: cover array_intersect routing configurations #5952 (rich7420)
test: guard native plan equality against omitted parameters #5953 (rich7420)
test: expand replace compatibility regression coverage #5409 (sam-1112)
feat: add native allocation accounting for memory observability #5934 (andygrove)
feat: run rlike natively by default for Java-equivalent literal patterns #5415 (sam-1112)
ci: write large actions/cache entries only on push to main #5973 (andygrove)
refactor: extract shared runtime filter components #5937 (pingzh)
test: skip the Spark RocksDB state-store suites when Comet is enabled #5987 (comphead)
chore: bench nondeterministic and json kernels #5989 (coderfender)
feat: add BinaryType support for SortMergeJoin #5928 (xhumanoid)
refactor: replace remaining hand-rolled loops with Arrow kernels #5367 (0lai0)
feat: support Spark 4 EmptyRelationExec as a native input #5821 (jeffw13)
chore: Bench agg (welford) stats #5988 (coderfender)
ci: rename the umbrella workflow from CI to Comet CI #6003 (andygrove)
ci: bootstrap Maven from the setup actions so every mvnw job is covered #5881 (andygrove)
ci: add dev/local-ci.sh to run the Spark SQL and Iceberg suites locally #5974 (comphead)
feat: narrow strict floating-point admission for scalar sort keys #5981 (0lai0)
feat: route compatible xxhash64 args through SparkXxhash64 #5960 (sam-1112)
feat: support string arrays in Spark-to-Comet conversion #5954 (rich7420)
feat: run length on binary input natively #5874 (dwsmith1983)
chore: drop the redundant width_bucket shim registrations and guard serde uniqueness #5873 (dwsmith1983)
feat: hook native Parquet writes into Spark’s WriteFilesExec seam on Spark 4.0+ #5763 (andygrove)
feat(iceberg): report native Iceberg scan planning metrics and scan time in the Spark UI #6027 (parthchandra)
test: guard page skipping in the native Iceberg scan #6040 (dwsmith1983)
test: stop claiming Spark agrees on signed-zero array literals #6055 (cestercian)
feat: trace Arrow memory held on the JVM side #6048 (andygrove)
feat: unix_timestamp codegen dispatch for strings, and fix pre-epoch fractional rounding #5789 (Satyr09)
chore: remove JaCoCo from the build #6084 (andygrove)
refactor: compose CometScanRule and CometExecRule into a single CometRule #6082 (andygrove)
chore: add a native Iceberg write benchmark #6038 (0lai0)
ci: retry the Spark test pre-compile step on resolution failures #6079 (andygrove)
ci: focus Miri on unsafe row and hash tests #6072 (rich7420)
test: run the Iceberg Spark tests with the native Iceberg writer enabled #5677 (andygrove)
feat: support listagg / string_agg aggregate (Spark 4.0+) #4816 (andygrove)
ci: run label-triggered CI as a separate workflow #6161 (andygrove)
feat: always count native allocations and log executor native memory usage #6162 (andygrove)
feat: use DataFusion unnest_outer instead of Comet ListEmptyToNullExpr #6132 (comphead)
chore: recommend –force-with-lease when updating PR branches #6171 (mizulun)
feat: positional round robin shuffle keyed on a row ordinal #6095 (andygrove)
chore: deprecate spark.comet.exec.memoryPool.fraction #6163 (andygrove)
feat: add spark.comet.explain.planOnly.enabled #5394 (andygrove)
feat: support direct Variant projection in native Parquet scans #5868 (peterxcli)
test: cover JSON and cast expression routing #6068 (rich7420)
feat: remove memory accounting from Comet’s on-heap mode #6066 (andygrove)
test: accept CometHashAggregateExec in the Spark 4.0 CollationSuite hash agg check #6220 (andygrove)
feat: per-location credentials for the S3 credential SPI #6031 (snmvaughan)
refactor: centralize data type support predicates #5025 (peterxcli)
chore: [branch-1.1] change version from 1.1.0-SNAPSHOT to 1.1.0 #6236 (andygrove)
feat: [branch-1.1] built-in S3 credential provider adapters for the native Parquet scan (#6023) #6318 (andygrove, parthchandra)
chore: [branch-1.1] Add 1.1.0 changelog #6282 (andygrove)
test: [branch-1.1] fix flaky mixed field id directory test on Spark 3.4 and 3.5 (#6312) #6527 (andygrove)
test: [branch-1.1] cover sliced boolean arrays in explode (#6473) #6492 (andygrove)
test: [branch-1.1] backport the Iceberg write report and the mid-write retry test (#6155, #6111) #6307 (andygrove, sam-1112)
feat: [branch-1.1] reuse S3 credentials until shortly before their reported expiry (#6509) #6545 (andygrove, snmvaughan)
Documentation updates:
docs: add note about run-iceberg-tests label in CI #5247 (mbutrovich)
docs: add 1.0.0 TPC-DS benchmark results, remove TPC-H #5284 (mbutrovich)
docs: add 1.0.0 changelog to main #5263 (andygrove)
docs: correct Spark 4.2 version and CI test status in installation guide #5315 (andygrove)
docs: add suggest-native-expression skill for assessing native expression candidates #5348 (andygrove)
docs: document C2R cost for wide/nested schemas in tuning guide #5458 (DebadityaHait)
docs: update stale interval multiplication note in expressions.md #5518 (peterxcli)
docs: add contributing guide link #5521 (Dharan-K)
doc: fix benchmark examples for MacOS #5522 (xhumanoid)
docs: add comet meeting link #5621 (coderfender)
docs: add a contributor guide page for CI and the merge queue #5863 (andygrove)
docs: add contributor guide page on memory management #5933 (andygrove)
docs: how to check the scheduled CI runs are actually running #6000 (andygrove)
docs: explain allocator hazards and diagram where memory is allocated #6014 (andygrove)
docs: pre-render mermaid diagrams to SVG at build time #6021 (andygrove)
docs: split the PR review skill by area and correct the shuffle contributor docs #6018 (andygrove)
docs: correct the stale range-partitioning strict floating-point rule #6049 (andygrove)
docs: recommend setting spark.executor.memoryOverhead alongside off-heap memory #6051 (andygrove)
docs: exempt testing-category and internal configs from the versioning policy #6089 (andygrove)
docs: make doc changes weekly comet sync #6160 (coderfender)
docs: correct the plugin and shuffle sections of the plugin overview #6197 (andygrove)
docs: fix stale and missing native Iceberg write details #6150 (andygrove)
docs: add an Iceberg writes contributor guide and review skill #6149 (andygrove)
docs: correct which operators run separate native plans in a task #6194 (andygrove)
docs: update the user guide for the 1.1.0 release #6168 (andygrove)
docs: [branch-1.1] backport the 1.1.0 upgrade notes and tuning guide updates (#6237, #6244, #6248) #6265 (andygrove)
ci: [branch-1.1] run every tier on release-branch pull requests, and the full suite before an RC (#6218) #6284 (andygrove)
docs: [branch-1.1] generate release docs for 1.1.0 #6344 (andygrove)
docs: [branch-1.1] credit co-authors in the 1.1.0 changelog #6359 (andygrove)
Other:
chore: start 1.1.0 development #5242 (andygrove)
refactor: replace hand-coded rollup of expression fallback reasons onto operators #5236 (andygrove)
test: rename CometCastSuite to CometNativeCastSuite #5268 (andygrove)
refactor: use arity helper for Int to Decimal128 reinterpretation #5193 (0lai0)
test: add guard for Iceberg version on Variant fallback test #5278 (mbutrovich)
chore: update documentation links for 1.0.0 release #5290 (andygrove)
chore(deps): bump reqsign-core from 3.2.0 to 3.2.1 in /native in the all-other-cargo-deps group #5287 (dependabot[bot])
chore(deps): bump object_store_opendal from 0.57.0 to 0.58.0 in /native #5289 (dependabot[bot])
chore(deps): bump the codeql-actions group with 2 updates #5286 (dependabot[bot])
bug: Revert “chore(deps): bump object_store_opendal from 0.57.0 to 0.58.0 in /native” #5332 (coderfender)
test: cover the narrowing direction of cast_and_stamp_schema #5285 (andygrove)
chore(deps): bump opendal from 0.57.0 to 0.58.1 in /native #5324 (manuzhang)
chore: respect Cargo parallelism settings for native release builds #5344 (pingzh)
test: make expression benchmark harness fair and reproducible #5371 (andygrove)
chore(deps): bump the all-other-cargo-deps group in /native with 5 updates #5360 (dependabot[bot])
test: fix vacuous signed-zero coverage in SQL file tests #5393 (sam-1112)
chore: fix clippy warnings for Rust 1.98 #5400 (ywskycn)
chore(deps): bump the codeql-actions group with 2 updates #5405 (dependabot[bot])
test: strengthen signed-zero SQL assertions #5404 (sunchao)
chore(deps): bump actions/checkout from 6 to 7 #5406 (dependabot[bot])
test: add expression fallback-invariance suite #5329 (4ktLuffy)
refactor: delegate ANSI integer arithmetic to arrow checked kernels #5280 (kazantsev-maksim)
test: docs and test-coverage hardening for native collect_list / collect_set #5055 (andygrove)
test: deduplicate Throwable cause-chain traversal #5441 (peterxcli)
refactor: use Arrow timezone type #5129 (Hashim1999164)
chore: add md formatting to
make format#5460 (comphead)test: fail unexpected vacuous fallback-invariance checks #5417 (sunchao)
ci: cache Maven distributions and retry bootstrap downloads #5422 (sunchao)
chore: add math benches #5479 (coderfender)
chore: Bench string functions #5492 (coderfender)
chore(deps): bump actions/setup-java from 5 to 6 #5524 (dependabot[bot])
chore(deps): bump the codeql-actions group with 2 updates #5523 (dependabot[bot])
chore: bench additional math function #5520 (coderfender)
chore: remove redundant condition and stray println in shuffle code #5562 (viirya)
test: cover struct data columns in native shuffle #5564 (viirya)
chore: rename .claude directory to vendor-neutral .ai #5594 (andygrove)
chore: bench additional scalar functions #5598 (coderfender)
test: verify Celeborn reflection compatibility #5604 (pingzh)
test: enable native path in lower/upper_enabled sql fixtures #5619 (cestercian)
chore: add benches datetime #5620 (coderfender)
chore: move dead and defensive serde guards out of convert #5595 (andygrove)
test: exercise the Iceberg write split-operator plan in Iceberg’s own suites #5640 (andygrove)
chore(deps): bump the codeql-actions group with 2 updates #5669 (dependabot[bot])
test: add explode operator microbenchmark #5381 (andygrove)
deps: bump DataFusion 55.0 and Arrow/Parquet 59.2 #5262 (mbutrovich)
chore: add benches array functions #5700 (coderfender)
test: restore Comet coverage for recursive HAVING and ORDER BY #5755 (rich7420)
test: restore Parquet V2 writer and delta encoding coverage #5760 (rich7420)
chore: Add benches for datetime funcs #5767 (coderfender)
bench: add a benchmark for the Spark hash kernels #5765 (viirya)
test: restore Spark 4.1 Variant shredding suites #5745 (rich7420)
refactor: share one helper for pushing a struct’s null mask into its children #5769 (viirya)
ci: label pull requests by changed paths and title prefix #5762 (dwsmith1983)
test: cover ambiguous exact nested Parquet field matches #5751 (peterxcli)
bench: measure nested types as native shuffle hash partitioning keys #5788 (viirya)
deps: bump to datafusion 55.1.0 #5865 (comphead)
bench: isolate map normalization and nested key hashing #5822 (viirya)
bench: add a shuffle read benchmark covering the per-block schema parse #5805 (peterxcli)
chore(deps): bump actions/setup-java from 4 to 6 #6012 (dependabot[bot])
chore(deps): bump the codeql-actions group with 2 updates #6011 (dependabot[bot])
deps: bump the iceberg-rust pin to bb1e4a4 and document why it is pinned #6094 (andygrove)
Credits#
Thank you to everyone who contributed to this release. Here is a breakdown of commits (PRs merged) per contributor. A PR with commits from more than one person counts for each of them.
147 Andy Grove
62 Peter Lee
26 Chao Sun
26 KUAN-HAO HUANG
19 Ping Zhang
15 Oleks V
14 dustin
13 Bhargava Vadlamani
13 Liang-Chi Hsieh
11 dependabot[bot]
10 ChenChen Lai
8 sam-1112
6 Matt Butrovich
5 Han-Yin Chang
5 Parth Chandra
4 Erik Bogado
4 Steve Vaughan
4 Wei Yan
3 Jordan Epstein
3 Manu Zhang
2 Alexey
2 Cestercian
2 Daipayan Mukherjee
2 Feng Zhang
2 Kazantsev Maksim
2 YuLun Mao
2 liupoyi-1031
1 Dharan-K
1 Hashim Khan
1 Henos D
1 Jeff Wang
1 LinSimon-901101
1 Michael Taranov
1 Scott Schenkein
1 Stefan Wang
1 Unik Dahal
1 Viacheslav Inozemtsev
1 Xuanyi Li
1 debaditya
1 divyank sameer shah
1 hsiang-c
Thank you also to everyone who contributed in other ways such as filing issues, reviewing PRs, and providing feedback on this release.