Comet Tuning Guide#

Comet provides some tuning options to help you get the best performance from your queries. Every deployment needs to configure how much memory Comet can use, so start with memory tuning. The guide is split into the following pages:

  • Memory Tuning: configuring Comet’s off-heap memory pool and the executor memory overhead, choosing a memory pool, batch size, and limiting spill disk usage.

  • Shuffle Tuning: enabling Comet shuffle, the native and columnar shuffle implementations, and shuffle compression.

  • Remote Shuffle with Celeborn: using Comet native shuffle with Apache Celeborn.

  • Scan Tuning: Parquet filter pushdown, Parquet split sizing, and Iceberg data file concurrency.

  • Operator Tuning: joins, adaptive partial aggregation, and sorting on floating-point values.

  • Reducing Row/Columnar Conversion Overhead: stages in which many operators fall back to Spark.

Configuring Tokio Runtime#

Comet uses a global tokio runtime per executor process. By default it starts one worker thread per executor core (spark.executor.cores, or the thread count of local[N] and local[*] masters) and allows up to 512 blocking threads, which is tokio’s default. If spark.executor.cores is not set outside local mode, Comet starts a single worker thread. These values can be overridden using the environment variables COMET_WORKER_THREADS and COMET_MAX_BLOCKING_THREADS.

Metrics Overhead#

The SQL metrics described in Metrics are always collected. Setting spark.comet.metrics.enabled=true additionally publishes plan-coverage counters (operators.native, operators.spark, queries.planned, transitions, and acceleration.ratio) through Spark’s metrics system under the comet source. It is disabled by default because it walks every executed plan on the driver after each query, and the counters are only useful with an external sink (for example Prometheus) configured. This setting must be applied before the SparkSession is created.

Explain Plan#

For an explanation of Comet plan output, the configs that control it, and how fallback to Spark works, see Understanding Comet Plans.