Deploying a standalone Ballista cluster using cargo install#

Another simple way to start a local cluster for testing purposes is to use cargo to install the scheduler and executor crates.

cargo install --locked ballista-scheduler
cargo install --locked ballista-executor

With these crates installed, it is now possible to start a scheduler process.

RUST_LOG=info ballista-scheduler

The scheduler will bind to port 50050 by default.

Next, start an executor processes in a new terminal session.

RUST_LOG=info ballista-executor

The executor will bind to port 50051 by default. Additional executors can be started by manually specifying a bind port. For example:

RUST_LOG=info ballista-executor --bind-port 50054 --bind-grpc-port 50055 --bind-health-port 50056

Each executor binds three ports — Arrow Flight (--bind-port, default 50051), gRPC (--bind-grpc-port, default 50052), and HTTP health (--bind-health-port, default 50053) — so every port must be moved, not just the first.

Installing with Optional Features#

Ballista supports optional features that can be enabled during installation using the --features flag.

Spark-Compatible Functions#

To enable Spark-compatible scalar, aggregate, and window functions from the datafusion-spark crate:

# Install scheduler with spark-compat feature
cargo install --locked --features spark-compat ballista-scheduler

# Install executor with spark-compat feature
cargo install --locked --features spark-compat ballista-executor

ballista-cli has no spark-compat feature, so a CLI installed with cargo install cannot call Spark functions: SQL is planned in the client, which has to know the functions too. To use them from the CLI, build it from source as described in Spark-Compatible Functions.

When the spark-compat feature is enabled, additional functions like sha1, expm1, sha2, and others become available in SQL queries.

Note: The spark-compat feature provides Spark-compatible expressions and functions only, not full Apache Spark API compatibility.

For more details about Spark-compatible functions, see Spark-Compatible Functions.