json_funcs Expression Audits#

Audit notes for expressions in this category that have been audited. Absence of an entry means the expression has not been audited yet, not that it is unsupported. See the user guide Spark Expression Support for current support status.

from_json#

  • Partial native support, marked Incompatible (requires explicit schema).

get_json_object#

  • Spark 3.4.3 (audited 2026-05-27): identical to 3.5.8.

  • Spark 3.5.8 (audited 2026-05-27): baseline. BinaryExpression with ExpectsInputTypes with CodegenFallback; inputTypes = Seq(StringType, StringType) -> StringType. Eval is inline and uses Jackson with RawStyle output. Foldable paths are parsed once. Returns NULL for invalid JSON, missing paths, or JsonProcessingException.

  • Spark 4.0.1 (audited 2026-05-27): the eval is extracted into a GetJsonObjectEvaluator helper (no behaviour change). The trait set now mixes in DefaultStringProducingExpression, and inputTypes is widened to StringTypeWithCollation(supportsTrimCollation = true) for both arguments.

  • Spark 4.1.1 (audited 2026-05-27): identical to 4.0.1.

  • Known incompatibility: the default codegen dispatcher uses Spark’s implementation. The native path is opt-in through spark.comet.expression.GetJsonObject.allowIncompatible=true; it rejects single-quoted JSON and unescaped control characters, and can differ for selected very large numbers, rare floating-point boundaries, duplicate keys inside returned containers, and Jackson’s recycled-buffer state near its number limit. Non-default Spark 4.0 string collations are not propagated (#2190).

  • Performance (tuned 2026-09-30, PR #4971): selected JSON and terminal-wildcard output reuse the serializer’s UTF-8 guarantee instead of validating every byte again. For 65,535-byte CJK strings in 64-row batches, terminal wildcards improve from 5.66 to 3.18 ms and selected arrays from 5.22 to 3.07 ms. Benchmark: benches/get_json_object.rs.

json_array_length#

  • LengthOfJsonArray: UnaryExpression with ExpectsInputTypes with CodegenFallback; inputTypes = Seq(StringType) -> IntegerType. Returns NULL for NULL input, invalid JSON, or non-array JSON; otherwise the number of top-level array elements.

  • Runs through the codegen dispatcher by default for byte-exact Spark compatibility.

  • Known incompatibility: the native path (built on serde_json) requires strict JSON, so single-quoted JSON, unescaped control characters, and trailing content require spark.comet.expression.LengthOfJsonArray.allowIncompatible=true and may still produce different results.

to_json#

  • Partial native support; options and map/array inputs fall back.

  • Performance (tuned 2026-07-13, PR #4902): escape_string returns Cow (zero-alloc borrow when nothing needs escaping) and bulk-copies unescaped byte runs instead of pushing char-by-char per value. 2x faster. Benchmark: benches/to_json.rs.