Add end-to-end Parquet VARIANT support - #19101
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## master #19101 +/- ##
============================================
+ Coverage 67.55% 67.81% +0.26%
- Complexity 1430 1512 +82
============================================
Files 3486 3501 +15
Lines 224100 226435 +2335
Branches 35370 35808 +438
============================================
+ Hits 151392 153564 +2172
+ Misses 60678 60599 -79
- Partials 12030 12272 +242
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
3e4dc18 to
acfba94
Compare
ca98faf to
b1a2872
Compare
There was a problem hiding this comment.
Pull request overview
This PR adds end-to-end support for the VARIANT logical type across Pinot ingestion (notably Parquet VARIANT(1)), query planning/execution (single-stage + multi-stage), response encoders, and Java/JDBC client surfaces, while explicitly rejecting unsupported raw-VARIANT operations (e.g., ordering/grouping/join keys) with actionable errors.
Changes:
- Introduces
VARIANTas a first-class Pinot data type, including schema/DDL handling, wire/proto support, and null-sentinel semantics. - Adds Parquet
VARIANT(1)ingestion support via the native Parquet reader and includes a packaged “variant batch” quickstart with resources and CI coverage. - Adds query-time VARIANT functions (parse/get/existence/type/json rendering) and enforces operator-level restrictions where raw VARIANT lacks equality/hash/ordering semantics.
Reviewed changes
Copilot reviewed 174 out of 175 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| pom.xml | Adds managed dependency for org.apache.parquet:parquet-variant. |
| pinot-tools/src/test/java/org/apache/pinot/tools/admin/command/QuickStartTest.java | Maps new quickstart command strings to VariantQuickStart. |
| pinot-tools/src/main/resources/examples/batch/variantEvents/* | Adds VARIANT batch quickstart schema/table config, ingestion spec, and documentation. |
| pinot-tools/pom.xml | Adds packaged quick-start-variant-batch program entry. |
| pinot-sql-ddl/src/main/, src/test/ | Adds DDL compile/reverse/roundtrip support for VARIANT and validates default emission rules. |
| pinot-spi/src/main/, src/test/ | Adds PinotDataType.VARIANT conversion/ingestion mapping and schema validation updates. |
| pinot-segment-spi/src/main/** | Handles VARIANT default-null sentinel persisted as empty string in segment metadata. |
| pinot-segment-local/src/main/, src/test/ | Adds VARIANT validation/sanitization rules; disables min/max, ordering, dictionary creation where unsupported; adds targeted tests. |
| pinot-query-planner/src/main/, src/test/ | Adds VARIANT type mapping (Calcite/plan nodes), plan validation, constant-folding guardrails, and serde coverage. |
| pinot-query-runtime/src/main/, src/test/ | Enforces VARIANT restrictions at runtime operators (sort/window/set/group-by/etc.) and adds operator tests. |
| pinot-core/src/main/, src/test/ | Enforces VARIANT restrictions in single-stage operators/aggregations/distinct/order-by and adds tests. |
| pinot-common/src/main/, src/test/ | Adds VARIANT to encoders, datablock equality, proto enum, function canonicalization, and introduces scalar VariantFunctions. |
| pinot-clients/pinot-jdbc-client/src/main/, src/test/ | Treats VARIANT as canonical JSON text (VARCHAR) while preserving encoded-variant-null vs SQL-null behavior. |
| pinot-clients/pinot-java-client/src/main/, src/test/ | Preserves legacy getString() null behavior while distinguishing VARIANT nulls. |
| pinot-plugins/pinot-input-format/pinot-parquet/src/main/** | Enhances Parquet reader selection/metadata helpers and initializes native reader with schema-bound converters for VARIANT. |
| pinot-controller/src/main/, src/test/ | Ensures stored-vs-compiled schema default comparisons handle VARIANT byte[] defaults by content. |
| pinot-compatibility-verifier/** | Adds file-contains operation and VARIANT mixed-version suite coverage including old-broker/new-servers phase. |
| .github/workflows/scripts/.pinot_quickstart.sh | Adds CI validation for the packaged VARIANT batch quickstart. |
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 174 out of 175 changed files in this pull request and generated no new comments.
Suppressed comments (1)
pinot-query-runtime/src/main/java/org/apache/pinot/query/runtime/operator/SortOperator.java:81
- The validation uses
supportsOrdering()but the error text always mentions raw VARIANT.supportsOrdering()is alsofalsefor other non-orderable logical types (e.g. OBJECT / array types), so this can surface a misleading VARIANT-specific message when the real unsupported type is something else.
Consider tailoring the message based on the actual column type (keep the VARIANT guidance when type == VARIANT, otherwise emit a generic "ORDER BY does not support values of type X" message).
cdcfc2c to
ae86bc2
Compare
xiangfu0
left a comment
There was a problem hiding this comment.
Reviewed the full VARIANT PR (checked out the branch; ran domain-focused passes over SPI/envelope, Parquet ingestion, the VariantUtils codec, single-stage + multi-stage enforcement, schema/table validation, and wire/client/compat). Overall this is very carefully engineered — the PVAR envelope, the four null-state handling, and the defense-in-depth opacity guards all check out, and I found no data-corruption or crash bugs in the core paths.
Two items I'd suggest fixing before merge (inline: an OBJECT-comparison regression and a partial-upsert validation gap), plus a few low-severity polish notes.
Verified-correct highlights: proto appends VARIANT=24 with no renumbering and ColumnDataType is name-serialized (no ordinal()-indexed tables exist repo-wide, so the mid-enum insertion is safe); Parquet buffer ownership / shredded-vs-unshredded selection / reader precedence / reinit safety; codec bounds + per-row cursor reset + thread-confinement; and the full set of schema/table restrictions (dictionary / secondary indexes / PK / partition / sorted / star-tree (default and explicit) / metrics-agg / upsert-comparison all rejected, min/max never published, never sorted).
…eys in planner Addresses the two Low review findings on apache#19101: - ORDER BY / sort / set-op / window / aggregate rejections previously emitted a raw-VARIANT-specific message for any non-orderable/non-hashable type (OBJECT, arrays, MAP). The messages now name the actual unsupported type and keep the raw-VARIANT wording plus variantGet guidance only for VARIANT. Applies to VariantTypeValidationVisitor, SortOperator, SortedMailboxReceiveOperator, and OrderByComparatorFactory. - VariantTypeValidationVisitor.validateAggregateInputs now also validates the AggregateNode GROUP BY keys, so the planner gate rejects GROUP BY on a raw VARIANT key consistently with sort/join/set-op/window (runtime already rejected it). Added a planner test.
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
|
Follow-up on my review: all four findings are now fixed and merged into this PR (head is now
Regression tests were added for each. spotless/checkstyle/license are clean on all four affected modules, and every affected unit-test class passes locally on JDK 25 ( On the current CI run, the Linter and all VARIANT-exercising tests pass. The three red checks are pre-existing flakes unrelated to these changes and to VARIANT: |
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
524e9c6 to
65600c4
Compare
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
65600c4 to
496b466
Compare
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
496b466 to
5be73f5
Compare
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
5be73f5 to
eabb992
Compare
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
eabb992 to
378dbc0
Compare
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
378dbc0 to
8a46db8
Compare
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
8a46db8 to
94333dc
Compare
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
94333dc to
29ff871
Compare
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
29ff871 to
ac23a5a
Compare
Introduce the VARIANT logical type and envelope, Parquet reconstruction, scalar and transform functions, and segment and query validation across the single-stage and multi-stage engines. Propagate VARIANT through schemas, wire formats, responses, clients, DDL, quickstart, and integration coverage while rejecting raw operations that require ordering, equality, or hashing.
Allow SQL NULL literals to participate in comparison validation without being classified as raw VARIANT values. Keep equality and DISTINCT FROM behavior aligned in the single-stage transform layer and the multi-stage filter operand, with regression coverage for both engines.
Exercise new-broker/old-server, mixed-server, and old-broker/new-server compatibility, including deterministic handling of the VARIANT wire type before server upgrade and rejection of partial mixed-fleet results. Centralize join-key validation, clarify reusable extraction ownership, and strengthen scalar, null-placeholder, window, and end-to-end coverage.
Raw VARIANT projection is only lossless with query null handling: otherwise the reserved empty-byte SQL-null placeholder cannot preserve null semantics. Enforce this contract at the multi-stage root plan, single-stage result block, and broker normal, empty, and materialized-view split response paths. Preserve legacy ResultSet.getString() behavior for established types while mapping only VARIANT SQL null to Java null across JSON, gRPC, and JDBC clients. Make Parquet VARIANT validation honor selected fields so an unselected nested VARIANT does not reject an unrelated projection, while selected and extract-all paths remain fail-closed. Reject raw VARIANT operands for IS DISTINCT FROM and IS NOT DISTINCT FROM through shared operand type resolution, add focused and end-to-end coverage, and normalize the touched documentation to the repository style.
…ert validation, and type-accurate error messages Addresses the four review findings on apache#19101 (2 Medium: OBJECT-comparison regression, partial-upsert default-strategy gap; 2 Low: type-accurate opacity messages, planner GROUP BY key validation). Tests added across pinot-core, pinot-query-runtime, pinot-query-planner, pinot-segment-local.
Restore the parquet-common runtime dependency required by parquet-variant and validate grouping-set keys against the expanded repeat schema. Reuse the query cursor for JSON rendering, cover every Parquet Variant type, prevent timestamp-derived-name validation bypasses, and clarify the type-capability validation APIs.
Honor producer-defined physical value ordering when object field offsets are non-monotonic, and preserve exact subtree envelopes through self-delimiting size parsing. Use allocation-free lookup for wide objects with a cross-producer ordering fallback, clarify materializing cursor APIs, and add focused, cross-engine, Parquet, and JMH coverage.
An audit of VARIANT against the recently landed BYTES array literal support (apache#19247) and the UUID logical type surfaced paths where the opacity contract could be bypassed or where VARIANT was missing wiring its ByteArray-backed sibling types have. Opacity guards: - Reject CAST from a raw VARIANT in both engines. Single-stage rejects in CastTransformFunction; multi-stage rejects in TransformOperandFactory as defense-in-depth behind Calcite validation, whose coercion rules for its native VARIANT type are not this contract's to rely on. - Stop single-stage compile-time folding of parseJson/tryParseJson: folding degraded the result to a plain BYTES literal, erasing JSON rendering, null-handling enforcement, and every opacity guard, and diverging from multi-stage which already skips the fold. The canonical name sets now derive from TransformFunctionType registrations so the engines cannot drift. - Reject ROLLUP/DEDUP merges for schemas with a VARIANT column: both merge types group rows by raw stored bytes, and equivalent Variant values can use different encodings. Enforced at the SegmentProcessorConfig choke point every merge execution flows through, with early submission-time validation in both task generators, including the deprecated collectorType alias. Parity wiring: - Map VARIANT to opaque Avro bytes in both schema converters so CONCAT segment processing round-trips the envelope; covered by the existing divergence-pin and round-trip tests. - Implement single-stage variantToJson natively; the design doc already promised it on both engines, and multi-stage had it. - Resolve VARIANT literals through PinotDataType.VARIANT in LiteralContext, mirroring the UUID arm. - Name the actual failing type in capability rejections instead of blaming VARIANT for MAP/STRUCT ordering gaps, and reject VARIANT explicitly in the data generator. Udf descriptor classes for the variant function family are deferred as a mechanical follow-up.
Rows in a segment overwhelmingly share one metadata dictionary and one object layout, yet the navigation cursor re-resolved every path key per row with string comparisons. The retained cursor now memoizes, per path element, the resolved dictionary id and object-entry index, and probes them first on the next row. The memo is self-validating on the current row's bytes: a hint is trusted only after re-checking that the hinted entry's id still spells the path key, so a stale memo degrades to the ordinary search and can never select a wrong field. The one verdict that relies on cross-row state - a key absent from the entire dictionary - is anchored to the exact envelope it was proven under and reused only when the current metadata region is byte-identical to that anchored region; validating against the previous navigated row instead would launder stale verdicts through rows that bypass the memo, which a regression test now covers. The dictionary-classification investment on a miss is additionally gated on the metadata having repeated across consecutive rows, so heterogeneous per-row dictionaries keep the pre-memo miss cost. Objects at or below four entries skip the memo entirely; a direct scan is already as cheap as validating it. Also vectorize long metadata key comparisons and render variantToJson through an unsynchronized commons-io StringBuilderWriter instead of StringWriter's synchronized StringBuffer. BenchmarkVariantGetMemo (avgt us per 1024 rows, before -> after): wide100Hit 423.4 -> 50.9 (8.3x) wide1000Hit 696.7 -> 53.5 (13.0x) wide100MissFromDictionary 863.7 -> 72.8 (11.9x) wide100MissFromObject 887.9 -> 219.6 (4.0x) narrow5Hit 63.1 -> 41.0 (1.5x) nestedHit 91.8 -> 88.1 (parity) wide100RotatingMetadataHit 427.9 -> 309.5 (1.4x) heterogeneousMetadataMiss parity with pre-memo cost (gated) The small-object/large-dictionary miss shape trades ~70us per 1024 rows for the anchored revalidation; every other shape improves or holds.
95945ce to
2bf0462
Compare
Summary
Add an end-to-end
VARIANTuser journey for Apache Pinot:VARIANTdimension retained in Pinot's raw forward indexVARIANT(1)values from unshredded and shredded filesRaw Variant values are deliberately opaque to operations that require equality, hashing, or ordering. Queries must extract a typed scalar before comparing, grouping, joining, sorting, using set operations, or applying general aggregates.
COUNT(raw_variant)remains supported.The persisted format, null semantics, ownership boundaries, and rollout contract are documented in
pinot-spi/VARIANT_DESIGN.md.End-to-end flow
Define a table column with
"dataType": "VARIANT", storage null handling enabled, dictionary disabled, and a raw forward index. The native Parquet reader reconstructsVARIANT(1)metadata/value pairs—including supported shredded layouts—into Pinot's versionedPVARenvelope.Frequently queried fields can be materialized with ingestion transforms while the full Variant remains available:
Supported scope
VARIANT(1)columnsVARIANTdimensions without dictionariesSET enableNullHandling=trueNested/repeated Variant columns, streaming ingestion, and quoted path keys are outside the initial scope.
A JSON index is intentionally not accepted on a
VARIANTcolumn. The existing JSON index expects Pinot JSON text and its query rewrite semantics; treating the binaryPVARenvelope as JSON would be incorrect. Indexing Variant paths should be introduced separately with an explicit physical/index contract. Materialized scalar columns are the supported indexed path for this PR.The automatic Parquet reader selection preserves existing Avro-metadata precedence. Files containing Avro metadata must explicitly select the native Parquet reader to ingest Variant.
Production safety
PVARframing and validation inVariantEnvelopeand freezes version-1 bytes with golden tests24to Variant while preserving UUID/UUID_ARRAY at22/23"null", missing paths, and conversion failures across query and client response formatsRolling upgrades
Variant remains an explicit full-fleet activation feature: upgrade controllers, brokers, servers, clients, and external ingestion jobs before registering a Variant schema, and do not roll back while Variant tables are active.
This remains one gated PR because splitting wire/schema registration from ingestion and query guards would create mergeable partial-activation states. The layered commits remain independently reviewable while activation stays fail-closed until every layer is present.
Compatibility coverage verifies:
Quickstart
The quickstart registers
variantEvents, ingests the committed five-row Parquet fixture, uploads the segment, and runs materialized-field, nested-extraction, JSON-rendering, and null-semantics queries. Its resources live underpinot-tools/src/main/resources/examples/batch/variantEvents.Verification
VariantTypeTestscenarios across both query engines in the full integration-test reactor, including an externally encoded Parquet object with non-monotonic field offsetsVariantUtilsTestcases covering every Variant physical encoding, wide-object thresholds, parquet-java/Arrow-RS Unicode ordering, malformed values, and reusable cursorsupstream/mastergit diff --checkPerformance benchmark
A 2,000,000-row workstation benchmark compared identical logical events represented as JSON text and Parquet
VARIANT(1), then ingested into raw Pinot columns with ZSTD compression. A third JSON scenario added Pinot's JSON index and is reported separately because it is an index-assisted comparison, not a format-only comparison. The reproducible harness is preserved in the benchmark snapshot.Ingestion and storage
Compared with raw JSON, the VARIANT source was 2.22% smaller and its Pinot segment was 1.61% larger. Adding the JSON index made the JSON segment 4.26× the size of the raw JSON segment.
Query latency
Each value is the median of four round-level p50 values. Every round used two warmups and 10 measured sequential executions.
eventTypecount (control)The raw comparison used four Parquet files/Pinot segments and four ingestion rounds with alternating JSON/VARIANT build order. Source generation was excluded from ingestion timing, and every query result was checked against independently generated ground truth.
These are paired, sequential, warm-cache, single-stage measurements from an Apple M2 Max workstation with 32 GB RAM, macOS 26.5.2, and Temurin 25 on AC power. They are directional workstation evidence, not distributed production-cluster throughput: network transfer, multi-stage execution, concurrent clients, distributed scheduling, and production hardware were outside the run.
Wide-object navigation microbenchmark
A focused JMH A/B compared the pre-fix PR head (
ae60e3de73) with the interoperability fix (bc5658c3a3) on JDK 25. Each case used one fork, two 500 ms warmups, three 500 ms measurements, flat integer values, and-prof gc.The first-field probe avoids the roughly 5.8× regression found in the initial binary-search implementation. Missing keys retain a small cost because a failed parquet-java-order binary search must fall back to byte equality for external producer ordering. Both versions are effectively allocation-free (normalized values below 0.02 B/op, no GC). These short microbenchmarks are directional; the large middle/last improvements are clear, while smaller deltas have wider uncertainty.
Review areas
This PR changes public SPI and query-wire behavior. The main ownership reviews are: