[fix](parquet) Fill columns absent from the physical parquet schema - #66850
Open
liutang123 wants to merge 1 commit into
Open
[fix](parquet) Fill columns absent from the physical parquet schema#66850liutang123 wants to merge 1 commit into
liutang123 wants to merge 1 commit into
Conversation
A table-format reader builds its schema-change tree from FE-supplied schema info (history_schema_info, field-id mapping), so it can map a table column to a file column that is not present in this file's physical parquet schema — for example a column added by schema change whose default was never materialized into the older data files. Such a column was neither read nor filled: _init_read_columns walks the file schema in physical order, so the column was silently dropped, and because children_column_exists() reported it as present it was never classified as fill-missing either. It then surfaced downstream as a 0-row column, tripping the filter.size() == offsets.size() check in filter_block_internal. Detect this case in _do_init_reader by checking the mapped file column against the physical schema, and demote the column to a missing column so it is filled with its default/null values.
Contributor
|
Thank you for your contribution to Apache Doris. Please clearly describe your PR:
|
Contributor
Author
|
run buildall |
Contributor
TPC-H: Total hot run time: 16851 ms |
Contributor
TPC-DS: Total hot run time: 80393 ms |
Contributor
TPC-DS: Total hot run time: 81181 ms |
Contributor
BE UT Coverage ReportIncrement line coverage Increment coverage report
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What problem does this PR solve?
A table-format reader builds its schema-change tree from FE-supplied schema info (history_schema_info, field-id mapping), so it can map a table column to a file column that is not present in this file's physical parquet schema — for example a column added by schema change whose default was never materialized into the older data files.
Such a column was neither read nor filled: _init_read_columns walks the file schema in physical order, so the column was silently dropped, and because children_column_exists() reported it as present it was never classified as fill-missing either. It then surfaced downstream as a 0-row column, tripping the filter.size() == offsets.size() check in filter_block_internal.
Detect this case in _do_init_reader by checking the mapped file column against the physical schema, and demote the column to a missing column so it is filled with its default/null values.
Stack:
Issue Number: close #xxx
Related PR: #xxx
Problem Summary:
Release note
None
Check List (For Author)
Test
Behavior changed:
Does this need documentation?
Check List (For Reviewer who merge this PR)