Repository navigation
Conversation
PFOR splits a page into vectors, 1024 values by default. Each vector subtracts its minimum, bit-packs the offsets at the width a cost model picks, and stores the values that need more bits as exceptions with their positions. An offset array locates every vector, so a reader decodes a vector on first access and skipping does not unpack the ones in between. The reader checks each header field and exception position against the page before using it. Tests cover widths from 0 up to the full value width, partial vectors and malformed pages.
The writer properties gain a PFOR switch, global or per column, which the V2 values writer factory checks before BYTE_STREAM_SPLIT for INT32 and INT64 columns. It is off by default. The format has no PFOR value yet, so the metadata converter refuses to write one with a message that says so instead of failing in an enum lookup.
|
This pull request has been automatically marked as stale because it has had no activity for at least 2 months. If you are still working on this change or plan to move it forward, please leave a comment or push a new commit so we know to keep it open. Otherwise, this PR will be closed automatically in about one month. Thank you for your contribution to Apache Parquet! |
parquet.enable.pfor turns PFOR on for jobs configured through a Hadoop Configuration, exposed the same way as parquet.enable.bytestreamsplit. It defaults to off.
Encodes and decodes INT32 and INT64 columns drawn from constant, sequential, clustered, outlier-heavy, random and TPC-DS-like distributions, and prints each one's compression ratio. Surefire's benchmark exclusion keeps it out of normal test runs.
prtkgaur
force-pushed
the
pfor-encoding
branch
from
October 8, 2026 03:33
1a4d062 to
6df1b09
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
What changes are included in this PR?
Are these changes tested?
Are there any user-facing changes?