Skip to content

[WIP][POC] Pfor encoding - #3595

Draft
prtkgaur wants to merge 4 commits into
apache:masterfrom
prtkgaur:pfor-encoding
Draft

prtkgaur wants to merge 4 commits into
apache:masterfrom
prtkgaur:pfor-encoding

Conversation

@prtkgaur

@prtkgaur prtkgaur commented Jun 3, 2026

Copy link
Copy Markdown

Rationale for this change

What changes are included in this PR?

Are these changes tested?

Are there any user-facing changes?

PFOR splits a page into vectors, 1024 values by default. Each vector
subtracts its minimum, bit-packs the offsets at the width a cost model
picks, and stores the values that need more bits as exceptions with
their positions. An offset array locates every vector, so a reader
decodes a vector on first access and skipping does not unpack the ones
in between. The reader checks each header field and exception position
against the page before using it. Tests cover widths from 0 up to the
full value width, partial vectors and malformed pages.
The writer properties gain a PFOR switch, global or per column, which
the V2 values writer factory checks before BYTE_STREAM_SPLIT for INT32
and INT64 columns. It is off by default. The format has no PFOR value
yet, so the metadata converter refuses to write one with a message that
says so instead of failing in an enum lookup.
@github-actions

github-actions Bot commented Aug 3, 2026

Copy link
Copy Markdown

This pull request has been automatically marked as stale because it has had no activity for at least 2 months. If you are still working on this change or plan to move it forward, please leave a comment or push a new commit so we know to keep it open. Otherwise, this PR will be closed automatically in about one month. Thank you for your contribution to Apache Parquet!

@github-actions github-actions Bot added the stale label Aug 3, 2026
parquet.enable.pfor turns PFOR on for jobs configured through a Hadoop
Configuration, exposed the same way as parquet.enable.bytestreamsplit.
It defaults to off.
Encodes and decodes INT32 and INT64 columns drawn from constant,
sequential, clustered, outlier-heavy, random and TPC-DS-like
distributions, and prints each one's compression ratio. Surefire's
benchmark exclusion keeps it out of normal test runs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants