Repository navigation
Conversation
| val isFloatComparison = (left.dataType == FloatType || left.dataType == DoubleType) && | ||
| (op == "Eq" || op == "NotEq" || op == "Lt" || op == "LtEq" || op == "Gt" || | ||
| op == "GtEq" || op == "IsNotDistinctFrom") | ||
| val lhs = if (isFloatComparison) NormalizeNaNAndZero(left) else left |
There was a problem hiding this comment.
Could we preserve predicate pruning when adding these normalization wrappers?
This code also runs with isPruningExpr = true, so even a simple floating-point filter such as d > 100.0 becomes a comparison between scalar function expressions.
The current Parquet pruning logic cannot rewrite the wrapped column expression, and the ORC converter requires a direct Column/Literal comparison.
There was a problem hiding this comment.
Thanks for catching this. I kept normalization in normal float comparison evaluation while preserving direct column/literal expressions for scan pruning. Testing also exposed an ORC case where d > 100.0 could incorrectly prune a row group containing NaN, so the ORC patch now keeps row groups when their float statistics are uncertain. I added Float/Double converter and Parquet/ORC scan regressions.
|
The Spark 3.0 and 3.1 failures came from the ORC Spark baseline dropping NaN for d > 100.0. Auron returned the expected NaN row. I updated the regression test to assert that result and native execution directly. |
Which issue does this PR close?
Closes #2517
Rationale for this change
Native floating-point comparisons and min/max aggregations return different results from Spark for NaN and signed zero. Spark treats NaN as greater than non-NaN values and considers -0.0 equal to 0.0.
What changes are included in this PR?
Are there any user-facing changes?
Yes. Native Float/Double comparisons and min/max results now match Spark for these edge cases.
How was this patch tested?
cargo test -p datafusion-ext-plans --lib --lockedcargo fmt --all --check./dev/reformat --check./build/mvn -pl spark-extension-shims-spark -am install -DskipTests -Ppre -Pspark-3.5 -Pscala-2.12./build/mvn -pl spark-extension-shims-spark test -DskipBuildNative -Ppre -Pspark-3.5 -Pscala-2.12 -Dsuites=org.apache.auron.AuronFunctionSuiteWas this patch authored or co-authored using generative AI tooling?
Generated-by: OpenAI Codex (GPT-6)