Execute tensor L2 norm with RowFn - #9768
Performance Regression: -5.28%
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 3 improved benchmarks
❌ 7 regressed benchmarks
✅ 2177 untouched benchmarks
⏩ 218 skipped benchmarks1
Warning
Please fix the performance issues or acknowledge them on CodSpeed.
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | column_x_extension_constant[32] |
170.8 µs | 233.4 µs | -26.82% |
| ❌ | Simulation | column_x_extension_constant[256] |
427.2 µs | 574.3 µs | -25.61% |
| ❌ | Simulation | column_x_column[256] |
88.9 µs | 112.5 µs | -20.97% |
| ❌ | Simulation | column_x_column[32] |
94.6 µs | 117.2 µs | -19.25% |
| ❌ | Simulation | column_x_constant[256] |
622.7 µs | 743.3 µs | -16.23% |
| ❌ | Simulation | column_x_constant[32] |
419.7 µs | 474.8 µs | -11.61% |
| ❌ | Simulation | column_x_extension_constant[2] |
246.4 µs | 276.4 µs | -10.85% |
| ⚡ | Simulation | nullable[2] |
630.1 µs | 417 µs | +51.1% |
| ⚡ | Simulation | non_nullable[2] |
630.3 µs | 419.2 µs | +50.34% |
| ⚡ | Simulation | allocate_drop_bytes[0] |
520.2 ns | 466 ns | +11.62% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ct/row-fn-tensor-l2-v2 (5974184) with ct/row-fn-tensor-rows (5eba5b2)
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩