CI: run each benchmark query in a separate process - #9385
Conversation
Signed-off-by: Mikhail Kot <mikhail@spiraldb.com>
Merging this PR will degrade performance by 20.38%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | decompress[u64, (10000, 4)] |
310.1 µs | 403.3 µs | -23.12% |
| ❌ | Simulation | cold_misaligned[(64, 256)] |
4.4 ms | 5.3 ms | -17.53% |
| 🆕 | WallTime | words_gather_dispatch[1024] |
N/A | 7 ns | N/A |
| 🆕 | WallTime | words_gather_dispatch[65536] |
N/A | 1.3 µs | N/A |
| 🆕 | WallTime | words_gather_scalar[1024] |
N/A | 100 ns | N/A |
| 🆕 | WallTime | words_gather_scalar[65536] |
N/A | 6.4 µs | N/A |
| 🆕 | WallTime | words_gather_dispatch[1024] |
N/A | 17 ns | N/A |
| 🆕 | WallTime | words_gather_dispatch[65536] |
N/A | 1.3 µs | N/A |
| 🆕 | WallTime | words_gather_scalar[1024] |
N/A | 147 ns | N/A |
| 🆕 | WallTime | words_gather_scalar[65536] |
N/A | 9.4 µs | N/A |
| 🆕 | WallTime | words_gather_dispatch[1024] |
N/A | 30 ns | N/A |
| 🆕 | WallTime | words_gather_dispatch[65536] |
N/A | 2.1 µs | N/A |
| 🆕 | WallTime | words_gather_scalar[1024] |
N/A | 360 ns | N/A |
| 🆕 | WallTime | words_gather_scalar[65536] |
N/A | 22.5 µs | N/A |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing myrrc/bench-process-per-query (eac735c) with develop (7ba790f)2
Footnotes
-
89 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
develop(204d1d4) during the generation of this report, so 7ba790f was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
Polar Signals Profiling ResultsLatest Run
Powered by Polar Signals Cloud |
Benchmarks: PolarSignals Profiling 📖Vortex (geomean): 2.055x ❌ datafusion / vortex-file-compressed / ns (2.055x ❌, 0↑ 10↓)
No file size changes detected. |
Benchmarks: FineWeb NVMe 📖Verdict: No clear signal (environment too noisy confidence) How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (2.162x ❌, 0↑ 9↓)
datafusion / parquet / ns (1.449x ❌, 0↑ 9↓)
duckdb / vortex-file-compressed / ns (1.697x ❌, 0↑ 9↓)
duckdb / parquet / ns (1.063x ➖, 1↑ 3↓)
File Size Changes (2 files changed, -46.3% overall, 0↑ 2↓)
Totals:
|
Benchmarks: TPC-H SF=1 on NVME 📖Verdict: Likely regression (environment too noisy confidence) How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (2.167x ❌, 0↑ 22↓)
datafusion / parquet / ns (1.329x ❌, 0↑ 21↓)
duckdb / vortex-file-compressed / ns (1.820x ❌, 0↑ 22↓)
duckdb / parquet / ns (1.054x ➖, 0↑ 2↓)
File Size Changes (9 files changed, -43.9% overall, 0↑ 9↓)
Totals:
|
Benchmarks: FineWeb S3 📖Verdict: Likely regression (environment too noisy confidence) How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.026x ➖, 0↑ 1↓)
datafusion / parquet / ns (1.193x ➖, 0↑ 2↓)
duckdb / vortex-file-compressed / ns (1.000x ➖, 0↑ 0↓)
duckdb / parquet / ns (0.168x ✅, 9↑ 0↓)
|
Benchmarks: Statistical and Population Genetics 📖Verdict: No clear signal (environment too noisy confidence) How to read Verdict and Engines
duckdb / vortex-file-compressed / ns (1.066x ➖, 3↑ 4↓)
duckdb / parquet / ns (1.047x ➖, 0↑ 2↓)
File Size Changes (2 files changed, -32.3% overall, 0↑ 2↓)
Totals:
|
Benchmarks: TPC-H SF=10 on NVME 📖Verdict: Likely regression (environment too noisy confidence) How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.423x ❌, 0↑ 22↓)
datafusion / parquet / ns (1.202x ❌, 0↑ 19↓)
duckdb / vortex-file-compressed / ns (1.363x ❌, 0↑ 19↓)
duckdb / parquet / ns (1.045x ➖, 1↑ 3↓)
File Size Changes (9 files changed, -44.0% overall, 0↑ 9↓)
Totals:
|
Benchmarks: Clickbench on NVME 📖Verdict: No clear signal (environment too noisy confidence) How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.722x ❌, 0↑ 38↓)
datafusion / parquet / ns (2.150x ❌, 0↑ 39↓)
duckdb / vortex-file-compressed / ns (1.761x ❌, 0↑ 41↓)
duckdb / parquet / ns (1.421x ❌, 0↑ 42↓)
File Size Changes (101 files changed, -39.2% overall, 0↑ 101↓)
Totals:
|
Benchmarks: Clickbench Sorted on NVME 📖Verdict: No clear signal (environment too noisy confidence) How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (2.359x ❌, 0↑ 10↓)
datafusion / parquet / ns (2.076x ❌, 0↑ 9↓)
duckdb / vortex-file-compressed / ns (2.188x ❌, 0↑ 10↓)
duckdb / parquet / ns (1.563x ❌, 0↑ 9↓)
File Size Changes (201 files changed, -42.8% overall, 49↑ 152↓)
Totals:
|
Benchmarks: TPC-H SF=1 on S3 📖Verdict: No clear signal (environment too noisy confidence) How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.634x ❌, 0↑ 20↓)
datafusion / parquet / ns (1.559x ❌, 0↑ 21↓)
duckdb / vortex-file-compressed / ns (1.120x ➖, 0↑ 0↓)
duckdb / parquet / ns (0.998x ➖, 0↑ 0↓)
|
Benchmarks: TPC-DS SF=1 on NVME 📖Verdict: No clear signal (environment too noisy confidence) How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.895x ❌, 0↑ 97↓)
datafusion / parquet / ns (1.682x ❌, 0↑ 99↓)
duckdb / vortex-file-compressed / ns (1.624x ❌, 0↑ 97↓)
duckdb / parquet / ns (1.163x ❌, 2↑ 74↓)
File Size Changes (25 files changed, -43.5% overall, 0↑ 25↓)
Totals:
|
|
Should we also clear the cache between runs? |
I think no, because then we'd measure cold runs only, and I find ClickBench's combined metric or similar metric of that kind) useful. In the end, in production you'd likely operate on a warm server |
Currently all queries in a benchmark suite are run in a same process.
This is an issue for #9381 since
memory pressure for earlier queries influences latter queries. This in turn
makes results very different.
The distinction between now and before is that cache was dropped per benchmark target (i.e. clickbench) and all queries per benchmark ran in same process. Now every query runs 10 times, each time in a separate process, and cache is dropped after every query (i.e. after every 10 runs) to make sure we don't benefit from cached pages from previous request.
Now we spawn a new process for every query. This also brings benchmark results
closer to ClickBench.
As result, duckdb-bench and datafusion-bench no longer need
--iterationsflag.Resolves: #9381