From 39d5904f7ef108afd3406bb3fde059ad307b33a2 Mon Sep 17 00:00:00 2001 From: "Joshua D. Drake" Date: Mon, 3 Aug 2026 11:31:51 -0600 Subject: [PATCH] docs: re-measure the cross-engine table on current main (#348) The published cross-engine query latency table was measured on 2026-08-02 at 12:53. Column projection merged at 17:33 the same day, 4.7 hours later. Every figure in it therefore came from a build where the reader loaded and decoded every column of every row group it visited, whatever the query referenced (#338, fixed in #339). It understated the project, which is the direction nobody checks. It published q6 at 81966 ms against TimescaleDB's 1771. Measured today the same shape is 1592 ms against 6956. It also showed pgColumnar losing to heap on q4 and q5, which does not reproduce. This is a new run rather than a row-by-row correction, because the two are not the same experiment. The fixture now carries indexes the earlier one did not, heap moving from 11045 ms to 6 ms on q1, and the earlier queries were never recorded. bench-tsdb/BENCHMARK_RESULTS.md, cited by #289 as the full data and method, has never existed in this repository. So the method is stated here and the SQL is published with the table. TimescaleDB is reported serial and labelled as such. Its parallel path fails on the bench host with "could not read blocks 0..0". That is not caused by pgColumnar: it persists with every pgColumnar planner hook disabled, and the chunks use the heap access method. Neither available workaround is usable, one still errors and the other silently returns zero rows. Citus is omitted rather than carried forward from a run whose configuration is unknown. q1 to q3 are published with their cause attached. They cost about 28.7 ms per row returned, flat across a twelve-fold change in row count, on the same index scan plan heap uses. That is #353, the default stripe_row_limit putting a table this wide over the 32 MB fetch cache, compounded by #355, the planner choosing an index scan for ordering without modelling the fetch cost. Publishing 123 seconds without that would read as a property of columnar storage rather than as one default and one cost model. The storage table is unchanged. Measured 6587 MB, 7975 MB and 22 GB against its published 6.4 GB, 7.8 GB and 22 GB. It was correct. The 2.67 GB figure in #289's body was the wrong one. Results were verified against heap per query. Counts and max match exactly, averages agree to 2.9e-15 relative difference, and no group appears in one engine and not the other. docs_style passes, and the STE line-length rule required splitting several sentences. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_011miCFRSatixeNRw3w5yNq8 --- docs/benchmarks.md | 144 +++++++++++++++++++++++++++++++++------------ 1 file changed, 108 insertions(+), 36 deletions(-) diff --git a/docs/benchmarks.md b/docs/benchmarks.md index 2f89cbb..25451b5 100644 --- a/docs/benchmarks.md +++ b/docs/benchmarks.md @@ -377,43 +377,115 @@ zstd compounds it, as the [Storage](#storage) section shows in detail. ### Cross-engine query latency -Warm latency, median of three, in milliseconds: - -| query | shape | pgColumnar | TimescaleDB | heap | Citus | +Measured on 2026-08-03 against `main` at commit `00290d7`, on the same +100,000,000-row fixture. This is a new run, not a correction of the previous +table row by row. The earlier figures were taken before column projection landed (issue #338, +fixed in #339). Every query then read and decoded every column of the table, +whatever it referenced. The queries behind those figures were also not recorded. +The two runs cannot be compared line by line. + +Method for this run, stated so it can be repeated: + +- All three engines hold the same 100,000,000 rows, verified before measuring. +- `cpu_pgc` and `cpu_heap` each carry a `(hostname, time DESC)` btree. + TimescaleDB carries the equivalent index on its chunks. +- pgColumnar and heap run with `max_parallel_workers_per_gather = 4`. + TimescaleDB runs serial, because its parallel path fails on this host with + `could not read blocks 0..0`. That fault is not caused by pgColumnar: it + persists with every pgColumnar planner hook disabled, and the chunks use the + heap access method. +- pgColumnar figures are given twice: with default settings, and with + `enable_ungrouped_vector_agg`, `enable_parallel_vector_agg` and + `enable_group_vectorization` all on. Those three default to off. +- Warm, `EXPLAIN (ANALYZE)` execution time, after one warm-up run. +- Citus was not re-measured and is omitted rather than carried forward from a + run whose configuration is unknown. + +Milliseconds: + +| query | shape | pgColumnar | pgColumnar, all options on | heap | TimescaleDB (serial) | | --- | --- | --- | --- | --- | --- | -| q1 | one host, 1 hour | 994 | 4 | 11045 | 315 | -| q2 | one host, 12 hours | 10532 | 5 | 11244 | 1935 | -| q3 | one host, 12 hours, 5 aggregates | 10462 | 6 | 11341 | 3630 | -| q4 | all hosts, 12 hours, group by host | 18856 | 5098 | 17290 | 6694 | -| q5 | all hosts, 12 hours, 10 aggregates | 22671 | 10437 | 21247 | 15407 | -| q6 | full scan, one value filter | 81966 | 1771 | 14699 | 8810 | -| q7 | last point per host | 785 | 299 | 47 | timeout | -| q8 | top 20 by max | 37397 | 7572 | 23970 | 12943 | - -Read this by query shape, not by a single winner. Three patterns hold. - -**TimescaleDB wins the host-filtered queries by a wide margin.** q1 to q3 filter -on one `hostname`. The columnstore segments by `hostname`, so it reads one segment -and skips the rest. pgColumnar stores in time order, so its zone maps do not skip -on `hostname` and it scans the range. This is a layout choice, not a ceiling. A -separate run stored the same table clustered on the filter key. q1 to q3 then fell -from about ten seconds to about one hundred milliseconds, because the zone maps -skipped. The cost is the all-host queries, which the clustered layout slows. - -**The full-scan aggregate is pgColumnar's weak shape today.** q6 reads every row -and filters on a value column. pgColumnar serial is slower than heap on it. The -decompression and aggregation path is not yet vectorized for this case. That work -is [issue #289](https://github.com/jdatcmd/pgcolumnar/issues/289). The parallel -scan below already brings q6 close to heap, and the vectorization will take it -further. - -**Citus and pgColumnar are within a small factor on the heavy grouped queries** -(q4, q5, q8). Citus times out on last point (q7) at the 120-second limit. - -The last-point query (q7) is not a scan number for pgColumnar or heap. Both tables -have a `(hostname, time)` index. The query reads one row per host, so the planner -walks the index rather than a full sort. TimescaleDB answers it from segment -order. +| q1 | one host, 1 hour | 10335 | 10253 | 6 | 0 | +| q2 | one host, 12 hours | 123546 | 119534 | 8 | 1 | +| q3 | one host, 12 hours, 5 aggregates | 161972 | 156701 | 8 | 2 | +| q4 | all hosts, 12 hours, group by host | 7322 | 7863 | 13912 | 4515 | +| q5 | all hosts, 12 hours, 10 aggregates | 16443 | 28494 | 17378 | 9576 | +| q6 | full scan, one value filter | 2292 | 1592 | 9083 | 6956 | +| q7 | last point per host | 368 | 365 | 31 | 196 | +| q8 | top 20 by max | 5153 | 5124 | 8715 | 10171 | + +The SQL is given at the end of this section. The previous table did not record +it, which is the main reason its numbers cannot be checked. + +Results were verified against heap per query. Counts and `max` aggregates match +exactly. Averages agree to within 2.9e-15 relative difference. The residual is float +reassociation across parallel workers. Every group present in one engine is +present in the other. + +**q1 to q3 are a defect, not a storage property.** They filter on one host and +return 360, 4,320 and 4,320 rows. Both engines take the same `Index Scan` plan on the equivalent index. +pgColumnar costs about 28.7 milliseconds per row returned. That cost is flat +across a twelve-fold change in row count. The cause is +[issue #353](https://github.com/jdatcmd/pgcolumnar/issues/353). The default +`stripe_row_limit` of 150,000 puts a table this wide over the 32 MB fetch cache +limit. Every fetch by row number then decodes the whole row group again. Lowering +`stripe_row_limit` to 100,000 on the same data takes the same query from 32.98 to +0.177 milliseconds per row. Until that is fixed, a table that serves selective +point queries through an index should be created with a smaller +`stripe_row_limit`. + +[Issue #355](https://github.com/jdatcmd/pgcolumnar/issues/355) compounds it. +The planner will choose an index scan over a columnar table to obtain ordering. +It does that because the per-row fetch cost is not modelled. That is why these queries should +not be read as a measure of the storage format. + +**pgColumnar leads on the scan-bound shapes.** q4, q6 and q8 read a large part of +the table. pgColumnar is ahead of heap on all three, and ahead of TimescaleDB on +q6 and q8. q6 in particular is 1592 ms against TimescaleDB's 6956 ms, where the +earlier table recorded 81966 ms against 1771 ms. Column projection accounts for +most of that change: on this fixture it alone takes q6 from 45094 ms to 6491 ms. + +**The optional vectorization is not uniformly a win.** It helps q6, at 2292 ms +down to 1592 ms. It leaves q1, q3, q7 and q8 unchanged. It costs on q5, 16443 ms +up to 28494 ms, which is not yet explained and is tracked in +[issue #349](https://github.com/jdatcmd/pgcolumnar/issues/349). This is why those +settings default to off. + +**TimescaleDB leads on the host-filtered queries** for the reason given before: +its columnstore segments by `hostname`, so it reads one segment. pgColumnar +stores in load order, so its zone maps do not skip on `hostname`. Clustering a +pgColumnar table on the filter key addresses that, at the cost of the all-host +queries. + +#### Queries + +```sql +-- q1, q2: one host, 1 hour and 12 hours +SELECT date_trunc('minute',time) m, max(usage_user) FROM cpu +WHERE hostname='host_1' AND time >= '2024-01-01' AND time < '2024-01-01' + interval '1 hour' +GROUP BY 1 ORDER BY 1; + +-- q3: as q2 with five aggregates +SELECT date_trunc('minute',time) m, max(usage_user), max(usage_system), + max(usage_idle), max(usage_nice), max(usage_iowait) FROM cpu +WHERE hostname='host_1' AND time >= '2024-01-01' AND time < '2024-01-01' + interval '12 hours' +GROUP BY 1 ORDER BY 1; + +-- q4, q5: all hosts over 12 hours, one metric and ten +SELECT date_trunc('hour',time) h, hostname, avg(usage_user) FROM cpu +WHERE time >= '2024-01-01' AND time < '2024-01-01' + interval '12 hours' +GROUP BY 1,2; + +-- q6: full scan, one value filter +SELECT count(*), avg(usage_system) FROM cpu WHERE usage_user > 90.0; + +-- q7: last point per host +SELECT DISTINCT ON (hostname) hostname, time, usage_user FROM cpu +ORDER BY hostname, time DESC; + +-- q8: top 20 by max +SELECT hostname, max(usage_user) mx FROM cpu GROUP BY 1 ORDER BY mx DESC LIMIT 20; +``` ### Parallel scan