The study compares graph databases by their ingest and query costs.
The authors compared Corvic AI, a columnar query engine, with seven specialised graph databases or databases with graph extensions at three graph scales.
No system in the sample was fastest at every task: Ladybug outperformed Corvic AI on queries over bounded neighbourhoods, while Corvic AI was faster on queries that scan or join a large part of the graph.
The authors call the main difference in their data the cost of preparing data for queries, rather than query latency: bulk-ingest throughput varied by three orders of magnitude, from 5.0 thousand to 4.3 million rows per second. Their crossover-point calculation shows that for fewer than roughly 100 thousand queries per data refresh, the ingest difference determines total cost.
Claim check:
- The authors compared Corvic AI with seven specialised graph databases or databases with graph extensions at three graph scales. (confirmed by the publication itself: evidence; «We benchmark Corvic AI - a purpose-built columnar query engine underlying Corvic’s ontology management layer ("memories")- against seven purpose-built or graph-extension database systems (LoraDB, Ladybug, DuckPGQ, Memgraph, Neo4j, HugeGraph, and FalkorDB) at three graph scales spanning three orders of magnitude.»)
- No system in the sample was fastest at every task: Ladybug outperformed Corvic AI on queries over bounded neighbourhoods, while Corvic AI was faster on queries that scan or join a large part of the graph. (confirmed by the publication itself: evidence; «Our central finding is that no system in this sample is categorically fastest: a native graph engine (Ladybug) outperforms Corvic AI on narrow, bounded-neighborhood shapes, while Corvic AI is faster on shapes that scan or join a large fraction of the graph»)
- The authors call the main difference in their data the cost of preparing data for queries, rather than query latency: bulk-ingest throughput varied by three orders of magnitude, from 5.0 thousand to 4.3 million rows per second. (confirmed by the publication itself: evidence; «The dominant cost differential in our data is not query latency but the cost of making data queryable at all: bulk-ingest throughput varies by three orders of magnitude across engines (5.0k-4.3M rows/s)»)
- Their crossover-point calculation shows that for fewer than roughly 100 thousand queries per data refresh, the ingest difference determines total cost. (confirmed by the publication itself: evidence; «a simple crossover-point calculation shows dominates total cost for any workload with fewer than roughly 105 queries per data refresh.»)
Primary sources:
score 55.5 out of 100 · kind: research