PostgreSQL was built for a different era. The original project dates back to the 1980s, when the main bottleneck for a database query was disk I/O. Most datasets now fit in RAM, analytics scans data in bulk, and fast NVMe drives have moved the bottleneck to CPU and memory bandwidth. A project called pgrust re-implements Postgres in Rust from the ground up, and its version 0.2 release shows what the engine can do when it targets today's hardware. The team reports a postgres rewritten in rust engine that is 300x faster than PostgreSQL on the ClickBench analytical benchmark, and roughly 30 percent faster on OLTP workloads.
The claim is not just headline numbers. pgrust is wire compatible with Postgres, meaning clients and drivers can speak the same protocol, and it is SQL dialect compatible. The project passes all 46,066 tests in the PostgreSQL regression suite. It is also AGPL-3.0 licensed. Version 0.2 is about 10x faster than the previous pgrust release, and the query engine rewrite alone drives roughly 10x of the 300x gain.
The problem: the Volcano model
Much of that speedup comes from the query engine, and the starting point is a model most database users have never heard of. Postgres converts a query into a plan built from nodes, then executes it with a style called the Volcano model. Each node exposes a next() method that returns exactly one row at a time. A sequential scan returns the next row in the table; an aggregate pulls rows from its child until it can produce the final result.
The Volcano model is elegant and simple to extend, which is why it survived for decades. But returning one row per virtual call is expensive. A function call per tuple stops CPU pipelining and prevents runtime optimizations. The pgrust team's post walks through a miniature version of this engine in Rust to show the cost: summing 500 million floating point numbers in Postgres takes about 20 seconds on an ARM Graviton4 instance with parallel queries disabled. A bare Rust loop over the same data finishes in 358 milliseconds. That 55x gap is a rough measure of engine overhead, storage format and locking included.
Optimization 1: batching
The first fix is to stop returning one row at a time. Instead of next(), each node gets a next_batch() that fills an array of 1024 values per call. That reduces dispatch overhead and, importantly, the batch buffer lives on the stack, so the aggregate does no heap allocation while computing. Allocation is one of the slowest things a hot loop can do.
Batching alone brings the miniature engine from 1.3 seconds down to 480 milliseconds, about 2.7x faster.
Optimization 2: operator fusion
Profiling the batched version shows the next hotspot: every scan row gets copied into the batch buffer and read back by the next operator. For common combinations of nodes, you can fuse them into one operator that can pass the data straight through, or even skip copying entirely.
In the example, fusing the sequential scan with the sum aggregate removes the intermediate copy and produces code that is essentially the same as the hand-written loop, at 358 milliseconds. Fusing every possible combination by hand does not scale, so the natural next step is JIT compilation, generating the fused code at query time. The pgrust team says a follow-up post will cover how it handles JIT.
Optimization 3: SIMD
The final trick uses the CPU as a small vector unit. SIMD instructions operate on multiple values at once, and summing eight floats per instruction instead of one is much faster. The example uses aarch64 NEON intrinsics to add four 128-bit accumulators in parallel, then combines them at the end.
The SIMD version runs the sum in 135 milliseconds, about 10x faster than the original Volcano code and almost 3x faster than the scalar Rust loop. Compilers usually do not auto-vectorize floating point sums, because floating point addition is not associative and reordering the addition changes results subtly. The developer has to opt in with explicit intrinsics.
The numbers
| implementation | time for 500M floats | speedup vs Volcano |
|---|---|---|
| Postgres | ~20 s | baseline |
| Volcano model | 1.3 s | 1x |
| + batching | 480 ms | 2.7x |
| + operator fusion | 358 ms | 3.6x |
| + SIMD | 135 ms | 9.6x |
Benchmark setup: AWS Graviton4 instance with 16 vCPUs, PostgreSQL 18.4 with parallel query disabled, warm data in shared buffers, median of 5 runs. A single query is not a workload, but the sequence shows where the cycles go.
Scaled up to full queries and real analytical tables, the same techniques lead to the ClickBench results: pgrust claims 300x over Postgres and the benchmark page shows it ahead of ClickHouse, a database purpose-built for analytical workloads.
What this means
Nothing here forces you to switch database. The point for developers is that the Postgres engine is not fundamentally maxed out. The C codebase, maintained for decades by thousands of contributors, will not be absorbed into Rust overnight. But engines like the one in pgrust show there is real headroom in query execution, and the techniques, batching, fusion, JIT, SIMD, are the same playbook used by DuckDB, DataFusion, and other fast analytical engines.
For teams running analytics on Postgres today, the practical options are the ones available since before this release: partition for scan elimination, keep working sets warm, use columnar storage extensions when long-range aggregates dominate, and check whether ClickHouse or a vectorized engine fits the workload. pgrust is a signal about where the ecosystem is heading, not a production migration target yet.
Frequently Asked Questions
- Is pgrust a fork of PostgreSQL? No, it is a compatible re-implementation in Rust. It speaks the wire protocol and SQL dialect of Postgres and passes the regression suite.
- Can I run my existing client against it? Wire compatible and SQL dialect compatible mean standard drivers and tooling will connect, though the project is early stage.
- Is it ready for production? Version 0.2 is a research-grade project, not a production database. Do not replace production Postgres with it today.
- Is Rust the future of Postgres? Postgres itself is written in C, and the C codebase is not going anywhere. Rust re-implementations like this one show what a memory-safe engine can look like, and they inform future design.
Key Takeaways
- The Postgres query engine is CPU bound on modern hardware workloads.
- Batching, operator fusion, and SIMD are the three structural changes with the largest return.
- A Rust rewrite can stay compatible with Postgres and still show 300x speedups on analytical benchmarks.
- Real Postgres deployments have migration paths that do not leave the C codebase, but the gap in engine efficiency is real.
Conclusion
The pgrust v0.2 release is a useful benchmark for expectations about how fast a Postgres-compatible engine can run. It does not turn Postgres into a fast analytics database by default, but it maps the trade-offs: executor overhead, memory, CPU, storage. If the JIT work delivers on the promise in the post, compatibility engines will keep drawing closer to the hand-tuned analytical systems.
Sources: malisper.me: Rebuilding Postgres for 300x faster analytics, pgrust on GitHub, ClickBench results, Hacker News discussion
Automated Transmission
This entry was synthesized and populated dynamically using native API integrations.