Do We Need Specialized Graph Databases? Re-evaluating the RDBMS for Social Networks

Do We Need Specialized Graph Databases? Benchmarking Real-Time Social Networking Applications

Anil Pacaci, Alice Zhou, Jimmy Lin, M Tamer, Özsu David, R Cheriton
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a benchmarking study of graph data management systems using a Kafka-integrated LDBC Social Network Benchmark (SNB). It compares specialized graph databases (Neo4j, TitanDB), RDF stores (Virtuoso), and traditional RDBMSes (Postgres) under real-time interactive workloads.

TL;DR

In the era of "Big Data," the industry often assumes that graph-structured problems necessitate specialized graph databases. However, this paper challenges that notion by benchmarking specialized systems against traditional RDBMSes. Using an enhanced LDBC benchmark that incorporates real-time streaming via Kafka, the researchers found that Postgres and Virtuoso (RDBMS) often outperform specialized graph engines like Neo4j and TitanDB, particularly when the latter are accessed through abstraction layers like Gremlin.

Problem & Motivation: The Specialized System Hype

The rise of online social networks (OSNs) created a massive demand for managing entities and their relationships. Specialized graph databases emerged, promising "index-free adjacency" and treating relationships as "first-class citizens."

The authors identify a critical gap: most benchmarks focus on batch processing (OLAP) or simple lookups, ignoring the real-time transactional nature (OLTP) of social media—where a stream of "likes" and "follows" must process concurrently with complex queries. They question whether the 30+ years of optimization in RDBMS technology can be so easily discarded.

Methodology: Bringing Real-Time Streams to Benchmarking

The core contribution is a revamped benchmarking architecture. The authors took the LDBC Social Network Benchmark (SNB) and integrated Apache Kafka.

Benchmarking Architecture

  • Kafka Integration: Instead of scheduled updates, a Kafka queue simulates the continuous, high-throughput update streams typical of platforms like Twitter or LinkedIn.
  • Unified Interfaces: They implemented queries in Gremlin (for TinkerPop3-compliant graph databases) and SQL (for RDBMSes) to ensure a fair comparison across systems including Neo4j, TitanDB, Postgres, and Virtuoso.

Experiments: Performance Showdown

The results were surprising. In single-node settings where the entire dataset could fit in memory, the RDBMSes held a clear edge.

1. The Cost of Abstraction

One of the most significant findings was the massive overhead of the Gremlin query language. In Neo4j, using native Cypher was up to 100x faster than Gremlin. This occurs because the Gremlin server often breaks down complex graph operations into many small, unoptimized requests to the underlying engine.

2. Read Latency Comparison

As shown in the table below, Postgres (SQL) and Virtuoso (SQL) dominated point lookups and one-hop traversals. While Neo4j showed stability across data scales (proving the benefit of index-free adjacency), it rarely beat the raw speed of a well-indexed relational table for interactive-scale queries.

Performance Comparison Table

3. Throughput under Real-Time Load

Under a concurrent workload of reads and writes, Postgres maintained the highest write throughput, significantly exceeding TitanDB and even Neo4j (Cypher). Neo4j's update performance suffered from periodic drops during checkpointing, while Postgres offered a more stable profile.

Aggregate Throughput

Deep Insight: Why RDBMS Still Wins

The effectiveness of RDBMSes in this study boils down to several factors:

  • Maturity: Decades of query optimization and robust transactional handling.
  • Query Planning: Declarative languages like SQL allow the engine to find the global optimum for a query, whereas procedural traversals (like Gremlin) can fall into suboptimal patterns.
  • Indexing: While graph databases tout "index-free" access, high-performance RDBMS B-Trees and Column-stores are remarkably efficient at the scale of typical interactive social queries.

Conclusion & Future Look

The paper concludes that TinkerPop3/Gremlin is not yet production-ready for high-performance scenarios due to its extreme overhead. More importantly, for many organizations, a specialized graph database may be unnecessary overhead; Postgres is often "good enough" and sometimes even faster.

Limitations: This study focused on single-node deployments. The authors acknowledge that the "scale-out" (distributed) capabilities of specialized graph databases might change the narrative for datasets that cannot fit on a single machine—a key area for future research.

Find Similar Papers

Try Our Examples

  • Search for recent studies comparing Relational Databases and Graph Databases in distributed or multi-node settings beyond single-node performance.
  • Which paper first introduced the LDBC Social Network Benchmark (SNB) and how has the interactive workload specification evolved since its inception?
  • Are there recent optimizations in the Apache TinkerPop/Gremlin stack that address the high overhead documented in this 2017 study?
Contents
Do We Need Specialized Graph Databases? Re-evaluating the RDBMS for Social Networks
1. TL;DR
2. Problem & Motivation: The Specialized System Hype
3. Methodology: Bringing Real-Time Streams to Benchmarking
4. Experiments: Performance Showdown
4.1. 1. The Cost of Abstraction
4.2. 2. Read Latency Comparison
4.3. 3. Throughput under Real-Time Load
5. Deep Insight: Why RDBMS Still Wins
6. Conclusion & Future Look