INDEX ARCHIVE
Google2006/ 4M STUDY TIME/ORIGINAL PDF

Bigtable: A Distributed Storage System for Structured Data

The sparse, distributed, persistent multidimensional sorted map powering Google Search, Gmail, and Google Maps.

AUTHORS: Fay Chang, Jeffrey Dean, Sanjay Ghemawat, et al.

CORE ARCHITECTURAL BREAKTHROUGH

"Model data as a sparse map indexed by (row:string, column:string, time:int64) -> string, backed by Log-Structured Merge (LSM) SSTables on top of GFS."

WHY MODERN SYSTEMS STILL DEPEND ON IT

Directly inspired Apache HBase and Cassandra. Introduced the industry to Column Families, MemTables, Bloom Filters, and SSTables.

KEY PROBLEMS SOLVED

01.Traditional relational databases could not scale to petabytes across thousands of machines.
02.Google needed to store the entire world wide web crawl history with timestamped revisions per URL without schema locks.
03.Workloads required sub-millisecond point lookups alongside massive parallel batch scans for MapReduce.

DIRECT MODERN SUCCESSORS

Apache HBaseOpen-source Java clone built on top of Hadoop HDFS and ZooKeeper.
Apache CassandraCombined Bigtable's data model and SSTables with Dynamo's peer-to-peer ring topology.
RocksDB / LevelDBStandalone single-node embedded storage engines using Bigtable's SSTable and MemTable architecture.