Lakebase Postgres branch-based restores for fast recovery at scale

Databricks introduces Lakebase Postgres, a managed database using decoupled compute and storage to enable branch-based restores. This architecture allows point-in-time recovery to complete in seconds rather than hours, even for 100 TB databases. The system treats restores as metadata operations instead of data copies, eliminating the need for expensive replicas or manual DBA intervention during recovery.

Cover image for Lakebase Postgres branch-based restores for fast recovery at scale

Managed OLTP restores have traditionally been very slow, especially for large databases, because compute and storage are bundled together and recovery requires provisioning a new instance, pulling a snapshot from object storage onto disk, and replaying WAL logs, all of which become slower and more costly as size grows. Conventional workarounds such as extra replicas or DBA‑led restores are expensive, risky, and still leave downtime of several hours, and they do not protect against bad writes that have already replicated. Lakebase Postgres changes this by decoupling compute from durable storage; the storage layer keeps the full history in object storage and a pageserver reconstructs pages from WAL, so a restore merely creates a timestamped branch as a metadata operation. This eliminates data copying and WAL replay, reducing restore time to seconds even for 100 TB databases and allowing automated agents to perform the recovery. A survey of developers with 1 TB+ production Postgres instances showed that 59 % had experienced critical failures, 30 % endured outages of three or more hours, and only 21 % recovered in under a minute, highlighting the business impact of slow restores.