-
Notifications
You must be signed in to change notification settings - Fork 0
Write‐Ahead Logging
Welcome to the write-ahead-logging wiki!
A Write-Ahead Log (WAL) is an append-only file where database systems record changes before applying them to the main storage files.
| Dimension | Matters |
|---|---|
| Durability and Atomicity: | It ensures that committed transactions are never lost, fulfilling key ACID properties even if the server crashes suddenly. |
| Crash Recovery: | If a power failure or system crash happens, the database reads the Write-ahead logging file and replays or rolls back the operations to restore data integrity. |
| Better Performance: | Bypassing immediate disk writes for every single commit speeds up database queries and transactions. |
-
Log First: The system writes every modification or transaction command to the sequential log on disk before touching the actual database tables or indexes.
-
Sequential I/O: Writing to an append-only log uses fast, sequential disk operations instead of slow, random writes across different data pages.
-
Delayed Updates: Actual data files are updated in memory first and flushed to disk asynchronously by background processes later.
The primary tradeoff of using a Write-Ahead Log (WAL) is balancing maximum data safety against the cost of writing data twice. Here is a direct comparison of the pros and cons:
- Fast Transaction Commits: Writing sequentially to the end of a log file is significantly faster than performing random disk writes across multiple database tables and indexes.
- Guaranteed Durability (ACID): If the system crashes or loses power, the database can fully recover by replaying the log (roll-forward) or undoing uncommitted changes (roll-back).
- Reduced Disk I/O Bottlenecks: Modifying data pages happens in memory first. The database can delay flushing these large, scattered data blocks to disk, batching them for background processes.
- Point-in-Time Recovery (PITR): By archiving WAL files alongside a periodic baseline backup, you can restore a database to its exact state at any specific microsecond in the past.
- Double Writing: Every change must be written twice—first to the WAL and later to the actual database storage files.
- Storage Overhead: WAL files can grow rapidly in write-heavy environments, requiring careful management, truncation, and log rotation to avoid running out of disk space.
- Recovery Delay: If a database crashes after a long period without saving its main data files to disk (checkpointing), the system startup time will be delayed while it processes a massive log file.
- Implementation Complexity: Designing a bulletproof WAL requires highly complex logic to handle edge cases like partial disk writes or corrupt log sectors during a crash.
How to use several advanced architectural solutions to mitigate double writing and storage overhead.