Surviving a restart
An RDB reader and writer, plus SAVE and BGSAVE.
You will build
an RDB reader and writer, plus SAVE and BGSAVE
You will understand
the RDB binary format, why you implement reading before writing, and how to convert an absolute deadline into a monotonic one
- roughly 280 lines across two new classes
- Stages 31 and 32
Where we are going
Your server loads a snapshot that real redis-server wrote:
Loaded 4 of 5 keys from /tmp/rdbtest/dump.rdb
And real redis-server loads a snapshot your server wrote:
$ redis-server --port 6402 --dir /tmp/rdbout
$ redis-cli -p 6402 keys '*'
1) "city"
2) "counter"
3) "later"
4) "name"
flowchart LR
A[Part 10<br/>config and CLI args] --> B[Stage 31<br/>read an RDB file]
B --> C[Stage 32<br/>write one]
C --> D([Part 12<br/>replication])
Stage 31, reading an RDB file
Goal. Start the server pointed at a snapshot and have the keys already there.
Why reading before writing
Writing your own format and reading it back proves nothing. A bug in the writer that the reader mirrors passes every test you can invent.
Reading a file that real Redis produced is a claim that can actually fail, and it forces you to implement the format as specified rather than as convenient.
So generate a snapshot with the real server first:
redis-server --port 6401 --dir /tmp/rdbtest
redis-cli -p 6401 set name Subash
redis-cli -p 6401 set counter 12345
redis-cli -p 6401 set later here ex 3600
redis-cli -p 6401 set soon gone px 100
redis-cli -p 6401 save
Keep that file. Checking it into your test resources means the parser stays honest as it changes.
The format
Binary, and mercifully simple at this level:
"REDIS0011" 9-byte magic and version
FA <str> <str> auxiliary field: redis-ver, ctime, used-mem
FE <length> select database
FB <len> <len> hash table size hints
FC <8 bytes LE> the next key expires at this unix millisecond
FD <4 bytes LE> or this unix second
00 <str> <str> a string key and value
FF <8 bytes> end of file, then a CRC64 checksum
Everything is a stream of opcodes, read until FF. Unknown opcodes are value types. 02 is a set,
04 a hash. This reader refuses them with a clear message rather than guessing.
All of this lives in one new class, RdbReader, which reads from a byte array and knows nothing
about sockets or files beyond opening one. Same rule as the RESP parser in Part 3: the thing that
parses should never be the thing that reads.
Length encoding is the interesting part
Every length and every string starts with a byte whose top two bits say how to read it:
00xxxxxx 6-bit length, 0 to 63
01xxxxxx yyyyyyyy 14-bit length
10000000 + 4 bytes 32-bit length, big endian
11xxxxxx not a length at all, a special encoding
That last case is why readLength returns a negative number. 11 means the low bits are an encoding
marker, so the caller has to switch behaviour rather than allocate a string.
C0 int8 follows stored as a number, returned as its text
C1 int16
C2 int32
C3 LZF-compressed
// the top two bits of the first byte say how the length itself is encoded
private long readLength() throws IOException {
int first = readByte();
int type = (first & 0xC0) >> 6;
return switch (type) {
case 0 -> first & 0x3F;
case 1 -> ((long) (first & 0x3F) << 8) | readByte();
case 2 -> first == 0x81 ? readBigEndian(8) : readBigEndian(4);
// 0xC0 marks a special encoding, which only readString knows how to handle
default -> -(first & 0x3F) - 1;
};
}
SET counter 12345 is stored as C1 plus two bytes, not as the five characters 12345. Redis does
that automatically for any value that looks like an integer, so a reader that only handles plain
strings appears to work until the first numeric key. The checked-in fixture contains one on purpose.
LZF compression is refused with an explicit message. Real Redis only compresses strings over 20 bytes, so short-valued files load fine. A documented limit beats a silent truncation.
Endianness is not consistent
Lengths are big endian. Expiry timestamps are little endian. Both in the same file.
That is not a mistake in the spec. The timestamps are memory dumps of a C long long on a
little-endian machine, while the lengths are written byte by byte. Worth knowing before you spend an
hour debugging a deadline in the year 55000.
Expiry is absolute, our clock is relative
The file stores an absolute unix millisecond. Our store deals in monotonic deadlines from
nanoTime, deliberately, so that clock changes cannot resurrect keys.
Converting is the bridge:
// used when loading an rdb file: the deadline arrives as an absolute wall-clock instant,
// which has to be turned into one on our monotonic clock
public void restore(String key, String value, long expiresAtEpochMillis) {
if (expiresAtEpochMillis == 0) {
set(key, value);
return;
}
long remaining = expiresAtEpochMillis - System.currentTimeMillis();
if (remaining <= 0) {
// already expired when the file was loaded, so it never enters the keyspace
return;
}
set(key, value, remaining);
}
The dropped case is why the log says “4 of 5”. A snapshot is not a promise that every key in it is still alive.
Failure is not fatal
A corrupt or unreadable snapshot logs and the server starts empty. A missing file is not even an error, because that is a first boot.
The alternative, refusing to start, means one bad file takes down a server that could have served traffic. Redis makes the same call.
Run it
$ java -cp target/classes com.example.redis.Main --dir /tmp/rdbtest --dbfilename dump.rdb
Loaded 4 of 5 keys from /tmp/rdbtest/dump.rdb
$ redis-cli keys '*'
1) "city"
2) "counter"
3) "later"
4) "name"
$ redis-cli get counter
"12345"
$ redis-cli ttl later
(integer) 3592
$ redis-cli get soon
(nil)
Notice that
Four of five. soon had a 100ms TTL and expired while the server was not running, so it was read
from the file and then dropped rather than restored.
counter came back as "12345" even though it was stored int-encoded. And later kept its
deadline across a process restart, converted from absolute to monotonic on the way in.
Try it yourself
- Delete the file and start the server. It starts empty with no error, because that is a first boot.
- Truncate the file to 20 bytes and start again. It logs and starts empty rather than refusing to run.
- Store a list in real Redis, save, then load the file with your server. It refuses the unsupported type with a clear message.
What usually goes wrong
A numeric value comes back wrong or throws. You are treating C0, C1, and C2 as lengths
rather than integer encodings.
Expiry lands in the far future. You read the timestamp big endian.
Stage 32, writing an RDB file
Goal. SAVE and BGSAVE, producing a file real Redis can load.
The claim that matters
$ redis-server --port 6402 --dir /tmp/rdbout --dbfilename dump.rdb
$ redis-cli -p 6402 keys '*'
1) "city"
2) "counter"
3) "later"
4) "name"
$ redis-cli -p 6402 ttl later
(integer) 3598
Another implementation refusing your file is a test that can fail. Your own reader accepting it is not.
The checksum shortcut
An RDB file ends with FF and eight bytes of CRC64. Computing it needs the Jones polynomial and a
256-entry table. Perfectly doable, and unnecessary here, because Redis skips verification when the
checksum is zero. That is documented behaviour for files written with checksums disabled.
So we write eight zero bytes and get a loadable file. The ponytail: note in the code says what was
skipped and why, which is the difference between a shortcut and a bug.
Writing atomically
The writer is its own class, RdbWriter, with one public method that returns bytes and one that
puts them on disk. Splitting those two matters more than it looks, because Part 12 needs exactly
those bytes without a file ever existing.
public static void write(Path file, List<RdbReader.Record> records) throws IOException {
byte[] snapshot = toBytes(records);
...
}
// the same bytes a replica receives during a full resync
public static byte[] toBytes(List<RdbReader.Record> records) {
...
}
A snapshot half-written when the process dies is worse than no snapshot, because the old good file is gone and the new one is unreadable.
Path temp = Files.createTempFile(directory, "dump", ".rdb");
Files.write(temp, snapshot);
Files.move(temp, file, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.ATOMIC_MOVE);
Same directory matters. A move across filesystems is a copy, and copies are not atomic. Either the old file is there or the new one is. A reader never sees a partial write.
SAVE and BGSAVE
SAVE blocks the calling client until the file is on disk. In real Redis it blocks the entire
server, which is why it carries a warning in the docs. Ours blocks one connection.
BGSAVE returns immediately and writes on another thread. One subtlety in the split:
if (background) {
// the snapshot is taken now, only the write is deferred
Thread.ofVirtual().start(() -> writeQuietly(file, records));
return RespWriter.simpleString("Background saving started");
}
The data is captured before the thread starts. Otherwise the background thread would read a keyspace that has moved on, and the file would hold a mixture of two moments. Real Redis forks the process for the same reason, getting a copy-on-write snapshot of memory at one instant.
Only strings are written
Lists, hashes, sets, sorted sets, and streams are skipped. That is a real limitation, and worth being blunt about rather than quietly dropping data.
Writing them in a form real Redis can read means implementing listpack and quicklist encodings, compact serialised forms with their own headers, element encodings, and length prefixes. That is a stage of its own, and it buys nothing conceptually over what the string path already taught.
The store’s snapshotStrings catches WrongTypeException per key and moves on, so a keyspace full of
lists saves cleanly as an empty file rather than failing.
Run it
$ redis-cli set name Subash
$ redis-cli set later here ex 3600
$ redis-cli rpush ignored a b
$ redis-cli save
OK
Restart your server pointed at the same directory:
Loaded 4 of 4 keys from /tmp/rdbout/dump.rdb
Then the real thing:
$ redis-server --port 6402 --dir /tmp/rdbout
$ redis-cli -p 6402 get name
"Subash"
$ redis-cli -p 6402 ttl later
(integer) 3598
Notice that
ignored is absent from both, as documented. And the TTL survived a full round trip: monotonic
deadline, to absolute epoch millisecond, to file, back to absolute, back to monotonic.
Try it yourself
- Save, then
xxdthe first 64 bytes of your file and findREDIS0011, theFAaux field, and theFE 00database selector. - Save with one key that has a TTL and one without. Find the
FCopcode and the eight little-endian bytes after it. - Load your file into real Redis, add a key there, save from Redis, and load it back into yours. A full round trip through both implementations.
What usually goes wrong
Real Redis says the file is corrupt. Usually a length written with the wrong form, or a missing
FF terminator.
The file loads but a TTL is wrong. Timestamps are little endian and lengths are big endian. It is easy to use one routine for both.
Go deeper: RDB against AOF
RDB is a point-in-time snapshot. Compact, fast to load, and you lose everything written since the last save.
AOF, the append-only file, logs every write command instead. Larger, slower to load, and it loses at most a second. Redis can run both.
Neither is implemented past RDB here. The concepts are in the snapshot path.
- Redis persistence
- RDB file format, annotated, the community reference everyone uses
- LZF compression, what we refuse to decode
What you built
A server whose data survives a restart, in a format that interoperates with the real thing in both directions.
Checkpoint
- Why does a round-trip test through your own reader prove so little?
- What do the top two bits of a length byte mean, and why can a length come back negative?
- Why is
SET counter 12345not stored as the text12345? - Why must an absolute expiry be converted rather than stored as-is?
- Why does
BGSAVEsnapshot the data before starting the thread?
Resources
Next
Two servers, one following the other.