Surviving a restart

An RDB reader and writer, plus SAVE and BGSAVE.

  • Part 11
  • advanced
  • about 80 minutes

You will build

an RDB reader and writer, plus SAVE and BGSAVE

You will understand

the RDB binary format, why you implement reading before writing, and how to convert an absolute deadline into a monotonic one

  • roughly 280 lines across two new classes
  • Stages 31 and 32

Where we are going

Your server loads a snapshot that real redis-server wrote:

Loaded 4 of 5 keys from /tmp/rdbtest/dump.rdb

And real redis-server loads a snapshot your server wrote:

$ redis-server --port 6402 --dir /tmp/rdbout
$ redis-cli -p 6402 keys '*'
1) "city"
2) "counter"
3) "later"
4) "name"
flowchart LR
    A[Part 10<br/>config and CLI args] --> B[Stage 31<br/>read an RDB file]
    B --> C[Stage 32<br/>write one]
    C --> D([Part 12<br/>replication])

Stage 31, reading an RDB file

Goal. Start the server pointed at a snapshot and have the keys already there.

Why reading before writing

Writing your own format and reading it back proves nothing. A bug in the writer that the reader mirrors passes every test you can invent.

Reading a file that real Redis produced is a claim that can actually fail, and it forces you to implement the format as specified rather than as convenient.

So generate a snapshot with the real server first:

redis-server --port 6401 --dir /tmp/rdbtest
redis-cli -p 6401 set name Subash
redis-cli -p 6401 set counter 12345
redis-cli -p 6401 set later here ex 3600
redis-cli -p 6401 set soon gone px 100
redis-cli -p 6401 save

Keep that file. Checking it into your test resources means the parser stays honest as it changes.

The format

Binary, and mercifully simple at this level:

"REDIS0011"          9-byte magic and version
FA <str> <str>       auxiliary field: redis-ver, ctime, used-mem
FE <length>          select database
FB <len> <len>       hash table size hints
FC <8 bytes LE>      the next key expires at this unix millisecond
FD <4 bytes LE>      or this unix second
00 <str> <str>       a string key and value
FF <8 bytes>         end of file, then a CRC64 checksum

Everything is a stream of opcodes, read until FF. Unknown opcodes are value types. 02 is a set, 04 a hash. This reader refuses them with a clear message rather than guessing.

All of this lives in one new class, RdbReader, which reads from a byte array and knows nothing about sockets or files beyond opening one. Same rule as the RESP parser in Part 3: the thing that parses should never be the thing that reads.

Length encoding is the interesting part

Every length and every string starts with a byte whose top two bits say how to read it:

00xxxxxx              6-bit length, 0 to 63
01xxxxxx yyyyyyyy     14-bit length
10000000 + 4 bytes    32-bit length, big endian
11xxxxxx              not a length at all, a special encoding

That last case is why readLength returns a negative number. 11 means the low bits are an encoding marker, so the caller has to switch behaviour rather than allocate a string.

C0  int8 follows        stored as a number, returned as its text
C1  int16
C2  int32
C3  LZF-compressed
// the top two bits of the first byte say how the length itself is encoded
private long readLength() throws IOException {
    int first = readByte();
    int type = (first & 0xC0) >> 6;

    return switch (type) {
        case 0 -> first & 0x3F;
        case 1 -> ((long) (first & 0x3F) << 8) | readByte();
        case 2 -> first == 0x81 ? readBigEndian(8) : readBigEndian(4);
        // 0xC0 marks a special encoding, which only readString knows how to handle
        default -> -(first & 0x3F) - 1;
    };
}

SET counter 12345 is stored as C1 plus two bytes, not as the five characters 12345. Redis does that automatically for any value that looks like an integer, so a reader that only handles plain strings appears to work until the first numeric key. The checked-in fixture contains one on purpose.

LZF compression is refused with an explicit message. Real Redis only compresses strings over 20 bytes, so short-valued files load fine. A documented limit beats a silent truncation.

Endianness is not consistent

Lengths are big endian. Expiry timestamps are little endian. Both in the same file.

That is not a mistake in the spec. The timestamps are memory dumps of a C long long on a little-endian machine, while the lengths are written byte by byte. Worth knowing before you spend an hour debugging a deadline in the year 55000.

Expiry is absolute, our clock is relative

The file stores an absolute unix millisecond. Our store deals in monotonic deadlines from nanoTime, deliberately, so that clock changes cannot resurrect keys.

Converting is the bridge:

// used when loading an rdb file: the deadline arrives as an absolute wall-clock instant,
// which has to be turned into one on our monotonic clock
public void restore(String key, String value, long expiresAtEpochMillis) {
    if (expiresAtEpochMillis == 0) {
        set(key, value);
        return;
    }

    long remaining = expiresAtEpochMillis - System.currentTimeMillis();
    if (remaining <= 0) {
        // already expired when the file was loaded, so it never enters the keyspace
        return;
    }
    set(key, value, remaining);
}

The dropped case is why the log says “4 of 5”. A snapshot is not a promise that every key in it is still alive.

Failure is not fatal

A corrupt or unreadable snapshot logs and the server starts empty. A missing file is not even an error, because that is a first boot.

The alternative, refusing to start, means one bad file takes down a server that could have served traffic. Redis makes the same call.

Run it

$ java -cp target/classes com.example.redis.Main --dir /tmp/rdbtest --dbfilename dump.rdb
Loaded 4 of 5 keys from /tmp/rdbtest/dump.rdb

$ redis-cli keys '*'
1) "city"
2) "counter"
3) "later"
4) "name"
$ redis-cli get counter
"12345"
$ redis-cli ttl later
(integer) 3592
$ redis-cli get soon
(nil)

Notice that

Four of five. soon had a 100ms TTL and expired while the server was not running, so it was read from the file and then dropped rather than restored.

counter came back as "12345" even though it was stored int-encoded. And later kept its deadline across a process restart, converted from absolute to monotonic on the way in.

Try it yourself

  1. Delete the file and start the server. It starts empty with no error, because that is a first boot.
  2. Truncate the file to 20 bytes and start again. It logs and starts empty rather than refusing to run.
  3. Store a list in real Redis, save, then load the file with your server. It refuses the unsupported type with a clear message.

What usually goes wrong

A numeric value comes back wrong or throws. You are treating C0, C1, and C2 as lengths rather than integer encodings.

Expiry lands in the far future. You read the timestamp big endian.


Stage 32, writing an RDB file

Goal. SAVE and BGSAVE, producing a file real Redis can load.

The claim that matters

$ redis-server --port 6402 --dir /tmp/rdbout --dbfilename dump.rdb
$ redis-cli -p 6402 keys '*'
1) "city"
2) "counter"
3) "later"
4) "name"
$ redis-cli -p 6402 ttl later
(integer) 3598

Another implementation refusing your file is a test that can fail. Your own reader accepting it is not.

The checksum shortcut

An RDB file ends with FF and eight bytes of CRC64. Computing it needs the Jones polynomial and a 256-entry table. Perfectly doable, and unnecessary here, because Redis skips verification when the checksum is zero. That is documented behaviour for files written with checksums disabled.

So we write eight zero bytes and get a loadable file. The ponytail: note in the code says what was skipped and why, which is the difference between a shortcut and a bug.

Writing atomically

The writer is its own class, RdbWriter, with one public method that returns bytes and one that puts them on disk. Splitting those two matters more than it looks, because Part 12 needs exactly those bytes without a file ever existing.

public static void write(Path file, List<RdbReader.Record> records) throws IOException {
    byte[] snapshot = toBytes(records);
    ...
}

// the same bytes a replica receives during a full resync
public static byte[] toBytes(List<RdbReader.Record> records) {
    ...
}

A snapshot half-written when the process dies is worse than no snapshot, because the old good file is gone and the new one is unreadable.

Path temp = Files.createTempFile(directory, "dump", ".rdb");
Files.write(temp, snapshot);
Files.move(temp, file, StandardCopyOption.REPLACE_EXISTING, StandardCopyOption.ATOMIC_MOVE);

Same directory matters. A move across filesystems is a copy, and copies are not atomic. Either the old file is there or the new one is. A reader never sees a partial write.

SAVE and BGSAVE

SAVE blocks the calling client until the file is on disk. In real Redis it blocks the entire server, which is why it carries a warning in the docs. Ours blocks one connection.

BGSAVE returns immediately and writes on another thread. One subtlety in the split:

if (background) {
    // the snapshot is taken now, only the write is deferred
    Thread.ofVirtual().start(() -> writeQuietly(file, records));
    return RespWriter.simpleString("Background saving started");
}

The data is captured before the thread starts. Otherwise the background thread would read a keyspace that has moved on, and the file would hold a mixture of two moments. Real Redis forks the process for the same reason, getting a copy-on-write snapshot of memory at one instant.

Only strings are written

Lists, hashes, sets, sorted sets, and streams are skipped. That is a real limitation, and worth being blunt about rather than quietly dropping data.

Writing them in a form real Redis can read means implementing listpack and quicklist encodings, compact serialised forms with their own headers, element encodings, and length prefixes. That is a stage of its own, and it buys nothing conceptually over what the string path already taught.

The store’s snapshotStrings catches WrongTypeException per key and moves on, so a keyspace full of lists saves cleanly as an empty file rather than failing.

Run it

$ redis-cli set name Subash
$ redis-cli set later here ex 3600
$ redis-cli rpush ignored a b
$ redis-cli save
OK

Restart your server pointed at the same directory:

Loaded 4 of 4 keys from /tmp/rdbout/dump.rdb

Then the real thing:

$ redis-server --port 6402 --dir /tmp/rdbout
$ redis-cli -p 6402 get name
"Subash"
$ redis-cli -p 6402 ttl later
(integer) 3598

Notice that

ignored is absent from both, as documented. And the TTL survived a full round trip: monotonic deadline, to absolute epoch millisecond, to file, back to absolute, back to monotonic.

Try it yourself

  1. Save, then xxd the first 64 bytes of your file and find REDIS0011, the FA aux field, and the FE 00 database selector.
  2. Save with one key that has a TTL and one without. Find the FC opcode and the eight little-endian bytes after it.
  3. Load your file into real Redis, add a key there, save from Redis, and load it back into yours. A full round trip through both implementations.

What usually goes wrong

Real Redis says the file is corrupt. Usually a length written with the wrong form, or a missing FF terminator.

The file loads but a TTL is wrong. Timestamps are little endian and lengths are big endian. It is easy to use one routine for both.

Go deeper: RDB against AOF

RDB is a point-in-time snapshot. Compact, fast to load, and you lose everything written since the last save.

AOF, the append-only file, logs every write command instead. Larger, slower to load, and it loses at most a second. Redis can run both.

Neither is implemented past RDB here. The concepts are in the snapshot path.


What you built

A server whose data survives a restart, in a format that interoperates with the real thing in both directions.

Checkpoint

  1. Why does a round-trip test through your own reader prove so little?
  2. What do the top two bits of a length byte mean, and why can a length come back negative?
  3. Why is SET counter 12345 not stored as the text 12345?
  4. Why must an absolute expiry be converted rather than stored as-is?
  5. Why does BGSAVE snapshot the data before starting the thread?

Resources

Next

Two servers, one following the other.

Part 12: Two Servers →