Two servers
Replication. A second instance that follows the first.
You will build
replication. A second instance that follows the first
You will understand
the handshake, why one connection carries a snapshot and then a stream, what inverts when a connection stops being request and response, and why commands are re-encoded
- roughly 300 lines across two new classes
Where we are going
$ java -cp target/classes com.example.redis.Main --port 6410
$ java -cp target/classes com.example.redis.Main --port 6411 --replicaof "localhost 6410"
Replicating from localhost:6410
$ redis-cli -p 6410 set live yes # write to the master
OK
$ redis-cli -p 6411 get live # read from the replica
"yes"
Nobody told the replica about that key. It arrived on its own.
flowchart LR
A[Part 11<br/>one server, persistent] --> B[Stage 33<br/>handshake]
B --> C[snapshot transfer]
C --> D[command stream]
Stage 33, replication
Goal. A replica that catches up and then keeps up.
Two problems, in order
A replica connecting to a master with a million keys needs those keys before anything else makes sense. That is a bulk transfer.
After that it needs every subsequent write, in order, forever. That is a stream.
Redis solves both with one connection that changes mode partway through. A handshake, then a snapshot, then an endless run of commands on the same socket. Nothing reconnects.
Two new classes carry it. Replication on the master side holds the replica registry, the
replication ID, and the offset. ReplicaClient on the replica side runs the handshake and then the
follow loop, on its own virtual thread so the server can still serve its own clients.
The handshake
Four exchanges, each waiting for its reply:
sequenceDiagram
participant R as Replica
participant M as Master
R->>M: PING
M-->>R: +PONG
R->>M: REPLCONF listening-port 6411
M-->>R: +OK
R->>M: REPLCONF capa psync2
M-->>R: +OK
R->>M: PSYNC ? -1
M-->>R: +FULLRESYNC <replid> 0
M-->>R: $<len>\r\n<rdb bytes>
Note over R,M: from here the connection is push-only
PSYNC ? -1 means “I have nothing, send me everything”. The ? is a replication ID the replica does
not know yet, and -1 an offset it does not have. A reconnecting replica sends real values and can
get a partial resync instead, which is not implemented here and is the reason the protocol carries
them.
The REPLCONF exchanges look pointless for a server that ignores them. They are not. The master
learns which port the replica listens on, which real Redis reports in INFO replication, and
capability negotiation is how the protocol evolved without breaking old replicas.
The snapshot is framed oddly
$<length>\r\n<raw rdb bytes>
Like a bulk string, and with no trailing CRLF. Parse it as a normal bulk string and you either hang waiting for two bytes that never come, or eat the first two bytes of the first replicated command.
That is why the replica reads the handshake and the snapshot by hand, byte at a time, and only
switches to RespParser once the stream turns into commands. Mode changes mid-connection are where
framing bugs live.
The payload is exactly what SAVE writes, so the RDB writer from Part 10 was already the hard part
of this stage.
Here is ReplicaClient.loadSnapshot, the part that reads it:
// the snapshot is framed like a bulk string but has no trailing CRLF
private void loadSnapshot(InputStream in) throws IOException {
String header = readLine(in);
if (!header.startsWith("$")) {
throw new IOException("expected an rdb payload, got '" + header + "'");
}
int length = Integer.parseInt(header.substring(1));
byte[] snapshot = in.readNBytes(length);
if (snapshot.length != length) {
throw new IOException("master closed the connection during the snapshot");
}
for (RdbReader.Record record : RdbReader.readBytes(snapshot)) {
store.restore(record.key(), record.value(), record.expiresAtEpochMillis());
}
}
After the handshake, the connection inverts
An ordinary connection is request and response. This one becomes push-only:
master → SET k v no reply expected
master → SET a b no reply expected
master → REPLCONF GETACK * ← the one exception, the replica must answer
Two consequences in the code.
On the master, a connection that sent PSYNC stops being a client and joins a replica list. It uses
the same ClientSession machinery pub/sub uses, a session the server writes into without being
asked, with the same synchronized send keeping writers from interleaving.
On the replica, commands from the master are executed and the replies thrown away. A replica that
answered every propagated SET with +OK would be talking to a master that is not listening, and
those bytes would arrive as garbage at the front of its next command.
String[] command;
while ((command = parser.next()) != null) {
byte[] reply = dispatcher.execute(command, session);
// a replica answers GETACK and stays silent about everything else
if (command.length > 1 && command[0].equalsIgnoreCase("REPLCONF")
&& command[1].equalsIgnoreCase("GETACK")) {
out.write(reply);
out.flush();
}
// the master re-encodes canonically, so re-encoding gives the same byte count
session.advanceReplicaOffset(Replication.encode(command).length);
}
What gets propagated
Writes only. GET, EXISTS, and TYPE change nothing, so sending them wastes bandwidth and replica
CPU.
Commands are re-encoded rather than forwarded verbatim. A client may write set k v in lowercase or
with odd spacing, and every replica should see one canonical form. It also gives a byte count for the
offset that both sides compute identically, which is what makes the replica’s REPLCONF ACK <offset>
comparable to the master’s.
SET k v propagates even when k already held v. Redis does the same, because deciding a write
was a no-op is more expensive than sending it.
Offsets
The master counts bytes it has sent. The replica counts bytes it has processed. Equal offsets mean caught up.
REPLCONF GETACK * asks a replica where it is, and it answers REPLCONF ACK <offset>. That is how
WAIT knows how many replicas have a given write.
My WAIT returns the connected replica count immediately rather than blocking for acknowledgements.
An honest simplification. It is correct when no writes are in flight, which is the case at a
redis-cli prompt, and wrong under load. Named as such rather than pretended otherwise.
INFO becomes meaningful
role:master | role:slave
connected_slaves:1
master_replid:8f3a... 40 hex characters
master_repl_offset:139
role is how tooling discovers topology, and it is the field Part 10’s INFO flagged as the first
thing replication would change.
Run it
Master first, with a write before the replica exists:
$ java -cp target/classes com.example.redis.Main --port 6410 --dir /tmp/repl/master
$ redis-cli -p 6410 set before handshake
OK
Then the replica:
$ java -cp target/classes com.example.redis.Main --port 6411 --replicaof "localhost 6410"
Replicating from localhost:6410
$ redis-cli -p 6411 get before
"handshake"
Live propagation:
$ redis-cli -p 6410 set live yes
$ redis-cli -p 6410 rpush items a b
$ redis-cli -p 6411 get live
"yes"
$ redis-cli -p 6411 lrange items 0 -1
1) "a"
2) "b"
Topology:
$ redis-cli -p 6410 info | grep -E 'role|slaves|offset'
role:master
connected_slaves:1
master_repl_offset:139
$ redis-cli -p 6411 info | grep role
role:slave
Notice that
before was written to the master before the replica existed, and it arrived in the snapshot.
live was written after, and it arrived in the stream. Two different mechanisms on one connection.
And items replicated even though RPUSH is a list command, while lists are not in the RDB snapshot
at all. Bulk transfer and streaming have different coverage. Worth understanding rather than being
surprised by.
Try it yourself
- Kill the replica and write ten keys to the master. Restart the replica. It gets all of them, in the snapshot, because a full resync is the only resync we implement.
- Run
redis-cli -p 6411 get livewhile watching the master’smaster_repl_offsetclimb after each write. - Write a list to the master, then kill and restart the replica. The list is gone, because it travelled in the stream and the snapshot cannot carry it. That is Part 11’s limitation showing through.
What usually goes wrong
The replica hangs after PSYNC. You parsed the RDB payload as a normal bulk string and are
waiting for a CRLF that never arrives.
The first replicated command fails to parse. Same cause, other direction. You consumed two bytes too many after the snapshot.
The master’s log fills with protocol errors. The replica is replying to propagated commands.
Go deeper: what a real replica does that this one does not
Partial resync. A reconnecting replica sends its last known replication ID and offset. If the master still has those bytes in a backlog buffer, it sends only the missing commands rather than the whole keyspace. A brief network blip then costs a few kilobytes rather than a gigabyte.
Blocking WAIT. Real WAIT numreplicas timeout sends GETACK and blocks until enough replicas
acknowledge the current offset.
Chained replication. A replica can itself have replicas.
Expiry handling. A replica does not expire keys on its own clock. The master sends an explicit
DEL when a key expires, so both sides agree.
What you built
Across twelve posts: a TCP server, the RESP protocol in both directions, six data types, key expiry, transactions with optimistic locking, pub/sub, blocking operations, streams, persistence that interoperates with real Redis, and replication between two instances.
67 commands. Roughly 2,500 lines of plain Java, no dependencies outside JUnit.
Checkpoint
- Why does the connection need a snapshot and a stream rather than one or the other?
- What breaks if you parse the RDB payload as an ordinary bulk string?
- Why must the replica discard replies to propagated commands?
- Why re-encode commands instead of forwarding the client’s bytes?
- Why does a list replicate even though lists are missing from the snapshot?
What is left, if you want more
Consumer groups. XGROUP, XREADGROUP, and XACK turn a stream into a work queue with
per-consumer acknowledgement.
Partial resync. The backlog buffer described above.
AOF. Log every write instead of snapshotting.
Set algebra. SINTER, SUNION, SDIFF.
Cluster mode. Hash slots, redirects, and gossip. A different project wearing the same name.
None of them teaches a mechanism this series has not already covered, which is a reasonable place to stop.