What redis-cli actually sends
A RESP parser, a RESP writer, and the project's first tests.
You will build
a RESP parser, a RESP writer, and the project's first tests
You will understand
the Redis wire protocol, why length prefixes beat delimiters, and why a parser must never do I/O
- roughly 200 lines across two new classes
- Stages 08 to 10
Where we are going
The real Redis client understands your server:
$ redis-cli ping
PONG
flowchart LR
A[Part 2<br/>bytes framed by newline] --> B[Stage 08<br/>read the protocol]
B --> C[Stage 09<br/>RespParser]
C --> D[Stage 10<br/>RespWriter]
D --> E([Part 4<br/>commands])
Stage 08, read the protocol off the wire
Goal. See exactly what a Redis client sends, using the server you already have.
The idea
Your echo server from Part 2 is the perfect instrument. It prints every request it frames, so pointing the real client at it shows you the protocol without reading a spec first.
Run it
Start your server, then in another terminal:
redis-cli -t 2 ping
redis-cli gets nonsense back, because it does not understand an echo server. Your server log
shows this:
Request: *1\r
Request: $4\r
Request: ping\r
Again with an argument:
redis-cli -t 2 echo "hey there"
Request: *2\r
Request: $4\r
Request: echo\r
Request: $9\r
Request: hey there\r
And with three:
redis-cli -t 2 set mykey myvalue
Request: *3\r
Request: $3\r
Request: set\r
Request: $5\r
Request: mykey\r
Request: $7\r
Request: myvalue\r
Notice that
Three things, and each one changes your code.
Every line ends with \r. Your framing splits on \n and leaves the carriage return behind. RESP
terminates every part with CRLF, written \r\n.
The command name arrives lowercase. redis-cli sent ping, not PING. Matching has to be
case-insensitive.
$9 for hey there counts the space. That number is a byte count, not a word or character count.
The format
Every command a client sends is an array of bulk strings. There are no exceptions.
*<number-of-arguments>\r\n
$<length-of-arg-1>\r\n<arg-1>\r\n
$<length-of-arg-2>\r\n<arg-2>\r\n
So ECHO "hey there" on the wire is literally:
*2\r\n$4\r\nECHO\r\n$9\r\nhey there\r\n
Five types, identified by the first byte:
| Byte | Type | Example | Used for |
|---|---|---|---|
+ | simple string | +OK\r\n | short status replies |
- | error | -ERR unknown command\r\n | failures |
: | integer | :42\r\n | counts |
$ | bulk string | $5\r\nhello\r\n | any value, binary safe |
* | array | *2\r\n... | lists of the above |
Two special forms matter and you cannot guess them:
$-1\r\n null bulk string, GET on a missing key
$0\r\n\r\n empty string, a real stored value of length zero
Those look similar and mean different things. Confusing them is the most common RESP bug, and it earns a test in every class that touches it.
Try it yourself
- Run
redis-cli lpush fruits apple bananaagainst your echo server and read the log. Predict the shape before you run it. - Send a command by hand:
printf '*1\r\n$4\r\nPING\r\n' | nc localhost 6379 | xxd - Store a value containing a newline and note that the
$length counts it as a byte.
Why length prefixes beat delimiters
Part 2 ended with two problems. A payload can never contain the delimiter, and a client that never sends one grows your buffer forever. Length prefixes fix both.
$7\r\nab\r\ncd\r\n
└── 7 bytes: a b \r \n c d
The payload contains \r\n and nothing breaks, because nobody is scanning for it. That is what
binary safe means, and it is why Redis can store a JPEG in a string.
It also kills the denial of service. The moment you read $7 you know exactly how many bytes to
expect, so you can reject an absurd length before allocating for it. A delimiter parser has to keep
buffering and hope.
Stage 09, the parser
Goal. Turn RESP bytes into a String[], correctly, including when only half of it has arrived.
The idea: what should a parser consume?
Three options, and the reasoning matters more than the answer.
String is wrong. Converting bytes to a String before parsing forces a charset decision on
data that may be arbitrary binary. Decode after the length has cut out the exact bytes, never
before.
InputStream is tempting and wrong differently. A stream parser calls read() when it needs
more bytes, so the parser blocks. That couples parsing to I/O: you cannot test it without a socket,
and “half a command arrived” becomes a parked thread instead of a return value.
byte[] with an explicit buffer works. The parser never does I/O. It is handed bytes and asked
whether a complete command is present.
The rule worth keeping past this project: the thing that parses should never be the thing that reads.
That gives a three-outcome contract:
flowchart LR
N["next()"] --> A["String[]<br/>a complete command"]
N --> B["null<br/>not enough bytes yet"]
N --> C["throws<br/>these bytes are garbage"]
null is not an error. It is the normal state of a stream that has delivered half a command.
Telling incomplete apart from invalid is the entire difficulty of this stage. Everything else is
bookkeeping.
The code
src/main/java/com/example/redis/RespParser.java, complete:
package com.example.redis;
import java.nio.charset.StandardCharsets;
import java.util.Arrays;
// Parses RESP bytes into commands. Owns no socket, callers feed it bytes.
public class RespParser {
// same cap real redis uses for proto-max-bulk-len
private static final int MAX_BULK_LENGTH = 512 * 1024 * 1024;
private static final int MAX_ARRAY_SIZE = 1024 * 1024;
private byte[] buffer = new byte[0];
// cursor into buffer, only meaningful during a next() call
private int pos;
public void append(byte[] data, int length) {
byte[] grown = Arrays.copyOf(buffer, buffer.length + length);
System.arraycopy(data, 0, grown, buffer.length, length);
buffer = grown;
}
// returns null when a complete command hasn't arrived yet
public String[] next() {
pos = 0;
String[] command = parseArray();
if (command == null) {
return null;
}
// consume exactly what we parsed, keep the rest for next time
buffer = Arrays.copyOfRange(buffer, pos, buffer.length);
return command;
}
private String[] parseArray() {
if (pos >= buffer.length) {
return null;
}
if (buffer[pos] != '*') {
throw new IllegalStateException("expected array, got byte " + buffer[pos]);
}
pos++;
Integer count = readNumber();
if (count == null) {
return null;
}
if (count < 0 || count > MAX_ARRAY_SIZE) {
throw new IllegalStateException("bad array size " + count);
}
String[] args = new String[count];
for (int i = 0; i < count; i++) {
String arg = parseBulkString();
if (arg == null) {
return null;
}
args[i] = arg;
}
return args;
}
private String parseBulkString() {
if (pos >= buffer.length) {
return null;
}
if (buffer[pos] != '$') {
throw new IllegalStateException("expected bulk string, got byte " + buffer[pos]);
}
pos++;
Integer length = readNumber();
if (length == null) {
return null;
}
if (length < 0 || length > MAX_BULK_LENGTH) {
throw new IllegalStateException("bad bulk length " + length);
}
// payload plus its trailing CRLF must all be here
if (pos + length + 2 > buffer.length) {
return null;
}
String value = new String(buffer, pos, length, StandardCharsets.UTF_8);
pos += length + 2;
return value;
}
// reads digits up to the next CRLF, null if that CRLF hasn't arrived
private Integer readNumber() {
int lineEnd = indexOfCrlf(pos);
if (lineEnd < 0) {
return null;
}
String digits = new String(buffer, pos, lineEnd - pos, StandardCharsets.UTF_8);
pos = lineEnd + 2;
try {
return Integer.parseInt(digits);
} catch (NumberFormatException e) {
throw new IllegalStateException("expected number, got '" + digits + "'");
}
}
private int indexOfCrlf(int from) {
for (int i = from; i + 1 < buffer.length; i++) {
if (buffer[i] == '\r' && buffer[i + 1] == '\n') {
return i;
}
}
return -1;
}
}
Setting pos = 0 at the top of next() and trimming the buffer only on success is what makes
partial input free. A failed parse leaves the buffer untouched, so the next attempt starts clean.
No rollback logic, no saved state.
The +2 in pos + length + 2 > buffer.length is the CRLF after the payload. Forget it and the
next command starts two bytes late, which surfaces as a bizarre parse error one command later.
new String(buffer, pos, length, UTF_8) decodes after the length has cut out the exact bytes.
Reverse that order and binary values break.
The length caps are the defence a delimiter parser could not have. $2000000000 from a hostile
client is rejected before anything is allocated.
Wire it in
handleClient loses all its framing:
RespParser parser = new RespParser();
byte[] buffer = new byte[1024];
int bytesRead;
while ((bytesRead = inputStream.read(buffer)) != -1) {
parser.append(buffer, bytesRead);
// one read may hold several commands, or none
String[] command;
while ((command = parser.next()) != null) {
System.out.println("Command: " + Arrays.toString(command));
outputStream.write(RespWriter.simpleString("PONG")); // stage 10 adds this
outputStream.flush();
}
}
The pending buffer, the newline scan, and the start cursor all disappear. They moved behind the
parser, along with the stray \r.
Tests start here
Stages 01 to 07 had no tests, deliberately. A test asserting ServerSocket.bind() binds tests the
JDK, not you. There was no branch, no transformation, and nc verified it faster than JUnit would.
The parser is the first pure function in the project. Bytes in, structure out, no sockets. It is also where bugs stop being visible by eye. Three tests justify the whole file:
@Test
void payloadMayContainCrlf() {
// the declared length decides where the value ends, not the delimiter
assertArrayEquals(new String[]{"ECHO", "a\r\nbc"},
parserWith("*2\r\n$4\r\nECHO\r\n$5\r\na\r\nbc\r\n").next());
}
@Test
void returnsNullUntilCommandComplete() {
RespParser parser = parserWith("*2\r\n$4\r\nECHO\r\n$5\r\nhel");
assertNull(parser.next());
byte[] rest = "lo\r\n".getBytes(StandardCharsets.UTF_8);
parser.append(rest, rest.length);
assertArrayEquals(new String[]{"ECHO", "hello"}, parser.next());
}
@Test
void parsesTwoCommandsFromOneAppend() {
RespParser parser = parserWith("*1\r\n$4\r\nPING\r\n*1\r\n$4\r\nPING\r\n");
assertArrayEquals(new String[]{"PING"}, parser.next());
assertArrayEquals(new String[]{"PING"}, parser.next());
assertNull(parser.next());
}
None of those is reachable by typing into nc, and each is a bug you would otherwise ship.
Run it
$ mvn test
[INFO] Tests run: 10, Failures: 0, Errors: 0
Then against the real client:
Command: [ping]
Command: [echo, hey there]
Command: [set, mykey, myvalue]
Notice that
No \r anywhere. If you still see one, something is splitting on \n.
Try it yourself
- Change
pos + length + 2topos + length + 1and run the tests. Watch which one fails and read the message. That is the “one command later” symptom in miniature. - Send a command in three separate writes with pauses. The parser returns
null,null, then the command. - Feed it
+OK\r\n, which is a reply rather than a command. It throws, because a client may only send arrays.
What usually goes wrong
StringIndexOutOfBoundsException on partial input. I hit exactly this:
int lineEnd = indexOfCrlf(pos);
if (lineEnd == 0) { // should be < 0
return null;
}
indexOfCrlf returns -1 when there is no CRLF yet. With == 0 the incomplete case fell through
to new String(buffer, pos, -1 - pos, ...), in the exact path the stage exists to handle. Five
tests failed at once and pointed straight at it.
The other one: everything works until a value contains a newline. That means you are scanning for
\r\n instead of using the declared length.
Go deeper: RESP2 against RESP3, and inline commands
Redis 6 introduced RESP3, adding maps, sets, doubles, booleans, and out-of-band push messages.
Clients opt in with HELLO 3. Without it, servers speak RESP2. Build RESP2. Replying to HELLO
with an error is legitimate, and clients fall back.
Redis also accepts inline commands, plain PING\r\n with no * or $, for typing into telnet.
Optional, and skipped here.
- Redis serialization protocol specification, read it once end to end
- RESP2 spec on GitHub
Stage 10, the writer
Goal. Turn Java values into RESP bytes.
The idea
The mirror image, in a separate class on purpose. After this stage, nothing outside these two files
touches a \r\n again.
flowchart LR
A[socket bytes] --> B[RespParser]
B --> C["String[]"]
C --> D[command logic]
D --> E[RespWriter]
E --> F[socket bytes]
style B fill:#e8f0fe
style E fill:#e8f0fe
When SET arrives in Part 5 it should decide "OK", not +OK\r\n. The moment protocol formatting
leaks into command code, every command becomes a place a missing CRLF can hide.
The code
src/main/java/com/example/redis/RespWriter.java:
package com.example.redis;
import java.io.ByteArrayOutputStream;
import java.nio.charset.StandardCharsets;
// Encodes Java values as RESP2 bytes. Owns no socket, callers write what it returns.
public class RespWriter {
private static final byte[] CRLF = "\r\n".getBytes(StandardCharsets.UTF_8);
private static final byte[] NULL_BULK = "$-1\r\n".getBytes(StandardCharsets.UTF_8);
// never build one of these from user data, it has no length prefix
public static byte[] simpleString(String value) {
return ("+" + value + "\r\n").getBytes(StandardCharsets.UTF_8);
}
public static byte[] error(String message) {
return ("-" + message + "\r\n").getBytes(StandardCharsets.UTF_8);
}
public static byte[] integer(long value) {
return (":" + value + "\r\n").getBytes(StandardCharsets.UTF_8);
}
public static byte[] bulkString(String value) {
if (value == null) {
return NULL_BULK;
}
// the declared length is bytes, not characters
byte[] payload = value.getBytes(StandardCharsets.UTF_8);
ByteArrayOutputStream out = new ByteArrayOutputStream();
out.writeBytes(("$" + payload.length + "\r\n").getBytes(StandardCharsets.UTF_8));
out.writeBytes(payload);
out.writeBytes(CRLF);
return out.toByteArray();
}
// elements are already encoded, so nesting costs nothing
public static byte[] array(byte[]... elements) {
ByteArrayOutputStream out = new ByteArrayOutputStream();
out.writeBytes(("*" + elements.length + "\r\n").getBytes(StandardCharsets.UTF_8));
for (byte[] element : elements) {
out.writeBytes(element);
}
return out.toByteArray();
}
}
Use payload.length after getBytes(UTF_8), never value.length(). "café" is 4 characters and
5 bytes. Declare $4 for 5 bytes and the client reads four, hits a stray byte where CRLF should
be, and every later reply on that connection is garbage. The failure surfaces one command after the
mistake.
array(byte[]...) takes already-encoded elements, so nesting is free. No recursive encoder, no
value-type hierarchy. That decision pays off in Part 8 when streams need arrays of arrays of
arrays.
Simple strings are never built from user data. They have no length prefix, so an embedded newline
ends the reply early and desynchronises the connection. +OK and +PONG are yours. A stored value
is not.
The test that earns its keep
@Test
void bulkStringLengthIsBytesNotCharacters() {
// café is 4 characters but 5 bytes in UTF-8
assertEquals("$5\r\ncafé\r\n", encoded(RespWriter.bulkString("café")));
}
It is the only test here that fails against the obvious String-based implementation, and the
resulting bug is invisible until a user stores a non-ASCII value.
Worth adding too, a round trip through both classes:
@Test
void writerOutputParsesBackAsACommand() {
byte[] encodedCommand = RespWriter.array(
RespWriter.bulkString("ECHO"), RespWriter.bulkString("a\r\nbc"));
RespParser parser = new RespParser();
parser.append(encodedCommand, encodedCommand.length);
String[] command = parser.next();
assertEquals("ECHO", command[0]);
assertEquals("a\r\nbc", command[1]);
}
That is the test that catches the two classes disagreeing about the protocol.
Run it
$ mvn test
[INFO] Tests run: 23, Failures: 0, Errors: 0
$ redis-cli -t 2 ping
PONG
$ printf '*1\r\n$4\r\nPING\r\n' | nc localhost 6379 | xxd
00000000: 2b50 4f4e 470d 0a +PONG..
Notice that
2b is + and 0d 0a is CRLF. That is the first time the real client understood your server. It
printed PONG because it decoded a valid RESP simple string. Emit PONG\r\n without the + and
it reports a protocol error instead.
Try it yourself
- Remove the
+fromsimpleStringand runredis-cli ping. Read the error it gives you. - Encode a nested array:
array(integer(1), array(bulkString("x"))). Print it. Note that you did not write a recursive function. - Write
bulkStringthe naive way usingvalue.length(), then run thecafétest.
What you built
Two classes, 23 tests, and a hard boundary. Protocol on one side, meaning on the other. Your server speaks Redis in both directions.
It also replies PONG to everything, including GET. Part 4 fixes that.
Checkpoint
- Why must you not convert bytes to a
Stringbefore parsing? - Why is
nullthe right signal for incomplete input rather than an exception? - Why is scanning for
\r\nto end a bulk string wrong, given every bulk string ends with one? - What is the difference on the wire between
$-1,$0\r\n\r\n, and+\r\n? - What breaks, and when, if a reply declares the wrong length?
Resources
- RESP specification, the primary source
- RESP2 in redis-specifications
StandardCharsetsJavadoc
Next
The server stops replying PONG to everything and starts caring what you asked.