What redis-cli actually sends

A RESP parser, a RESP writer, and the project's first tests.

  • Part 3
  • intermediate
  • about 60 minutes

You will build

a RESP parser, a RESP writer, and the project's first tests

You will understand

the Redis wire protocol, why length prefixes beat delimiters, and why a parser must never do I/O

  • roughly 200 lines across two new classes
  • Stages 08 to 10

Where we are going

The real Redis client understands your server:

$ redis-cli ping
PONG
flowchart LR
    A[Part 2<br/>bytes framed by newline] --> B[Stage 08<br/>read the protocol]
    B --> C[Stage 09<br/>RespParser]
    C --> D[Stage 10<br/>RespWriter]
    D --> E([Part 4<br/>commands])

Stage 08, read the protocol off the wire

Goal. See exactly what a Redis client sends, using the server you already have.

The idea

Your echo server from Part 2 is the perfect instrument. It prints every request it frames, so pointing the real client at it shows you the protocol without reading a spec first.

Run it

Start your server, then in another terminal:

redis-cli -t 2 ping

redis-cli gets nonsense back, because it does not understand an echo server. Your server log shows this:

Request: *1\r
Request: $4\r
Request: ping\r

Again with an argument:

redis-cli -t 2 echo "hey there"
Request: *2\r
Request: $4\r
Request: echo\r
Request: $9\r
Request: hey there\r

And with three:

redis-cli -t 2 set mykey myvalue
Request: *3\r
Request: $3\r
Request: set\r
Request: $5\r
Request: mykey\r
Request: $7\r
Request: myvalue\r

Notice that

Three things, and each one changes your code.

Every line ends with \r. Your framing splits on \n and leaves the carriage return behind. RESP terminates every part with CRLF, written \r\n.

The command name arrives lowercase. redis-cli sent ping, not PING. Matching has to be case-insensitive.

$9 for hey there counts the space. That number is a byte count, not a word or character count.

The format

Every command a client sends is an array of bulk strings. There are no exceptions.

*<number-of-arguments>\r\n
$<length-of-arg-1>\r\n<arg-1>\r\n
$<length-of-arg-2>\r\n<arg-2>\r\n

So ECHO "hey there" on the wire is literally:

*2\r\n$4\r\nECHO\r\n$9\r\nhey there\r\n

Five types, identified by the first byte:

ByteTypeExampleUsed for
+simple string+OK\r\nshort status replies
-error-ERR unknown command\r\nfailures
:integer:42\r\ncounts
$bulk string$5\r\nhello\r\nany value, binary safe
*array*2\r\n...lists of the above

Two special forms matter and you cannot guess them:

$-1\r\n      null bulk string, GET on a missing key
$0\r\n\r\n   empty string, a real stored value of length zero

Those look similar and mean different things. Confusing them is the most common RESP bug, and it earns a test in every class that touches it.

Try it yourself

  1. Run redis-cli lpush fruits apple banana against your echo server and read the log. Predict the shape before you run it.
  2. Send a command by hand: printf '*1\r\n$4\r\nPING\r\n' | nc localhost 6379 | xxd
  3. Store a value containing a newline and note that the $ length counts it as a byte.

Why length prefixes beat delimiters

Part 2 ended with two problems. A payload can never contain the delimiter, and a client that never sends one grows your buffer forever. Length prefixes fix both.

$7\r\nab\r\ncd\r\n
      └── 7 bytes: a b \r \n c d

The payload contains \r\n and nothing breaks, because nobody is scanning for it. That is what binary safe means, and it is why Redis can store a JPEG in a string.

It also kills the denial of service. The moment you read $7 you know exactly how many bytes to expect, so you can reject an absurd length before allocating for it. A delimiter parser has to keep buffering and hope.


Stage 09, the parser

Goal. Turn RESP bytes into a String[], correctly, including when only half of it has arrived.

The idea: what should a parser consume?

Three options, and the reasoning matters more than the answer.

String is wrong. Converting bytes to a String before parsing forces a charset decision on data that may be arbitrary binary. Decode after the length has cut out the exact bytes, never before.

InputStream is tempting and wrong differently. A stream parser calls read() when it needs more bytes, so the parser blocks. That couples parsing to I/O: you cannot test it without a socket, and “half a command arrived” becomes a parked thread instead of a return value.

byte[] with an explicit buffer works. The parser never does I/O. It is handed bytes and asked whether a complete command is present.

The rule worth keeping past this project: the thing that parses should never be the thing that reads.

That gives a three-outcome contract:

flowchart LR
    N["next()"] --> A["String[]<br/>a complete command"]
    N --> B["null<br/>not enough bytes yet"]
    N --> C["throws<br/>these bytes are garbage"]

null is not an error. It is the normal state of a stream that has delivered half a command. Telling incomplete apart from invalid is the entire difficulty of this stage. Everything else is bookkeeping.

The code

src/main/java/com/example/redis/RespParser.java, complete:

package com.example.redis;

import java.nio.charset.StandardCharsets;
import java.util.Arrays;

// Parses RESP bytes into commands. Owns no socket, callers feed it bytes.
public class RespParser {

    // same cap real redis uses for proto-max-bulk-len
    private static final int MAX_BULK_LENGTH = 512 * 1024 * 1024;
    private static final int MAX_ARRAY_SIZE = 1024 * 1024;

    private byte[] buffer = new byte[0];

    // cursor into buffer, only meaningful during a next() call
    private int pos;

    public void append(byte[] data, int length) {
        byte[] grown = Arrays.copyOf(buffer, buffer.length + length);
        System.arraycopy(data, 0, grown, buffer.length, length);
        buffer = grown;
    }

    // returns null when a complete command hasn't arrived yet
    public String[] next() {
        pos = 0;
        String[] command = parseArray();
        if (command == null) {
            return null;
        }

        // consume exactly what we parsed, keep the rest for next time
        buffer = Arrays.copyOfRange(buffer, pos, buffer.length);
        return command;
    }

    private String[] parseArray() {
        if (pos >= buffer.length) {
            return null;
        }
        if (buffer[pos] != '*') {
            throw new IllegalStateException("expected array, got byte " + buffer[pos]);
        }
        pos++;

        Integer count = readNumber();
        if (count == null) {
            return null;
        }
        if (count < 0 || count > MAX_ARRAY_SIZE) {
            throw new IllegalStateException("bad array size " + count);
        }

        String[] args = new String[count];
        for (int i = 0; i < count; i++) {
            String arg = parseBulkString();
            if (arg == null) {
                return null;
            }
            args[i] = arg;
        }
        return args;
    }

    private String parseBulkString() {
        if (pos >= buffer.length) {
            return null;
        }
        if (buffer[pos] != '$') {
            throw new IllegalStateException("expected bulk string, got byte " + buffer[pos]);
        }
        pos++;

        Integer length = readNumber();
        if (length == null) {
            return null;
        }
        if (length < 0 || length > MAX_BULK_LENGTH) {
            throw new IllegalStateException("bad bulk length " + length);
        }

        // payload plus its trailing CRLF must all be here
        if (pos + length + 2 > buffer.length) {
            return null;
        }

        String value = new String(buffer, pos, length, StandardCharsets.UTF_8);
        pos += length + 2;
        return value;
    }

    // reads digits up to the next CRLF, null if that CRLF hasn't arrived
    private Integer readNumber() {
        int lineEnd = indexOfCrlf(pos);
        if (lineEnd < 0) {
            return null;
        }
        String digits = new String(buffer, pos, lineEnd - pos, StandardCharsets.UTF_8);
        pos = lineEnd + 2;
        try {
            return Integer.parseInt(digits);
        } catch (NumberFormatException e) {
            throw new IllegalStateException("expected number, got '" + digits + "'");
        }
    }

    private int indexOfCrlf(int from) {
        for (int i = from; i + 1 < buffer.length; i++) {
            if (buffer[i] == '\r' && buffer[i + 1] == '\n') {
                return i;
            }
        }
        return -1;
    }
}

Setting pos = 0 at the top of next() and trimming the buffer only on success is what makes partial input free. A failed parse leaves the buffer untouched, so the next attempt starts clean. No rollback logic, no saved state.

The +2 in pos + length + 2 > buffer.length is the CRLF after the payload. Forget it and the next command starts two bytes late, which surfaces as a bizarre parse error one command later.

new String(buffer, pos, length, UTF_8) decodes after the length has cut out the exact bytes. Reverse that order and binary values break.

The length caps are the defence a delimiter parser could not have. $2000000000 from a hostile client is rejected before anything is allocated.

Wire it in

handleClient loses all its framing:

RespParser parser = new RespParser();
byte[] buffer = new byte[1024];
int bytesRead;

while ((bytesRead = inputStream.read(buffer)) != -1) {
    parser.append(buffer, bytesRead);

    // one read may hold several commands, or none
    String[] command;
    while ((command = parser.next()) != null) {
        System.out.println("Command: " + Arrays.toString(command));

        outputStream.write(RespWriter.simpleString("PONG"));  // stage 10 adds this
        outputStream.flush();
    }
}

The pending buffer, the newline scan, and the start cursor all disappear. They moved behind the parser, along with the stray \r.

Tests start here

Stages 01 to 07 had no tests, deliberately. A test asserting ServerSocket.bind() binds tests the JDK, not you. There was no branch, no transformation, and nc verified it faster than JUnit would.

The parser is the first pure function in the project. Bytes in, structure out, no sockets. It is also where bugs stop being visible by eye. Three tests justify the whole file:

@Test
void payloadMayContainCrlf() {
    // the declared length decides where the value ends, not the delimiter
    assertArrayEquals(new String[]{"ECHO", "a\r\nbc"},
            parserWith("*2\r\n$4\r\nECHO\r\n$5\r\na\r\nbc\r\n").next());
}

@Test
void returnsNullUntilCommandComplete() {
    RespParser parser = parserWith("*2\r\n$4\r\nECHO\r\n$5\r\nhel");
    assertNull(parser.next());

    byte[] rest = "lo\r\n".getBytes(StandardCharsets.UTF_8);
    parser.append(rest, rest.length);
    assertArrayEquals(new String[]{"ECHO", "hello"}, parser.next());
}

@Test
void parsesTwoCommandsFromOneAppend() {
    RespParser parser = parserWith("*1\r\n$4\r\nPING\r\n*1\r\n$4\r\nPING\r\n");
    assertArrayEquals(new String[]{"PING"}, parser.next());
    assertArrayEquals(new String[]{"PING"}, parser.next());
    assertNull(parser.next());
}

None of those is reachable by typing into nc, and each is a bug you would otherwise ship.

Run it

$ mvn test
[INFO] Tests run: 10, Failures: 0, Errors: 0

Then against the real client:

Command: [ping]
Command: [echo, hey there]
Command: [set, mykey, myvalue]

Notice that

No \r anywhere. If you still see one, something is splitting on \n.

Try it yourself

  1. Change pos + length + 2 to pos + length + 1 and run the tests. Watch which one fails and read the message. That is the “one command later” symptom in miniature.
  2. Send a command in three separate writes with pauses. The parser returns null, null, then the command.
  3. Feed it +OK\r\n, which is a reply rather than a command. It throws, because a client may only send arrays.

What usually goes wrong

StringIndexOutOfBoundsException on partial input. I hit exactly this:

int lineEnd = indexOfCrlf(pos);
if (lineEnd == 0) {     // should be < 0
    return null;
}

indexOfCrlf returns -1 when there is no CRLF yet. With == 0 the incomplete case fell through to new String(buffer, pos, -1 - pos, ...), in the exact path the stage exists to handle. Five tests failed at once and pointed straight at it.

The other one: everything works until a value contains a newline. That means you are scanning for \r\n instead of using the declared length.

Go deeper: RESP2 against RESP3, and inline commands

Redis 6 introduced RESP3, adding maps, sets, doubles, booleans, and out-of-band push messages. Clients opt in with HELLO 3. Without it, servers speak RESP2. Build RESP2. Replying to HELLO with an error is legitimate, and clients fall back.

Redis also accepts inline commands, plain PING\r\n with no * or $, for typing into telnet. Optional, and skipped here.


Stage 10, the writer

Goal. Turn Java values into RESP bytes.

The idea

The mirror image, in a separate class on purpose. After this stage, nothing outside these two files touches a \r\n again.

flowchart LR
    A[socket bytes] --> B[RespParser]
    B --> C["String[]"]
    C --> D[command logic]
    D --> E[RespWriter]
    E --> F[socket bytes]
    style B fill:#e8f0fe
    style E fill:#e8f0fe

When SET arrives in Part 5 it should decide "OK", not +OK\r\n. The moment protocol formatting leaks into command code, every command becomes a place a missing CRLF can hide.

The code

src/main/java/com/example/redis/RespWriter.java:

package com.example.redis;

import java.io.ByteArrayOutputStream;
import java.nio.charset.StandardCharsets;

// Encodes Java values as RESP2 bytes. Owns no socket, callers write what it returns.
public class RespWriter {

    private static final byte[] CRLF = "\r\n".getBytes(StandardCharsets.UTF_8);
    private static final byte[] NULL_BULK = "$-1\r\n".getBytes(StandardCharsets.UTF_8);

    // never build one of these from user data, it has no length prefix
    public static byte[] simpleString(String value) {
        return ("+" + value + "\r\n").getBytes(StandardCharsets.UTF_8);
    }

    public static byte[] error(String message) {
        return ("-" + message + "\r\n").getBytes(StandardCharsets.UTF_8);
    }

    public static byte[] integer(long value) {
        return (":" + value + "\r\n").getBytes(StandardCharsets.UTF_8);
    }

    public static byte[] bulkString(String value) {
        if (value == null) {
            return NULL_BULK;
        }

        // the declared length is bytes, not characters
        byte[] payload = value.getBytes(StandardCharsets.UTF_8);
        ByteArrayOutputStream out = new ByteArrayOutputStream();
        out.writeBytes(("$" + payload.length + "\r\n").getBytes(StandardCharsets.UTF_8));
        out.writeBytes(payload);
        out.writeBytes(CRLF);
        return out.toByteArray();
    }

    // elements are already encoded, so nesting costs nothing
    public static byte[] array(byte[]... elements) {
        ByteArrayOutputStream out = new ByteArrayOutputStream();
        out.writeBytes(("*" + elements.length + "\r\n").getBytes(StandardCharsets.UTF_8));
        for (byte[] element : elements) {
            out.writeBytes(element);
        }
        return out.toByteArray();
    }
}

Use payload.length after getBytes(UTF_8), never value.length(). "café" is 4 characters and 5 bytes. Declare $4 for 5 bytes and the client reads four, hits a stray byte where CRLF should be, and every later reply on that connection is garbage. The failure surfaces one command after the mistake.

array(byte[]...) takes already-encoded elements, so nesting is free. No recursive encoder, no value-type hierarchy. That decision pays off in Part 8 when streams need arrays of arrays of arrays.

Simple strings are never built from user data. They have no length prefix, so an embedded newline ends the reply early and desynchronises the connection. +OK and +PONG are yours. A stored value is not.

The test that earns its keep

@Test
void bulkStringLengthIsBytesNotCharacters() {
    // café is 4 characters but 5 bytes in UTF-8
    assertEquals("$5\r\ncafé\r\n", encoded(RespWriter.bulkString("café")));
}

It is the only test here that fails against the obvious String-based implementation, and the resulting bug is invisible until a user stores a non-ASCII value.

Worth adding too, a round trip through both classes:

@Test
void writerOutputParsesBackAsACommand() {
    byte[] encodedCommand = RespWriter.array(
            RespWriter.bulkString("ECHO"), RespWriter.bulkString("a\r\nbc"));

    RespParser parser = new RespParser();
    parser.append(encodedCommand, encodedCommand.length);

    String[] command = parser.next();
    assertEquals("ECHO", command[0]);
    assertEquals("a\r\nbc", command[1]);
}

That is the test that catches the two classes disagreeing about the protocol.

Run it

$ mvn test
[INFO] Tests run: 23, Failures: 0, Errors: 0
$ redis-cli -t 2 ping
PONG

$ printf '*1\r\n$4\r\nPING\r\n' | nc localhost 6379 | xxd
00000000: 2b50 4f4e 470d 0a       +PONG..

Notice that

2b is + and 0d 0a is CRLF. That is the first time the real client understood your server. It printed PONG because it decoded a valid RESP simple string. Emit PONG\r\n without the + and it reports a protocol error instead.

Try it yourself

  1. Remove the + from simpleString and run redis-cli ping. Read the error it gives you.
  2. Encode a nested array: array(integer(1), array(bulkString("x"))). Print it. Note that you did not write a recursive function.
  3. Write bulkString the naive way using value.length(), then run the café test.

What you built

Two classes, 23 tests, and a hard boundary. Protocol on one side, meaning on the other. Your server speaks Redis in both directions.

It also replies PONG to everything, including GET. Part 4 fixes that.

Checkpoint

  1. Why must you not convert bytes to a String before parsing?
  2. Why is null the right signal for incomplete input rather than an exception?
  3. Why is scanning for \r\n to end a bulk string wrong, given every bulk string ends with one?
  4. What is the difference on the wire between $-1, $0\r\n\r\n, and +\r\n?
  5. What breaks, and when, if a reply declares the wrong length?

Resources

Next

The server stops replying PONG to everything and starts caring what you asked.

Part 4: Bytes Become Commands →