TCP before Redis

A server that listens on port 6379, accepts clients, and echoes back whatever they send.

  • Part 1
  • beginner
  • about 45 minutes

You will build

a server that listens on port 6379, accepts clients, and echoes back whatever they send

You will understand

what binding a port actually does, why ServerSocket and Socket are different things, and why TCP has no messages

  • roughly 40 lines
  • Stages 01 to 05

Where we are going

By the end of this post, this works:

$ printf 'hello\n' | nc localhost 6379
hello

Your program received those bytes over a network connection and sent them back. Every Redis command in the rest of the series is that loop with meaning attached.

flowchart LR
    A[Stage 01<br/>empty project] --> B[Stage 02<br/>bind a port]
    B --> C[Stage 03<br/>accept a client]
    C --> D[Stage 04<br/>read bytes]
    D --> E[Stage 05<br/>write bytes]
    E --> F([Part 2<br/>framing])

There is no Redis in this post. No commands, no keys, no protocol. That is deliberate. Most of the protocol bugs I hit later turned out to be misunderstandings I could have caught here.

If you have never written network code, go slowly, and do the experiments in Stage 04.


Stage 01, an empty project

Goal. A Maven project that compiles, runs, and is committed.

The idea

This stage exists so the build-run-commit loop is boring before anything interesting depends on it. If your build is shaky, you will spend Stage 04 debugging Maven instead of debugging TCP.

Two words you need.

Maven is the build tool. It compiles your code, runs tests, and expects sources in src/main/java. Follow that layout and it needs almost no configuration.

pom.xml is Maven’s config file. Ours is 20 lines and stays that way.

The code

src/main/java/com/example/redis/Main.java, complete:

package com.example.redis;

public class Main {
    public static void main(String[] args) {
        System.out.println("Redis server starting...");
    }
}

The package is whatever you chose in Part 0. Every code block here uses com.example.redis. The pom.xml is there too.

Run it

$ mvn clean test
[INFO] BUILD SUCCESS

$ mvn -q clean compile
$ java -cp target/classes com.example.redis.Main
Redis server starting...

Those last two commands run every stage in this series. mvn -q clean compile wipes old output and compiles. java -cp target/classes runs the class from where Maven put it.

Notice that

BUILD SUCCESS came from a project with zero tests. It proved your code compiles and the layout is valid. Nothing executed.

That distinction stops being academic in Stage 02, where compiling and working become different things.

Try it yourself

  1. Delete a semicolon, run mvn clean test. Read the error: file, line, column. You will see that shape a lot.
  2. Run ls target/classes/com/example/redis/. That .class file is what java actually runs.
  3. Change the message, recompile, run. Make the loop muscle memory now, because you will run it after every change from here on.

What usually goes wrong

Main method not found in class Main. You wrote void main() without String[] args. Java 25 allows it. Earlier versions do not.

class file has wrong version. Maven and java are on different JDKs. Compare mvn -v with java -version, then set JAVA_HOME.


Stage 02, bind to a port

Goal. Your program claims port 6379 and stays running.

The idea

Some vocabulary first, used for the rest of the series.

A port is a number from 0 to 65535 that identifies one program on a machine. An IP address says which computer. A port says which program on it. Redis conventionally uses 6379.

A socket is the operating system’s handle for one end of a network connection. You get one from the OS, read and write through it, then close it.

Binding is asking the OS to reserve an address and port for your process.

sequenceDiagram
    participant App as Your JVM
    participant OS as Operating system
    App->>OS: new ServerSocket()
    OS-->>App: an unbound socket
    App->>OS: bind(0.0.0.0:6379)
    Note over OS: port reserved for this process,<br/>accept queue created
    OS-->>App: ok
    Note over App,OS: the OS now completes client<br/>handshakes on your behalf

The code

package com.example.redis;

import java.io.IOException;
import java.net.BindException;
import java.net.InetSocketAddress;
import java.net.ServerSocket;

public class Main {
    public static void main(String[] args) throws IOException, InterruptedException {

        try (ServerSocket serverSocket = new ServerSocket()) {

            // lets us rebind immediately after a restart instead of waiting out TIME_WAIT
            serverSocket.setReuseAddress(true);
            serverSocket.bind(new InetSocketAddress(6379));

            System.out.println("Listening on port " + serverSocket.getLocalPort());

            // nothing accepts connections yet, so park main here
            // otherwise the JVM exits and takes the socket down with it
            Thread.currentThread().join();

        } catch (BindException e) {
            System.err.println("Port 6379 already in use. Is another Redis running?");
            System.exit(1);
        }
    }
}

new ServerSocket() with no arguments creates an unbound socket. The one-argument constructor binds immediately, and setReuseAddress has to be called before binding, so there would be no gap to call it in. Splitting also shows what the shortcut hides. Create, configure, and bind are separate operations.

setReuseAddress(true) matters more than it looks. After you kill a server the OS holds its address in a state called TIME_WAIT for up to a couple of minutes. Without the flag, restarting immediately can fail with “Address already in use” while nothing is running. You will restart this server hundreds of times.

getLocalPort() prints what you actually bound rather than what you meant to. It costs nothing, and it stops the log lying to you the day the port becomes configurable.

Thread.currentThread().join() blocks forever. join() waits for a thread to finish, and main is waiting on itself. Crude on purpose. Stage 03 deletes it.

The try-with-resources block closes the socket on exit. Irrelevant today. The habit matters later.

Run it

$ java -cp target/classes com.example.redis.Main
Listening on port 6379

It does not exit. That is correct.

Second terminal. Does the OS agree?

$ lsof -nP -iTCP:6379 -sTCP:LISTEN
COMMAND   PID    USER   FD   TYPE  NODE NAME
java    15253 isubash    5u  IPv6   TCP *:6379 (LISTEN)

Can a client reach it?

$ nc -vz localhost 6379
Connection to localhost port 6379 [tcp/*] succeeded!

Notice that

That connection succeeded against a program doing nothing. No accept(), no read loop, no reply. It is parked on join().

The operating system completed the whole TCP handshake for you, and the connection is sitting in a queue your program has never looked at. Stage 03 reaches into that queue.

Binding is not communication. Hold on to that.

Try it yourself

  1. Comment out Thread.currentThread().join(). The program prints and exits, and nc now fails. A bound socket does not keep the JVM alive. It is a file handle, not a thread.
  2. Start the server twice in two terminals. The second prints your BindException message. That is exclusivity, demonstrated.
  3. Change the port to 6380. Confirm 6379 now fails and 6380 succeeds.

What usually goes wrong

Address already in use while nothing seems to be running. Something owns it. Check with lsof -nP -iTCP:6379 -sTCP:LISTEN. On my machine it was a Homebrew Redis service, and later a Docker container.

Your test passes but you are talking to someone else’s server. This one cost me an hour. A real Redis had bound the specific addresses 127.0.0.1 and ::1. I bound the wildcard 0.0.0.0. macOS permits that combination with SO_REUSEADDR, so there was no error, and nc localhost 6379 reached Redis rather than me. On Linux it would have thrown.

The habit that outlives the bug: verify which server answered, not just that something did.

Go deeper: what bind and listen actually do

new ServerSocket() plus bind() maps onto three C syscalls. socket() creates the descriptor, bind() attaches it to an address, listen() creates the queue of completed connections. Java collapses the last two into one call.

The queue has a depth called the backlog. ServerSocket lets you set it as a constructor argument and defaults to 50. When it fills, further connection attempts are refused by the OS before your code sees them.


Stage 03, accept a connection

Goal. Take a waiting connection off the queue and print who connected.

The idea

The two socket types are not interchangeable, and confusing them causes strange bugs later.

flowchart TD
    SS["ServerSocket<br/><i>one per server</i><br/>a door, never carries data"]
    SS -->|accept| S1["Socket, client A<br/><i>a conversation</i>"]
    SS -->|accept| S2["Socket, client B"]
    SS -->|accept| S3["Socket, client C"]

accept() does not answer a connection. The OS already did that. It hands your program a brand new Socket for a connection that already exists, and leaves the ServerSocket listening.

accept() blocks, meaning the thread stops and waits until there is something to return. Blocking here is how you wait without burning CPU.

The code

Replace the join() line with:

while (true) {
    // blocks until the kernel has a completed connection waiting for us
    Socket clientSocket = serverSocket.accept();

    System.out.println("Client connected: " + clientSocket.getRemoteSocketAddress());

    // nothing to say to it yet, so hang up
    clientSocket.close();
}

Add import java.net.Socket;.

The while (true) is not optional. Without it you accept one client, main returns, the JVM exits, and the port dies with it.

getRemoteSocketAddress() prints something like /127.0.0.1:59532. That number is the client’s ephemeral port, a random high port its OS picked for this connection.

clientSocket.close() is deliberate. You have nothing to send, so holding the socket only leaks a file handle. The visible effect is that nc connects and exits instantly. That is your server hanging up.

Run it

Listening on port 6379
Client connected: /127.0.0.1:59532
Client connected: /127.0.0.1:59533
nc localhost 6379     # then again
nc localhost 6379

Notice that

The port number changed between connections. 59532, then 59533. Your server has one port, and many clients can be connected at once, because a connection is identified by all four of client IP, client port, server IP, and server port.

Also, deleting join() did not make the program exit. accept() blocks too. The problem of keeping the process alive and the problem of waiting for a client were always the same problem, and you just replaced the fake solution with the real one.

Try it yourself

  1. Connect five times and watch the ephemeral port climb.
  2. Remove clientSocket.close(). nc now stays connected instead of exiting, because nobody hung up.
  3. Remove the while (true), connect twice. The second connection is never reported and the server exits.

What usually goes wrong

Nothing prints when you connect. Your accept loop is unreachable, or you are looking at a server that never called accept(). Remember Stage 02: connecting succeeds without your involvement, so “it connected fine” is not evidence your code ran.


Stage 04, read bytes

Goal. Print what a client sends.

This is the most important stage in the post.

The idea

Almost everyone who struggles with a protocol parser struggles because they believe TCP delivers messages. It delivers a stream of bytes, in order, with no marks between them.

An InputStream is the read side of a socket. You ask for bytes and get however many have arrived.

A buffer is a byte array you reuse as a landing area. Its size is not a message size. It is how much you are willing to take at once.

The code

Inside the accept loop, replacing clientSocket.close():

InputStream inputStream = clientSocket.getInputStream();
byte[] buffer = new byte[1024];
int bytesRead;

// read() returns -1 when the client closes its end
while ((bytesRead = inputStream.read(buffer)) != -1) {
    String data = new String(buffer, 0, bytesRead, StandardCharsets.UTF_8);

    // escape the invisible bytes, they matter once the protocol arrives
    System.out.println("Received: " + bytesRead + " bytes: "
            + data.replace("\r", "\\r").replace("\n", "\\n"));
}

System.out.println("Client disconnected");
clientSocket.close();

Imports: java.io.InputStream and java.nio.charset.StandardCharsets.

read(buffer) blocks until at least one byte arrives, writes some into the buffer, and returns how many. It does not fill the buffer and it does not wait for more.

The 0, bytesRead bounds on the String constructor are mandatory. Convert the whole array and you print 1024 bytes including leftovers from earlier reads. That is the classic first bug here.

Passing StandardCharsets.UTF_8 explicitly matters because the platform default charset varies by machine. Identical code gives different output in a different locale.

-1 means end of stream. The client closed. That is normal, not an error, and there is no exception to catch.

Run it

nc localhost 6379
hello⏎
Received: 6 bytes: hello\n

Notice that

Six bytes, not five. The Enter key is a byte like any other. From here to the end of the series the bugs live in bytes you cannot see, which is why the log escapes them.

Try it yourself

Do all four. This set makes the rest of the series make sense.

1. Two messages, with a pause. Type hello⏎ then world⏎.

Received: 6 bytes: hello\n
Received: 6 bytes: world\n

Two reads. Feels obvious. It is about to stop being.

2. Two messages, no pause.

printf 'hello\nworld\n' | nc localhost 6379
Received: 12 bytes: hello\nworld\n

One read, two messages. Your program cannot tell where one ends.

3. One message, split.

python3 -c "
import socket, time
s = socket.create_connection(('127.0.0.1', 6379))
s.sendall(b'hel'); time.sleep(0.3); s.sendall(b'lo\n')"
Received: 3 bytes: hel
Received: 3 bytes: lo\n

Two reads, one message.

4. A second client. While one nc is connected, open another in a third terminal and type something. Nothing happens until the first leaves. Do not fix this. Part 2 does.

Notice that, again

Experiment 1 worked only because you paused. The pause did the splitting, not TCP. All of these are legal for a client that sent hello world:

one read   → "hello world"
two reads  → "hello" then " world"
six reads  → "h" "e" "l" "l" "o" " world"

Your code has to be correct for every one of them. Nothing in TCP promises which you get.

What usually goes wrong

Garbage after your text. You forgot the 0, bytesRead bounds and are printing stale buffer.

You reached for BufferedReader.readLine(). It looks like the obvious tool and it hides the exact problem this stage exists to show. It is also wrong for Redis later, because bulk strings carry a declared length and may contain \r, \n, or raw binary. Reading until a newline corrupts them. Stay on raw bytes.

Go deeper: why two writes can arrive as one read

TCP buffers small writes and may coalesce them into one segment. That is Nagle’s algorithm, designed to avoid flooding a network with tiny packets. Routers may also split a large write across segments.

Neither is something you can rely on or configure away. Every protocol over TCP defines its own framing, and Part 3 shows the two ways to do it.


Stage 05, write bytes back

Goal. Echo whatever arrives, byte for byte.

The idea

One socket carries two independent directions.

flowchart LR
    C[Client] -->|getInputStream| S[Your server]
    S -->|getOutputStream| C

That independence matters later, when a client stops sending while still listening.

The code

Inside the read loop, after the print:

// echo back exactly what arrived, same bounds discipline as the read
outputStream.write(buffer, 0, bytesRead);
outputStream.flush();

with OutputStream outputStream = clientSocket.getOutputStream(); beside the input stream.

Same bounds rule as reading. Drop 0, bytesRead and you ship 1024 bytes of stale data.

flush() does nothing here, because a raw socket stream writes straight to the OS. Call it anyway. The day you wrap this in a BufferedOutputStream or a PrintWriter, its absence becomes a reply sitting in a Java buffer while the client waits forever.

Run it

$ printf 'hello\n' | nc localhost 6379 | xxd
00000000: 6865 6c6c 6f0a    hello.

$ printf 'abc' | nc localhost 6379 | xxd
00000000: 6162 63           abc

xxd prints hex. 0a is the newline that came in and went back out.

Notice that

No newline was invented for abc, and none was stripped from hello\n. Byte-exact in both directions, which is the precision the protocol will demand in Part 3.

One more thing worth knowing: write() returning does not mean the client received anything. It means the OS took custody of your bytes.

Try it yourself

  1. Send three messages down one connection, reading the reply between each. All three come back and the connection stays usable.
  2. Delete flush(). Nothing changes. Now wrap the stream in new BufferedOutputStream(...) without flushing and watch the client hang forever. Put both back. That is the failure the habit prevents.
  3. printf '\x00\x01\x02' | nc localhost 6379 | xxd. Raw binary round-trips untouched, because nothing is interpreting it.

What usually goes wrong

The client hangs waiting for a reply. Missing flush() behind a buffered stream, nine times out of ten.

Extra bytes come back. Missing bounds on the write.


What you built

flowchart LR
    subgraph Your server
        B[bind 6379] --> A[accept]
        A --> R[read bytes]
        R --> W[write bytes]
        W --> R
    end
    C[nc or any TCP client] <--> A

Forty lines, no Redis, and a handful of facts the rest of the series rests on.

Binding is not communication. The OS completes handshakes without you.

ServerSocket listens and Socket converses. One of the first, many of the second.

TCP has no message boundaries. One read is not one request, and never will be.

Bytes are bytes. Decode late, count exactly, print the invisible ones.

Checkpoint

Answer these without looking anything up.

  1. What does bind() ask the operating system for, and what changes on the machine when it succeeds?
  2. Why did nc connect successfully to a server that never called accept()?
  3. Your server has one port. How can many clients be connected at once?
  4. Why is read() returning fewer bytes than the client sent not a bug?
  5. What does write() returning actually guarantee?

If number 4 is not obvious yet, redo the Stage 04 experiments. Everything in Part 2 rests on it.

Resources

Next

Your server answers two requests sent in one packet with a single reply. Fixing that means building the framing TCP refuses to provide, and then discovering that one client can block every other one.

Part 2: One Read Is Not One Request →