Skip to the piece

167 KB of free heap decided the architecture

TARS-CASE · ESP32-WROOM-32 · architecture note, May 2026

Three constraints run this robot’s voice pipeline: a recording window of a second and a half, an upload that is a single HTTP POST rather than a stream, and a proxy in the middle whose only job is to change one audio format into another. All three come out of one number that nobody chose. This is that number, what it bought, and the architecture it settled.

167 KB
free heap after WiFi and TLS
48 KB
the one buffer it has to hold
1.5 s
what that buys at 16 kHz
2–15 ms
the jitter that ruled out Linux

01


The number

An ESP32-WROOM-32 that has brought up WiFi and negotiated TLS has about 167 KB of heap left. That figure is a leftover, not a design target. Every decision below it is downstream of what the radio stack did not take.

It is written down twice in this project. Once in an architecture note, as the first sentence of the argument. And once where it actually bites, as a comment above a malloc:

// 1.5s at 16 kHz = 24000 samples = 48 KB
// fits in ESP32-WROOM-32's ~167 KB free heap
#define RECORD_SAMPLES 24000
static int16_t* audioBuf = NULL;

One allocation, taken once at startup and never given back, holding a little under a third of everything left. If it fails the firmware prints the free heap and stops, because there is no version of this robot that talks without it.

02


What a second and a half is

Sixteen kilohertz, sixteen bits, one channel: 32 KB per second of speech. 48 KB is 1.5 seconds, and 1.5 seconds is the whole utterance.

Not 1.5 seconds of headroom. The recording window is fixed. When the wake word fires, the firmware records exactly 24000 samples and then stops, whether or not the person in front of it has finished a sentence:

while (totalRecorded < RECORD_SAMPLES) {
size_t remaining = RECORD_SAMPLES - totalRecorded;
size_t toRead = min(remaining, (size_t)CHUNK_SAMPLES);
totalRecorded += audioInput.read(audioBuf + totalRecorded, toRead);
}

The right behavior is endpointing: record until the speaker stops, then send. That needs a buffer sized for the longest thing anyone might say, which is a buffer sized for a number you do not know in advance, which is the one shape a fixed 48 KB allocation cannot take. The firmware’s own notes list silence-based endpointing under next steps, and it is still there.

03


Why there is a proxy in the middle

The conversational model this robot talks to takes Opus over a socket. The ESP32 sends raw 16-bit PCM over a plain HTTP POST. Something has to sit between them, so something does: a bridge running in a Colab notebook whose entire job is to transcode one format into the other and pace the upload into the model in real time.

That proxy is not an architectural preference. It is a memory artefact. Encoding Opus on the device and holding a socket open for the length of an exchange are both things the remaining heap will not pay for, so the transcoding moved to the only machine in the system with room for it.

You can read the cost of that hop off a single line in the client. The bridge does not answer until its own pipeline has finished. Paced upload, a second of trailing silence, then up to fifteen seconds of collecting the reply:

http.setConnectTimeout(5000);
http.setTimeout(30000); // the bridge runs to completion first

Thirty seconds of patience in a device whose control loop is not allowed to be late by ten microseconds. Those two numbers living in the same binary is the whole problem stated in one line.

04


Before there were cores

The balance sketch came first, and it has no concurrency in it at all. It is one loop that spins waiting for the IMU’s interrupt and does the control work in the gaps:

while (!mpuInterrupt && fifoCount < packetSize)
{
pid.Compute();
Serial.print(input); Serial.print(" =>"); Serial.println(output);
if (input > 150 && input < 200) { ... }
}

The rate is not held by a timer. The PID library is told SetSampleTime(10) and refuses to recompute more often than every 10 ms, so the busy-wait can run as fast as it likes and the controller still updates at 100 Hz. It works, and it works because nothing else is competing for the processor.

Two details in that loop are worth naming because they are the difference between a model and a machine. The setpoint is 182, not 180. The angle at which this chassis is actually upright is not the angle the maths says it is, and the number is calibrated per robot off the serial monitor. And the PID output is limited to ±255 and then halved on the way to the driver, so the authority the controller really has is half of what the limits say.

The other detail is the print. A serial write on every pass of the balance loop is time spent on the balance path, and it is fine exactly as long as the balance path owns the chip.

05


The concurrency that made it fit

Adding voice ends that. Audio capture, wake-word checks, a WiFi POST and a thirty-second wait cannot share a thread with a controller that has to answer every 10 ms. So the integrated firmware stops being a loop and becomes two pinned tasks:

xTaskCreatePinnedToCore(balanceTask, "BalanceTask",
10000, NULL, 2, NULL, 1); // core 1
xTaskCreatePinnedToCore(audioTask, "AudioTask",
20000, NULL, 1, NULL, 0); // core 0

Balance takes core 1 at priority 2 with a 10 KB stack. Audio takes core 0 at priority 1 with 20 KB, twice the stack for the work that is allowed to be slow. The `loop()` that used to hold everything is now empty.

What matters more than the pinning is what the two tasks are allowed to share, which is one queue, ten entries deep, holding nothing but a motor mode. Audio pushes; balance reads with a zero timeout and carries on if there is nothing there:

if (xQueueReceive(motorCommandQueue, &voiceCommand, 0) == pdTRUE) {
balanceControl.setMotorMode(voiceCommand);
}

That zero is the design. There is no lock the balance task can wait on, no buffer it shares with the network, and no call in its path that can block on audio. The voice side can stall for thirty seconds and the only thing that happens to the chassis is that it does not receive a new command.

06


The half of the cap that came off

For a while the reply had the same 48 KB problem as the recording, for the same reason: the firmware downloaded the answer into a buffer before playing it, so the answer could only be as long as the buffer, and it was cut off at a second and a half.

That half was removable, and removing it removed the buffer rather than growing it. The reply is now read off the HTTP stream in 512-byte pieces and written straight to the I2S output as it arrives:

static int16_t pcmBuf[256];
size_t got = stream->readBytes(raw + leftover, toRead);
audioOut->write(pcmBuf, wholeSamples);

There is no clock in that loop and it does not need one. The write blocks on the I2S DMA buffer, and the DMA drains at the speaker’s sample rate, so the download paces itself to real time for free. The reply is now unbounded in length while costing 512 bytes of RAM.

The recording window did not move. That direction is still a single POST of the whole 1.5 seconds, because streaming the upload means holding a socket to the model for the length of an utterance, and that is the thing the heap will not buy. One of the two caps was an implementation detail. The other one is the budget.

07


Where the budget ran out

The architecture note that came out of this is blunt about which kind of problem is left: these are RAM and protocol constraints that cannot be optimized away. Not tuning, not a smaller buffer, not a better codec. The fixed window, the round-trip upload and the transcoding proxy are all one constraint wearing three costumes.

Which leaves two ways out, and they are not close. Put a Linux machine in the robot and run everything on it. Or put a Linux machine in the robot and leave the balance loop where it is.

The first one loses on the same measurement the rest of this site keeps coming back to. The interrupt-driven loop on the microcontroller measures under 10 µs of jitter. The same control code under a general-purpose kernel, driving the motors through software PWM in userspace, gets 2–15 ms. Three orders of magnitude, on a chassis whose entire recovery budget is a fraction of a degree. The failure that buys is not a wobble: if any process stalls, the robot falls over.

So the decision is the split. The microcontroller keeps attitude, PID and the motor driver, and the note’s own estimate is that about 80 % of the existing balance sketch runs unchanged. The single-board computer takes the microphone, the speaker, the wake word and the network. Because it can hold a socket and encode Opus, it talks to the model directly and the transcoding proxy stops existing.

The part that is easy to miss is the fault isolation. Two processors is not only more room; it is a boundary. The voice half can crash, hang, or be killed outright and the chassis stays upright, because the thing keeping it upright is not running on that computer.

08


What the decision costs, and what it deletes

None of that is free. It is another board and another thirty-five dollars, roughly 1.8 A against 1.2 A, a filesystem on an SD card that dies if you cut power at the wrong moment, and a boot that takes fifteen to thirty seconds. That last one is only survivable because of the split: the microcontroller is balancing the moment it has power, and the conversation joins when it is ready.

And the first line of the migration plan is the part I did not expect to write:

Remove FreeRTOS dual-core task pinning
(single core is sufficient)

The two-task design is what made voice possible on one chip, and it is the first thing the migration deletes. Once audio is not on the microcontroller, there is nothing left to protect the balance loop from, and the concurrency goes back to being a single loop that owns the processor, which is what the March sketch was before any of this started.

It is worth saying plainly what state this is in. The architecture is decided and written down; it is not built. The pipeline above it is not finished either: the wake-word detector is a placeholder that triggers on any loud sound, and the local command classifier returns “complex” for everything, which means every utterance goes to the network whether it needed to or not. Both of those are cheap to fix and neither of them changes the number.

Which is the thing worth keeping. The interesting constraint in an embedded system is rarely the one in the datasheet. It is the one left over after the parts you did not choose have taken their share. 167 KB is not a specification. It is a remainder, and it decided the shape of the whole machine.