CASE: Voice-Controlled Self-Balancing Robot
2026 · ESP32 · FreeRTOS · C/C++ · PID control · MPU6050 DMP · I²S audio
Design and build a two-wheeled self-balancing robot that holds vertical equilibrium via real-time PID control on a dual-core ESP32, integrating on-device wake-word voice input with cloud-based conversational AI so the balance loop is never interrupted by audio processing.
- 100 Hz
- Balance Loop Rate
- 2
- Dedicated Cores
First look

What it is made of
Hardware
- Ran the balance loop on an ESP32-WROOM-32's Core 1 under FreeRTOS, so WiFi and TLS never share a core with it. The loop holds 100 Hz with under 10 µs of jitter.
- Read attitude from an MPU6050 through its on-chip DMP quaternion fusion, so the controller receives a drift-corrected angle without spending MCU cycles on filtering.
- Drove two DC motors through a TB6612FNG off a 7.4 V pack, so PID output becomes PWM with the current headroom the chassis needs to catch itself.
- Captured voice on an INMP441 I2S microphone and answered through a MAX98357A into a 4 Ω 3 W speaker, so the robot hears and replies without an external codec.
Software
- Held vertical equilibrium with a PID controller (Kp 15 · Ki 140 · Kd 0.9) fed by the DMP at 100 Hz, so the robot stays upright while it is talking.
- Detected “Hey CASE” on-device from the I2S stream, so the network is only touched once someone has actually spoken.
- Bridged audio to Nvidia Personaplex (Moshi) through a Google Colab endpoint transcoding PCM ↔ Opus, so conversational AI runs off-board inside a 167 KB heap.
Why the loop stays on the microcontroller
The same gains, the same shove, two schedulers. On the left the control update runs when the timer fires; on the right it runs when a general-purpose operating system gets round to it: mostly a few milliseconds late, occasionally 140 ms late. Watch the right-hand chassis wander.
- outside the band
- 0.00 s
- now
- 4.98°
- rests at
- 0.16°
- outside the band
- 0.00 s
- now
- 4.98°
- rests at
- 0.32°
Lean drawn ×6. The difference between these two is fractions of a degree, which is the whole point and invisible at life size.
Neither falls, which is the honest result. The reason to keep the loop on the microcontroller is margin, not survival. Under the scheduler the same gains spend roughly three times as long outside the half-degree band and come to rest further off upright, and every one of those excursions is time the chassis is spending its recovery budget on the operating system instead of on the next shove.
What came of it
- The robot holds vertical equilibrium while handling live voice input.
- The balance loop is never interrupted by audio or network work.
The trade it forced
Why the ESP32 alone wasn’t enough
167 KB of free heap after WiFi and TLS caps the board at 1.5 s recordings and HTTP round-trips, with no streaming anywhere. But the interrupt-driven DMP loop holds <10 µs jitter where a Linux scheduler gives 2–15 ms, so balance stays on the MCU and only the voice pipeline moves to a Raspberry Pi. About 80% of self_balance.ino survives that migration.
The other plates


Self-balancing · Wake-word activation · Conversational AI · Voice-commanded motion · Audio capture & playback