Skip to content

RTSP Server

Phil Schatzmann edited this page Sep 14, 2026 · 1 revision

This tutorial walks through streaming audio and video from an Arduino/ESP32 board (or the desktop, via the Arduino-Emulator) using arduino-audio-tools' RTSP server. It covers the building blocks, a working example for each media type, and a set of gotchas that are easy to hit and hard to diagnose from the symptom alone.

For the class-level API reference, see the headers in this directory (RTSPServer.h, RTSPFormat.h, RTSPMediaSource.h, ...) and the Communication wiki page. Runnable examples referenced throughout live in examples/examples-communication/rtsp/ (hardware) and tests-cmake/network/rtsp-* (desktop, no hardware needed).

1. The building blocks

Every RTSP stream in this library is assembled from four pieces:

  • RTSPFormat - describes one media stream for SDP purposes (codec, payload type, sample rate/dimensions) and its timing (timerPeriodUs(), fragmentSize()). Concrete formats: RTSPFormatPCM, RTSPFormatMP3, RTSPFormatAAC, RTSPFormatMJPEG, RTSPFormatH264, RTSPFormatMTS, ...
  • IMediaSource - hands the streamer the actual bytes to send.
    • RTSPMediaSource wraps any Arduino Stream (or AudioStream) - use this for byte-oriented sources: PCM, a codec's encoded output, a file, a generated tone.
    • RTSPMediaCallbackSource wraps a plain C read-callback function instead of a Stream - convenient when the data doesn't already live behind a Stream (e.g. a camera driver, or a packetized video encoder).
  • RTSPMediaStreamer<Platform> - owns the RTP session: builds RTP headers, paces sending via a timer (RTSPFormat::timerPeriodUs()), and writes to the client over TCP-interleaved or UDP transport.
  • RTSPServer<Platform> - accepts RTSP connections (DESCRIBE/SETUP/ PLAY/TEARDOWN) and drives one RTSPMediaStreamer.

Platform is a template parameter binding these to a concrete network stack - almost always RTSPPlatformWiFi (RTSPPlatformWiFi.h), which is just RTSPPlatform<WiFiServer, WiFiClient, WiFiUDP>.

RTSPOutput<Platform> is a convenience wrapper that bundles an RTSPFormat + an AudioEncoder + an internal RTSPMediaSource + RTSPMediaStreamer into one AudioOutput-shaped object, so it can sit at the end of an AudioPlayer/StreamCopy pipeline like any other output. Reach for it for audio; for video, wire RTSPMediaCallbackSource/ RTSPMediaStreamer/RTSPServer together directly (see §3).

2. Audio: a generated sine tone (PCM)

The simplest possible server - no encoder, no file, just a live signal:

#include "AudioTools.h"
#include "AudioTools/Communication/RTSP/RTSPPlatformWiFi.h"
#include "AudioTools/Communication/RTSP.h"

AudioInfo info(16000, 1, 16);
SineFromTable<int16_t> sineWave(32000);
GeneratedSoundStream<int16_t> sound(sineWave);
RTSPMediaSource source(sound, info);          // Stream + AudioInfo
RTSPMediaStreamer<RTSPPlatformWiFi> streamer(source);
RTSPServer<RTSPPlatformWiFi> rtsp(streamer, 8554);

void setup() {
  auto cfg = sineWave.defaultConfig();
  cfg.copyFrom(info);
  sineWave.begin(cfg, N_B4);

  rtsp.begin();   // on ESP32 you can instead call begin(ssid, password)
                  // to have it join WiFi for you first
}

void loop() { delay(1000); }    // RTSP + RTP run on their own tasks

Connect with ffplay -rtsp_transport tcp rtsp://<device-ip>:8554/. A full, desktop-runnable copy of this lives at tests-cmake/network/rtsp-sine/rtsp-sine.cpp.

3. Audio: streaming MP3 files from disk

#include "AudioTools.h"
#include "AudioTools/Disk/AudioSourceSTD.h"       // or AudioSourceSD on hardware
#include "AudioTools/AudioCodecs/MP3Parser.h"
#include "AudioTools/Communication/RTSP/RTSPPlatformWiFi.h"
#include "AudioTools/Communication/RTSP.h"

MP3ParserEncoder enc;                 // repackages existing MP3 frames - no re-encoding
RTSPFormatMP3 mp3format(enc);
MetaDataFilterEncoder filter(enc);    // strips ID3 tags from the stream
RTSPOutput<RTSPPlatformWiFi> rtsp_out(mp3format, filter);
AudioSourceSTD source("/music/", ".mp3");
CopyDecoder decoder;                  // no decoding - MP3 goes out as MP3
AudioPlayer player(source, rtsp_out, decoder);
RTSPServer<RTSPPlatformWiFi> rtsp(rtsp_out.streamer(), 8554);

void setup() {
  mp3format.setUseRfc2250Header(true);   // see gotcha #1 below
  player.begin();
  rtsp_out.begin();
  rtsp.begin();
}

void loop() {
  if (rtsp_out && rtsp) player.copy();
}

A desktop-runnable copy lives at tests-cmake/network/rtsp-player/rtsp-player.cpp.

Gotcha #1: the RFC 2250 header (payload type 14 / MPA)

RTSPFormatMP3 supports two wire formats for RTP payload type 14 (MPEG audio): raw MP3 frames back-to-back, or frames prefixed with the optional RFC 2250 4-byte MPEG-audio header (2 reserved bytes + a 2-byte fragment offset). VLC expects raw frames; ffmpeg/ffplay expect the 4-byte header. The default is off (raw frames). If you only test with VLC this is invisible; against ffplay every frame boundary silently shifts by 4 bytes and decodes as pure noise, with no error printed on either the server or client side to point at the cause. If your stream "plays" but sounds like static, check this first:

mp3format.setUseRfc2250Header(true);

Gotcha #2: fragment size must match one encoder frame

RTSPMediaStreamer's generic byte-stream send path reads exactly RTSPFormat::fragmentSize() bytes every timerPeriodUs() tick and sends whatever it got as one RTP packet. If fragmentSize() is bigger than one actual encoded frame, each packet silently carries several frames' worth of data, sent at one frame's timing - i.e. several times faster than real time. TCP hides this (it just buffers ahead, so playback timing still looks right to a casual test), but UDP has no flow control: bursting several-times-realtime traffic over loopback or a real network overruns the receive buffer and the client reports massive RTP sequence gaps ("missed N packets") followed by decode errors - exactly the "TCP works, UDP is garbage" symptom. RTSPFormatMP3 now sizes its fragments dynamically from AudioEncoder::frameSize() (backed by MP3ParserEncoder, which knows the real parsed frame size) once the encoder has seen at least one frame, instead of a fixed guess - if you write a new RTSPFormat/encoder pairing for a frame-oriented codec, make sure fragmentSize() tracks the codec's real frame size the same way, or you will hit this exact failure mode again.

4. Video: MJPEG

Video sources are packetized (RFC 2435 for JPEG, RFC 6184 for H.264): the encoder itself builds each RTP payload, fragment offsets and all. Unlike a byte-stream codec, JPEGRtpEncoder/H264RtpEncoder (RTSPVideoEncoder) are themselves IMediaSources (RTSPVideoEncoder : AudioEncoder, IMediaSource), so they can go straight into RTSPMediaStreamer - no RTSPMediaCallbackSource/callback wrapper needed:

#include "AudioTools.h"
#include "AudioTools/Communication/RTSP/RTSPPlatformWiFi.h"
#include "AudioTools/Communication/RTSP.h"

RTSPFormatMJPEG mjpegFormat(640, 480, 15.0f);
JPEGRtpEncoder jpegEncoder;

RTSPMediaStreamer<RTSPPlatformWiFi> rtspStreamer(jpegEncoder);
RTSPServer<RTSPPlatformWiFi> rtspServer(rtspStreamer, 8554);

void setup() {
  jpegEncoder.setFormat(mjpegFormat);
  jpegEncoder.setMaxFragmentSize(1400);
  rtspServer.begin();
}

void loop() {
  // gate capture on rtspServer.clientCount(), then:
  // jpegEncoder.write(jpegFrameBytes, jpegFrameLen);
}

(RTSPMediaCallbackSource still earns its keep for a source that isn't already an IMediaSource/Stream - e.g. wiring a plain C camera driver's read function straight in without an adapter object.)

See examples/examples-communication/rtsp/video-server-mjpeg/ for the full ESP32-CAM version, and tests-cmake/network/rtsp-video/rtsp-video.cpp for a desktop-runnable copy that loops a static test frame instead of a live camera.

Gotcha #3: RFC 2435 needs standard Huffman tables

RFC 2435 never transmits JPEG Huffman tables at all - it hardcodes the default ones from the JPEG standard (Annex K) and expects every source JPEG to use them. Most software JPEG encoders (including ffmpeg's mjpeg encoder) default to optimized, per-image Huffman tables to save a few hundred bytes, which is exactly what you don't want here: a receiver decodes with the wrong tables from the very first symbol, producing garbage from block 0 on, deterministically, every single frame. If you generate test frames with ffmpeg, force the standard tables explicitly:

ffmpeg -i input.png -pix_fmt yuvj420p -huffman default out.jpg

A real camera (ESP32-CAM, OV2640/OV3660 etc.) already uses the standard tables, so this only bites synthetic/test JPEGs, not real capture hardware.

Gotcha #4: RFC 2435 needs exactly two quantization tables

RFC 2435's Q=255 (custom-table) path assumes a standard baseline JPEG layout with two quantization tables: one for luma, one for chroma. Some encoders - again, easy to hit with synthetic/low-detail test images

  • emit a single table shared by all three components instead (SOF's Tq field is 0 for every component). JPEGRtpEncoder handles this for you (it duplicates the one table it finds into both slots so the RFC 2435 wire format's two-table expectation is still satisfied), but if you're debugging a different JPEG/RTP implementation and see decode errors localized at y=0 x=0 on every frame with an otherwise-plausible bitstream, this is the first thing to check.

Video: H.264

Same shape as MJPEG, with RTSPFormatH264 + H264RtpEncoder (RFC 6184, packetization-mode=1: single-NAL-unit packets plus FU-A fragmentation for large NALs). SPS/PPS are captured automatically from the first encoded frame and advertised in the SDP (sprop-parameter-sets/ profile-level-id) as soon as they're known, and re-sent in-band before every IDR frame so clients that join mid-stream still get them. See examples/examples-communication/rtsp/video-server-h264/ (needs the external codec-h264-ESP32S3 library) and video-server-wireframe-cube-h264/ (needs TinyGPU + codec-h264-ESP32S3, but no camera - renders a synthetic scene, so it's a good template if you don't have camera hardware handy).

5. Combined audio + video

There is currently no multi-track SDP support - RTSPServer/ RTSPMediaStreamer carry exactly one IMediaSource/RTSPFormat, so one server instance is one m=audio or one m=video line, never both. Two ways to actually get both tracks to a client:

  • Two servers, two URLs (.../audio, .../video, different ports) - simplest, but the client must open two independent sessions with no guaranteed sync between them.
  • Mux to MPEG-TS first, stream the container as one track - RFC 2250 defines a static RTP payload type for MPEG-2 Transport Stream (MP2T, payload type 33). Since an MPEG-TS stream already interleaves and PTS-syncs audio and video internally, sending it as a single m=video ... RTP/AVP 33 track gives every RTSP client (VLC, ffplay, most IP-camera viewers) a fully synced two-track stream from one session/port, no protocol-level RTSP work required. This is what RTSPFormatMTS (RTSPFormat.h) implements - pair it with MuxerMTS (AudioTools/AudioCodecs/ContainerMTS.h) to build the TS bytes, and front the result with RTSPMediaSource (it's a plain Stream, MP2T needs no extra per-fragment header):
RTSPFormatMTS mtsFormat(25.0f);   // fps only paces the RTP send timer
RTSPMediaSource mtsSource(mtsStream, mtsFormat);  // mtsStream: any Stream of .ts bytes
RTSPMediaStreamer<RTSPPlatformWiFi> rtspStreamer(mtsSource);
RTSPServer<RTSPPlatformWiFi> rtspServer(rtspStreamer, 8554);

mtsSource.setFragmentSize(N) should be a whole multiple of MTS_PACKET_SIZE (188 bytes) - RFC 2250 requires each RTP payload to hold a whole number of TS packets; 7 * 188 = 1316 is a common choice that stays under a typical 1500-byte MTU. See tests-cmake/network/rtsp-mts/rtsp-mts.cpp for a full working example, including a small LoopingFileStream wrapper that streams a .ts file from disk on repeat.

Gotcha #5: don't rewind a shared Stream/File from loop()

If your Stream needs to loop (replay a file, replay a test buffer), rewind it inside that Stream's own read()/readBytes() - not from a separate check in Arduino's loop(). RTSPServer sends RTP on its own background task/thread, so a File::seek(0) called from loop() races the streaming thread's concurrent reads on the same non-thread-safe File object: intermittent, hard-to-reproduce decode corruption right around the loop point. AudioTools/Disk/FileLoop.h's FileLoop class already does this correctly for a real File on hardware. On the desktop/Arduino-Emulator specifically, FileLoop::setFile()/its constructor take File by value, which the emulator's File (wraps a non-copyable std::fstream directly) can't support - see rtsp-mts/rtsp-mts.cpp's LoopingFileStream for a minimal from-scratch version of the same pattern that works around it.

6. Testing without hardware

Every example above has a desktop-runnable counterpart in tests-cmake/network/rtsp-*, built against the Arduino-Emulator's WiFiServer/WiFiClient/WiFiUDP/SD/File shims - no board needed:

cd tests-cmake && mkdir build && cd build && cmake .. && make rtsp-sine
./network/rtsp-sine/rtsp-sine &
ffplay -rtsp_transport tcp rtsp://127.0.0.1:8554/

Useful verification commands:

# Confirm the SDP/handshake and reported codec without playing audio/video:
ffprobe -rtsp_transport tcp rtsp://127.0.0.1:8554/

# Play it:
ffplay -rtsp_transport tcp rtsp://127.0.0.1:8554/
# ... or over UDP, matching a real client's default transport:
ffplay rtsp://127.0.0.1:8554/

# Capture a few seconds to a file for offline inspection/diffing against
# a known-good source (e.g. sample-rate/byte-exact comparison):
ffmpeg -rtsp_transport tcp -i rtsp://127.0.0.1:8554/ -t 5 -c copy out.ext

ffprobe alone only proves the SDP/codec identification is correct - it does not exercise full decode, so it will not catch the Huffman-table or quantization-table issues from gotchas #3/#4. Always verify with ffplay (or an actual decode + comparison, as above) before considering a new format/encoder pairing done. And always test UDP, not just TCP - TCP's buffering masks pacing bugs (gotcha #2) that only manifest as real packet loss over UDP, which is what most real RTSP clients use by default.

Clone this wiki locally