-
-
Notifications
You must be signed in to change notification settings - Fork 377
RTSP Server
This tutorial walks through streaming audio and video from an Arduino/ESP32 board (or the desktop, via the Arduino-Emulator) using arduino-audio-tools' RTSP server. It covers the building blocks, a working example for each media type, and a set of gotchas that are easy to hit and hard to diagnose from the symptom alone.
For the class-level API reference, see the headers in this directory
(RTSPServer.h, RTSPFormat.h, RTSPMediaSource.h, ...) and the
Communication wiki page.
Runnable examples referenced throughout live in
examples/examples-communication/rtsp/ (hardware) and
tests-cmake/network/rtsp-* (desktop, no hardware needed).
Every RTSP stream in this library is assembled from four pieces:
-
RTSPFormat- describes one media stream for SDP purposes (codec, payload type, sample rate/dimensions) and its timing (timerPeriodUs(),fragmentSize()). Concrete formats:RTSPFormatPCM,RTSPFormatMP3,RTSPFormatAAC,RTSPFormatMJPEG,RTSPFormatH264,RTSPFormatMTS, ... -
IMediaSource- hands the streamer the actual bytes to send.-
RTSPMediaSourcewraps any ArduinoStream(orAudioStream) - use this for byte-oriented sources: PCM, a codec's encoded output, a file, a generated tone. -
RTSPMediaCallbackSourcewraps a plain C read-callback function instead of aStream- convenient when the data doesn't already live behind aStream(e.g. a camera driver, or a packetized video encoder).
-
-
RTSPMediaStreamer<Platform>- owns the RTP session: builds RTP headers, paces sending via a timer (RTSPFormat::timerPeriodUs()), and writes to the client over TCP-interleaved or UDP transport. -
RTSPServer<Platform>- accepts RTSP connections (DESCRIBE/SETUP/ PLAY/TEARDOWN) and drives oneRTSPMediaStreamer.
Platform is a template parameter binding these to a concrete network
stack - almost always RTSPPlatformWiFi (RTSPPlatformWiFi.h), which is
just RTSPPlatform<WiFiServer, WiFiClient, WiFiUDP>.
RTSPOutput<Platform> is a convenience wrapper that bundles an
RTSPFormat + an AudioEncoder + an internal RTSPMediaSource +
RTSPMediaStreamer into one AudioOutput-shaped object, so it can sit at
the end of an AudioPlayer/StreamCopy pipeline like any other output.
Reach for it for audio; for video, wire RTSPMediaCallbackSource/
RTSPMediaStreamer/RTSPServer together directly (see §3).
The simplest possible server - no encoder, no file, just a live signal:
#include "AudioTools.h"
#include "AudioTools/Communication/RTSP/RTSPPlatformWiFi.h"
#include "AudioTools/Communication/RTSP.h"
AudioInfo info(16000, 1, 16);
SineFromTable<int16_t> sineWave(32000);
GeneratedSoundStream<int16_t> sound(sineWave);
RTSPMediaSource source(sound, info); // Stream + AudioInfo
RTSPMediaStreamer<RTSPPlatformWiFi> streamer(source);
RTSPServer<RTSPPlatformWiFi> rtsp(streamer, 8554);
void setup() {
auto cfg = sineWave.defaultConfig();
cfg.copyFrom(info);
sineWave.begin(cfg, N_B4);
rtsp.begin(); // on ESP32 you can instead call begin(ssid, password)
// to have it join WiFi for you first
}
void loop() { delay(1000); } // RTSP + RTP run on their own tasksConnect with ffplay -rtsp_transport tcp rtsp://<device-ip>:8554/. A full,
desktop-runnable copy of this lives at
tests-cmake/network/rtsp-sine/rtsp-sine.cpp.
#include "AudioTools.h"
#include "AudioTools/Disk/AudioSourceSTD.h" // or AudioSourceSD on hardware
#include "AudioTools/AudioCodecs/MP3Parser.h"
#include "AudioTools/Communication/RTSP/RTSPPlatformWiFi.h"
#include "AudioTools/Communication/RTSP.h"
MP3ParserEncoder enc; // repackages existing MP3 frames - no re-encoding
RTSPFormatMP3 mp3format(enc);
MetaDataFilterEncoder filter(enc); // strips ID3 tags from the stream
RTSPOutput<RTSPPlatformWiFi> rtsp_out(mp3format, filter);
AudioSourceSTD source("/music/", ".mp3");
CopyDecoder decoder; // no decoding - MP3 goes out as MP3
AudioPlayer player(source, rtsp_out, decoder);
RTSPServer<RTSPPlatformWiFi> rtsp(rtsp_out.streamer(), 8554);
void setup() {
mp3format.setUseRfc2250Header(true); // see gotcha #1 below
player.begin();
rtsp_out.begin();
rtsp.begin();
}
void loop() {
if (rtsp_out && rtsp) player.copy();
}A desktop-runnable copy lives at
tests-cmake/network/rtsp-player/rtsp-player.cpp.
RTSPFormatMP3 supports two wire formats for RTP payload type 14
(MPEG audio): raw MP3 frames back-to-back, or frames prefixed with the
optional RFC 2250 4-byte MPEG-audio header (2 reserved bytes + a 2-byte
fragment offset). VLC expects raw frames; ffmpeg/ffplay expect the
4-byte header. The default is off (raw frames). If you only test with
VLC this is invisible; against ffplay every frame boundary silently
shifts by 4 bytes and decodes as pure noise, with no error printed on
either the server or client side to point at the cause. If your stream
"plays" but sounds like static, check this first:
mp3format.setUseRfc2250Header(true);RTSPMediaStreamer's generic byte-stream send path reads exactly
RTSPFormat::fragmentSize() bytes every timerPeriodUs() tick and sends
whatever it got as one RTP packet. If fragmentSize() is bigger than one
actual encoded frame, each packet silently carries several frames'
worth of data, sent at one frame's timing - i.e. several times faster
than real time. TCP hides this (it just buffers ahead, so playback
timing still looks right to a casual test), but UDP has no flow
control: bursting several-times-realtime traffic over loopback or a
real network overruns the receive buffer and the client reports massive
RTP sequence gaps ("missed N packets") followed by decode errors -
exactly the "TCP works, UDP is garbage" symptom. RTSPFormatMP3 now
sizes its fragments dynamically from AudioEncoder::frameSize() (backed
by MP3ParserEncoder, which knows the real parsed frame size) once the
encoder has seen at least one frame, instead of a fixed guess - if you
write a new RTSPFormat/encoder pairing for a frame-oriented codec,
make sure fragmentSize() tracks the codec's real frame size the same
way, or you will hit this exact failure mode again.
Video sources are packetized (RFC 2435 for JPEG, RFC 6184 for H.264): the
encoder itself builds each RTP payload, fragment offsets and all. Unlike a
byte-stream codec, JPEGRtpEncoder/H264RtpEncoder (RTSPVideoEncoder)
are themselves IMediaSources (RTSPVideoEncoder : AudioEncoder, IMediaSource), so they can go straight into RTSPMediaStreamer - no
RTSPMediaCallbackSource/callback wrapper needed:
#include "AudioTools.h"
#include "AudioTools/Communication/RTSP/RTSPPlatformWiFi.h"
#include "AudioTools/Communication/RTSP.h"
RTSPFormatMJPEG mjpegFormat(640, 480, 15.0f);
JPEGRtpEncoder jpegEncoder;
RTSPMediaStreamer<RTSPPlatformWiFi> rtspStreamer(jpegEncoder);
RTSPServer<RTSPPlatformWiFi> rtspServer(rtspStreamer, 8554);
void setup() {
jpegEncoder.setFormat(mjpegFormat);
jpegEncoder.setMaxFragmentSize(1400);
rtspServer.begin();
}
void loop() {
// gate capture on rtspServer.clientCount(), then:
// jpegEncoder.write(jpegFrameBytes, jpegFrameLen);
}(RTSPMediaCallbackSource still earns its keep for a source that isn't
already an IMediaSource/Stream - e.g. wiring a plain C camera driver's
read function straight in without an adapter object.)
See examples/examples-communication/rtsp/video-server-mjpeg/ for the
full ESP32-CAM version, and
tests-cmake/network/rtsp-video/rtsp-video.cpp for a desktop-runnable
copy that loops a static test frame instead of a live camera.
RFC 2435 never transmits JPEG Huffman tables at all - it hardcodes the
default ones from the JPEG standard (Annex K) and expects every source
JPEG to use them. Most software JPEG encoders (including ffmpeg's
mjpeg encoder) default to optimized, per-image Huffman tables to
save a few hundred bytes, which is exactly what you don't want here: a
receiver decodes with the wrong tables from the very first symbol,
producing garbage from block 0 on, deterministically, every single
frame. If you generate test frames with ffmpeg, force the standard
tables explicitly:
ffmpeg -i input.png -pix_fmt yuvj420p -huffman default out.jpgA real camera (ESP32-CAM, OV2640/OV3660 etc.) already uses the standard tables, so this only bites synthetic/test JPEGs, not real capture hardware.
RFC 2435's Q=255 (custom-table) path assumes a standard baseline JPEG layout with two quantization tables: one for luma, one for chroma. Some encoders - again, easy to hit with synthetic/low-detail test images
- emit a single table shared by all three components instead (SOF's
Tqfield is 0 for every component).JPEGRtpEncoderhandles this for you (it duplicates the one table it finds into both slots so the RFC 2435 wire format's two-table expectation is still satisfied), but if you're debugging a different JPEG/RTP implementation and see decode errors localized aty=0 x=0on every frame with an otherwise-plausible bitstream, this is the first thing to check.
Same shape as MJPEG, with RTSPFormatH264 + H264RtpEncoder (RFC 6184,
packetization-mode=1: single-NAL-unit packets plus FU-A fragmentation
for large NALs). SPS/PPS are captured automatically from the first
encoded frame and advertised in the SDP (sprop-parameter-sets/
profile-level-id) as soon as they're known, and re-sent in-band before
every IDR frame so clients that join mid-stream still get them. See
examples/examples-communication/rtsp/video-server-h264/ (needs the
external codec-h264-ESP32S3 library) and
video-server-wireframe-cube-h264/ (needs TinyGPU +
codec-h264-ESP32S3, but no camera - renders a synthetic scene, so it's
a good template if you don't have camera hardware handy).
There is currently no multi-track SDP support - RTSPServer/
RTSPMediaStreamer carry exactly one IMediaSource/RTSPFormat, so one
server instance is one m=audio or one m=video line, never both. Two
ways to actually get both tracks to a client:
-
Two servers, two URLs (
.../audio,.../video, different ports) - simplest, but the client must open two independent sessions with no guaranteed sync between them. -
Mux to MPEG-TS first, stream the container as one track - RFC 2250
defines a static RTP payload type for MPEG-2 Transport Stream (MP2T,
payload type 33). Since an MPEG-TS stream already interleaves and
PTS-syncs audio and video internally, sending it as a single
m=video ... RTP/AVP 33track gives every RTSP client (VLC, ffplay, most IP-camera viewers) a fully synced two-track stream from one session/port, no protocol-level RTSP work required. This is whatRTSPFormatMTS(RTSPFormat.h) implements - pair it withMuxerMTS(AudioTools/AudioCodecs/ContainerMTS.h) to build the TS bytes, and front the result withRTSPMediaSource(it's a plainStream, MP2T needs no extra per-fragment header):
RTSPFormatMTS mtsFormat(25.0f); // fps only paces the RTP send timer
RTSPMediaSource mtsSource(mtsStream, mtsFormat); // mtsStream: any Stream of .ts bytes
RTSPMediaStreamer<RTSPPlatformWiFi> rtspStreamer(mtsSource);
RTSPServer<RTSPPlatformWiFi> rtspServer(rtspStreamer, 8554);mtsSource.setFragmentSize(N) should be a whole multiple of
MTS_PACKET_SIZE (188 bytes) - RFC 2250 requires each RTP payload to
hold a whole number of TS packets; 7 * 188 = 1316 is a common choice
that stays under a typical 1500-byte MTU. See
tests-cmake/network/rtsp-mts/rtsp-mts.cpp for a full working example,
including a small LoopingFileStream wrapper that streams a .ts file
from disk on repeat.
If your Stream needs to loop (replay a file, replay a test buffer),
rewind it inside that Stream's own read()/readBytes() - not from
a separate check in Arduino's loop(). RTSPServer sends RTP on its own
background task/thread, so a File::seek(0) called from loop() races
the streaming thread's concurrent reads on the same non-thread-safe
File object: intermittent, hard-to-reproduce decode corruption right
around the loop point. AudioTools/Disk/FileLoop.h's FileLoop class
already does this correctly for a real File on hardware. On the
desktop/Arduino-Emulator specifically, FileLoop::setFile()/its
constructor take File by value, which the emulator's File (wraps a
non-copyable std::fstream directly) can't support - see
rtsp-mts/rtsp-mts.cpp's LoopingFileStream for a minimal from-scratch
version of the same pattern that works around it.
Every example above has a desktop-runnable counterpart in
tests-cmake/network/rtsp-*, built against the
Arduino-Emulator's
WiFiServer/WiFiClient/WiFiUDP/SD/File shims - no board needed:
cd tests-cmake && mkdir build && cd build && cmake .. && make rtsp-sine
./network/rtsp-sine/rtsp-sine &
ffplay -rtsp_transport tcp rtsp://127.0.0.1:8554/Useful verification commands:
# Confirm the SDP/handshake and reported codec without playing audio/video:
ffprobe -rtsp_transport tcp rtsp://127.0.0.1:8554/
# Play it:
ffplay -rtsp_transport tcp rtsp://127.0.0.1:8554/
# ... or over UDP, matching a real client's default transport:
ffplay rtsp://127.0.0.1:8554/
# Capture a few seconds to a file for offline inspection/diffing against
# a known-good source (e.g. sample-rate/byte-exact comparison):
ffmpeg -rtsp_transport tcp -i rtsp://127.0.0.1:8554/ -t 5 -c copy out.extffprobe alone only proves the SDP/codec identification is correct - it
does not exercise full decode, so it will not catch the Huffman-table or
quantization-table issues from gotchas #3/#4. Always verify with
ffplay (or an actual decode + comparison, as above) before considering
a new format/encoder pairing done. And always test UDP, not just
TCP - TCP's buffering masks pacing bugs (gotcha #2) that only manifest
as real packet loss over UDP, which is what most real RTSP clients use
by default.