DDS Benchmarking Suite
Comparable DDS Measurements
Published measurements for tin-dds, Retina, CycloneDDS, iceoryx, Fast DDS, and RTI Connext. Network results use separate publisher and subscriber processes on one host. Local fanout results measure one participant delivering to as many as 16 readers. Internal shared-memory mechanisms and full DDS paths are shown separately. Network latency is reported as a one-way estimate (RTT ÷ 2). Local fanout latency is measured directly.
Complete reliable DDS exchanges at three application payload sizes.
The charts keep typical and tail behavior visible on one scale.
Every five-implementation run completed all 10,000 measured exchanges without reported loss.
Complete DDS measurement boundary
The complete DDS charts time every application stage from
write() to take(). Discovery and matching
happen before the timer starts.
- 1Application writes
write() - 2Serialize dataCDR encoding
- 3Track deliveryWriter history + reliability
- 4Build messageRTPS DATA framing
- 5Move the dataSHM handoff or UDP
- 6Parse messageRTPS parsing
- 7Order and storeReader reliability + history
- 8Decode dataCDR deserialization
- 9Application takes
take()
Current Tin DDS and Retina build
This current-source cohort exchanged the same 64-byte application sample
between two CPU-pinned processes over reliable UDP. Each implementation
completed five interleaved runs of 10,000 measured round trips after 500
warm-up exchanges. The timer covers the complete public
write()-to-take() path, and all 100,000
measured exchanges arrived without loss, duplication, or rejection.
Loading current-source latency chart…
Loading current-source completion-rate chart…
| Implementation | p50 | p99 | p99.9 | Exchanges/s |
|---|---|---|---|---|
| Tin DDS 0.1.0 | 5.75 µs | 7.96 µs | 11.63 µs | 82,262 |
| Retina 0.1.0 | 6.54 µs | 9.24 µs | 11.59 µs | 74,051 |
The median statistics come from one ARM64 host session. Its scheduler and power state can move absolute latency, so this cohort remains separate from the earlier five-implementation comparison. The machine-readable matrix links to all ten raw observation artifacts and their source and compiler metadata.
Five DDS implementations, one workload
Tin DDS, Retina, CycloneDDS, Fast DDS, and RTI Connext each exchange the same 64-byte application sample between two processes over reliable UDP. Discovery and 500 warm-up exchanges finish before 10,000 observations are retained. The charts report the median statistic from five interleaved, CPU-pinned runs on one ARM64 Linux host. Batching is off and durability is volatile for every implementation. Each row is a homogeneous pair of that implementation; this chart measures performance, not mixed-vendor interoperability.
Loading head-to-head latency chart…
Loading head-to-head completion-rate chart…
Exact numeric values
| Implementation | p50 | p99 | p99.9 | Exchanges/s |
|---|---|---|---|---|
| Tin DDS 0.1.0 | 5.20 µs | 6.92 µs | 10.50 µs | 93,223 |
| Retina 0.1.0 | 6.43 µs | 9.09 µs | 10.91 µs | 75,052 |
| CycloneDDS 0.10.4 | 16.36 µs | 18.65 µs | 77.42 µs | 30,006 |
| Fast DDS 2.14.6 | 19.26 µs | 26.05 µs | 40.56 µs | 24,692 |
| RTI Connext 7.7.0 | 22.00 µs | 34.00 µs | 156.00 µs | 22,464 |
Every run retained exactly 10,000 observations and completed without a reported lost exchange. The machine-readable matrix links to all 25 run artifacts and their raw latency observations.
Reliable DDS latency as payload size grows
Tin DDS and Retina completed 90,000 reliable round trips through their public writer and reader APIs with zero lost samples. Each run establishes discovery and endpoint matching before 500 warm-up exchanges, then measures serialization, reliability, history, UDP transport, and application delivery. The table reports the median of three interleaved runs on a provenance-complete, slim Tinverse Linux Guix GCP runner image.
Loading latency chart…
Loading completion-rate chart…
| Payload | Tin DDS p50 | Tin DDS p99 | Retina p50 | Retina p99 |
|---|---|---|---|---|
| 64 B | 3.97 µs | 5.75 µs | 2.85 µs | 4.58 µs |
| 1 KiB | 4.28 µs | 5.97 µs | 3.16 µs | 4.94 µs |
| 4 KiB | 6.47 µs | 8.54 µs | 8.78 µs | 11.21 µs |
These runs used the Tinverse Linux v4-lto policy on a
c3-standard-4 VM, with publishers pinned to CPU 1 and
subscribers to CPU 3.
Complete DDS publish-to-receive latency
These Retina results measure from the application calling
write() until the receiving application obtains the sample
with take(). Each direction includes CDR serialization,
writer history and reliability, RTPS DATA framing, shared memory or UDP,
RTPS parsing, reader reliability and history, and CDR deserialization.
SPDP/SEDP discovery, endpoint matching, and warm-up finish before timing
begins.
Loading transport-latency chart…
Loading transport-rate chart…
| Architecture | Transport | p50 | p99 | p99.9 | Maximum | Round trips/s |
|---|---|---|---|---|---|---|
| x86-64 | UDP | 5.401 µs | 6.980 µs | 8.193 µs | 138.893 µs | 90,043 |
| x86-64 | Shared memory | 0.971 µs | 1.521 µs | 2.188 µs | 71.543 µs | 465,394 |
| ARM64 | UDP | 8.838 µs | 12.157 µs | 15.972 µs | 3.408 ms | 59,084 |
| ARM64 | Shared memory | 3.217 µs | 3.921 µs | 5.102 µs | 4.226 ms | 149,227 |
Each row is the median of five interleaved one-million-sample runs. The latency is a one-way estimate obtained by dividing a complete ping-pong round trip by two. Every run delivered every sample without loss or retransmission. The machine-readable result records the measurement boundary, aggregation, and delivery counters.
Reliable when delivery is disrupted
Tin DDS and Retina each recovered every one of 15,000 trials for four RTPS reliability behaviors: rejecting duplicate data, accepting data that arrives out of order, requesting a missing sequence, and requesting a missing fragment. Across both implementations, all 120,000 trials completed successfully.
| Reliability behavior | Tin DDS | Retina |
|---|---|---|
| Duplicate data rejected | 15,000 / 15,000 | 15,000 / 15,000 |
| Out-of-order data accepted | 15,000 / 15,000 | 15,000 / 15,000 |
| Missing sequence requested | 15,000 / 15,000 | 15,000 / 15,000 |
| Missing fragment requested | 15,000 / 15,000 | 15,000 / 15,000 |
These deterministic checks isolate the reliability mechanisms from the network. Their internal timings cover different amounts of work, so the result reported here is recovery correctness, not a speed ranking. The summary and complete trial records are available for inspection.
Reliable and best-effort local fanout
One write was delivered to 1, 4, or 16 readers through each implementation's DDS history and matching logic. Tin DDS and Retina completed 180,000 measured write-and-take iterations with no lost, duplicate, or rejected deliveries. The Release runs alternate implementation order and retain every latency sample.
- 1 Write
- 2 Serialize
- 3 Writer state
- 4 Frame
- 5 Transport
- 6 Parse
- 7 Reader state
- 8 Deserialize
- 9 Take
Loading best-effort fanout chart…
Loading reliable fanout chart…
Exact numeric values
| Implementation | QoS | Readers | p50 | p99 | Deliveries/s |
|---|---|---|---|---|---|
| Tin DDS | Best effort | 1 | 287 ns | 324 ns | 2.65M |
| Tin DDS | Best effort | 4 | 574 ns | 620 ns | 5.92M |
| Tin DDS | Best effort | 16 | 1.70 µs | 1.93 µs | 8.91M |
| Retina | Best effort | 1 | 426 ns | 473 ns | 1.87M |
| Retina | Best effort | 4 | 602 ns | 981 ns | 5.61M |
| Retina | Best effort | 16 | 1.72 µs | 3.04 µs | 8.58M |
| Tin DDS | Reliable | 1 | 296 ns | 445 ns | 2.49M |
| Tin DDS | Reliable | 4 | 583 ns | 982 ns | 5.94M |
| Tin DDS | Reliable | 16 | 1.67 µs | 2.50 µs | 9.00M |
| Retina | Reliable | 1 | 315 ns | 565 ns | 2.42M |
| Retina | Reliable | 4 | 621 ns | 1.05 µs | 4.88M |
| Retina | Reliable | 16 | 1.69 µs | 1.89 µs | 8.98M |
Latency is one complete logical iteration: one write followed by a take from every reader. Delivery rate counts reader deliveries, so it rises with fanout. This is local DDS dispatch, not UDP latency. See the summary and raw child artifacts.
Performance with the DDS behavior included
The reliable UDP results exercise public typed writers and readers in separate processes. Discovery, endpoint matching, serialization, reliability, history, transport, and application delivery are part of the path. A run passes only when its delivery accounting is complete; latency alone is not enough.
Typed end to end
The application writes and receives concrete data types, not transport test packets.
Delivery is measured
Artifacts retain sent, received, lost, duplicate, and rejected sample counts when the runner provides them.
Scopes stay separate
Shared-memory mechanism timings are never presented as though they were a vendor's complete DDS path.
Split-process cohort (July–August 2026)
Every row uses a publisher and subscriber in separate persistent processes on one host. Measurements were collected on an ARM64 system and an x86-64 system. The linked artifacts record the workload and measurement method; duration-based vendor tools may complete a different number of round trips than the fixed-iteration tin-dds and Retina runners.
Stage 5 only: shared-memory handoff
These rows time direct payload-slot and notification work. They do not include the public writer and reader APIs, serialization, DDS history, RTPS framing and parsing, or deserialization.
- 1 Write
- 2 Serialize
- 3 Writer state
- 4 Frame
- 5 Transport
- 6 Parse
- 7 Reader state
- 8 Deserialize
- 9 Take
| Implementation | p50 | p99 |
|---|---|---|
| tin-dds (daemon segment, loaned payload) | 412 ns | 486 ns |
| Retina (daemon loan and SHM notification rings) | 434 ns | 490 ns |
Full DDS shared-memory references
| Implementation | p50 | p99 |
|---|---|---|
| RTI Connext 7.7 (full DDS over SHMEM) | 5,000 ns | 6,000 ns |
| Fast DDS 2.6 (data-sharing zero-copy) | 9,030 ns | 10,813 ns |
Reliable UDP, two processes
Loading x86-64 implementation comparison…
| Implementation | p50 | p99 | Round trips/s |
|---|---|---|---|
| tin-dds | 3,171 ns | 3,318 ns | 157,105 |
| Retina | 1,997 ns | 3,128 ns | 243,099 |
| RTI Connext 7.7 | 6,000 ns | 11,000 ns | 56,204 |
| CycloneDDS 0.10.5 (ddsperf) | 2,378 ns | 3,044 ns | — |
| Fast DDS 2.6 | 9,336 ns | 12,409 ns | — |
Stage 5 only: shared-memory handoff
These rows time direct payload-slot and notification work. They do not include the public writer and reader APIs, serialization, DDS history, RTPS framing and parsing, or deserialization.
- 1 Write
- 2 Serialize
- 3 Writer state
- 4 Frame
- 5 Transport
- 6 Parse
- 7 Reader state
- 8 Deserialize
- 9 Take
| Implementation | p50 | p99 |
|---|---|---|
| iceoryx RouDi 2.0.5 (loaned samples) | 2,467 ns | 2,861 ns |
| Retina (daemon loan and SHM notification rings) | 2,153 ns | 2,473 ns |
Full DDS shared-memory reference
| Implementation | p50 | p99 |
|---|---|---|
| Fast DDS 2.11 (data-sharing zero-copy) | 38,157 ns | 45,148 ns |
Reliable UDP, two processes
| Implementation | p50 | p99 | Exchanges/s |
|---|---|---|---|
| Tin DDS 0.1.0 | 5.20 µs | 6.92 µs | 93,223 |
| Retina 0.1.0 | 6.43 µs | 9.09 µs | 75,052 |
| CycloneDDS 0.10.4 | 16.36 µs | 18.65 µs | 30,006 |
| Fast DDS 2.14.6 | 19.26 µs | 26.05 µs | 24,692 |
| RTI Connext 7.7.0 | 22.00 µs | 34.00 µs | 22,464 |
The rows share a process topology and payload size. Mechanism tables time loan and notification work; full DDS tables include their implementation's wider DDS path. Those scopes are deliberately separated and are not ranked against each other. Treat these as same-host reference points, not universal rankings: CPU scheduling, power mode, configuration, and payload size all affect the result.
The ARM64 UDP rows were refreshed together on August 18, 2026 using the matched five-implementation workload above. The x86-64 reference rows and ARM64 shared-memory mechanism rows come from their separately linked cohorts and are not mixed into that ranking.