audio-branding-and-storytelling
The Evolution of Network Audio Protocols: From Cobranet to Aes67
Table of Contents
The landscape of professional audio has been reshaped by the shift from analog point-to-point wiring to digital audio networking. Early systems relied on bulky multicore cables and dedicated analog consoles, limiting flexibility and scalability. As Ethernet became ubiquitous, engineers sought ways to transport high-quality audio over standard network infrastructure. This transition gave rise to a series of network audio protocols, each addressing the growing demands for channel count, latency, and interoperability. From the pioneering CobraNet to the open standard AES67, the evolution of these protocols reflects the industry's relentless pursuit of flexibility, reliability, and seamless integration across diverse platforms.
Early Network Audio Protocols: CobraNet and EtherSound
CobraNet: The First Ethernet Audio Standard
Introduced in the late 1990s by Peak Audio (later acquired by Cirrus Logic), CobraNet was one of the first commercially successful protocols for transmitting uncompressed digital audio over standard Ethernet. It operated at 100 Mbps and could carry up to 64 channels at 48 kHz, 20-bit resolution. CobraNet used a proprietary packet structure and relied on a unique "CobraNet" switch to manage clock synchronization and data flow. While it reduced the complexity of analog wiring for installed sound and live events, its proprietary nature meant that only hardware from licensed manufacturers could interoperate. This vendor lock-in limited adoption and paved the way for more open solutions.
CobraNet’s synchronization mechanism used a master clock that sent beacon packets; all devices locked to this clock, ensuring deterministic latency. However, the protocol was not designed for dynamic routing or scalability across multiple subnets. As network demands grew, the 64-channel limit and the need for dedicated switches became significant drawbacks.
EtherSound: A Contemporaneous Contender
Around the same time, EtherSound emerged from Digigram and later evolved under the leadership of Lawo and others. EtherSound offered lower latency than CobraNet—down to 125 µs—and supported up to 64 channels in one direction. It required dedicated switches and used a daisy-chain topology that simplified cabling but complicated redundancy. EtherSound found a home in broadcast and recording studios where low latency was paramount, but like CobraNet, it was proprietary and lacked the interoperability that the industry soon demanded.
The Rise of Open Standards: Dante and Ravenna
Dante: Scalability and Ubiquity
Developed by Audinate, Dante debuted in the mid-2000s and quickly became the dominant network audio protocol for installed sound, live events, and fixed installations. Dante runs over standard Gigabit Ethernet (and now 10GbE) and supports up to 512 channels per link at 48 kHz, with lower resolutions allowing even higher channel counts. Its key innovation was the use of a self-managing network topology: Dante devices automatically discover each other and route audio using a software controller, removing the need for complex switch configuration.
Dante employs IEEE 1588 Precision Time Protocol (PTP) for clock synchronization, achieving extremely low jitter and enabling sample‐accurate playback across multiple channels. It also supports redundancy via Dante Dual Redundant (primary and secondary networks). Dante’s widespread adoption—over 3,000 product families—makes it a de facto standard for many manufacturers, including Yamaha, Shure, Bose, and Allen & Heath. However, Dante remains a proprietary protocol (though Audinate licenses it openly), which means that native Dante devices cannot directly communicate with devices using other protocols without a bridge or converter.
Ravenna: Open and Modular
Ravenna was developed by ALC NetworX (now part of the Lawo group) as an open, IP-based audio networking solution. Unlike Dante, Ravenna is built entirely on open standards: RTP (Real‐time Transport Protocol) for audio payload, PTP (IEEE 1588) for synchronization, and SAP (Session Announcement Protocol) for discovery. This open architecture allows any manufacturer to implement Ravenna without licensing fees, fostering interoperability across brands.
Ravenna supports up to 1024 channels per stream at up to 192 kHz sample rate, with sub-millisecond latency. It is widely used in broadcast, broadcast production, and live sound environments where high channel counts and precise timing are essential. Major adopters include Lawo, Neumann, Merging Technologies, and DirectOut. Ravenna also integrates with Audio over IP standards like AES67 and SMPTE ST 2110, making it future-proof for media networks.
AES67: The Standard for Interoperability
In 2013, the Audio Engineering Society published AES67-2013, a standard designed to enable interoperability between different IP‐based audio networking protocols. It defines a common framework for synchronization, media transport, and device discovery, allowing devices from different ecosystems (Dante, Ravenna, Livewire, Q‑LAN, etc.) to exchange audio without needing custom gateways. AES67 is not a full protocol stack but a profile of existing standards that any compliant device must follow.
Key Features of AES67
- High‑precision synchronization using IEEE 1588‑2008 (PTPv2). Devices must support the profile defined in AES67, which specifies a clock accuracy of ±0.05 ppm and a synchronization domain with sub‑microsecond timing.
- Media transport via RTP over UDP, using 16‑, 20‑, or 24‑bit PCM audio, with sample rates of 44.1, 48, 96, and 192 kHz. The standard mandates support for at least 8 audio channels per stream.
- Discovery and connection management using SAP (Session Announcement Protocol) for advertising streams and mDNS for device discovery. Some implementations also use RTSP for session setup.
- QoS (Quality of Service) hints using DiffServ to prioritize audio traffic, ensuring consistent performance on managed networks.
How AES67 Bridges Ecosystems
The primary achievement of AES67 is that it defines a common “language” that Dante, Ravenna, Livewire, and other protocols can all speak. For example, a Dante device can output an AES67‑compliant stream that a Ravenna device can receive natively, provided both implement the AES67 profile. This interoperability is critical for large‑scale installations where equipment from multiple vendors must work together—such as stadiums, broadcast centers, and convention venues.
However, AES67 does not cover every feature of each native protocol. For instance, Dante’s automatic routing and controller software are not part of AES67; manual or third‑party management tools are needed to set up AES67 streams. Similarly, Ravenna’s redundant networking and certain advanced clocking modes go beyond the AES67 profile. Despite these limitations, AES67 has become a crucial layer for “the last mile” of compatibility.
Comparison of Major Network Audio Protocols
| Protocol | Year Introduced | Channel Count (per link) | Latency (typical) | Synchronization | Proprietary/Open |
|---|---|---|---|---|---|
| CobraNet | 1996 | 64 | 1.33 ms or 5.33 ms | Proprietary beacon | Proprietary |
| EtherSound | 2001 | 64 | 125 µs | Master/slave | Proprietary |
| Dante | 2006 | 512 (1 GbE) | 150 µs – 1 ms | IEEE 1588 PTP | Proprietary (open license) |
| Ravenna | 2010 | 1024 | 125 µs – 1 ms | IEEE 1588 PTP | Open standard |
| AES67 | 2013 | Varies (>=8) | Depends on implementation | IEEE 1588 PTP | Open standard |
Note: Channel counts and latencies assume typical configurations; actual performance depends on network design and hardware capabilities.
Implementation Considerations for AES67 Networks
Network Infrastructure
Deploying AES67 requires a managed Ethernet network that supports PTP transparency or boundary clocks. Switches must handle multicast traffic efficiently (IGMP snooping) and prioritize audio packets using DiffServ code points (DSCP). For larger systems, a dedicated VLAN for audio traffic is recommended to isolate control and data streams from other network congestion.
Latency Management
AES67 defines three latency profiles: low (sub‑millisecond), medium (1–3 ms), and high (up to 10 ms). The low profile is essential for live monitoring and broadcast intercom, while higher latencies are acceptable for playback systems. When mixing AES67 with native protocol streams, the system designer must ensure that all devices operate within the same latency domain to prevent echo or timing drift.
Clock Redundancy
While AES67 specifies a single grandmaster clock, many implementations support redundant clock sources via PTP. A secondary grandmaster can take over if the primary fails, ensuring uninterrupted operation. In multi‑vendor environments, careful clock hierarchy planning is required to avoid loops or instability.
Applications Across Industries
Live Sound and Touring
Large concerts and festivals rely on Dante and Ravenna for distributing hundreds of channels between FOH, monitors, broadcast trucks, and recording rigs. AES67 compatibility allows mixing consoles from different manufacturers (e.g., Yamaha and DiGiCo) to share audio without custom analog or MADI bridges. Touring companies value the ability to interconnect gear that supports AES67, reducing truck space and setup time.
Broadcast and Production
Television and radio studios have adopted AoIP extensively. AES67 forms the backbone of modern “all‑IP” broadcast plants, integrating with SMPTE ST 2110 for video transport. This convergence simplifies cabling, enables remote production, and reduces operational costs. Newsrooms and control rooms benefit from seamless audio routing across multiple studios.
Installed Sound and Commercial AV
Corporate boardrooms, houses of worship, and airports use Dante and Ravenna for paging systems, background music, and distributed audio. AES67 ensures that devices from different manufacturers—such as amplifiers, DSPs, and microphones—can communicate without requiring a single vendor’s ecosystem. This flexibility is essential for large projects that are built over multiple phases.
Recording and Post‑Production
High‑end studios employ Ravenna for its support of PCM audio up to 192 kHz and its low jitter. Merging Technologies’ Ovation and Pyramix systems are native Ravenna, while other DAWs can connect via AES67 interfaces. The ability to transport multichannel mixes over a single cable simplifies patching and reduces signal degradation.
The Future of Network Audio Protocols
Convergence with AVB/TSN
Audio Video Bridging (AVB) and Time‑Sensitive Networking (TSN) are IEEE standards that extend the capabilities of standard Ethernet with deterministic latency and guaranteed bandwidth. While AVB was initially promoted by the AVnu Alliance, its adoption in professional audio has been slower than Dante or Ravenna. However, TSN is now a key component of SMPTE ST 2110 for video, and audio protocols are increasingly integrating TSN features (e.g., gPTP). AES67 can already run on networks that support TSN, and future revisions may include TSN profiles for even tighter synchronization.
SMPTE ST 2110 and IP‑Based Media
The broadcast industry is moving toward fully IP‑based production using SMPTE ST 2110, which separates video, audio, and ancillary data into individual RTP streams. AES67 is the audio component of ST 2110‑30 (PCM audio) and ST 2110‑31 (AES3 transport). As more broadcast facilities adopt ST 2110, AES67’s role as the common audio transport will only grow.
Higher Sample Rates and Immersive Audio
Demand for high‑resolution audio (up to 384 kHz) and object‑based immersive formats (Dolby Atmos, MPEG‑H) requires protocols that can handle larger payloads and more channels. Ravenna already supports 384 kHz; Dante is extending its capabilities via AES67 profiles. Future AES67 revisions may standardize sample rates beyond 192 kHz and support for multichannel objects.
Security and Authentication
Network audio protocols have historically paid little attention to security, assuming trusted physical networks. With the rise of remote production and cloud‑based workflows, encryption and authentication are becoming critical. Dante and Ravenna have added basic security features, and the AES67 standard is likely to include security recommendations in future updates. Implementation of 802.1X, IPsec, or DTLS may become mandatory for certain applications.
Conclusion
The evolution of network audio protocols from CobraNet to AES67 mirrors the broader journey of digital audio from proprietary islands to a unified, interoperable ecosystem. CobraNet and EtherSound proved the concept of audio over Ethernet but suffered from vendor lock‑in. Dante and Ravenna brought scalability and openness, with Dante dominating the commercial market and Ravenna leading in broadcast and high‑end recording. AES67 emerged as the critical bridge, enabling devices from different ecosystems to coexist and communicate.
Today, audio professionals can choose from a rich palette of protocols, each with strengths tailored to specific use cases. The trend toward IP‑based media standards like SMPTE ST 2110 and the integration of TSN promise even tighter integration across audio, video, and control systems. For educators and students, understanding this evolution is essential—not only to appreciate past innovations but to anticipate the future of networked audio. As bandwidth increases and latency decreases, the boundaries between local and remote, between studio and live sound, will continue to blur. Open standards such as AES67 will remain the foundation upon which the next generation of audio experiences is built.