Thank you! You’re right, I didn’t discuss the MIDI implementation in the blog post. Right now it’s a little primitive but functional. Each MIDI message (e.g 3 byte note-on message) maps to 1 UDP packet. I have a simple script on the host side that interfaces these UDP packets with an ALSA loopback MIDI interface.
As someone who uses a USB equivalent to this piece of hardware (in my case, a Behringer UMC1820): having MIDI and audio inputs on the same interface is a nice convenience feature. Avoids needing to use up multiple ports on the host machine, and avoids a lot of the time synchronization issues that arise with trying to use multiple interfaces at once.
They're basically interchangeable. But yeah TRS is more compact and more normative when there is no preamp. Some monitors I've seen have xlr inputs though, even though they expect line level input.
Another option is to just use XLR/TRS combo jacks if you want to support both, but that’s not really any different than the user using an XLR to TRS cord or converter
Well, it's certainly a dumb situation that I don't like that's created by my employer (and many others with similar setups), but it takes only two clicks to pick a category to make it available to folks behind these firewalls.
I went ahead and chose "blog" because it sounds like it's a blog. Apparently the "newly seen domains" category goes away on its own after 30 days, so that would have fixed itself, though then it would have been "uncategorized" and still blocked.
Yes. And no. The affected users usually have no control over the situation, so if the site owner cares at all that those people can't access it, then it is effectively their problem as they can potentially do something about it.
If neither party cares enough, then it is nobody's problem.
That's pretty much it. I wouldn't have mentioned it if not for the fact that it's easy to fix and for users it just looks like the site is down, which is not uncommon for small projects posted to HN.
I guess some people thought I liked cloudflare or something... that's definitely not the case.
"easy to fix" is relative. This "new domain equals bad" thinking has spread in the IT security world. My corporate anti-virus blocks them too and I see it a lot while browsing HN on the job.
The only real solution, as a "holder of a young domain" is to wait 30 days before you use it.
Is this open source? I’ve been working on https://codeberg.org/olpad/openmic which is in the more traditional domain of USB audio interfaces, but I’d be interested in taking a look at the internals of the ETH-68
I’m still working towards a first prototype, development has been slow due to my own schedule. But I’d expect round trip latency near what you’d expect from a USB interface. It would be cool to add ethernet support in the future though!
What's your use case where you want more outputs than inputs? At home I run a Audiofuse 16Rig with a lot of ADAT expansion to get 32i/32o, but I only use like 6o, while every input is used.
The project just exposes everything the hardware can do, the audio codec chip used for the conversion is Burr-Brown PCM3168A 24-Bit, 96-kHz/192-kHz, 6-In/8-Out.
Some I can imagine are surround sound, multiple monitor pairs (including dedicated monitoring channels for different musicians or simply different outputs for mixing).
tech nerd/musician here, I would be very interested in this! do you have an email list or something I can sign up for? this is fantastic, and if the $$$ is not crazy, I would want this for sure.
Thank you for taking the time to write up the work properly. It's quite refreshing to read content like this after what's been the standard fare that's been posted here at HN.
Thank you very much for this comment. TBH when I made the post on r/linuxaudio, I figured there was a chance it would get drowned out by all the AI generated VST posts but fortunately that wasn’t the case
A problem with USB MIDI interfaces I've tried has been latency.
If I'm playing on a musical keyboard trying to add a new track while listening to existing ones, the tempo of the two doesn't quite line up and it's jarring. I've had to switch to an old, dedicated PCI audio card.
Would this resolve that?
Being able to use a spare Ethernet jack and throw away the card would be great.
And yes, I know there are newer USB adapters which are supposed to be able be able to help, like FPT (which I may have tried, can't recall), along with playing with ASIO settings. Since the problem was "solved", I never really got far down that rabbit hole. Also wary of getting into troubleshooting/isolating effects of my USB driver/software stacks (that I suspect introduced some unrelated known issues on certain older motherboards), the generation of the USB standard it's plugged into, extra levels of PCIe expanders gating the ports I'm using, contention from competing devices on the same bus, etc. Maybe these are groundless concerns but I figured there may be more variables involved vs. an Ethernet jack that has a comparatively direct hardwired connection to my bus (maybe through 1 expander). I can accept I might just be an old fart who needs to try again on more modern hardware/stacks.
How does the receive side recover the transmit side's sample clock? There is a BNC for clock sharing between "multiple units", but I'm not sure if that's used / required between transmit and receive.
Or is there no such synchronization, in which case there would be long-term drift?
I think I'm going to have to make a blog post addressing some of the clocking questions that come up. My brief answer for now is that there is only one clock, and that is the ETH-68 clock. The overall data flow is push based; the arrival of a new buffer of data at the Linux host IS the clocking event that schedules an audio graph evaluation. There is no need for another clock on the Linux host.
That is sort of true, although it should be possible to take advantage of the resampler in PipeWire or JACK (zalsa_in/zalsa_out) to effectively get multiple interfaces in the same graph. I would expect some hit to the audio latency when going through the resampler.
Unless audio is sampled at ETH-68 clock this won't work for long periods. Except if you are tuning a loop continuously doing continuous resampling using a Farrow filter for example.
Seem like a link to a Pipewire repo. Instead it is a Pipewire implementation of Milan-AVB.
Milan-AVB is a deterministic, standards-based media networking protocol built on top of Audio Video Bridging (AVB) and Time-Sensitive Networking (TSN) IEEE standards. Maintained by the Avnu Alliance, it is designed for the professional audio, video, and live event industries to guarantee plug-and-play interoperability across devices from different manufacturers
Far be it from me to discourage an H7 build, but I do question the codec choice. It's far from top of the line and it's not like the design is tight on space. TI offers much better (almost 20 dB SNR more on the ADC). Maybe a gen 2 could benefit from a better codec.
Yes this is an older codec but it has been reliably in production for around 20 years, and has a high channel count to cost ratio. The specifications are not cutting edge but I get comparable performance to my trusty Saffire Pro 40.
I have my eye on some other codecs for the next project.
I'm curious what you mean about the STM32H7. My experience has been that it is a challenging part, at least partly due to bugs in the HAL. This is especially true for the ethernet implementation. But once it is working the performance is quite good.
Oh just the I used it in uni, made a design with it (though never assembled it-parts are still in a bag), and I've used it professionally for some years. The HAL certainly has bugs but they were easy to work around for me. That said, I've never used the Ethernet peripheral on it. And everything else has been easy enough for me to go direct register access if needed.
It seems like, at 16000TbaseT anyways, you’re adding an overhead of about 150% on top of copper/fiber latency plus transmit time latency to process data at that bandwidth. I wonder if the same holds true at 100 vs 1000, 2500, 10000? Certainly this is a known tradeoff for DDR performance tuning — if you don’t mind spiking response times greatly, you can get the advertised maximum speeds, else you accept less bandwidth for somewhat less latency — and they’re both effectively using the same strategies to talk over copper.
> The typical default latency for a Dante audio device is 1 msec.
(emphasis mine)
The latency depends on the device. Hardware implementations of Dante commonly support latencies of 1ms or less, but software implementations are higher. The minimum latency of Dante Virtual Soundcard running on a PC is 4ms.
That's the appropriate number to compare against here (since the PC is using a software driver to interface with the network). However, that 4ms number is one-way latency, and the OP's 3.6ms number is round-trip. So this is already half the latency of DVS. (That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.)
> That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.
Hm, the audio latency doesn't fluctuate so it's not like the testing conditions affect the measurement. The audio latency is a fixed quantity that depends completely on the number of storage elements in the data path which isn't variable.
Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting. The histograms on my page (which should be animated BTW) show real time processing latency measurements while running the audio all the way through Bitwig with a moderate DSP load (multiple instances of Pianoteq, samplers, live MIDI input). On my system (details at the bottom of my page), I can do this at 48 kHz and 64 sample buffers with zero underruns. If I drop down to 32 samples per buffer, I do start getting underruns.
All I can do from the hardware side is try to minimize the processing latency of a typical cycle so that there is more head room to absorb jitter. The vast majority of the jitter comes from the Linux host. It's up to the end user to tune the system for low jitter. This is usually the case for audio on Linux, and the rabbit hole can go pretty deep on system tuning.
> Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting.
Right, that's what I meant. There will be some jitter depending on the scheduler, system load, network performance and traffic, and the quality of hardware/drivers; so more latency gives you headroom to absorb the jitter without underruns. The "ideal configurations" I was referring to were the low-traffic network and high-quality NIC. On a setup with more jitter, you might have to increase the buffer size (and thus latency) for reliable operation.
Mostly, I was just trying to contextualize the numbers for readers who aren't super familiar with low-latency audio networking. Sure, this project may not achieve the sub-1-ms roundtrip latencies that you can get with dedicated Dante hardware (like a Yamaha mixer and stagebox); but Dante can't do better than 8ms when one of the ends is a PC (although they were probably aiming for reliability on setups not aggressively tuned for minimum jitter.)
Gotcha thanks. Just to be clear, the "high-quality NIC" is $20 on Amazon. The crappy TP-Link switch I'm using is also about $25. Achieving the ideal configuration takes little effort and money.
1G would certainly decrease the processing latency (not the audio latency) by quite a bit. STM32H7 doesn't have a 1G MAC. 1G MAC is kind of rare on "friendly" microcontrollers, although there are at least two that I'm evaluating for the next project.
gigabit have throughput, but not latency. For small packets and on low load link there will be almost no difference (i had better source: https://serverfault.com/questions/276651/network-latency-100... , ethernet cnc controllers have same problems). It needs to go 10gbps or even 25gpbs to see lower latency, where electrical signal uses much higher electrical signaling bandwidth.
But the packets aren’t small when streaming audio. The latency improvement is substantial for 1000 byte packets as the accepted answer in your SO post demonstrates
Ah, then my blunder. I saw comments about 64/128 sample buffers automatically tough that most packets will be in 128/256 bytes range where difference not that big.
If there were open source Dante that actually worked it would be very cool. Don’t know how much usage it would get since the hardware will always lock it in. But for small studios and research it might be cool. Interesting project nevertheless.
There is AES67 support in PipeWire as of 1-2 years ago. I've tried it with some Dante hardware that I own, running the hardware in AES67 mode. It didn't work very well for me, even after lots of fiddling. If the experience would have been better, I probably would not have started working on ETH-68. It might be better now, I hope that it is.
ETH-68 doesn't have nearly as broad of scope as AES67. It is a much simpler system, so easier to realize for a 1-person team working on weekends.
Just a few other comments:
1. AES67 usually means Dante hardware which is notoriously expensive.
2. Dante devices running in AES67 mode often (always?) have their capabilities reduced. At least in some cases you are limited to 48 kHz, and I believe you are also limited on latency compensation settings. Someone with more experience could chime in.
It's a hardware audio interface with a very low latency (buffer of 64 samples, 3.6 ms), which supports both PipeWire and JACK over Ethernet, as shown on the pictures in the post. It's audio over CAT6, and a lot of it, 6 inputs + 8 outputs (TRS balanced) at 48 or 96 kHz.
Imagine having this on the stage right next to your analog gear, and a computer 50m away.
(Opening the page under discussion was actually helpful, it lists all this right at the top.)
Afaik all pro audio standards more or less require ptp for timing, and that might add requirements for nics and switches? otoh the i210 nic, that is mentioned in the post, does have full (g)ptp/phc/tsn/etc support
I currently have a setup with a Raspberry Pi (alpine+pipewire+dac) to stream audio over the network.
This enables me the watch video with no latency problems, because it's embedded in the audio stack (buffered). That's why I was wondering which problem gets solved with this project.
But now I understand that this is for concerts not for some audiophile multi room setup.
Pro audio systems frequently (and increasingly) use networked audio. Some obvious uses are distributing the sound from the instruments and people on a stage over to the front-of-house mix position that's usually somewhere mid-crowd, and also to the monitor mix position that's usually in a vaguely-quieter area off to the side of the stage, and to the broadcast truck.
The old tried-and-true method also still works: Analog splits. Take a bunch of audio sources (eg, microphones) and plug them into passive stage boxes that output over a thick-ass cable. Those thick-ass cables go to larger passive split boxes (often on wheels by this point), with two or more outputs for even-thicker cables, with one pair of wires for every individual signal -- often with individual shielding and jacketing.
Eventually, these splits can deliver audio to the different places that need it -- where it's ultimately broken back out into a bazillion individual cables that get plugged into things like mixers.
The cables can be very long (hundreds of meters) in length, and extremely heavy. They're expensive to produce, they're expensive to maintain, and they're expensive to wrangle. They often get transported in their own dedicated wooden trunks. But at least it's simple: A bunch of different audio devices scattered all over a venue, wired in parallel, listening to the signals that are directly produced by microphones on a stage.
---
But with networked audio, it can be more like this: A few boxes on a stage that accept analog audio on one side and emit network frames (often Ethernet or Ethernet-adjacent, and actually using IP isn't a rule at all) that contain digital audio on the other side. Those frames go to a network switch. One or more tiny-ass network cable comes out of the switch and goes wherever it needs to go, and switches can be cascaded, and more audio channels can be added downstream. Because Ethernet(ish) is a many-to-many network, it's bidirectional, too: Audio signals can go upstream just as easily as they go downstream.
It's tidy. It works. It's still expensive because the endpoints are expensive, but the cables themselves can be fairly inexpensive (think robustly-built Cat6 or fiber patch cords instead of giant cable trunks). If the venue's infrastructure goes to the right places and can be trusted, then it can also be used: Plug the stuff from the stage switch into a fiber patch panel on the building, and plug the broadcast truck outside into the same building, tie them together in some MDF or IDF somewhere, and send it. (And in a pure and just world where dedicated fiber links both exist and are easy: Patch another into the studio downtown. Or route it over an IP link that is shared with other purposes, if appropriate and also feeling brave.)
But with the tidiness comes complexity. Like... Latency is kind of a big deal here in ways that aren't a practical issue with analog audio. Putting too much delay between a vocalist and the monitors that they hear themselves with is actively deleterious of their ability to sing, for example.
And buffers are still required (they're ~always required when packet-switched network frames get converted to continuous analog signals). Keeping the buffers small requires very tight timing signals that get shared between all points. Pre-existing systems often achieve that with things like Precision Timing Protocol (though variations exist).
That all conspires to mean that the heavy lifting at the endpoints is often in the realm of FPGAs.
But, again: We get many channels over some bog-standard network cabling. Dante, for example, can be used to transport hundreds of 48KHz 24-bit audio channels on one gigabit ethernet link.
---
Anyway: This is a cheaper, smaller method. It uses an STM32H7 microcontroller to convert betwixt the network transport stuff and the DACs and ADCs of the analog world. It's designed to be used with the open-source Jack system that is commonly-used internally whenever Linux gets involved in recording or stage use, so it's simple to integrate with a Linux PC running software like Reaper. And at the end of the day, it's transportable over the Ethernet networks we all have.
And despite being built around an STM32, it achieves quite usable latency: The stated 3.620ms is about the same as a 1.2 meters of distance for sound in air.
Neat stuff. I'll probably never use it, but it's neat. :)
Distributing music != recording music != processing music.
There ARE good reasons for recording and processing music at higher bitrates and sample sizes. Effects, especially those with positive feedback loops, can go more unstable and clip and lose information with fewer bits and samples.
There's just no good reason to DISTRIBUTE music at the higher rates.
Don't such effects already upsample and downsample as needed internally ? You don't need to waste cpu on the effects that don't need higher sampling rate, right ?
They could. But then you get clipping effects and sampling effects when you convert back and forth.
Better to upsample once (generally a non-lossy operation) and operate on single precision FP (24-bit mantissa and an exponent) at higher sampling from that point forward and then downsample once (generally a LOSSY operation).
Stereo pair, pilot, RDS (the scrolling artist/song text), and the 67 kHz SCA. It's not quite 192kHz as I recall, but that's nearest fit really. AES192/192kHz MPX composite on the back end of every exciter/transmitter when I was last near them.
That's for the FM baseband signal, not the audio. Among other things you're doing here, you're basically treating a stereo signal (plus a pilot that effectively contains no information, plus an extremely low bitrate RDS stream in an extremely inefficient way) as one monaural signal.
But maybe there's a use case for replacing an AES192 signal carrying the full FM baseband signal over ETH-67? I'm not sure it's the correct fit, though.
That's how radio works now, it's owned by like 3 companies. Automation/playout for large groups of stations is all centralized, and they carry the baseband over IP to transmitter. Radio solved it awhile ago.
That's may be how the back-end of a modern broadcast FM radio transmitter works, but the good old stereo radio transmission itself still the same as it's been for many decades: It is analog, and mid-side encoded, and it remains completely compatible with monophonic receiver implementations.
Sorry, I was reading this late at night and must have misread the microcontroller. Then, the limitation lies in the codec and it would not be a minor modification or a matter of trading bit rate for latency.
I need to record ultrasound for work onboard construction vessels, currently we use either specialized equipment for PAM, which is limited in some aspects, or a USB sound card with a SBC and a hacky setup for sending PCM over TCP.
I would love this and am actively looking for a unit for my linux setup but: why the limit on sample rate? why not also support 44.1 kHz ?
How am I supposed to master for CD, which is still something people do?
Great to hear that you will support it in the future!
Sure, resampling works, but i only want to do it when it can't be avoided, not for basic playback.
If this had open source firmware, I think I know a lot of people who would be interested, but I can't seem to find out if that's true or not and where to buy one.
Thanks for your interest! I made a single reddit post about ETH-68 this week and this has been copied around various forums. My intent was to figure out if anyone thought this would be cool enough to produce. I'm trying to figure out if I should do a production run, open source it, or some combo of the two.
I think this would be a cool project to do as a DIY.
The market for a linux native Audio Interface is there I think.
At least I for one would appreciate this!
I'm currently running 4x Echo Audiofire 12s in my Linux based Studio and something akin to what you are doing would be a much more fun and less expensive option than RME. Next version is 24 channel?? ;)
Lol, this wasn't my choice really! The JACK server's default UDP listen port is 3000 with the netone backend, so that is default port that ETH-68 sends to. This can be adjusted on both ends
The choice of port wasn't the complaint so much as the choice of IP address. 12.12.12.10 is a real IP address on the Internet, so if you plug this device in somewhere with an Internet connection, it will begin spamming nuisance traffic to some random device out there. It would be better to choose an IP address in one of the private spaces for this purpose.
Nice device. Might be interesting if you thought outside the Jack. On MacOS there is BlackHole (open source) and Loopback (by one of my favorite companies, Rogue Amoeba). And creating a CoreAudio virtual audio driver on MacOS is easier than ever if you take a frontier LLM for a spin.
I have also thought some about whether or not I could make BlackHole work.
My takeaway so far is that macOS is more of a hassle for developing this stuff. It doesn't help that CoreAudio is not open source and the docs are hard to navigate. Hopefully someone will fix JACK-router on macOS one of these days.
They quote a roundtrip of 3.6ms, so one-way 1.8ms. In the plots on their page, it looks like the processing latency is centered on around 1ms:
> The LATMON pulse width is therefore an accurate measure of the total processing latency of each cycle and is affected by every element in the data path: processing delay in the microcontroller, network transmission delay, host OS delays, signal processing delay in the DAW, etc
You are confusing the audio latency with the processing latency. I see this mistake a lot.
One way audio latency is about 3.6 ms divided by two = 1.8 ms. This type of latency is audible.
Processing latency is just the amount of time it takes to complete all processing for each audio cycle. From the histograms, the typical processing latency is about 625 us and of course there is some jitter (almost all of the jitter comes from the Linux host BTW). Processing latency is not audible. However if the processing latency exceeds the deadline on a given audio cycle, there will be an underrun which will cause an audible glitch.
As the owner of the PCIe card that measurement was taken from, yes, I’d say it’s quite impressive. The average round-trip latency for a USB audio interface at 48 kHz/128 samples would usually fall somewhere around 8 ms, which is a bit much if you’re monitoring post-FX.
On Linux, we don’t have much in the way of Thunderbolt support for audio interfaces, so the only way to achieve this sort of latency has traditionally been with PCIe or PCI audio interfaces. Having a low-cost, infinitely more portable solution would be very welcome.
Hi, I'm Alex, I made ETH-68. I didn't create the post here on HN but I will answer some of the questions that have come up in the comments
First off, very cool project and great article showcasing it! Why does it also have MIDI?
Thank you! You’re right, I didn’t discuss the MIDI implementation in the blog post. Right now it’s a little primitive but functional. Each MIDI message (e.g 3 byte note-on message) maps to 1 UDP packet. I have a simple script on the host side that interfaces these UDP packets with an ALSA loopback MIDI interface.
As someone who uses a USB equivalent to this piece of hardware (in my case, a Behringer UMC1820): having MIDI and audio inputs on the same interface is a nice convenience feature. Avoids needing to use up multiple ports on the host machine, and avoids a lot of the time synchronization issues that arise with trying to use multiple interfaces at once.
Probably for the same reason MADI also has MIDI, for sending control messages.
Might you do clock synchronization via PTP in the future?
Possibly. The STM32H7 Ethernet MAC does have IEEE 1588 support, although I've heard that it is not straightforward to use.
Unrelated question: what is the keyboard on the last picture of the post?
Keychron, Q1 I think?
Is it actually for sale or otherwise obtainable?
Not yet but actively figuring out what to do about that!
Why not XLR?
XLR is generally for microphones, TRS is typically used for line level and instruments (hi-z)
They're basically interchangeable. But yeah TRS is more compact and more normative when there is no preamp. Some monitors I've seen have xlr inputs though, even though they expect line level input.
Another option is to just use XLR/TRS combo jacks if you want to support both, but that’s not really any different than the user using an XLR to TRS cord or converter
XLR and combo XLR/TRS tend to be significantly pricier, which matters since there are 14x in the BOM.
FYI your domain is marked "newly seen" on cloudflare, which results in it being blocked by Ubiquiti Unifi domain filtering.
You can submit to have it re-categorized using the link below.
https://radar.cloudflare.com/domains/feedback/naturalsystems...
That's an user problem, isn't it? Why a niche site owner consider caring about proprietary Ubiquiti systems?
Well, it's certainly a dumb situation that I don't like that's created by my employer (and many others with similar setups), but it takes only two clicks to pick a category to make it available to folks behind these firewalls.
I went ahead and chose "blog" because it sounds like it's a blog. Apparently the "newly seen domains" category goes away on its own after 30 days, so that would have fixed itself, though then it would have been "uncategorized" and still blocked.
If only one of us had access to better ways to address this problem, oh well
> That's an user problem, isn't it?
Yes. And no. The affected users usually have no control over the situation, so if the site owner cares at all that those people can't access it, then it is effectively their problem as they can potentially do something about it.
If neither party cares enough, then it is nobody's problem.
The affected users can stop using brain damaged systems.
P.S. By the way, to paraphrase an OS developer, Cloudflare, fuck you for "Sorry, you have been blocked. You are unable to access thunderbird.net".
That's pretty much it. I wouldn't have mentioned it if not for the fact that it's easy to fix and for users it just looks like the site is down, which is not uncommon for small projects posted to HN.
I guess some people thought I liked cloudflare or something... that's definitely not the case.
"easy to fix" is relative. This "new domain equals bad" thinking has spread in the IT security world. My corporate anti-virus blocks them too and I see it a lot while browsing HN on the job.
The only real solution, as a "holder of a young domain" is to wait 30 days before you use it.
Is this open source? I’ve been working on https://codeberg.org/olpad/openmic which is in the more traditional domain of USB audio interfaces, but I’d be interested in taking a look at the internals of the ETH-68
Nice project. Have you tried measuring the round trip audio latency on Linux?
I’m still working towards a first prototype, development has been slow due to my own schedule. But I’d expect round trip latency near what you’d expect from a USB interface. It would be cool to add ethernet support in the future though!
What's your use case where you want more outputs than inputs? At home I run a Audiofuse 16Rig with a lot of ADAT expansion to get 32i/32o, but I only use like 6o, while every input is used.
The project just exposes everything the hardware can do, the audio codec chip used for the conversion is Burr-Brown PCM3168A 24-Bit, 96-kHz/192-kHz, 6-In/8-Out.
Some I can imagine are surround sound, multiple monitor pairs (including dedicated monitoring channels for different musicians or simply different outputs for mixing).
hello Alex,
tech nerd/musician here, I would be very interested in this! do you have an email list or something I can sign up for? this is fantastic, and if the $$$ is not crazy, I would want this for sure.
Thank you for taking the time to write up the work properly. It's quite refreshing to read content like this after what's been the standard fare that's been posted here at HN.
Thank you very much for this comment. TBH when I made the post on r/linuxaudio, I figured there was a chance it would get drowned out by all the AI generated VST posts but fortunately that wasn’t the case
A problem with USB MIDI interfaces I've tried has been latency.
If I'm playing on a musical keyboard trying to add a new track while listening to existing ones, the tempo of the two doesn't quite line up and it's jarring. I've had to switch to an old, dedicated PCI audio card.
Would this resolve that?
Being able to use a spare Ethernet jack and throw away the card would be great.
And yes, I know there are newer USB adapters which are supposed to be able be able to help, like FPT (which I may have tried, can't recall), along with playing with ASIO settings. Since the problem was "solved", I never really got far down that rabbit hole. Also wary of getting into troubleshooting/isolating effects of my USB driver/software stacks (that I suspect introduced some unrelated known issues on certain older motherboards), the generation of the USB standard it's plugged into, extra levels of PCIe expanders gating the ports I'm using, contention from competing devices on the same bus, etc. Maybe these are groundless concerns but I figured there may be more variables involved vs. an Ethernet jack that has a comparatively direct hardwired connection to my bus (maybe through 1 expander). I can accept I might just be an old fart who needs to try again on more modern hardware/stacks.
How does the receive side recover the transmit side's sample clock? There is a BNC for clock sharing between "multiple units", but I'm not sure if that's used / required between transmit and receive.
Or is there no such synchronization, in which case there would be long-term drift?
I think I'm going to have to make a blog post addressing some of the clocking questions that come up. My brief answer for now is that there is only one clock, and that is the ETH-68 clock. The overall data flow is push based; the arrival of a new buffer of data at the Linux host IS the clocking event that schedules an audio graph evaluation. There is no need for another clock on the Linux host.
> There is no need for another clock on the Linux host
surely that works only as long as eth68 is the only interface? but often you might want other interfaces too which inherently have their own clocks
That is sort of true, although it should be possible to take advantage of the resampler in PipeWire or JACK (zalsa_in/zalsa_out) to effectively get multiple interfaces in the same graph. I would expect some hit to the audio latency when going through the resampler.
Unless audio is sampled at ETH-68 clock this won't work for long periods. Except if you are tuning a loop continuously doing continuous resampling using a Farrow filter for example.
In my comment that you replied to, I outlined a scenario where the only clock is the ETH-68 clock, so yes audio is sampled at the ETH-68 clock.
See my other comment in this thread about resampling with PipeWire or JACK.
Looks neat. I'm curious, does using ethernet offer anything over USB in terms of obtaining low latency?
May also be of interest: https://github.com/kebag-logic/pipewire
Seem like a link to a Pipewire repo. Instead it is a Pipewire implementation of Milan-AVB.
Milan-AVB is a deterministic, standards-based media networking protocol built on top of Audio Video Bridging (AVB) and Time-Sensitive Networking (TSN) IEEE standards. Maintained by the Avnu Alliance, it is designed for the professional audio, video, and live event industries to guarantee plug-and-play interoperability across devices from different manufacturers
If ETH-68 support Milan-AVB that would indeed be cool (especially for users on other platfroms).
Far be it from me to discourage an H7 build, but I do question the codec choice. It's far from top of the line and it's not like the design is tight on space. TI offers much better (almost 20 dB SNR more on the ADC). Maybe a gen 2 could benefit from a better codec.
Yes this is an older codec but it has been reliably in production for around 20 years, and has a high channel count to cost ratio. The specifications are not cutting edge but I get comparable performance to my trusty Saffire Pro 40.
I have my eye on some other codecs for the next project.
I'm curious what you mean about the STM32H7. My experience has been that it is a challenging part, at least partly due to bugs in the HAL. This is especially true for the ethernet implementation. But once it is working the performance is quite good.
Oh just the I used it in uni, made a design with it (though never assembled it-parts are still in a bag), and I've used it professionally for some years. The HAL certainly has bugs but they were easy to work around for me. That said, I've never used the Ethernet peripheral on it. And everything else has been easy enough for me to go direct register access if needed.
After the whole drama about TI opamp, i would never trust any components from them.
Edit: For the context: https://www.eevblog.com/forum/chat/ti-ne5532-audio-opamp-cha...
I’m aware they botched the 5532. But it’s gonna be a hard life boycotting all of TIs components
Jfc. Never thought TI would do this, this is insane.
They really found a way to enshittify an OpAmp?
Really nice!!!
I wonder if gigabit would make a difference on latency, faster packet transmissions. Could packets drop from 64 to 32b?
4ms is pretty good but I feel like sub 2ms would be nicer.
https://semiengineering.com/latency-considerations-for-1-6t-...
It seems like, at 16000TbaseT anyways, you’re adding an overhead of about 150% on top of copper/fiber latency plus transmit time latency to process data at that bandwidth. I wonder if the same holds true at 100 vs 1000, 2500, 10000? Certainly this is a known tradeoff for DDR performance tuning — if you don’t mind spiking response times greatly, you can get the advertised maximum speeds, else you accept less bandwidth for somewhat less latency — and they’re both effectively using the same strategies to talk over copper.
Dante’s default latency compensation is 1ms: https://dev.audinate.com/GA/dante-controller/userguide/webhe...
> The typical default latency for a Dante audio device is 1 msec.
(emphasis mine)
The latency depends on the device. Hardware implementations of Dante commonly support latencies of 1ms or less, but software implementations are higher. The minimum latency of Dante Virtual Soundcard running on a PC is 4ms.
That's the appropriate number to compare against here (since the PC is using a software driver to interface with the network). However, that 4ms number is one-way latency, and the OP's 3.6ms number is round-trip. So this is already half the latency of DVS. (That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.)
Yep, thanks for clarifying about this!
> That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.
Hm, the audio latency doesn't fluctuate so it's not like the testing conditions affect the measurement. The audio latency is a fixed quantity that depends completely on the number of storage elements in the data path which isn't variable.
Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting. The histograms on my page (which should be animated BTW) show real time processing latency measurements while running the audio all the way through Bitwig with a moderate DSP load (multiple instances of Pianoteq, samplers, live MIDI input). On my system (details at the bottom of my page), I can do this at 48 kHz and 64 sample buffers with zero underruns. If I drop down to 32 samples per buffer, I do start getting underruns.
All I can do from the hardware side is try to minimize the processing latency of a typical cycle so that there is more head room to absorb jitter. The vast majority of the jitter comes from the Linux host. It's up to the end user to tune the system for low jitter. This is usually the case for audio on Linux, and the rabbit hole can go pretty deep on system tuning.
> Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting.
Right, that's what I meant. There will be some jitter depending on the scheduler, system load, network performance and traffic, and the quality of hardware/drivers; so more latency gives you headroom to absorb the jitter without underruns. The "ideal configurations" I was referring to were the low-traffic network and high-quality NIC. On a setup with more jitter, you might have to increase the buffer size (and thus latency) for reliable operation.
Mostly, I was just trying to contextualize the numbers for readers who aren't super familiar with low-latency audio networking. Sure, this project may not achieve the sub-1-ms roundtrip latencies that you can get with dedicated Dante hardware (like a Yamaha mixer and stagebox); but Dante can't do better than 8ms when one of the ends is a PC (although they were probably aiming for reliability on setups not aggressively tuned for minimum jitter.)
Gotcha thanks. Just to be clear, the "high-quality NIC" is $20 on Amazon. The crappy TP-Link switch I'm using is also about $25. Achieving the ideal configuration takes little effort and money.
Brooklyn, IP Core can do 125us or 250us depending on number of switch hops.
Sadly the H7 doesn't have a gigabit MAC. And most likely doing one over USB high speed host wouldn't help? But it would be an interesting experiment
One can fit a nice (eg, not 1-channel, but actually useable) LTE base station frontend in a gigabit eth on 2014-s tech.
1G would certainly decrease the processing latency (not the audio latency) by quite a bit. STM32H7 doesn't have a 1G MAC. 1G MAC is kind of rare on "friendly" microcontrollers, although there are at least two that I'm evaluating for the next project.
gigabit have throughput, but not latency. For small packets and on low load link there will be almost no difference (i had better source: https://serverfault.com/questions/276651/network-latency-100... , ethernet cnc controllers have same problems). It needs to go 10gbps or even 25gpbs to see lower latency, where electrical signal uses much higher electrical signaling bandwidth.
But the packets aren’t small when streaming audio. The latency improvement is substantial for 1000 byte packets as the accepted answer in your SO post demonstrates
Ah, then my blunder. I saw comments about 64/128 sample buffers automatically tough that most packets will be in 128/256 bytes range where difference not that big.
If there were open source Dante that actually worked it would be very cool. Don’t know how much usage it would get since the hardware will always lock it in. But for small studios and research it might be cool. Interesting project nevertheless.
There's also AVB but it also requires some specific functionality from network switches...
And there’s AES67 which I think PipeWire does support.
I got AES67 running and released on a cheap ESP32-P4 MCU via Ethernet.
There is AES67 support in PipeWire as of 1-2 years ago. I've tried it with some Dante hardware that I own, running the hardware in AES67 mode. It didn't work very well for me, even after lots of fiddling. If the experience would have been better, I probably would not have started working on ETH-68. It might be better now, I hope that it is.
ETH-68 doesn't have nearly as broad of scope as AES67. It is a much simpler system, so easier to realize for a 1-person team working on weekends.
Just a few other comments: 1. AES67 usually means Dante hardware which is notoriously expensive. 2. Dante devices running in AES67 mode often (always?) have their capabilities reduced. At least in some cases you are limited to 48 kHz, and I believe you are also limited on latency compensation settings. Someone with more experience could chime in.
Very cool. However, I think I'm missing the usecase. It needs a separate LAN, so basically it's audio over CAT6?
Don't get me wrong, I'm genuinely trying to understand.
How does this compare to pipewire over ethernet? Is it realtime vs buffered?
It's a hardware audio interface with a very low latency (buffer of 64 samples, 3.6 ms), which supports both PipeWire and JACK over Ethernet, as shown on the pictures in the post. It's audio over CAT6, and a lot of it, 6 inputs + 8 outputs (TRS balanced) at 48 or 96 kHz.
Imagine having this on the stage right next to your analog gear, and a computer 50m away.
(Opening the page under discussion was actually helpful, it lists all this right at the top.)
Yeah I've been waiting until traditional usb audio interfaces show up that use the USB4 pcie support for ultra low latency.
But this might be even more effective.
Besides audio dsp for live use, there is also the use case of visualizers.
How does it differ to standards based systems like
https://github.com/bondagit/aes67-linux-daemon
Afaik all pro audio standards more or less require ptp for timing, and that might add requirements for nics and switches? otoh the i210 nic, that is mentioned in the post, does have full (g)ptp/phc/tsn/etc support
If you want low and deterministic latency between different devices then you'll need ptp
ptp support in nics is becoming more widespread, nics not yet as far as I can see.
It doesn’t need to be a separate LAN. But using a dedicated LAN helps to keep extraneous traffic from delaying time sensitive audio packets.
ETH-68 isn’t running PipeWire, or Linux. Maybe I don’t understand your question.
Not sure what you mean by realtime vs buffered. Care to elaborate?
I currently have a setup with a Raspberry Pi (alpine+pipewire+dac) to stream audio over the network.
This enables me the watch video with no latency problems, because it's embedded in the audio stack (buffered). That's why I was wondering which problem gets solved with this project.
But now I understand that this is for concerts not for some audiophile multi room setup.
Some background on the utility:
Pro audio systems frequently (and increasingly) use networked audio. Some obvious uses are distributing the sound from the instruments and people on a stage over to the front-of-house mix position that's usually somewhere mid-crowd, and also to the monitor mix position that's usually in a vaguely-quieter area off to the side of the stage, and to the broadcast truck.
The old tried-and-true method also still works: Analog splits. Take a bunch of audio sources (eg, microphones) and plug them into passive stage boxes that output over a thick-ass cable. Those thick-ass cables go to larger passive split boxes (often on wheels by this point), with two or more outputs for even-thicker cables, with one pair of wires for every individual signal -- often with individual shielding and jacketing.
Eventually, these splits can deliver audio to the different places that need it -- where it's ultimately broken back out into a bazillion individual cables that get plugged into things like mixers.
The cables can be very long (hundreds of meters) in length, and extremely heavy. They're expensive to produce, they're expensive to maintain, and they're expensive to wrangle. They often get transported in their own dedicated wooden trunks. But at least it's simple: A bunch of different audio devices scattered all over a venue, wired in parallel, listening to the signals that are directly produced by microphones on a stage.
---
But with networked audio, it can be more like this: A few boxes on a stage that accept analog audio on one side and emit network frames (often Ethernet or Ethernet-adjacent, and actually using IP isn't a rule at all) that contain digital audio on the other side. Those frames go to a network switch. One or more tiny-ass network cable comes out of the switch and goes wherever it needs to go, and switches can be cascaded, and more audio channels can be added downstream. Because Ethernet(ish) is a many-to-many network, it's bidirectional, too: Audio signals can go upstream just as easily as they go downstream.
It's tidy. It works. It's still expensive because the endpoints are expensive, but the cables themselves can be fairly inexpensive (think robustly-built Cat6 or fiber patch cords instead of giant cable trunks). If the venue's infrastructure goes to the right places and can be trusted, then it can also be used: Plug the stuff from the stage switch into a fiber patch panel on the building, and plug the broadcast truck outside into the same building, tie them together in some MDF or IDF somewhere, and send it. (And in a pure and just world where dedicated fiber links both exist and are easy: Patch another into the studio downtown. Or route it over an IP link that is shared with other purposes, if appropriate and also feeling brave.)
But with the tidiness comes complexity. Like... Latency is kind of a big deal here in ways that aren't a practical issue with analog audio. Putting too much delay between a vocalist and the monitors that they hear themselves with is actively deleterious of their ability to sing, for example.
And buffers are still required (they're ~always required when packet-switched network frames get converted to continuous analog signals). Keeping the buffers small requires very tight timing signals that get shared between all points. Pre-existing systems often achieve that with things like Precision Timing Protocol (though variations exist).
That all conspires to mean that the heavy lifting at the endpoints is often in the realm of FPGAs.
But, again: We get many channels over some bog-standard network cabling. Dante, for example, can be used to transport hundreds of 48KHz 24-bit audio channels on one gigabit ethernet link.
---
Anyway: This is a cheaper, smaller method. It uses an STM32H7 microcontroller to convert betwixt the network transport stuff and the DACs and ADCs of the analog world. It's designed to be used with the open-source Jack system that is commonly-used internally whenever Linux gets involved in recording or stage use, so it's simple to integrate with a Linux PC running software like Reaper. And at the end of the day, it's transportable over the Ethernet networks we all have.
And despite being built around an STM32, it achieves quite usable latency: The stated 3.620ms is about the same as a 1.2 meters of distance for sound in air.
Neat stuff. I'll probably never use it, but it's neat. :)
How difficult would it be to extend it to 192kHz, or even 384kHz? Is it limited by the ESP32 hardware?
For what application do you need 96kHz or 192kHz of bandwidth?
Don't forget about cats and bats, they like music too
See also: https://news.ycombinator.com/item?id=3668310
Distributing music != recording music != processing music.
There ARE good reasons for recording and processing music at higher bitrates and sample sizes. Effects, especially those with positive feedback loops, can go more unstable and clip and lose information with fewer bits and samples.
There's just no good reason to DISTRIBUTE music at the higher rates.
Don't such effects already upsample and downsample as needed internally ? You don't need to waste cpu on the effects that don't need higher sampling rate, right ?
They could. But then you get clipping effects and sampling effects when you convert back and forth.
Better to upsample once (generally a non-lossy operation) and operate on single precision FP (24-bit mantissa and an exponent) at higher sampling from that point forward and then downsample once (generally a LOSSY operation).
Good old FM radio (in stereo) is a 192kHz mux, annoyingly.
How so?
Stereo pair, pilot, RDS (the scrolling artist/song text), and the 67 kHz SCA. It's not quite 192kHz as I recall, but that's nearest fit really. AES192/192kHz MPX composite on the back end of every exciter/transmitter when I was last near them.
That's for the FM baseband signal, not the audio. Among other things you're doing here, you're basically treating a stereo signal (plus a pilot that effectively contains no information, plus an extremely low bitrate RDS stream in an extremely inefficient way) as one monaural signal.
But maybe there's a use case for replacing an AES192 signal carrying the full FM baseband signal over ETH-67? I'm not sure it's the correct fit, though.
That's how radio works now, it's owned by like 3 companies. Automation/playout for large groups of stations is all centralized, and they carry the baseband over IP to transmitter. Radio solved it awhile ago.
https://www.telosalliance.com/radio-processing/audio-interfa... etc.
That's may be how the back-end of a modern broadcast FM radio transmitter works, but the good old stereo radio transmission itself still the same as it's been for many decades: It is analog, and mid-side encoded, and it remains completely compatible with monophonic receiver implementations.
Tuning into time radio signals like DCF77?
Ultrasonic recordings
https://www.youtube.com/watch?v=hCQCP-5g5bo
After ridiculing 96 kHz sampling for years, I finally found a use for it.
A niche use, mind you.
Basically, I found myself doing real-to-complex baseband conversion for audio signals, and doing so efficiently halved my Nyquist rate.
Increasing my sample rate to 96 kHz let me construct the 48 kHz analytic signal that I wanted.
Hoisted on my own petard :)
Recording music?
Ultrasound. I need to record and process signals up to 80kHz, so I need 192kHz sampling. No way around it.
With stage audio equipment?
What ESP32?
I'm using an STM32H7, not ESP32. The DAC on the PCM3168A can go up to 192 kHz, but the ADC can only be clocked up to 96 kHz.
Sorry, I was reading this late at night and must have misread the microcontroller. Then, the limitation lies in the codec and it would not be a minor modification or a matter of trading bit rate for latency.
I need to record ultrasound for work onboard construction vessels, currently we use either specialized equipment for PAM, which is limited in some aspects, or a USB sound card with a SBC and a hacky setup for sending PCM over TCP.
I would love this and am actively looking for a unit for my linux setup but: why the limit on sample rate? why not also support 44.1 kHz ? How am I supposed to master for CD, which is still something people do?
The next Rev will have an additional oscillator to support 44.1/88.2.
Just curious though, since I don’t work at these sample rates. Is resampling not good enough? Or maybe I don’t understand the mastering flow.
Great to hear that you will support it in the future! Sure, resampling works, but i only want to do it when it can't be avoided, not for basic playback.
> The next Rev will have an additional oscillator to support 44.1/88.2.
Are you sure that's necessary? The STM32H7 has a decent fractional PLL.
If this had open source firmware, I think I know a lot of people who would be interested, but I can't seem to find out if that's true or not and where to buy one.
how do I buy this?
Or even build it. Strange there is zero info...
Thanks for your interest! I made a single reddit post about ETH-68 this week and this has been copied around various forums. My intent was to figure out if anyone thought this would be cool enough to produce. I'm trying to figure out if I should do a production run, open source it, or some combo of the two.
I think this would be a cool project to do as a DIY. The market for a linux native Audio Interface is there I think. At least I for one would appreciate this!
I'd definitely be interested. Maybe put an interest form on the website? Or a link to one.
How much do you think it'd cost if you were to sell and ship to the US?
I'm currently running 4x Echo Audiofire 12s in my Linux based Studio and something akin to what you are doing would be a much more fun and less expensive option than RME. Next version is 24 channel?? ;)
> eth68 sends capture packets to 12.12.12.10:3000 by default.
Um.
At least it’s AT&T, not Huawei :)
Riiip…. I hate private projects pooping on public spaces.
Lol, this wasn't my choice really! The JACK server's default UDP listen port is 3000 with the netone backend, so that is default port that ETH-68 sends to. This can be adjusted on both ends
The choice of port wasn't the complaint so much as the choice of IP address. 12.12.12.10 is a real IP address on the Internet, so if you plug this device in somewhere with an Internet connection, it will begin spamming nuisance traffic to some random device out there. It would be better to choose an IP address in one of the private spaces for this purpose.
See, for example, Cloudflare's stories of issues caused by misconfigured systems appropriating the "1.1.1.1" address for their own purposes: https://blog.cloudflare.com/fixing-reachability-to-1-1-1-1-g...
Oh thanks for that feedback, easy change.
each user of ULAs has a unique address range, whereas IPv4 private addressing is common to many users
https://en.wikipedia.org/wiki/Unique_local_address
Nice device. Might be interesting if you thought outside the Jack. On MacOS there is BlackHole (open source) and Loopback (by one of my favorite companies, Rogue Amoeba). And creating a CoreAudio virtual audio driver on MacOS is easier than ever if you take a frontier LLM for a spin.
I have done a small amount of work on macOS but it is not a priority. I did some preliminary work with this nice library:
https://github.com/gavv/libASPL
I have also thought some about whether or not I could make BlackHole work.
My takeaway so far is that macOS is more of a hassle for developing this stuff. It doesn't help that CoreAudio is not open source and the docs are hard to navigate. Hopefully someone will fix JACK-router on macOS one of these days.
Is the hardware and software open source? Is there a link to a repo that I have missed?
> Very low latency: 3.620 milliseconds round trip at 48 kHz with 64 sample buffer
"very low latency" in audio is <=1ms. 3.6ms is good but not special.
They quote a roundtrip of 3.6ms, so one-way 1.8ms. In the plots on their page, it looks like the processing latency is centered on around 1ms:
> The LATMON pulse width is therefore an accurate measure of the total processing latency of each cycle and is affected by every element in the data path: processing delay in the microcontroller, network transmission delay, host OS delays, signal processing delay in the DAW, etc
You are confusing the audio latency with the processing latency. I see this mistake a lot.
One way audio latency is about 3.6 ms divided by two = 1.8 ms. This type of latency is audible.
Processing latency is just the amount of time it takes to complete all processing for each audio cycle. From the histograms, the typical processing latency is about 625 us and of course there is some jitter (almost all of the jitter comes from the Linux host BTW). Processing latency is not audible. However if the processing latency exceeds the deadline on a given audio cycle, there will be an underrun which will cause an audible glitch.
But it’s matching AND exceeding the performance of a $999 PCIe card designed for a similar purpose.
I would love to know why this is hand-wavy and not very cool? Even as not-audiophile, I’ve got ideas of things to use this for.
> RME HDSPe AIO Pro PCIe
> eth68 matches the latency performance of the RME card at 48 kHz and surpasses it by 0.33 ms at 96 kHz.
As the owner of the PCIe card that measurement was taken from, yes, I’d say it’s quite impressive. The average round-trip latency for a USB audio interface at 48 kHz/128 samples would usually fall somewhere around 8 ms, which is a bit much if you’re monitoring post-FX.
On Linux, we don’t have much in the way of Thunderbolt support for audio interfaces, so the only way to achieve this sort of latency has traditionally been with PCIe or PCI audio interfaces. Having a low-cost, infinitely more portable solution would be very welcome.
Awesome. Love to see it!
This is just the answer I was looking for. I definitely don’t know anything about the audio space, but am into networking hardcode.
I love seeing audiophile projects that don’t involve a rebranded tp-link switch and marked up 1000%
Round trip latency for RME devices (widely considered the best in the industry) is around 3ms, so this is indeed "very low latency".
You might be thinking of latency of the converters, which is normally sub-ms.
Check out the round trip latency measurements for audio interfaces on Linux here:
https://interfacinglinux.com/linux-compatible-audio-interfac...
AFAIK the best one is RME AIO Pro PCIe card. ETH-68 matches the round trip audio latency of this card at 48 kHz and surpasses it be 0.33 ms at 96 kHz.