Rendered at 05:49:41 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
landgenoot 1 hours ago [-]
Very cool. However, I think I'm missing the usecase.
It needs a separate LAN, so basically it's audio over CAT6?
Don't get me wrong, I'm genuinely trying to understand.
How does this compare to pipewire over ethernet? Is it realtime vs buffered?
nine_k 51 minutes ago [-]
It's a hardware audio interface with a very low latency (buffer of 64 samples, 3.6 ms), which supports both PipeWire and JACK over Ethernet, as shown on the pictures in the post. It's audio over CAT6, and a lot of it, 6 inputs + 8 outputs (TRS balanced) at 48 or 96 kHz.
Imagine having this on the stage right next to your analog gear, and a computer 50m away.
(Opening the page under discussion was actually helpful, it lists all this right at the top.)
NewJazz 42 minutes ago [-]
Yeah I've been waiting until traditional usb audio interfaces show up that use the USB4 pcie support for ultra low latency.
But this might be even more effective.
Besides audio dsp for live use, there is also the use case of visualizers.
alowell 48 minutes ago [-]
It doesn’t need to be a separate LAN. But using a dedicated LAN helps to keep extraneous traffic from delaying time sensitive audio packets.
ETH-68 isn’t running PipeWire, or Linux. Maybe I don’t understand your question.
Not sure what you mean by realtime vs buffered. Care to elaborate?
alowell 7 hours ago [-]
Hi, I'm Alex, I made ETH-68. I didn't create the post here on HN but I will answer some of the questions that have come up in the comments
NewJazz 26 minutes ago [-]
Is it actually for sale or otherwise obtainable?
2 hours ago [-]
hommelix 2 hours ago [-]
Unrelated question: what is the keyboard on the last picture of the post?
alowell 54 minutes ago [-]
Keychron, Q1 I think?
zbrozek 3 hours ago [-]
Might you do clock synchronization via PTP in the future?
iFreilicht 7 hours ago [-]
First off, very cool project and great article showcasing it! Why does it also have MIDI?
yellowapple 3 hours ago [-]
As someone who uses a USB equivalent to this piece of hardware (in my case, a Behringer UMC1820): having MIDI and audio inputs on the same interface is a nice convenience feature. Avoids needing to use up multiple ports on the host machine, and avoids a lot of the time synchronization issues that arise with trying to use multiple interfaces at once.
alowell 6 hours ago [-]
Thank you! You’re right, I didn’t discuss the MIDI implementation in the blog post. Right now it’s a little primitive but functional. Each MIDI message (e.g 3 byte note-on message) maps to 1 UDP packet. I have a simple script on the host side that interfaces these UDP packets with an ALSA loopback MIDI interface.
jhallenworld 7 hours ago [-]
How does the receive side recover the transmit side's sample clock? There is a BNC for clock sharing between "multiple units", but I'm not sure if that's used / required between transmit and receive.
Or is there no such synchronization, in which case there would be long-term drift?
alowell 7 hours ago [-]
I think I'm going to have to make a blog post addressing some of the clocking questions that come up. My brief answer for now is that there is only one clock, and that is the ETH-68 clock. The overall data flow is push based; the arrival of a new buffer of data at the Linux host IS the clocking event that schedules an audio graph evaluation. There is no need for another clock on the Linux host.
zokier 6 hours ago [-]
> There is no need for another clock on the Linux host
surely that works only as long as eth68 is the only interface? but often you might want other interfaces too which inherently have their own clocks
alowell 6 hours ago [-]
That is sort of true, although it should be possible to take advantage of the resampler in PipeWire or JACK (zalsa_in/zalsa_out) to effectively get multiple interfaces in the same graph. I would expect some hit to the audio latency when going through the resampler.
Far be it from me to discourage an H7 build, but I do question the codec choice. It's far from top of the line and it's not like the design is tight on space. TI offers much better (almost 20 dB SNR more on the ADC). Maybe a gen 2 could benefit from a better codec.
mrheosuper 3 hours ago [-]
After the whole drama about TI opamp, i would never trust any components from them.
I’m aware they botched the 5532. But it’s gonna be a hard life boycotting all of TIs components
MrBuddyCasino 15 minutes ago [-]
Jfc. Never thought TI would do this, this is insane.
alowell 7 hours ago [-]
Yes this is an older codec but it has been reliably in production for around 20 years, and has a high channel count to cost ratio. The specifications are not cutting edge but I get comparable performance to my trusty Saffire Pro 40.
I have my eye on some other codecs for the next project.
alowell 5 hours ago [-]
I'm curious what you mean about the STM32H7. My experience has been that it is a challenging part, at least partly due to bugs in the HAL. This is especially true for the ethernet implementation. But once it is working the performance is quite good.
inatreecrown2 6 hours ago [-]
I would love this and am actively looking for a unit for my linux setup but: why the limit on sample rate? why not also support 44.1 kHz ?
How am I supposed to master for CD, which is still something people do?
alowell 6 hours ago [-]
The next Rev will have an additional oscillator to support 44.1/88.2.
Just curious though, since I don’t work at these sample rates. Is resampling not good enough? Or maybe I don’t understand the mastering flow.
inatreecrown2 6 hours ago [-]
Great to hear that you will support it in the future!
Sure, resampling works, but i only want to do it when it can't be avoided, not for basic playback.
duskwuff 5 hours ago [-]
> The next Rev will have an additional oscillator to support 44.1/88.2.
Are you sure that's necessary? The STM32H7 has a decent fractional PLL.
dasv 9 hours ago [-]
How difficult would it be to extend it to 192kHz, or even 384kHz? Is it limited by the ESP32 hardware?
hackingonempty 8 hours ago [-]
For what application do you need 96kHz or 192kHz of bandwidth?
an_d_rew 14 minutes ago [-]
After ridiculing 96 kHz sampling for years, I finally found a use for it.
A niche use, mind you.
Basically, I found myself doing real-to-complex baseband conversion for audio signals, and doing so efficiently halved my Nyquist rate.
Increasing my sample rate to 96 kHz let me construct the 48 kHz analytic signal that I wanted.
Don't forget about cats and bats, they like music too
tacos8me 8 hours ago [-]
Good old FM radio (in stereo) is a 192kHz mux, annoyingly.
ssl-3 43 minutes ago [-]
That's may be how the back-end of a modern broadcast FM radio transmitter works, but the good old stereo radio transmission itself still the same as it's been for many decades: It is analog, and mid-side encoded, and it remains completely compatible with monophonic receiver implementations.
anyfoo 8 hours ago [-]
How so?
tacos8me 7 hours ago [-]
Stereo pair, pilot, RDS (the scrolling artist/song text), and the 67 kHz SCA. It's not quite 192kHz as I recall, but that's nearest fit really. AES192/192kHz MPX composite on the back end of every exciter/transmitter when I was last near them.
anyfoo 7 hours ago [-]
That's for the FM baseband signal, not the audio. Among other things you're doing here, you're basically treating a stereo signal (plus a pilot that effectively contains no information, plus an extremely low bitrate RDS stream in an extremely inefficient way) as one monaural signal.
But maybe there's a use case for replacing an AES192 signal carrying the full FM baseband signal over ETH-67? I'm not sure it's the correct fit, though.
tacos8me 7 hours ago [-]
That's how radio works now, it's owned by like 3 companies. Automation/playout for large groups of stations is all centralized, and they carry the baseband over IP to transmitter. Radio solved it awhile ago.
Distributing music != recording music != processing music.
There ARE good reasons for recording and processing music at higher bitrates and sample sizes. Effects, especially those with positive feedback loops, can go more unstable and clip and lose information with fewer bits and samples.
There's just no good reason to DISTRIBUTE music at the higher rates.
miguelnegrao 6 hours ago [-]
Don't such effects already upsample and downsample as needed internally ? You don't need to waste cpu on the effects that don't need higher sampling rate, right ?
bsder 5 hours ago [-]
They could. But then you get clipping effects and sampling effects when you convert back and forth.
Better to upsample once (generally a non-lossy operation) and operate on single precision FP (24-bit mantissa and an exponent) at higher sampling from that point forward and then downsample once (generally a LOSSY operation).
alowell 7 hours ago [-]
I'm using an STM32H7, not ESP32. The DAC on the PCM3168A can go up to 192 kHz, but the ADC can only be clocked up to 96 kHz.
Neywiny 8 hours ago [-]
What ESP32?
purpleidea 6 hours ago [-]
If this had open source firmware, I think I know a lot of people who would be interested, but I can't seem to find out if that's true or not and where to buy one.
phaserphile 9 hours ago [-]
how do I buy this?
askvictor 8 hours ago [-]
Or even build it. Strange there is zero info...
alowell 7 hours ago [-]
Thanks for your interest! I made a single reddit post about ETH-68 this week and this has been copied around various forums. My intent was to figure out if anyone thought this would be cool enough to produce. I'm trying to figure out if I should do a production run, open source it, or some combo of the two.
NewJazz 22 minutes ago [-]
I'd definitely be interested. Maybe put an interest form on the website? Or a link to one.
How much do you think it'd cost if you were to sell and ship to the US?
inatreecrown2 5 hours ago [-]
I think this would be a cool project to do as a DIY.
The market for a linux native Audio Interface is there I think.
At least I for one would appreciate this!
jauntywundrkind 1 days ago [-]
Really nice!!!
I wonder if gigabit would make a difference on latency, faster packet transmissions. Could packets drop from 64 to 32b?
4ms is pretty good but I feel like sub 2ms would be nicer.
> The typical default latency for a Dante audio device is 1 msec.
(emphasis mine)
The latency depends on the device. Hardware implementations of Dante commonly support latencies of 1ms or less, but software implementations are higher. The minimum latency of Dante Virtual Soundcard running on a PC is 4ms.
That's the appropriate number to compare against here (since the PC is using a software driver to interface with the network). However, that 4ms number is one-way latency, and the OP's 3.6ms number is round-trip. So this is already half the latency of DVS. (That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.)
alowell 5 hours ago [-]
> That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.
Hm, the audio latency doesn't fluctuate so it's not like the testing conditions affect the measurement. The audio latency is a fixed quantity that depends completely on the number of storage elements in the data path which isn't variable.
Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting. The histograms on my page (which should be animated BTW) show real time processing latency measurements while running the audio all the way through Bitwig with a moderate DSP load (multiple instances of Pianoteq, samplers, live MIDI input). On my system (details at the bottom of my page), I can do this at 48 kHz and 64 sample buffers with zero underruns. If I drop down to 32 samples per buffer, I do start getting underruns.
All I can do from the hardware side is try to minimize the processing latency of a typical cycle so that there is more head room to absorb jitter. The vast majority of the jitter comes from the Linux host. It's up to the end user to tune the system for low jitter. This is usually the case for audio on Linux, and the rabbit hole can go pretty deep on system tuning.
alowell 7 hours ago [-]
Yep, thanks for clarifying about this!
lukeh 5 hours ago [-]
Brooklyn, IP Core can do 125us or 250us depending on number of switch hops.
It seems like, at 16000TbaseT anyways, you’re adding an overhead of about 150% on top of copper/fiber latency plus transmit time latency to process data at that bandwidth. I wonder if the same holds true at 100 vs 1000, 2500, 10000? Certainly this is a known tradeoff for DDR performance tuning — if you don’t mind spiking response times greatly, you can get the advertised maximum speeds, else you accept less bandwidth for somewhat less latency — and they’re both effectively using the same strategies to talk over copper.
Neywiny 8 hours ago [-]
Sadly the H7 doesn't have a gigabit MAC. And most likely doing one over USB high speed host wouldn't help? But it would be an interesting experiment
alowell 7 hours ago [-]
1G would certainly decrease the processing latency (not the audio latency) by quite a bit. STM32H7 doesn't have a 1G MAC. 1G MAC is kind of rare on "friendly" microcontrollers, although there are at least two that I'm evaluating for the next project.
lstodd 7 hours ago [-]
One can fit a nice (eg, not 1-channel, but actually useable) LTE base station frontend in a gigabit eth on 2014-s tech.
eqvinox 1 days ago [-]
> Very low latency: 3.620 milliseconds round trip at 48 kHz with 64 sample buffer
"very low latency" in audio is <=1ms. 3.6ms is good but not special.
k_roy 9 hours ago [-]
But it’s matching AND exceeding the performance of a $999 PCIe card designed for a similar purpose.
I would love to know why this is hand-wavy and not very cool? Even as not-audiophile, I’ve got ideas of things to use this for.
> RME HDSPe AIO Pro PCIe
> eth68 matches the latency performance of the RME card at 48 kHz and surpasses it by 0.33 ms at 96 kHz.
Venn1 9 hours ago [-]
As the owner of the PCIe card that measurement was taken from, yes, I’d say it’s quite impressive. The average round-trip latency for a USB audio interface at 48 kHz/128 samples would usually fall somewhere around 8 ms, which is a bit much if you’re monitoring post-FX.
On Linux, we don’t have much in the way of Thunderbolt support for audio interfaces, so the only way to achieve this sort of latency has traditionally been with PCIe or PCI audio interfaces. Having a low-cost, infinitely more portable solution would be very welcome.
k_roy 8 hours ago [-]
Awesome. Love to see it!
This is just the answer I was looking for. I definitely don’t know anything about the audio space, but am into networking hardcode.
I love seeing audiophile projects that don’t involve a rebranded tp-link switch and marked up 1000%
tern 9 hours ago [-]
Round trip latency for RME devices (widely considered the best in the industry) is around 3ms, so this is indeed "very low latency".
You might be thinking of latency of the converters, which is normally sub-ms.
alowell 7 hours ago [-]
Check out the round trip latency measurements for audio interfaces on Linux here:
AFAIK the best one is RME AIO Pro PCIe card. ETH-68 matches the round trip audio latency of this card at 48 kHz and surpasses it be 0.33 ms at 96 kHz.
h3lp 9 hours ago [-]
They quote a roundtrip of 3.6ms, so one-way 1.8ms. In the plots on their page, it looks like the processing latency is centered on around 1ms:
> The LATMON pulse width is therefore an accurate measure of the total processing latency of each cycle and is affected by every element in the data path: processing delay in the microcontroller, network transmission delay, host OS delays, signal processing delay in the DAW, etc
alowell 7 hours ago [-]
You are confusing the audio latency with the processing latency. I see this mistake a lot.
One way audio latency is about 3.6 ms divided by two = 1.8 ms. This type of latency is audible.
Processing latency is just the amount of time it takes to complete all processing for each audio cycle. From the histograms, the typical processing latency is about 625 us and of course there is some jitter (almost all of the jitter comes from the Linux host BTW). Processing latency is not audible. However if the processing latency exceeds the deadline on a given audio cycle, there will be an underrun which will cause an audible glitch.
Polizeiposaune 9 hours ago [-]
> eth68 sends capture packets to 12.12.12.10:3000 by default.
Um.
alowell 7 hours ago [-]
Lol, this wasn't my choice really! The JACK server's default UDP listen port is 3000 with the netone backend, so that is default port that ETH-68 sends to. This can be adjusted on both ends
NobodyNada 5 hours ago [-]
The choice of port wasn't the complaint so much as the choice of IP address. 12.12.12.10 is a real IP address on the Internet, so if you plug this device in somewhere with an Internet connection, it will begin spamming nuisance traffic to some random device out there. It would be better to choose an IP address in one of the private spaces for this purpose.
Don't get me wrong, I'm genuinely trying to understand.
How does this compare to pipewire over ethernet? Is it realtime vs buffered?
Imagine having this on the stage right next to your analog gear, and a computer 50m away.
(Opening the page under discussion was actually helpful, it lists all this right at the top.)
But this might be even more effective.
Besides audio dsp for live use, there is also the use case of visualizers.
ETH-68 isn’t running PipeWire, or Linux. Maybe I don’t understand your question.
Not sure what you mean by realtime vs buffered. Care to elaborate?
Or is there no such synchronization, in which case there would be long-term drift?
surely that works only as long as eth68 is the only interface? but often you might want other interfaces too which inherently have their own clocks
Edit: For the context: https://www.eevblog.com/forum/chat/ti-ne5532-audio-opamp-cha...
I have my eye on some other codecs for the next project.
Just curious though, since I don’t work at these sample rates. Is resampling not good enough? Or maybe I don’t understand the mastering flow.
Are you sure that's necessary? The STM32H7 has a decent fractional PLL.
A niche use, mind you.
Basically, I found myself doing real-to-complex baseband conversion for audio signals, and doing so efficiently halved my Nyquist rate.
Increasing my sample rate to 96 kHz let me construct the 48 kHz analytic signal that I wanted.
Hoisted on my own petard :)
https://www.youtube.com/watch?v=hCQCP-5g5bo
But maybe there's a use case for replacing an AES192 signal carrying the full FM baseband signal over ETH-67? I'm not sure it's the correct fit, though.
https://www.telosalliance.com/radio-processing/audio-interfa... etc.
There ARE good reasons for recording and processing music at higher bitrates and sample sizes. Effects, especially those with positive feedback loops, can go more unstable and clip and lose information with fewer bits and samples.
There's just no good reason to DISTRIBUTE music at the higher rates.
Better to upsample once (generally a non-lossy operation) and operate on single precision FP (24-bit mantissa and an exponent) at higher sampling from that point forward and then downsample once (generally a LOSSY operation).
How much do you think it'd cost if you were to sell and ship to the US?
I wonder if gigabit would make a difference on latency, faster packet transmissions. Could packets drop from 64 to 32b?
4ms is pretty good but I feel like sub 2ms would be nicer.
(emphasis mine)
The latency depends on the device. Hardware implementations of Dante commonly support latencies of 1ms or less, but software implementations are higher. The minimum latency of Dante Virtual Soundcard running on a PC is 4ms.
That's the appropriate number to compare against here (since the PC is using a software driver to interface with the network). However, that 4ms number is one-way latency, and the OP's 3.6ms number is round-trip. So this is already half the latency of DVS. (That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.)
Hm, the audio latency doesn't fluctuate so it's not like the testing conditions affect the measurement. The audio latency is a fixed quantity that depends completely on the number of storage elements in the data path which isn't variable.
Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting. The histograms on my page (which should be animated BTW) show real time processing latency measurements while running the audio all the way through Bitwig with a moderate DSP load (multiple instances of Pianoteq, samplers, live MIDI input). On my system (details at the bottom of my page), I can do this at 48 kHz and 64 sample buffers with zero underruns. If I drop down to 32 samples per buffer, I do start getting underruns.
All I can do from the hardware side is try to minimize the processing latency of a typical cycle so that there is more head room to absorb jitter. The vast majority of the jitter comes from the Linux host. It's up to the end user to tune the system for low jitter. This is usually the case for audio on Linux, and the rabbit hole can go pretty deep on system tuning.
It seems like, at 16000TbaseT anyways, you’re adding an overhead of about 150% on top of copper/fiber latency plus transmit time latency to process data at that bandwidth. I wonder if the same holds true at 100 vs 1000, 2500, 10000? Certainly this is a known tradeoff for DDR performance tuning — if you don’t mind spiking response times greatly, you can get the advertised maximum speeds, else you accept less bandwidth for somewhat less latency — and they’re both effectively using the same strategies to talk over copper.
"very low latency" in audio is <=1ms. 3.6ms is good but not special.
I would love to know why this is hand-wavy and not very cool? Even as not-audiophile, I’ve got ideas of things to use this for.
> RME HDSPe AIO Pro PCIe
> eth68 matches the latency performance of the RME card at 48 kHz and surpasses it by 0.33 ms at 96 kHz.
On Linux, we don’t have much in the way of Thunderbolt support for audio interfaces, so the only way to achieve this sort of latency has traditionally been with PCIe or PCI audio interfaces. Having a low-cost, infinitely more portable solution would be very welcome.
This is just the answer I was looking for. I definitely don’t know anything about the audio space, but am into networking hardcode.
I love seeing audiophile projects that don’t involve a rebranded tp-link switch and marked up 1000%
You might be thinking of latency of the converters, which is normally sub-ms.
https://interfacinglinux.com/linux-compatible-audio-interfac...
AFAIK the best one is RME AIO Pro PCIe card. ETH-68 matches the round trip audio latency of this card at 48 kHz and surpasses it be 0.33 ms at 96 kHz.
> The LATMON pulse width is therefore an accurate measure of the total processing latency of each cycle and is affected by every element in the data path: processing delay in the microcontroller, network transmission delay, host OS delays, signal processing delay in the DAW, etc
One way audio latency is about 3.6 ms divided by two = 1.8 ms. This type of latency is audible.
Processing latency is just the amount of time it takes to complete all processing for each audio cycle. From the histograms, the typical processing latency is about 625 us and of course there is some jitter (almost all of the jitter comes from the Linux host BTW). Processing latency is not audible. However if the processing latency exceeds the deadline on a given audio cycle, there will be an underrun which will cause an audible glitch.
Um.
See, for example, Cloudflare's stories of issues caused by misconfigured systems appropriating the "1.1.1.1" address for their own purposes: https://blog.cloudflare.com/fixing-reachability-to-1-1-1-1-g...