spaceywilly 11 hours ago

That’s very interesting, thank you for sharing. Looks like it could be a very useful tool for testing high performance networking.

I wonder if something similar could be done using TC BPF instead of AF_XDP? My only reservation about AF XDP is that it requires a special NIC to support it, so it may not be useful for a “regular Joe” user. I wonder if TC BPF would also work since it similarly bypasses the Kernel networking stack, I believe you can put packets directly into the NIC TX queue for transmission

  • tptacek 11 hours ago

    You can, but the interesting thing about AF_XDP is that you've got a userland path to writing directly to the card's DMA buffers; TC BPF still allocates an skbuff for every packet you send.

  • Palomides 10 hours ago

    it seems like every NIC on the market that can do 100Gb has support in its linux kernel driver, so probably not a big deal in practice

  • bgpdude 11 hours ago

    you can use generic af_xdp which sits at the TC layer. Just get a bit less performance.

    • bgpdude 11 hours ago

      That's what the veth examples are using

barryvand 9 hours ago

ooh no more DPDK, this will make it a lot easier. I just started working with trex but it's all really complicated. Going to give this a try

  • tptacek 7 hours ago

    I think, and I'm saying this in part to get someone to correct me, that post-XDP (so 5 years or so now) DPDK is basically obsolete. Is there a circumstance where it would make sense to start from DPDK rather than XDP?

    • pstavirs an hour ago

      At this point there's a much bigger ecosystem for DPDK than AF_XDP I think. Also more people know about DPDK than AF_XDP right now e.g. the Ostinato traffic generator's line-rate Turbo functionality uses AF_XDP but most customers assume it uses DPDK.

      Disclosure: Ostinato creator here.

    • bgpdude 4 hours ago

      I think you're correct. Not really aware of any real limitations, other than it's slightly slower than dpdk (it's not a complete bypass), but at a much easier ease of use.

      • tptacek 4 hours ago

        AF_XDP kind of is a complete bypass, right? RX get scooped right off the DMA buffer for the card, and TX get shoved right back in.

        • bgpdude 4 hours ago

          Yeah fair point. I should have been more precise. AF_XDP with ZC bypasses the kernel networking stack and the NIC DMA's directly into UMEM and so in that sense it absolutely is a kernel bypass.

          The difference I was getting at vs DPDK is that the kernel NIC driver/NAPI/XDP path is still involvd. With DPDK the userspace PMD is effectively driving the NIC and accessing the queues directly.

          either way, it's great and everyone should use it :) that is assuming they have a use case for it. The use-cases are perhaps somewhat limited as it also bypasses the kernel tcp-ip stack, so you gotta do a lot yourself.