Linux Storage Networking, Part 2 of 4.
Part 1: ARP Flux and NIC Binding
Two NICs, one bond, one IP. That sounds easier than dealing with two interfaces on the same subnet. And for some things it is.
For iSCSI and NVMe/TCP, I would rather let multipath use separate interfaces. A bond changes how traffic spreads across the links, and it hides those links from the storage stack when something goes wrong. That’s worth knowing before you decide the second cable solved the problem. They work at different layers. LACP picks an Ethernet link for each flow. Storage multipathing, often called MPIO, selects among paths to the storage device. In Linux that commonly means DM-Multipath for iSCSI and native NVMe multipath for NVMe/TCP. One doesn’t replace the other.
In part one on Linux iSCSI interface binding and ARP flux, the example was a Linux host with two storage NICs and four reachable array ports. Each NIC connected to all four ports, giving us eight paths per volume. This time, put the NICs in an LACP bond and connect from one IP to each array port. That gives us four sessions over the bond. The extra cable is still there. How much of it do we actually use?
How LACP hashing distributes iSCSI connections
A normal LACP bond distributes TCP flows across its member links. It doesn’t take one connection and give it the combined bandwidth of both NICs. Keeping packets in order is part of the reason. The Linux bond chooses a member using a hash of packet headers. xmit_hash_policy determines which fields go into that calculation. For ordinary unfragmented TCP traffic, the useful distinction is:
| Policy | Fields that distinguish traffic |
|---|---|
layer2 | MAC addresses and Ethernet protocol |
layer2+3 | MAC addresses and IP addresses |
layer3+4 | IP addresses and transport ports |
On a host with four iSCSI sessions each one gets its own TCP connection. The host’s MAC and IP are the same. The destination TCP port is normally 3260. The array portal changes, and each connection gets a source port.

With layer2, each destination MAC gives you a fixed result. If the MACs happen to hash to the same member, that is where the traffic goes. If the array is behind a router, the next-hop MAC may be the same for every portal. Now all those different destination IPs can still produce the same link choice.
layer2+3 includes IP addresses, so routing doesn’t collapse everything to the gateway’s MAC. But a small number of host/portal pairs is still a small number of hash results. You don’t get a promise that two go left and two go right.
layer3+4 adds the TCP ports. That gives separate connections to the same destination more opportunities to use different members. It is usually the useful starting point when a supported storage design calls for LACP. The bonding driver documentation also describes its caveat around fragmented traffic and packet ordering.
One last thing about these. Unless you know the has algorithms specifics you don’t have a deterministic placement policy. You don’t know where you traffic goes until it’s there. You didn’t do the math on both the host side and switch side. Non deterministic behavior during a failure is much harder to troubleshoot than deterministic behavior.
Four sessions isn’t much to spread around
Here is a simple way to look at the distribution. Assume four equally busy connections, with each independently landing on either of two links with equal probability:

Under those assumptions, an even split happens 37.5% of the time. All four landing on one link happens 12.5% of the time. The rest is a three-and-one split. That’s a simplified model. A real hash is deterministic, source ports aren’t independent coin flips, and the sessions probably aren’t equally busy. So don’t read the chart as a prediction. The point is that four connections aren’t enough to count on balanced links. When a connection is recreated, a new source port can change its hash. So the distribution can change after a reconnect or reboot even when the storage configuration hasn’t changed. If you’re looking at throughput, look at the individual bond members too.
NVMe/TCP has more connections to work with
NVMe/TCP is a little different. A controller connection has an admin queue and I/O queues, with separate TCP connections for those queues. How many you get depends on the host, target and connection settings. More TCP connections give layer3+4 more chances to distribute the traffic. That can spread a lot better than a handful of single-connection iSCSI sessions. It doesn’t give native NVMe multipath visibility into the bond members, though. A queue affected by a bad member can trigger transport recovery for its controller connection. The exact recovery depends on the failure and driver behavior, but the storage stack still can’t select a healthy Ethernet member directly. That is the bond’s job.
The switch chooses where reads come back
xmit_hash_policy controls traffic leaving the host. The switch independently chooses the bond member used to send traffic back.

For storage, most of the payload coming back is your reads. You can change the host’s transmit hash and still have reads arriving on one NIC because the switch is using a MAC-based policy. If you use LACP, review the switch’s port-channel hashing with the network team. Check actual member counters in both directions. Good write numbers don’t tell you anything about the reads.
LACP failure detection: link up doesn’t mean it works
A cable pull is easy to detect. With miimon=100, the bond checks carrier about every 100 milliseconds. That’s how often it polls. Your application isn’t going to recover in a tenth of a second. A link that remains up while dropping traffic is harder. If LACP control packets stop arriving, LACP can expire the partner. Fast LACP normally uses a roughly three-second timeout; slow LACP can take roughly 90 seconds.

The LACP numbers in this picture apply when the control packets stop. If data forwarding is broken but LACP packets continue to get through, LACP stays perfectly happy. It’s only checking that the partner is still talking to it. The iSCSI example in the picture assumes a five-second idle interval and a five-second response timeout. Your values may differ. And ten seconds isn’t the whole story, since transport recovery and multipath timeouts come after it.
If the bond still considers a member usable, reconnecting a session may not help. With a MAC/IP hash, the recreated connection can go straight back to the same member. With a port-based hash, it might land somewhere else. That can look like a path that fixes itself and then breaks again later. Also, the bonding driver’s ARP monitor is not supported in 802.3ad mode. Fast LACP helps, but only with the failures it can see.
What about the other bond modes?
They exist for different problems. Here’s the short version from a storage point of view:
| Mode | What I would keep in mind for storage |
|---|---|
balance-rr | Sends packets round-robin. Reordering within a TCP flow can cause retransmissions and poor throughput. |
active-backup | Predictable failover, but only one member carries traffic at a time. Multipath still sees the logical interface. |
balance-xor | Uses a static hash without LACP negotiation. Distribution and failure detection still need attention. |
broadcast | Sends every frame on every member. That is not extra useful storage bandwidth. |
802.3ad | LACP, with the flow distribution and monitoring behavior above. |
balance-tlb | Balances transmit traffic; receive traffic normally uses one member. |
balance-alb | Also balances IPv4 receive traffic by manipulating ARP replies. Distribution depends on the peers and learned mappings. |
In particular, an array with several target IPs may appear as several peers for adaptive receive balancing. That still doesn’t guarantee evenly spread reads, and multipath still can’t see the individual links.
What storage multipath can see through a bond
In our example, four sessions over one bond all depend on the same logical interface. Multipath can fail a storage path when that path becomes unhealthy. It can’t say “stop using the second bond member” because that member is below the layer it controls. The four-versus-eight count comes from this example’s connection layout. Bonding doesn’t impose a universal four-path limit. You can log in more sessions, but more sessions still won’t show multipath the physical links.
There is a switch consideration too. A conventional LACP bundle across two switches requires them to present a common aggregation system, usually through stacking or MLAG. That ties the two switches together, and they can fail together in ways two separate fabrics wouldn’t. I am fine with bonding for management and cluster traffic when the platform calls for it. It is also common with NFS, where extra connections such as nconnect can give the hash more flows. NFS has its own multipathing and trunking options on supported clients and servers, so that needs its own design discussion.
For block storage, separate storage NICs and storage multipathing are easier for me to follow and troubleshoot. If the supported design requires a bond anyway, start with mode=802.3ad, xmit_hash_policy=layer3+4, miimon=100 and lacp_rate=fast on the host, and L3/L4 hashing with fast LACP on the switch port-channel. Then check the member counters in both directions, and the failure behavior under I/O. Test Test Test. Don’t assume you did it right.
Storage network checks: MTU, boot ordering and failover
Bonded or not, these are easy to miss.
First, MTU. A successful login doesn’t prove that a large data packet can get through. If you use a 9000-byte IPv4 MTU, test each intended path with a full-sized packet and fragmentation disabled:
ping -c 3 -M do -s 8972 -I ens1f0 10.10.1.100ping -c 3 -M do -s 8972 -I ens1f1 10.10.1.102
Those commands are for the separate-NIC topology from part one. 8972 leaves room for the usual IPv4 and ICMP headers. Using the interface name selects the device; supplying only a source IP is not the same test. On a bond, a ping follows the bond’s selection, so one successful ping doesn’t exercise every member. Host NICs, VLANs, switch ports, inter-switch links and target interfaces all need compatible MTUs. Routed paths also depend on working path-MTU discovery. Test real I/O after the ping test.
Second, host- and array-facing switch ports should have the appropriate edge-port configuration. Classic spanning-tree delays after link-up can leave the host trying to use a port that isn’t forwarding yet. Apply that setting to actual endpoint ports, not switch-to-switch links.
Third, give each interface one configuration owner. That’s NetworkManager on the RHEL family and SLES 16, wicked on SLES 15, netplan with systemd-networkd on Ubuntu, and ifupdown on Debian and Proxmox. Use the one your installation uses, and check for cloud-init, old DHCP profiles or provisioning tools that also touch the NIC. When two tools fight over an interface, the winner can change from boot to boot.
Fourth, make interface names stable too, by pinning each storage NIC’s name to its MAC address with a systemd .link file or a NetworkManager match. If something happens that make an interface name change like a firmware update or a new PCI device installation an iSCSI binding to the old interface name never logs in.
Then check boot ordering. The storage interfaces need addresses and routes before login or autoconnect. network-online.target doesn’t necessarily mean every storage interface is ready. Network-backed mounts need the appropriate dependencies, including _netdev where applicable, and the initiator needs a retry strategy.
For NVMe/TCP, use the multipath implementation supported by your distribution and array. With native NVMe multipath, confirm it is enabled and inspect the subsystem policy:
cat /sys/module/nvme_core/parameters/multipathcat /sys/class/nvme-subsystem/nvme-subsys*/iopolicynvme list-subsys
Don’t configure a second, competing multipath layer around the same devices. If DM-Multipath is also running for iSCSI, exclude the NVMe devices from it:
blacklist { devnode "^nvme"}
Merge that into the existing blacklist section rather than adding a second one, and check the distribution’s exclusions and array guidance instead of assuming one multipath.conf applies everywhere.
For iSCSI, remember that existing node records retain settings copied when they were created. Editing iscsid.conf doesn’t rewrite those records or retune every running session. Compare the records and live session settings with what you intended.
Finally, test both cables while I/O is running. Let everything recover between tests. Then test the switch or fabric failure you claim to tolerate. A cable pull is a good test, but it’s the easy failure. A switch that keeps link up while dropping data is the one that hurts.
That is enough bonding for now. Part three moves the array onto a routed network, which gives us a different set of ways to send storage down the wrong interface.
One thought on “LACP vs Multipath: Linux iSCSI and NVMe/TCP”