zed and udev thrashing with repeated online events #7366

abrasive · 2018-03-29T10:59:41Z

System information

Type	Version/Name
Distribution Name	Gentoo
Distribution Version	(up to date)
Linux Kernel	4.9.76-gentoo-r1
Architecture	x86_64
ZFS Version	0.7.6
SPL Version	0.7.6
udev version	217

Describe the problem you're observing

In this recently built system, I often find the disks thrashing with no load on the system. iotop shows zed and udevd are the only things touching the disk. I haven't yet worked out what triggers it, but stopping and restarting zed stops the thrashing. This sounds like it might be related to #6667 and potentially also #7132, but I don't understand why the udev events are being generated.

udevadm monitor shows a constant stream of changes on the ZFS partition of three of the four disks in my (2x2 mirror) array:

KERNEL[800477.625559] change   /devices/pci0000:00/0000:00:01.0/0000:01:00.0/host8/port-8:2/end_device-8:2/target8:0:2/8:0:2:0/block/sdd/sdd2 (block)
UDEV  [800477.641848] change   /devices/pci0000:00/0000:00:01.0/0000:01:00.0/host8/port-8:2/end_device-8:2/target8:0:2/8:0:2:0/block/sdd/sdd2 (block)
KERNEL[800477.674069] change   /devices/pci0000:00/0000:00:01.0/0000:01:00.0/host8/port-8:0/end_device-8:0/target8:0:0/8:0:0:0/block/sdb/sdb2 (block)
UDEV  [800477.686767] change   /devices/pci0000:00/0000:00:01.0/0000:01:00.0/host8/port-8:0/end_device-8:0/target8:0:0/8:0:0:0/block/sdb/sdb2 (block)
KERNEL[800477.717713] change   /devices/pci0000:00/0000:00:01.0/0000:01:00.0/host8/port-8:3/end_device-8:3/target8:0:3/8:0:3:0/block/sde/sde2 (block)
UDEV  [800477.729357] change   /devices/pci0000:00/0000:00:01.0/0000:01:00.0/host8/port-8:3/end_device-8:3/target8:0:3/8:0:3:0/block/sde/sde2 (block)

Meanwhile zed is logging away about the same devices coming online, eg.:

Mar 29 21:40:53 obtainium zed[6393]: zfsdle_vdev_online: setting device '/dev/sdd2' to ONLINE state in pool 'base'
Mar 29 21:40:53 obtainium zed[6393]: zfs_slm_event: EC_dev_status.dev_dle
Mar 29 21:40:53 obtainium zed[6393]: zfsdle_vdev_online: searching for 'pci-0000:01:00.0-sas-phy2-lun-0' in 'base'
Mar 29 21:40:53 obtainium zed[6393]: zed_disk_event:
Mar 29 21:40:53 obtainium zed[6393]: ^Iclass: EC_dev_status
Mar 29 21:40:53 obtainium zed[6393]: ^Isubclass: dev_dle
Mar 29 21:40:53 obtainium zed[6393]: ^Idev_name: /dev/sdd2
Mar 29 21:40:53 obtainium zed[6393]: ^Ipath: /devices/pci0000:00/0000:00:01.0/0000:01:00.0/host8/port-8:2/end_device-8:2/target8:0:2/8:0:2:0/block/sdd/sdd2
Mar 29 21:40:53 obtainium zed[6393]: ^Idevid: scsi-[redacted]-part2
Mar 29 21:40:53 obtainium zed[6393]: ^Iphys_path: pci-0000:01:00.0-sas-phy4-lun-0
Mar 29 21:40:53 obtainium zed[6393]: ^Idev_size: 5983994712064
Mar 29 21:40:53 obtainium zed[6393]: ^Ipool_guid: [redacted]
Mar 29 21:40:53 obtainium zed[6393]: ^Ivdev_guid: [redacted]

There are no other zed log entries in my syslog during this period (eg config_sync).

Describe how to reproduce the problem

Wait. Does not happen immediately at boot.

The text was updated successfully, but these errors were encountered:

abrasive · 2018-03-29T11:36:06Z

It seems to be possible to trigger this by starting and stopping a scrub on the pool.

abrasive · 2018-03-29T23:15:51Z

Another reliable way to trigger this is to issue:

sg_logs -t  /dev/disk/by-path/pci-0000:01:00.0-sas-phy0-lun-0

This then starts spamming events on whichever disk was queried. This action causes a udev change event for the block device and for each of its partitions.

abrasive · 2018-03-29T23:23:04Z

strace shows that zed is doing a BLKFLSBUF ioctl which is causing the udev change event.

Its stack traces seem to indicate that this is due to zpool_relabel_disk being called from zpool_vdev_online.

Disabling autoexpand on the pool prevents the problem from occurring.

Flushing the device and invaliding the page cache to ensure a consistent view of the device is only needed when the device was successfully expanded. Unconditionally performing these operations is safe but it has the side effect of triggering another udev change event which will be handled by the ZED. Signed-off-by: Brian Behlendorf <behlendorf1@llnl.gov> Issue openzfs#7366

behlendorf · 2018-04-05T19:34:07Z

@abrasive thank you for identifying the root cause here. If possible would you mind verifying the proposed fix in #7393.

abrasive · 2018-04-06T11:09:20Z

Thanks for the patch, @behlendorf! Unfortunately it does not work. The change events seem to be generated upon the close(fd).

I note, after some experimentation, that if you close a write filehandle to a disk partition, you get a udev change event even if you didn't write anything. This doesn't happen with read handles. For example sudo dd if=/dev/null of=/dev/sda will do the job.

behlendorf · 2018-06-19T21:48:10Z

I believe PR #7629 may resolve this by filtering out all udev change event for partitions.

abrasive · 2018-06-20T02:14:36Z

Thanks for the update, @behlendorf. I've reenabled autoexpand on my pool but I can no longer reproduce the problem! zfs, udev and kernel all remain at the original versions. No disks have changed either.

I'm happy to dig further if you still want to pursue this one but if so I could do with some help figuring out what triggering factor might have changed.

While the autoexpand property may seem like a small feature it depends on a significant amount of system infrastructure. Enough of that infrastructure is now in place with a few modifications for Linux it can be supported. Auto-expand works as follows; when a block device is modified (re-sized, closed after being open r/w, etc) a change uevent is generated for udev. The ZED, which is monitoring udev events, passes the change event along to zfs_deliver_dle() if the disk or partition contains a zfs_member as identified by blkid. From here the device is matched against all imported pool vdevs using the vdev_guid which was read from the label by blkid. If a match is found the ZED reopens the pool vdev. This re-opening is important because it allows the vdev to be briefly closed so the disk partition table can be re-read. Otherwise, it wouldn't be possible to report thee maximum possible expansion size. Finally, if the property autoexpand=on a vdev expansion will be attempted. After performing some sanity checks on the disk to verify that it is safe to expand, the primary partition (-part1) will be expanded and the partition table updated. The partition is then re-opened (again) to detect the updated size which allows the new capacity to be used. In order to make all of the above possible the following changes were required: * Updated the zpool_expand_001_pos and zpool_expand_003_pos tests. These tests now create a pool which is layered on a loopback, scsi_debug, and file vdev. This allows for testing of non- partitioned block device (loopback), a partition block device (scsi_debug), and a file which does not receive udev change events. This provided for better test coverage, and by removing the layering on ZFS volumes there issues surrounding layering one pool on another are avoided. * zpool_find_vdev_by_physpath() updated to accept a vdev guid. This allows for matching by guid rather than path which is a more reliable way for the ZED to reference a vdev. * Fixed zfs_zevent_wait() signal handling which could result in the ZED spinning when a signal was not handled. * Removed vdev_disk_rrpart() functionality which can be abandoned in favor of kernel provided blkdev_reread_part() function. * Added a rwlock which is held as a writer while a disk is being reopened. This is important to prevent errors from occurring for any configuration related IOs which bypass the SCL_ZIO lock. The zpool_reopen_007_pos.ksh test case was added to verify IO error are never observed when reopening. This is not expected to impact IO performance. Additional fixes which aren't critical but were discovered and resolved in the course of developing this functionality. * Added PHYS_PATH="/dev/zvol/dataset" to the vdev configuration for ZFS volumes. This is as good as a unique physical path, while the volumes are not used in the test cases anymore for other reasons this improvement was included. Signed-off-by: Brian Behlendorf <behlendorf1@llnl.gov> Issue openzfs#120 Issue openzfs#2437 Issue openzfs#5771 Issue openzfs#7366 Issue openzfs#7582

While the autoexpand property may seem like a small feature it depends on a significant amount of system infrastructure. Enough of that infrastructure is now in place with a few modifications for Linux it can be supported. Auto-expand works as follows; when a block device is modified (re-sized, closed after being open r/w, etc) a change uevent is generated for udev. The ZED, which is monitoring udev events, passes the change event along to zfs_deliver_dle() if the disk or partition contains a zfs_member as identified by blkid. From here the device is matched against all imported pool vdevs using the vdev_guid which was read from the label by blkid. If a match is found the ZED reopens the pool vdev. This re-opening is important because it allows the vdev to be briefly closed so the disk partition table can be re-read. Otherwise, it wouldn't be possible to report thee maximum possible expansion size. Finally, if the property autoexpand=on a vdev expansion will be attempted. After performing some sanity checks on the disk to verify that it is safe to expand, the primary partition (-part1) will be expanded and the partition table updated. The partition is then re-opened (again) to detect the updated size which allows the new capacity to be used. In order to make all of the above possible the following changes were required: * Updated the zpool_expand_001_pos and zpool_expand_003_pos tests. These tests now create a pool which is layered on a loopback, scsi_debug, and file vdev. This allows for testing of non- partitioned block device (loopback), a partition block device (scsi_debug), and a file which does not receive udev change events. This provided for better test coverage, and by removing the layering on ZFS volumes there issues surrounding layering one pool on another are avoided. * zpool_find_vdev_by_physpath() updated to accept a vdev guid. This allows for matching by guid rather than path which is a more reliable way for the ZED to reference a vdev. * Fixed zfs_zevent_wait() signal handling which could result in the ZED spinning when a signal was not handled. * Removed vdev_disk_rrpart() functionality which can be abandoned in favor of kernel provided blkdev_reread_part() function. * Added a rwlock which is held as a writer while a disk is being reopened. This is important to prevent errors from occurring for any configuration related IOs which bypass the SCL_ZIO lock. The zpool_reopen_007_pos.ksh test case was added to verify IO error are never observed when reopening. This is not expected to impact IO performance. Additional fixes which aren't critical but were discovered and resolved in the course of developing this functionality. * Added PHYS_PATH="/dev/zvol/dataset" to the vdev configuration for ZFS volumes. This is as good as a unique physical path, while the volumes are not used in the test cases anymore for other reasons this improvement was included. Signed-off-by: Sara Hartse <sara.hartse@delphix.com> Signed-off-by: Brian Behlendorf <behlendorf1@llnl.gov> Issue openzfs#120 Issue openzfs#2437 Issue openzfs#5771 Issue openzfs#7366 Issue openzfs#7582

shodanshok · 2018-12-31T11:42:47Z

@behlendorf I can consistently hit that problem by using the reproducer from @abrasive - ie: autoexpand=on and sg_logs -t <disk>. From my understading, PR #7629 was reverted due to corruption issues. Can you confirm that we should not use autoexpand to properly avoid the issue described here? Or do other workarounds exist? Thanks.

behlendorf · 2019-01-03T00:14:49Z

@shenyan1 that's right for the 0.7.x releases. For now, I'd suggest leaving autoexpand=off. The required fix has been integrated to master to resolve this issue. If you're in a position where you can verify the correct behavior for your environment using latest 0.8.0-rc2 release candidate I'd recommend giving it a try.

xgiovio · 2022-02-06T21:17:17Z

strace shows that zed is doing a BLKFLSBUF ioctl which is causing the udev change event.

Its stack traces seem to indicate that this is due to zpool_relabel_disk being called from zpool_vdev_online.

Disabling autoexpand on the pool prevents the problem from occurring.

this is the solution. Setting the autoexpand property to off removes the problems and no more config sync messages

tonyhutter · 2022-07-29T19:25:25Z

reopening and self-assigning so I can track this with #7132

Users were seeing floods of `config_sync` events when autoexpand was enabled. This happened because all "disk status change" udev events invoke the autoexpand codepath, which calls zpool_relabel_disk(), which in turn cause another "disk status change" event to happen, in a feedback loop. Note that "disk status change" happens every time a user calls close() on a block device. This commit breaks the feedback loop by only allowing an autoexpand to happen if the disk actually changed size. Fixes: openzfs#7132 Fixes: openzfs#7366 Signed-off-by: Tony Hutter <hutter2@llnl.gov>

Users were seeing floods of `config_sync` events when autoexpand was enabled. This happened because all "disk status change" udev events invoke the autoexpand codepath, which calls zpool_relabel_disk(), which in turn cause another "disk status change" event to happen, in a feedback loop. Note that "disk status change" happens every time a user calls close() on a block device. This commit breaks the feedback loop by only allowing an autoexpand to happen if the disk actually changed size. Reviewed-by: Brian Behlendorf <behlendorf1@llnl.gov> Signed-off-by: Tony Hutter <hutter2@llnl.gov> Closes: openzfs#7132 Closes: openzfs#7366 Closes openzfs#13729

behlendorf mentioned this issue Apr 5, 2018

Fix zed/udev thrashing due to autoexpand=on #7393

Closed

13 tasks

abrasive mentioned this issue Jun 19, 2018

Random reads by 9 z_wr_iss threads crippling sequential write performance #7594

Closed

behlendorf mentioned this issue Jun 28, 2018

Add support for autoexpand property #7629

Merged

13 tasks

behlendorf closed this as completed in d441e85 Jul 23, 2018

ericdaltman mentioned this issue Sep 25, 2018

Huge performance drop (30%~60%) after upgrading to 0.7.9 from 0.6.5.11 #7834

Closed

xgiovio mentioned this issue Feb 6, 2022

zed excessive logging with class=config_sync #7132

Closed

tonyhutter self-assigned this Jul 29, 2022

tonyhutter reopened this Jul 29, 2022

tonyhutter mentioned this issue Aug 3, 2022

zed: Fix config_sync autoexpand flood #13729

Merged

13 tasks

behlendorf closed this as completed in e27e692 Sep 8, 2022

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

zed and udev thrashing with repeated online events #7366

zed and udev thrashing with repeated online events #7366

abrasive commented Mar 29, 2018

abrasive commented Mar 29, 2018

abrasive commented Mar 29, 2018 •

edited

abrasive commented Mar 29, 2018 •

edited

behlendorf commented Apr 5, 2018

abrasive commented Apr 6, 2018

behlendorf commented Jun 19, 2018

abrasive commented Jun 20, 2018

shodanshok commented Dec 31, 2018

behlendorf commented Jan 3, 2019

xgiovio commented Feb 6, 2022

tonyhutter commented Jul 29, 2022

zed and udev thrashing with repeated online events #7366

zed and udev thrashing with repeated online events #7366

Comments

abrasive commented Mar 29, 2018

System information

Describe the problem you're observing

Describe how to reproduce the problem

abrasive commented Mar 29, 2018

abrasive commented Mar 29, 2018 • edited

abrasive commented Mar 29, 2018 • edited

behlendorf commented Apr 5, 2018

abrasive commented Apr 6, 2018

behlendorf commented Jun 19, 2018

abrasive commented Jun 20, 2018

shodanshok commented Dec 31, 2018

behlendorf commented Jan 3, 2019

xgiovio commented Feb 6, 2022

tonyhutter commented Jul 29, 2022

abrasive commented Mar 29, 2018 •

edited

abrasive commented Mar 29, 2018 •

edited