btrfs: zoned: reset meta_write_pointer on zone reset

btrfs_reset_unused_block_groups() resets a block group's zone and sets
alloc_offset back to 0 so the space can be reused, but it leaves
meta_write_pointer pointing at the previous end of the zone.

Once the block group is reactivated and reused for metadata, newly
allocated tree blocks live before that stale write pointer.
btrfs_check_meta_write_pointer() then sees them behind the write pointer,
so they can never be written out in sequential order: the dirty extent
buffers are stranded and pin their btree_inode folios until unmount.

Reset meta_write_pointer back to the start of the block group for
metadata and system block groups.

Fixes: 453a73c306 ("btrfs: zoned: reclaim unused zone by zone resetting")
Reviewed-by: Naohiro Aota <naohiro.aota@wdc.com>
Signed-off-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Signed-off-by: David Sterba <dsterba@suse.com>
This commit is contained in:
Johannes Thumshirn 2026-07-03 07:54:45 +02:00 committed by David Sterba
parent 1ebe51c29f
commit 5fabb1cf25

View File

@ -3189,6 +3189,17 @@ int btrfs_reset_unused_block_groups(struct btrfs_space_info *space_info, u64 num
reclaimed = bg->alloc_offset;
bg->zone_unusable = bg->length - bg->zone_capacity;
bg->alloc_offset = 0;
/*
* The zone was just reset to empty, so alloc_offset went back to
* the start of the zone. For metadata/system block groups the
* write pointer must follow it back to the start of the zone;
* otherwise it stays stale at the previous (finished) zone end,
* and metadata written into the reused zone would sit behind the
* write pointer, could never be written out in sequential order,
* and would be stranded (pinning its folio) until unmount.
*/
if (bg->flags & (BTRFS_BLOCK_GROUP_METADATA | BTRFS_BLOCK_GROUP_SYSTEM))
bg->meta_write_pointer = bg->start;
/*
* This holds because we currently reset fully used then freed
* block group.