mirror of
https://github.com/torvalds/linux.git
synced 2026-10-08 11:36:02 +02:00
98fc57d167
108301 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
97be98b94d |
- Serialize truncate, fallocate, and mmap fault paths with
invalidate_lock, avoiding mmap failures during concurrent size changes
and exposure of uninitialized data during allocation.
- Correct fallocate signal and zeroing error handling.
- Fix FITRIM range alignment to prevent discard requests from extending
into allocated clusters.
- Fix free-cluster accounting when cluster-freeing rollback or bitmap
clearing fails.
- Keep volumes marked dirty when ntfs errors have been recorded.
- Compute bi_sector in 512-byte units, preventing silent corruption on
4Kn devices.
- Validate sectors_per_cluster values and prevent undefined shifts when
parsing MFT and index record sizes.
- Bound $AttrDef traversal to the loaded table size.
- Fix MFT record resizing, memmove overlap, and kmap_local cleanup issues.
- Improve error propagation across attribute, EA, and reparse operations,
including returning -ERANGE for undersized xattr buffers.
- Avoid modifying the HasEA flag when setxattr fails and return
DT_UNKNOWN when directory inode lookup fails.
- Reduce contention in WOF decompression by performing block reads outside
the decompression lock.
-----BEGIN PGP SIGNATURE-----
iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqZGasWHGxpbmtpbmpl
b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCIdnD/9OEohX3GvIqwHT90GubLIGunJr
S2D1MSJz0AwNF0sNQhTawAc3fbjwI77B2H3mI/Xghkd4IvgtzcY/L/jYfaZ3M7sn
Grctto0BypHI5DuBbArfjTQdW/NkPR0IpXGyBLQ8sO6aYVUPGAG0lvL9tT1Zm52N
JQU1mtjEihE5ZpD79gx8PexuDJHIg0uuok4EANk9Vu+Ub68bDBsnl/Zyxm4spIEA
976QAdboGDvo+71IdpPSaMuSAMytOf7LDJqxECqZXN5aUOoz9wrJnjELVg+xRE6c
AFM9hHZ4tZ0zs5A0EpR835URaB/bxGWpbGdkCyDDBm+QMHiTGNp5nGFl/Sd/NO5B
NcSaj0Tc2+7DbcTLU2hk1FhUsEk8eTwBZK05gxE6OajAIfAyXqxTNeB1th5smLKG
PjKWjQh9F2okIB71D6jkdntAs/0RPyuu37bTl0EtJeuRoWYZooHkgeA+tV350BQz
vjOwQuVUnDoNRQ0z1egrqAZalgjoNG7xLu0fI+n7eXZ5a4XZvQpYlKWKsodXYVwX
TBwnQhst8zEx49fe5dIBGnLiZhhQMe2zxmh7lxAOu6VcVPE4Xv6jSwzcilyiPtQF
YrdZGOEgtByOLsD+m1PRYZvukvQDDpn+NX6w4dJ+YL3MjGB7X1e2c7imqonJJYM1
0hByGPAn3g+5XdGSig==
=gAcy
-----END PGP SIGNATURE-----
Merge tag 'ntfs-for-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs
Pull ntfs fixes from Namjae Jeon:
- Serialize truncate, fallocate, and mmap fault paths with
invalidate_lock, avoiding mmap failures during concurrent size
changes and exposure of uninitialized data during allocation
- Correct fallocate signal and zeroing error handling
- Fix FITRIM range alignment to prevent discard requests from extending
into allocated clusters
- Fix free-cluster accounting when cluster-freeing rollback or bitmap
clearing fails
- Keep volumes marked dirty when ntfs errors have been recorded
- Compute bi_sector in 512-byte units, preventing silent corruption on
4Kn devices
- Validate sectors_per_cluster values and prevent undefined shifts when
parsing MFT and index record sizes
- Bound $AttrDef traversal to the loaded table size
- Fix MFT record resizing, memmove overlap, and kmap_local cleanup
issues
- Improve error propagation across attribute, EA, and reparse
operations, including returning -ERANGE for undersized xattr buffers
- Avoid modifying the HasEA flag when setxattr fails and return
DT_UNKNOWN when directory inode lookup fails
- Reduce contention in WOF decompression by performing block reads
outside the decompression lock
* tag 'ntfs-for-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs: (23 commits)
ntfs: take invalidate_lock in ntfs_filemap_page_mkwrite()
ntfs: take invalidate_lock in ntfs_setattr_size()
ntfs: handle signal interruption in fallocate
ntfs: fix FITRIM range alignment
ntfs: read WOF chunks outside the decompression lock
ntfs: leave HasEA flag untouched on setxattr failure
ntfs: fix race between fallocate and mmap reads
ntfs: fix memmove overlap in ntfs_new_attr_flags
ntfs: compute bi_sector in 512-byte units
ntfs: reject invalid sectors_per_cluster in the boot sector
ntfs: bound $AttrDef table walk to the loaded table size
ntfs: fix undefined behavior in mft/index record size calculation
ntfs: treat any nonzero dio zero-range return as an error
ntfs: fix incorrect MFT record pointer passed to ntfs_attr_record_resize
ntfs: do not mark the volume clean in sync_fs when errors were recorded
ntfs: skip free cluster decrement when rollback fails
ntfs: only count successfully cleared runs when freeing clusters
ntfs: fix kmap_local leak in write_mft_record_nolock() error paths
ntfs: return real error from ntfs_non_resident_attr_record_add()
ntfs: preserve error code in ntfs_resident_attr_record_add()
...
|
||
|
|
89a312991d |
SMB client fixes for v7.3-rc2
A batch of bug fixes for the SMB client:
- Fixes for fallocate range operations (insert, collapse, zero, punch
hole): the insert range implementation copied overlapping chunks in
the wrong direction, corrupting file data on every server except
Windows. Several related issues in the same area are also
addressed — stale page cache and FS-Cache readback, an integer
truncation on large files, missing RLIMIT_FSIZE validation and
missing sparse file marking.
- Data corruption fixes in the O_TRUNC open path: one where i_size
was zeroed before the server confirmed the truncate and another
where the lack of locking allowed concurrent buffered writes to be
silently discarded.
- Heap overflow fixes in legacy SMB1 paths: one in extended attribute
writes and one in POSIX ACL handling, both exploitable via
unprivileged setxattr(2).
- Fix for multiuser mount with krb5 failing because the username
option was not propagated to new per-user connections.
- Fix for split debug message in __release_mid() after a printk
conversion.
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQTcqRusfSdYROJQwGkpVtNKoQNdYwUCapcv6wAKCRApVtNKoQNd
Y0dgAQDvlnpdCsg1SZZN7T/wSy08fP7GEl2lUoCb8m6LQHlDlwEA/RG+FeY8sRkb
iJexqIGT85a48SHmpSavzBO3qkuqIQo=
=+Qbr
-----END PGP SIGNATURE-----
Merge tag 'cifs-fixes-7.3-rc2' of https://git.manguebit.org/linux
Pull smb client fixes from Paulo Alcantara:
- Fixes for fallocate range operations (insert, collapse, zero, punch
hole)
The insert range implementation copied overlapping chunks in the
wrong direction, corrupting file data on every server except Windows.
Several related issues in the same area are also addressed — stale
page cache and FS-Cache readback, an integer truncation on large
files, missing RLIMIT_FSIZE validation and missing sparse file
marking.
- Data corruption fixes in the O_TRUNC open path: one where i_size was
zeroed before the server confirmed the truncate and another where the
lack of locking allowed concurrent buffered writes to be silently
discarded
- Heap overflow fixes in legacy SMB1 paths: one in extended attribute
writes and one in POSIX ACL handling, both exploitable via
unprivileged setxattr(2)
- Fix for multiuser mount with krb5 failing because the username option
was not propagated to new per-user connections
- Fix for split debug message in __release_mid() after a printk
conversion
* tag 'cifs-fixes-7.3-rc2' of https://git.manguebit.org/linux:
smb: client: reject SetEA requests that do not fit the request buffer
smb: client: fix data corruption with concurrent writes and O_TRUNC
cifs: don't update i_size in cifs_do_truncate without a cached handle
smb: client: fix heap overflow in cifs_do_set_acl()
smb: client: fix multiuser mount with krb5
smb: client: transport: Fix debug printing in __release_mid()
smb/client: invalidate fscache for fallocate range operations
smb/client: fix stale page cache in insert/collapse range
smb/client: fix integer truncation in collapse range
smb/client: fix data corruption in emulated insert range
smb/client: mark file sparse before emulating insert range
smb/client: validate new EOF for zero range
smb/client: validate new EOF for insert range
cifs: add revalidation on FSCTL failure in smb2_duplicate_extents()
|
||
|
|
9a58da8005 |
- Prevent unintended data exposure by clearing pipe compound padding and
the response buffer.
- Initialize missing fields in FS_OBJECT_ID_INFORMATION,
FS_CONTROL_INFORMATION, and FS_POSIX_INFORMATION.
- Propagate DACL parsing and allocation failures so malformed security
descriptors are rejected.
- Rate-limit errors for unmapped SIDs to prevent kernel log flooding.
- Drain multichannel sessions during LOGOFF, wake deferred locks and
cancellable requests, and ensure cancellation callbacks run only once.
- Fix listener kthread reference handling and teardown ordering during
netdevice events.
- Validate normalized-name and IPC share configuration response lengths.
- Update the KSMBD MAINTAINERS entry and add Paulo Alcantara as
an SMBDIRECT co-maintainer.
-----BEGIN PGP SIGNATURE-----
iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqWnz0WHGxpbmtpbmpl
b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCJK8EACCE2K2p9CH6kiy9VnMjEqTbIBF
ZRCmxrspoPAMuTbK6529dXHUVTsXlUdJ/FVzGwNLtvXwEIVjNaQDqBEFWCdPElE+
8grKsC1S3gH3t8Z1wT6eNh5cpDoA+rWJDbNK4DsmHdoVagyjd9dd7fkMi7nq0WJS
NO7BTHaTuTaZDul8UXc1gqkVLviZZWkrtkGVVnsJV1z5cFls6P81cVmtzP0836cU
kVDYSI0EZnX+1P5CtOxL3r5LDBex6lRHU+rj1ypJRJDM2nR+bYIeJk+XMjylKCHT
liPj7dwI/ptVzp+n3dbcTyhLZayDhZ0/GeJanX2/midtiNSKhao9h94BymPU91jV
JugPlkAO8Vqwo7xojWRqudz4Kg/vgr66NexQ/3W2tuRXXFN4kEWmQG0N5+kH0K3d
sJ5xA9uLj24+d29fjylkdSGpuRLR8XcR01he2CaqLRopXZrCxFChwzZwbads1rI/
kXtYrORB0u99ScwTRQeW90dzeZ+1R3aHOyf8H86zyJ07l2NxG8t5L/49vuaiaiEZ
5r4hhPVumlmDQdPoOcugOmkJL68+W4TzS7UfcOSgq43W31BE1dYsaX7u4MHNk/Nq
UiWJJPArJ3ry8e4GLqQXx4ylZJnykGS9s676gFCO1GAI+eXoQ4R0k7RCKy/TW0ap
t3oa0Yj+Y4QAZCD3Rw==
=UL/J
-----END PGP SIGNATURE-----
Merge tag 'ksmbd-for-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb
Pull smb server fixes from Namjae Jeon:
- Prevent unintended data exposure by clearing pipe compound padding
and the response buffer
- Initialize missing fields in FS_OBJECT_ID_INFORMATION,
FS_CONTROL_INFORMATION, and FS_POSIX_INFORMATION
- Propagate DACL parsing and allocation failures so malformed security
descriptors are rejected
- Rate-limit errors for unmapped SIDs to prevent kernel log flooding
- Drain multichannel sessions during LOGOFF, wake deferred locks and
cancellable requests, and ensure cancellation callbacks run only once
- Fix listener kthread reference handling and teardown ordering during
netdevice events
- Validate normalized-name and IPC share configuration response lengths
- Update the KSMBD MAINTAINERS entry and add Paulo Alcantara as an
SMBDIRECT co-maintainer
* tag 'ksmbd-for-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb:
ksmbd: validate normalized name response length
ksmbd: fix listener task lifetime on netdev events
ksmbd: prevent out-of-bounds reads in share config responses
ksmbd: rate limit unmapped SID errors
ksmbd: propagate DACL parsing errors
ksmbd: zero pipe read compound padding
ksmbd: safely drain sessions during logoff
MAINTAINERS: Update the KSMBD entry
MAINTAINERS: Add Paulo Alcantara as an SMBDIRECT co-maintainer
ksmbd: fill in FileSysIdentifier in FS_POSIX_INFORMATION
ksmbd: initialize FileSystemControlFlags in FS_CONTROL_INFORMATION
ksmbd: zero the FS_OBJECT_ID_INFORMATION buffer before filling it in
|
||
|
|
a7f25dc23f |
xfs: fixes for 7.3-rc2
Signed-off-by: Carlos Maiolino <cem@kernel.org> -----BEGIN PGP SIGNATURE----- iJUEABMJAB0WIQSmtYVZ/MfVMGUq1GNcsMJ8RxYuYwUCapUNuQAKCRBcsMJ8RxYu Y/QVAX9SDXNSP3dw04wAuYgwSH5Ftm+WAnwusAsSvJkQdTvU0nEpAHyjb6WokS5a EbOGy5UBfRyqJFOmOw6wF5Ax0Aoxrt+lN8CuoDoh6aEhtYlh0jvd50ustYX8QSas W2R9B6IFIw== =JWP4 -----END PGP SIGNATURE----- Merge tag 'xfs-fixes-7.3-rc2' of gitolite.kernel.org:/pub/scm/fs/xfs/xfs-linux Pull xfs fixes from Carlos Maiolino: "This contains a few fixes for the zoned storage support, a possible deadlock vector fix, some code refactoring patches and a quota evasion fix on XFS while exporting it via NFS. Please note that for the quota evasion fix, a couple patches for the capability subsystem are included in the pull request. Those have been ack'ed by the respective maintainer which also agreed to have them going through the xfs tree. This also includes a patch for the quota subsystem to stop issuing audit messages during quota enforcing. Quota maintainer also ack'ed and agreed with this going through xfs tree" * tag 'xfs-fixes-7.3-rc2' of gitolite.kernel.org:/pub/scm/fs/xfs/xfs-linux: capability: unexport has_capability_noaudit xfs: replace ns_capable_noaudit quota: Don't issue audit messages on quota enforcing capability: Add new capable_noaudit xfs: fix capability check in xfs xfs: restore bi_bdev in xfs_zone_gc_write_chunk xfs: split ioend handling into a separate source file xfs: factor out a xfs_iomap_set_anon_write helper xfs: fix zoned write iomap flags assignments xfs: fix racy open zone caching xfs: handle NULL open_zone for merged ioends in xfs_ioend_put_open_zones xfs: use inode_init_always_gfp with __GFP_NOFAIL in xfs_inode_alloc xfs: remove kmem_to_page() xfs: don't flush and invalidate internal RT device twice in xfs_shutdown_devices xfs: split an assert in xfs_trans_log_buf xfs: don't hold buffer locks across sync transaction commit in xfs_sync_sb_buf |
||
|
|
4aa2c106ae |
smb: client: reject SetEA requests that do not fit the request buffer
CIFSSMBSetEA() copies the caller's extended attribute value into the
SMB request buffer without checking that it fits. The requirement is
stated in the source but was never implemented:
/*BB add length check to see if it would fit in
negotiated SMB buffer size BB */
/* if (ea_value_len > buffer_size - 512 (enough for header)) */
if (ea_value_len)
memcpy(parm_data->list.name + name_len + 1,
ea_value, ea_value_len);
The only bound applied on the way in is in cifs_xattr_set():
#define MAX_EA_VALUE_SIZE CIFSMaxBufSize
...
if (size > MAX_EA_VALUE_SIZE)
CIFSMaxBufSize is the full payload capacity of the buffer, so a value
of exactly that size leaves no room for the SMB header, the TRANS2
parameter block, the fealist header and the EA name that are written
ahead of it in the same object.
SendReceive() already enforces the correct limit on this very length:
if (in_len > CIFSMaxBufSize + MAX_CIFS_HDR_SIZE)
but it is called after the copy has taken place. An unprivileged
setxattr(2) on an SMB1 mount with a 250-byte name and a 16384-byte
value writes 16384 bytes starting 345 bytes into a 16588-byte
cifs_request object, ending 141 bytes past it:
BUG: KASAN: slab-out-of-bounds in CIFSSMBSetEA+0xabc/0xde0
Write of size 16384 at addr ffff888003aa0159 by task init/68
__asan_memcpy+0x3c/0x60
CIFSSMBSetEA+0xabc/0xde0
cifs_xattr_set+0xd3a/0xff0
__vfs_setxattr+0x13e/0x1a0
The buggy address is located 345 bytes inside of
allocated 16588-byte region
Apply SendReceive()'s limit to the assembled request before the copy
rather than after it, and widen the byte counters so the sum cannot
wrap before it is tested.
byte_count is also tested against U16_MAX, because it is stored in the
16-bit pSMB->ByteCount. That becomes reachable when CIFSMaxBufSize is
raised at module load, where it may be set as high as 1024*127: with a
5-byte EA name and a 65521-byte value, count is exactly U16_MAX while
byte_count is 65556, and cpu_to_le16() would truncate it to 20 and
transmit a frame whose ByteCount does not match its length. Testing
byte_count covers count as well, since byte_count is the larger of the
two and count's only 16-bit consumer is written after this point.
check_add_overflow() is evaluated first so that total_len is assigned
before it is reported.
Fixes:
|
||
|
|
a8603b52b3 |
smb: client: fix data corruption with concurrent writes and O_TRUNC
cifs_do_truncate() flushes dirty pages with filemap_write_and_wait()
and truncates the file on the server, but in the old code both
operations ran without holding i_rwsem or invalidate_lock. A
concurrent buffered write via netfs_perform_write() -- which only
needs i_rwsem shared -- could dirty new pages after the flush but
before the local truncation, and those pages would be silently
discarded by cifs_setsize() -> truncate_pagecache().
Fix by acquiring inode_lock (exclusive i_rwsem) and
filemap_invalidate_lock at the top of cifs_do_truncate(), so the
entire flush-truncate-resize sequence is atomic with respect to:
- buffered writes (blocked by exclusive i_rwsem, since
netfs_start_io_write takes i_rwsem shared),
- read page faults (blocked by exclusive invalidate_lock, since
filemap_fault takes it shared),
- writeback collection (blocked by netfs_wb_begin/netfs_wb_end
around the server truncate and local resize, since
netfs_writepages also acquires the wb lock).
Fixes:
|
||
|
|
0fecc393f2 |
ntfs: take invalidate_lock in ntfs_filemap_page_mkwrite()
ntfs_filemap_page_mkwrite() calls iomap_page_mkwrite() without holding
mapping->invalidate_lock, so a concurrent truncate or fallocate can be
in the middle of invalidating pagecache and rewriting the runlist while
the write fault maps blocks and dirties the folio. This races with
ntfs_attr_fallocate(), which merges clusters into the in-memory
runlist, drops the runlist lock, and only afterwards zeroes the newly
allocated clusters on disk; and with the punch-hole/insert/collapse
paths that free clusters after truncating the cache.
Per Documentation/filesystems/locking.rst, ->page_mkwrite() must ensure
there are no truncate/invalidate races, "usually mapping->invalidate_lock
is suitable for proper serialization". xfs takes its mmaplock (= the
invalidate_lock rwsem) shared in exactly this path.
Take invalidate_lock shared around iomap_page_mkwrite(). The read-only
fault path is already covered because filemap_fault() itself grabs
invalidate_lock shared on instantiation/read paths; only page_mkwrite
was bypassing it in this driver.
Fixes:
|
||
|
|
9cc5761b8f |
ntfs: take invalidate_lock in ntfs_setattr_size()
ntfs_setattr_size() updates i_size and resizes the on-disk attribute
without holding mapping->invalidate_lock. Page faults take the lock
shared, so a fault racing the resize can resolve a VCN against the
transient runlist state of ntfs_non_resident_attr_expand() and fail
with a spurious SIGBUS, and can interleave with the size-change
epilogue (truncate_pagecache(), i_size_write(),
pagecache_isize_extended()).
Take invalidate_lock exclusively around the whole resize after
inode_dio_wait(), matching the fallocate path and other filesystems
such as xfs, which wraps truncate in its mmaplock (= invalidate_lock).
Fixes:
|
||
|
|
4dc8f4ee2d |
ntfs: handle signal interruption in fallocate
The ntfs_attr_fallocate() function checks for pending signals during
allocation loops and exits early via 'out' label. However, when a signal
interrupts the operation with err == 0, the function returns 0 (success)
instead of -EINTR.
The signal_pending() checks at the allocation loops jump to 'out' without
setting err = -EINTR, so the function returns success even when interrupted
by a signal.
Set err = -EINTR when jumping to the signal exit path, and only override
when no other error is pending. This ensures:
- Allocation interrupted by signal returns -EINTR
- Allocation that completed successfully before signal arrived returns 0
- Other errors are preserved and not overwritten by -EINTR
Fixes:
|
||
|
|
ba9572bc43 |
ksmbd: validate normalized name response length
FILE_NORMALIZED_NAME_INFORMATION converts the open file path to UTF-16.
smb2_allocate_rsp_buf() leaves these responses in the 448-byte small
buffer, and get_file_normalized_name_info() converts the path without
checking the remaining space.
An authenticated client can query a long path and make
smbConvertToUTF16() write beyond work->response_buf.
Use the large response buffer for normalized-name queries. Before
conversion, verify that the response has room for the worst-case UTF-16
output and its terminator.
Fixes:
|
||
|
|
a506290f59 |
ksmbd: fix listener task lifetime on netdev events
The listener thread exits when its listening socket is shutdown. The
netdevice notifier shuts down the socket before calling kthread_stop(), so
the task_struct can be freed before kthread_stop() gets its reference.
Create the listener in a stopped state and hold an extra task_struct
reference until kthread_stop_put() completes. Also stop and release
listeners before freeing their interface records during TCP teardown.
Fixes:
|
||
|
|
f25e93768f |
ksmbd: prevent out-of-bounds reads in share config responses
Validate IPC share configuration payload sizes before consuming
variable-length fields. Bound veto list parsing and account for
the separator byte when deriving the path length.
Fixes:
|
||
|
|
feca5e70fc |
ksmbd: rate limit unmapped SID errors
A client can include many structurally valid but unmapped SIDs in a DACL.
Logging every mapping failure lets one request generate hundreds of kernel
error messages.
Rate limit the message to prevent an authenticated client from flooding
the kernel log.
Fixes:
|
||
|
|
c61dc7b1b4 |
ksmbd: propagate DACL parsing errors
parse_dacl() silently accepts truncated ACEs and allocation failures,
allowing set_info_sec() to continue with an incomplete ACL conversion.
Return parsing and allocation errors to parse_sec_desc() so malformed
security descriptors are rejected before inode attributes or ACL xattrs
are updated.
Fixes:
|
||
|
|
73f860489e |
ksmbd: zero pipe read compound padding
Compound response handling extends the last response iov to an eight-byte
boundary.
smb2_read_pipe() allocates only the payload size, so the alignment padding
can expose up to seven bytes of uninitialized kernel heap memory.
Allocate the aligned size and clear the unused tail before pinning the
response buffer.
Fixes:
|
||
|
|
d12168084c |
ksmbd: safely drain sessions during logoff
SMB3 multichannel allows requests for one session to run on multiple
connections. Wait for all channels bound to a session before freeing
shared session objects.
A deferred byte-range lock remains counted as a running request and only
wakes when its file closes. Wake blocked locks during the drain without
unpublishing or modifying their file objects. Synchronous CANCEL requests
must invoke their cancellation callback to wake pending operations, while
CHANGE_NOTIFY completion remains specific to the asynchronous path.
Serialize session teardown with channel registration and previous-session
cleanup, and use atomic work-state transitions so LOGOFF, CANCEL, and
connection teardown invoke cancellation callbacks only once.
Fixes:
|
||
|
|
db2267b27c |
ksmbd: fill in FileSysIdentifier in FS_POSIX_INFORMATION
smb2_get_info_filesystem() reports 56 bytes for FS_POSIX_INFORMATION,
that is the whole of FILE_SYSTEM_POSIX_INFO, but never assigns
FileSysIdentifier. Those eight bytes go to the client as they are found
in the response buffer.
The buffer is zeroed on allocation, so a standalone request leaks
nothing. A compound request can leak: the offset of the next response
is advanced by the length pinned for the previous one, so a reply that
was written into the buffer and then dropped in favour of the short
error response of smb2_set_err_rsp() stays there, and the next reply is
laid over it with only the header cleared.
Report the file system id statfs() returned, which is what the field is
for. FileSysIdentifier is __le64 and f_fsid is a pair of ints, so
assemble the value first, val[0] as the low half, and convert it on the
way out.
Fixes:
|
||
|
|
c0cd3fc682 |
ksmbd: initialize FileSystemControlFlags in FS_CONTROL_INFORMATION
smb2_get_info_filesystem() reports 48 bytes for FS_CONTROL_INFORMATION,
that is the whole of struct smb2_fs_control_info, but never assigns
FileSystemControlFlags. Those four bytes go to the client as they are
found in the response buffer.
The buffer is zeroed on allocation, so a standalone request leaks
nothing. A compound request can leak: the offset of the next response
is advanced by the length pinned for the previous one, so a reply that
was written into the buffer and then dropped in favour of the short
error response of smb2_set_err_rsp() stays there, and the next reply is
laid over it with only the header cleared.
ksmbd does not implement quota tracking, so report no control flags.
Fixes:
|
||
|
|
399aa12450 |
ksmbd: zero the FS_OBJECT_ID_INFORMATION buffer before filling it in
smb2_get_info_filesystem() reports 64 bytes for FS_OBJECT_ID_INFORMATION,
that is the whole of struct object_id_info, but writes only 46 of them:
- objid[] is 16 bytes, and when the volume UUID is not available only
sizeof(stfs.f_fsid) (8) bytes are copied into it;
- extended_info.version_string[] is STRING_LENGTH (28) bytes, and only
strlen("1.1.0") (5) bytes are copied into it.
The response buffer is zeroed on allocation (kvzalloc() in
smb2_allocate_rsp_buf()), so for a standalone request the remaining 31
bytes are zero. In a compound request they need not be. The offset of
the next response is advanced by the length pinned for the previous one,
so if a preceding command wrote its reply into the buffer and then
failed, smb2_set_err_rsp() pins only the short error response and the
next reply lands inside the area that has already been written. Only
the header is cleared there:
memset((char *)rsp_hdr, 0, sizeof(struct smb2_hdr) + 2);
The client then receives up to 31 bytes of a response it was not meant
to see, including one that failed with an access denied error.
Clear the structure before filling it in. As a side effect
version_string is now NUL terminated.
Fixes:
|
||
|
|
fe39cd9d48 |
cifs: don't update i_size in cifs_do_truncate without a cached handle
If find_writable_file() returns null, cifs_file_flush will return
0 without issuing set_file_size, and the outer 'if (!rc)' block
will set i_size to 0 before telling the server to truncate. If
the cifs_open() then fails, the inode will have size 0, while
the server file is unchanged.
Move the netfs_resize_file() and cifs_setsize() into the 'if
(cfile)', so they only run after a successful set_file_size.
In the no-handle else branch, evict stale pages with
truncate_inode_pages before the O_TRUNC open to dispose of old
cache pages, and let the open response set the i_size.
Fixes:
|
||
|
|
1dac61e2c2 |
smb: client: fix heap overflow in cifs_do_set_acl()
cifs_set_acl() validates ACL size using posix_acl_xattr_size():
4 + (count * 8) // 4-byte header + 8 bytes per ACE
cifs_do_set_acl() then calls posix_acl_to_cifs() to write the CIFS
wire format into the same buffer:
6 + (count * 10) // 6-byte header + 10 bytes per ACE
An ACL that passes the xattr-based check in cifs_set_acl() can
overflow the heap when posix_acl_to_cifs() writes the larger CIFS
format.
Validate the CIFS format size against the remaining buffer space and
USHRT_MAX before converting--data_count is __u16, so sizes above
USHRT_MAX truncate the on-wire packet length, causing the server to
apply a partial ACL. Replace MaxDataCount = 1000 with
min(CIFSMaxBufSize, USHRT_MAX).
Fixes:
|
||
|
|
6949939586 |
smb: client: fix multiuser mount with krb5
Customer reported that they could no longer mount their SMB shares
with multiuser mount option and krb5. Turned out that the client
wasn't duplicating username option when creating multiuser
connections, therefore failing to retrieve credentials as
cifs.upcall(8) couldn't find them in keytab.
Fix this by duplicating username option (if set) from original fs
context before creating multiuser connections with krb5.
Reproducer:
```
$ ktutil
ktutil: add_entry -password -p testuser -k 1 -e aes256-cts
Password for testuser@ZELDA.TEST:
ktutil: write_kt /etc/krb5.keytab
ktutil: quit
$ klist -ke
Keytab name: FILE:/etc/krb5.keytab
KVNO Principal
---- ----------------------------------------------------------------
1 testuser@ZELDA.TEST (aes256-cts-hmac-sha1-96)
$ mount.cifs //w22-root2/scratch /mnt/1 -o \
uid=1000,sec=krb5,username=testuser@ZELDA.TEST,multiuser
mount error(13): Permission denied
Refer to the mount.cifs(8) manual page (e.g. man mount.cifs) and
kernel log messages (dmesg)
```
Reported-by: Jacob Shivers <jshivers@redhat.com>
Fixes:
|
||
|
|
d83a21bb26 |
smb: client: transport: Fix debug printing in __release_mid()
Long time ago during upgrading printk():s to the respective pr_<level>()
calls one misconversion happened and nobody has noticed that. So,
previously printk(KERN_DEBUG) + printk() worked as one long debug print
since the trailing '\n' is only present in the followup printk() format
string. The culprit change missed that and split the message to two on
the different levels. Restore the original behaviour to make users be
less confused in the most likely never happen cases of partially getting
that message.
Fixes:
|
||
|
|
448ba0ae65 |
smb/client: invalidate fscache for fallocate range operations
smb3_zero_range(), smb3_punch_hole(), smb3_insert_range(), and
smb3_collapse_range() modify file contents through server-side range
operations. These operations discard the affected page cache, but leave
the FS-Cache cookie valid, so a later read may return data cached before
the range operation.
Fix this by invalidating FS-Cache after outstanding I/O has completed
and before modifying the file on the server.
Run the following as root on a CIFS mount with fsc enabled and an active
CacheFiles backend:
bash -c '
MNT=/mnt/cifs
FILE="$MNT/repro"
# Generate four 1 MiB random blocks: [A][B][C][D].
dd if=/dev/urandom of=/tmp/src bs=1M count=4 status=none
# Expected contents after zeroing B: [A][zero][C][D].
cp /tmp/src /tmp/expected
dd if=/dev/zero of=/tmp/expected bs=1M seek=1 count=1 \
conv=notrunc status=none
cp /tmp/src "$FILE"
# Populate FS-Cache, then discard the page cache.
sync
echo 1 > /proc/sys/vm/drop_caches
cat "$FILE" > /dev/null
sync
echo 1 > /proc/sys/vm/drop_caches
fallocate --zero-range -o 1M -l 1M "$FILE"
if cmp -s /tmp/expected "$FILE"; then
echo "readback: OK"
else
echo "readback: STALE DATA"
fi
'
Before this change, the readback differs from /tmp/expected:
readback: STALE DATA
After this change, it matches:
readback: OK
Fixes:
|
||
|
|
01261a6fa4 |
smb/client: fix stale page cache in insert/collapse range
smb3_insert_range() and smb3_collapse_range() use
truncate_pagecache_range() to invalidate the affected page cache.
However, if off or old_eof is not page-aligned, the boundary pages are
only partially zeroed and remain uptodate. As a result, the client may
return stale data after a successful insert/collapse range operation.
For example, with 4K pages:
page 0 page 1 page 2
0------4K 4K------8K 8K------12K
^ ^
off=2K old_eof=10K
Page 1 is removed from the page cache, while the boundary pages are
only partially zeroed. After COPYCHUNK moves the data on the server,
these cached pages may still return stale data.
This can be reproduced on a CIFS mount:
bash -c '
FILE=/mnt/scratch/repro
# Use a 6 KiB file so EOF is not page-aligned.
dd if=/dev/urandom of=/tmp/src bs=1K count=6 status=none
# Expected: a 4 KiB hole followed by the original data.
rm -f /tmp/expected
truncate -s 4K /tmp/expected
cat /tmp/src >> /tmp/expected
cp /tmp/src "$FILE"
# Prime the page cache before moving data on the server.
cat "$FILE" > /dev/null
fallocate --insert-range -o 0 -l 4K "$FILE"
if cmp -s /tmp/expected "$FILE"; then
echo "readback: OK"
else
echo "readback: STALE DATA"
fi
'
Fix this by writing back dirty data and discarding the page cache from
the start of the page containing off to EOF before moving data on the
server.
Fixes:
|
||
|
|
7811701d6a |
smb/client: fix integer truncation in collapse range
smb3_collapse_range() stores the ssize_t return value of
smb2_copychunk_range() in an int. A successful copy larger than
INT_MAX is truncated to a negative value and treated as an error.
Reproducer:
MNT=/mnt/scratch
truncate -s 2056M "$MNT/file"
fallocate --collapse-range -o 1M -l 1M "$MNT/file"
Fix this by using __smb2_copychunk_range(), which reports success as
zero instead of returning the copied byte count.
Before this change, the reproducer fails with:
fallocate: fallocate failed: Success
and the file size remains unchanged at 2056 MiB. After this change, the
reproducer succeeds and the file size becomes the expected 2055 MiB.
Fixes:
|
||
|
|
0923ae9f23 |
smb/client: fix data corruption in emulated insert range
smb3_insert_range() shifts [off, EOF) right with COPYCHUNK, copying from
low to high offsets. When the ranges overlap, the copy can overwrite
source data that has not yet been copied. For a 1 MiB insert at offset 0:
offset: 0 1M 2M 3M 4M 5M
before: | A | B | C | D |
expected: | hole | A | B | C | D |
current: | hole | A | A | A | A | (corrupted)
Let x be the insertion offset, L the total length to move, delta the
insert length, and C the normal chunk size allowed by the server.
Insert range maps
[x, x + L) -> [x + delta, x + delta + L).
When delta >= L, the complete source and target ranges are disjoint, so
the normal copy order and chunk size are safe:
offset: 0 4 8 12 16 20 24 28 32
source: [--S0--][--S1--][--S2--][--S3--]
target: [--T0--][--T1--][--T2--][--T3--]
When delta < L, the complete source and target ranges overlap, so the
copy must proceed from EOF backwards. There are two subcases.
If delta >= C, each corresponding source and target chunk is disjoint.
The 1 MiB example has L = 4 MiB and delta = C = 1 MiB:
offset: 0 1M 2M 3M 4M 5M
source: [--S0--][--S1--][--S2--][--S3--]
target: [--T0--][--T1--][--T2--][--T3--]
Copying S0 from [0, 1M) to [1M, 2M) overwrites S1 before it is copied.
Processing chunks from EOF backwards prevents this inter-chunk
overwrite.
If delta < C, the source and target ranges of a normal chunk also
overlap. For example, with L = 16, delta = 2 and C = 4:
offset: 0 2 4 6 8 10 12 14 16 18
source: [--S0--][--S1--][--S2--][--S3--]
target: [--T0--][--T1--][--T2--][--T3--]
Here S0 and T0 overlap over [2,4), S1 and T1 over [6,8), and so on.
Backward ordering cannot control how the server copies bytes inside one
descriptor, so the chunk size must be limited to delta.
Fix this by copying overlapping right shifts from EOF backwards. Limit
the chunk size to delta when delta < C so that each chunk's source and
target ranges do not overlap. Using larger chunks would require a way to
identify servers that safely handle overlapping COPYCHUNK descriptors.
Therefore:
delta >= L:
keep the normal copy order and chunk size
delta < L:
delta >= C: copy backwards and keep the normal chunk size
delta < C: copy backwards and limit the chunk size to delta
Only the delta < C subcase requires reducing the chunk size for data
integrity.
Reproducer:
bash -c '
MNT=/mnt/scratch
# Generate four 1 MiB random blocks: [A][B][C][D].
dd if=/dev/urandom of=/tmp/src bs=1M count=4 status=none
# With C = 1 MiB, test delta = C and delta < C.
for delta in 1M 1K; do
truncate -s 0 /tmp/expected
truncate -s "$delta" /tmp/expected
cat /tmp/src >> /tmp/expected
cp /tmp/src "$MNT/file"
fallocate --insert-range -o 0 -l "$delta" "$MNT/file"
if cmp -s /tmp/expected "$MNT/file"; then
echo "delta=$delta: OK"
else
echo "delta=$delta: CORRUPTED"
fi
done
'
The corruption reproduces with Samba and ksmbd, while Windows handles
the overlapping COPYCHUNK ranges safely.
The 1 MiB case tests delta >= C, while the 1 KiB case tests delta < C.
Before this change, the reproducer reports:
delta=1M: CORRUPTED
delta=1K: CORRUPTED
After this change, it passes against both ksmbd and Samba:
delta=1M: OK
delta=1K: OK
Fixes:
|
||
|
|
cd03ce4950 |
smb/client: mark file sparse before emulating insert range
The SMB client emulates FALLOC_FL_INSERT_RANGE with SET_EOF, COPYCHUNK
and SET_ZERO_DATA.
SET_ZERO_DATA creates a hole only when the file is sparse. On a
non-sparse file, it clears the inserted range but leaves its blocks
allocated, causing the extent count check in xfstests generic/064 to
fail.
Fix this by marking the file sparse before modifying it.
This patch produces the expected sparse extents in xfstests generic/064
only when the server-reported block size is compatible with the server's
deallocation granularity.
For ksmbd, the reported block size follows the backing filesystem,
and the test passes. For Samba, the test passes with a block size
matching the backend granularity, for example, 4 KiB on Btrfs, but not
with the default 1 KiB value. For Windows Server 2022, 4 KiB inserts do
not generate holes, while aligned inserts of 64 KiB or larger do.
Fixes:
|
||
|
|
88972e3575 |
smb/client: validate new EOF for zero range
When FALLOC_FL_ZERO_RANGE is used without FALLOC_FL_KEEP_SIZE,
smb3_zero_range() may extend EOF without checking RLIMIT_FSIZE, allowing
the file to grow beyond the caller's file-size limit.
Fix this by calling inode_newsize_ok() before sending the zero-range
request when the operation would extend EOF.
Reproducer, using a file on a CIFS mount:
bash -c '
FILE=/mnt/cifs/repro
trap "" SIGXFSZ
ulimit -f 3072
truncate -s 2M "$FILE"
fallocate --zero-range -o 0 -l 4M "$FILE"
echo "fallocate rc=$?"
stat -c "file size=%s" "$FILE"
'
Before this change, the operation succeeds despite the 3 MiB limit:
fallocate rc=0
file size=4194304
After this change, fallocate fails and leaves the file at 2 MiB.
Fixes:
|
||
|
|
1519dc88c8 |
smb/client: validate new EOF for insert range
smb3_insert_range() does not check if the new file size
(i_size + len) is valid. This allows FALLOC_FL_INSERT_RANGE to bypass
RLIMIT_FSIZE, exceed s_maxbytes, or produce a size outside the loff_t
range.
Use check_add_overflow() to calculate the new EOF. Validate it with
inode_newsize_ok() before modifying the file.
Reproducer, using a file on a CIFS mount:
bash -c '
FILE=/mnt/cifs/repro
trap "" SIGXFSZ
ulimit -f 3072 # RLIMIT_FSIZE = 3 MiB
# A regular write is stopped at 3 MiB.
dd if=/dev/zero of="$FILE" bs=1M count=4 status=none
stat -c "size after write: %s" "$FILE"
# Insert 2 MiB into a 2 MiB file.
truncate -s 2M "$FILE"
fallocate -i -o 0 -l 2M "$FILE"
stat -c "size after insert: %s" "$FILE"
'
Before this change, the regular write stops at the 3 MiB limit, but
insert range grows the file to 4 MiB:
dd: error writing '/mnt/cifs/repro': File too large
size after write: 3145728
size after insert: 4194304
After this change, insert range also fails at the limit and leaves the
2 MiB file unchanged:
dd: error writing '/mnt/cifs/repro': File too large
size after write: 3145728
fallocate: fallocate failed: File too large
size after insert: 2097152
Fixes:
|
||
|
|
034dd340b0 |
tracing fixes for v7.3:
- Fix error output of boot instance creation failure Currently if a boot instance creation fails, instead of printing out the name of the instance that failed, it prints "(null)". That is because it prints "cur_str" that had already been processed by strsep(). Print the saved name instead. While at it, print the error code of the failure. - Fix use-after-free for same named historgrams Histograms can be named so that they can be used in multiple events. But if the named histogram has a variable attached, the second event that uses the named histogram which duplicates it and needs to free the original after duplication leaves the old variable in place and still visible. If another histogram uses than variable, it will use the stale one which will try to reference the freed duplicate histogram and crash the kernel. Free the duplicate variables along with the duplicated histogram data. - Check return value of kthread_run() in event self test The events self tests uses a kthread for testing but does not check if it succeeded in creating a kthread. If the kthread creation were to fail, the code will still try to call kthread_stop() on the error returned. - Fix race between reading trace_pipe and updating subbuffer size If a user is reading the trace_pipe file at the same time they update the ring buffer sub-buffer size, can cause the trace_pipe read to read stale data. Add trace_access_lock() around updating the ring buffer sub-buffer size. - Fix eventfs_inode on failure path in creation of the events directory In the creation of the "events" directory, if after allocating the eventfs_inode a failure is detected, it calls cleanup_ei() which calls free_ei(). The free_ei() will test if eventfs_inode being freed has no children. It is a bug if it does. But on the failure case of the creation of the "events" directory, the children lists have not yet been initialized and the free will trigger a warning because list_empty() on an uninitialized list returns false. Move the initialization into init_ei() where it makes more sense and makes sure that a created eventfs_inode has its lists initialized upon creation. - Check return value of kthread_run() in ftrace direct sample code The sample code that shows how to use the ftrace direct calls does not test the return of kthread_run() to see if it succeeds. Return a failure if the kthread_run() doesn't succeed. - Clear user events state on fork in case of alloc failure On fork, the child gets a pointer to the parent's user events state. It makes a copy of it then updates the child's pointer to it. But if the allocation fails, the duplication function leaves the child with a pointer to its parent's descriptor. When the child cleans up its data, it will free the parent's descriptor while the parent is still using it. In the duplication function, set the child's user_event_mm to NULL before testing if the allocation succeeded, and when it exits it will not free the parent's descriptor. - Fix retry exhaustion in simple ring buffer reader swap simple_ring_buffer_swap_reader_page() starts with retry set to 8 and post-decrements it only after a failed link replacement. On the final attempt, a successful replacement leaves retry at zero, while a failed replacement leaves it at -1. But the check for success expects the retry value to be non-zero and exits with an error on zero. This is the opposite result. Fix it. - Fail nicely when the remote swap_reader_page() returns an error Currently, if the swap_reader_page() of a remote buffer fails, it triggers a WARN_ON_ONCE() and continues normally. Instead, have it exit with an error and a pr_warn() print instead of a full WARNING. -----BEGIN PGP SIGNATURE----- iIoEABYKADIWIQRRSw7ePDh/lE+zeZMp5XQQmuv6qgUCapOC3hQccm9zdGVkdEBn b29kbWlzLm9yZwAKCRAp5XQQmuv6qvjkAQCGVuyK980rwiBnfenWLpeB3QjfHA8B mV0mJSlGWm1t1gEA9WWzMGbp+OHeRV2xyA+xW7OS1S58VO9OIGrzXCGqbAM= =TrF5 -----END PGP SIGNATURE----- Merge tag 'trace-v7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace Pull tracing fixes from Steven Rostedt: - Fix error output of boot instance creation failure Currently if a boot instance creation fails, instead of printing out the name of the instance that failed, it prints "(null)". That is because it prints "cur_str" that had already been processed by strsep(). Print the saved name instead. While at it, print the error code of the failure. - Fix use-after-free for same named historgrams Histograms can be named so that they can be used in multiple events. But if the named histogram has a variable attached, the second event that uses the named histogram which duplicates it and needs to free the original after duplication leaves the old variable in place and still visible. If another histogram uses than variable, it will use the stale one which will try to reference the freed duplicate histogram and crash the kernel. Free the duplicate variables along with the duplicated histogram data. - Check return value of kthread_run() in event self test The events self tests uses a kthread for testing but does not check if it succeeded in creating a kthread. If the kthread creation were to fail, the code will still try to call kthread_stop() on the error returned. - Fix race between reading trace_pipe and updating subbuffer size If a user is reading the trace_pipe file at the same time they update the ring buffer sub-buffer size, can cause the trace_pipe read to read stale data. Add trace_access_lock() around updating the ring buffer sub-buffer size. - Fix eventfs_inode on failure path in creation of the events directory In the creation of the "events" directory, if after allocating the eventfs_inode a failure is detected, it calls cleanup_ei() which calls free_ei(). The free_ei() will test if eventfs_inode being freed has no children. It is a bug if it does. But on the failure case of the creation of the "events" directory, the children lists have not yet been initialized and the free will trigger a warning because list_empty() on an uninitialized list returns false. Move the initialization into init_ei() where it makes more sense and makes sure that a created eventfs_inode has its lists initialized upon creation. - Check return value of kthread_run() in ftrace direct sample code The sample code that shows how to use the ftrace direct calls does not test the return of kthread_run() to see if it succeeds. Return a failure if the kthread_run() doesn't succeed. - Clear user events state on fork in case of alloc failure On fork, the child gets a pointer to the parent's user events state. It makes a copy of it then updates the child's pointer to it. But if the allocation fails, the duplication function leaves the child with a pointer to its parent's descriptor. When the child cleans up its data, it will free the parent's descriptor while the parent is still using it. In the duplication function, set the child's user_event_mm to NULL before testing if the allocation succeeded, and when it exits it will not free the parent's descriptor. - Fix retry exhaustion in simple ring buffer reader swap simple_ring_buffer_swap_reader_page() starts with retry set to 8 and post-decrements it only after a failed link replacement. On the final attempt, a successful replacement leaves retry at zero, while a failed replacement leaves it at -1. But the check for success expects the retry value to be non-zero and exits with an error on zero. This is the opposite result. Fix it. - Fail nicely when the remote swap_reader_page() returns an error Currently, if the swap_reader_page() of a remote buffer fails, it triggers a WARN_ON_ONCE() and continues normally. Instead, have it exit with an error and a pr_warn() print instead of a full WARNING. * tag 'trace-v7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: ring-buffer: Stop remote reader update when page swap fails tracing: Fix retry exhaustion in simple ring buffer reader swap tracing/user_events: Clear copied tracing state before fork duplication samples/ftrace: Fix kthread_stop() on ERR_PTR in ftrace-direct-multi-modify samples/ftrace: Fix kthread_stop() on ERR_PTR in ftrace-direct-modify eventfs: Initialize ei->children and ei->list in init_ei() tracing: Fix use-after-free in trace_pipe read on sub-buffer order change tracing: Fix crash passing ERR_PTR to kthread_stop() tracing: Fix use-after-free with same-name named triggers tracing: Fix logged instance name on creation failure |
||
|
|
03c6ecc4b4 |
ntfs: fix FITRIM range alignment
ntfs_trim_fs() aligns the start of a free extent up to the device discard
granularity, but derives the discard length by aligning the original extent
length down. When the free extent start is not discard-aligned, adding that
length to the aligned start can extend the discard past the free extent and
into allocated clusters.
For example, with 4 KiB clusters and 32 KiB discard granularity, the free
extent [4 KiB, 36 KiB) becomes the discard range [32 KiB, 64 KiB), so
28 KiB beyond the free extent may be discarded.
Align the absolute end of the free extent down and derive the length from
the two aligned endpoints. Skip extents that contain no full discard unit.
Reproduced with a 4 KiB-cluster NTFS filesystem on scsi_debug configured
for 32 KiB discard granularity and read-zero-after-trim. Before this
change, FITRIM zeroed seven allocated 4 KiB clusters following an unaligned
32 KiB hole. With this change, the same data remains intact across FITRIM
and remount.
Fixes:
|
||
|
|
41a52ba4a5 |
ntfs: read WOF chunks outside the decompression lock
WOF decompression uses four module-global workspaces, one per compression
format, each with a static mutex. ntfs_read_wof_compressed_block() takes
that mutex once and holds it across the whole chunk loop, so both block
reads run inside it:
mutex_lock(ws->lock);
for each chunk {
parse_wof_chunk_table(..., ws->input, ...); /* reads disk */
ntfs_read_wof_chunk(..., ws->input, ...); /* reads disk */
decompress into ws->output;
}
mutex_unlock(ws->lock);
Readers of system-compressed files then serialise system-wide on the disk
waits, not just on the decompressor scratch the lock exists for. One
reader sleeping in submit_bio_wait() blocks all the rest.
The waits dominate. Reading an 8 MiB xpress4k file (2048 chunks at a 48%
compressed ratio, so 2048 acquisitions and 4096 block reads) and timing
ws->lock against the part of it spent in ntfs_bdev_read():
backing store held of that in I/O held after
virtio, host page cache 348 ms 321 ms (92%) 24.6 ms
virtio, throttled 100 MB/s 978 ms 948 ms (96%) 36.6 ms
The page-cache row is a lower bound, having no seek cost at all, and the
share still grows with slower storage because only the wait scales while
decompression stays near 26 ms.
The reads are inside the lock only because they land in ws->input, a
buffer shared through the workspace. Nothing else requires it:
parse_wof_chunk_table() and ntfs_read_wof_chunk() already take the buffer
as a parameter and both set *chunk_mem to a pointer inside it, so a
caller-owned buffer works unchanged.
Allocate that buffer per call, do both reads without the lock, and take
the lock only around decompression, which is the step needing ws->output
and ws->scratch. squashfs is arranged this way already: its
squashfs_decompress() is handed a bio that has been read, and locks only
for the CPU work.
Block reads are unchanged in number, they just no longer run under the
lock, and hold time stops tracking device speed.
This also unnests two per-inode locks from the global one, runlist->lock
taken by both reads and base_ni->mrec_lock taken for a resident stream.
A resident chunk needs no I/O at all, yet used to queue behind a reader
blocked in submit_bio_wait() and then take mrec_lock inside the global
mutex.
The buffer is 4608 bytes for xpress4k and at most 33280 for lzx32k. This
path already does GFP_NOFS allocations per call in ntfs_attr_iget(), and
in ntfs_attr_get_search_ctx() for a resident stream, so one more does not
change how it behaves under memory pressure. The workspace keeps output
and scratch, 4 KiB to 32 KiB and 6224 bytes (xpress) or 10240 (lzx), and
its "already allocated" test moves from ws->input to ws->output.
The lock is now taken per chunk rather than per call, which differs only
for a folio spanning several chunks: a few more uncontended mutex
operations in exchange for not holding it across the reads between them.
Verified under QEMU against an uncompressed copy of the same data, on an
8 MiB file and a 100000 byte one, the latter covering the tail chunk that
is not a full comp_unit.
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
|
||
|
|
548e7bcd0c |
A wide variety of mostly CephFS fixes and cleanups, split between
changes that address edge cases (Sam, Xiubo, Matthew), efficiency improvements (Max) and AI-assisted hardening (Michael, Jeremy). One thing that stands out is Alex's change to how CephFS behaves in NEARFULL scenarios: the long-standing "make all writes synchronous" behavior has become opt-in. It was always somewhat controversial and doesn't make much sense for modern deployments; the new default is to continue normal operation (i.e. buffer writes as MDS allows, etc). The behavior in case the cluster reaches any FULL state remains the same as before. -----BEGIN PGP SIGNATURE----- iQFHBAABCgAxFiEEydHwtzie9C7TfviiSn/eOAIR84sFAmqRz9QTHGlkcnlvbW92 QGdtYWlsLmNvbQAKCRBKf944AhHzi/63CACpEmwY/3lOZ4M0IQV2UJqSWzNNtUDI Hdq7hosk5gRXP/gG1bV63i935Ibe/Sp6Cb/XkTRcrPIxy/1eky8PDZN3knPlPocM TMAdLKUOzzpmehqORWdVsEGSYIXuIfVhrey30pfHVLQc86orTj7worDZydYl8r3L K6nAM8gfcT9l9Sd4jtquaT61kqCcjXKPANlvUtt8oqniMRdpL63GnFHaU33n3XTE 5Dalh4YHtIL4gTA6xZLbZqOq+99QbmmqlqlMiwFNtrfpVtPO7HWHrEy61mYrKW7U Nr7HRF6X+MeUngZVI5AgrH5K6HtlE0SeHZt2XSKuMKGMG+I4dcygUcSr =F8+z -----END PGP SIGNATURE----- Merge tag 'ceph-for-7.3-rc1' of https://github.com/ceph/ceph-client Pull ceph updates from Ilya Dryomov: "A wide variety of mostly CephFS fixes and cleanups, split between changes that address edge cases (Sam, Xiubo, Matthew), efficiency improvements (Max) and AI-assisted hardening (Michael, Jeremy). One thing that stands out is Alex's change to how CephFS behaves in NEARFULL scenarios: the long-standing "make all writes synchronous" behavior has become opt-in. It was always somewhat controversial and doesn't make much sense for modern deployments; the new default is to continue normal operation (i.e. buffer writes as MDS allows, etc). The behavior in case the cluster reaches any FULL state remains the same as before" * tag 'ceph-for-7.3-rc1' of https://github.com/ceph/ceph-client: (32 commits) ceph: force a cap message when a deferred revoke can't be acked immediately libceph: reject buckets with mismatched CRUSH ids ceph: reject export_targets ranks >= CEPH_MAX_MDS in mdsmap decode ceph: fix leaked inode reference on writeback abort at umount libceph: remove ceph_put_page_vector() libceph: validate banner payload length ceph: make nearfull sync writes opt-in ceph: do not repeat ceph_trim_dentries() if no progress possible ceph: drop mdsc->mutex before decoding the MDS reply ceph: fix UAF in check_new_map() on session freed during unlock ceph: fix UAF in __kick_flushing_caps() on cf entry freed during unlock ceph: pass inode pointer around instead of reloading it ceph: mark cap remove with RB_CLEAR_NODE() instead of setting ci=NULL ceph: add helper function ceph_cap_is_removed() ceph: make __ceph_remove_cap() static ceph: cap delegated inode count in ceph_parse_deleg_inos() ceph: bound num_export_targets array for mds info v2/v3 ceph: bound MDSCapAuth path and fs_name decode in handle_session() ceph: bound xattr value length in __build_xattrs() ceph: bound copied dentry name length in NFS export get_name ... |
||
|
|
ce727a090b |
This pull request contains updates for UBI and UBIFS:
UBI: - Support for a per-device wear-leveling threshold - Various fixes and cleanups of error paths - Correctly preserve torture flag up wear-leveling UBIFS: - Various fixes and cleanups of error paths and kernel-doc -----BEGIN PGP SIGNATURE----- iQJmBAABCABQFiEEdgfidid8lnn52cLTZvlZhesYu8EFAmqRnEsbFIAAAAAABAAO bWFudTIsMi41KzEuMTEsMiwyFhxyaWNoYXJkQHNpZ21hLXN0YXIuYXQACgkQZvlZ hesYu8HaKhAAuU11eCVhmQk9jpIdaOnFKwokE/PIj23SRz13jc7PWmFbAIHiaVt4 0p1Xd59kPfdFqsfBo6ZgSr+f0cQu2N1MjzZYSmaklG5iJk2+IuMOIvNdve/GFzNg X86J7yxLtrCq+ULNNgGv0m89G/uYoFP27Su0rAid4D2T5gYEOisXpPw5AhAL6+bS FLRMlt0QWCAtb66FmSeDgTW042NPoSCZNOsxF35X9hQ6RxvftB8mbggmTummmlgb K9Nsuwkarterq7JhS4X+RL6aZG7yDPfHVpdCDD6Ui4W0SP59W0oSokikm1myw6Yg 9NH67jUD03s0y/z3QaNiTVPuQXz5dpUyxGzK/FCT8y2aiycCDmwWF0VG8MVMhC0j TmNiuSWAljweaPSsgF176ISRG63++zwMGtz/JclVlkUwkVKYFbpEv5lcT0EzfCqh ahBlpM2aCLgJaYExd6LSVKEu0tT3+x+2MJDnRWcZM2OPMuL65IQw5V2lbdJN78Gb XFX6bfolhneptc2JwQKJx8E72iW7zuFovprTyS/J+TNcceFbfQTTcFt9EFZaXWnh +wiWm7rL7Xp9t8MA4Sc979oKFJMkAMl3Z6WTUHJWUZicX2hK+3OQaCVNO8LO/izu nPanrTeBnRiBKuAc5skQN5ywaAD3GgG/5ERWP/tdHkdsx34pcvuPnHQ= =+fCs -----END PGP SIGNATURE----- Merge tag 'ubifs-for-linus-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/rw/ubifs Pull UBI and UBIFS updates from Richard Weinberger: "UBI: - Support for a per-device wear-leveling threshold - Various fixes and cleanups of error paths - Correctly preserve torture flag up wear-leveling UBIFS: - Various fixes and cleanups of error paths and kernel-doc" * tag 'ubifs-for-linus-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/rw/ubifs: UBI: support per-device wear-leveling threshold UBI: fix two issues in the ubi.mtd MODULE_PARM_DESC mtd: ubi: Release device reference on busy detach ubi: Fix rollback for explicit UBI device numbers ubifs: fix out-of-bounds read in signature length check UBI: fastmap: Pass to_be_tortured when reusing old fastmap PEBs UBI: Preserve torture flag when rescheduling failed erasures ubifs: ubifs.h: clean up kernel-doc comments ubifs: key.h: use correct function parameter name ubifs: debug.h: fix kernel-doc struct prototypes |
||
|
|
115bd364ab |
f2fs-for-7.3-rc1
In this round, key enhancements focus on reducing inode management memory
overhead, introducing resizable tail sections with unified pinned allocation,
and boosting I/O throughput via parallel multi-device flushes and asynchronous
f2fs_write_end_io() execution. We also add dynamic device alias reservations to
allow on-the-fly space donation from user partitions.
Alongside these features, critical bug fixes resolve folio race conditions,
lingering dirty flags, dentry and block counter leaks, and potential deadloops
in f2fs_fsync_node_pages(). Additional stability patches address error-path
handling across symlink, sync, and rename/unlink operations, prevent pinned file
fragmentation, and correct segment migration and free section accounting in
free_segment_range.
Enhancement:
- reduce memory footprint of ino management
- support dynamic reserve/release for device aliasing
- issue multi-device flushes in parallel
- add a way to run f2fs_write_end_io() asynchronously
- support resizable tail section and unify pinned allocation
Bug fix:
- fix to pass folio->index to f2fs_sanity_check_node_footer()
- fix folio_nr_pages() race after put in large folio invalidate
- fix to clear dirty flag on folio in error path
- accurately adjust free_sections during free_segment_range
- fix to avoid potential deadloop in f2fs_fsync_node_pages()
- fix the error path in symlink, device alias in rename/unlink,
f2fs_sync_fs,
- fix to migrate all curseg types during free_segment_range
- fix to avoid pinfile fragment on fragment:{block, segment} mode
- fix valid block count leak on data block allocation failure
- fix dentry folio leak in find_in_level
- reject overlapping move range after len expansion
- fix some bugs related to file pinning, GC functions, i_size.
And, the series includes a number of minor bug fixes.
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEE00UqedjCtOrGVvQiQBSofoJIUNIFAmqPybMACgkQQBSofoJI
UNKp2g/+OP6XZi56hNTqscnKyKrDdVJnOcS/YOe7d1BR+070qmPTpyFjLgjng05K
exu55rz9vJ3DlpFLsjMEo60DRlDEc5rR4AqymMjqFJH9424ZlxPpdDn6ofCVT0Ck
D6RTf3y1HFSi4x7//gPQofR9y4MlDrH2Q7NPDriipqbymuNXEjrx/vdr2nq/kHUu
2lbf7QQs08qYiyDxQcxOFdCdxUTrsEW/tkYZiwgbU2nCJ/eG2R59amgtYJg3SlVt
xdrf+IaSS7kE5+mGCoBm0WooPpB507kHaoQpZYDj2uueFvEw7nFcSfOCXapOvg8U
wFkRR0F/rZ4+AW/u4n8Ye7N4a7WjWMTBwfR3WJ2j+arhfn87vJZK6LQQlYwq16l4
tRcQFcCrKsXHh5HY2OGj8DzTd40zryXujH566YioBCAXU312My1yFjeTJzjobbUW
TclkfMl689iTr9pqBhIjKT2tTvZFLROYSLk5UBNFNSfA4PtsAAhqlUrE1ck3AuTA
8kIjLmppfZEBqVuZCF0T1z9TXk0Bg0eM8qHbl/8SdDavmf1pF1BupLqchfdZp6Jr
4iV1weK4dzmCF6++YDsNlvnBGyTbhiFdN7C5cMlv89PTSTUQZ45Me6JmNwrRDqyU
MSTyiSSpI+3rAMO0h58PHi+QAOtM0f+S7C24ySkBBerBCbMTK6A=
=xQ77
-----END PGP SIGNATURE-----
Merge tag 'f2fs-for-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs
Pull f2fs updates from Jaegeuk Kim:
"In this round, key enhancements focus on reducing inode management
memory overhead, introducing resizable tail sections with unified
pinned allocation, and boosting I/O throughput via parallel
multi-device flushes and asynchronous f2fs_write_end_io() execution.
We also add dynamic device alias reservations to allow on-the-fly
space donation from user partitions.
Alongside these features, critical bug fixes resolve folio race
conditions, lingering dirty flags, dentry and block counter leaks, and
potential deadloops in f2fs_fsync_node_pages(). Additional stability
patches address error-path handling across symlink, sync, and
rename/unlink operations, prevent pinned file fragmentation, and
correct segment migration and free section accounting in
free_segment_range.
Enhancements:
- reduce memory footprint of ino management
- support dynamic reserve/release for device aliasing
- issue multi-device flushes in parallel
- add a way to run f2fs_write_end_io() asynchronously
- support resizable tail section and unify pinned allocation
Bug fixes:
- fix to pass folio->index to f2fs_sanity_check_node_footer()
- fix folio_nr_pages() race after put in large folio invalidate
- fix to clear dirty flag on folio in error path
- accurately adjust free_sections during free_segment_range
- fix to avoid potential deadloop in f2fs_fsync_node_pages()
- fix the error path in symlink, device alias in rename/unlink,
f2fs_sync_fs
- fix to migrate all curseg types during free_segment_range
- fix to avoid pinfile fragment on fragment:{block, segment} mode
- fix valid block count leak on data block allocation failure
- fix dentry folio leak in find_in_level
- reject overlapping move range after len expansion
- fix some bugs related to file pinning, GC functions, i_size
And, the series includes a number of minor bug fixes"
* tag 'f2fs-for-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs: (51 commits)
f2fs: support resizable tail section and unify pinned allocation
f2fs: don't leave the hashed inode while it's unlinked
f2fs: accurately adjust free_sections during free_segment_range
f2fs: fix to avoid potential deadloop in f2fs_fsync_node_pages()
f2fs: use adjusted write range after f2fs_write_checks()
f2fs: fix to propagate error from f2fs_sync_fs()
f2fs: return symlink writeback errors
f2fs: fix error handling on device alias check in rename and unlink
f2fs: fix to reset all pinned status during fggc
f2fs: use f2fs_{down, up}_(read, write}_trace() for nat_tree_lock
f2fs: reduce memory footprint of ino management
f2fs: fix i_size when pinned fallocate partially fails
f2fs: fix to migrate all curseg types during free_segment_range
f2fs: avoid setting SBI_NEED_FSCK on transient resize failure
f2fs: fix to avoid pinfile fragment on fragment:{block, segment} mode
f2fs: cleanup w/ f2fs_need_rand_{blk, seg, seg_blk}
f2fs: fix to shrink gc_lock coverage in f2fs_gc_range()
f2fs: fix to reclaim space in f2fs_allocate_pinning_section()
f2fs: unify add/remove ino entry API for all ino types
f2fs: fix to zero post-EOF data when extending file size
...
|
||
|
|
18fbf5151d |
mm.git review status for linus..mm-stable
Everything: Total patches: 171 Reviews/patch: 1.83 Reviewed rate: 82% Excluding selftests: Total patches: 149 Reviews/patch: 1.77 Reviewed rate: 80% Excluding selftests and maple_tree: Total patches: 129 Reviews/patch: 1.99 Reviewed rate: 89% Summary of patch series in this merge: - "mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff" (Lorenzo Stoakes): Index MAP_PRIVATE file-backed folios by their anonymous page offset to resolve confusion around reverse mapping for zeroed and CoW'd file-backed memory. Use this new VMA anonymous page offset tracking to eliminate index conflicts and lay the foundation for scalable CoW performance improvements. - "promote mapped executable folios after first usage for MGLRU" (Baolin Wang): Make MGLRU's protection of mapped executable file folios more reliable. Follow the classical LRU's logic, promoting mapped executable file folios after their first usage to give executable code a better chance to stay in memory and improve workload performance. - "mm: vmscan: fix node reclaim ignoring swappiness parameter" (Ridong Chen): Fix per-node proactive reclaim interface's ignoring the swappiness parameter when CONFIG_MEMCG is disabled by consolidating sc_swappiness() into a single function that checks proactive_swappiness regardless of kernel configuration. - "mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance cost" (Usama Arif): Reduce lru_lock contention in the reclaim path by deriving scan-balance costs from vmstat counters rather than lock-acquired producer updates. Read and decay these cost signals on the reclaim side under a dedicated per-lruvec lock, reducing total LRU lock wait time by over 60% without impacting scan throughput. - "zram: fix zram issues reported by sashiko" (Sergey Senozhatsky): Fix two low-risk zram bugs which Sashiko spotted in drive-by review. - "Honor XA_FLAGS_ACCOUNT in xas_split_alloc() and charge to folio's memcg" (Zi Yan): Fix xas_split_alloc() by enabling target folio memcg charging during splits and adding the missing __GFP_ACCOUNT flag for proper XArray node memory accounting. - "selftests/mm: use pattern matching in .gitignore" (Pratyush Mallick): Replace hardcoded binary names in selftests/mm/.gitignore with a generic pattern-matching rule to automatically ignore generated test files and avoid manual updates when adding new tests. - "mm/page_ext: remove pgdat_page_ext_init()" (Sang-Heon Jeon): Make the incompatibility between FLATMEM and NUMA explicit in mm/Kconfig and remove the unused pgdat_page_ext_init() function. - "zram: fix zstd error paths and add parameter validation" (Haoqin Huang): Clean up zram compression backends by removing redundant error cleanup, adding parameter and dictionary validation, auto-prefixing algorithm error logs, and resetting parameters prior to reinitialization. - "zram: fix stale scan bounds after reinitialization" (Longlong Xia): Prevent out-of-bounds slot accesses during concurrent zram resets by moving table scan bound calculations under dev_lock in writeback_store() and read_block_state(). - "add anon mTHP collapse test cases" (Baolin Wang): Extend selftests helper functions to support arbitrary page orders and add new test cases and options for mTHP collapse in khugepaged. - "selftests/mm: Handle unsupported and transient test conditions" (Muhammad Usama Anjum): Update MM selftests to report a SKIP status instead of a failure when required kernel or filesystem features are unsupported, while adding retry logic for transient page migration errors. - "mm/zswap: Fixes and improves the zswap shrink" (Hao Jia): Fix the missing zswap global shrinker when CONFIG_MEMCG is disabled and extend shrink_memcg() to support batch writeback for improved writeback efficiency. - "alloc_tag: introduce IOCTL-based filtering for MAP" (Suren Baghdasaryan): Introduce an IOCTL-based binary interface for memory allocation profiling that enables kernel-side filtering before per-CPU counter aggregation. This eliminates the text-parsing overhead of /proc/allocinfo and provides up to a 20x speedup by transferring only filtered allocation data to userspace. - "better block swap batching and a different take on swap_ops v5" (Christoph Hellwig): Refactor block swap I/O to use swap_iocb for batching instead of single-bio requests and rebase the swap_ops interface, achieving faster swap throughput during kernel builds. - "mm: kmemleak: reduce transient false positives by confirming leaks" (Catalin Marinas): Reduce false-positive kmemleak reports by combining two kmemleak enhancements that add a second confirmation scan and a configurable minimum unreferenced scan count module parameter. - "mm: kmemleak: default min_unref_scans to 2 for verbose kernels" (Breno Leitao): Auto-scanning kernels can generate false-positive memory leak reports on single scans, so this patch defaults min_unref_scans to 2 when CONFIG_DEBUG_KMEMLEAK_VERBOSE is enabled to require a second confirming scan. - "swap_ops updates" (Christoph Hellwig): Batching I/O for synchronous swap devices causes performance regressions and filesystem-based swap suffers from double-indirection overhead. This series resolves both issues by reintroducing per-folio writes for synchronous swap and allowing filesystems to directly export their own swap_ops. - "mm/khugepaged: several cleanups" (Nico Pache): khugepaged accumulated redundant state-checking patterns and outdated comments following mTHP integration. Introduce dedicated helpers for PTE validation and event counting while refreshing the internal documentation. - "maple_tree: lock checking and clean ups" (Liam Howlett): Syzbot reports incorrectly blame memory management exit paths for locking bugs, maple tree erase operations risk allocation failures without gfp flags and internal documentation lacks clarity. Improve lock error detection, update docs, fix race and allocation edge cases and optimize erase allocations using a fallback to GFP_KERNEL | GFP_NOFAIL. -----BEGIN PGP SIGNATURE----- iHUEABYKAB0WIQTTMBEPP41GrTpTJgfdBJ7gKXxAjgUCao9nJQAKCRDdBJ7gKXxA jk/9AQDlfevYJuSJmzAI8bt8ISG+/TfXMtIZC/MdbHqtQVYWPQD8Cvm3DUZsdGB/ Gloq/HBFuMPgE8p2pwUIthdgnTPNvAc= =c+Nb -----END PGP SIGNATURE----- Merge tag 'mm-stable-2026-08-26-15-22' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm Pull more MM updates from Andrew Morton: - "mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff" (Lorenzo Stoakes) Index MAP_PRIVATE file-backed folios by their anonymous page offset to resolve confusion around reverse mapping for zeroed and CoW'd file-backed memory. Use this new VMA anonymous page offset tracking to eliminate index conflicts and lay the foundation for scalable CoW performance improvements. - "promote mapped executable folios after first usage for MGLRU" (Baolin Wang) Make MGLRU's protection of mapped executable file folios more reliable. Follow the classical LRU's logic, promoting mapped executable file folios after their first usage to give executable code a better chance to stay in memory and improve workload performance. - "mm: vmscan: fix node reclaim ignoring swappiness parameter" (Ridong Chen) Fix per-node proactive reclaim interface's ignoring the swappiness parameter when CONFIG_MEMCG is disabled by consolidating sc_swappiness() into a single function that checks proactive_swappiness regardless of kernel configuration. - "mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance cost" (Usama Arif) Reduce lru_lock contention in the reclaim path by deriving scan-balance costs from vmstat counters rather than lock-acquired producer updates. Read and decay these cost signals on the reclaim side under a dedicated per-lruvec lock, reducing total LRU lock wait time by over 60% without impacting scan throughput. - "zram: fix zram issues reported by sashiko" (Sergey Senozhatsky) Fix two low-risk zram bugs which Sashiko spotted in drive-by review. - "Honor XA_FLAGS_ACCOUNT in xas_split_alloc() and charge to folio's memcg" (Zi Yan) Fix xas_split_alloc() by enabling target folio memcg charging during splits and adding the missing __GFP_ACCOUNT flag for proper XArray node memory accounting. - "selftests/mm: use pattern matching in .gitignore" (Pratyush Mallick) Replace hardcoded binary names in selftests/mm/.gitignore with a generic pattern-matching rule to automatically ignore generated test files and avoid manual updates when adding new tests. - "mm/page_ext: remove pgdat_page_ext_init()" (Sang-Heon Jeon) Make the incompatibility between FLATMEM and NUMA explicit in mm/Kconfig and remove the unused pgdat_page_ext_init() function. - "zram: fix zstd error paths and add parameter validation" (Haoqin Huang) Clean up zram compression backends by removing redundant error cleanup, adding parameter and dictionary validation, auto-prefixing algorithm error logs, and resetting parameters prior to reinitialization. - "zram: fix stale scan bounds after reinitialization" (Longlong Xia) Prevent out-of-bounds slot accesses during concurrent zram resets by moving table scan bound calculations under dev_lock in writeback_store() and read_block_state(). - "add anon mTHP collapse test cases" (Baolin Wang) Extend selftests helper functions to support arbitrary page orders and add new test cases and options for mTHP collapse in khugepaged. - "selftests/mm: Handle unsupported and transient test conditions" (Muhammad Usama Anjum) Update MM selftests to report a SKIP status instead of a failure when required kernel or filesystem features are unsupported, while adding retry logic for transient page migration errors. - "mm/zswap: Fixes and improves the zswap shrink" (Hao Jia) Fix the missing zswap global shrinker when CONFIG_MEMCG is disabled and extend shrink_memcg() to support batch writeback for improved writeback efficiency. - "alloc_tag: introduce IOCTL-based filtering for MAP" (Suren Baghdasaryan) Introduce an IOCTL-based binary interface for memory allocation profiling that enables kernel-side filtering before per-CPU counter aggregation. This eliminates the text-parsing overhead of /proc/allocinfo and provides up to a 20x speedup by transferring only filtered allocation data to userspace. - "better block swap batching and a different take on swap_ops v5" (Christoph Hellwig) Refactor block swap I/O to use swap_iocb for batching instead of single-bio requests and rebase the swap_ops interface, achieving faster swap throughput during kernel builds. - "mm: kmemleak: reduce transient false positives by confirming leaks" (Catalin Marinas) Reduce false-positive kmemleak reports by combining two kmemleak enhancements that add a second confirmation scan and a configurable minimum unreferenced scan count module parameter. - "mm: kmemleak: default min_unref_scans to 2 for verbose kernels" (Breno Leitao) Auto-scanning kernels can generate false-positive memory leak reports on single scans, so this patch defaults min_unref_scans to 2 when CONFIG_DEBUG_KMEMLEAK_VERBOSE is enabled to require a second confirming scan. - "swap_ops updates" (Christoph Hellwig) Batching I/O for synchronous swap devices causes performance regressions and filesystem-based swap suffers from double-indirection overhead. This series resolves both issues by reintroducing per-folio writes for synchronous swap and allowing filesystems to directly export their own swap_ops. - "mm/khugepaged: several cleanups" (Nico Pache) khugepaged accumulated redundant state-checking patterns and outdated comments following mTHP integration. Introduce dedicated helpers for PTE validation and event counting while refreshing the internal documentation. - "maple_tree: lock checking and clean ups" (Liam Howlett) Syzbot reports incorrectly blame memory management exit paths for locking bugs, maple tree erase operations risk allocation failures without gfp flags and internal documentation lacks clarity. Improve lock error detection, update docs, fix race and allocation edge cases and optimize erase allocations using a fallback to GFP_KERNEL | GFP_NOFAIL. * tag 'mm-stable-2026-08-26-15-22' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (172 commits) selftests/proc: make proc-maps-race work with READ_IMPLIES_EXEC memcg: move LRU size accounting on reparenting instead of copying it mm/vmscan: fix comment logic in balance_pgdat maple_tree: add helper mas_make_walkable() maple_tree: avoid extra gap calculation maple_tree: fix argument name in header maple_tree: change two GFP flags in tests maple_tree: document erase and allocations better maple_tree: avoid mas_erase() and mtree_erase() failures maple_tree: document that erase may use GFP_KERNEL for allocations maple_tree: catch race in mas_alloc_cyclic() maple_tree: add bulk parent set helper maple_tree: micro optimisation of mas_wr_store_type() maple_tree: optimise mas_wr_node_store() when not in rcu mode maple_tree: use prefetched value in mas_wr_store_type() maple_tree: clarify comments on mas_nomem() maple_tree: drop MAPLE_ALLOC_SLOTS maple_tree: drop dead code from mas_extend_spanning_null() maple_tree: documentation fix maple_tree: add write lock checking with lockdep sequence numbers ... |
||
|
|
ac727d86fb |
ntfs: leave HasEA flag untouched on setxattr failure
In ntfs_set_ea(), the exit path unconditionally updates the HasEA
flag based on ea_info_qsize. When an error occurs before
ea_info_qsize is updated, NInoClearHasEA() hides existing on-disk
EAs until the inode is evicted.
Only update the flag on success.
Fixes:
|
||
|
|
67aded1da1 |
ntfs: fix race between fallocate and mmap reads
The fallocate implementation only takes invalidate_lock for punch hole,
collapse range, and insert range operations. For standard allocation modes
(mode == 0, FALLOC_FL_KEEP_SIZE), the lock is not held.
During ntfs_attr_fallocate(), new clusters are mapped to the runlist via
ntfs_attr_map_cluster() before being zeroed by ntfs_dio_zero_range(). This
creates a window where concurrent mmap page faults can read uninitialized
disk data.
Since mmap uses filemap_fault() which takes invalidate_lock in shared mode,
it can fault in pages during this window and expose old disk contents to
userspace. This is an information leak and data integrity issue.
Fix by taking invalidate_lock for all fallocate operations, not just for
punch/collapse/insert modes. This prevents concurrent page faults from
accessing unzeroed clusters during the allocation window.
Fixes:
|
||
|
|
acb1095fd2 |
ntfs: fix memmove overlap in ntfs_new_attr_flags
When the record shrinks while the payload offsets increase (e.g., enabling
compression reduces padding, making arec_size < old_arec_size, but the header
grows by 8 bytes), moving the name first can overwrite the old mapping_pairs
before they are copied. Move mapping_pairs first in this case.
Since mp_ofs is derived from name_ofs, they always change in the same
direction. Checking name_ofs alone is sufficient.
Fixes:
|
||
|
|
6faa235a64 |
ntfs: compute bi_sector in 512-byte units
bi_sector counts in 512 byte sectors and not in multiples of the
volume's sector size. Under "normal" circumstances (with 512 byte
sectors in NTFS) the current code works as is; however, when we have
a 4k sector size on the volume the current usage of NTFS_B_TO_SECTOR()
and ntfs_bytes_to_sector() end up converting to the number of 4k
sectors after mount.
Reads work today on 4k volumes as bdev-io.c as performing the shift
correctly inline. With writes, we end up with significant silent disk
corruption on these volumes.
This fixes changes to use the new ntfs_bytes_to_bio_sector() function
everywhere we're performing this calculation (including the existing
read path). For the change in inode.c it removes a dead code block
rather than updating.
Fixes:
|
||
|
|
c966d29e01 |
f2fs: support resizable tail section and unify pinned allocation
Currently, zoned block devices restrict pinned file allocations to conventional zones at the beginning of the storage (before first_seq_zone_segno), triggering range GC when conventional space is exhausted. On regular block devices, when preparing for future online filesystem resizing (e.g. partition shrinking), pinned files must not be allocated in the tail area that will be truncated, as pinned files cannot be relocated by GC. Specifying the resizable tail area size (in sections) allows uniform mount configuration across devices of different storage capacities. To support this, introduce a unified `pinned_area_max_secno` boundary abstraction in `f2fs_sb_info`: 1. Add `-o resizable_tail_secno=%u` mount option to specify the number of sections at the tail of the filesystem reserved for resizing. 2. In `f2fs_fill_super()`, initialize `sbi->pinned_area_max_secno` as: min(MAIN_SECS(sbi) - resizable_tail_sec, zoned_max_sec). 3. In `get_new_segment()`, restrict segment allocation for pinned files (`pinning == true`) to `0 .. sbi->pinned_area_max_secno - 1`. If no free section is available in the pinned area, return -EAGAIN. 4. In `f2fs_allocate_pinning_section()`, unify the range GC trigger to run `f2fs_gc_range()` up to `sbi->pinned_area_max_secno` whenever `sbi->pinned_area_max_secno < MAIN_SECS(sbi)` and allocation returns -EAGAIN. 5. Expose `/sys/fs/f2fs/<dev>/pinned_area_max_secno` as a read-only sysfs node. Signed-off-by: Daeho Jeong <daehojeong@google.com> Signed-off-by: Sunmin Jeong <s_min.jeong@samsung.com> Reviewed-by: Wenjie Qi <qiwenjie@xiaomi.com> Reviewed-by: Chao Yu <chao@kernel.org> Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org> |
||
|
|
53676a5e28 |
cifs: add revalidation on FSCTL failure in smb2_duplicate_extents()
smb2_duplicate_extents() has no handling for
FSCTL_DUPLICATE_EXTENTS_TO_FILE failure: when the FSCTL fails, local
inode metadata may be stale from the pre-extension or from concurrent
remote writes, but is never refreshed.
Force revalidation on FSCTL failure and use i_size_read() for the
pre-extension check.
Fixes:
|
||
|
|
73e3f07100 |
NFS client updates for Linux 7.3
Highlights include:
Stable fixes:
- SunRPC: Use-after-free fixes for the sunrpc client code
- NFSv4: Delegation hash table leak
- lockd: NULL dereference on lockowner allocation failure
- SunRPC: Fix a handshake completion race in the TLS code
- NFSv4.1/pNFS: Fix an error sign checking issue when deciding whether
the layout is still in use, or can be returned.
- NFSv4.1: Fix a layout segment leak in pnfs_layout_process()
Other bugfixes:
- SunRPC: Fix a missing NULL check in the rpcbind client
- SunRPC: annotate shared socket callbacks with READ_ONCE/WRITE_ONCE
- NFSv4: nfs_inode_set_delegation() error paths should return the delegation
- NFSv4: Use clear_and_wake_up_bit() in nfs_clear_invalid_mapping() and
the pNFS code.
- NFSv4: Fix the nfs4_alloc_client() error paths to free the IDR
allocation
- NFS: fix folio dereference before NULL check in nfs_inode_remove_request()
- NFS: Fix delayed delegation return
- NFSv4: Fix another state manager race with umount
- pNFS/blocklayout: Fix device leaks on parse failure
- pNFS: Avoid cancelling in-flight I/O during a layout recall if the
server doesn't require it
- NFSv4/flexfiles: report cancelled I/O as a layout error
- NFSv4/flexfiles: fix NULL dereference for NFSv4.0 data servers
- NFSv4: Fix incorrect argument passed to nfs4_delete_lease()
- NFSv3: Fix several symlink issues resulting from nfs_atomic_open_v23()
- NFSv4.1: Fix an uninitialised variable issue in the callback code
- NFSv4.2: fix LAYOUTSTATS send buffer exhaustion
Features and cleanups:
- NFSv4.2: Allow the server to specify that file data may not be cached
- NFS/localio: optimise I/O submission when when not doing memory reclaim
- NFS/localio: Remove duplicate wait code in nfs_local_commit
- NFSv4/flexfiles: support loosely coupled NFSv4.x data servers
- NFSv4/pnfs: key the data server cache on the NFS version
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQR8xgHcVzJNfOYElJo6EXfx2a6V0QUCao9SMwAKCRA6EXfx2a6V
0VxpAP9KSFbBnHU/DTq6zJ0xNeatZLBssrdkD1aPbHGsJPXukgEAgmo9tk0AgdJo
gxPeuVJIepg9PEIxI6jd6TxwpUV8NQI=
=k59o
-----END PGP SIGNATURE-----
Merge tag 'nfs-for-7.3-1' of git://git.linux-nfs.org/projects/trondmy/linux-nfs
Pull NFS client updates from Trond Myklebust:
"Highlights include:
Stable fixes:
- Use-after-free fixes for the sunrpc client code
- Delegation hash table leak
- NULL dereference on lockowner allocation failure
- Fix a handshake completion race in the TLS code
- Fix an error sign checking issue when deciding whether the pNFS
layout is still in use, or can be returned
- Fix a layout segment leak in pnfs_layout_process()
Other bugfixes:
- Fix a missing NULL check in the rpcbind client
- annotate shared socket callbacks with READ_ONCE/WRITE_ONCE
- nfs_inode_set_delegation() error paths should return the delegation
- Use clear_and_wake_up_bit() in nfs_clear_invalid_mapping() and the
pNFS code.
- Fix the nfs4_alloc_client() error paths to free the IDR allocation
- fix folio dereference before NULL check in
nfs_inode_remove_request()
- Fix delayed delegation return
- Fix another state manager race with umount
- Fix device leaks on parse failure
- Avoid cancelling in-flight I/O during a layout recall if the server
doesn't require it
- flexfiles: report cancelled I/O as a layout error
- flexfiles: fix NULL dereference for NFSv4.0 data servers
- Fix incorrect argument passed to nfs4_delete_lease()
- Fix several symlink issues resulting from nfs_atomic_open_v23()
- Fix an uninitialised variable issue in the NFSv4.1 callback code
- fix LAYOUTSTATS send buffer exhaustion
Features and cleanups:
- NFSv4.2: Allow the server to specify that file data may not be cached
- localio: optimise I/O submission when when not doing memory reclaim
- localio: Remove duplicate wait code in nfs_local_commit
- flexfiles: support loosely coupled NFSv4.x data servers
- pNFS: key the data server cache on the NFS version"
* tag 'nfs-for-7.3-1' of git://git.linux-nfs.org/projects/trondmy/linux-nfs: (33 commits)
NFSv4.1: fix layout segment leak on the pnfs_layout_process() forget path
NFSv4/pnfs: key the data server cache on the NFS version
NFSv4.2: fix LAYOUTSTATS send buffer exhaustion
pNFS: Fix EBUSY check in pnfs_layout_need_return
NFSv4.1: zero referring call lists before decoding
nfs: fix ENXIO on O_CREAT open of existing symlink over NFSv3
SUNRPC: wait for in-flight client TLS handshake callback
NFSv4: Fix incorrect argument passed to nfs4_delete_lease() in nfs4_add_lease()
lockd: fix NULL dereference on lockowner allocation failure
NFS: fix delegation_hash_table leak when nfs4_server_common_setup() fails
NFSv4/flexfiles: support loosely coupled data servers
NFSv4/flexfiles: fix NULL dereference for NFSv4.0 data servers
NFSv4: pin the superblock for active state owners
sunrpc: fix use-after-free in __rpc_clnt_handle_event and __rpc_clnt_remove_pipedir
NFS/localio: issue commit inline when not in a memory-reclaim context
NFS/localio: remove dead FLUSH_SYNC handling from nfs_local_commit
NFS/localio: issue IO inline when not in a memory-reclaim context
NFS: Fix delayed delegation return list handling
NFS: Verify symlink inode before caching target
NFS: fix folio dereference before NULL check in nfs_inode_remove_request()
...
|
||
|
|
8fdf946445 |
ceph: force a cap message when a deferred revoke can't be acked immediately
When the MDS revokes capabilities, handle_cap_grant() normally guarantees a response by setting `CHECK_CAPS_FLUSH_FORCE` (see commit |
||
|
|
aedc9053d9 |
ceph: reject export_targets ranks >= CEPH_MAX_MDS in mdsmap decode
MDSMap export_targets entries are monitor controlled. check_new_map()
uses each entry as a bit number in a fixed stack bitmap, so a rank
outside the protocol namespace can make set_bit() write past the end of
the array.
Reject ranks outside CEPH_MAX_MDS while decoding the map. Do not
validate against possible_max_rank here because maps may legitimately
reference ranks beyond a temporarily reduced max_mds.
Cc: stable@vger.kernel.org
Fixes:
|
||
|
|
c25aee9c63 |
ceph: fix leaked inode reference on writeback abort at umount
ceph_dirty_folio() takes a wrbuffer claim on each newly dirtied folio: it
bumps i_wrbuffer_ref (taking an ihold() on the 0->1 transition) and
attaches the snap_context to folio->private. That claim is released only
by ceph_put_wrbuffer_cap_refs(), which for a submitted write runs from
writepages_finish().
In ceph_submit_write(), if ceph_inc_osd_stopping_blocker() fails -- which
happens during umount -- the request is aborted before submission: the
already-collected folios are only redirtied and unlocked, so
writepages_finish() never runs and the claim is leaked.
redirty_page_for_writepage() -> folio_redirty_for_writepage() ->
filemap_dirty_folio() sets PG_dirty directly and does not go through
->dirty_folio, so ceph_dirty_folio() is not re-entered to rebalance it.
Because every subsequent writeback also fails the osd_stopping_blocker,
i_wrbuffer_ref never returns to 0, the ihold() is never dropped, and the
inode cannot be evicted:
VFS: Busy inodes after unmount of ceph
kernel BUG at fs/super.c:650!
Release the orphaned claim in the abort path before redirtying, via
ceph_undo_wrbuffer_claim(): detach the snap_context, drop the wrbuffer
reference (letting i_wrbuffer_ref reach 0 and iput() the inode), and drop
the snap_context reference -- i.e. do what writepages_finish() would have
done for these never-submitted folios.
Only the locked_pages entries are undone; folios still in the fbatch were
never dirty-cleared by this call (folio_clear_dirty_for_io() is the
ownership-transfer point, and a successful move NULLs the fbatch slot), so
they hold no claim this call owns.
Cc: stable@vger.kernel.org
Fixes:
|
||
|
|
2a2f98e17e |
libceph: remove ceph_put_page_vector()
ceph_put_page_vector() was paired with ceph_get_direct_page_vector(),
which was removed in commit
|
||
|
|
5f074d7f29 |
ceph: make nearfull sync writes opt-in
The kernel CephFS client has historically treated a cluster or pool NEARFULL condition as a request to force successful writes through generic_write_sync(). That effectively turns otherwise buffered writes into synchronous writes and can cause a severe throughput drop as soon as a single OSD or the file data pool crosses the nearfull threshold. On modern large clusters, NEARFULL is primarily an operator health signal rather than an immediate client-side capacity failure. Operators can still have substantial usable capacity while a cluster is rebalancing, splitting PGs, or expanding onto new devices. RBD, RGW and the userspace CephFS client do not impose this extra client-side sync-write throttle, so the kernel client behavior is surprising and operationally painful. Change the default behavior so NEARFULL no longer changes normal write-sync semantics. FULL and pool FULL still fail with -ENOSPC, and explicitly synchronous writes continue to be synced by generic_write_sync(). Add a nearfull_sync mount option for deployments that want the legacy backpressure behavior. When this option is set, successful writes are promoted to IOCB_DSYNC if the cluster or file data pool is marked NEARFULL, preserving the old behavior for conservative deployments. Link: https://tracker.ceph.com/issues/74849 Signed-off-by: Alex Markuze <amarkuze@redhat.com> Reviewed-by: Xiubo Li <xiubo.li@clyso.com> Signed-off-by: Ilya Dryomov <idryomov@gmail.com> |
||
|
|
e7d7aa7b73 |
ceph: do not repeat ceph_trim_dentries() if no progress possible
ceph_cap_reclaim_work() re-queues itself for as long as
ceph_trim_dentries() returns -EAGAIN, which happens whenever a lease
walk exhausts its `nr_to_scan` budget. This creates a busy loop that
consumes CPU without making any progress when there is nothing to
reclaim: with no cap pressure (`count==0`) and every scanned lease
still valid, each pass runs the full scan budget down to zero and
returns `-EAGAIN`, only to be queued again immediately.
The dir-lease walk made this worse. When `expire_dir_lease` is
`false` (i.e. we have no intention of reclaiming dir leases),
__dir_lease_check() returned `TOUCH` for every valid lease. `TOUCH`
moves the dentry to the tail of the list and resets `di->time` via
__dentry_dir_lease_touch(), so a walk over N valid leases pointlessly
rewrote the list, refreshed the timestamps (preventing them from ever
aging out) and always drained `nr_to_scan`, guaranteeing the `-EAGAIN`
requeue.
Fix this in three steps:
- Return `KEEP` instead of `TOUCH` when `expire_dir_lease` is
`false`. If we are not going to reclaim the lease, leave it in
place instead of churning the list and resetting its timestamp; the
walk then terminates naturally (or via `STOP` at the first fresh
lease).
- Only return `-EAGAIN` from the first (dentry-lease) walk when something
was actually freed. A full batch that frees nothing means retrying
the same list immediately is futile; fall through to the dir-lease
walk instead.
- After both walks, bail out with success (0) when nothing was freed
and there is no cap pressure (`count==0`). There is no reason to
keep retrying when we are not over the cap limit and made no
progress.
Under real cap pressure (`count>0`) the reclaim path is unchanged and
still retries via `-EAGAIN`.
Without this patch, I saw 500 ceph_trim_dentries() calls per second on
our web servers. This is very visible in `/proc/lock_stat` (5 minute
capture):
class name con-bounces contentions waittime-min waittime-max waittime-total waittime-avg acq-bounces acquisitions holdtime-min holdtime-max holdtime-total holdtime-avg
&mdsc->dentry_list_lock: 126180 128218 0.04 8063.44 15986965.20 124.69 1573354
|