Five are DAMON fixes. One fixes an arm64 contpte bug where DAMON can
write past the end of a page-table page, resulting in memory corruption
and possible crashes.
Two are hugetlb fixes. One fixes an mremap() address calculation bug
which can panic x86-64.
There's also a missing anon_vma publication barrier which can result in
hung tasks, and a writeback fix to keep long cgroup writeback drains from
delaying Tasks-RCU grace periods.
The remainder are smaller fixes and maintenance changes.
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQTTMBEPP41GrTpTJgfdBJ7gKXxAjgUCarHHNAAKCRDdBJ7gKXxA
jhb7AP4yE/k7RrZC6zWg4M9ejI3fVlHI9+EG1mCLGiV57jZ4mAEA9LLSZryOD6Nc
zsaeCtZhHcEdxW6EhO18hMfH4b7oJA0=
=alhC
-----END PGP SIGNATURE-----
Merge tag 'mm-hotfixes-stable-2026-09-21-17-08' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Pull MM fixes from Andrew Morton:
"14 hotfixes. 10 are cc:stable. 11 are for MM.
Five DAMON fixes: one fixes an arm64 contpte bug where DAMON can write
past the end of a page-table page, resulting in memory corruption and
possible crashes.
Two hugetlb fixes: one fixes an mremap() address calculation bug which
can panic x86-64.
There's also a missing anon_vma publication barrier which can result
in hung tasks, and a writeback fix to keep long cgroup writeback
drains from delaying Tasks-RCU grace periods.
The remainder are smaller fixes and maintenance changes"
* tag 'mm-hotfixes-stable-2026-09-21-17-08' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
MAINTAINERS: update Xu Xin's email
writeback: report a Tasks-RCU quiescent state per cgwb drain pass
mm/damon/core: reset invalid quota->charge_target_from
MAINTAINERS: add Baoquan and Baolin as MGLRU reviewers
mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish
mm/hugetlb: preserve mremap address delta when skipping page tables
mm/damon/core: fix unconditionally skip last region
mm/damon/vaddr: avoid hw-driven pte updates during damon_hugetlb_mkold()
mm/damon/core: allow esz to be set to zero
mm/damon/ops-common: use a page-aligned address in damon_ptep_mkold()
ocfs2: make ocfs2_calc_xattr_init() return void
mailmap: update Haowen Bai's email address
selftests/cgroup: account for zswap shrinker writeback
mm/hugetlb: do not dissolve gigantic pages without runtime support
Signed-off-by: Carlos Maiolino <cem@kernel.org>
-----BEGIN PGP SIGNATURE-----
iJUEABMJAB0WIQSmtYVZ/MfVMGUq1GNcsMJ8RxYuYwUCarEtCAAKCRBcsMJ8RxYu
YwmHAX4+ML995OtOeKIN2Cpiu/0RYW/bHPrBBCHsiZN2BFZ3Ap2W7nxm4/OFewA9
5oB6eKgBgL8Zi2K0ZPsiF+IxAEkWUBkvCgiQtByO3dSS6aMR6KzfrQBj0Pftlcgy
z+C74fVlyA==
=USzk
-----END PGP SIGNATURE-----
Merge tag 'xfs-fixes-7.3-rc5' of gitolite.kernel.org:/pub/scm/fs/xfs/xfs-linux
Pull xfs fixes from Carlos Maiolino:
"This mostly contain 'random' bugfixes found by LLM tools on different
xfs subsystems. A few code cleanups and a NULL ptr deref on zoned
support. Those 'random' bugfixes include possible buf overrus, UAFs,
block leaks, etc..."
* tag 'xfs-fixes-7.3-rc5' of gitolite.kernel.org:/pub/scm/fs/xfs/xfs-linux: (26 commits)
xfs: fix wild memcpy access when formatting ondisk rtrefcount btree roots
xfs: don't let hidden_space go negative in xfs_metafile_resv_init
xfs: drop dquot flush lock when we can't find a buffer to flush
xfs: fix cursor and pointer handling when recovering iunlink buckets
xfs: fix blockgc group quota scanning when usrquota isn't enforced
xfs: don't merge different file IO error types
xfs: don't let memory failures leak blocks and kill repairs
xfs: don't cross reference rmapbt with bitmaps if they're incomplete
xfs: fix rtgroup repair estimations
xfs: call xfs_dquot_set_prealloc_limits if we installed default rtb limits
xfs: fix typos and repeated words in comments
xfs: remove unused xfs_reflink_remap_range declaration
xfs: remove duplicate INO1_WRITTEN check
xfs: don't try to get a reference to a NULL oz in xfs_get_cached_zone
xfs: check di_forkoff correctly in scrub
xfs: only flag zero padding for dir3 data blocks, not dir3 block blocks
xfs: fix integer overflows in xbitmap set functions
xfs: use the correct reservations for rtrmap/refcount recovery
xfs: don't call xfs_exchange_range_finish for a dry run
xfs: check padding field in xfs_ioc_commit_range
...
- Fix timer signal <-> exec() race, to prevent UAF (Thomas Gleixner)
- Clean up POSIX CPU timers right after de_thread(), to prevent UAF
(Hyunwoo Kim)
- Fix POSIX CPU timers race between expiry and timer_settime(),
to prevent UAF (Thomas Gleixner)
Signed-off-by: Ingo Molnar <mingo@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqvrMYRHG1pbmdvQGtl
cm5lbC5vcmcACgkQEnMQ0APhK1gqUQ/+NkruN984bFynF/eZ0/2DFv91AAUP8zgH
/S3PBlwuSbYFN9JVhngDMwxQkamE56weJbFc+0QuvVT5UVw/vX9BS4QOvvzN+f8D
FEN3UqD0d1B8OwlPNTw0sFPwJDdPctTinfKOhNjNQe6RLFsNARvGyaKDIDroWTfV
dxuJ/7Ecs+5m1bmGJnPEC+IH/OnV9BEEl1NdZb+INKpBlui9LCsw4rRIj/8dPK/H
UNhvXpykKrJCDftbCzAFSNryuzcJgq4kHtMbsqiUL6y50AB69eHGi/Y0xYBAEr1h
NiDPq2PAMmH1NCCMsTtqbJZMqgCr+7DSZiCFn7bZPwg0V5tV4PFZD484q0sCbiej
Fwg+arHd0icnceIcWMsBWPUVOSLxZaWdp9a2Tj3Ill06//b5bEDBJBbpecS+so3t
8W6IvdoCYm7sz50mohnjOdx7biHPu0yhwgj+EoAV3nZKoALQAAcI7+HJzSWpGnJi
HIO0zylRAZCjk9H3QNWO+LdWgifc8DysAZOWpmbuwGgp8q483IDRDtme/kMt3+D1
1qTHa1TD/tPo8UmmgyVJQ7e1hCxBkGuuBBu5Y3/qkUEOQM6B/H2Ji7stxsLpW3JL
HLzC3kL2SBBVBO2ljqiH5IhVAL10Qm5vPxaCOjExdBt1vjMxN1DowZqD4rT/tw58
ArpD4zmr8VQ=
=KhWJ
-----END PGP SIGNATURE-----
Merge tag 'timers-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull timer race fixes from Ingo Molnar:
- Fix timer signal <-> exec() race, to prevent UAF (Thomas Gleixner)
- Clean up POSIX CPU timers right after de_thread(), to prevent UAF
(Hyunwoo Kim)
- Fix POSIX CPU timers race between expiry and timer_settime(),
to prevent UAF (Thomas Gleixner)
* tag 'timers-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
exec: Cleanup POSIX timers right after de_thread()
signal: Prevent exec() race
-----BEGIN PGP SIGNATURE-----
iQJPBAABCgA5FiEE8rQSAMVO+zA4DBdWxWXV+ddtWDsFAmqu3wQbFIAAAAAABAAO
bWFudTIsMi41KzEuMTIsMiwyAAoJEMVl1fnXbVg7Y5kQAJ1ANaQWod+AUKhgbtA6
IcEFH8AAVrrTqqN0SxOl/tqHL6Xt/5mKlQcTpYPxeccUkp72M9BeGik4Gp7OwN1D
vkVTgrmhHT4r+4ae4GJP0yCtNJBa0fRXjsMFbNYTKGWcz9yOsW1OkBsq8VtRo7ym
CSVBc57buUi0mzbnLNDh69G/YA7NCTyaxXKjPARNYy+cy0LMIDmgZCByhgePQvzB
aqJUakRKpaeXEQIf0nT/70XcGhXNtfK1GOnLZM1ySFqLufMzdTsEzFsT8dVugqU1
wr1HMaMq1iNKwE4sJgNWS84wRT1zxZopKIOsTufFu98Zb2C0kvUSjWeJSqtFM88G
ggj/jmaoRhCU1cXE1jvZWy4Fe5zQH8deSQZ7zUB4ZzqCaDRwOEAj4IpuQQ9Ok91E
J/07SCvKPk/QUPbA2e5lwRL3aximsD1LfRxWIOZ+xg/HVX+vYfHmy5gzkl/hhfrm
RHkHLz8iXNHV7qhIVzZ/GYvTxR6U6nbLVUbkl5n/8eFePM7YHnoWtvMHT71qWwrh
mPx8uTK+4BI2+rAHTKqxnwEQ27OHj5odlGKTlKiaMrQR9eqnnhJLTMPDJMzqQ7TI
ZiibxG8UqSTsNbvOXXtd7MykSRpSb/VWyJImKuw30kx3f7f2yd1aCUyU4T7o5N/3
FRI+2AbkoSeype3Rw+VcEPoi
=WxSA
-----END PGP SIGNATURE-----
Merge tag 'for-7.3-rc3-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux
Pull btrfs fixes from David Sterba:
"Among the regular fixes, there are two that were reported recently and
have user impact:
- filesystem id is now stable again on the default and common case
(it broke openconnect key derivation, while this is not secure,
it's still in use), the intention was to change id for the
temp_fsid use case
- fix detection of /dev/root and rename it after device scan, this
broke booting of initramdisk-less system with grub2 as the probe
needs the real device
Regular fixes:
- don't store compressed inline extent if the size is larger than
uncompressed
- in zoned mode, handle activation of zones for all supported block
group profiles in case there are still free ones left
- check space for a chunk item when reading sys array from superblock
- allow using space reserves when removing verity items fails
- fix error handling after free space tree rebuild fails
- abort transaction if reflink or hole punching fails and it's not
possible to update the inode
- properly protect block group iteration during device replace start
- error message fixups"
* tag 'for-7.3-rc3-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux:
btrfs: derive f_fsid with dev_t only when temp_fsid is active
btrfs: add "/dev/root" exception for device path update
btrfs: check if there is space for chunk item when validating sys chunk array
btrfs: abort transaction on failure to update inode for hole punching and reflinking
btrfs: clear free space tree creation state on rebuild failure
btrfs: handle lack of space when cleaning up verity items
btrfs: fix creation of compressed inline extents that don't save space
btrfs: tree-checker: fix error message regarding free space extent items
btrfs: tree-checker: print dev extent offset in error message
btrfs: take commit root semaphore when iterating in mark_block_group_to_copy()
btrfs: zoned: handle RAID profiles in btrfs_can_activate_zone()
A batch of bug fixes for the smb client:
- Fix multiple out-of-bounds reads and use-after-frees in the SMB2/3
receive path that are reachable from a malicious or compromised
server: a stale next_buffer pointer and an integer overflow in
compound encrypted frame handling, missing minimum-PDU-size and
per-sub-PDU length validation before parsing command-specific
response fields, missing bounds checks in DFS referral, server
interface list, EA list, POSIX SID, snapshot enumeration and SMB1
reparse point parsing
- Fix use-after-frees and races in multichannel and connection
teardown, including an interface freed while still in use when
adding channels, a server used after its channel reference was
dropped, a reconnect work item left queued after the server is
freed and an uninitialized reconnect list node
- Fix a heap overflow in the native symlink parser: an absolute
target without an NT drive prefix caused out-of-bounds writes and a
u16 length underflow leading to a 64K memcpy into a small buffer,
triggerable by a user with write access to a mounted share under
default settings
- Fix WSL reparse point parsing: use unaligned accessors for the
packed extended-attribute payload to avoid alignment faults on some
architectures and stop leaving partially mutated fattr fields on
parse failure
- Fix lease break ACKs being sent through the wrong session on
multiuser mounts, which caused read failures (e.g. on NetApp
ONTAP/Azure Files) when copying files
- Fix an smbd_connection leak when cifs_get_tcp_session() fails after
an RDMA connection was already established
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQTcqRusfSdYROJQwGkpVtNKoQNdYwUCaq2frwAKCRApVtNKoQNd
Y1mcAQDcDTkep03jzghyJG6xWJ3S7KNbeYpjkOPnPyR+Et7HmAD/eXLFvgkJ3wC7
tBUDjDTLeyP6/DOBmDb/fIKEw2vfBQs=
=+Vsw
-----END PGP SIGNATURE-----
Merge tag 'cifs-fixes-7.3-rc4' of https://git.manguebit.org/linux
Pull smb client fixes from Paulo Alcantara:
"A batch of bug fixes for the smb client:
- Fix multiple out-of-bounds reads and use-after-frees in the SMB2/3
receive path that are reachable from a malicious or compromised
server: a stale next_buffer pointer and an integer overflow in
compound encrypted frame handling, missing minimum-PDU-size and
per-sub-PDU length validation before parsing command-specific
response fields, missing bounds checks in DFS referral, server
interface list, EA list, POSIX SID, snapshot enumeration and SMB1
reparse point parsing
- Fix use-after-frees and races in multichannel and connection
teardown, including an interface freed while still in use when
adding channels, a server used after its channel reference was
dropped, a reconnect work item left queued after the server is
freed and an uninitialized reconnect list node
- Fix a heap overflow in the native symlink parser: an absolute
target without an NT drive prefix caused out-of-bounds writes and a
u16 length underflow leading to a 64K memcpy into a small buffer,
triggerable by a user with write access to a mounted share under
default settings
- Fix WSL reparse point parsing: use unaligned accessors for the
packed extended-attribute payload to avoid alignment faults on some
architectures and stop leaving partially mutated fattr fields on
parse failure
- Fix lease break ACKs being sent through the wrong session on
multiuser mounts, which caused read failures (e.g. on NetApp
ONTAP/Azure Files) when copying files
- Fix an smbd_connection leak when cifs_get_tcp_session() fails after
an RDMA connection was already established"
* tag 'cifs-fixes-7.3-rc4' of https://git.manguebit.org/linux:
cifs: Fix server use-after-free in cifs_chan_skip_or_disable()
smb: client: fix reparse buffer bounds in cifs_query_reparse_point()
smb: client: fix potential OOB read in smb3_enum_snapshots()
smb: client: fix missing iov bounds check in parse_posix_sids()
smb: client: fix OOB struct field reads in move_smb2_ea_to_cifs()
smb: client: reject short Next offsets in parse_server_interfaces()
smb: client: fix missing lower-bound check on DFS referral string offsets
smb: client: fix server->total_read for compound encrypted PDUs
smb: client: validate minimum PDU size before smb2_get_data_area_len()
smb: client: fix next_buffer UAF and NextCommand bounds in compound PDUs
smb: client: fix use-after-free of iface in cifs_try_adding_channels()
smb: client: fix fattr leaking on wsl_to_fattr() failure
smb: client: fix unaligned access in WSL reparse point parser
smb: client: fix smbd_connection leak on cifs_get_tcp_session() error
smb: client: fix rlist race and missing initialization
smb: client: cancel reconnect work in clean_demultiplex_info()
smb/client: send lease break ACKs thru correct session for multiuser mounts
smb: client: validate absolute native symlink targets before NT fixups
after ten seconds of inactivity when a new session setup request is
received. Sessions now expire only after credential expiration, while
stale unauthenticated sessions are cleaned up after a 45-second timeout.
- Keep earlier responses in compound requests when Query Info fails
because the output buffer is too small. The error response is appended
without truncating preceding responses.
- Return STATUS_BUFFER_OVERFLOW for partial
FILE_NORMALIZED_NAME_INFORMATION responses instead of incorrectly
returning STATUS_INFO_LENGTH_MISMATCH.
-----BEGIN PGP SIGNATURE-----
iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqtTBMWHGxpbmtpbmpl
b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCD2GD/9OzkmZ+/kIMum8IQi9upOM5zUD
+2nyFMebJ6qXmrhOuUgLn2ptUYxO3q1B9VvzTjKrM1Vun0VuvHmHQbGUp9a/3F7c
VTncl7E3FEHqlPkLWQJFv2dS+TYYcOhjoc2TDAY9093xktcMHK5zjv4E/DV2o0bf
NN/GcrrcWSxdJU8WI9JvY2kzmPQMDfM8uVbj2RqHSZ8m9mF5cajtDnzCCS0K6UOR
qzMrekzewIskUtcQNqU8hJWl1sgiYdD+16LmKmwLd3uOZISc3Miy5Bg8VpNN+B1g
XHhr49G4Fb8PkEYjncjxQH7zop1ID5UyC2xN63NuoH8+mkm4qn6uG2NrLhmVQO+f
Z81Ov3g77/jK3Z2fw00H3A7VGSPs931BaRTNj+lPkHTVVq3PyP1XZsfXkPUoRu1Q
xeaJydScGNYE+kaYbseXwN8haJGawd1Dd+Afn4W2zikUU5tKZxO9d2tBr4Fhl/Fm
0OBcAssTArhrY7PX2fAOQ4sUwAC4nMXSEIfdWIvcgU2OoyuFltQ55ooCoM0uLs6+
7R4rxZLipdZmKhuE2mlAWcFQJn8nS0EHXeyEDjO0u8KYsV6C8d+jDNpvtS+/5QVF
bKE4/uUX/lV0IZ55nFSPT7XdI+gxeiUql3+8bqOVqkw7wR4VjaC0xIQybyEBTMYH
IJEKfjx4tZcWGTzmdA==
=9nNl
-----END PGP SIGNATURE-----
Merge tag 'ksmbd-for-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb
Pull smb server fixes from Namjae Jeon:
- Fix session expiration so that valid sessions are no longer removed
after ten seconds of inactivity when a new session setup request is
received.
Sessions now expire only after credential expiration, while stale
unauthenticated sessions are cleaned up after a 45-second timeout.
- Keep earlier responses in compound requests when Query Info fails
because the output buffer is too small. The error response is
appended without truncating preceding responses.
- Return STATUS_BUFFER_OVERFLOW for partial
FILE_NORMALIZED_NAME_INFORMATION responses instead of incorrectly
returning STATUS_INFO_LENGTH_MISMATCH.
* tag 'ksmbd-for-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb:
ksmbd: keep compound responses on query info errors
ksmbd: fix partial normalized name responses
ksmbd: follow SMB2 session expiration semantics
dynamically reserving MFT tail records, accounting for records added
during allocation, and avoiding false -ENOSPC failures.
- Repack non-resident $MFT/$ATTRIBUTE_LIST when its mapping pairs no longer
fit in the base MFT record, while propagating allocation and writeback
errors.
- Serialize runlist updates with the runlist lock and restore both the
in-memory runlist and on-disk mapping pairs when allocation rollback is
required.
- Propagate folio errors and harden inode failure handling by treating
interrupted reads as transient failures and discarding and unhashing
inodes whose initialization fails.
- Fix the $MFTMirr write offset when mirror records span multiple folios,
preventing mirror records from overwriting the first record with large
MFT record sizes.
-----BEGIN PGP SIGNATURE-----
iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqtRIEWHGxpbmtpbmpl
b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCBK8EAChFdTxig3sYog3dL27SX+1yNC4
S1QtSQMvPJzcvLG2/mESC57snx1u6AXZ6enaQ/m8zQQvkWHFdwD+odgjOQ3474Yc
MS7xQn5pCsAo3LmSWiDsQfHmvxgDMYlIRU1vmqr7fG2pj+W6BR2LB1PD3RI9exI7
0WWAXBPZH5w5C9GE1Zo7TF9Xwby5Or31RS8+R57PXA/PJ1ivpWnlkpfLi4M/YkZK
GIpunZafSpcKbEsuWcjhdz11bR4G9Qlwuuq0MDguLC/qsqsobHCeSbdx+4IsEAq5
02yBl5hYm2E4u2KBedpe7oRwFvlPN0uakEGYS8SA1ad9XamjGIw6T0tkzkDL48Pq
dhVAeX2oa8O9u+VK+qF/HIUylh/UbmHQJW8iSiZWO8WdULGBG8oCHI1hcSnMguwJ
njyK75UXz4fMsKW6ZpRu0sRGqtKKcbg8IrCvLslPIOS2A9OAwSzytDKI+x1Kbgu0
SVPYjf6XeOz83tvE+2OhfTT1hWkeKezMiUe4E/y9rgEDxv5vE4C5pvpDfRlj3oyn
2JTXXjjUQGSQw/9cKbLsbElDH/FLEojLAsIFgM+2FbcG+x7PUUgSGdC4R4qZ9MUF
cuMzfhi8vKyCYmJc8nNE4J6b6UkBNOFJMYBcArtByFW3LqfIzmqwW81Xx0Hz4cxi
dmOhD6gXXYoi9m3lBQ==
=rT1L
-----END PGP SIGNATURE-----
Merge tag 'ntfs-for-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs
Pull ntfs fixes from Namjae Jeon:
- Make MFT extension work on existing Windows-created volumes by
dynamically reserving MFT tail records, accounting for records added
during allocation, and avoiding false -ENOSPC failures
- Repack non-resident $MFT/$ATTRIBUTE_LIST when its mapping pairs no
longer fit in the base MFT record, while propagating allocation and
writeback errors
- Serialize runlist updates with the runlist lock and restore both the
in-memory runlist and on-disk mapping pairs when allocation rollback
is required
- Propagate folio errors and harden inode failure handling by treating
interrupted reads as transient failures and discarding and unhashing
inodes whose initialization fails
- Fix the $MFTMirr write offset when mirror records span multiple
folios, preventing mirror records from overwriting the first record
with large MFT record sizes
* tag 'ntfs-for-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs:
ntfs: fix $MFTMirr write offset when it spans multiple folios
ntfs: unhash failed inode reads
ntfs: discard inodes that fail initialization
ntfs: ignore interrupted inode reads as corruption
ntfs: propagate folio errors
ntfs: protect runlist updates with the runlist lock
ntfs: account for MFT records added during allocation
ntfs: repack $MFT/$ATTRIBUTE LIST
ntfs: use dynamic MFT tail reservation
ocfs2_calc_xattr_init() used to read the default ACL off the parent inode
itself, so it could return an error from ocfs2_xattr_get_nolock(). Commit
bd7c05fb4a ("ocfs2: fix circular locking dependency in
ocfs2_init_acl()") moved that lookup before the transaction starts and
deleted the error path, but left the now vestigial 'int ret = 0'
declaration and both 'return ret' statements behind, along with an
unreachable error branch in ocfs2_mknod().
Drop the leftover variable and convert the return type to void, so the
callee states that it always succeeds and the caller no longer carries a
check that can never trigger.
No functional change.
Link: https://lore.kernel.org/20260904023751.3703334-1-joseph.qi@linux.alibaba.com
Fixes: bd7c05fb4a ("ocfs2: fix circular locking dependency in ocfs2_init_acl()")
Signed-off-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202609040247.8B3lmoqX-lkp@intel.com/
Cc: Mark Fasheh <mark@fasheh.com>
Cc: Joel Becker <jlbec@evilplan.org>
Cc: Junxiao Bi <junxiao.bi@oracle.com>
Cc: Changwei Ge <gechangwei@live.cn>
Cc: Jun Piao <piaojun@huawei.com>
Cc: Heming Zhao <heming.zhao@suse.com>
When a secondary channel is no longer supported by the server,
cifs_chan_skip_or_disable() drops the channel reference with
cifs_put_tcp_session() and then continues to use the server pointer by
calling cifs_signal_cifsd_for_reconnect() on it and reading its
primary_server pointer. cifs_put_tcp_session() can drop the last
reference of the channel and tear it down, so both the channel and the
primary server (whose reference is also dropped by
cifs_put_tcp_session()) can be freed before they are signaled for
reconnect.
Signal the channel and the primary server and capture the primary
server pointer before dropping the channel reference with
cifs_put_tcp_session().
Fixes: f591062bdb ("cifs: handle servers that still advertise multichannel after disabling")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
In cifs_query_reparse_point(), the start >= end check before casting to
struct reparse_data_buffer * only ensures the start pointer is within the
response. It fails to verify that there is enough space remaining for the
fixed 8-byte header of the structure.
If a server provides a DataOffset that leaves less than 8 bytes remaining,
the check passes, but subsequent reads of ReparseTag and ReparseDataLength
will occur out-of-bounds.
Fix this by ensuring the remaining space is at least the size of the
reparse_data_buffer structure before accessing its fields.
Fixes: 56e84c64fc ("cifs: Fix validation of SMB1 query reparse point response")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
If snapshot_array_size is smaller than GMT_TOKEN_SIZE,
smb3_enum_snapshots() sets ret_data_len to
sizeof(struct smb_snapshot_array) without verifying the actual length
of the server's reply.
Because SMB2_ioctl() places no lower bound on the server-supplied
OutputCount and allocates retbuf to exactly that length, a short reply
results in ret_data_len exceeding the size of retbuf. The subsequent
copy_to_user() then reads past the end of retbuf, leaking adjacent slab
memory to userspace. The subsequent clamp check is ineffective as it
only reduces ret_data_len.
Fix this by rejecting replies shorter than
sizeof(struct smb_snapshot_array) with -EIO. Note that the bound is set
to the 12-byte struct size rather than the 16-byte
MIN_SNAPSHOT_ARRAY_SIZE defined in MS-SMB2 3.3.5.15.1, because 12 bytes
is exactly what copy_to_user() attempts to read.
Fixes: e02789a53d ("smb3: enumerating snapshots was leaving part of the data off end")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
In parse_posix_sids(), sidsbuf_end is calculated using the server-supplied
out_len without being validated against the actual length of the received
iov (iov_len).
If a server provides an inflated out_len, sidsbuf_end will point past the
end of the iov. This defeats the bounds guards in posix_info_sid_size(),
allowing out-of-bounds reads into adjacent kernel memory.
Fix this by rejecting responses where the calculated sidsbuf_end would
exceed the received iov boundaries or cause pointer wraparound.
Fixes: a90f37e3d7ac ("smb: client: parse owner/group when creating reparse points")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
In move_smb2_ea_to_cifs(), the while (src_size > 0) loop condition is
insufficient. It allows iteration to continue even if the remaining
src_size is too small to contain a complete smb2_ea_info structure.
Consequently, reads of ea_name_length and ea_value_length can occur
out-of-bounds.
Fix this by ensuring src_size >= sizeof(*src) before attempting to read
any structure fields. Additionally, reject any next_entry_offset that is
smaller than sizeof(*src) or that would advance the pointer beyond the
available buffer.
Note that for calls where the server returns a malformed EA list, the
error returned to userspace changes from -ENODATA (getxattr) or
-ERANGE (listxattr) to -EIO. This correctly signals a server protocol
error rather than misleadingly indicating "attribute not present" or
"output buffer too small".
Fixes: 95907fea4f ("cifs: Add support for reading attributes on SMB2+")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
In parse_server_interfaces(), the server-supplied Next offset is
validated against bytes_left, but not against the size of the interface
structure itself.
A small, non-zero Next value can pass the bounds check but advance the
pointer by less than sizeof(*p). This causes the next iteration of the
loop to read misaligned, overlapping structure fields.
Fix this by ensuring the Next offset is at least sizeof(*p).
Fixes: 7d34ec36ab ("smb3: fix for slab out of bounds on mount to ksmbd")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
parse_dfs_referrals() checks that DfsPathOffset and NetworkAddressOffset
do not exceed the buffer end, but fails to check that they don't point
inside the referral header itself.
If a server provides an offset smaller than
sizeof(struct dfs_referral_level_3), the derived string pointer overlaps
with the struct fields, causing cifs_strndup_from_utf16() to interpret
header data as UTF-16 strings.
Fix this by enforcing that string offsets are at least sizeof(*ref).
Fixes: 4ecce920e1 ("CIFS: move DFS response parsing out of SMB1 code")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
In receive_encrypted_standard(), server->total_read is left at the
full decrypted frame size when walking sub-PDUs of a compound encrypted
frame. As a result, cifs_handle_standard() passes this full size
to smb2_check_message(), causing the PDU length guards to incorrectly
validate the entire compound frame instead of the current sub-PDU.
This allows truncated non-last sub-PDUs to bypass length validation,
leading to out-of-bounds reads in smb2_get_data_area_len().
Fix this by setting server->total_read to the true length of the
current sub-PDU: next_cmd for non-last sub-PDUs, and the remaining
pdu_length for the last one.
Fixes: b24df3e30c ("cifs: update receive_encrypted_standard to handle compounded responses")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
__smb2_calc_size() calls smb2_get_data_area_len(), which reads
command-specific struct fields to locate the data area. However,
smb2_check_message() only validates StructureSize2, meaning a truncated
response could cause smb2_get_data_area_len() to read out-of-bounds.
Replace has_smb2_data_area[] with smb2_min_pdu_len[], which is now
used to indicate both whether a command's response has a data area
and the size of that fixed response struct. A non-zero entry means
the command has a data area, and is the minimum length required
before the struct is read.
For each command with a data area, PDUs shorter than this minimum size
are rejected instead of parsed.
The minimum is not applied to SMB2 error responses, which carry only
the 9-byte error body, the same exemption the StructureSize2 check
above it already makes. STATUS_MORE_PROCESSING_REQUIRED is
treated as a normal reply, since an in-progress SESSION_SETUP
response carries a full body and a security blob.
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Fix several related bounds checking and pointer lifecycle issues in
receive_encrypted_standard()'s handling of compound encrypted frames:
- Clear next_buffer after assigning it to server->bigbuf. A stale
next_buffer pointer can lead to a use-after-free on subsequent
error paths.
- Update pdu_length to the decrypted plaintext size (buf_size). Using
the pre-decryption length allows NextCommand to point into stale
ciphertext residue.
- Reject next_cmd values smaller than MID_HEADER_SIZE(server).
- Fix an integer overflow in the upper bound check by verifying
pdu_length - next_cmd < MID_HEADER_SIZE(server), ensuring the
trailing slice is large enough for a header.
Fixes: b24df3e30c ("cifs: update receive_encrypted_standard to handle compounded responses")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Commit c2a74ed049 ("btrfs: derive f_fsid from on-disk fsid and dev_t")
mixed dev_t into f_fsid for all single-device setups to avoid f_fsid
collisions with cloned filesystems.
However, doing this unconditionally breaks backward compatibility.
statfs(2) f_fsid changes after a kernel upgrade, and also can shift
across reboots or dev re-attaches as dev_t values change.
Fix this by only mixing dev_t when temp_fsid is active. This means for
non-temp_fsid setups or the original mount, we use the old method of
deriving fsid based on the UUID.
So in the case of a cloned Btrfs filesystem, we won't be able to
maintain the same fsid across mount recycle if the mount order changes.
Reported-by: Dave Hansen <dave.hansen@intel.com>
Link: https://lore.kernel.org/linux-btrfs/be0c08f5-2f31-40f5-8a3b-f2f58b3e00ff@intel.com
Fixes: c2a74ed049 ("btrfs: derive f_fsid from on-disk fsid and dev_t")
CC: stable@vger.kernel.org # 7.2
Signed-off-by: Anand Jain <asj@kernel.org>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
[BEHAVIOR CHANGE]
Since commit 108cc87339 ("btrfs: fix a lockdep caused by path
resolution during device scan"), users with btrfs rootfs but without an
initramfs are complaining that grub2 can no longer detect the rootfs
device:
/usr/sbin/grub-probe: error: cannot find a device for / (is /dev mounted?).
[CAUSE]
Although using btrfs without an initramfs is not recommended (if a new
device is added to the rootfs, the system can no longer boot, as there
is no way to register all devices), there is still a minority of users
doing this.
If there is no initramfs but the rootfs is on a block-device-based
filesystem, the kernel boot sequence initializes a minimal ramfs/tmpfs,
creates "/dev/root" with the proper device number for the rootfs, and
then invokes mount using "/dev/root".
That's why the end user will get the mount output:
/dev/root on / rw
To be honest, this is a user space problem: no one should trust the
device path shown in mount, only the device number.
E.g. one can even use "/proc/self/fd/*" to mount an fs, and that proc
path will be registered, and no one else can mount that fs using that
path.
Before commit 108cc87339 ("btrfs: fix a lockdep caused by path
resolution during device scan"), btrfs had an internal path lookup
workaround to address such weird paths, it works by checking if the
existing device path can still resolve to the device number.
But that path resolution is deadlock prone, thus it's replaced by a
simple devt check.
This works fine in most cases, as a btrfs device is registered by udev at
boot time, thus all paths are sane.
However this will not work for systems without an initramfs, causing the
unreachable "/dev/root" path to exist forever without a way to rename it.
[WORKAROUND]
Add an exception to the device path rename requirement.
If the device has the name "/dev/root", we know it's booted without an
initramfs, and only for that case we allow device path update.
And if someone intentionally created "/dev/root" after boot, the
existing devt checks will reject that weird name as usual.
This should satisfy the minority of users, and still keep most of the
existing guards preventing unexpected/unnecessary device path updates.
Fixes: 108cc87339 ("btrfs: fix a lockdep caused by path resolution during device scan")
Link: https://lore.kernel.org/linux-btrfs/CAKLYgeL7nrA4nXcewdv9Fqg_s=3GS=vmoypnEiZBKQ7rySZFuQ@mail.gmail.com/
Link: https://lore.kernel.org/linux-btrfs/dfbe1e27-dab8-4d55-8cf3-0b28eeac5df4@gmail.com/
Signed-off-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
We checked if have enough remaining space for a key before dereferencing a
key, but we then dereference a chunk item, to get the number of stripes,
without checking if there is space for the item. So add a check to see if
there is enough space for a chunk item before dereferencing the item to
extract the stripe count.
Fixes: 2a9bb78cfd ("btrfs: validate system chunk array at btrfs_validate_super()")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
If we fail to update the inode we error out without aborting the
transaction, which can result in a persistent inconsistency if after
the failure the transaction is committed, as we have dropped file
extent items from a range and either punched a hole or insert a new file
extent item for that range (for reflinks).
So add the missing transaction abort.
Fixes: 2aaa665581 ("Btrfs: add hole punching")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
btrfs_rebuild_free_space_tree() sets BTRFS_FS_CREATING_FREE_SPACE_TREE
before rebuilding the free space tree. Several error paths return
without clearing this flag.
The transaction restart failure path can leave the flag set on a live
filesystem, causing delayed reference processing to be skipped. Clear it
on all free space tree rebuild failure paths. Keep
BTRFS_FS_FREE_SPACE_TREE_UNTRUSTED set, since a failed rebuild leaves
the free space tree untrusted. Callers must fall back to extent-tree
caching.
Fixes: 882af9f13e ("btrfs: handle free space tree rebuild in multiple transactions")
CC: stable@vger.kernel.org # 6.14+
Assisted-by: LLM
Reviewed-by: Boris Burkov <boris@bur.io>
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Guanghui Yang <3497809730@qq.com>
Signed-off-by: David Sterba <dsterba@suse.com>
LOLLM noticed that the inode btree root formatting methods copy too many
bytes -- there's only one set of keys in node blocks, not two. This
causes memory corruption of whatever's beyond the buffers.
Cc: stable@vger.kernel.org # v6.14
Fixes: f0415af60f ("xfs: wire up a new metafile type for the realtime refcount")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM points out that if the amount of fdblocks that we can reserve for
a metadata btree file goes below the space already used by that file,
then the hidden_space subtraction can underflow, causing
xfs_dec_fdblocks to subtract a huge amount of space. We never want the
target to be less than the used sapce, so fix the logic that adjusts
dblocks_avail downwards.
Also fix an error in the adjacent comment.
Cc: stable@vger.kernel.org # v6.15
Fixes: 1df8d75030 ("xfs: make metabtree reservations global")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM noticed that xfs_qm_flush_one fails to drop the dquot flush lock
if it can't grab the buffer associated with the dquot. Since there's no
buffer, nobody else is going to drop the dqflock, so we need to do it
ourselves.
Cc: stable@vger.kernel.org # v6.13
Fixes: ca378189fd ("xfs: convert quotacheck to attach dquot buffers")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM pointed out a bug in xlog_recover_iunlink_bucket:
1. We don't null out prev_ip after releasing it, which can lead to UAF
problems if the inodegc flush call in the loop fails.
at which point I noticed even more bugs:
2. If the inodegc flush inside the loop fails, we also leak @ip.
3. We set prev_agino to agino having already advanced agino, which
results in inodes with i_prev_unlinked set to itself.
4. If we exit the bottom of the loop with prev_ip set, then prev_ip
aliases ip and we also set its i_prev_unlinked to itself.
Bugs 3 and 4 introduce loops into the unlinked list, though these loops
don't surface because we immediately flush each unlinked inode after
loading it.
Fix all of these issues.
Cc: stable@vger.kernel.org # v6.0
Fixes: 04755d2e58 ("xfs: refactor xlog_recover_process_iunlinks()")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM noticed the copy-paste error here -- if user quotas aren't
enforced but we're near the group quota limit, we fail to set FLAG_GID
and hence we might not actually free any preallocations, causing
unnecessary EDQUOT. Fix that.
Cc: stable@vger.kernel.org # v5.12
Fixes: c237dd7c70 ("xfs: flush eof/cowblocks if we can't reserve quota for inode creation")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM noticed that we can accidentally merge file range health
monitoring events even if they have different errors. We shouldn't do
that.
Cc: stable@vger.kernel.org # v7.0
Fixes: dfa8bad3a8 ("xfs: convey file I/O errors to the health monitor")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM complains that a memory allocation failure in
xrep_newbt_add_blocks results in online repair leaking blocks that were
previously allocated to write a new btree, but the problem is worse than
that -- a limitation of the codebase is that the callers cannot undo the
transaction /and/ return the error -- either you undo all changes and
commit the transaction, or you error out and the filesystem goes down.
However, the new btree space reservation object isn't that big (~48
bytes). Let's just do a NOFAIL allocation and the problem goes away.
Cc: stable@vger.kernel.org # v6.8
Fixes: be40841763 ("xfs: implement block reservation accounting for btrees we're staging")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM points out that runtime errors (e.g. ENOMEM) when we're trying to
compute space usag bitmaps are silently dropped by the rmapbt scrubber.
We ought to flag that as an incomplete scrub instead of reporting
cross-referencing errors based on faulty data.
Cc: stable@vger.kernel.org # v6.4
Fixes: fed050f345 ("xfs: cross-reference rmap records with ag btrees")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
When I added online fsck for realtime reflink, I forgot to update
xrep_calc_rtgroup_resblks to factor in the size of the refcount btree
when it guesses how much space we need to start a repair. This hasn't
been a huge problem in practice because there are few filesystems with
(a) realtime, (b) rtgroups, (c) reflink, and (d) no rmap. But let's fix
this before someone stumbles upon it, especially since LOLLM flagged
this for me.
Cc: stable@vger.kernel.org # v6.14
Fixes: 83ccffc489 ("xfs: online repair of the realtime refcount btree")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
Now that we have quotas for the realtime volume, we also have
precomputed watermark limits for the realtime block counts. These
precomputations should be done any time we change the rtb limits, which
means that xfs_qm_adjust_dqlimits needs to ensure that if we installed
a default rtb limit.
Cc: stable@vger.kernel.org # v6.13
Fixes: 5dd70852b0 ("xfs: create quota preallocation watermarks for realtime quota")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
cifs_try_adding_channels() iterates ses->iface_list with
list_for_each_entry_safe_from(), which captures the next entry
(niface) under iface_lock. The loop body then drops iface_lock for
the whole duration of cifs_ses_add_channel().
A concurrent interface refresh (SMB3_request_interfaces() ->
parse_server_interfaces()) marks all ifaces inactive and removes and
frees any that are not re-advertised via list_del() + kref_put(),
where release_iface() is a bare kfree(). Since niface typically has
no channel holding a reference, the list reference is its last and it
can be freed inside the unlocked window. On continue, the iterator
advance step then dereferences niface->iface_head.next, and the loop
body reads iface->rdma_capable/is_active, both on freed memory.
Fix this by never keeping an unreferenced list pointer across the
unlocked window. Each channel attempt now re-scans the list from the
head under iface_lock, takes a kref on the selected candidate, and
passes only that referenced candidate to cifs_ses_add_channel().
weight_fulfilled still tracks selection progress, so restarting the
scan preserves the original weighted distribution and the
weight_fulfilled-before-kref_put ordering on the failure path.
Add a per-pass attempts cap so a flapping interface refresh cannot
keep the inner loop spinning within a single tries increment.
Fixes: aa45dadd34 ("cifs: change iface_list from array to sorted linked list")
Cc: stable@vger.kernel.org
Assisted-by: Qoder:Qwen3.8-Max
Signed-off-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Acked-by: Shyam Prasad N <sprasad@microsoft.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
A per-thread CPU timer holds a reference to the PID of the thread it is
attached to and, while it is armed, its node is queued in that thread's
posix_cputimers. The task is looked up by that PID.
When a non-leader thread exec()s, de_thread() changes which task owns
that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL,
but the node is still queued on tsk, which is alive. timer_lock_sighand()
takes a failed lookup to mean that the node is already dequeued, so it
has nothing to undo.
begin_new_exec() calls posix_cpu_timers_exit(me) right after
exec_task_namespaces() and that removes the leftover node, so the state
normally stays invisible. But bprm->point_of_no_return is set before
de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or
exec_task_namespaces() fails, the task dies before it gets there.
exit_itimers() then frees the k_itimer while its node is still queued,
and reaping tsk later erases that freed node from the rbtree.
In short:
the non-leader thread B the parent
timer_create(CLOCK_THREAD_CPUTIME_ID)
timer_settime()
arm_timer() // the node is queued on B
execve()
de_thread(B)
exchange_tids(B, leader) // B's PID now belongs to the leader
release_task(leader)
__exit_signal(leader)
posix_cpu_timers_exit(leader) // cleans leader's queue, not B's
__unhash_process(leader) // that PID has no task anymore
exec_mmap()
mmap_read_lock_killable(old_mm)
kill(B, SIGKILL)
// -EINTR
get_signal()
do_exit()
exit_itimers()
posix_timer_delete()
posix_cpu_timer_del()
posix_timer_unhash_and_free() // freed while still queued
wait4()
release_task(B)
posix_cpu_timers_exit(B)
cleanup_timerqueue()
timerqueue_del() // use-after-free
Move the POSIX timer cleanup right after de_thread() before any of the
later failure conditions brings the task into do_exit().
[ tglx: Move the cleanup right after de_thread() ]
Fixes: 55e8c8eb2c ("posix-cpu-timers: Store a reference to a pid not a task")
Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Tested-by: Kijo Park <red993688@gmail.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/ao7Q8miiuLAPVnWv@v4bel
Link: https://patch.msgid.link/20260911090541.627712075@kernel.org
When enable_verity() hits the qgroup limit, rollback_verity() needs its
own metadata reservation. When the qgroup limit or lack of space refuses
the rollback, the whole filesystem is forced read-only even though the
qgroup limit was for one subvolume only. Also orphan cleanup at the next
mount fails the same way, so the leftover items are never removed: with
-EDQUOT the subvolume stays unreachable, and with -ENOSPC on a full
filesystem the next read-write mount fails.
Start transactions with btrfs_start_transaction_fallback_global_rsv() in
btrfs_orphan_cleanup(), drop_verity_items() and rollback_verity(). Those
calls only delete items and free the space in the end, so they may use
the global reserve and skip the qgroup limit, which avoids -ENOSPC and
-EDQUOT.
Fixes: 146054090b ("btrfs: initial fsverity support")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Daniel Linjama <daniel@dev.linjama.com>
Signed-off-by: David Sterba <dsterba@suse.com>
If the compressed data of an inline extent is larger than or equals to the
size of the uncompressed data, we are still allowing the creation of the
compressed inline extent, which does not result in any benefits, quite the
contrary as we waste metadata space and have to decompress when reading.
This is a recent regression introduced in commit 3eaf5f082c ("btrfs:
extract inlined creation into a dedicated delalloc helper").
It happens because we are passing the block size to btrfs_compress_bio(),
so we don't get -E2BIG from the compression code anymore, but we can not
pass i_size either, because if i_size is smaller than sector size, we
end up never creating lzo compressed inline extent for such small i_size
values. So refuse the compressed result at run_delalloc_inline() if
its size is not smaller than the uncompressed size (i_size).
Reported-by: Hanabishi <i.r.e.c.c.a.k.u.n+kernel.org@gmail.com>
Link: https://lore.kernel.org/linux-btrfs/c97652a5-ac6b-4de6-aa23-3cdebc01d00b@gmail.com/
Fixes: 3eaf5f082c ("btrfs: extract inlined creation into a dedicated delalloc helper")
CC: stable@vger.kernel.org # 7.1+
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
Do not reset the RFC1002 length of the complete response when a query
info buffer is too small. The current command will add its error response
through ksmbd_iov_pin_rsp(), while resetting the base length can truncate
earlier responses in a compound request.
This lets ksmbd return the earlier responses and the query-info error
response together. Remove the now-unused rsp_org parameter from the pipe
query-info helpers.
Fixes: e2b76ab8b5 ("ksmbd: add support for read compound")
Reported-by: Mobin Aydinfar <mobin@mobintestserver.ir>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
Windows may request FILE_NORMALIZED_NAME_INFORMATION with an output
buffer that only fits the fixed portion of the variable-length response.
Treat the fixed portion as FILE_NORMALIZED_NAME_INFORMATION_SIZE so ksmbd
returns STATUS_BUFFER_OVERFLOW instead of STATUS_INFO_LENGTH_MISMATCH.
This avoids rejecting valid partial normalized-name responses.
Fixes: 6b8b79226b ("ksmbd: fix partial file information responses")
Reported-by: Mobin Aydinfar <mobin@mobintestserver.ir>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
Correct ten misspellings and eight accidentally doubled words in
comments. No code changes.
The doubled word in xfs_zone_alloc.c was not a duplicate: "so that is is
reused" is "it is" misspelt, so that one reads "so that it is reused"
rather than dropping a word.
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
The implementation was inlined into xfs_file_remap_range() by commit
3fc9f5e409 ("xfs: remove xfs_reflink_remap_range"), leaving this
declaration orphaned.
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
Commit a23eca8844 ("xfs: fix exchange-range reflink flag clearing
issue with INO1_WRITTEN") duplicated commit b2d5a81dae ("xfs: fix
exchange-range reflink flag clearing issue with INO1_WRITTEN"), so
xmi_can_exchange_reflink_flags() ended up with two identical
XFS_EXCHMAPS_INO1_WRITTEN checks. The second one is dead code,
since the first one already returned false. Remove it.
Signed-off-by: Jiangshan Yi <yijiangshan@kylinos.cn>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
oz can be NULL when we resample it after taking i_flags_lock, so account
for that.
Fixes: 2d829cc767 ("xfs: fix racy open zone caching")
Signed-off-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Hans Holmberg <hans.holmberg@wdc.com>
Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
The di_forkoff check in xchk_dinode is incorrect, according to LOLLM.
XFS_DFORK_BOFF returns a byte count relative to the start of the literal
area, not the start of the inode. Therefore, this check won't flag
di_forkoff values that are larger than the literal area but not the
inode size itself. Fix this check; sadly the old APTR code was correct.
Cc: stable@vger.kernel.org # v6.8
Fixes: 6b5d917780 ("xfs: dont cast to char * for XFS_DFORK_*PTR macros")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM complains that xchk_directory_data_bestfree can be passed a
directory block that is either in "block" or "data" format, but the
check here unconditionally treats the dir3_block and dir3_data blocks as
if they have the same header format (they don't). Consequently, we can
incorrectly set the preen state on dir3_block blocks, which of course
we can't preen away because dir3_block blocks do not have a padding
field. Fix this.
Cc: stable@vger.kernel.org # v7.1-rc4
Fixes: 939919ccdd ("xfs: check directory data block header padding in scrub")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM complains that the xbitmap set functions can suffer an integer
underflow or overflow and thereby return the wrong left and right
pointers. Fix that logic bomb, even though (AFAICT) we never actually
try to set the *entire* bitmap.
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM noticed that we might reserve the wrong number of blocks for
recovering rtrmap and rtrefcount updates after a crash. Fix that.
Cc: stable@vger.kernel.org # v6.14
Fixes: 5e0679d1c6 ("xfs: support recovering rmap intent items targetting realtime extents")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM noticed that we strip file privileges and whatnot even for a dry
run. We also shouldn't flush dirty data to disk or trim COW staging
events for a dry run. Neither of those behaviors are allowed by the
manpage, so fix that by exiting early on DRY_RUN in various functions.
Cc: stable@vger.kernel.org # v6.10
Fixes: 42672471f9 ("xfs: bind together the front and back ends of the file range exchange code")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM points out that we don't check the ioctl padding field here, so
let's do that. I don't think there are many users yet since exchrange
requires a new feature flag, so it's a good time to try to plug this
hole.
Cc: stable@vger.kernel.org # v6.12
Fixes: 398597c3ef ("xfs: introduce new file range commit ioctls")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
LOLLM points out that we're supposed to use time_after_eq, not a raw >=
operation here, or else jiffies wraps can go unnoticed. Fix this.
Cc: stable@vger.kernel.org # v6.10
Fixes: 271557de7c ("xfs: reduce the rate of cond_resched calls inside scrub")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>