Commit Graph

108549 Commits

Author SHA1 Message Date
Linus Torvalds
a52a93358a 14 hotfixes. 10 are cc:stable. 11 are for MM.
Five are DAMON fixes.  One fixes an arm64 contpte bug where DAMON can
 write past the end of a page-table page, resulting in memory corruption
 and possible crashes.
 
 Two are hugetlb fixes.  One fixes an mremap() address calculation bug
 which can panic x86-64.
 
 There's also a missing anon_vma publication barrier which can result in
 hung tasks, and a writeback fix to keep long cgroup writeback drains from
 delaying Tasks-RCU grace periods.
 
 The remainder are smaller fixes and maintenance changes.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTTMBEPP41GrTpTJgfdBJ7gKXxAjgUCarHHNAAKCRDdBJ7gKXxA
 jhb7AP4yE/k7RrZC6zWg4M9ejI3fVlHI9+EG1mCLGiV57jZ4mAEA9LLSZryOD6Nc
 zsaeCtZhHcEdxW6EhO18hMfH4b7oJA0=
 =alhC
 -----END PGP SIGNATURE-----

Merge tag 'mm-hotfixes-stable-2026-09-21-17-08' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm

Pull MM fixes from Andrew Morton:
 "14 hotfixes.  10 are cc:stable.  11 are for MM.

  Five DAMON fixes: one fixes an arm64 contpte bug where DAMON can write
  past the end of a page-table page, resulting in memory corruption and
  possible crashes.

  Two hugetlb fixes: one fixes an mremap() address calculation bug which
  can panic x86-64.

  There's also a missing anon_vma publication barrier which can result
  in hung tasks, and a writeback fix to keep long cgroup writeback
  drains from delaying Tasks-RCU grace periods.

  The remainder are smaller fixes and maintenance changes"

* tag 'mm-hotfixes-stable-2026-09-21-17-08' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
  MAINTAINERS: update Xu Xin's email
  writeback: report a Tasks-RCU quiescent state per cgwb drain pass
  mm/damon/core: reset invalid quota->charge_target_from
  MAINTAINERS: add Baoquan and Baolin as MGLRU reviewers
  mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish
  mm/hugetlb: preserve mremap address delta when skipping page tables
  mm/damon/core: fix unconditionally skip last region
  mm/damon/vaddr: avoid hw-driven pte updates during damon_hugetlb_mkold()
  mm/damon/core: allow esz to be set to zero
  mm/damon/ops-common: use a page-aligned address in damon_ptep_mkold()
  ocfs2: make ocfs2_calc_xattr_init() return void
  mailmap: update Haowen Bai's email address
  selftests/cgroup: account for zswap shrinker writeback
  mm/hugetlb: do not dissolve gigantic pages without runtime support
2026-09-22 10:03:33 -07:00
Linus Torvalds
f0100363d8 xfs: fixes for v7.3-rc5
Signed-off-by: Carlos Maiolino <cem@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iJUEABMJAB0WIQSmtYVZ/MfVMGUq1GNcsMJ8RxYuYwUCarEtCAAKCRBcsMJ8RxYu
 YwmHAX4+ML995OtOeKIN2Cpiu/0RYW/bHPrBBCHsiZN2BFZ3Ap2W7nxm4/OFewA9
 5oB6eKgBgL8Zi2K0ZPsiF+IxAEkWUBkvCgiQtByO3dSS6aMR6KzfrQBj0Pftlcgy
 z+C74fVlyA==
 =USzk
 -----END PGP SIGNATURE-----

Merge tag 'xfs-fixes-7.3-rc5' of gitolite.kernel.org:/pub/scm/fs/xfs/xfs-linux

Pull xfs fixes from Carlos Maiolino:
 "This mostly contain 'random' bugfixes found by LLM tools on different
  xfs subsystems. A few code cleanups and a NULL ptr deref on zoned
  support. Those 'random' bugfixes include possible buf overrus, UAFs,
  block leaks, etc..."

* tag 'xfs-fixes-7.3-rc5' of gitolite.kernel.org:/pub/scm/fs/xfs/xfs-linux: (26 commits)
  xfs: fix wild memcpy access when formatting ondisk rtrefcount btree roots
  xfs: don't let hidden_space go negative in xfs_metafile_resv_init
  xfs: drop dquot flush lock when we can't find a buffer to flush
  xfs: fix cursor and pointer handling when recovering iunlink buckets
  xfs: fix blockgc group quota scanning when usrquota isn't enforced
  xfs: don't merge different file IO error types
  xfs: don't let memory failures leak blocks and kill repairs
  xfs: don't cross reference rmapbt with bitmaps if they're incomplete
  xfs: fix rtgroup repair estimations
  xfs: call xfs_dquot_set_prealloc_limits if we installed default rtb limits
  xfs: fix typos and repeated words in comments
  xfs: remove unused xfs_reflink_remap_range declaration
  xfs: remove duplicate INO1_WRITTEN check
  xfs: don't try to get a reference to a NULL oz in xfs_get_cached_zone
  xfs: check di_forkoff correctly in scrub
  xfs: only flag zero padding for dir3 data blocks, not dir3 block blocks
  xfs: fix integer overflows in xbitmap set functions
  xfs: use the correct reservations for rtrmap/refcount recovery
  xfs: don't call xfs_exchange_range_finish for a dry run
  xfs: check padding field in xfs_ioc_commit_range
  ...
2026-09-21 08:22:53 -07:00
Linus Torvalds
0a15ba6b0c Timer race fixes:
- Fix timer signal <-> exec() race, to prevent UAF (Thomas Gleixner)
 
  - Clean up POSIX CPU timers right after de_thread(), to prevent UAF
    (Hyunwoo Kim)
 
  - Fix POSIX CPU timers race between expiry and timer_settime(),
    to prevent UAF (Thomas Gleixner)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqvrMYRHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1gqUQ/+NkruN984bFynF/eZ0/2DFv91AAUP8zgH
 /S3PBlwuSbYFN9JVhngDMwxQkamE56weJbFc+0QuvVT5UVw/vX9BS4QOvvzN+f8D
 FEN3UqD0d1B8OwlPNTw0sFPwJDdPctTinfKOhNjNQe6RLFsNARvGyaKDIDroWTfV
 dxuJ/7Ecs+5m1bmGJnPEC+IH/OnV9BEEl1NdZb+INKpBlui9LCsw4rRIj/8dPK/H
 UNhvXpykKrJCDftbCzAFSNryuzcJgq4kHtMbsqiUL6y50AB69eHGi/Y0xYBAEr1h
 NiDPq2PAMmH1NCCMsTtqbJZMqgCr+7DSZiCFn7bZPwg0V5tV4PFZD484q0sCbiej
 Fwg+arHd0icnceIcWMsBWPUVOSLxZaWdp9a2Tj3Ill06//b5bEDBJBbpecS+so3t
 8W6IvdoCYm7sz50mohnjOdx7biHPu0yhwgj+EoAV3nZKoALQAAcI7+HJzSWpGnJi
 HIO0zylRAZCjk9H3QNWO+LdWgifc8DysAZOWpmbuwGgp8q483IDRDtme/kMt3+D1
 1qTHa1TD/tPo8UmmgyVJQ7e1hCxBkGuuBBu5Y3/qkUEOQM6B/H2Ji7stxsLpW3JL
 HLzC3kL2SBBVBO2ljqiH5IhVAL10Qm5vPxaCOjExdBt1vjMxN1DowZqD4rT/tw58
 ArpD4zmr8VQ=
 =KhWJ
 -----END PGP SIGNATURE-----

Merge tag 'timers-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull timer race fixes from Ingo Molnar:

 - Fix timer signal <-> exec() race, to prevent UAF (Thomas Gleixner)

 - Clean up POSIX CPU timers right after de_thread(), to prevent UAF
   (Hyunwoo Kim)

 - Fix POSIX CPU timers race between expiry and timer_settime(),
   to prevent UAF (Thomas Gleixner)

* tag 'timers-urgent-2026-09-20' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
  exec: Cleanup POSIX timers right after de_thread()
  signal: Prevent exec() race
2026-09-20 09:41:00 -07:00
Linus Torvalds
518e5b794c for-7.3-rc3-tag
-----BEGIN PGP SIGNATURE-----
 
 iQJPBAABCgA5FiEE8rQSAMVO+zA4DBdWxWXV+ddtWDsFAmqu3wQbFIAAAAAABAAO
 bWFudTIsMi41KzEuMTIsMiwyAAoJEMVl1fnXbVg7Y5kQAJ1ANaQWod+AUKhgbtA6
 IcEFH8AAVrrTqqN0SxOl/tqHL6Xt/5mKlQcTpYPxeccUkp72M9BeGik4Gp7OwN1D
 vkVTgrmhHT4r+4ae4GJP0yCtNJBa0fRXjsMFbNYTKGWcz9yOsW1OkBsq8VtRo7ym
 CSVBc57buUi0mzbnLNDh69G/YA7NCTyaxXKjPARNYy+cy0LMIDmgZCByhgePQvzB
 aqJUakRKpaeXEQIf0nT/70XcGhXNtfK1GOnLZM1ySFqLufMzdTsEzFsT8dVugqU1
 wr1HMaMq1iNKwE4sJgNWS84wRT1zxZopKIOsTufFu98Zb2C0kvUSjWeJSqtFM88G
 ggj/jmaoRhCU1cXE1jvZWy4Fe5zQH8deSQZ7zUB4ZzqCaDRwOEAj4IpuQQ9Ok91E
 J/07SCvKPk/QUPbA2e5lwRL3aximsD1LfRxWIOZ+xg/HVX+vYfHmy5gzkl/hhfrm
 RHkHLz8iXNHV7qhIVzZ/GYvTxR6U6nbLVUbkl5n/8eFePM7YHnoWtvMHT71qWwrh
 mPx8uTK+4BI2+rAHTKqxnwEQ27OHj5odlGKTlKiaMrQR9eqnnhJLTMPDJMzqQ7TI
 ZiibxG8UqSTsNbvOXXtd7MykSRpSb/VWyJImKuw30kx3f7f2yd1aCUyU4T7o5N/3
 FRI+2AbkoSeype3Rw+VcEPoi
 =WxSA
 -----END PGP SIGNATURE-----

Merge tag 'for-7.3-rc3-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux

Pull btrfs fixes from David Sterba:
 "Among the regular fixes, there are two that were reported recently and
  have user impact:

   - filesystem id is now stable again on the default and common case
     (it broke openconnect key derivation, while this is not secure,
     it's still in use), the intention was to change id for the
     temp_fsid use case

   - fix detection of /dev/root and rename it after device scan, this
     broke booting of initramdisk-less system with grub2 as the probe
     needs the real device

  Regular fixes:

   - don't store compressed inline extent if the size is larger than
     uncompressed

   - in zoned mode, handle activation of zones for all supported block
     group profiles in case there are still free ones left

   - check space for a chunk item when reading sys array from superblock

   - allow using space reserves when removing verity items fails

   - fix error handling after free space tree rebuild fails

   - abort transaction if reflink or hole punching fails and it's not
     possible to update the inode

   - properly protect block group iteration during device replace start

   - error message fixups"

* tag 'for-7.3-rc3-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux:
  btrfs: derive f_fsid with dev_t only when temp_fsid is active
  btrfs: add "/dev/root" exception for device path update
  btrfs: check if there is space for chunk item when validating sys chunk array
  btrfs: abort transaction on failure to update inode for hole punching and reflinking
  btrfs: clear free space tree creation state on rebuild failure
  btrfs: handle lack of space when cleaning up verity items
  btrfs: fix creation of compressed inline extents that don't save space
  btrfs: tree-checker: fix error message regarding free space extent items
  btrfs: tree-checker: print dev extent offset in error message
  btrfs: take commit root semaphore when iterating in mark_block_group_to_copy()
  btrfs: zoned: handle RAID profiles in btrfs_can_activate_zone()
2026-09-19 13:48:20 -07:00
Linus Torvalds
17e7b8eacf smb client fixes for v7.3-rc4
A batch of bug fixes for the smb client:
 
  - Fix multiple out-of-bounds reads and use-after-frees in the SMB2/3
    receive path that are reachable from a malicious or compromised
    server: a stale next_buffer pointer and an integer overflow in
    compound encrypted frame handling, missing minimum-PDU-size and
    per-sub-PDU length validation before parsing command-specific
    response fields, missing bounds checks in DFS referral, server
    interface list, EA list, POSIX SID, snapshot enumeration and SMB1
    reparse point parsing
 
  - Fix use-after-frees and races in multichannel and connection
    teardown, including an interface freed while still in use when
    adding channels, a server used after its channel reference was
    dropped, a reconnect work item left queued after the server is
    freed and an uninitialized reconnect list node
 
  - Fix a heap overflow in the native symlink parser: an absolute
    target without an NT drive prefix caused out-of-bounds writes and a
    u16 length underflow leading to a 64K memcpy into a small buffer,
    triggerable by a user with write access to a mounted share under
    default settings
 
  - Fix WSL reparse point parsing: use unaligned accessors for the
    packed extended-attribute payload to avoid alignment faults on some
    architectures and stop leaving partially mutated fattr fields on
    parse failure
 
  - Fix lease break ACKs being sent through the wrong session on
    multiuser mounts, which caused read failures (e.g. on NetApp
    ONTAP/Azure Files) when copying files
 
  - Fix an smbd_connection leak when cifs_get_tcp_session() fails after
    an RDMA connection was already established
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTcqRusfSdYROJQwGkpVtNKoQNdYwUCaq2frwAKCRApVtNKoQNd
 Y1mcAQDcDTkep03jzghyJG6xWJ3S7KNbeYpjkOPnPyR+Et7HmAD/eXLFvgkJ3wC7
 tBUDjDTLeyP6/DOBmDb/fIKEw2vfBQs=
 =+Vsw
 -----END PGP SIGNATURE-----

Merge tag 'cifs-fixes-7.3-rc4' of https://git.manguebit.org/linux

Pull smb client fixes from Paulo Alcantara:
 "A batch of bug fixes for the smb client:

   - Fix multiple out-of-bounds reads and use-after-frees in the SMB2/3
     receive path that are reachable from a malicious or compromised
     server: a stale next_buffer pointer and an integer overflow in
     compound encrypted frame handling, missing minimum-PDU-size and
     per-sub-PDU length validation before parsing command-specific
     response fields, missing bounds checks in DFS referral, server
     interface list, EA list, POSIX SID, snapshot enumeration and SMB1
     reparse point parsing

   - Fix use-after-frees and races in multichannel and connection
     teardown, including an interface freed while still in use when
     adding channels, a server used after its channel reference was
     dropped, a reconnect work item left queued after the server is
     freed and an uninitialized reconnect list node

   - Fix a heap overflow in the native symlink parser: an absolute
     target without an NT drive prefix caused out-of-bounds writes and a
     u16 length underflow leading to a 64K memcpy into a small buffer,
     triggerable by a user with write access to a mounted share under
     default settings

   - Fix WSL reparse point parsing: use unaligned accessors for the
     packed extended-attribute payload to avoid alignment faults on some
     architectures and stop leaving partially mutated fattr fields on
     parse failure

   - Fix lease break ACKs being sent through the wrong session on
     multiuser mounts, which caused read failures (e.g. on NetApp
     ONTAP/Azure Files) when copying files

   - Fix an smbd_connection leak when cifs_get_tcp_session() fails after
     an RDMA connection was already established"

* tag 'cifs-fixes-7.3-rc4' of https://git.manguebit.org/linux:
  cifs: Fix server use-after-free in cifs_chan_skip_or_disable()
  smb: client: fix reparse buffer bounds in cifs_query_reparse_point()
  smb: client: fix potential OOB read in smb3_enum_snapshots()
  smb: client: fix missing iov bounds check in parse_posix_sids()
  smb: client: fix OOB struct field reads in move_smb2_ea_to_cifs()
  smb: client: reject short Next offsets in parse_server_interfaces()
  smb: client: fix missing lower-bound check on DFS referral string offsets
  smb: client: fix server->total_read for compound encrypted PDUs
  smb: client: validate minimum PDU size before smb2_get_data_area_len()
  smb: client: fix next_buffer UAF and NextCommand bounds in compound PDUs
  smb: client: fix use-after-free of iface in cifs_try_adding_channels()
  smb: client: fix fattr leaking on wsl_to_fattr() failure
  smb: client: fix unaligned access in WSL reparse point parser
  smb: client: fix smbd_connection leak on cifs_get_tcp_session() error
  smb: client: fix rlist race and missing initialization
  smb: client: cancel reconnect work in clean_demultiplex_info()
  smb/client: send lease break ACKs thru correct session for multiuser mounts
  smb: client: validate absolute native symlink targets before NT fixups
2026-09-18 13:44:59 -07:00
Linus Torvalds
c3d85c669d - Fix session expiration so that valid sessions are no longer removed
after ten seconds of inactivity when a new session setup request is
    received. Sessions now expire only after credential expiration, while
    stale unauthenticated sessions are cleaned up after a 45-second timeout.
 
  - Keep earlier responses in compound requests when Query Info fails
    because the output buffer is too small. The error response is appended
    without truncating preceding responses.
 
  - Return STATUS_BUFFER_OVERFLOW for partial
    FILE_NORMALIZED_NAME_INFORMATION responses instead of incorrectly
    returning STATUS_INFO_LENGTH_MISMATCH.
 -----BEGIN PGP SIGNATURE-----
 
 iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqtTBMWHGxpbmtpbmpl
 b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCD2GD/9OzkmZ+/kIMum8IQi9upOM5zUD
 +2nyFMebJ6qXmrhOuUgLn2ptUYxO3q1B9VvzTjKrM1Vun0VuvHmHQbGUp9a/3F7c
 VTncl7E3FEHqlPkLWQJFv2dS+TYYcOhjoc2TDAY9093xktcMHK5zjv4E/DV2o0bf
 NN/GcrrcWSxdJU8WI9JvY2kzmPQMDfM8uVbj2RqHSZ8m9mF5cajtDnzCCS0K6UOR
 qzMrekzewIskUtcQNqU8hJWl1sgiYdD+16LmKmwLd3uOZISc3Miy5Bg8VpNN+B1g
 XHhr49G4Fb8PkEYjncjxQH7zop1ID5UyC2xN63NuoH8+mkm4qn6uG2NrLhmVQO+f
 Z81Ov3g77/jK3Z2fw00H3A7VGSPs931BaRTNj+lPkHTVVq3PyP1XZsfXkPUoRu1Q
 xeaJydScGNYE+kaYbseXwN8haJGawd1Dd+Afn4W2zikUU5tKZxO9d2tBr4Fhl/Fm
 0OBcAssTArhrY7PX2fAOQ4sUwAC4nMXSEIfdWIvcgU2OoyuFltQ55ooCoM0uLs6+
 7R4rxZLipdZmKhuE2mlAWcFQJn8nS0EHXeyEDjO0u8KYsV6C8d+jDNpvtS+/5QVF
 bKE4/uUX/lV0IZ55nFSPT7XdI+gxeiUql3+8bqOVqkw7wR4VjaC0xIQybyEBTMYH
 IJEKfjx4tZcWGTzmdA==
 =9nNl
 -----END PGP SIGNATURE-----

Merge tag 'ksmbd-for-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb

Pull smb server fixes from Namjae Jeon:

 - Fix session expiration so that valid sessions are no longer removed
   after ten seconds of inactivity when a new session setup request is
   received.

   Sessions now expire only after credential expiration, while stale
   unauthenticated sessions are cleaned up after a 45-second timeout.

 - Keep earlier responses in compound requests when Query Info fails
   because the output buffer is too small. The error response is
   appended without truncating preceding responses.

 - Return STATUS_BUFFER_OVERFLOW for partial
   FILE_NORMALIZED_NAME_INFORMATION responses instead of incorrectly
   returning STATUS_INFO_LENGTH_MISMATCH.

* tag 'ksmbd-for-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/smb:
  ksmbd: keep compound responses on query info errors
  ksmbd: fix partial normalized name responses
  ksmbd: follow SMB2 session expiration semantics
2026-09-18 11:05:55 -07:00
Linus Torvalds
bfda5a01aa - Make MFT extension work on existing Windows-created volumes by
dynamically reserving MFT tail records, accounting for records added
    during allocation, and avoiding false -ENOSPC failures.
 
  - Repack non-resident $MFT/$ATTRIBUTE_LIST when its mapping pairs no longer
    fit in the base MFT record, while propagating allocation and writeback
    errors.
 
  - Serialize runlist updates with the runlist lock and restore both the
    in-memory runlist and on-disk mapping pairs when allocation rollback is
    required.
 
  - Propagate folio errors and harden inode failure handling by treating
    interrupted reads as transient failures and discarding and unhashing
    inodes whose initialization fails.
 
  - Fix the $MFTMirr write offset when mirror records span multiple folios,
    preventing mirror records from overwriting the first record with large
    MFT record sizes.
 -----BEGIN PGP SIGNATURE-----
 
 iQJKBAABCgA0FiEE6NzKS6Uv/XAAGHgyZwv7A1FEIQgFAmqtRIEWHGxpbmtpbmpl
 b25Aa2VybmVsLm9yZwAKCRBnC/sDUUQhCBK8EAChFdTxig3sYog3dL27SX+1yNC4
 S1QtSQMvPJzcvLG2/mESC57snx1u6AXZ6enaQ/m8zQQvkWHFdwD+odgjOQ3474Yc
 MS7xQn5pCsAo3LmSWiDsQfHmvxgDMYlIRU1vmqr7fG2pj+W6BR2LB1PD3RI9exI7
 0WWAXBPZH5w5C9GE1Zo7TF9Xwby5Or31RS8+R57PXA/PJ1ivpWnlkpfLi4M/YkZK
 GIpunZafSpcKbEsuWcjhdz11bR4G9Qlwuuq0MDguLC/qsqsobHCeSbdx+4IsEAq5
 02yBl5hYm2E4u2KBedpe7oRwFvlPN0uakEGYS8SA1ad9XamjGIw6T0tkzkDL48Pq
 dhVAeX2oa8O9u+VK+qF/HIUylh/UbmHQJW8iSiZWO8WdULGBG8oCHI1hcSnMguwJ
 njyK75UXz4fMsKW6ZpRu0sRGqtKKcbg8IrCvLslPIOS2A9OAwSzytDKI+x1Kbgu0
 SVPYjf6XeOz83tvE+2OhfTT1hWkeKezMiUe4E/y9rgEDxv5vE4C5pvpDfRlj3oyn
 2JTXXjjUQGSQw/9cKbLsbElDH/FLEojLAsIFgM+2FbcG+x7PUUgSGdC4R4qZ9MUF
 cuMzfhi8vKyCYmJc8nNE4J6b6UkBNOFJMYBcArtByFW3LqfIzmqwW81Xx0Hz4cxi
 dmOhD6gXXYoi9m3lBQ==
 =rT1L
 -----END PGP SIGNATURE-----

Merge tag 'ntfs-for-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs

Pull ntfs fixes from Namjae Jeon:

 - Make MFT extension work on existing Windows-created volumes by
   dynamically reserving MFT tail records, accounting for records added
   during allocation, and avoiding false -ENOSPC failures

 - Repack non-resident $MFT/$ATTRIBUTE_LIST when its mapping pairs no
   longer fit in the base MFT record, while propagating allocation and
   writeback errors

 - Serialize runlist updates with the runlist lock and restore both the
   in-memory runlist and on-disk mapping pairs when allocation rollback
   is required

 - Propagate folio errors and harden inode failure handling by treating
   interrupted reads as transient failures and discarding and unhashing
   inodes whose initialization fails

 - Fix the $MFTMirr write offset when mirror records span multiple
   folios, preventing mirror records from overwriting the first record
   with large MFT record sizes

* tag 'ntfs-for-7.3-rc4' of git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/ntfs:
  ntfs: fix $MFTMirr write offset when it spans multiple folios
  ntfs: unhash failed inode reads
  ntfs: discard inodes that fail initialization
  ntfs: ignore interrupted inode reads as corruption
  ntfs: propagate folio errors
  ntfs: protect runlist updates with the runlist lock
  ntfs: account for MFT records added during allocation
  ntfs: repack $MFT/$ATTRIBUTE LIST
  ntfs: use dynamic MFT tail reservation
2026-09-18 11:02:08 -07:00
Joseph Qi
525c0edc03 ocfs2: make ocfs2_calc_xattr_init() return void
ocfs2_calc_xattr_init() used to read the default ACL off the parent inode
itself, so it could return an error from ocfs2_xattr_get_nolock().  Commit
bd7c05fb4a ("ocfs2: fix circular locking dependency in
ocfs2_init_acl()") moved that lookup before the transaction starts and
deleted the error path, but left the now vestigial 'int ret = 0'
declaration and both 'return ret' statements behind, along with an
unreachable error branch in ocfs2_mknod().

Drop the leftover variable and convert the return type to void, so the
callee states that it always succeeds and the caller no longer carries a
check that can never trigger.

No functional change.

Link: https://lore.kernel.org/20260904023751.3703334-1-joseph.qi@linux.alibaba.com
Fixes: bd7c05fb4a ("ocfs2: fix circular locking dependency in ocfs2_init_acl()")
Signed-off-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202609040247.8B3lmoqX-lkp@intel.com/
Cc: Mark Fasheh <mark@fasheh.com>
Cc: Joel Becker <jlbec@evilplan.org>
Cc: Junxiao Bi <junxiao.bi@oracle.com>
Cc: Changwei Ge <gechangwei@live.cn>
Cc: Jun Piao <piaojun@huawei.com>
Cc: Heming Zhao <heming.zhao@suse.com>
2026-09-18 08:20:13 -07:00
Wentao Liang
717e0a2503 cifs: Fix server use-after-free in cifs_chan_skip_or_disable()
When a secondary channel is no longer supported by the server,
cifs_chan_skip_or_disable() drops the channel reference with
cifs_put_tcp_session() and then continues to use the server pointer by
calling cifs_signal_cifsd_for_reconnect() on it and reading its
primary_server pointer. cifs_put_tcp_session() can drop the last
reference of the channel and tear it down, so both the channel and the
primary server (whose reference is also dropped by
cifs_put_tcp_session()) can be freed before they are signaled for
reconnect.

Signal the channel and the primary server and capture the primary
server pointer before dropping the channel reference with
cifs_put_tcp_session().

Fixes: f591062bdb ("cifs: handle servers that still advertise multichannel after disabling")
Cc: stable@vger.kernel.org
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 19:29:25 -03:00
Frank Sorenson
5f0306e731 smb: client: fix reparse buffer bounds in cifs_query_reparse_point()
In cifs_query_reparse_point(), the start >= end check before casting to
struct reparse_data_buffer * only ensures the start pointer is within the
response. It fails to verify that there is enough space remaining for the
fixed 8-byte header of the structure.

If a server provides a DataOffset that leaves less than 8 bytes remaining,
the check passes, but subsequent reads of ReparseTag and ReparseDataLength
will occur out-of-bounds.

Fix this by ensuring the remaining space is at least the size of the
reparse_data_buffer structure before accessing its fields.

Fixes: 56e84c64fc ("cifs: Fix validation of SMB1 query reparse point response")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:04:26 -03:00
Frank Sorenson
4775c3b7a5 smb: client: fix potential OOB read in smb3_enum_snapshots()
If snapshot_array_size is smaller than GMT_TOKEN_SIZE,
smb3_enum_snapshots() sets ret_data_len to
sizeof(struct smb_snapshot_array) without verifying the actual length
of the server's reply.

Because SMB2_ioctl() places no lower bound on the server-supplied
OutputCount and allocates retbuf to exactly that length, a short reply
results in ret_data_len exceeding the size of retbuf. The subsequent
copy_to_user() then reads past the end of retbuf, leaking adjacent slab
memory to userspace.  The subsequent clamp check is ineffective as it
only reduces ret_data_len.

Fix this by rejecting replies shorter than
sizeof(struct smb_snapshot_array) with -EIO. Note that the bound is set
to the 12-byte struct size rather than the 16-byte
MIN_SNAPSHOT_ARRAY_SIZE defined in MS-SMB2 3.3.5.15.1, because 12 bytes
is exactly what copy_to_user() attempts to read.

Fixes: e02789a53d ("smb3: enumerating snapshots was leaving part of the data off end")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:04:20 -03:00
Frank Sorenson
b09d092eb2 smb: client: fix missing iov bounds check in parse_posix_sids()
In parse_posix_sids(), sidsbuf_end is calculated using the server-supplied
out_len without being validated against the actual length of the received
iov (iov_len).

If a server provides an inflated out_len, sidsbuf_end will point past the
end of the iov. This defeats the bounds guards in posix_info_sid_size(),
allowing out-of-bounds reads into adjacent kernel memory.

Fix this by rejecting responses where the calculated sidsbuf_end would
exceed the received iov boundaries or cause pointer wraparound.

Fixes: a90f37e3d7ac ("smb: client: parse owner/group when creating reparse points")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:04:12 -03:00
Frank Sorenson
eeb5ef6083 smb: client: fix OOB struct field reads in move_smb2_ea_to_cifs()
In move_smb2_ea_to_cifs(), the while (src_size > 0) loop condition is
insufficient. It allows iteration to continue even if the remaining
src_size is too small to contain a complete smb2_ea_info structure.
Consequently, reads of ea_name_length and ea_value_length can occur
out-of-bounds.

Fix this by ensuring src_size >= sizeof(*src) before attempting to read
any structure fields. Additionally, reject any next_entry_offset that is
smaller than sizeof(*src) or that would advance the pointer beyond the
available buffer.

Note that for calls where the server returns a malformed EA list, the
error returned to userspace changes from -ENODATA (getxattr) or
-ERANGE (listxattr) to -EIO. This correctly signals a server protocol
error rather than misleadingly indicating "attribute not present" or
"output buffer too small".

Fixes: 95907fea4f ("cifs: Add support for reading attributes on SMB2+")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:04:07 -03:00
Frank Sorenson
1b3221bb12 smb: client: reject short Next offsets in parse_server_interfaces()
In parse_server_interfaces(), the server-supplied Next offset is
validated against bytes_left, but not against the size of the interface
structure itself.

A small, non-zero Next value can pass the bounds check but advance the
pointer by less than sizeof(*p). This causes the next iteration of the
loop to read misaligned, overlapping structure fields.

Fix this by ensuring the Next offset is at least sizeof(*p).

Fixes: 7d34ec36ab ("smb3: fix for slab out of bounds on mount to ksmbd")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:04:03 -03:00
Frank Sorenson
e83330c55e smb: client: fix missing lower-bound check on DFS referral string offsets
parse_dfs_referrals() checks that DfsPathOffset and NetworkAddressOffset
do not exceed the buffer end, but fails to check that they don't point
inside the referral header itself.

If a server provides an offset smaller than
sizeof(struct dfs_referral_level_3), the derived string pointer overlaps
with the struct fields, causing cifs_strndup_from_utf16() to interpret
header data as UTF-16 strings.

Fix this by enforcing that string offsets are at least sizeof(*ref).

Fixes: 4ecce920e1 ("CIFS: move DFS response parsing out of SMB1 code")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:03:59 -03:00
Frank Sorenson
f73726b83e smb: client: fix server->total_read for compound encrypted PDUs
In receive_encrypted_standard(), server->total_read is left at the
full decrypted frame size when walking sub-PDUs of a compound encrypted
frame. As a result, cifs_handle_standard() passes this full size
to smb2_check_message(), causing the PDU length guards to incorrectly
validate the entire compound frame instead of the current sub-PDU.

This allows truncated non-last sub-PDUs to bypass length validation,
leading to out-of-bounds reads in smb2_get_data_area_len().

Fix this by setting server->total_read to the true length of the
current sub-PDU: next_cmd for non-last sub-PDUs, and the remaining
pdu_length for the last one.

Fixes: b24df3e30c ("cifs: update receive_encrypted_standard to handle compounded responses")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:03:53 -03:00
Frank Sorenson
b4694f269e smb: client: validate minimum PDU size before smb2_get_data_area_len()
__smb2_calc_size() calls smb2_get_data_area_len(), which reads
command-specific struct fields to locate the data area. However,
smb2_check_message() only validates StructureSize2, meaning a truncated
response could cause smb2_get_data_area_len() to read out-of-bounds.

Replace has_smb2_data_area[] with smb2_min_pdu_len[], which is now
used to indicate both whether a command's response has a data area
and the size of that fixed response struct.  A non-zero entry means
the command has a data area, and is the minimum length required
before the struct is read.

For each command with a data area, PDUs shorter than this minimum size
are rejected instead of parsed.

The minimum is not applied to SMB2 error responses, which carry only
the 9-byte error body, the same exemption the StructureSize2 check
above it already makes.  STATUS_MORE_PROCESSING_REQUIRED is
treated as a normal reply, since an in-progress SESSION_SETUP
response carries a full body and a security blob.

Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:03:49 -03:00
Frank Sorenson
05762c5bc1 smb: client: fix next_buffer UAF and NextCommand bounds in compound PDUs
Fix several related bounds checking and pointer lifecycle issues in
receive_encrypted_standard()'s handling of compound encrypted frames:

- Clear next_buffer after assigning it to server->bigbuf. A stale
  next_buffer pointer can lead to a use-after-free on subsequent
  error paths.
- Update pdu_length to the decrypted plaintext size (buf_size). Using
  the pre-decryption length allows NextCommand to point into stale
  ciphertext residue.
- Reject next_cmd values smaller than MID_HEADER_SIZE(server).
- Fix an integer overflow in the upper bound check by verifying
  pdu_length - next_cmd < MID_HEADER_SIZE(server), ensuring the
  trailing slice is large enough for a header.

Fixes: b24df3e30c ("cifs: update receive_encrypted_standard to handle compounded responses")
Cc: stable@vger.kernel.org
Signed-off-by: Frank Sorenson <sorenson@redhat.com>
Reviewed-by: David Howells <dhowells@redhat.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-17 15:03:27 -03:00
Anand Jain
72de4807ba btrfs: derive f_fsid with dev_t only when temp_fsid is active
Commit c2a74ed049 ("btrfs: derive f_fsid from on-disk fsid and dev_t")
mixed dev_t into f_fsid for all single-device setups to avoid f_fsid
collisions with cloned filesystems.

However, doing this unconditionally breaks backward compatibility.
statfs(2) f_fsid changes after a kernel upgrade, and also can shift
across reboots or dev re-attaches as dev_t values change.

Fix this by only mixing dev_t when temp_fsid is active.  This means for
non-temp_fsid setups or the original mount, we use the old method of
deriving fsid based on the UUID.

So in the case of a cloned Btrfs filesystem, we won't be able to
maintain the same fsid across mount recycle if the mount order changes.

Reported-by: Dave Hansen <dave.hansen@intel.com>
Link: https://lore.kernel.org/linux-btrfs/be0c08f5-2f31-40f5-8a3b-f2f58b3e00ff@intel.com
Fixes: c2a74ed049 ("btrfs: derive f_fsid from on-disk fsid and dev_t")
CC: stable@vger.kernel.org # 7.2
Signed-off-by: Anand Jain <asj@kernel.org>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-09-17 19:45:12 +02:00
Qu Wenruo
b797b52e88 btrfs: add "/dev/root" exception for device path update
[BEHAVIOR CHANGE]
Since commit 108cc87339 ("btrfs: fix a lockdep caused by path
resolution during device scan"), users with btrfs rootfs but without an
initramfs are complaining that grub2 can no longer detect the rootfs
device:

  /usr/sbin/grub-probe: error: cannot find a device for / (is /dev mounted?).

[CAUSE]
Although using btrfs without an initramfs is not recommended (if a new
device is added to the rootfs, the system can no longer boot, as there
is no way to register all devices), there is still a minority of users
doing this.

If there is no initramfs but the rootfs is on a block-device-based
filesystem, the kernel boot sequence initializes a minimal ramfs/tmpfs,
creates "/dev/root" with the proper device number for the rootfs, and
then invokes mount using "/dev/root".

That's why the end user will get the mount output:

  /dev/root on / rw

To be honest, this is a user space problem: no one should trust the
device path shown in mount, only the device number.

E.g. one can even use "/proc/self/fd/*" to mount an fs, and that proc
path will be registered, and no one else can mount that fs using that
path.

Before commit 108cc87339 ("btrfs: fix a lockdep caused by path
resolution during device scan"), btrfs had an internal path lookup
workaround to address such weird paths, it works by checking if the
existing device path can still resolve to the device number.

But that path resolution is deadlock prone, thus it's replaced by a
simple devt check.

This works fine in most cases, as a btrfs device is registered by udev at
boot time, thus all paths are sane.

However this will not work for systems without an initramfs, causing the
unreachable "/dev/root" path to exist forever without a way to rename it.

[WORKAROUND]
Add an exception to the device path rename requirement.

If the device has the name "/dev/root", we know it's booted without an
initramfs, and only for that case we allow device path update.

And if someone intentionally created "/dev/root" after boot, the
existing devt checks will reject that weird name as usual.

This should satisfy the minority of users, and still keep most of the
existing guards preventing unexpected/unnecessary device path updates.

Fixes: 108cc87339 ("btrfs: fix a lockdep caused by path resolution during device scan")
Link: https://lore.kernel.org/linux-btrfs/CAKLYgeL7nrA4nXcewdv9Fqg_s=3GS=vmoypnEiZBKQ7rySZFuQ@mail.gmail.com/
Link: https://lore.kernel.org/linux-btrfs/dfbe1e27-dab8-4d55-8cf3-0b28eeac5df4@gmail.com/
Signed-off-by: Qu Wenruo <wqu@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-09-17 19:45:08 +02:00
Filipe Manana
aeab4c6287 btrfs: check if there is space for chunk item when validating sys chunk array
We checked if have enough remaining space for a key before dereferencing a
key, but we then dereference a chunk item, to get the number of stripes,
without checking if there is space for the item. So add a check to see if
there is enough space for a chunk item before dereferencing the item to
extract the stripe count.

Fixes: 2a9bb78cfd ("btrfs: validate system chunk array at btrfs_validate_super()")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-09-17 19:44:58 +02:00
Filipe Manana
97fcd34aa9 btrfs: abort transaction on failure to update inode for hole punching and reflinking
If we fail to update the inode we error out without aborting the
transaction, which can result in a persistent inconsistency if after
the failure the transaction is committed, as we have dropped file
extent items from a range and either punched a hole or insert a new file
extent item for that range (for reflinks).

So add the missing transaction abort.

Fixes: 2aaa665581 ("Btrfs: add hole punching")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-09-17 19:39:41 +02:00
Guanghui Yang
3565893cc7 btrfs: clear free space tree creation state on rebuild failure
btrfs_rebuild_free_space_tree() sets BTRFS_FS_CREATING_FREE_SPACE_TREE
before rebuilding the free space tree.  Several error paths return
without clearing this flag.

The transaction restart failure path can leave the flag set on a live
filesystem, causing delayed reference processing to be skipped. Clear it
on all free space tree rebuild failure paths. Keep
BTRFS_FS_FREE_SPACE_TREE_UNTRUSTED set, since a failed rebuild leaves
the free space tree untrusted. Callers must fall back to extent-tree
caching.

Fixes: 882af9f13e ("btrfs: handle free space tree rebuild in multiple transactions")
CC: stable@vger.kernel.org # 6.14+
Assisted-by: LLM
Reviewed-by: Boris Burkov <boris@bur.io>
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Guanghui Yang <3497809730@qq.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-09-17 19:32:11 +02:00
Darrick J. Wong
fe2f9135df xfs: fix wild memcpy access when formatting ondisk rtrefcount btree roots
LOLLM noticed that the inode btree root formatting methods copy too many
bytes -- there's only one set of keys in node blocks, not two.  This
causes memory corruption of whatever's beyond the buffers.

Cc: stable@vger.kernel.org # v6.14
Fixes: f0415af60f ("xfs: wire up a new metafile type for the realtime refcount")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
476582d754 xfs: don't let hidden_space go negative in xfs_metafile_resv_init
LOLLM points out that if the amount of fdblocks that we can reserve for
a metadata btree file goes below the space already used by that file,
then the hidden_space subtraction can underflow, causing
xfs_dec_fdblocks to subtract a huge amount of space.  We never want the
target to be less than the used sapce, so fix the logic that adjusts
dblocks_avail downwards.

Also fix an error in the adjacent comment.

Cc: stable@vger.kernel.org # v6.15
Fixes: 1df8d75030 ("xfs: make metabtree reservations global")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
ffb48dccce xfs: drop dquot flush lock when we can't find a buffer to flush
LOLLM noticed that xfs_qm_flush_one fails to drop the dquot flush lock
if it can't grab the buffer associated with the dquot.  Since there's no
buffer, nobody else is going to drop the dqflock, so we need to do it
ourselves.

Cc: stable@vger.kernel.org # v6.13
Fixes: ca378189fd ("xfs: convert quotacheck to attach dquot buffers")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
65f39d09d7 xfs: fix cursor and pointer handling when recovering iunlink buckets
LOLLM pointed out a bug in xlog_recover_iunlink_bucket:

1. We don't null out prev_ip after releasing it, which can lead to UAF
   problems if the inodegc flush call in the loop fails.

at which point I noticed even more bugs:

2. If the inodegc flush inside the loop fails, we also leak @ip.

3. We set prev_agino to agino having already advanced agino, which
   results in inodes with i_prev_unlinked set to itself.

4. If we exit the bottom of the loop with prev_ip set, then prev_ip
   aliases ip and we also set its i_prev_unlinked to itself.

Bugs 3 and 4 introduce loops into the unlinked list, though these loops
don't surface because we immediately flush each unlinked inode after
loading it.

Fix all of these issues.

Cc: stable@vger.kernel.org # v6.0
Fixes: 04755d2e58 ("xfs: refactor xlog_recover_process_iunlinks()")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
f8f6382ff1 xfs: fix blockgc group quota scanning when usrquota isn't enforced
LOLLM noticed the copy-paste error here -- if user quotas aren't
enforced but we're near the group quota limit, we fail to set FLAG_GID
and hence we might not actually free any preallocations, causing
unnecessary EDQUOT.  Fix that.

Cc: stable@vger.kernel.org # v5.12
Fixes: c237dd7c70 ("xfs: flush eof/cowblocks if we can't reserve quota for inode creation")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
d80993655f xfs: don't merge different file IO error types
LOLLM noticed that we can accidentally merge file range health
monitoring events even if they have different errors.  We shouldn't do
that.

Cc: stable@vger.kernel.org # v7.0
Fixes: dfa8bad3a8 ("xfs: convey file I/O errors to the health monitor")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
ab1c416d23 xfs: don't let memory failures leak blocks and kill repairs
LOLLM complains that a memory allocation failure in
xrep_newbt_add_blocks results in online repair leaking blocks that were
previously allocated to write a new btree, but the problem is worse than
that -- a limitation of the codebase is that the callers cannot undo the
transaction /and/ return the error -- either you undo all changes and
commit the transaction, or you error out and the filesystem goes down.

However, the new btree space reservation object isn't that big (~48
bytes).  Let's just do a NOFAIL allocation and the problem goes away.

Cc: stable@vger.kernel.org # v6.8
Fixes: be40841763 ("xfs: implement block reservation accounting for btrees we're staging")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
d7b92cbe56 xfs: don't cross reference rmapbt with bitmaps if they're incomplete
LOLLM points out that runtime errors (e.g. ENOMEM) when we're trying to
compute space usag bitmaps are silently dropped by the rmapbt scrubber.
We ought to flag that as an incomplete scrub instead of reporting
cross-referencing errors based on faulty data.

Cc: stable@vger.kernel.org # v6.4
Fixes: fed050f345 ("xfs: cross-reference rmap records with ag btrees")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
41c4c41cf6 xfs: fix rtgroup repair estimations
When I added online fsck for realtime reflink, I forgot to update
xrep_calc_rtgroup_resblks to factor in the size of the refcount btree
when it guesses how much space we need to start a repair.  This hasn't
been a huge problem in practice because there are few filesystems with
(a) realtime, (b) rtgroups, (c) reflink, and (d) no rmap.  But let's fix
this before someone stumbles upon it, especially since LOLLM flagged
this for me.

Cc: stable@vger.kernel.org # v6.14
Fixes: 83ccffc489 ("xfs: online repair of the realtime refcount btree")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Darrick J. Wong
065f3ce593 xfs: call xfs_dquot_set_prealloc_limits if we installed default rtb limits
Now that we have quotas for the realtime volume, we also have
precomputed watermark limits for the realtime block counts.  These
precomputations should be done any time we change the rtb limits, which
means that xfs_qm_adjust_dqlimits needs to ensure that if we installed
a default rtb limit.

Cc: stable@vger.kernel.org # v6.13
Fixes: 5dd70852b0 ("xfs: create quota preallocation watermarks for realtime quota")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-17 11:27:41 +02:00
Joseph Qi
d034e836ee smb: client: fix use-after-free of iface in cifs_try_adding_channels()
cifs_try_adding_channels() iterates ses->iface_list with
list_for_each_entry_safe_from(), which captures the next entry
(niface) under iface_lock.  The loop body then drops iface_lock for
the whole duration of cifs_ses_add_channel().

A concurrent interface refresh (SMB3_request_interfaces() ->
parse_server_interfaces()) marks all ifaces inactive and removes and
frees any that are not re-advertised via list_del() + kref_put(),
where release_iface() is a bare kfree().  Since niface typically has
no channel holding a reference, the list reference is its last and it
can be freed inside the unlocked window.  On continue, the iterator
advance step then dereferences niface->iface_head.next, and the loop
body reads iface->rdma_capable/is_active, both on freed memory.

Fix this by never keeping an unreferenced list pointer across the
unlocked window.  Each channel attempt now re-scans the list from the
head under iface_lock, takes a kref on the selected candidate, and
passes only that referenced candidate to cifs_ses_add_channel().
weight_fulfilled still tracks selection progress, so restarting the
scan preserves the original weighted distribution and the
weight_fulfilled-before-kref_put ordering on the failure path.

Add a per-pass attempts cap so a flapping interface refresh cannot
keep the inner loop spinning within a single tries increment.

Fixes: aa45dadd34 ("cifs: change iface_list from array to sorted linked list")
Cc: stable@vger.kernel.org
Assisted-by: Qoder:Qwen3.8-Max
Signed-off-by: Joseph Qi <joseph.qi@linux.alibaba.com>
Acked-by: Shyam Prasad N <sprasad@microsoft.com>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
2026-09-16 15:39:06 -03:00
Hyunwoo Kim
acb03d3881 exec: Cleanup POSIX timers right after de_thread()
A per-thread CPU timer holds a reference to the PID of the thread it is
attached to and, while it is armed, its node is queued in that thread's
posix_cputimers. The task is looked up by that PID.

When a non-leader thread exec()s, de_thread() changes which task owns
that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL,
but the node is still queued on tsk, which is alive. timer_lock_sighand()
takes a failed lookup to mean that the node is already dequeued, so it
has nothing to undo.

begin_new_exec() calls posix_cpu_timers_exit(me) right after
exec_task_namespaces() and that removes the leftover node, so the state
normally stays invisible. But bprm->point_of_no_return is set before
de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or
exec_task_namespaces() fails, the task dies before it gets there.
exit_itimers() then frees the k_itimer while its node is still queued,
and reaping tsk later erases that freed node from the rbtree.

In short:

      the non-leader thread B           the parent

  timer_create(CLOCK_THREAD_CPUTIME_ID)
  timer_settime()
    arm_timer()            // the node is queued on B
  execve()
    de_thread(B)
      exchange_tids(B, leader)  // B's PID now belongs to the leader
      release_task(leader)
        __exit_signal(leader)
          posix_cpu_timers_exit(leader)  // cleans leader's queue, not B's
          __unhash_process(leader)  // that PID has no task anymore
    exec_mmap()
      mmap_read_lock_killable(old_mm)
                                kill(B, SIGKILL)
      // -EINTR
  get_signal()
    do_exit()
      exit_itimers()
        posix_timer_delete()
          posix_cpu_timer_del()
        posix_timer_unhash_and_free()  // freed while still queued
                                wait4()
                                  release_task(B)
                                    posix_cpu_timers_exit(B)
                                      cleanup_timerqueue()
                                        timerqueue_del()  // use-after-free

Move the POSIX timer cleanup right after de_thread() before any of the
later failure conditions brings the task into do_exit().

[ tglx: Move the cleanup right after de_thread() ]

Fixes: 55e8c8eb2c ("posix-cpu-timers: Store a reference to a pid not a task")
Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Tested-by: Kijo Park <red993688@gmail.com>
Reviewed-by: Oleg Nesterov <oleg@redhat.com>
Reviewed-by: Frederic Weisbecker <frederic@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/ao7Q8miiuLAPVnWv@v4bel
Link: https://patch.msgid.link/20260911090541.627712075@kernel.org
2026-09-16 18:44:03 +02:00
Daniel Linjama
76bf149cd0 btrfs: handle lack of space when cleaning up verity items
When enable_verity() hits the qgroup limit, rollback_verity() needs its
own metadata reservation. When the qgroup limit or lack of space refuses
the rollback, the whole filesystem is forced read-only even though the
qgroup limit was for one subvolume only. Also orphan cleanup at the next
mount fails the same way, so the leftover items are never removed: with
-EDQUOT the subvolume stays unreachable, and with -ENOSPC on a full
filesystem the next read-write mount fails.

Start transactions with btrfs_start_transaction_fallback_global_rsv() in
btrfs_orphan_cleanup(), drop_verity_items() and rollback_verity(). Those
calls only delete items and free the space in the end, so they may use
the global reserve and skip the qgroup limit, which avoids -ENOSPC and
-EDQUOT.

Fixes: 146054090b ("btrfs: initial fsverity support")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Daniel Linjama <daniel@dev.linjama.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-09-16 15:56:10 +02:00
Filipe Manana
cc337324cc btrfs: fix creation of compressed inline extents that don't save space
If the compressed data of an inline extent is larger than or equals to the
size of the uncompressed data, we are still allowing the creation of the
compressed inline extent, which does not result in any benefits, quite the
contrary as we waste metadata space and have to decompress when reading.

This is a recent regression introduced in commit 3eaf5f082c ("btrfs:
extract inlined creation into a dedicated delalloc helper").

It happens because we are passing the block size to btrfs_compress_bio(),
so we don't get -E2BIG from the compression code anymore, but we can not
pass i_size either, because if i_size is smaller than sector size, we
end up never creating lzo compressed inline extent for such small i_size
values. So refuse the compressed result at run_delalloc_inline() if
its size is not smaller than the uncompressed size (i_size).

Reported-by: Hanabishi <i.r.e.c.c.a.k.u.n+kernel.org@gmail.com>
Link: https://lore.kernel.org/linux-btrfs/c97652a5-ac6b-4de6-aa23-3cdebc01d00b@gmail.com/
Fixes: 3eaf5f082c ("btrfs: extract inlined creation into a dedicated delalloc helper")
CC: stable@vger.kernel.org # 7.1+
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-09-16 15:56:02 +02:00
Namjae Jeon
9fa26285ae ksmbd: keep compound responses on query info errors
Do not reset the RFC1002 length of the complete response when a query
info buffer is too small. The current command will add its error response
through ksmbd_iov_pin_rsp(), while resetting the base length can truncate
earlier responses in a compound request.

This lets ksmbd return the earlier responses and the query-info error
response together. Remove the now-unused rsp_org parameter from the pipe
query-info helpers.

Fixes: e2b76ab8b5 ("ksmbd: add support for read compound")
Reported-by: Mobin Aydinfar <mobin@mobintestserver.ir>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
2026-09-15 22:28:52 +09:00
Namjae Jeon
f4fafaf021 ksmbd: fix partial normalized name responses
Windows may request FILE_NORMALIZED_NAME_INFORMATION with an output
buffer that only fits the fixed portion of the variable-length response.
Treat the fixed portion as FILE_NORMALIZED_NAME_INFORMATION_SIZE so ksmbd
returns STATUS_BUFFER_OVERFLOW instead of STATUS_INFO_LENGTH_MISMATCH.

This avoids rejecting valid partial normalized-name responses.

Fixes: 6b8b79226b ("ksmbd: fix partial file information responses")
Reported-by: Mobin Aydinfar <mobin@mobintestserver.ir>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Namjae Jeon <linkinjeon@kernel.org>
2026-09-15 22:28:51 +09:00
Hemanth Selam
2f3c2a6f96 xfs: fix typos and repeated words in comments
Correct ten misspellings and eight accidentally doubled words in
comments.  No code changes.

The doubled word in xfs_zone_alloc.c was not a duplicate: "so that is is
reused" is "it is" misspelt, so that one reads "so that it is reused"
rather than dropping a word.

Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:47:28 +02:00
Anuj Gupta
c16b885ad6 xfs: remove unused xfs_reflink_remap_range declaration
The implementation was inlined into xfs_file_remap_range() by commit
3fc9f5e409 ("xfs: remove xfs_reflink_remap_range"), leaving this
declaration orphaned.

Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:20:05 +02:00
Jiangshan Yi
ce2b91bebc xfs: remove duplicate INO1_WRITTEN check
Commit a23eca8844 ("xfs: fix exchange-range reflink flag clearing
issue with INO1_WRITTEN") duplicated commit b2d5a81dae ("xfs: fix
exchange-range reflink flag clearing issue with INO1_WRITTEN"), so
xmi_can_exchange_reflink_flags() ended up with two identical
XFS_EXCHMAPS_INO1_WRITTEN checks.  The second one is dead code,
since the first one already returned false.  Remove it.

Signed-off-by: Jiangshan Yi <yijiangshan@kylinos.cn>
Reviewed-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:19:14 +02:00
Christoph Hellwig
14e379600d xfs: don't try to get a reference to a NULL oz in xfs_get_cached_zone
oz can be NULL when we resample it after taking i_flags_lock, so account
for that.

Fixes: 2d829cc767 ("xfs: fix racy open zone caching")
Signed-off-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Hans Holmberg <hans.holmberg@wdc.com>
Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:12:38 +02:00
Darrick J. Wong
e9193f2f1c xfs: check di_forkoff correctly in scrub
The di_forkoff check in xchk_dinode is incorrect, according to LOLLM.
XFS_DFORK_BOFF returns a byte count relative to the start of the literal
area, not the start of the inode.  Therefore, this check won't flag
di_forkoff values that are larger than the literal area but not the
inode size itself.  Fix this check; sadly the old APTR code was correct.

Cc: stable@vger.kernel.org # v6.8
Fixes: 6b5d917780 ("xfs: dont cast to char * for XFS_DFORK_*PTR macros")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:11:11 +02:00
Darrick J. Wong
c54110d814 xfs: only flag zero padding for dir3 data blocks, not dir3 block blocks
LOLLM complains that xchk_directory_data_bestfree can be passed a
directory block that is either in "block" or "data" format, but the
check here unconditionally treats the dir3_block and dir3_data blocks as
if they have the same header format (they don't).  Consequently, we can
incorrectly set the preen state on dir3_block blocks, which of course
we can't preen away because dir3_block blocks do not have a padding
field.  Fix this.

Cc: stable@vger.kernel.org # v7.1-rc4
Fixes: 939919ccdd ("xfs: check directory data block header padding in scrub")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:11:11 +02:00
Darrick J. Wong
46c1b6674a xfs: fix integer overflows in xbitmap set functions
LOLLM complains that the xbitmap set functions can suffer an integer
underflow or overflow and thereby return the wrong left and right
pointers.  Fix that logic bomb, even though (AFAICT) we never actually
try to set the *entire* bitmap.

Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:11:11 +02:00
Darrick J. Wong
471e0b6e2d xfs: use the correct reservations for rtrmap/refcount recovery
LOLLM noticed that we might reserve the wrong number of blocks for
recovering rtrmap and rtrefcount updates after a crash.  Fix that.

Cc: stable@vger.kernel.org # v6.14
Fixes: 5e0679d1c6 ("xfs: support recovering rmap intent items targetting realtime extents")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:11:11 +02:00
Darrick J. Wong
8fc18580ec xfs: don't call xfs_exchange_range_finish for a dry run
LOLLM noticed that we strip file privileges and whatnot even for a dry
run.  We also shouldn't flush dirty data to disk or trim COW staging
events for a dry run.  Neither of those behaviors are allowed by the
manpage, so fix that by exiting early on DRY_RUN in various functions.

Cc: stable@vger.kernel.org # v6.10
Fixes: 42672471f9 ("xfs: bind together the front and back ends of the file range exchange code")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:11:11 +02:00
Darrick J. Wong
3083ba8dde xfs: check padding field in xfs_ioc_commit_range
LOLLM points out that we don't check the ioctl padding field here, so
let's do that.  I don't think there are many users yet since exchrange
requires a new feature flag, so it's a good time to try to plug this
hole.

Cc: stable@vger.kernel.org # v6.12
Fixes: 398597c3ef ("xfs: introduce new file range commit ioctls")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:11:11 +02:00
Darrick J. Wong
984aab2d90 xfs: use correct jiffies comparison function in xchk_maybe_relax
LOLLM points out that we're supposed to use time_after_eq, not a raw >=
operation here, or else jiffies wraps can go unnoticed.  Fix this.

Cc: stable@vger.kernel.org # v6.10
Fixes: 271557de7c ("xfs: reduce the rate of cond_resched calls inside scrub")
Signed-off-by: Darrick J. Wong <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-09-15 10:09:57 +02:00