Commit Graph

106934 Commits

Author SHA1 Message Date
Linus Torvalds
51f247c4b2 for-7.2-rc4-tag
-----BEGIN PGP SIGNATURE-----
 
 iQJPBAABCgA5FiEE8rQSAMVO+zA4DBdWxWXV+ddtWDsFAmpdukAbFIAAAAAABAAO
 bWFudTIsMi41KzEuMTIsMiwyAAoJEMVl1fnXbVg70hwQAJ78AUVnvTSwMi2w6TEx
 KFYZSMPdBi9TwGC3lDCcTA/U5PqGTkKACWBLf5z+9YXAcb3By7OiSAQeko7sib0X
 k0jz5SmJjMHrHe1qGQvaztLin9/Ow9K8/QUQW0qhunwF6FbmgkfWdXw4au5+zV5R
 rwdyDt74nlhf3pCOpJb1JL2BFy23PXbfQIty9I+2qKUMPSgqhC6l7CBYsWZoG5ub
 QC2+Go5Ygz61iKDpAZ3TKjULI58+7K3FBq+76Jr9UajPVD75637Wze3zsrgcb2LV
 THxITsWTxx5CwiQF1Sf0JGo/rU+V0wU5oJqwmn4FrGeZXhnGXXcnXmG70VzvAJYX
 4so+p0aejefqQnwcFbQE0+dFjrEuD+Dl8PVzlDSMygz/y1xpSlxVz+83kQFH+MV6
 1bxrOJNS6sLwvqtwRa8INjqvMfk443Ub5sFUS+bav0ZUlOg3BmitkEP0H0NlQ+5p
 CAMCXFSfb/1UDXMrMHAcZn3LZtDLNaVVkneCz97A72D5C6Ae4rL3/FeHvmDk1KUh
 StO3ReF4JXeN3FYTzWpMIw07jGxR9ZuFVn78DAW9Gwjk658wsXBqcGPgg/GAsy6R
 bvrFRkWiJGSSdUeipNwxjfTeM8NwkA7fv+DnLbp6HzpOLDKIC8pID1Eoh1nwTlas
 R073SeoNHO0sC6Lz+2NcvVMb
 =1jNx
 -----END PGP SIGNATURE-----

Merge tag 'for-7.2-rc4-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux

Pull btrfs fixes from David Sterba:
 "I'm catching up with the fix backlog in the development branch, so
  here's a number of them and will probably send one more for this or
  the next rc:

   - relocation fixes:
     - skip attempting compression on reloc inodes
     - exclude inline extents from file extent offset checks
     - fix minor memory leak after error when adding reloc root
     - fix root cleanup after inserting and merging

   - fix clearing folio tags after writeback

   - clear logging flag of extent map before splitting

   - fix unsigned 32/64 type conversions when accounting dirty metadata,
     leading to continually exceeding threshold

   - fix regression in 32bit compat ioctl for subvolume info

   - fix type of SEARCH_TREE ioctl buffer in UAPI header

   - fix expression in ASSERT expression which can be unconditionally
     evaluated on some compilers

   - only account delalloc bytes for regular inodes"

* tag 'for-7.2-rc4-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux:
  btrfs: fix GET_SUBVOL_INFO after compat refactor
  btrfs: free mapping node on duplicate reloc root insert
  btrfs: fix a regression where PAGECACHE_TAG_DIRTY is never cleared
  btrfs: don't propagate EXTENT_FLAG_LOGGING to split extent maps
  btrfs: fix u32 to s64 type conversion in dirty_metadata_bytes accounting
  btrfs: fix NULL pointer deref during assertion in btrfs_backref_free_node()
  btrfs: only account delalloc bytes for regular file inodes in btrfs_getattr()
  btrfs: reject inline file extents item in get_new_location()
  btrfs: do not try compression for data reloc inodes
  btrfs: declare btrfs_ioctl_search_args_v2::buf as __u8
  btrfs: fix reloc root cleanup in merge_reloc_roots()
  btrfs: fix use-after-free on reloc root after error in insert_dirty_subvol()
2026-07-21 08:06:24 -07:00
Linus Torvalds
b95f03f04d 12 hotfixes. 8 are cc:stable and the remainder address post-7.1 issues or
aren't considered appropriate for backporting.  10 are for MM.
 
 All are singletons - please see the relevant changelogs for details.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTTMBEPP41GrTpTJgfdBJ7gKXxAjgUCal5rBQAKCRDdBJ7gKXxA
 jtgtAQCuWNUTCR7u+MzAuO3Nh46DxHXeb27OTHZL8JcazQTEQgD+NZfwqVYnNNX/
 4CVqqZvrXJQDg0aiWtIP4VdLirNh/Ac=
 =Gslp
 -----END PGP SIGNATURE-----

Merge tag 'mm-hotfixes-stable-2026-07-20-11-37' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm

Pull misc fixes from Andrew Morton:
 "12 hotfixes. 8 are cc:stable and the remainder address post-7.1 issues
  or aren't considered appropriate for backporting. 10 are for MM.

  All are singletons - please see the relevant changelogs for details"

* tag 'mm-hotfixes-stable-2026-07-20-11-37' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm:
  mm/memory-failure: trace: change memory_failure_event to ras subsystem
  mm: page_reporting: allow driver to set batch capacity
  mm/kmemleak: fix checksum computation for per-cpu objects
  mm/damon/core: disallow overlapping input ranges for damon_set_regions()
  MAINTAINERS: add Usama as a THP reviewer
  fat: avoid stack overflow warning
  mm/damon/core: validate ranges in damon_set_regions()
  m68k: avoid -Wunused-but-set-parameter in clear_user_page()
  mm/huge_memory: set PG_has_hwpoisoned only after new folio head is established
  mm/page_vma_mapped: fix device-private PMD handling
  MAINTAINERS: s/SeongJae/SJ/
  userfaultfd: prevent registration of special VMAs
2026-07-20 13:04:47 -07:00
Linus Torvalds
1229e2e57a eight ksmbd server fixes
-----BEGIN PGP SIGNATURE-----
 
 iQGzBAABCgAdFiEE6fsu8pdIjtWE/DpLiiy9cAdyT1EFAmpa3JoACgkQiiy9cAdy
 T1EP7wv9FiM0oUOekxfydVFTMdpzkMpvBHcLoVSpnFjvgJZVh4gOmvyY0fRuXKRB
 o40SAoBvZP5KRe6qjBxo2o+z4T6NpRY+K1ftv0xxFvgDZ3+wvyV6jEf80cLnoDAl
 xqOji4PeeZfrKVwPYAFTVWhImQZ9IZuP3cBpZKqf5EWYovk43Ex/VQCqnVdAMY33
 fvdUEY1u8HVVILOzHRWvo7QAlyw6SEBjdmTLqCB2mNDwBcfg1C/zuvbf8KYS4LXu
 +C8F137JlSz6tplrTBGIq3FVwLxhGunZwanO3IVM8B0+FhSsedWT/qS5CbgijZxn
 GpT9CBKWN/KU3YPHEFDKxcSkJAc0MA20cb9LRUdIy7uv+OnYVV2oY209dukj8RpI
 05F8RC1frrdhY9vZZhjIw/x5jbrHb4cGEohGyB6pH8C1gJc1AOuFBEZ7X7dHgGqA
 jjJ3bBD4f8mL7Pim37gubyCYvO/MtjEkKYWvoWxtjWewHxpJdvH+JefnOz8Cyp94
 BXR4hd1r
 =K5Ks
 -----END PGP SIGNATURE-----

Merge tag 'v7.2-rc3-smb3-server-fixes' of git://git.samba.org/ksmbd

Pull smb server fixes from Steve French:
 "ksmbd server fixes, mostly addressing malformed SMB request
  handling and connection/session lifetime issues, including
  two information-disclosure or memory-safety bugs in the SMB2
  request/response paths.

   - validate FILE_ALLOCATION_INFORMATION before block rounding to
     prevent a client-controlled overflow from truncating a file.

   - pin connections while asynchronous oplock and lease-break
     notifications are pending.

   - initialize compound SMB2 READ alignment padding, preventing
     disclosure of uninitialized heap bytes.

   - release the allocated alternate-stream xattr name after rename.

   - size multichannel binding session-key buffers for the largest
     permitted key, avoiding a stack buffer overflow.

   - remove a disconnecting connection's channels from every session,
     including channels whose binding state has since changed.

   - serialize binding preauthentication-session lookup and update
     against its teardown.

   - check that every compound request element contains StructureSize2
     before reading it"

* tag 'v7.2-rc3-smb3-server-fixes' of git://git.samba.org/ksmbd:
  ksmbd: validate compound request size before reading StructureSize2
  ksmbd: lock the binding preauth session in smb3_preauth_hash_rsp
  ksmbd: remove stale channels from all sessions on teardown
  ksmbd: fix stack buffer overflow in multichannel session-key copy
  ksmbd: fix memory leak of xattr_stream_name in smb2_rename()
  ksmbd: zero the smb2_read alignment tail to avoid an infoleak
  ksmbd: pin conn during async oplock break notification
  ksmbd: fix integer overflow in set_file_allocation_info()
2026-07-17 21:41:54 -07:00
Linus Torvalds
fce2dfa773 ten client fixes
-----BEGIN PGP SIGNATURE-----
 
 iQGzBAABCgAdFiEE6fsu8pdIjtWE/DpLiiy9cAdyT1EFAmpZZLcACgkQiiy9cAdy
 T1HsUwv/SMAAxV/8OlEiSfhg+YBjgHio1MRG8eB8yyaJJxpbt5AfYKR5yDc57GUT
 +CdVKxc+ZBFKhaNDFF+w0Z+nzTC5KGRRiixNVperet8gqzwuL5ZE1BrSJJDbf7CY
 uPuUTlTol380U7mV/F+3elV+1oFiyNcZFIPPnTKldJdrVOf+tHKIoP9ffVbnVw4X
 eY5h1EbmV+VCDNI+crkFPl49VeeYhOGeUx+zhW2gwkMae0/CBKeYVa2ztvWOPZeq
 k8tLdL/OU4ZT85XFPPrjdJ8RY39kNqv6ZKE4k5LScWjm4Cz/kNvMrVbACoUfR6Vr
 a7kt64rCDmpfcgpB5uIMsQ7TkPPULwT4qic41uBq8LxmsDKvYJD+Fc1mJkl2b4gW
 6T0K0jieBleECc4Mj/XLwoZlfQux1F/2801aHo3QxcfY5zjjRY+FlEhrDGJ4dgXA
 wGVXwvZF4od6Yp+dTn+JxsJs8x3PrLV5uON09yetoPaptIfINlCtfPc8noFOI+aa
 3ApU8cuF
 =MgQD
 -----END PGP SIGNATURE-----

Merge tag 'v7.2-rc3-smb3-client-fixes' of git://git.samba.org/sfrench/cifs-2.6

Pull smb client fixes from Steve French:

 - fallocate fixes

 - unit test fixes

 - fix allocation size after duplicate extents

 - fix check for overlapping data areas

* tag 'v7.2-rc3-smb3-client-fixes' of git://git.samba.org/sfrench/cifs-2.6:
  smb/client: flush dirty data before punching a hole
  smb/client: Use EXPORT_SYMBOL_IF_KUNIT() to export symbols in SMB2
  smb/client: Use EXPORT_SYMBOL_IF_KUNIT() to export symbols
  smb: client: reject overlapping data areas in SMB2 responses
  smb/client: refresh allocation after EOF-extending fallocate
  smb/client: emulate small EOF-extending mode 0 fallocate ranges
  smb/client: reduce fallocate zero buffer allocation
  smb/client: handle overlapping allocated ranges in fallocate
  smb/client: refresh allocation size after duplicate extents
  smb: client: use kvzalloc() for megabyte buffer in simple fallocate
2026-07-16 16:46:26 -07:00
Linus Torvalds
e22254e9dd xfs: fixes for v7.2-rc4
Signed-off-by: Carlos Maiolino <cem@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iJUEABMJAB0WIQSmtYVZ/MfVMGUq1GNcsMJ8RxYuYwUCalj/DwAKCRBcsMJ8RxYu
 Y3UMAYCrPzc41SGZdOkoxvO5g3BwGISXYCiCCIFFvnnDXztuiJytgZ+VPH+C7X4J
 tlGjRXUBfi8PcoePIbDfPXxay+T/g5Pz29ADyl2UAkxPYRRWeYYxKTyCTBvBdUXS
 lMCnv0pSPg==
 =Hcay
 -----END PGP SIGNATURE-----

Merge tag 'xfs-fixes-7.2-rc4' of git://git.kernel.org/pub/scm/fs/xfs/xfs-linux

Pull xfs fixes from Carlos Maiolino:
 "This contains mostly a series of bug fixes found by different LLM
  models"

* tag 'xfs-fixes-7.2-rc4' of git://git.kernel.org/pub/scm/fs/xfs/xfs-linux: (21 commits)
  xfs: don't zap bmbt forks if they are MAXLEVELS tall
  xfs: clamp timestamp nanoseconds correctly
  xfs: fully check the parent handle when it points to the rootdir
  xfs: handle non-inode owners for rtrmap record checking
  xfs: fix off-by-one error when calling xchk_xref_has_rt_owner
  xfs: set xfarray killable sort correctly
  xfs: grab rtrmap btree when checking rgsuper
  xfs: write the rg superblock when fixing it
  xfs: use the rt version of the cow staging checker
  xfs: use rtrefcount btree cursor in xchk_xref_is_rt_cow_staging
  xfs: don't wrap around quota ids in dqiterate
  xfs: move cow_replace_mapping to xfs_bmap_util.c
  xfs: make cow repair somewhat flaky when debugging knob enabled
  xfs: don't replace the wrong part of the cow fork
  xfs: resample the data fork mapping after cycling ILOCK
  xfs: fix null pointer dereference in tracepoint
  xfs: use xfs_csn_t for xlog_cil_push_now() push_seq parameter
  xfs: tie zoned sysfs lifetime to zone info
  xfs: fail recovery on a committed log item with no regions
  xfs: splice unsorted log items back to the transaction after the loop
  ...
2026-07-16 09:40:03 -07:00
Linus Torvalds
133fc7aacc Changes since last update:
- Fix sanity checks for ztailpacking tail pclusters to avoid
    false corruption reports
 
  - Use more informative s_id for file-backed mounts
 
  - Hide the meaningless "cache_strategy=" mount option on plain
    (uncompressed) filesystems
 
  - Remove the unneeded erofs_is_ishare_inode() helper
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEQ0A6bDUS9Y+83NPFUXZn5Zlu5qoFAmpY47QRHHhpYW5nQGtl
 cm5lbC5vcmcACgkQUXZn5Zlu5qoPzQ//Rjcp/c9wAegDZG+Y5qwquAqIVtiO582a
 DlimqGXP/+XIOx8IqhIzqXx9R1GpIjIT6xpWrqhXr+Xcics2W48CE6G/0xg+Mh0h
 GDtf836rolHeeqJNG6gmEaWik3C91MDbZY3hwYJIQmS+gTVUBx1hwWvi22hauPZI
 qGJ9wVEKlv6yAOC5KORx5hQidMeNblAI+z9MWWV/lm5Zf4w1Ve2wp0298utBzcnR
 xVk6r+sEarqnHZ6O/ncM72S0OmyD/PrdfJkHUqF3RPjMVI4PSnlfQD+1IOGAmmy+
 m0JnBUUylWI7c3Lg3e7RjgVT7Qbytn3iBp9UeMUNXsWHy6bG3zbuTmYTuWXckL79
 YMOHdRpZ33scOts7+7D0hsp63dDErs1ALyNv9a8Mw2pO1ITxCs9D+omhzjwwL2Y/
 XchOFuLMCBqfgigwGOvKOX5IBwdLpQxCMJVZsBGfpWNXcE0YqGJ77XaatoUt5gdu
 FdUP7zwFTjTiMhXlBSPCYm1fID/e0+rqcEeztSsvIGlu0DClrXItCYTlhThSU3nr
 eu/YmMZw/6vsXI8cPSX1Elkr52Q6bp2jlL2mCfvSEKgRTmNXrxgGWHLVsEb9i1aa
 lON+F5B/lDMt9hV7RL7eINRG1Ov7ADm1wyt14oVSBSgh2dx2PzAUL2YpWsmpVsCz
 K8HBfmru4lg=
 =3hh+
 -----END PGP SIGNATURE-----

Merge tag 'erofs-for-7.2-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/xiang/erofs

Pull erofs fixes from Gao Xiang:

 - Fix sanity checks for ztailpacking tail pclusters to avoid
   false corruption reports

 - Use more informative s_id for file-backed mounts

 - Hide the meaningless "cache_strategy=" mount option on plain
   (uncompressed) filesystems

 - Remove the unneeded erofs_is_ishare_inode() helper

* tag 'erofs-for-7.2-rc4-fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/xiang/erofs:
  erofs: hide "cache_strategy=" for plain filesystems
  erofs: get rid of erofs_is_ishare_inode() helper
  erofs: relax sanity check for tail pclusters due to ztailpacking
  erofs: use more informative s_id for file-backed mounts
2026-07-16 09:30:42 -07:00
Xiang Mei (Microsoft)
15b38176fd ksmbd: validate compound request size before reading StructureSize2
When ksmbd validates a compound (chained) SMB2 request,
ksmbd_smb2_check_message() reads pdu->StructureSize2 without first
checking that the compound element is large enough to contain it.
StructureSize2 is a 2-byte field at offset 64
(__SMB2_HEADER_STRUCTURE_SIZE) from the start of each element.

The compound-walking logic only guarantees that a full 64-byte SMB2
header is present for the trailing element: when NextCommand is 0, len is
reduced to the number of bytes remaining after next_smb2_rcv_hdr_off. A
remote client can craft a compound request whose last element has exactly
64 bytes, so the 2-byte StructureSize2 read at offset 64 extends one byte
past the receive buffer, producing a slab-out-of-bounds read.

  BUG: KASAN: slab-out-of-bounds in ksmbd_smb2_check_message (fs/smb/server/smb2misc.c:402)
  Read of size 2 at addr ffff888012ae31ac by task kworker/0:1/14
  The buggy address is located 172 bytes inside of allocated 173-byte region
  Workqueue: ksmbd-io handle_ksmbd_work
  Call Trace:
   ...
   kasan_report (mm/kasan/report.c:595)
   ksmbd_smb2_check_message (fs/smb/server/smb2misc.c:402)
   handle_ksmbd_work (fs/smb/server/server.c:119)
   process_one_work (kernel/workqueue.c:3314)
   worker_thread (kernel/workqueue.c:3397)
   kthread (kernel/kthread.c:436)
   ret_from_fork (arch/x86/kernel/process.c:158)
   ret_from_fork_asm (arch/x86/entry/entry_64.S:245)

Reject any compound element that is too small to hold StructureSize2
before dereferencing it.

Fixes: e2f34481b2 ("cifsd: add server-side procedures for SMB3")
Reported-by: AutonomousCodeSecurity@microsoft.com
Signed-off-by: Xiang Mei (Microsoft) <xmei5@asu.edu>
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-16 10:18:25 -05:00
Gil Portnoy
35754558f6 ksmbd: lock the binding preauth session in smb3_preauth_hash_rsp
smb3_preauth_hash_rsp() computes the SMB3.1.1 preauth integrity hash on
the response path. For a binding SESSION_SETUP it looks up the
per-connection preauth_session and reads its Preauth_HashValue.

smb2_sess_setup() frees that preauth_session under ksmbd_conn_lock().
Two SMB2 requests on one connection can run concurrently, so an unlocked
lookup and hash can use a preauth_session after another worker frees it.

Take ksmbd_conn_lock() before selecting conn->binding and hold it across
the selected preauth hash lookup and update. This preserves the existing
hash selection while preventing the lookup-to-use lifetime race.

Fixes: 1c5daa2ea9 ("ksmbd: handle channel binding with a different user")
Signed-off-by: Gil Portnoy <dddhkts1@gmail.com>
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-16 10:18:25 -05:00
Gil Portnoy
b10665730f ksmbd: remove stale channels from all sessions on teardown
ksmbd_sessions_deregister() removes a connection's channels from other
sessions' channel lists only while conn->binding is still set:

	if (conn->binding) {
		hash_for_each_safe(sessions_table, ...)
			ksmbd_chann_del(conn, sess);
	}

conn->binding is a transient flag: it is cleared once a binding
SESSION_SETUP completes, and also by a subsequent non-binding
SESSION_SETUP on the same connection (a reauthentication on a bound
channel, or a new SessionId==0 setup). A connection that has bound a
channel into another session's ksmbd_chann_list and then clears
conn->binding leaves that channel behind when it disconnects: the
channel, whose chann->conn points at the now freed struct ksmbd_conn,
stays on the owner session's list.

When the owning connection later tears down, the second loop
dereferences the stale channel:

	xa_for_each(&sess->ksmbd_chann_list, chann_id, chann)
		if (chann->conn != conn)
			ksmbd_conn_set_exiting(chann->conn);   /* freed */

which is a use-after-free write into the freed ksmbd_conn (the same
stale channel is also walked by show_proc_session() through /proc). The
session is leaked as well, because its channel list never empties.

Remove the conn->binding gate so a connection always removes its
channels from every session on teardown.

Fixes: faf8578c77 ("ksmbd: find bound sessions during reauthentication")
Signed-off-by: Gil Portnoy <dddhkts1@gmail.com>
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-16 10:18:25 -05:00
Gil Portnoy
610346149d ksmbd: fix stack buffer overflow in multichannel session-key copy
Commit 4b706360ff ("ksmbd: fix multichannel binding and enforce channel
limit") moved the binding-path session key out of the session-wide
sess->sess_key (CIFS_KEY_SIZE = 40) into a new per-channel buffer, and
sized both that buffer and the on-stack copy used during binding with
SMB2_NTLMV2_SESSKEY_SIZE (16):

	struct channel {
		char	sess_key[SMB2_NTLMV2_SESSKEY_SIZE];	/* 16 */
		...
	};

	ntlm_authenticate() / krb5_authenticate():
		char channel_key[SMB2_NTLMV2_SESSKEY_SIZE] = {};	/* 16 */
		char *auth_key = conn->binding ? channel_key : sess->sess_key;

The two writers that fill this destination still bound the copy length
against CIFS_KEY_SIZE (40), not against the 16-byte buffer:

	ksmbd_decode_ntlmssp_auth_blob() (NTLM key exchange):
		if (sess_key_len > CIFS_KEY_SIZE)	/* 40 */
			return -EINVAL;
		arc4_crypt(ctx_arc4, sess_key,
			   (char *)authblob + sess_key_off, sess_key_len);

	ksmbd_krb5_authenticate():
		if (resp->session_key_len > sizeof(sess->sess_key))	/* 40 */
			...
		memcpy(sess_key, resp->payload, resp->session_key_len);

On a binding SESSION_SETUP, auth_key points at the 16-byte channel_key,
so a client that supplies an NTLM EncryptedRandomSessionKey of up to 40
bytes (with NTLMSSP_NEGOTIATE_KEY_EXCH), or a Kerberos ticket whose
session key is longer than 16 bytes (a normal AES256 key is 32), writes
past the 16-byte stack buffer -- up to a 24-byte kernel stack overflow.
KASAN reports it as a stack-out-of-bounds write in arc4_crypt() called
from ksmbd_decode_ntlmssp_auth_blob().

The destinations must be able to hold the full session key the length
checks already permit. Size the per-channel key buffer and the two
on-stack channel_key buffers with CIFS_KEY_SIZE, matching sess->sess_key.

Fixes: 4b706360ff ("ksmbd: fix multichannel binding and enforce channel limit")
Signed-off-by: Gil Portnoy <dddhkts1@gmail.com>
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-16 10:18:25 -05:00
Gil Portnoy
08b9452a33 ksmbd: fix memory leak of xattr_stream_name in smb2_rename()
On an SMB2 SET_INFO(FileRenameInformation) whose target names an alternate
data stream, smb2_rename() obtains a formatted stream-name string from
ksmbd_vfs_xattr_stream_name(), which allocates it with kasprintf() and
returns it through an out-param:

	rc = ksmbd_vfs_xattr_stream_name(stream_name, &xattr_stream_name, ...);
	if (rc)
		goto out;
	rc = ksmbd_vfs_setxattr(..., xattr_stream_name, ...);
	if (rc < 0) {
		...
		goto out;
	}
	goto out;

xattr_stream_name is declared inside the alternate-data-stream block, but
the out: label is outside that block and frees only new_name, so it cannot
release xattr_stream_name. ksmbd_vfs_setxattr() takes a const char * and
only reads the name, so it does not take ownership either. Both the
setxattr-failure and the success path therefore leak the kasprintf()'d
string. An authenticated client with a writable share can leak kernel
memory on every stream rename, exhausting kernel memory over time.

Free xattr_stream_name after its use, before the block's goto out. The
two earlier goto out paths never assign the variable, so there is no
double-free.

Signed-off-by: Gil Portnoy <dddhkts1@gmail.com>
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-16 10:18:25 -05:00
Gil Portnoy
b078f39af3 ksmbd: zero the smb2_read alignment tail to avoid an infoleak
Commit 6b9a2e09d4 ("ksmbd: avoid zeroing the read buffer in smb2_read()")
switched the SMB2 READ payload buffer from kvzalloc() to kvmalloc(), on the
premise that only the nbytes actually read are ever transmitted, so the
ALIGN(length, 8) tail need not be initialized.

That premise does not hold for a compound response. ksmbd_vfs_read() fills
only nbytes, leaving [nbytes, ALIGN(length, 8)) uninitialized. The aux
payload is pinned as the last response iov with iov_len == nbytes, but when
the READ is a member of a compound, init_chained_smb2_rsp() 8-byte-aligns
the previous member by extending that same iov:

	new_len = ALIGN(len, 8);
	work->iov[work->iov_idx].iov_len += (new_len - len);
	inc_rfc1001_len(work->response_buf, new_len - len);

so up to 7 uninitialized bytes of the kvmalloc()'d slab tail are sent
to the client. When the read length is small the buffer is served from
a general kmalloc slab, so those bytes can be stale kernel-heap
contents, including pointer values -- an information leak usable to
defeat KASLR.

An authenticated client triggers it with a compound request containing a
READ whose returned nbytes is not 8-aligned (for example [READ, CLOSE] with
a 1-byte read).

Zero only the alignment tail after the read, preserving the bulk
no-zeroing optimization of 6b9a2e09d4.

Fixes: 6b9a2e09d4 ("ksmbd: avoid zeroing the read buffer in smb2_read()")
Cc: stable@vger.kernel.org
Signed-off-by: Gil Portnoy <dddhkts1@gmail.com>
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-16 10:18:25 -05:00
Qihang
aa5d8f3f96 ksmbd: pin conn during async oplock break notification
smb2_oplock_break_noti() and smb2_lease_break_noti() store a ksmbd_conn
pointer in an async ksmbd_work and then queue that work on ksmbd-io.  The
work only increments conn->r_count, which prevents teardown from passing
the pending-request wait after the increment, but it does not pin the
struct ksmbd_conn object.

If connection teardown races with an oplock break notification, the last
conn reference can be dropped before the queued worker finishes.  The
worker then uses the freed conn in ksmbd_conn_write() and
ksmbd_conn_r_count_dec().

Take a real conn reference when publishing the conn pointer to the async
work item, and drop it after the notification work has decremented
r_count.  Apply the same lifetime rule to lease break notification, which
uses the same work->conn pattern.

Fixes: 3aa660c059 ("ksmbd: prevent connection release during oplock break notification")
Signed-off-by: Qihang <q.h.hack.winter@gmail.com>
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-16 10:18:25 -05:00
Ibrahim Hashimov
1c0ae3df69 ksmbd: fix integer overflow in set_file_allocation_info()
set_file_allocation_info() converts the client-supplied
FILE_ALLOCATION_INFORMATION::AllocationSize into a 512-byte block
count with:

	alloc_blks = (le64_to_cpu(file_alloc_info->AllocationSize) + 511) >> 9;

AllocationSize is a fully client-controlled __le64 field; the only
validation performed by the caller (smb2_set_info_file(), case
FILE_ALLOCATION_INFORMATION) is that the fixed buffer is at least
sizeof(struct smb2_file_alloc_info) == 8 bytes. The value itself is
never range-checked before this arithmetic.

When AllocationSize is close to U64_MAX (e.g. 0xffffffffffffffff),
"AllocationSize + 511" wraps around mod 2^64 to a small number
(0xffffffffffffffff + 511 = 510), so alloc_blks becomes 0. Since any
existing regular file has stat.blocks > 0, the function then takes
the "shrink" branch and calls:

	ksmbd_vfs_truncate(work, fp, alloc_blks * 512);   /* == 0 */

silently truncating the file to size 0, even though the client asked
to grow the allocation to (what looks like) the maximum possible
size. The trailing "if (size < alloc_blks * 512) i_size_write(inode,
size);" restore is guarded by a comparison that is never true once
alloc_blks == 0, so the truncation is not undone. This lets an
authenticated SMB client that already holds an open handle with
FILE_WRITE_DATA on a file silently truncate that same file to size 0
via a single crafted SET_INFO(FILE_ALLOCATION_INFORMATION) request
advertising a near-U64_MAX AllocationSize, even though the request
asks to grow the file's allocation rather than shrink it. This is a
functional/data-loss bug, not a privilege-boundary
violation: the same client could already truncate the file via
FILE_END_OF_FILE_INFORMATION or a plain write.

Fix it by validating AllocationSize against MAX_LFS_FILESIZE, the
same upper bound the VFS itself uses to reject unrepresentable file
sizes, before doing the "+511" rounding, and rejecting oversized
values with -EINVAL. Bounding AllocationSize to
MAX_LFS_FILESIZE - 511 guarantees the "+511" addition cannot wrap,
and that the subsequent "alloc_blks * 512" values passed to
vfs_fallocate() and ksmbd_vfs_truncate() stay within a representable
loff_t as well.

No legitimate SMB client asks for an allocation size anywhere near
2^64 bytes, so this only rejects a value that was previously
silently misinterpreted as zero.

Runtime-verified on a v6.19 KASAN test stand: sending SET_INFO
(FILE_ALLOCATION_INFORMATION) with AllocationSize = 0xffffffffffffffff
against ksmbd now returns -EINVAL and leaves the target file's size
unchanged, where the unpatched kernel truncated it from 4096 to 0
bytes.

Fixes: e2f34481b2 ("cifsd: add server-side procedures for SMB3")
Cc: stable@vger.kernel.org
Signed-off-by: Ibrahim Hashimov <security@auditcode.ai>
Assisted-by: AuditCode-AI:2026.07
Acked-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-16 10:18:25 -05:00
Huiwen He
d7d2adcd02 smb/client: flush dirty data before punching a hole
Punching a hole after a large buffered write may leave the range
reported as data. Reproduce it with:

  xfs_io -f \
    -c "pwrite -b 3m -S 0x61 0 3m" \
    -c "fpunch 1m 1m" \
    -c "seek -h 0" \
    -c "seek -d 1m" \
    /mnt/test/repro

Punching 1 MiB at offset 1 MiB should produce:

  0          1 MiB       2 MiB       3 MiB
  |  DATA    |   HOLE    |   DATA    | EOF

Instead, the entire file is reported as data. SEEK_HOLE(0) returns EOF,
and SEEK_DATA(1M) returns 1M.

This happens because a dirty folio spanning the punched range can be
written back after the punch and refill the hole.

Fix this by flushing and waiting for dirty data in the punched range
before invalidating the page cache and issuing FSCTL_SET_ZERO_DATA.

The xfstests generic/539 pass against Samba/ksmbd with this change.

Signed-off-by: Huiwen He <hehuiwen@kylinos.cn>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-14 21:51:52 -05:00
Andy Shevchenko
e2e08effef smb/client: Use EXPORT_SYMBOL_IF_KUNIT() to export symbols in SMB2
Replace EXPORT_SYMBOL_FOR_MODULES() with EXPORT_SYMBOL_IF_KUNIT()
to mark the symbols as visible only if CONFIG_KUNIT is enabled.

Kunit test should import the namespace EXPORTED_FOR_KUNIT_TESTING to
use these marked symbols. This is the standard way for all KUnit
tests.

Signed-off-by: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-14 21:51:19 -05:00
Andy Shevchenko
b8de511f0e smb/client: Use EXPORT_SYMBOL_IF_KUNIT() to export symbols
Replace EXPORT_SYMBOL_FOR_MODULES() with EXPORT_SYMBOL_IF_KUNIT()
to mark the symbols as visible only if CONFIG_KUNIT is enabled.

Kunit test should import the namespace EXPORTED_FOR_KUNIT_TESTING to
use these marked symbols. This is the standard way for all KUnit
tests.

Signed-off-by: Andy Shevchenko <andriy.shevchenko@linux.intel.com>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-14 21:48:00 -05:00
Linus Torvalds
7059bdf4f0 for-7.2-rc3-tag
-----BEGIN PGP SIGNATURE-----
 
 iQJPBAABCgA5FiEE8rQSAMVO+zA4DBdWxWXV+ddtWDsFAmpVwFMbFIAAAAAABAAO
 bWFudTIsMi41KzEuMTIsMiwyAAoJEMVl1fnXbVg70fgP/AqgKPhUBx5SAmTvFsno
 9vqa1g68BZtpsGeA6Djo2v3cCaRlFRwW9kcnKCiY25sNZEAtcsC+z3oT5ft1zRZE
 ZVIMFE4zQTGmjGNo/4AV2R7utXud+pq5HjQCiu5Ihjkl7/Qf8x1szI/tQzfXoHes
 KZfx97/F3oaC11b2+v/euey25jnRdrzus/5GvE2gMEbTGdogpIsBArLgKMecZeoZ
 NISD5STVU9TKUe6Dw40ivC0CAb5RdiSJabF4vN2/u/pXnzxFTnJ336LebLljgzzE
 c4PPFDYavGo2CRRyBSsxD6FLSN+RJUdZT4+L9GcPMnx20YcbhUF9NiW1QvQzmYcd
 +uvvCjtbBZSc/pyu6FOZcD5Ib9WkWjGkYEnq2knXDUMt5MDNSh4yo1D76zrb+yuR
 ez+YvcDBpYjmqOzlShjp1AhbpvaKYIEC0TW+1grS8vZVVy1vlVwqyC3ubh6ZpAJp
 Q/y50oS9nQkBUfnOdhQ1jhO+xHFzX3xnSzImqoX+ogVQKDfnEFeezaYcna/CDAeL
 jDP81abheSaO06+2gzlkSUXPJKh0TVwbyd34EQzs4Zft20UtpeHUfi5fCPzYZRGZ
 rRY/HXaztMBtbEOEzbjS/J0y1Y3t/DXcKRonftpudXdKP92DsNdOLK7y0vtI1A4i
 al79WpL+VaEw2+ByIZ3uanl/
 =A35I
 -----END PGP SIGNATURE-----

Merge tag 'for-7.2-rc3-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux

Pull btrfs fixes from David Sterba:

 - fix root structure leak after relocation error

 - fix optimization when checksums are read from commit root, fall back
   to checksum root during relocation

 - in tree-checker, validate length of inode reference in items

 - validate properties before setting them

 - validate free space cache entries on load

 - transaction abort fixes

 - fix printing of internal trees as signed numbers

 - add error messages after critical lzo compression errors

* tag 'for-7.2-rc3-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux:
  btrfs: print-tree: print header owner as signed
  btrfs: decentralize transaction aborts in create_reloc_root()
  btrfs: tree-checker: validate INODE_REF's namelen
  btrfs: lzo: add error message for invalid headers
  btrfs: fallback to transaction csum tree on a commit root csum miss
  btrfs: fix root leak if its reloc root is unexpected in merge_reloc_roots()
  btrfs: reject free space cache with more entries than pages
  btrfs: fix transaction abort logic in btrfs_fileattr_set()
  btrfs: validate properties before setting them
2026-07-14 09:10:27 -07:00
Darrick J. Wong
59c462b0f5 xfs: don't zap bmbt forks if they are MAXLEVELS tall
LOLLM noticed a discrepancy between the bmbt level checks in the libxfs
bmbt code vs. the inode repair code.  We do actually allow a bmbt root
that proclaims to have a height of XFS_BM_MAXLEVELS.

Cc: stable@vger.kernel.org # v6.8
Fixes: e744cef206 ("xfs: zap broken inode forks")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
15e38a9366 xfs: clamp timestamp nanoseconds correctly
LOLLM noticed an off-by-one error in the nsec clamping; fix that so that
we never have tv_nsec == 1e9.

Cc: stable@vger.kernel.org # v6.8
Fixes: 2d295fe657 ("xfs: repair inode records")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
ba150ce634 xfs: fully check the parent handle when it points to the rootdir
LOLLM noticed that the directory tree path checking declares the path to
be ok if the inumber in the parent pointer reaches the root directory.
Unfortunately, it neglects to check that the generation is correct.  Fix
that by moving the generation check up.

Cc: stable@vger.kernel.org # v6.10
Fixes: 928b721a11 ("xfs: teach online scrub to find directory tree structure problems")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
353a5900bc xfs: handle non-inode owners for rtrmap record checking
LOLLM noticed that two helper functions in the rtrmapbt scrub code don't
actually handle non-inode owners correctly -- CoW staging extents and
rgsuperblock extents are not shareable, but they are mergeable.  Fix
these two helpers.

Cc: stable@vger.kernel.org # v6.14
Fixes: 2d9a3e9805 ("xfs: allow overlapping rtrmapbt records for shared data extents")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
5d72a68f20 xfs: fix off-by-one error when calling xchk_xref_has_rt_owner
LOLLM noticed an off-by-one error when computing the length of the
rtrmap to cross-check.

Cc: stable@vger.kernel.org # v6.14
Fixes: 037a44d827 ("xfs: cross-reference the realtime rmapbt")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
540ddc6262 xfs: set xfarray killable sort correctly
LOLLM noticed that we *disable* interruptible sorts when the KILLABLE
flag is set.  This is backwards.  Fix the incorrect logic, and rename
the variable to make the connection more obvious.

Cc: stable@vger.kernel.org # v6.10
Fixes: 271557de7c ("xfs: reduce the rate of cond_resched calls inside scrub")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
ea6e2d9de2 xfs: grab rtrmap btree when checking rgsuper
LOLLM noticed that we aren't grabbing the rtrmap btree when we check the
realtime group superblock.  As a result, none of the cross-referencing
checks have ever run.  Fix this.

Cc: stable@vger.kernel.org # v6.14
Fixes: 428e488465 ("xfs: allow queued realtime intents to drain before scrubbing")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
9af789fa27 xfs: write the rg superblock when fixing it
The rtgroup superblock fixer should write the rtgroup superblock.
LOLLM noticed this, oops. :/

Cc: stable@vger.kernel.org # v6.13
Fixes: 1433f8f9ce ("xfs: repair realtime group superblock")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
881f2eb0fc xfs: use the rt version of the cow staging checker
LOLLM also noticed that xchk_rtrmapbt_xref ought to be using the rtdev
version of the "is this a cow extent?" helper function, not the datadev
one.

Cc: stable@vger.kernel.org # v6.14
Fixes: 91683bb3f2 ("xfs: cross-reference checks with the rt refcount btree")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
ee248157da xfs: use rtrefcount btree cursor in xchk_xref_is_rt_cow_staging
LOLLM points out that we pass the wrong btree cursor here.  We want the
rtrefcount btree cursor, not the non-rt one.  This is fairly benign
since it only affects tracing data.

Cc: stable@vger.kernel.org # v6.14
Fixes: 91683bb3f2 ("xfs: cross-reference checks with the rt refcount btree")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
d766e4e5e8 xfs: don't wrap around quota ids in dqiterate
LOLLM noticed that q_id is an unsigned 32-bit variable.  If it happens
to be set to XFS_DQ_ID_MAX due to a filesystem that actually has a dquot
for ID_MAX, then this addition will truncate to zero and the iteration
starts over.  Fix this by casting to u64.

Cc: stable@vger.kernel.org # v6.8
Fixes: 21d7500929 ("xfs: improve dquot iteration for scrub")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
60a1dde9d2 xfs: move cow_replace_mapping to xfs_bmap_util.c
Move the actual details of (partially) replacing a COW fork mapping to
xfs_bmap_util.c so that all the code doing hairy operations on subsets
of bmbt_irecs are kept together.

Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
bcb0621204 xfs: make cow repair somewhat flaky when debugging knob enabled
Introduce a new behavior for the cow fork repair code: if the debugging
knob is enabled, we'll pick a random subrange of each cow fork mapping
to mark as bad.  This will exercise the xrep_cow_replace_mapping more
thoroughly.

Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
a1caeeadbf xfs: don't replace the wrong part of the cow fork
LOLLM points out that xfs_iext_lookup_extent can return a @got where
got->br_startoff < startoff.  In this case, xrep_cow_replace_range
replaces the entire mapping instead of just the part that had been
marked bad in the bitmap, but advances the bitmap cursor in
xrep_cow_replace by the amount replaced.  As a result, we fail to
replace the end of the bad range, and replace part of the good range.

Fix this by rewriting the replace method to handle replacing the middle
of a cow fork mapping.  This we do by returning both the current mapping
as @got, and the subset of the mapping that we want to replace as @rep,
using @rep to store the results of the new allocation, and comparing
@rep to @got to figure out the exact transformations needed.

Cc: stable@vger.kernel.org # v6.8
Fixes: dbbdbd0086 ("xfs: repair problems in CoW forks")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Assisted-by: LOLLM # finding obvious bugs
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:01:47 +02:00
Darrick J. Wong
2f4acd0fcd xfs: resample the data fork mapping after cycling ILOCK
xfs_reflink_fill_{cow_hole,delalloc} are both presented with an inode,
a data fork mapping, and a cow fork mapping.  Unfortunately, these two
helpers cycle the ILOCK to grab a transaction, which means that the
mappings are stale as soon as we reacquire the ILOCK.  Currently we
refresh the cow fork mapping by re-calling xfs_find_trim_cow_extent, but
we don't refresh the data fork mapping beforehand, which means that the
xfs_bmap_trim_cow in that function queries the refcount btree about the
wrong physical blocks and returns an inaccurate value in *shared.

If *shared is now false, the directio write proceeds with a stale data
fork mapping.  Fix this by querying the data fork mapping if the
sequence counter changes across the ILOCK cycle.

Cc: hch@lst.de
Cc: stable@vger.kernel.org # v4.11
Fixes: 3c68d44a2b ("xfs: allocate direct I/O COW blocks in iomap_begin")
Signed-off-by: "Darrick J. Wong" <djwong@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
2026-07-14 11:00:14 +02:00
Daan De Meyer
75a41e3e51 btrfs: fix GET_SUBVOL_INFO after compat refactor
btrfs_search_slot() returns a positive value when the search key does
not exactly match an item. This is expected here, since offset 0 is used
to find the first ROOT_BACKREF for the subvolume and the actual key has
the parent root ID as its offset.

Before the compat ioctl refactoring, the native handler still copied the
filled structure to userspace when the search returned 1. After the
lookup was moved to a shared helper, both native and compat callers
treat the positive return value as a failure and skip copy_to_user(),
leaving BTRFS_IOC_GET_SUBVOL_INFO unusable for non-top-level
subvolumes.

Reset ret after successfully validating and reading the ROOT_BACKREF so
the helper reports success and both callers copy the result to
userspace.

Fixes: 538e5bdbc8 ("btrfs: add 32-bit compat ioctl for BTRFS_IOC_GET_SUBVOL_INFO")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Daan De Meyer <daan@amutable.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:05:10 +02:00
Guanghui Yang
6a8269b645 btrfs: free mapping node on duplicate reloc root insert
__add_reloc_root() allocates a mapping_node before inserting it into
rc->reloc_root_tree.  If rb_simple_insert() finds an existing entry, it
returns the existing rb_node and leaves the newly allocated node unlinked.

The error path then returns -EEXIST without freeing the new node.  Since
the node was never inserted into reloc_root_tree, the later cleanup in
put_reloc_control() cannot find it either.

Free the newly allocated node before returning -EEXIST.

The callers currently assert that -EEXIST should not happen, so this is a
defensive cleanup for an unexpected duplicate insert path.  If the path is
ever reached, the local allocation should still be released.

Fixes: 57a304cfd4 ("btrfs: do not panic in __add_reloc_root")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Guanghui Yang <3497809730@qq.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:05:00 +02:00
Qu Wenruo
9b73625a4f btrfs: fix a regression where PAGECACHE_TAG_DIRTY is never cleared
[BUG]
The following script (already submitted as generic/798) will report
incorrect dirty page numbers, with 64K page size systems and 4K fs block
size:

 # mkfs.btrfs -s 4k -f $dev
 # mount $dev $mnt
 # xfs_io -f -c "pwrite 0 64K" -c fsync -c "cachestat 0 64K" $mnt/foobar
 Cached: 1, Dirty: 1, Writeback: 0, Evicted: 0, Recently Evicted: 0

Note that the dirtied page number is still 1.

[CAUSE]
The cachestat() goes through the XArray of the page cache, but
instead of checking each folio's flag, it uses the
PAGECACHE_TAG_DIRTY tag to report dirty pages.

Since commit 095be159f3 ("btrfs: unify folio dirty flag clearing"),
btrfs replaced a folio_clear_dirty_for_io() call inside
extent_write_cache_pages() with folio_test_dirty().

This will cause the following call sequence for the folio at file offset
0:

 extent_write_cache_pages()
 |- folio_test_dirty()
 |  The folio is still dirty, continue to writeback.
 |
 |- extent_writepage()
    |- extent_writepage_io()
       |- submit_one_sector() for range [0, 4K)
       |  |- btrfs_folio_clear_dirty()
       |  |- btrfs_folio_set_writeback()
       |     |- folio_start_writeback()
       |        It's the first writeback block, we set the writeback
       |	flag for the folio.
       |	But the folio is still dirty, PAGECACHE_TAG_DIRTY is
       |	kept
       |
       |- submit_one_sector() for range [4K, 8K)
       |  |- btrfs_folio_clear_dirty()
       |  |- btrfs_folio_set_writeback()
       |     The folio already has writeback flag, no need to call
       |     folio_start_writeback()
       |
       | ...
       |- submit_one_sector() for range [60K, 64K)
	  |- btrfs_folio_clear_dirty()
	  |- btrfs_folio_set_writeback()
             The folio already has writeback flag, no need to call
             folio_start_writeback()

So the PAGECACHE_TAG_DIRTY is never cleared.

Meanwhile for the old code, before that commit, the sequence looks
like:

 extent_write_cache_pages()
 |- folio_clear_dirty_for_io()
 |  The folio is still dirty, so continue to writeback.
 |  But the folio dirty flag is cleared now.
 |
 |- extent_writepage()
    |- extent_writepage_io()
       |- submit_one_sector() for range [0, 4K)
       |  |- btrfs_folio_clear_dirty()
       |  |- btrfs_folio_set_writeback()
       |     |- folio_start_writeback()
       |        |- xas_clear(PAGECACHE_TAG)
       |
       |        It's the first writeback block, we set the writeback
       |	flag for the folio.
       |	And the folio is not dirty, PAGECACHE_TAG_DIRTY is
       |        cleared
       |
       |- submit_one_sector() for range [4K, 8K)
       |  |- btrfs_folio_clear_dirty()
       |  |- btrfs_folio_set_writeback()
       |     The folio already has writeback flag, no need to call
       |     folio_start_writeback()
       |
       | ...
       |- submit_one_sector() for range [60K, 64K)
	  |- btrfs_folio_clear_dirty()
	  |- btrfs_folio_set_writeback()
             The folio already has writeback flag, no need to call
             folio_start_writeback()

Unlike the new code, old code will clear PAGECACHE_TAG_DIRTY for the
first writeback block.

There is a deeper problem, dirty and writeback folio flags are updated
at very different timing.
The dirty flag is only cleared when the last sub-folio block has dirty
flag cleared.
But the writeback flag is set when the first block starts writeback, and
later blocks that go through writeback will not call
folio_start_writeback() again.

If we rely on folio_start_writeback() to update the
PAGECACHE_TAG_DIRTY and PAGECACHE_TAG_TOWRITE, it will always be
incorrect in one way or another.

[FIX]
Do not let folio_start_writeback() do any PAGECACHE_TAG_TOWRITE
handling.

Instead, manually clear both PAGECACHE_TAG_TOWRITE and
PAGECACHE_TAG_DIRTY flags when the folio is no longer dirty during
btrfs_subpage_set_writeback().

However this is only a hot-fix, for the long term solution we will
follow iomap, by calling folio_start_writeback() immediately for the
whole folio, and folio_end_writeback() after all writeback finished
for the folio.

Fixes: 095be159f3 ("btrfs: unify folio dirty flag clearing")
Reviewed-by: Boris Burkov <boris@bur.io>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:04:48 +02:00
Leo Martins
5eff4d5b17 btrfs: don't propagate EXTENT_FLAG_LOGGING to split extent maps
When btrfs_drop_extent_map_range() splits an extent map, the new split
maps inherit the original map's flags through a local 'flags' variable.
Commit f86f7a75e2 ("btrfs: use the flags of an extent map to identify
the compression type") changed the EXTENT_FLAG_LOGGING clearing to
operate on em->flags instead of that local 'flags' copy, so a split of
an extent map that is currently being logged wrongly inherits
EXTENT_FLAG_LOGGING.

The flag is then never cleared on the split, and when it is freed while
still on the inode's modified_extents list (for example by the extent
map shrinker) it trips the WARN_ON(!list_empty(&em->list)) in
btrfs_free_extent_map() and leads to a use-after-free.

Clear EXTENT_FLAG_LOGGING from the local 'flags' copy used for the
splits and only clear EXTENT_FLAG_PINNED from em->flags, restoring the
behaviour prior to f86f7a75e2.

CC: Jeff Layton <jlayton@kernel.org>
Link: https://lore.kernel.org/all/20260629-btrfs-skip-logging-v1-1-4e3a28c1acaf@kernel.org/
Fixes: f86f7a75e2 ("btrfs: use the flags of an extent map to identify the compression type")
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Leo Martins <loemra.dev@gmail.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:04:22 +02:00
Dave Chen
8b5a09ceb6 btrfs: fix u32 to s64 type conversion in dirty_metadata_bytes accounting
The percpu_counter dirty_metadata_bytes is updated by negating eb->len
and passing it to percpu_counter_add_batch(), whose amount parameter is
s64.  Since commit 84cda1a608 ("btrfs: cache folio size and shift in
extent_buffer"), eb->len is u32.  The u32 result of -eb->len, when
widened to the s64 parameter, becomes a large positive value instead of
the intended negative value.  For eb->len == 16384 the counter adds
+4294950912 instead of subtracting 16384.

The counter therefore grows on every metadata writeback instead of
shrinking by the extent buffer size, permanently exceeding
BTRFS_DIRTY_METADATA_THRESH and causing __btrfs_btree_balance_dirty()
to trigger balance_dirty_pages_ratelimited() unconditionally, adding
unnecessary writeback pressure.

Cast eb->len to s64 before negation at both call sites so the
subtraction is performed in signed 64-bit arithmetic.

Reviewed-by: Filipe Manana <fdmanana@suse.com>
Fixes: 84cda1a608 ("btrfs: cache folio size and shift in extent_buffer")
Signed-off-by: Dave Chen <davechen@synology.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:04:19 +02:00
Filipe Manana
f0c1f14cc1 btrfs: fix NULL pointer deref during assertion in btrfs_backref_free_node()
In btrfs_backref_free_node() we have the following assertion:

  ASSERT(node->eb == NULL, "node->eb->start=%llu", node->eb->start);

and a user reported the following crash:

  Oops: general protection fault, probably for non-canonical address 0xdffffc0000000000: 0000 [#1] SMP KASAN NOPTI
  KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007]
  CPU: 0 UID: 0 PID: 10422 Comm: syz.0.17 Not tainted 7.1.0-02765-g6b5a2b7d9bc1-dirty #44 PREEMPT(full)
  Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.15.0-1 04/01/2014
  RIP: 0010:btrfs_backref_free_node fs/btrfs/backref.c:3057 [inline]
  RIP: 0010:btrfs_backref_free_node+0xb9/0x200 fs/btrfs/backref.c:3051
  Code: 00 fc ff (...)
  RSP: 0018:ffa0000006b0f3c0 EFLAGS: 00010246
  RAX: dffffc0000000000 RBX: 0000000000000000 RCX: ffffffff840eb78b
  RDX: 0000000000000000 RSI: ffffffff840eafa5 RDI: ff110000742ab768
  RBP: ff110000742ab700 R08: 0000000000000000 R09: 0000000000000000
  R10: ff110000742ab700 R11: 00000000000a81f9 R12: ff11000107a92020
  R13: ff1100005c182ea8 R14: 0000000000000000 R15: dffffc0000000000
  FS:  0000555575536500(0000) GS:ff11000183985000(0000) knlGS:0000000000000000
  CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
  CR2: 00007fa3d0e9d580 CR3: 000000002232a000 CR4: 0000000000753ef0
  PKRU: 00000000
  Call Trace:
   <TASK>
   btrfs_backref_cleanup_node+0x27/0x30 fs/btrfs/backref.c:3133
   relocate_tree_block fs/btrfs/relocation.c:2604 [inline]
   relocate_tree_blocks+0x11b0/0x1a20 fs/btrfs/relocation.c:2707
   relocate_block_group+0x499/0xf30 fs/btrfs/relocation.c:3635
   do_nonremap_reloc fs/btrfs/relocation.c:5323 [inline]
   btrfs_relocate_block_group+0x1749/0x5fb0 fs/btrfs/relocation.c:5490
   btrfs_relocate_chunk+0x12b/0x950 fs/btrfs/volumes.c:3647
   __btrfs_balance fs/btrfs/volumes.c:4586 [inline]
   btrfs_balance+0x1c7f/0x55c0 fs/btrfs/volumes.c:4973
   btrfs_ioctl_balance fs/btrfs/ioctl.c:3474 [inline]
   btrfs_ioctl+0x38a4/0x5d20 fs/btrfs/ioctl.c:5570
   vfs_ioctl fs/ioctl.c:51 [inline]
   __do_sys_ioctl fs/ioctl.c:597 [inline]
   __se_sys_ioctl fs/ioctl.c:583 [inline]
   __x64_sys_ioctl+0x18f/0x210 fs/ioctl.c:583
   do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
   do_syscall_64+0x11f/0x860 arch/x86/entry/syscall_64.c:94
   entry_SYSCALL_64_after_hwframe+0x77/0x7f
   RIP: 0033:0x7fb38e3b56dd
   Code: 02 b8 ff (...)
   RSP: 002b:00007fff04115788 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
   RAX: ffffffffffffffda RBX: 00007fb38f6b0020 RCX: 00007fb38e3b56dd
   RDX: 00002000000003c0 RSI: 00000000c4009420 RDI: 0000000000000004
   RBP: 00007fb38e451b48 R08: 0000000000000000 R09: 0000000000000000
   R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
   R13: 0000000000000000 R14: 00007fb38f6b0020 R15: 00007fb38f6b002c
   </TASK>

It seems that this happens on some systems for some reason, when the
ASSERT() macro calls the inline function verify_assert_printk_format()
to evaluate the format string and arguments, causing the NULL pointer
dereference on node->eb.

So change the assertion to check for a NULL node->eb before dereferencing
it. Also, while at it, make the assertion more useful by printing the
owner of the extent buffer as well as its level.

Reported-by: Yue Sun <samsun1006219@gmail.com>
Link: https://lore.kernel.org/linux-btrfs/20260626065542.38413-1-samsun1006219@gmail.com/
Fixes: c4e7778580 ("btrfs: use verbose assertions in backref.c")
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Reviewed-by: David Sterba <dsterba@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:04:10 +02:00
Dave Chen
9411aafdf3 btrfs: only account delalloc bytes for regular file inodes in btrfs_getattr()
btrfs_getattr() unconditionally reads BTRFS_I(inode)->new_delalloc_bytes
and adds it (sector-aligned) to stat->blocks for every inode type.
However, new_delalloc_bytes lives in a union with last_dir_index_offset:

    union {
        u64 new_delalloc_bytes;     /* files only */
        u64 last_dir_index_offset;  /* directories only */
    };

For a directory inode this memory holds last_dir_index_offset, which is
set during directory logging (e.g. flush_dir_items_batch()) to the
offset of the last logged BTRFS_DIR_INDEX_KEY.  That offset grows with
the number of entries ever created in the directory (dir indexes are
monotonic and never reused), so it can be arbitrarily large.

As a result, after a directory has been logged (e.g. via an fsync that
triggers directory logging), btrfs_getattr() reports inflated st_blocks
for that directory.  The inflation is purely in-core and disappears
after the inode is evicted and reloaded (btrfs_alloc_inode() zeroes the
union), e.g. after a remount.

Reproducer (on a btrfs filesystem):

    D=/mnt/btrfs/d
    mkdir -p $D
    for i in $(seq 1 20000); do touch $D/f$i; done
    sync                      # commit, push dir index high
    touch $D/trigger          # dirty the dir in a new transaction
    xfs_io -c fsync $D        # log the directory -> sets last_dir_index_offset
    stat -c '%b' $D           # st_blocks is now inflated (e.g. 40)
    # umount + mount -> st_blocks drops back to the correct value

The evict path already knows this union is type-dependent and guards the
corresponding WARN_ON with !S_ISDIR() in btrfs_destroy_inode(); only
btrfs_getattr() was missing the equivalent check.

Only read new_delalloc_bytes for regular files, which are the only
inodes that ever set it.

Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Dave Chen <davechen@synology.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:03:26 +02:00
Qu Wenruo
800b519602 btrfs: reject inline file extents item in get_new_location()
Commit a6908f88c9 ("btrfs: validate data reloc tree file extent item
members") introduced extra checks on file extent items for data reloc
inodes, but it checked the file extent offset without checking if the file
extent is inlined.

This can lead to either false alerts (as the offset member is inside the
inlined data) or even reading beyond the item range.

This has already triggered a warning in a syzbot report.
Although the root fix is to avoid compression for data reloc inodes, for
the sake of consistency, reject inlined file extents first.

Fixes: a6908f88c9 ("btrfs: validate data reloc tree file extent item members")
CC: stable@vger.kernel.org
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:03:15 +02:00
Qu Wenruo
ae4316f332 btrfs: do not try compression for data reloc inodes
[BUG]
There is a syzbot report that the check inside get_new_location()
triggered:

  BTRFS info (device loop0): found 31 extents, stage: move data extents
  BTRFS info (device loop0): leaf 8908800 gen 16 total ptrs 28 free space 1676 owner 18446744073709551607
         item 0 key (256 INODE_ITEM 0) itemoff 3835 itemsize 160
                 inode generation 5 transid 0 size 0 nbytes 0
                 block group 0 mode 40755 links 1 uid 0 gid 0
                 rdev 0 sequence 0 flags 0x0
                 atime 1669132761.0
                 ctime 1669132761.0
                 mtime 1669132761.0
                 otime 0.0
         item 1 key (256 INODE_REF 256) itemoff 3823 itemsize 12
                 index 0 name_len 2
         item 2 key (258 INODE_ITEM 0) itemoff 3663 itemsize 160
                 inode generation 1 transid 16 size 733184 nbytes 106496
                 block group 0 mode 100600 links 0 uid 0 gid 0
                 rdev 0 sequence 24 flags 0x18
         item 3 key (258 EXTENT_DATA 0) itemoff 3595 itemsize 68
                 generation 16 type 0
                 inline extent data size 47 ram_bytes 4096 compression 1
  [...]
         item 27 key (18446744073709551611 ORPHAN_ITEM 258) itemoff 2376 itemsize 0
  BTRFS error (device loop0): unexpected non-zero offset in file extent item for data reloc inode 258 key offset 0 offset 9277520992061368337
  ------------[ cut here ]------------
  btrfs_abort_should_print_stack(__error)

[CAUSE]
The above dump tree shows the first file extent item is inlined, which
should make no sense for data reloc inodes, as such inodes just
represent where the data extents are in the relocation destination chunk.

However the relocation path preallocates space for each block,
then dirties them, cluster by cluster.
It's possible to have a single block at the beginning of the block
group, and no other block in the same cluster.

So relocation will preallocate a file extent for that block and dirty
the first block.  Then memory pressure forces the data reloc inode to be
written back, before any other blocks are dirtied/allocated.

Finally commit 3eaf5f082c ("btrfs: extract inlined creation into a dedicated
delalloc helper") changed the sequence of delalloc. Before that commit we
always tried NOCOW first, so that dirtied block would be written back into
the preallocated space, and appear as a regular extent.

But with that commit, we always try inline first, and since compression
is forced, we try compressing the first block, and then inline the
compressed data, resulting in the above inlined file extent in the data
reloc tree.

Then the check in get_new_location() will check the file offset, without
checking if the file extent is inlined or not, resulting in the above
failure.

[FIX]
Do not allow compression for data reloc inodes.

Since data reloc inode sizes are always block aligned, as long as we do
not compress, @data_len will always be at least one block, and
that will cause can_cow_file_range_inline() to return false, thus no
inlined extent will be created.

Reported-by: syzbot+d950c6ba09b79f6e1864@syzkaller.appspotmail.com
Link: https://lore.kernel.org/linux-btrfs/6a373dc5.764cf64f.168fbe.0001.GAE@google.com/
Fixes: 3eaf5f082c ("btrfs: extract inlined creation into a dedicated delalloc helper")
CC: stable@vger.kernel.org
Reviewed-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:03:02 +02:00
Filipe Manana
b78fe9563e btrfs: fix reloc root cleanup in merge_reloc_roots()
If the root we got has zero root refs in its root item, we are resetting
the root's ->reloc_root without using barriers like we do everywhere else.
Sashiko complained about this while reviewing another patch, and it's
correct (see the Link tag below).

Also, we should not clear BTRFS_ROOT_DEAD_RELOC_TREE from the root unless
the root points to the reloc root we have.

Fix this by using clear_reloc_root(), which issues the memory barrier
after setting the root's ->reloc_root to NULL and before clearing the bit
BTRFS_ROOT_DEAD_RELOC_TREE from the root.

Link: https://sashiko.dev/#/patchset/cf84f1a217c719e25b6b69e4298dd7afd36c9427.1781194426.git.fdmanana%40suse.com
Reviewed-by: Boris Burkov <boris@bur.io>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:02:13 +02:00
Filipe Manana
83201804ef btrfs: fix use-after-free on reloc root after error in insert_dirty_subvol()
If during relocation we fail in insert_dirty_subvol() because
btrfs_update_reloc_root() returned an error, we will leave a root's
reloc_root field pointing to a reloc root that was freed instead of NULL,
resulting later in a use-after-free, or double free attempt during
unmount.

The sequence of steps is this:

1) During relocation the call to btrfs_update_reloc_root() in
   insert_dirty_subvol() fails, so insert_dirty_subvol() returns the
   error to merge_reloc_root() without adding the root to the list
   rc->dirty_subvol_roots;

2) Then merge_reloc_root() aborts the current transaction because
   insert_dirty_subvol() returned an error;

3) Up the call chain, merge_reloc_roots() gets the error, adds the
   reloc root for root X to the local reloc_roots list and jumps to the
   'out' label, where it calls free_reloc_roots() to free all the reloc
   roots in the local reloc_roots list. This frees the reloc root for
   root X;

4) We go up the call chain to relocate_block_group() which calls
   clean_dirty_subvols() to go over dirty roots and set their
   ->reloc_root field to NULL, but root X is not in the dirty_subvol_roots
   list, so its ->reloc_root still points to a reloc root;

5) Relocation finishes, with an error and a transaction abort, but the
   ->reloc_root field for root X still points to the reloc root that was
   freed in step 3;

6) When unmounting the fs we end up calling:

     btrfs_free_fs_roots()
        btrfs_drop_and_free_fs_root()
           --> calls btrfs_put_root() against root X's ->reloc_root
               which is not NULL and points to the already freed
               reloc root in step 4 above

  Resulting in a use-after-free to a double free attempt.

Syzbot reported this with the following dmesg/syslog:

   [  106.004389][ T5339] BTRFS error (device loop0 state A): Transaction aborted (error -5)
   [  106.014266][ T5339] BTRFS: error (device loop0 state A) in merge_reloc_root:1655: errno=-5 IO failure
   [  106.021891][ T1061] BTRFS error (device loop0 state A): error while writing out transaction: -5
   [  106.026964][ T1061] BTRFS warning (device loop0 state A): Skipping commit of aborted transaction.
   [  106.033807][ T5340] BTRFS error (device loop0 state A): bdev /dev/loop0 errs: wr 3, rd 0, flush 0, corrupt 0, gen 0
   [  106.039265][ T1061] BTRFS: error (device loop0 state A) in cleanup_transaction:2067: errno=-5 IO failure
   [  106.044382][ T5339] BTRFS info (device loop0 state EA): forced readonly
   [  106.074329][ T5339] BTRFS: error (device loop0 state EA) in merge_reloc_roots:1887: errno=-5 IO failure
   [  106.081004][ T5356] BTRFS info (device loop0 state EA): scrub: started on devid 1
   [  106.085611][ T5339] BTRFS info (device loop0 state EA): balance: ended with status: -30
   [  106.089517][ T5356] BTRFS info (device loop0 state EA): scrub: not finished on devid 1 with status: -30
   [  106.662365][ T5338] BTRFS info (device loop0 state EA): last unmount of filesystem 3a375e4e-b156-4d76-a2ad-16e198ce1409
   [  106.682946][ T5338] ==================================================================
   [  106.686574][ T5338] BUG: KASAN: slab-use-after-free in btrfs_put_root+0x2f/0x250
   [  106.690090][ T5338] Write of size 4 at addr ffff88803f978630 by task syz.0.0/5338
   [  106.693173][ T5338]
   [  106.694279][ T5338] CPU: 0 UID: 0 PID: 5338 Comm: syz.0.0 Not tainted syzkaller #0 PREEMPT(full)
   [  106.694293][ T5338] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
   [  106.694300][ T5338] Call Trace:
   [  106.694308][ T5338]  <TASK>
   [  106.694314][ T5338]  dump_stack_lvl+0xe8/0x150
   [  106.694331][ T5338]  print_address_description+0x55/0x1e0
   [  106.694343][ T5338]  ? btrfs_put_root+0x2f/0x250
   [  106.694358][ T5338]  print_report+0x58/0x70
   [  106.694368][ T5338]  kasan_report+0x117/0x150
   [  106.694384][ T5338]  ? btrfs_put_root+0x2f/0x250
   [  106.694399][ T5338]  kasan_check_range+0x264/0x2c0
   [  106.694416][ T5338]  btrfs_put_root+0x2f/0x250
   [  106.694430][ T5338]  btrfs_drop_and_free_fs_root+0x160/0x210
   [  106.694447][ T5338]  btrfs_free_fs_roots+0x2f9/0x3c0
   [  106.694464][ T5338]  ? __pfx_btrfs_free_fs_roots+0x10/0x10
   [  106.694479][ T5338]  ? free_root_pointers+0x5bf/0x5f0
   [  106.694494][ T5338]  close_ctree+0x798/0x12d0
   [  106.694511][ T5338]  ? __pfx_close_ctree+0x10/0x10
   [  106.694526][ T5338]  ? _raw_spin_unlock_irqrestore+0x74/0x80
   [  106.694599][ T5338]  ? rcu_preempt_deferred_qs_irqrestore+0x906/0xbc0
   [  106.694620][ T5338]  ? __rcu_read_unlock+0x83/0xe0
   [  106.694636][ T5338]  ? btrfs_put_super+0x48/0x1c0
   [  106.694652][ T5338]  ? __pfx_btrfs_put_super+0x10/0x10
   [  106.694667][ T5338]  generic_shutdown_super+0x13d/0x2d0
   [  106.694682][ T5338]  kill_anon_super+0x3b/0x70
   [  106.694695][ T5338]  btrfs_kill_super+0x41/0x50
   [  106.694710][ T5338]  deactivate_locked_super+0xbc/0x130
   [  106.694722][ T5338]  cleanup_mnt+0x437/0x4d0
   [  106.694736][ T5338]  ? _raw_spin_unlock_irq+0x23/0x50
   [  106.694752][ T5338]  task_work_run+0x1d9/0x270
   [  106.694769][ T5338]  ? __pfx_task_work_run+0x10/0x10
   [  106.694784][ T5338]  ? do_raw_spin_unlock+0x4d/0x210
   [  106.694802][ T5338]  do_exit+0x70f/0x22c0
   [  106.694817][ T5338]  ? trace_irq_disable+0x3b/0x140
   [  106.694835][ T5338]  ? __pfx_do_exit+0x10/0x10
   [  106.694848][ T5338]  ? preempt_schedule_thunk+0x16/0x30
   [  106.694863][ T5338]  ? preempt_schedule_common+0x82/0xd0
   [  106.694878][ T5338]  ? preempt_schedule_thunk+0x16/0x30
   [  106.694892][ T5338]  do_group_exit+0x21b/0x2d0
   [  106.694906][ T5338]  ? entry_SYSCALL_64_after_hwframe+0x77/0x7f
   [  106.694918][ T5338]  __x64_sys_exit_group+0x3f/0x40
   [  106.694932][ T5338]  x64_sys_call+0x221a/0x2240
   [  106.694944][ T5338]  do_syscall_64+0x174/0x580
   [  106.694954][ T5338]  ? clear_bhb_loop+0x40/0x90
   [  106.694967][ T5338]  entry_SYSCALL_64_after_hwframe+0x77/0x7f
   [  106.694978][ T5338] RIP: 0033:0x7f958ef9ce59
   [  106.694988][ T5338] Code: Unable to access opcode bytes at 0x7f958ef9ce2f.
   [  106.694994][ T5338] RSP: 002b:00007fffd4058318 EFLAGS: 00000246 ORIG_RAX: 00000000000000e7
   [  106.695008][ T5338] RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 00007f958ef9ce59
   [  106.695015][ T5338] RDX: 00007f958c3f8000 RSI: 0000000000000000 RDI: 0000000000000000
   [  106.695022][ T5338] RBP: 0000000000000003 R08: 0000000000000000 R09: 00007f958f1e73e0
   [  106.695028][ T5338] R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
   [  106.695034][ T5338] R13: 00007f958f1e73e0 R14: 0000000000000003 R15: 00007fffd40583d0
   [  106.695046][ T5338]  </TASK>
   [  106.695050][ T5338]
   [  106.821635][ T5338] Allocated by task 1061:
   [  106.823446][ T5338]  kasan_save_track+0x3e/0x80
   [  106.825498][ T5338]  __kasan_kmalloc+0x93/0xb0
   [  106.827381][ T5338]  __kmalloc_cache_noprof+0x31c/0x660
   [  106.829525][ T5338]  btrfs_alloc_root+0x75/0x930
   [  106.831458][ T5338]  read_tree_root_path+0x127/0xb00
   [  106.833556][ T5338]  btrfs_read_tree_root+0x34/0x60
   [  106.835553][ T5338]  create_reloc_root+0x6b3/0xcb0
   [  106.837556][ T5338]  btrfs_init_reloc_root+0x2ec/0x4b0
   [  106.839557][ T5338]  record_root_in_trans+0x2ab/0x350
   [  106.841685][ T5338]  btrfs_record_root_in_trans+0x15c/0x180
   [  106.844237][ T5338]  start_transaction+0x39c/0x1820
   [  106.846638][ T5338]  btrfs_finish_one_ordered+0x88e/0x2680
   [  106.849436][ T5338]  btrfs_work_helper+0x37b/0xc20
   [  106.851549][ T5338]  process_scheduled_works+0xb5d/0x1860
   [  106.853807][ T5338]  worker_thread+0xa53/0xfc0
   [  106.855773][ T5338]  kthread+0x389/0x470
   [  106.857548][ T5338]  ret_from_fork+0x514/0xb70
   [  106.859493][ T5338]  ret_from_fork_asm+0x1a/0x30
   [  106.861504][ T5338]
   [  106.862527][ T5338] Freed by task 5339:
   [  106.864224][ T5338]  kasan_save_track+0x3e/0x80
   [  106.866180][ T5338]  kasan_save_free_info+0x46/0x50
   [  106.868371][ T5338]  __kasan_slab_free+0x5c/0x80
   [  106.870462][ T5338]  kfree+0x1c5/0x640
   [  106.872180][ T5338]  __del_reloc_root+0x341/0x3b0
   [  106.874290][ T5338]  free_reloc_roots+0x5f/0x90
   [  106.876282][ T5338]  merge_reloc_roots+0x73f/0x8a0
   [  106.878489][ T5338]  relocate_block_group+0xbcc/0xe70
   [  106.880742][ T5338]  do_nonremap_reloc+0xa8/0x5b0
   [  106.882885][ T5338]  btrfs_relocate_block_group+0x7e6/0xc40
   [  106.885336][ T5338]  btrfs_relocate_chunk+0x115/0x820
   [  106.887502][ T5338]  __btrfs_balance+0x1db0/0x2ae0
   [  106.889543][ T5338]  btrfs_balance+0xaf3/0x11b0
   [  106.891456][ T5338]  btrfs_ioctl_balance+0x3d3/0x610
   [  106.893672][ T5338]  __se_sys_ioctl+0xfc/0x170
   [  106.895530][ T5338]  do_syscall_64+0x174/0x580
   [  106.897518][ T5338]  entry_SYSCALL_64_after_hwframe+0x77/0x7f
   [  106.900101][ T5338]
   [  106.901123][ T5338] The buggy address belongs to the object at ffff88803f978000
   [  106.901123][ T5338]  which belongs to the cache kmalloc-4k of size 4096
   [  106.906907][ T5338] The buggy address is located 1584 bytes inside of
   [  106.906907][ T5338]  freed 4096-byte region [ffff88803f978000, ffff88803f979000)
   [  106.912980][ T5338]
   [  106.914022][ T5338] The buggy address belongs to the physical page:
   [  106.916716][ T5338] page: refcount:0 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x3f978
   [  106.920390][ T5338] head: order:3 mapcount:0 entire_mapcount:0 nr_pages_mapped:0 pincount:0
   [  106.923834][ T5338] flags: 0x4fff00000000040(head|node=1|zone=1|lastcpupid=0x7ff)
   [  106.927104][ T5338] page_type: f5(slab)
   [  106.928898][ T5338] raw: 04fff00000000040 ffff88801ac42140 dead000000000122 0000000000000000
   [  106.932507][ T5338] raw: 0000000000000000 0000000800040004 00000000f5000000 0000000000000000
   [  106.936193][ T5338] head: 04fff00000000040 ffff88801ac42140 dead000000000122 0000000000000000
   [  106.939856][ T5338] head: 0000000000000000 0000000800040004 00000000f5000000 0000000000000000
   [  106.943601][ T5338] head: 04fff00000000003 fffffffffffffe01 00000000ffffffff 00000000ffffffff
   [  106.947268][ T5338] head: ffffffffffffffff 0000000000000000 00000000ffffffff 0000000000000008
   [  106.950988][ T5338] page dumped because: kasan: bad access detected
   [  106.953710][ T5338] page_owner tracks the page as allocated
   [  106.956198][ T5338] page last allocated via order 3, migratetype Unmovable, gfp_mask 0xd2820(GFP_ATOMIC|__GFP_NOWARN|__GFP_NORETRY|__GFP_COMP|__GFP_NOMEMALLOC), pid 24, tgid 24 (kworker/u4:2), ts 105728970387, free_ts 29540875453
   [  106.964984][ T5338]  post_alloc_hook+0x22d/0x280
   [  106.966956][ T5338]  get_page_from_freelist+0x2593/0x2610
   [  106.969307][ T5338]  __alloc_frozen_pages_noprof+0x18d/0x380
   [  106.971839][ T5338]  allocate_slab+0x77/0x660
   [  106.973709][ T5338]  refill_objects+0x339/0x3d0
   [  106.975696][ T5338]  __pcs_replace_empty_main+0x321/0x720
   [  106.978136][ T5338]  __kmalloc_node_track_caller_noprof+0x572/0x7b0
   [  106.981009][ T5338]  __alloc_skb+0x2c1/0x7d0
   [  106.982983][ T5338]  nsim_dev_trap_report_work+0x29a/0xb90
   [  106.985356][ T5338]  process_scheduled_works+0xb5d/0x1860
   [  106.987710][ T5338]  worker_thread+0xa53/0xfc0
   [  106.989847][ T5338]  kthread+0x389/0x470
   [  106.991727][ T5338]  ret_from_fork+0x514/0xb70
   [  106.993722][ T5338]  ret_from_fork_asm+0x1a/0x30
   [  106.995900][ T5338] page last free pid 77 tgid 77 stack trace:
   [  106.998479][ T5338]  __free_frozen_pages+0xc1c/0xd30
   [  107.000819][ T5338]  vfree+0x1d1/0x2f0
   [  107.002631][ T5338]  delayed_vfree_work+0x55/0x80
   [  107.004848][ T5338]  process_scheduled_works+0xb5d/0x1860
   [  107.007366][ T5338]  worker_thread+0xa53/0xfc0
   [  107.009388][ T5338]  kthread+0x389/0x470
   [  107.011177][ T5338]  ret_from_fork+0x514/0xb70
   [  107.013313][ T5338]  ret_from_fork_asm+0x1a/0x30
   [  107.015454][ T5338]
   [  107.016460][ T5338] Memory state around the buggy address:
   [  107.019052][ T5338]  ffff88803f978500: fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb
   [  107.022691][ T5338]  ffff88803f978580: fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb
   [  107.026264][ T5338] >ffff88803f978600: fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb
   [  107.029721][ T5338]                                      ^
   [  107.032062][ T5338]  ffff88803f978680: fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb
   [  107.035547][ T5338]  ffff88803f978700: fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb
   [  107.038865][ T5338] ==================================================================

Fix this by resetting a root's ->reloc_root if we get an error while
trying to merge a reloc root.

Reported-by: syzbot+b3d472d13f9d7bf20669@syzkaller.appspotmail.com
Link: https://lore.kernel.org/linux-btrfs/6a1ebde9.c1435f33.112120.0176.GAE@google.com/
Reviewed-by: Qu Wenruo <wqu@suse.com>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-07-14 07:01:55 +02:00
Linus Torvalds
0f26556c5e nfsd-7.2 fixes:
Issues reported with v7.2-rc:
 - Fix a UAF when unlocking a filesystem
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEKLLlsBKG3yQ88j7+M2qzM29mf5cFAmpUTYQACgkQM2qzM29m
 f5fluA//Ttq+/8JVZnAwxCYyMTM+ix8S1qYAP9eXxB4OwHVjVGdm/aAHbm6fmnuw
 IiMdJw0ZB2RQEFrD+jLJxWipYil/oHYUhUsS1sxSLTE9RMeAVnD4SbGEO1Hg2imV
 zISIpM7kNkh7NcJkKgzb8mCh04C6TSLwUsq36sOf6H6JlWJsQjM9j6t343pET3NY
 IDxjCIz/C7wfPRHsg+ycsXCvbrF3tE27Grq2HX0/0tPIo8JoqR3sCgIWz00UN3NT
 QE0FfPU0fUmn0TW6WupMmibC/cSBjUkEayEIWOuRdo2HfjdHFJKKVEJhbPfueHZU
 NiTTyMaIC2bKXclAZR9CEggyZmdfJTBMkmQSIvMKr2XWYKcLFTrVfcXpU3iuYgLq
 R97keupMX5YrBA1HZa+pWozuBdBxYD4ykJFFYy3hF+84X8JY4om3g40mtIG1P52/
 lc61qZGqMGddkPPL7pGVxJ0l6husKeCd3T/vo1HupeRg1swOB0VzCYK9NNo1tBO4
 yf7ZdqOE2r7crmYpIE2woNdXs9fGpkpZriLq8iWvlVWIuHHfjYolk0a62QIO/GV1
 Tgvu442+gJYcf+QLOZIf+YRA1LvNxXyENR/ZKXDhQ85h/nTEULKL5Ne9Fqq2qOTk
 5YCNE6/XkB5sGNA/aF06RPDnUjuBNX7PTEMt5a+t2ikB5nkq/YY=
 =NRRN
 -----END PGP SIGNATURE-----

Merge tag 'nfsd-7.2-1' of git://git.kernel.org/pub/scm/linux/kernel/git/cel/linux

Pull nfsd fix from Chuck Lever:

 - Fix a use-after-free when unlocking a filesystem

* tag 'nfsd-7.2-1' of git://git.kernel.org/pub/scm/linux/kernel/git/cel/linux:
  NFSD: Prevent post-shutdown use-after-free in NFSD_CMD_UNLOCK_FILESYSTEM
2026-07-13 08:38:01 -07:00
Shoichiro Miyamoto
8986c93290 smb: client: reject overlapping data areas in SMB2 responses
Commit 53b7c271f0 ("smb: client: restrict implied bcc[0] exemption to
responses without data area") restricted the implied bcc[0] length
exception to responses without a data area. However, the overlap
handling in __smb2_calc_size() clears data_length, which can make an
invalid response appear to have no data area and so qualify for the
exception.

Track data area overlap separately and reject such responses before
applying the length compatibility exceptions.

Fixes: 53b7c271f0 ("smb: client: restrict implied bcc[0] exemption to responses without data area")
Cc: stable@vger.kernel.org
Signed-off-by: Shoichiro Miyamoto <shoichiro.miyamoto@gmail.com>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-13 10:37:55 -05:00
Huiwen He
5bd1d3dcc2 smb/client: refresh allocation after EOF-extending fallocate
Before this change, xfstests generic/496 was not supported on ksmbd:

        generic/496 ... [not run] fallocated swap not supported here

ksmbd handles SetEOF as truncate, so EOF extension alone does not
allocate backing blocks. A fallocated swapfile can therefore still
look sparse to swapon.

Request allocation for EOF-extending fallocate ranges that can be
represented by FILE_ALLOCATION_INFORMATION, and refresh the allocation
state afterwards.

With this change, xfstests generic/496 and generic/701 pass on ksmbd.

However, Samba "strict allocate = no" now exposes the real generic/701
failure: the old pass came from inflated local i_blocks, not from
server allocation. generic/213 also fails in that case because an
oversized allocation request may not return ENOSPC.

Signed-off-by: Huiwen He <hehuiwen@kylinos.cn>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-13 10:37:55 -05:00
Huiwen He
7a06d3b816 smb/client: emulate small EOF-extending mode 0 fallocate ranges
When a mode 0 fallocate extends EOF from 1G to 2G + 1M, the client
currently sends SetEOF for 2G + 1M. This can make fallocate return
success without allocating the requested range, or allocate extra
space before that range.

For example, on a fresh file:

        xfs_io -f \
          -c "falloc 0 1G" \
          -c "falloc 2G 1M" \
          -c "truncate 3G" test

The second fallocate should allocate [2G, 2G + 1M), leaving [1G, 2G)
as a hole.

Before this change, the result depended on the server allocation policy.
With Samba "strict allocate = no", SetEOF could return success without
allocating [2G, 2G + 1M). With "strict allocate = yes":

	# filefrag -v test
        [0, 1G)             allocated
        [1G, 2G)            allocated unexpectedly
        [2G, 2G + 1M)       allocated

SMB cannot allocate that arbitrary range, so write zeroes to small
EOF-extending ranges instead. Limit this to 1 MiB to bound the
client-side I/O cost.

With "strict allocate = no", the requested range [2G, 2G + 1M) is
allocated by the writes. With "strict allocate = yes":

	# filefrag -v test
        [0, 1G)             allocated
        [1G, 2G)            hole
        [2G, 2G + 1M)       allocated

This fixes the small EOF-extending range case exercised by generic/213.

Signed-off-by: Huiwen He <hehuiwen@kylinos.cn>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-13 10:37:55 -05:00
Huiwen He
9e4ec3be67 smb/client: reduce fallocate zero buffer allocation
The fallocate emulation allocates a 1 MiB zero-filled buffer even
though each SMB2_write request is limited to SMB2_MAX_BUFFER_SIZE,
which is 64 KiB. A high-order 1 MiB allocation is more likely to
fail on a fragmented system.

Allocate only the smaller of the requested range and SMB2_MAX_BUFFER_SIZE,
and reuse that zero-filled buffer for every write request. Also reject
a successful write that makes no progress to avoid looping indefinitely.

This reduces the contiguous allocation required by fallocate emulation
without changing the written data or range semantics.

Signed-off-by: Huiwen He <hehuiwen@kylinos.cn>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-13 10:37:55 -05:00
Huiwen He
b09ae45d85 smb/client: handle overlapping allocated ranges in fallocate
smb3_simple_fallocate_range() can skip holes when an allocated range
returned by the server starts before the current fallocate offset. The
skipped hole is not zero-filled, but fallocate still returns success. A
later write to that hole may therefore fail with ENOSPC.

The function queries allocated ranges so that it can preserve existing
contents and write zeroes only into holes. However, the server may return
a range that starts before the current fallocate offset.

For example, assume the fallocate request is [100, 400) and the only
allocated range returned by the server is [0, 200):

        Request:      [100, 400)
        Server range: [  0, 200)  allocated

        Correct:
        [100, 200)    allocated data, skip
        [200, 400)    hole, zero-fill

        Current:
        [100, 300)    skipped
        [300, 400)    zero-filled afterwards

The current code adds the full server range length, 200, to the current
offset 100 and moves to 300. As a result, the hole in [200, 300) is
skipped without being zero-filled.

Fix this by advancing only over the part of the allocated range that
overlaps the current fallocate offset.  Ignore ranges that end before the
current offset and reject ranges whose end offset overflows.

This also prevents a malformed range length from causing an out-of-bounds
zero-buffer read.

Fixes: 966a3cb7c7 ("cifs: improve fallocate emulation")
Signed-off-by: Huiwen He <hehuiwen@kylinos.cn>
Reviewed-by: ChenXiaoSong <chenxiaosong@kylinos.cn>
Signed-off-by: Steve French <stfrench@microsoft.com>
2026-07-13 10:37:54 -05:00