Commit Graph

1447776 Commits

Author SHA1 Message Date
Sean Christopherson
d151ca6e12 KVM: x86/hyperv: Get target FIFO in hv_tlb_flush_enqueue(), not caller
When handling Hyper-V PV TLB flushes, retrieve the to-be-used FIFO in
hv_tlb_flush_enqueue() instead of having the caller pass in the FIFO.  This
will make it easier to fix a cross-vCPU race where KVM can access a vCPU's
FIFO before it's fully initialized.

No functional change intended.

Link: https://patch.msgid.link/20260630225619.511632-2-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:07 -07:00
Sean Christopherson
7a642d8dcf KVM: x86: Read CR4.DE in emulator if and only if accessing DR4 or DR5
Micro-optimize emulation of MOV DR instructions by checking CR4.DE if and
only if DR4 or DR5 is being accessed.

No functional change intended.

Reviewed-by: Jim Mattson <jmattson@google.com>
Link: https://patch.msgid.link/20260612230113.684301-9-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:07 -07:00
Sean Christopherson
b077cfed52 KVM: x86: WARN if MOV DR emulation hits a "too late" #GP
WARN if ->set_dr() => kvm_set_dr() fails when emulating a MOV DR write,
as the emulator _must_ pre-check for #GPs in order to get the event
priority right when emulating MOV DR for L2 on SVM (all exceptions have
higher priority than the instruction intercept).

Opportunistically update the comment as the blurb about "#UD" being
checked is incomplete and misleading.

Reviewed-by: Jim Mattson <jmattson@google.com>
Link: https://patch.msgid.link/20260612230113.684301-8-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:06 -07:00
Sean Christopherson
1f077338d1 KVM: x86: Use kvm_dr{6,7}_valid() to check DR{4,5,6,7} write values in emulator
Use kvm_dr{6,7}_valid() to validate the incoming DR{4,5,6,7} value in the
emulator instead of open coding an equivalent check.  In the unlikely event
that the behavior of DR6/7 (and their aliases) changes in the future, using
common helpers will hopefully make it less likely the emulator logic will
be overlooked.

No functional change intended.

Reviewed-by: Jim Mattson <jmattson@google.com>
Link: https://patch.msgid.link/20260612230113.684301-7-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:05 -07:00
Sean Christopherson
48512697c0 KVM: VMX: Prioritize DR7.GD=1 #DB over CPL>0 #GP on Intel
When emulating a MOV DR on Intel with DR7.GD=1 at CPL>0, prioritize the #DB
due to DR7.GD over the #GP due to CPL>0, as empirical testing shows that
Intel CPUs (Skylake, Icelake and Emerald Rapids) prioritize the DR7.GD #DB
over all #GPs, whereas AMD CPUs prioritize the CPL>0 #GP (but not illegal
value #GPs) over the #DB.

Outside of the emulator, don't bother trying to provide the "correct"
priority based on the virtual CPU model, as it's simply impossible to do
so without intercepting *all* MOV DR accesses, which would result in a
massive, unacceptable performance hit.  Note, getting the priority right
when advertising Intel on AMD would also require intercepting #GP, as SVM
prioritizes all exceptions over the instruction intercept.

Note, neither Intel's SDM nor AMD's APM says anything about the relative
priority, hence the empirical testing.  Arguably Intel's description of
DR7.GD:

  causes a debug exception to be generated prior to any MOV instruction
  that accesses a debug register.

implies that DR7.GD has higher priority.  But that's a fairly weak argument
as the statement would still hold true if the #GP due to CPL>0 had higher
priority, as the #GP would prevent any access to a DR.

Fixes: 3b88e41a41 ("KVM: SVM: Add intercept check for accessing dr registers")
Link: https://patch.msgid.link/20260612230113.684301-6-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:05 -07:00
Sean Christopherson
32a7188a66 KVM: x86: Prioritize #UD on MOV DR over #GP due to non-zero CPL
Manually handle the CPL check for MOV DR instructions instead of using the
Priv flag, *after* checking for #UD scenarios, as #GP due to CPL>0 has
lower priority than all #UDs.

Fixes: 1e470be5a1 ("KVM: x86 emulator: fix mov dr to inject #UD when needed.")
Reviewed-by: Jim Mattson <jmattson@google.com>
Link: https://patch.msgid.link/20260612230113.684301-5-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:04 -07:00
Sean Christopherson
e27ca3dfbb KVM: x86: Manually check DR4/5 write values to fix SVM intercept priority
Manually (pre)check the values being written to DR4/5, i.e. the DR6/DR7
aliases, instead of relying on ->set_dr() => kvm_set_dr() to signal a #GP.
SVM unfortunately prioritizes all exceptions over an instruction intercept,
i.e. nSVM is relying on the emulator to perform *all* exception checks
prior to attempting to execute the instruction.

Fixes: 3b88e41a41 ("KVM: SVM: Add intercept check for accessing dr registers")
Reviewed-by: Jim Mattson <jmattson@google.com>
Link: https://patch.msgid.link/20260612230113.684301-4-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:03 -07:00
Sean Christopherson
55ac576a9b KVM: x86: Prioritize DR7.GD #DB over #GP due to illegal DR6/7 value
When emulating a MOV DR, specifically a write to DR6 or DR7, treat a #DB
due to DR7.GD (General Detect) as higher priority than a #GP due to an
illegal value.  While neither Intel's SDM nor AMD's APM says anything
about the relative priority, empirical testing on Intel and AMD shows that
the #DB has higher priority.  And for VMX, where the instruction intercept
has priority over *all* exceptions, KVM already treats the #DB as having
higher priority.

Cc: Maciej W. Rozycki <macro@orcam.me.uk>
Fixes: 3b88e41a41 ("KVM: SVM: Add intercept check for accessing dr registers")
Reviewed-by: Jim Mattson <jmattson@google.com>
Link: https://patch.msgid.link/20260612230113.684301-3-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:03 -07:00
Carlos López
8f683a4dc5 KVM: x86: Treat any non-zero return from set_dr() as a faulting condition
When emulating a MOV to a debug register, em_dr_write() calls
@ctxt->ops->set_dr(), which is forwarded to emulator_set_dr() and
then kvm_set_dr(). The latter checks that the written value is valid,
otherwise returning an error, in which case the emulator is supposed to
inject a #GP fault into the guest.

Commit 996ff5429e ("KVM: x86: move kvm_inject_gp up from kvm_set_dr
to callers") changed the contract of kvm_set_dr() (and thus
emulator_set_dr()), returning 1 as an error instead of -1, but the
caller in em_dr_write() was never updated, checking only if the returned
value is negative. The end result is that em_dr_write() does not detect
the error, so an invalid write does not generate a #GP, but at the same
time the register value is not updated.

The practical impact is limited, as check_dr_write() already checks DR6
and DR7 manually. However, it misses DR4/DR5, which alias DR6/DR7 when
CR4.DE=0.

Fix the bug by treating any non-zero return from set_dr() as a reason to
inject #GP.

Note, the manual checks on DR6 and DR7 are flawed, as they incorrectly
prioritize the #GP over a DR7.GD=1 #DB (the General Detect #DB has
priority on both Intel and AMD).

Note #2, relying on ->set_dr() to detect #GP is also flawed as all
exceptions have higher priority than the instruction intercept on SVM,
i.e. the manual checks need to be extended to DR4 and DR5 (after the
priority bug is fixed).

Fixes: 996ff5429e ("KVM: x86: move kvm_inject_gp up from kvm_set_dr to callers")
Signed-off-by: Carlos López <clopez@suse.de>
Link: https://patch.msgid.link/20260601133320.91479-2-clopez@suse.de
[sean: drop explicit "!= 0", massage changelog]
Reviewed-by: Jim Mattson <jmattson@google.com>
Link: https://patch.msgid.link/20260612230113.684301-2-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:02 -07:00
Binbin Wu
8835a24a56 KVM: x86: Fix emulated CPUID features being applied to wrong sub-leaf
Pass the CPUID index into cpuid_func_emulated() and return no emulated
features for indexed CPUID leaves with a non-zero index.

KVM currently emulates CPUID features only for index 0, but
kvm_vcpu_after_set_cpuid() looks up emulated features by function alone.
As a result, reverse_cpuid[] entries that share a function but use a
non-zero index, e.g. CPUID.7.1:ECX, can inherit emulated features that
belong to index 0.  For example, RDPID, which is CPUID.7.0:ECX[22], can
be incorrectly OR'd into CPUID.7.1:ECX.

This is benign today because the affected bits do not correspond to
features KVM cares about, but it can become a real bug as new CPUID
features are defined.  Make the helper index-aware so emulated features
are applied only to the CPUID entry they actually describe.

Fixes: e592ec657d ("KVM: x86: Initialize guest cpu_caps based on KVM support")
Suggested-by: Sean Christopherson <seanjc@google.com>
Signed-off-by: Binbin Wu <binbin.wu@linux.intel.com>
Reviewed-by: Xiaoyao Li <xiaoyao.li@intel.com>
Link: https://patch.msgid.link/20260609075748.612704-1-binbin.wu@linux.intel.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:02 -07:00
Sean Christopherson
eb606a2438 KVM: TDX: Return EINVAL, not EOPNOTSUPP, for NULL INIT_MEM_REGION source
Return EINVAL instead of EOPNOTSUPP if userspace attempts to pass a NULL
pointer for the source page of INIT_MEM_REGION, so that KVM's ABI is
consistent between TDX and SNP (for LAUNCH_UPDATE).  EOPNOTSUPP was chosen
to be a forward-looking error code for when guest_memfd supports in-place
conversion, but even when in-place conversion comes along, it's an awkward
error code as KVM is deliberately choosing to disallow virtual address '0',
which is technically a legal userspace address.  I.e. it's not so much a
lack of support as it is that KVM reserves address '0' to simplify KVM's
internal implementation.

Opportunistically move the check so that it's co-located with the other
checks on the userspace address, and so that it's more obvious that a NULL
source address is explicitly disallowed.

Fixes: 2a62345b30 ("KVM: guest_memfd: GUP source pages prior to populating guest memory")
Cc: Yan Zhao <yan.y.zhao@intel.com>
Cc: Ackerley Tng <ackerleytng@google.com>
Reviewed-by: Xiaoyao Li <xiaoyao.li@intel.com>
Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
Reviewed-by: Binbin Wu <binbin.wu@linxu.intel.com>
Reviewed-by: Yan Zhao <yan.y.zhao@intel.com>
Tested-by: Yan Zhao <yan.y.zhao@intel.com>
Reviewed-by: Ackerley Tng <ackerleytng@google.com>
Link: https://patch.msgid.link/20260630213711.479692-3-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:01 -07:00
Joerg Roedel
2abe1ff201 KVM: SEV: Explicitly disallow NULL user address for SNP_LAUNCH_UPDATE
Explicitly reject a NULL userspace virtual address for the source page of
SNP_LAUNCH_UPDATE instead of relying on the post-populate callback to do
the check, and don't WARN on failure, as the scenario is blatantly user-
triggerable, as reported by Sashiko.  Waiting until post-populate to check
the address "works", but makes it unnecessarily difficult to see that KVM's
ABI is to disallow a NULL source page for non-ZERO pages.

Note, several existing VMMs pass a valid userspace address for the ZERO
case, i.e. KVM can't *require* the userspace address to be NULL for ZERO
pages, at least not without breaking userspace.

Fixes: dee5a47cc7 ("KVM: SEV: Add KVM_SEV_SNP_LAUNCH_UPDATE command")
Reported-by: Sashiko Bot <sashiko-bot@kernel.org>
Closes: https://lore.kernel.org/all/20260611125849.9ED631F00893@smtp.kernel.org
Signed-off-by: Joerg Roedel <joerg.roedel@amd.com>
Co-developed-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Ackerley Tng <ackerleytng@google.com>
Link: https://patch.msgid.link/20260630213711.479692-2-seanjc@google.com
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:00 -07:00
Yosry Ahmed
6d00e67326 KVM: nVM: Ensure INVVPID is emulated on the correct physical CPU
When emulating INVVPID, KVM executes INVVPID on the physical CPU using
vpid02 (instead of the L1 assigned VPID), after doing some validations
on the operands. However, it is possible that the physical CPU KVM
executes INVVPID on is different from the CPU L2 is running on.

For example, in the following scenario:
- L2 runs on CPU #1 and exits to L1 (vmx->nested.vmcs02.cpu=1)
- L1 migrates to CPU #2 and executes INVVPID
- KVM executes INVVPID on CPU #2
- L1 migrates back to CPU #1 and runs L2 (vmx->nested.vmcs02.cpu=1)

The TLB entries on CPU #1 are never invalidated, because INVVPID was
executed on CPU #2, and vmcs02 never ran on a different pCPU (i.e.
vmx_vcpu_load_vmcs() will *not* request KVM_REQ_TLB_FLUSH).

Ensure that INVVPID is being executed on the same pCPU that L2 last ran
on, and if not, fallback to clearing last_vpid=0 to trigger a full VPID
flush on the next nested VM-Enter (as KVM will detect L1 using a
different VPID for L2). If L2 ends up running on a different pCPU, KVM
will flush the TLB anyway through vmx_vcpu_load_vmcs().

Cc: stable@vger.kernel.org
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Link: https://patch.msgid.link/20260616214652.2157032-4-yosry@kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:41:00 -07:00
Sean Christopherson
32912404b4 KVM: nVMX: Decouple INVVPID operand checks from flushing of vpid02
Separate the INVVPID operand checks from the actual flushing of vpid02 so
the flushing can be adjusted to do the right thing when vmcs02  was last
loaded on a different pCPU, without having to duplicate the logic across
multiple case-statements.

Opportunistically let the VM-Fail paths poke out past 80 chars.

No functional change intended.

Cc: stable@vger.kernel.org
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Link: https://patch.msgid.link/20260616214652.2157032-3-yosry@kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:40:59 -07:00
Yosry Ahmed
f077238941 KVM: nVMX: Always flush vpid02 on first use
Make sure vpid02 is always flushed on first use by setting last_vpid=0
when allocating vpid02.  nested_vmx_transition_tlb_flush() will always
detect a VPID change on first VM-Enter after VMXON, because VPID=0 in
vmcs12 is not allowed if L1 enables VPID.

This avoids using stale TLB entries from a previous lifetime of the
VPID, that might have been associated with a different vCPU (or a
completely different VM).

Note that last_vpid is already being initialized as 0 when the vCPU is
created, but it is not reset when vpid02 is freed on VMXOFF. Hence, the
problem can only occur if L1 does VMXOFF -> VMXON, runs an L2, and KVM
happens to reuse a VPID that has TLB entries on the physical CPU.

Cc: stable@vger.kernel.org
Signed-off-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Jim Mattson <jmattson@google.com>
Link: https://patch.msgid.link/20260616214652.2157032-2-yosry@kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
2026-07-08 13:40:58 -07:00
Paolo Bonzini
a204badd84 Merge branch 'kvm-chainsaw' into HEAD
The kvm_mmu is a "god data structure" that includes three different
tasks: describing the guest page table's format, walking the guest
page tables and building the page tables.  This means that the
(already poorly named) nested_mmu is only used in part, since it
has no page tables to construct.

Furthermore, some parts are reused across guest and host page
tables (such as the reserved bits detector) but others are not;
for example permission_fault is replaced by simplified code such as
is_executable_pte().

This series cleans this up by splitting kvm_mmu in three parts:

- kvm_pagewalk is the page table walker.  There are two of them
  per vCPU, gva_walk and ngpa_walk.  walk_mmu is *always* replaced
  by a single gva_walk no matter if running an L1 or L2 guest,
  unlike in the current code that moves it between root_mmu and
  nested_mmu.

- kvm_mmu retains the page table building functionality.  It uses
  a page table walker to build shadow pages; that is always gva_walk
  for root_mmu or ngpa_walk for guest_mmu.

- kvm_page_format allows KVM to operate on PTEs that already exist,
  and merges the code around permission_mask() with the pre-existing
  struct rsvd_bits_validate.  Both kvm_pagewalk and kvm_mmu have their
  own kvm_page_format, just like struct kvm_mmu had two instances of
  struct rsvd_bits_validate for gPTE and SPTE reserved bit checks.

The cleanup alone already does something useful, which is to reduce
the confusion between guest_mmu and nested_mmu.  nested_mmu came to
exist long before the introduction of guest_mmu and stole the obvious
name, resulting in comments like "Exempt nested MMUs" where the code
actually exempts guest_mmu.  Renaming guest_mmu could be the next
step, though the RFC had multiple opinions about how to do this.

However, the last patch also shows the code reuse benefits can be used
for new features too.  By adapting the permission_fault() machinery and
using it to test SPTEs against struct kvm_page_fault, it makes it possible
to support SPTEs that have XS!=XU; these were not supported yet by KVM,
but could now be added via memory attributes.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:32:09 +02:00
Paolo Bonzini
216b1a47e8 KVM: x86/mmu: use kvm_page_format to test SPTEs
is_access_allowed(), and is_executable_pte() within it, are effectively a
special version of permission_fault() that only supports a subset of roles.
In particular it does not allow SMEP, SMAP and PKE; while SMAP and PKE are
not a problem (they are not supported by EPT/NPT and is_access_allowed() is
only used for either EPT/NPT or CR0.PG=0), lack of ring-0 execution checks
support means that KVM can only use MBEC and GMET with shadow paging.

Replace is_access_allowed() with a modified version of permission_fault(),
looking at the same bitmasks but using the struct kvm_mmu's fmt member;
the new version supports MBEC/GMET for free, just by virtue of filling
in the ACC_* masks from the SPTE.

This prepares for a possible future where TDP entries could have XS!=XU,
for example as part of implementing Hyper-V VSM natively inside KVM.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:32:00 +02:00
Paolo Bonzini
6bb13c2eb7 KVM: x86/mmu: parameterize update_permission_bitmask()
Make it possible to apply the computation loop to both guest
and shadow PTEs formats; the latter do not have an extended role, so
pass the four parameters to the function one by one.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:31:59 +02:00
Paolo Bonzini
c71239d463 KVM: x86/mmu: merge struct rsvd_bits_validate into struct kvm_page_format
This remove one level of indirection, and adds the data for the permission
bitmask machinery to struct kvm_mmu.  This way, it will be possible to
reuse the permission bitmasks for SPTEs as well.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:31:59 +02:00
Paolo Bonzini
72f734a435 KVM: x86/mmu: pull page format to a new struct
Structs kvm_mmu ("build page tables") and kvm_pagewalk ("walk page
tables") share a piece of common functionality, namely "looking at PTEs".
Both of them have code to check is a specific access (described by
PFERR_* constants) is allowed by a PTE, and both of them also validate
that reserved bits are zero on the respective page tables page tables;
for SPTEs the code is only there to check internal consistency, but in
this case it is indeed shared via struct rsvd_bits_validate.

In preparation for sharing more PTE parsing code between struct
kvm_pagewalk and struct kvm_mmu, create a new struct that contains all
precalculated tables, including the data that is extracted from the
CPU role.  For now only struct kvm_pagewalk uses it.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:31:59 +02:00
Paolo Bonzini
7ccb75f2bf KVM: x86/mmu: cleanup functions that initialize shadow MMU
Now that the GVA->GPA page walker is initialized independently,
init_kvm_softmmu() does not do anything more than calling
kvm_init_shadow_mmu() so eliminate it from the call chain.
At the same time, rename kvm_init_shadow_mmu() to
init_kvm_shadow_mmu() for consistency with init_kvm_tdp_mmu().

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:31:59 +02:00
Paolo Bonzini
148fe965df KVM: x86/mmu: unify root_gva_walk and ngva_walk
At this point, vcpu->arch.ngva_walk and vcpu->arch.root_gva_walk contain
the same information; compare init_kvm_page_walk() on one side with
init_kvm_softmmu() + shadow_mmu_init_context() on the other.  They only
differ in when each is active, and root_gva_walk is also used by shadow
paging, via FNAME(walk_addr) and its callers.

Always use the same instance of kvm_pagewalk to do GVA->GPA translations,
for both guest emulation and shadow paging, instead of flipping the
gva_walk pointer back and forth.  After all the page walking does behave
the same no matter if you are in guest mode or not; the difference lies
in the behavior of kvm_translate_gpa and thus in vcpu->arch.mmu, not in
the page walker itself.

This completes the transition from walk_mmu/nested_mmu as the page
walking entry points to gva_walk/ngpa_walk, and removes duplicated
code between the initialization of root_mmu.w and ngva_walk (The
Struct Formerly Known As nested_mmu).

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:31:59 +02:00
Paolo Bonzini
0a59a8b8a2 KVM: x86/mmu: pull struct kvm_pagewalk out of struct kvm_mmu
Replace kvm_mmu's w field with a pointer to an external instance of
struct kvm_pagewalk.  This is the first step towards using a single
kvm_pagewalk struct for all GVA walks, whether nested or not.

With this patch, non-MMU code basically does not use kvm_mmu anymore:
it does care about page walks, but it funnels (almost) all interactions
with the TLB to mmu.c.

kvm_mmu_invalidate_addr() still needs to go from kvm_pagewalk to kvm_mmu,
and it cannot anymore use container_of() to do it; the mapping is hardcoded
based on the provided kvm_pagewalk, since there are just three of them.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:31:59 +02:00
Paolo Bonzini
b99416dd99 KVM: x86/mmu: change nested_mmu.w to ngva_walk
nested_mmu is now only used for its w member.  While there is still a
single container_of() going from (possibly) gva_walk to its containing
struct kvm_mmu, it is never reached for nested_mmu and therefore it is
safe to strip nested_mmu.w out of its containing struct kvm_mmu.

So do it, and rename it following the model of gva_walk itself.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:30:25 +02:00
Paolo Bonzini
d42e93ee58 KVM: x86/mmu: change walk_mmu to struct kvm_pagewalk
Now that walk_mmu is only accessed for its "w" member, store
directly the pointer to it.  Since it is only used to convert
guest GVAs or nGVAs, call it gva_walk.

Note that there is still one container_of() going from (possibly)
walk_mmu.w to its containing struct kvm_mmu, but for now all instances
of struct kvm_pagewalk do live within a kvm_mmu.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:29:38 +02:00
Paolo Bonzini
9262ed5a03 KVM: x86/mmu: pass struct kvm_pagewalk to kvm_mmu_invalidate_addr
kvm_mmu_invalidate_addr()'s callers only want to tell it whether to
invalidate a GVA or GPA.  This will ultimately be represented by two
different kvm_pagewalk structs, so adjust the type of the parameter.

As of this patch, the GVA case is represented by both root_mmu.w and
nested_mmu.w.  Since nested_mmu never has a sync_spte callback, it would
exit at its check, but really nested_mmu should not be a kvm_mmu in the
first place: it is only used as the walk_mmu, and walk_mmu is only used
for its struct kvm_pagewalk member.

Since calling container_of() on the nested_mmu would be bogus after it is
turned into a struct kvm_pagewalk, introduce a separate check to check
if no work is needed beyond kvm_x86_call(flush_tlb_gva).  Implement it
so that it is as similar as possible to what was happening until now; that
is, do nothing if the invalidation is happening for a nested GVA.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:20:38 +02:00
Paolo Bonzini
7c70bf8f93 KVM: x86/mmu: move remaining permission fields to struct kvm_pagewalk
As promised, this removes the remaining instances of
container_of(w, struct kvm_mmu, w), meaning that struct
kvm_pagewalk's definition is pretty much complete.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:50:15 +02:00
Paolo Bonzini
173777d428 KVM: x86/mmu: change CPU-role accessor fields to take struct kvm_pagewalk
With this change, walk_addr_generic and its callees do not need to use
container_of() anymore.  There are only two remaining occurrences of
container_of, in permission_fault() and kvm_mmu_refresh_passthrough_bits().

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:47:49 +02:00
Paolo Bonzini
d869ff104e KVM: x86/mmu: move CPU-related fields to struct kvm_pagewalk
struct kvm_pagewalk's behavior depends on the CPU state and its
page format.  Move related fields so that walk_mmu remains
self contained.

Note that for now, some of the accessors still use kvm_mmu
to split the churn.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:47:44 +02:00
Paolo Bonzini
cb6d93b80c KVM: x86/mmu: move inject_page_fault to struct kvm_pagewalk
Injection of page faults is also part of accesses to guest
page tables; in particular, __kvm_inject_emulated_page_fault()
calls it on walk_mmu.  Move it to struct kvm_pagewalk as
part of converting walk_mmu to a struct kvm_pagewalk.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:46:41 +02:00
Paolo Bonzini
4da3f68249 KVM: x86/mmu: move get_pdptr to struct kvm_pagewalk
Continue with yet another callback used in FNAME(walk_addr_generic),
as another step towards removing container_of() from there.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:46:21 +02:00
Paolo Bonzini
9b91462829 KVM: x86/mmu: move gva_to_gpa to struct kvm_pagewalk
gva_to_gpa is the main entry point into walk_mmu, which
is only used for guest page table walking (as opposed to building
the page tables).  Moving gva_to_gpa to struct kvm_pagewalk
is a step towards making walk_mmu a struct kvm_pagewalk, and
removes several uses of struct kvm_mmu in x86.c.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:45:12 +02:00
Paolo Bonzini
d9af934170 KVM: x86/mmu: move get_guest_pgd to struct kvm_pagewalk
Start moving page walking functionality out of kvm_mmu; the easiest
target is the callbacks.

Change the kvm_mmu_get_guest_pgd() wrapper to take a struct kvm_pagewalk
too, avoiding the MMU indirection (and associated container_of) whenever
the caller already has one.  All container_of uses need to go before
nested_mmu can be changed to a struct kvm_pagewalk, signifying that it
is not used to build page tables.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:42:54 +02:00
Paolo Bonzini
ac889bdd23 KVM: x86/mmu: introduce struct kvm_pagewalk
In preparation for separating walking and building of page tables,
introduce a dummy struct kvm_pagewalk and pass it around instead of
its containing kvm_mmu to functions that do not build the page tables.
Outermost functions retrieve the mmu via container_of, while internal
functions can pass around the struct kvm_pagewalk pointer.

x86.c is still mostly oblivious to the existence of struct kvm_pagewalk,
with are only a couple exceptions for now, but the plan is for it to
use struct kvm_pagewalk whenever dealing with guest page tables and have
only limited knowledge of struct kvm_mmu.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:41:48 +02:00
Paolo Bonzini
9a82ec076e KVM: x86/hyperv: remove unnecessary mmu_is_nested() check
Just always go through kvm_translate_gpa(), which will either invoke
the vendor check or just return hc->ingpa back.

This is a better way to fix the issue of commit 464af6fc2b ("KVM:
x86: check for nEPT/nNPT in slow flush hypercalls", 2026-05-03).

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:41:36 +02:00
Paolo Bonzini
50406d35f5 KVM selftests for 7.3, early edition
- Automatically allocate a full page for L2 guest stacks on x86 instead of
    requiring test-specific L1 guest code to carve out a portion of the L1
    stack for L2 usage, and to ensure the L2 stack also adheres to the x86-64
    calling convention ABI.
 
  - Add a selftest to verify {Guest,Host}-Only behavior in x86's mediated PMU.
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEKTobbabEP7vbhhN9OlYIJqCjN/0FAmo7/RAACgkQOlYIJqCj
 N/36TQ//U6IBajlUQzQ9ChjNq7WBNS0hBzVbDSRYgoqtowgVyVafQ9YBbpg4VMoP
 WFfU7jptiPNY1+7Kxz2xylIn33YUjdk7qhDy/G0jsRWyLgkxMLFJwqAGbXbG1nSK
 btny8CPwhH5vg3t/a/AEvobtu5m4EWsO8mgGOQphIuh5Dw8GScGxEtcQUuP4Abz9
 fl9IkBfUqUNsw22Ddv/F5yLR/6SuyHezRFc/4NYJWckOXYdpDYxgXruh7JS/GE7a
 Rh6gpw/TLZXJ0fZUMjqy8i/+VCxG64K+65YqH1vsBiZ5scGJWC7KdUA5kM7coBmA
 Ri49LHSP4jROeb/Gu8+v+0JH3Y+Fwl1cRR9qOi8vhc39V8D+vbJyOfO4Hgv2cmOo
 xn3qVRA1c2utLZ304oyL2mSJhYKLNp83XFG/PB3BhF9PVgE7X3H8wGlkEVc3tm0q
 dhaGPztPi862/C4Wk+XDLecwaPRMR9ypqUD2GvyLG50MeXu8ekG9aJweIqR/orij
 qveCAwTiGUtVl4PoUOFcSdRoV5jDVBXGJr8/KRoYNgH+AWNzFslYSZi18F/zlLux
 4JZr21w0qR8FWv//6EKfpdkH/g7R+8jhlQ4nuFBid/WYl58r6NFJ++viGCPfeqom
 d5f4kZ7buQv38ldR5ORuuSmqDEYMM/8PDMIR6Kmke9WiAJqY6DM=
 =AYtV
 -----END PGP SIGNATURE-----

Merge tag 'kvm-x86-selftests_l2_stacks-7.3' of https://github.com/kvm-x86/linux into HEAD

KVM selftests for 7.3, early edition

 - Automatically allocate a full page for L2 guest stacks on x86 instead of
   requiring test-specific L1 guest code to carve out a portion of the L1
   stack for L2 usage, and to ensure the L2 stack also adheres to the x86-64
   calling convention ABI.

 - Add a selftest to verify {Guest,Host}-Only behavior in x86's mediated PMU.
2026-06-24 12:01:00 -04:00
Paolo Bonzini
458dbb64b9 Merge branch 'kvm-spring-clean' into HEAD
It's still technically spring!

Perform spring cleaning on x86.{c,h} and asm/kvm_host.h, by adding regs.c
(the kvm_cache_regs.h => regs.h is already applied) and msrs.{c,h}, and moving
relevant code out of x86.c.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:29:26 -04:00
Sean Christopherson
6f8ec95fdb KVM: x86: Move a pile of stuff from kvm_host.h => x86.h
Move the majority of remaining KVM-internal declarations and defines in
kvm_host.h to x86.h, so that kvm_host.h only holds structure and function
definitions that need to be visible to arch-neutral KVM.

Land the emulator interfaces in x86.h, even though kvm_emulate.h *seems*
like a good home, as the interfaces and defines being moved are provided by
x86.c.  I.e. keep kvm_emulate.h as an interface to the emulator proper.

Note, any "misses" are likely unintentional.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260613000329.732085-31-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Paolo Bonzini
35fdfa632e KVM: move TSS constants from kvm_host.h to tss.h
Suggested-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
20fe925246 KVM: x86/mmu: Move kvm_mmu_do_page_fault() from mmu_internal.h => mmu.c
Move kvm_mmu_do_page_fault() into mmu.c, as there are no users outside of
mmu.c, and the function typically isn't inlined by the compiler anyways.
This will allow moving the EMULTYPE_xxx definitions into x86.h without
having to include x86.h in mmu_internal.h, i.e. will help preserve the
goal of making x86.h KVM x86's "top-level" include.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-30-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
31a2cf735c KVM: x86/mmu: Move kvm_arch_async_page_ready() below kvm_tdp_page_fault()
Move the implementation of kvm_arch_async_page_ready() "down" in mmu.c so
that it lives below kvm_tdp_page_fault().  This will allow moving
kvm_mmu_do_page_fault() into mmu.c without needing a forward declaration.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-29-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
5b9585bc5b KVM: x86: Rework kvm_arch_interrupt_allowed() into kvm_is_interrupt_allowed()
Rename kvm_arch_interrupt_allowed() to kvm_is_interrupt_allowed() and
change its return type to a boolean, as the purpose of the helper is purely
to check if interrupts are architecturally allowed.  I.e. whether or not
interrupts are temporarily disallowed due a pending nested VM-Enter is
irrelevant (and callers are most definitely not supposed to care).

Opportunistically bury the helper in x86.c, as it hasn't been referenced by
arch-neutral code since commit a1b37100d9 ("KVM: Reduce runnability
interface with arch support code"), and has long since gained _very_
x86-specific semantics (see above).

Opportunistically add a comment to call out that treating -EBUSY as
"allowed" is intentional.

For all intents and purposes, no functional change intended (KVM treats
-EBUSY as "allowed" before and after).

Cc: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260613000329.732085-28-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
4f1f1ffbdd KVM: x86: Don't treat interrupts as allowed just because a nested run is pending
When querying whether or not interrupts (IRQs) are allowed, check for a
pending nested run _after_ checking whether or not interrupts are blocked.
If L1 is running L2 _without_ nested_exit_on_intr(), i.e. if L1 IRQs can
be blocked while running L2, and interrupts will indeed be blocked once the
nested VM-Enter to L2 is completed, then KVM should treat interrupts as not
being allowed.

For injection, this avoids an unnecessary (forced) VM-Exit, as KVM can
immediately request an IRQ window, instead of forcing an exit and _then_
requesting an IRQ window (because after the forced exit, KVM will see that
interrupts are blocked).

For non-injection usage, only kvm_vcpu_ready_for_interrupt_injection() is
affected in practice.  Barring KVM bugs or misbehaving userspace (at which
point all architectural guarantees are off), kvm_vcpu_has_events() is
unreachable when a nested run is pending.  To reach kvm_vcpu_has_events(),
kvm_vcpu_running() needs to return false, i.e. vcpu->arch.mp_state needs
to be something other than RUNNABLE.  If nested_run_pending is true, then
mp_state *must* be RUNNABLE (again barring bugs or stupid userspace),
because KVM shouldn't emulate VMRUN/VMLAUNCH/VMRESUME while the vCPU is
!RUNNABLE.

The one "near miss" is VMX's GUEST_ACTIVITY_STATE field, which allows L1 to
put the vCPU into HLT or WFS as part of nested VMLAUNCH/VMRESUME.  However,
KVM clears nested_run_pending prior to calling kvm_emulate_halt_noskip()
when putting L2 into HLT via GUEST_ACTIVITY_HLT, and also clears the flag
before setting mp_state to INIT_RECEIVED.  SVM has no equivalent to
GUEST_ACTIVITY_STATE.

I.e. the vCPU will always be runnable if a nested run is pending, and thus
kvm_arch_vcpu_runnable() => kvm_vcpu_has_events() is effectively dead code,
as is __kvm_emulate_halt() => kvm_vcpu_has_events().  Oh, and TDX doesn't
support nested VMX.  Similarly, kvm_can_do_async_pf() is unreachable as
KVM shouldn't be faulting in memory with a pending nested VM-Enter.

As for kvm_vcpu_ready_for_interrupt_injection(), KVM's current behavior of
incorrectly treating interrupts as being allowed could result in KVM
prematurely exiting to userspace to accept an ExtINT.  But, KVM will still
hold/block the ExtINT and request its own IRQ window.  I.e. the net effect
is more or less the same as the for-injection case, the unnecessary exit
just happens at a different boundary.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260613000329.732085-27-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
ee67344af1 KVM: x86: Move kvm_pv_send_ipi() declaration from kvm_host.h => lapic.h
Move the declaration of kvm_pv_send_ipi() into lapic.h, as its
implementation is provided by lapic.c (sending PV IPIs relies on the
optimized APIC map provided by the in-kernel local APIC), and it's only
used by KVM x86 code.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-26-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
0bdd2d6d73 KVM: x86: Move IRQ-related helper declarations from kvm_host.h => irq.h
Move the function declaration for APIs to get/query pending IRQs from
kvm_host.h to irq.h, as the APIs are only used by KVM x86 code.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-25-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
bc61fbab31 KVM: x86: Move __kvm_irq_line_state() from kvm_host.h => ioapic.h
Bury __kvm_irq_line_state() in CONFIG_KVM_IOAPIC=y code, as it's only used
by PIC and I/O APIC code.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-24-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
462474588b KVM: x86: Move misc "VALID MASK" defines from kvm_host.h => x86.c
Move a variety of "VALID MASK" defines, e.g. that capture which flags in
a given ioctl are supported by KVM, from kvm_host.h to x86.c.  The set of
valid flags/bits is very much a KVM-internal detail, as the values from the
hardcoded #defines are often captured and massaged by KVM's setup code, i.e.
*directly* using the macros outside of KVM x86 would be actively dangerous.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-23-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
781ee15bd9 KVM: x86: Move LLDT assembly wrappers into VMX
Move kvm_{load,read}_ldt() into vmx.c, as vmx_{load,store}_ldt(), as they
are exclusively used by VMX to save/restore host state, and have no
business being globally visible.

Ideally, KVM-specific helpers wouldn't exist at all, as they are nothing
more than assembly wrappers for SLDT and LLDT, i.e. should be provided by
the kernel, not by KVM.  Punt that cleanup to the future, as
arch/x86/include/asm/desc.h _does_ provide helpers, but load_ldt() is only
available for CONFIG_PARAVIRT_XXL=n builds, and both {load,store}_ldt()
unnecessarily constrain the operands to memory.

No functional change intended.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-22-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:04 -04:00
Sean Christopherson
77036f577b KVM: x86: Move MMU helper declarations from kvm_host.h => mmu.h
Move a pile of MMU helper declarations and macros into mmu.h, as they are
very much KVM x86 internal APIs and details, and not intended to be exposed
to arch-neutral KVM, and certainly not to the broader kernel.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-21-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:04 -04:00
Sean Christopherson
95ac6a63d2 KVM: x86/pmu: Move "struct kvm_x86_pmu_event_filter" definition to pmu.c
Move the definition of "struct kvm_x86_pmu_event_filter" to pmu.c, as the
as the details of the filters are very much implementation details that can
and should be buried in pmu.c.  While the _existence_ of filters is public
knowledge, almost by definition, the contents don't need to be exposed
outside of the PMU code as the filter data is provided by userspace, i.e.
it pretty much has to be dynamically allocated, and thus never should be
fully embedded in a globally visible structure.

No functional change intended

Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-20-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:03 -04:00