Commit Graph

1447747 Commits

Author SHA1 Message Date
Paolo Bonzini
b99416dd99 KVM: x86/mmu: change nested_mmu.w to ngva_walk
nested_mmu is now only used for its w member.  While there is still a
single container_of() going from (possibly) gva_walk to its containing
struct kvm_mmu, it is never reached for nested_mmu and therefore it is
safe to strip nested_mmu.w out of its containing struct kvm_mmu.

So do it, and rename it following the model of gva_walk itself.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:30:25 +02:00
Paolo Bonzini
d42e93ee58 KVM: x86/mmu: change walk_mmu to struct kvm_pagewalk
Now that walk_mmu is only accessed for its "w" member, store
directly the pointer to it.  Since it is only used to convert
guest GVAs or nGVAs, call it gva_walk.

Note that there is still one container_of() going from (possibly)
walk_mmu.w to its containing struct kvm_mmu, but for now all instances
of struct kvm_pagewalk do live within a kvm_mmu.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:29:38 +02:00
Paolo Bonzini
9262ed5a03 KVM: x86/mmu: pass struct kvm_pagewalk to kvm_mmu_invalidate_addr
kvm_mmu_invalidate_addr()'s callers only want to tell it whether to
invalidate a GVA or GPA.  This will ultimately be represented by two
different kvm_pagewalk structs, so adjust the type of the parameter.

As of this patch, the GVA case is represented by both root_mmu.w and
nested_mmu.w.  Since nested_mmu never has a sync_spte callback, it would
exit at its check, but really nested_mmu should not be a kvm_mmu in the
first place: it is only used as the walk_mmu, and walk_mmu is only used
for its struct kvm_pagewalk member.

Since calling container_of() on the nested_mmu would be bogus after it is
turned into a struct kvm_pagewalk, introduce a separate check to check
if no work is needed beyond kvm_x86_call(flush_tlb_gva).  Implement it
so that it is as similar as possible to what was happening until now; that
is, do nothing if the invalidation is happening for a nested GVA.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 11:20:38 +02:00
Paolo Bonzini
7c70bf8f93 KVM: x86/mmu: move remaining permission fields to struct kvm_pagewalk
As promised, this removes the remaining instances of
container_of(w, struct kvm_mmu, w), meaning that struct
kvm_pagewalk's definition is pretty much complete.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:50:15 +02:00
Paolo Bonzini
173777d428 KVM: x86/mmu: change CPU-role accessor fields to take struct kvm_pagewalk
With this change, walk_addr_generic and its callees do not need to use
container_of() anymore.  There are only two remaining occurrences of
container_of, in permission_fault() and kvm_mmu_refresh_passthrough_bits().

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:47:49 +02:00
Paolo Bonzini
d869ff104e KVM: x86/mmu: move CPU-related fields to struct kvm_pagewalk
struct kvm_pagewalk's behavior depends on the CPU state and its
page format.  Move related fields so that walk_mmu remains
self contained.

Note that for now, some of the accessors still use kvm_mmu
to split the churn.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:47:44 +02:00
Paolo Bonzini
cb6d93b80c KVM: x86/mmu: move inject_page_fault to struct kvm_pagewalk
Injection of page faults is also part of accesses to guest
page tables; in particular, __kvm_inject_emulated_page_fault()
calls it on walk_mmu.  Move it to struct kvm_pagewalk as
part of converting walk_mmu to a struct kvm_pagewalk.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:46:41 +02:00
Paolo Bonzini
4da3f68249 KVM: x86/mmu: move get_pdptr to struct kvm_pagewalk
Continue with yet another callback used in FNAME(walk_addr_generic),
as another step towards removing container_of() from there.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:46:21 +02:00
Paolo Bonzini
9b91462829 KVM: x86/mmu: move gva_to_gpa to struct kvm_pagewalk
gva_to_gpa is the main entry point into walk_mmu, which
is only used for guest page table walking (as opposed to building
the page tables).  Moving gva_to_gpa to struct kvm_pagewalk
is a step towards making walk_mmu a struct kvm_pagewalk, and
removes several uses of struct kvm_mmu in x86.c.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:45:12 +02:00
Paolo Bonzini
d9af934170 KVM: x86/mmu: move get_guest_pgd to struct kvm_pagewalk
Start moving page walking functionality out of kvm_mmu; the easiest
target is the callbacks.

Change the kvm_mmu_get_guest_pgd() wrapper to take a struct kvm_pagewalk
too, avoiding the MMU indirection (and associated container_of) whenever
the caller already has one.  All container_of uses need to go before
nested_mmu can be changed to a struct kvm_pagewalk, signifying that it
is not used to build page tables.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:42:54 +02:00
Paolo Bonzini
ac889bdd23 KVM: x86/mmu: introduce struct kvm_pagewalk
In preparation for separating walking and building of page tables,
introduce a dummy struct kvm_pagewalk and pass it around instead of
its containing kvm_mmu to functions that do not build the page tables.
Outermost functions retrieve the mmu via container_of, while internal
functions can pass around the struct kvm_pagewalk pointer.

x86.c is still mostly oblivious to the existence of struct kvm_pagewalk,
with are only a couple exceptions for now, but the plan is for it to
use struct kvm_pagewalk whenever dealing with guest page tables and have
only limited knowledge of struct kvm_mmu.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:41:48 +02:00
Paolo Bonzini
9a82ec076e KVM: x86/hyperv: remove unnecessary mmu_is_nested() check
Just always go through kvm_translate_gpa(), which will either invoke
the vendor check or just return hc->ingpa back.

This is a better way to fix the issue of commit 464af6fc2b ("KVM:
x86: check for nEPT/nNPT in slow flush hypercalls", 2026-05-03).

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-25 10:41:36 +02:00
Paolo Bonzini
458dbb64b9 Merge branch 'kvm-spring-clean' into HEAD
It's still technically spring!

Perform spring cleaning on x86.{c,h} and asm/kvm_host.h, by adding regs.c
(the kvm_cache_regs.h => regs.h is already applied) and msrs.{c,h}, and moving
relevant code out of x86.c.

Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:29:26 -04:00
Sean Christopherson
6f8ec95fdb KVM: x86: Move a pile of stuff from kvm_host.h => x86.h
Move the majority of remaining KVM-internal declarations and defines in
kvm_host.h to x86.h, so that kvm_host.h only holds structure and function
definitions that need to be visible to arch-neutral KVM.

Land the emulator interfaces in x86.h, even though kvm_emulate.h *seems*
like a good home, as the interfaces and defines being moved are provided by
x86.c.  I.e. keep kvm_emulate.h as an interface to the emulator proper.

Note, any "misses" are likely unintentional.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260613000329.732085-31-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Paolo Bonzini
35fdfa632e KVM: move TSS constants from kvm_host.h to tss.h
Suggested-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
20fe925246 KVM: x86/mmu: Move kvm_mmu_do_page_fault() from mmu_internal.h => mmu.c
Move kvm_mmu_do_page_fault() into mmu.c, as there are no users outside of
mmu.c, and the function typically isn't inlined by the compiler anyways.
This will allow moving the EMULTYPE_xxx definitions into x86.h without
having to include x86.h in mmu_internal.h, i.e. will help preserve the
goal of making x86.h KVM x86's "top-level" include.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-30-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
31a2cf735c KVM: x86/mmu: Move kvm_arch_async_page_ready() below kvm_tdp_page_fault()
Move the implementation of kvm_arch_async_page_ready() "down" in mmu.c so
that it lives below kvm_tdp_page_fault().  This will allow moving
kvm_mmu_do_page_fault() into mmu.c without needing a forward declaration.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-29-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
5b9585bc5b KVM: x86: Rework kvm_arch_interrupt_allowed() into kvm_is_interrupt_allowed()
Rename kvm_arch_interrupt_allowed() to kvm_is_interrupt_allowed() and
change its return type to a boolean, as the purpose of the helper is purely
to check if interrupts are architecturally allowed.  I.e. whether or not
interrupts are temporarily disallowed due a pending nested VM-Enter is
irrelevant (and callers are most definitely not supposed to care).

Opportunistically bury the helper in x86.c, as it hasn't been referenced by
arch-neutral code since commit a1b37100d9 ("KVM: Reduce runnability
interface with arch support code"), and has long since gained _very_
x86-specific semantics (see above).

Opportunistically add a comment to call out that treating -EBUSY as
"allowed" is intentional.

For all intents and purposes, no functional change intended (KVM treats
-EBUSY as "allowed" before and after).

Cc: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260613000329.732085-28-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
4f1f1ffbdd KVM: x86: Don't treat interrupts as allowed just because a nested run is pending
When querying whether or not interrupts (IRQs) are allowed, check for a
pending nested run _after_ checking whether or not interrupts are blocked.
If L1 is running L2 _without_ nested_exit_on_intr(), i.e. if L1 IRQs can
be blocked while running L2, and interrupts will indeed be blocked once the
nested VM-Enter to L2 is completed, then KVM should treat interrupts as not
being allowed.

For injection, this avoids an unnecessary (forced) VM-Exit, as KVM can
immediately request an IRQ window, instead of forcing an exit and _then_
requesting an IRQ window (because after the forced exit, KVM will see that
interrupts are blocked).

For non-injection usage, only kvm_vcpu_ready_for_interrupt_injection() is
affected in practice.  Barring KVM bugs or misbehaving userspace (at which
point all architectural guarantees are off), kvm_vcpu_has_events() is
unreachable when a nested run is pending.  To reach kvm_vcpu_has_events(),
kvm_vcpu_running() needs to return false, i.e. vcpu->arch.mp_state needs
to be something other than RUNNABLE.  If nested_run_pending is true, then
mp_state *must* be RUNNABLE (again barring bugs or stupid userspace),
because KVM shouldn't emulate VMRUN/VMLAUNCH/VMRESUME while the vCPU is
!RUNNABLE.

The one "near miss" is VMX's GUEST_ACTIVITY_STATE field, which allows L1 to
put the vCPU into HLT or WFS as part of nested VMLAUNCH/VMRESUME.  However,
KVM clears nested_run_pending prior to calling kvm_emulate_halt_noskip()
when putting L2 into HLT via GUEST_ACTIVITY_HLT, and also clears the flag
before setting mp_state to INIT_RECEIVED.  SVM has no equivalent to
GUEST_ACTIVITY_STATE.

I.e. the vCPU will always be runnable if a nested run is pending, and thus
kvm_arch_vcpu_runnable() => kvm_vcpu_has_events() is effectively dead code,
as is __kvm_emulate_halt() => kvm_vcpu_has_events().  Oh, and TDX doesn't
support nested VMX.  Similarly, kvm_can_do_async_pf() is unreachable as
KVM shouldn't be faulting in memory with a pending nested VM-Enter.

As for kvm_vcpu_ready_for_interrupt_injection(), KVM's current behavior of
incorrectly treating interrupts as being allowed could result in KVM
prematurely exiting to userspace to accept an ExtINT.  But, KVM will still
hold/block the ExtINT and request its own IRQ window.  I.e. the net effect
is more or less the same as the for-injection case, the unnecessary exit
just happens at a different boundary.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260613000329.732085-27-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
ee67344af1 KVM: x86: Move kvm_pv_send_ipi() declaration from kvm_host.h => lapic.h
Move the declaration of kvm_pv_send_ipi() into lapic.h, as its
implementation is provided by lapic.c (sending PV IPIs relies on the
optimized APIC map provided by the in-kernel local APIC), and it's only
used by KVM x86 code.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-26-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
0bdd2d6d73 KVM: x86: Move IRQ-related helper declarations from kvm_host.h => irq.h
Move the function declaration for APIs to get/query pending IRQs from
kvm_host.h to irq.h, as the APIs are only used by KVM x86 code.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-25-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
bc61fbab31 KVM: x86: Move __kvm_irq_line_state() from kvm_host.h => ioapic.h
Bury __kvm_irq_line_state() in CONFIG_KVM_IOAPIC=y code, as it's only used
by PIC and I/O APIC code.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-24-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
462474588b KVM: x86: Move misc "VALID MASK" defines from kvm_host.h => x86.c
Move a variety of "VALID MASK" defines, e.g. that capture which flags in
a given ioctl are supported by KVM, from kvm_host.h to x86.c.  The set of
valid flags/bits is very much a KVM-internal detail, as the values from the
hardcoded #defines are often captured and massaged by KVM's setup code, i.e.
*directly* using the macros outside of KVM x86 would be actively dangerous.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-23-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:05 -04:00
Sean Christopherson
781ee15bd9 KVM: x86: Move LLDT assembly wrappers into VMX
Move kvm_{load,read}_ldt() into vmx.c, as vmx_{load,store}_ldt(), as they
are exclusively used by VMX to save/restore host state, and have no
business being globally visible.

Ideally, KVM-specific helpers wouldn't exist at all, as they are nothing
more than assembly wrappers for SLDT and LLDT, i.e. should be provided by
the kernel, not by KVM.  Punt that cleanup to the future, as
arch/x86/include/asm/desc.h _does_ provide helpers, but load_ldt() is only
available for CONFIG_PARAVIRT_XXL=n builds, and both {load,store}_ldt()
unnecessarily constrain the operands to memory.

No functional change intended.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-22-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:04 -04:00
Sean Christopherson
77036f577b KVM: x86: Move MMU helper declarations from kvm_host.h => mmu.h
Move a pile of MMU helper declarations and macros into mmu.h, as they are
very much KVM x86 internal APIs and details, and not intended to be exposed
to arch-neutral KVM, and certainly not to the broader kernel.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-21-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:04 -04:00
Sean Christopherson
95ac6a63d2 KVM: x86/pmu: Move "struct kvm_x86_pmu_event_filter" definition to pmu.c
Move the definition of "struct kvm_x86_pmu_event_filter" to pmu.c, as the
as the details of the filters are very much implementation details that can
and should be buried in pmu.c.  While the _existence_ of filters is public
knowledge, almost by definition, the contents don't need to be exposed
outside of the PMU code as the filter data is provided by userspace, i.e.
it pretty much has to be dynamically allocated, and thus never should be
fully embedded in a globally visible structure.

No functional change intended

Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-20-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:24:03 -04:00
Sean Christopherson
765ae5dc6f KVM: x86: Move "struct kvm_x86_msr_filter" definition to msrs.c
Move the definition of "struct kvm_x86_msr_filter" and its associate,
"struct msr_bitmap_range", to msrs.c, as the details of the filters are
very much implementation details that can and should be buried in msrs.c.
While the _existence_ of filters is public knowledge, almost by definition,
the contents don't need to be exposed outside of the MSR code as the filter
data is provided by userspace, i.e. it pretty much has to be dynamically
allocated, and thus never should be fully embedded in a globally visible
structure.

Note, this creates a discrepancy with the PMU event filter structure; that
will be remedied shortly.

No functional change intended.

Suggested-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-19-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:23:43 -04:00
Sean Christopherson
159a51a038 KVM: x86: Move MSR helper declarations from kvm_host.h => msrs.h
Relocate declarations of MSR helpers (and kvm_nr_uret_msrs) from x86's
kvm_host.h to msrs, to continue trimming down kvm_host.h.

Deliberately leave the funky read_msr() where it is, as it will hopefully
be removed entirely as part of a broader kernel-API cleanup.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-18-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:09 -04:00
Sean Christopherson
dd164f5ce8 KVM: x86: Move kvm_{g,s}et_segment() to inline helpers in regs.h
Define kvm_{g,s}et_segment() as inline functions in regs.h, as they are
literally one-line wrappers to invoke vendor code.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-17-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:09 -04:00
Sean Christopherson
8d4707af57 KVM: x86: Move register helper declarations from kvm_host.h => regs.h
Relocate declarations of Control/Debug Register, EFLAGS and RIP helpers
from x86's kvm_host.h to regs.h, to continue trimming down kvm_host.h.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-16-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:09 -04:00
Sean Christopherson
7a26830801 KVM: x86: Move the bulk of MSR specific code from x86.c to msrs.{c,h}
Introduce msrs.{c,h}, and move the vast majority of MSR specific code out
of x86.{c,h}.  Use a plural "msrs" instead of just "msr" to be consistent
with regs.{c,h}, and to make it easier to differentiate KVM's code from the
other 5+ msr.c files in the kernel.

Opportunistically drop the "x86.h" include from mtrr.c, mostly as proof
that the bulk of the MSR code is indeed being relocated to msrs.c.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-15-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:09 -04:00
Sean Christopherson
9d13832ced KVM: x86: Expose several TSC helpers via x86.h for use by MSR code
Begrudgingly move adjust_tsc_offset_{guest,host}() to x86.h as inlines,
and expose several other TSC helpers in anticipation of moving KVM's MSR
code to a dedicated msrs.c.  Unfortunately for KVM, several MSRs that KVM
emulates can affect TSC state.

Opportunistically drop a superfluous local "tsc_offset" variable, whose
existence causes checkpatch to complain about lack of a blank line.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260613000329.732085-14-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:09 -04:00
Sean Christopherson
57412a311f KVM: x86: Extract get/set MSR (list) ioctl logic to helpers
Extract the code for getting/setting MSRs and MSR lists to dedicated
helpers in anticipation of moving the MSR code to a new msrs.c.

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-13-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:09 -04:00
Sean Christopherson
c392e508ff KVM: x86: Move kvm_{load,put}_guest_fpu() to fpu.h
Move the kvm_{load,put}_guest_fpu() helpers to fpu.h in anticipation of
moving the bulk of KVM's MSR handling code out of x86.c.  KVM needs to
load and put the FPU when accessing MSRs that are managed via XSTATE,
a.k.a. the so called "FPU".

No functional change intended.

Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-12-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:09 -04:00
Sean Christopherson
494c42a409 KVM: x86/hyperv: Eliminate an unnecessary include of x86.h in hyperv.h
Drop an mostly unused include of x86.h from hyperv.h, and instead pull in
regs.h, which is need for at least is_guest_mode().  This eliminates the
last include of x86.h from a common x86 header, i.e. solidifies that x86.h
is the top of the pyramid.

Add a missing x86.h include in cpuid.c to avoid build breakage.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-11-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:09 -04:00
Sean Christopherson
cde542b652 KVM: x86: Move eager_page_split to mmu.{c,h}
Move KVM's eager_page_split module param to the MMU, as it is very much an
MMU knob.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-10-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
af0dbcaaa2 KVM: x86: Move tdp_enabled from kvm_host.h to mmu.h
Relocated the declaration of tdp_enabled into mmu.h, and opportunistically
hoist tdp_mmu_enabled up to the top so that the two are co-located.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-9-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
febae924e8 KVM: x86: Swap the include order between x86.h and mmu.h
Invert the include ordering between x86.h and mmu.h, and move
mmu_is_nested() to mmu.h where it belongs (mmu_is_nested()'s placement in
x86.h was solely responsible for the existing ordering), so that x86.h is
not included by most KVM internal headers (but includes them).

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-8-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
9d2180ac56 KVM: x86: Move kvm_caps and kvm_host_values to asm/kvm_host.h
Relocate the kvm_caps and kvm_host_values struct definitions and their
associated global variable declarations to asm/kvm_host.h to allow for a
variety of cleanups in x86.h and mmu.h, and to establish a (hopefully)
maintainable rule that asm/kvm_host.h's role is to define common
structures (and declare any associated globals), and anything needed by
arch-neutral KVM.

While it would be lovely to trim kvm_host.h down to the point where it
*only* holds things needed by arch-neutral and/or non-KVM code, multiple
attempts to do just that have failed miserably.  Trying to "hide" code
from arch-neutral KVM is too restrictive (and ultimately pointless), and
KVM x86 itself also needs a place to define common structures and their
globals, e.g. to avoid inconsistent header include chains and/or misplaced
helpers.

E.g. as pointed out by Kai, it's weird that x86.h, which is a kitchen sink
of sorts, includes regs.h, but not mmu.h.  Literally the only reason that
x86.h doesn't include mmu.h is that mmu.h references kvm_host, which is
currently defined in x86.h.  As a result of odd include ordering, the
very clearly MMU-specific helper mmu_is_nested() lives in x86.h, not mmu.h

"Fix" the kvm_host dependency so that x86.h can be the "central" include
everyone expects it to be, and set KVM x86 on the path to having somewhat
sensible "rules" for what goes where:

  - asm/kvm_host.h holds "common" structure definitions and associated key
    global variables, and things that are referenced by arch-neutral KVM.
  - <thing>.{c,h} holds relevant declarations and definitions.
  - x86.{c,h} is the kitchen sink for everything else.

Cc: Kai Huang <kai.huang@intel.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260613000329.732085-7-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
c71b025cae KVM: x86: Move local APIC specific helpers out of asm/kvm_host.h
Move local APIC IRQ helpers out of asm/kvm_host.h so that they
are co-located with the structs they use, and exposed to a smaller
portion of the broader world.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-6-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
2f5bb3fe58 KVM: x86: Move the bulk of register specific code from x86.c to regs.c
Introduce regs.c, and move the vast majority of register specific code out
of x86.c and into regs.c.  Deliberately leave behind MSR code, as KVM's MSR
support is complex enough to warrant its own compilation unit, and doesn't
have much in common with the other register code.

Note, "struct kvm_sregs" has fields for EFER and MSR_IA32_APICBASE, and so
the {G,S}ET_REGS flows technically contain a tiny amount of MSR code.
MSR_IA32_APICBASE is already managed by lapic.c, and so doesn't require a
"placement decision".  As for EFER, leave all other EFER handling in x86.c
(later to be moved to msrs.c).  The primary interface to EFER, set_efer(),
is very much MSR specific, even though EFER is arguably more of a Control
Register than an MSR.

No functional change intended.

Reviewed-by: Kai Huang <kai.huang@intel.com>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-5-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
bd130c8d72 KVM: x86: Rename __{g,s}et_sregs2() => kvm_vcpu_ioctl_x86_{g,s}et_sregs2()
Rename the KVM_{G,S}ET_SREGS2 helpers in anticipation of moving them out of
x86.c (while leaving the ioctl dispatch behind).  Having globally visible
APIs named __{g,s}et_sregs2() would be "fine", but ugly, given that
__{g,s}et_sregs() will NOT be globally visible.  As a bonus, this makes it
a bit more obvious that the helpers implement newer versions of
kvm_arch_vcpu_ioctl_set_sregs().

No functional change intended.

Cc: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Message-ID: <20260613000329.732085-4-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
0217de7343 KVM: x86: Move get_segment_base() to regs.h, as kvm_get_segment_base()
Move get_segment_base() to regs.h, as kvm_get_segment_base(), so that the
bulk of the register code can be moved from x86.c to a new regs.c, without
simultaneously needing to rename "public" helpers to explicitly scope them
to KVM.

No functional change intended.

Cc: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
6a8a98aa9c KVM: x86: Extract REGS and SREGS runtime sync code to helpers
Extract the REGS and SREGS portions of {store,sync}_regs() into separate
helpers in anticipation of moving the register specific code out of x86.c
and into regs.c.

No functional change intended.

Cc: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Binbin Wu <binbin.wu@linux.intel.com>
Message-ID: <20260613000329.732085-2-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 08:13:08 -04:00
Sean Christopherson
098e32cba3 x86/apic: KVM: Use cpu_physical_id() to get APIC ID of running vCPU for AVIC
Use cpu_physical_id() instead of default_cpu_present_to_apicid() when
getting the APIC ID of the pCPU on which a vCPU is running/loaded, as the
kernel has gone way off the rails if a vCPU is loaded on a pCPU that has
been physically removed from the system.  Even if the impossible were to
happen, the absolutely worst case scenario is that hardware will ring the
AIVC doorbell on the wrong pCPU, i.e. a severely broken system will
experience mild performance issues.

Kill off KVM's superfluous kvm_cpu_get_apicid() wrapper along with the
for-KVM export of default_cpu_present_to_apicid(), as they existed purely
for the wonky AVIC usage.

Cc: Kai Huang <kai.huang@intel.com>
Cc: Yosry Ahmed <yosry@kernel.org>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Acked-by: Naveen N Rao (AMD) <naveen@kernel.org>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Reviewed-by: Yosry Ahmed <yosry@kernel.org>
Message-ID: <20260612185459.591892-1-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 07:52:24 -04:00
Sean Christopherson
02953418a1 KVM: x86/mmu: Expose number of shadow MMU shadow pages as a stat
Turn arch.n_used_mmu_pages into a stat, mmu_shadow_pages, as the number of
live shadow pages is arguably _the_ most critical datapoint when it comes
to analyzing the shadow MMU.  Before the TDP MMU came along, i.e. when the
shadow MMU was the only MMU, explicitly tracking the number of shadow pages
wasn't as interesting, because the same information could more or less be
gleaned from the pages_{1g,2m,4k} stats.  But with the TDP MMU, where the
shadow MMU is only used for nested TDP, it becomes extremely difficult, if
not impossible, to determine which SPTEs are coming from the TDP MMU, and
which are coming from the shadow MMU.

E.g. when triaging/debugging shadow MMU performance issues due to "too many
shadow pages", being able to observe that 99%+ of all shadow pages are
unsync is critical to being able to deduce that KVM is effectively leaking
shadow pages.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260612133727.411902-1-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 07:52:23 -04:00
Paolo Bonzini
91b16b53a0 KVM: s390: Fix S390_USER_OPEREXEC and more gmap fixes
-----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEoWuZBM6M3lCBSfTnuARItAMU6BMFAmo7nlMACgkQuARItAMU
 6BPr+A//VLw9/5C2pgKVWRvJRVNcf8kIgAF1feONXJVoJiDGrbyiFcy4oCrppIZi
 OxAB2GvqV7Dbyu+tWH/iqPp3gxLf86sNh8JLkba8puWXP5SElyvVpEADYw8T97pX
 dL6j+9vYptMSTIVrhEYgiS5ghgCW6NymM2/6d+uWOCcPjuCaa2owO7K9fCsaLfAK
 O2D6HeRKEQdDszNKdplloSP5FjNn/t/zdPINclhfHdDNl7zSCF5Z3Y9cWGf+L2vL
 fYUiCPRPGANOreBcGMEBtDif667/U7nwiq7daC1rB+Q0MwRoq4h5jk0130aTTgXs
 hOA2EypipNO/ELO1hPJBNYTaSF1XjJMlrq5FHGkXXC+byWTsoKJgFXYkC9pQVd+r
 dcpu7S9nh9jkjSgc8F1+sNQ3+GHc9XrI214ALYgDr1PIZBevSlXaE3dtcU98+qtH
 IfKkonDEeRc8oZUTVLiINkQtZwtHqDKhqdhn5z508xJw5n+qr2oLEe/76NMjwH+E
 jteqA4WFaS2TGnOVf6IsYVdOQFMuOKrl+Se/M7fE1ZQ41IL4+aWT4gnpz3O+8sg9
 XmgiZJHXRtcH5nXfLbGryoO4HZPvwoV2mv0GmbmMVHBkve5qGqdElCRIQLAVImGc
 77/1DL73Kln7XawIik/MFDkCdt7GG3Qbqu1ALbKObCdJRSiFhzk=
 =OiBj
 -----END PGP SIGNATURE-----

Merge tag 'kvm-s390-next-7.2-2' of https://git.kernel.org/pub/scm/linux/kernel/git/kvms390/linux into HEAD

* Fix S390_USER_OPEREXEC so it can now be enabled regardless of other
  unrelated capabilities

* Fix handling of the _PAGE_UNUSED pte bit that could lead to guest
  memory corruption in some scenarios

* A bunch of misc gmap fixes (locking, behaviour under memory pressure)

* Fix CMMA dirty tracking
2026-06-24 13:41:41 +02:00
Carlos López
bb365a506b KVM: x86: Unconditionally recompute CR8 intercept on PPR update
The TPR_THRESHOLD field in the VMCS is used by VMX to induce VM exits
when the guest's virtual TPR falls under the specified threshold,
allowing KVM to inject previously masked interrupts.

KVM handles these VM exits in handle_tpr_below_threshold().
Commit eb90f3417a ("KVM: vmx: speed up TPR below threshold vmexits")
optimized this function by calling apic_update_ppr() instead of raising
KVM_REQ_EVENT. apic_update_ppr() then raises KVM_REQ_EVENT if there is
a pending, deliverable interrupt.

However, if there are no new interrupts pending, apic_update_ppr() does
not issue the request. Thus, kvm_lapic_update_cr8_intercept() and
vmx_update_cr8_intercept() are not called before VM entry, which results
in a high, stale TPR_THRESHOLD. This is problematic due to the following
sentence in 28.2.1.1 "VM-Execution Control Fields" in the SDM:

  The following check is performed if the “use TPR shadow” VM-execution
  control is 1 and the “virtualize APIC accesses” and “virtual-interrupt
  delivery” VM-execution controls are both 0: the value of bits 3:0 of
  the TPR threshold VM-execution control field should not be greater
  than the value of bits 7:4 of VTPR.

This error condition is typically not observed when KVM runs on a bare
metal system because modern processors support APICv, which enables
virtual-interrupt delivery, and which KVM uses when possible. This
causes the processor to no longer generate TPR-below-threshold exits
and to no longer check TPR_THRESHOLD on entry. However, when running
on older platforms, or under nested virtualization on a hypervisor that
does not support virtual-interrupt delivery and enforces this check
(like Hyper-V) this can cause a VM entry failure with hardware error
0x7, as seen in [1].

Call kvm_lapic_update_cr8_intercept() if apic_update_ppr() does not
find a deliverable interrupt (and thus does not raise KVM_REQ_EVENT).
Remove calls to kvm_lapic_update_cr8_intercept() on paths that end up in
apic_update_ppr(), as they now become redundant. This ensures that any
path that updates the guest's PPR also figures out if KVM needs to wait
for a TPR change (using TPR_THRESHOLD on VMX or CR8 intercepts on SVM).

Link: https://github.com/coconut-svsm/svsm/issues/1081 [1]
Tested-by: Stefano Garzarella <sgarzare@redhat.com>
Cc: stable@vger.kernel.org
Fixes: eb90f3417a ("KVM: vmx: speed up TPR below threshold vmexits")
Signed-off-by: Carlos López <clopez@suse.de>
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260618174347.1981064-3-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 11:55:14 +02:00
Sean Christopherson
7ef78d71ca KVM: VMX: Grab vmcs12 on CR8 interception update iff vCPU is in guest mode
When updating CR8 intercepts, get vmcs12 if and only if the vCPU is in
guest mode so that a future change can have update CR8 intercepts during
vCPU creation, without running afoul of get_vmcs12()'s lockdep assertion.

  ------------[ cut here ]------------
  debug_locks && !(lock_is_held(&(&vcpu->mutex)->dep_map) || !refcount_read(&vcpu->kvm->users_count))
  WARNING: arch/x86/kvm/vmx/nested.h:61 at get_vmcs12 arch/x86/kvm/vmx/nested.h:60 [inline], CPU#0: syz.2.19/5879
  WARNING: arch/x86/kvm/vmx/nested.h:61 at vmx_update_cr8_intercept+0x3de/0x4e0 arch/x86/kvm/vmx/vmx.c:6879, CPU#0: syz.2.19/5879
  Modules linked in:
  CPU: 0 UID: 0 PID: 5879 Comm: syz.2.19 Not tainted syzkaller #0 PREEMPT(full)
  Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014
  RIP: 0010:get_vmcs12 arch/x86/kvm/vmx/nested.h:60 [inline]
  RIP: 0010:vmx_update_cr8_intercept+0x3de/0x4e0 arch/x86/kvm/vmx/vmx.c:6879
  Call Trace:
   <TASK>
   apic_update_ppr arch/x86/kvm/lapic.c:984 [inline]
   kvm_lapic_reset+0x1c24/0x2980 arch/x86/kvm/lapic.c:3023
   kvm_vcpu_reset+0x44c/0x1bf0 arch/x86/kvm/x86.c:12986
   kvm_arch_vcpu_create+0x746/0x8b0 arch/x86/kvm/x86.c:12847
   kvm_vm_ioctl_create_vcpu+0x428/0x930 virt/kvm/kvm_main.c:4201
   kvm_vm_ioctl+0x893/0xd50 virt/kvm/kvm_main.c:5159
   vfs_ioctl fs/ioctl.c:51 [inline]
   __do_sys_ioctl fs/ioctl.c:597 [inline]
   __se_sys_ioctl+0xfc/0x170 fs/ioctl.c:583
   do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
   do_syscall_64+0x174/0x580 arch/x86/entry/syscall_64.c:94
   entry_SYSCALL_64_after_hwframe+0x77/0x7f
   </TASK>

No functional change intended.

Reported-by: syzbot ci <syzbot+ci493c6d734b63e050@syzkaller.appspotmail.com>
Closes: https://lore.kernel.org/all/6a2adf3b.3b0a2d4e.8c8d1.0012.GAE@google.com
Cc: stable@vger.kernel.org
Signed-off-by: Sean Christopherson <seanjc@google.com>
Message-ID: <20260618174347.1981064-2-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 11:55:13 +02:00
Sean Christopherson
cc3d0e1afd KVM: x86: WARN (once) if RTC pending EOI tracking goes off the rails
WARN once if KVM's tracking for pending EOIs for Real-Time Clock IRQs goes
off the rails, as there's no reason to bug the host or risk a DoS due to
spamming dmesg with endless WARNs.  Absolute worst case scenario, guest
time will go awry.

Signed-off-by: Sean Christopherson <seanjc@google.com>
Reviewed-by: Kai Huang <kai.huang@intel.com>
Message-ID: <20260618174527.1982333-1-seanjc@google.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
2026-06-24 11:54:29 +02:00