- When MSI is enabled but iMSI-RX is not used, configure AXIINTC to allow
GIT ITS to handle MSI (Marek Vasut)
- Refactor GIC600 implementation to make it easier to add platforms that
only support 32-bit addressing (Marek Vasut)
- Add Renesas R-Car Gen4 S4/V4H/V4M to the list of GIC600 integrations that
only support 32-bit addressing (Marek Vasut)
* pci/controller/dwc-rcar-gen4:
irqchip/gic-v3: Add Renesas R-Car Gen4 erratum workaround
irqchip/gic-v3: Refactor GIC600 limited to 32bit PA erratum handling
PCI: rcar-gen4: Configure AXIINTC if iMSI-RX is not used
PCI: dwc: Move iMSI-RX check before calling 'pp->ops->init()'
- Add DT binding and driver support for Hawi SoC (Matthew Leung)
- Skip PERST# GPIOs provided by downstream PCIe devices, which should be
handled by drivers of those devices (Manivannan Sadhasivam)
- Stop advertising Attention Button Present (no Qcom SoCs support Attention
Buttons) so pciehp can use Presence Detect Changed events (Qiang Yu)
* pci/controller/dwc-qcom:
PCI: qcom: Clear Attention Button Present in Slot Capabilities
PCI: qcom: Rename qcom_pcie_set_slot_nccs() to qcom_pcie_set_slot_cap()
PCI: qcom: Skip PERST# GPIOs provided by downstream PCIe devices
PCI: qcom: Add support for Hawi
dt-bindings: PCI: qcom: Document Hawi and Maili PCIe Controllers
- Correct the PERST# GPIO state so it remains asserted until power and
REFCLK become stable to fix enumeration failure (Ronald Claveau)
* pci/controller/dwc-meson:
PCI: meson: Fix GPIO state while requesting PERST#
- Remove PERST# checking from pci_host_common_parse_port() so callers can
decide whether to fall back to legacy DT binding with PERST# in the host
bridge (Sherry Sun)
- Fix build issues when PCI_PWRCTRL_GENERIC or PCI_HOST_COMMON is a module
(Arnd Bergmann)
- Create pwrctrl devices only once by doing it from imx_pcie_probe()
instead of imx_pcie_host_init(), which is used during both probe and
resume (Sherry Sun)
- Use 'dw_pcie_rp->skip_pwrctrl_off' to avoid powering off devices during
suspend to preserve wakeup capability (Sherry Sun)
- Add runtime PM support for i.MX95 to allow dynamic power management when
the link is idle (Richard Zhu)
* pci/controller/dwc-imx6:
PCI: imx6: Add runtime PM support for i.MX95
PCI: imx6: Add 'skip_pwrctrl_off' flag support
PCI: imx6: Move pci_pwrctrl_create_devices() to imx_pcie_probe()
PCI: imx6: Fix building against PCI_PWRCTRL_GENERIC
PCI: imx6: Fix building against PCI_HOST_COMMON
PCI: host-generic: Move legacy DT binding fallback decision to caller of pci_host_common_parse_ports()
- Factor pcie_valid_speed() and pci_bus_speed2lnkctl2() out of bwctrl so
they can be shared by the DWC core (Hans Zhang)
- Flush MSI writes from endpoint before unmapping the iATU, as we already
do for MSI-X writes (Niklas Cassel)
- Unmap MSI iATU window before mapping MSI-X window, to avoid a subsequent
MSI write using a disabled aperture and losing the interrupt (Niklas
Cassel)
- Change endpoint .pre_init() and .init() callbacks to return errors and
handle them (Marek Vasut)
* pci/controller/dwc:
PCI: dwc: Handle return value from endpoint .pre_init callback
PCI: dwc: Handle return value from endpoint .init callback
PCI: dwc: ep: Fix unmap potentially unmapping the wrong iATU
PCI: dwc: ep: Flush cached MSI write before unmapping the iATU
PCI: dwc: Use common speed conversion function
PCI: Move pci_bus_speed2lnkctl2() to public header
PCI: Add public pcie_valid_speed() for shared validation
- Add missing MODULE_DEVICE_TABLE to generate module aliases for OF-based
module autoloading (Pengpeng Hou)
- Add debugfs 'ltssm_status' file for LGA- and HPA-based Cadence
controllers (Hans Zhang)
- Support up to x4 (not x2) lanes for J200 (Takuma Fujiwara)
- Fix host/endpoint dependencies for cadence-plat driver to fix link error
when cadence-plat is built-in but the host or endpoint driver is modular
(Aksh Garg)
* pci/controller/cadence:
PCI: cadence: Fix host/endpoint dependencies for cadence-plat driver
PCI: j721e: Fix incorrect max_lanes for J7200
PCI: cadence: Add LGA IP debugfs for LTSSM status
PCI: cadence: Add HPA IP debugfs for LTSSM status
PCI: cadence: Add HPA architecture flag
PCI: cadence: Add missing MODULE_DEVICE_TABLE()
- Switch to irq_domain_create_linear() so we can obsolete
irq_domain_add_linear() (Jiri Slaby)
* pci/controller/aspeed:
PCI: aspeed: Switch to irq_domain_create_linear()
* pci/controller/root-port-reset:
misc: pci_endpoint_test: Add AER error handlers
PCI: dw-rockchip: Implement .reset_root_port() and use for link down
PCI: qcom: Implement .reset_root_port() and use for link down
PCI: host-common: Add link down handling for Root Ports
PCI/ERR: Add support for resetting the Root Ports in a platform-specific way
PCI: dwc: ep: Clear MSI iATU mapping in dw_pcie_ep_cleanup()
- Check doorbell SUCCESS bit in pci_endpoint_test to avoid treating some
failures as successes (Niklas Cassel)
- Fail doorbell test when the trigger IRQ is missed (Niklas Cassel)
* pci/endpoint:
misc: pci_endpoint_test: Fail doorbell test when the trigger IRQ is missed
misc: pci_endpoint_test: Check SUCCESS bit for doorbell status
- Add ACS quirk for Pericom PI7C9X2G608 switches (Tim Harvey)
- Fix a long-standing bug in the Intel PCH Root Port MPC ACS quirk that
didn't update the intended INTEL_MPC_REG_IRBNCE bit because it used a
16-bit config write when a 32-bit write was intended (Mohamad Raizudeen)
* pci/virtualization:
PCI: Fix 32-bit config write in Intel PCH Root Port MPC ACS quirk
PCI: Add ACS quirk for Pericom PI7C9X2G608 switches [12d8:2608]
- In pci_write_legacy_io(), avoid out-of-bounds reads from the user buffer
and fix incorrect ioport write data (1-byte writes on little-endian
powerpc, 2- and 4-byte writes on big-endian powerpc) (Krzysztof
Wilczyński)
- In pci_read_legacy_io(), fix incorrect ioport read data for 2- and 4-byte
reads on big-endian powerpc (Krzysztof Wilczyński)
- Fix I/O port accessor argument order in Alpha pci_legacy_write()
(Krzysztof Wilczyński)
- Avoid spurious runtime PM wakeup on config space accesses that are
outside config space and fail before reaching PCI (Krzysztof Wilczyński)
- Return -EINVAL, not -ENODEV, for mmap of I/O BAR that fails because the
arch doesn't support it, as we do for procfs (Krzysztof Wilczyński)
- Check for LOCKDOWN_PCI_ACCESS for legacy_io and legacy_mem, as we do for
other config space accessors (Krzysztof Wilczyński)
- Tidy static resource attribute names and macros that build them
(Krzysztof Wilczyński)
* pci/sysfs:
PCI/sysfs: Add legacy I/O and memory attribute macros
alpha/PCI: Make the suffix the first __pci_dev_resource_attr() parameter
PCI/sysfs: Add pci_ prefix to static PCI resource attribute names
PCI/sysfs: Add lockdown checks to legacy I/O and memory handlers
PCI/sysfs: Return -EINVAL for unsupported I/O BAR mmap
PCI/sysfs: Avoid spurious runtime PM wakeup on config space accesses
alpha/PCI: Fix I/O port accessor argument order in pci_legacy_write()
PCI/sysfs: Fix read byte order in pci_read_legacy_io()
PCI/sysfs: Fix out-of-bounds read in pci_write_legacy_io()
- Add Microchip PCI1008 device ID and include it in NTB DMA alias quirk
(Logan Gunthorpe)
* pci/switchtec:
PCI: switchtec: Add Microchip PCI1008 to NTB DMA alias quirk
PCI: switchtec: Add Microchip PCI1008 device ID
dmaengine: switchtec-dma: Add PCI1008 device ID
- Add hotplug reservation only once (not at each level of the hierarchy) so
bridge windows don't grow more than necessary (Ilpo Järvinen)
* pci/resource:
PCI: Do not add hotplug reservation multiple times
- Take a reference on the I2C adapter in tc9563 to avoid uninterruptible
hang when unloading an I2C module while in-use (Johan Hovold)
- Restrict tc9563 Tx Amplitude, DFE and N_FTS to USP, DSP1 and DSP2 in DT
binding (Manivannan Sadhasivam)
- Fix parsing tc9563 integrated Ethernet MAC Endpoint node (Manivannan
Sadhasivam)
- Power off only tc9563 external-facing ports (DSP1, DSP2), leaving USP and
DSP3 powered up (Manivannan Sadhasivam)
- Skip tc9563 Tx amplitude and DFE tuning for DSP3, which don't support
them (Manivannan Sadhasivam)
- Move integrated MAC Endpoint out of the list of internal ports and
configure it separately (Manivannan Sadhasivam)
* pci/pwrctrl:
PCI/pwrctrl: tc9563: Move Integrated MAC Endpoint out of 'tc9563_pwrctrl_ports' enum
PCI/pwrctrl: tc9563: Rename DSP3 to VDSP
PCI/pwrctrl: tc9563: Skip Tx amplitude and DFE tuning for DSP3
PCI/pwrctrl: tc9563: Power off only the external ports in tc9563_pwrctrl_disable_port()
PCI/pwrctrl: tc9563: Fix parsing the integrated Ethernet MAC Endpoint node
dt-bindings: PCI: toshiba,tc9563: Restrict Tx Amplitude, DFE and N_FTS to USP, DSP1 and DSP2
PCI/pwrctrl: tc9563: Take i2c adapter module reference
- Avoid spurious runtime PM wakeup on config space accesses that are
outside config space and fail before reaching PCI (Krzysztof Wilczyński)
- Warn on user-space writes to kernel-exclusive config space regions, as we
already do for sysfs (Krzysztof Wilczyński)
- Check credentials of opener, not reader, for config space reads, as we
already to for sysfs (Krzysztof Wilczyński)
* pci/procfs:
PCI/proc: Use file_ns_capable() when checking config space read access
PCI/proc: Warn on writes to kernel-exclusive config space regions
PCI/proc: Avoid spurious runtime PM wakeup on config space accesses
- Allow probing even without child services so it can do power management
(Brian Norris)
* pci/portdrv:
PCI/portdrv: Allow probing even without child services
- Allow D3 for native hotplug-capable Root Ports on non-x86 platforms (we
avoid D3 for these ports on x86 because some old platforms didn't
validate it) (Manivannan Sadhasivam)
* pci/pm:
PCI: Allow D3 for native hotplug-capable Root Ports on non-x86 platforms
- Pass empty string, not an uninitialized device_class string, to
acpi_bus_generate_netlink_event(), so we can remove device_class
completely in the future (Rafael J. Wysocki)
* pci/hotplug:
PCI: acpiphp_ibm: Do not use uninitialized device_class
- Don't store pci_device_id in agp amd-k7 and via, ata, scsi nsp32, ipack
tpci200, mlxsw since the dynamic ID feature means the ID is only
guaranteed to live during probe (Gary Guo)
- Add pci_match_one_id() to match an ID directly so dynamic ID insertion
doesn't need to make a temporary device for matching (Gary Guo)
- Check for existing ID inside the dynamic ID addition critical section to
avoid a time-of-check vs time-of-use race (Gary Guo)
- Copy device ID to avoid use-after-free when match races with sysfs
dynamic ID removal (Gary Guo)
* pci/enumeration:
PCI: Fix UAF when probe runs concurrent to dyn ID removal
PCI: Fix dyn_id add TOCTOU
PCI: Make pci_match_one_device() match on ID instead of device
agp/amd-k7: Don't rely on address of pci_device_id
agp/via: Don't rely on address of pci_device_id
mlxsw: pci: Don't store pci_device_id
ipack: tpci200: Don't store pci_device_id
scsi: nsp32: Don't store pci_device_id
ata: ata_generic: Don't store pci_device_id
- Allow DPC on all Downstream Ports, not just Root Ports, when OS controls
AER (Darshit Shah)
* pci/dpc:
PCI/DPC: Allow DPC on all Downstream Ports when OS controls AER
- Avoid L0s for Realtek RTS525A, where it causes an AER interrupt storm
(Max Lee)
- Program the same ASPM Control values for every function of multi-function
devices, as recommended by the PCIe spec (Krishna Chaitanya Chundru)
- Avoid ASPM L0s, L1, and L1 PM Substates based on 'aspm-no-l0s',
'aspm-no-l1' [1], and 'aspm-no-l1ss' DT properties (Krishna Chaitanya
Chundru)
* pci/aspm:
PCI/ASPM: Mask ASPM states based on Devicetree properties
PCI/ASPM: Disable/restore ASPM on every function for multi-function devices
PCI/ASPM: Use pcie_capability_clear_and_set_word() for ASPM disable/restore
PCI/ASPM: Avoid L0s for Realtek RTS525A
- Update mappings of errors to agent & layer (Lukas Wunner)
- Log agent & layer of each individual error when multiple errors detected
(Lukas Wunner)
- Log Error Source only once, not twice in separate messages (Lukas Wunner)
- Emit TLP Log only for unmasked errors (Lukas Wunner)
- Support Advisory Non-Fatal Errors (Lukas Wunner)
* pci/aer:
PCI/AER: Support Advisory Non-Fatal Errors
PCI/AER: Move retrieval of FEP and TLP Log into helper
PCI/AER: Emit TLP Log only for unmasked errors
PCI/AER: Deduplicate logging of Error Source Identification
PCI/AER: Log agent & layer for each individual error
PCI/AER: Fix mapping of errors to agent & layer
Per PCIe r7.0 sec 6.2.4.3, certain Non-Fatal Errors may be signaled using
ERR_COR instead of ERR_NONFATAL. These "Advisory Non-Fatal Errors" are
listed in sec 6.2.7 and explained in detail in sec 6.2.3.2.4.
Advisory Non-Fatal Errors set bits in the Uncorrectable Error Status
Register as well as one bit in the Correctable Error Status Register
(Advisory Non-Fatal Error Status, bit 13). The latter is masked by
default, hence these errors are currently not signaled at all (except on
non-compliant products which choose to unmask the bit).
Unmask Advisory Non-Fatal Errors on device enumeration.
Some Non-Fatal Errors are always Advisory, others may be Advisory at the
discretion of the detecting agent. If multiple errors occur, the agent
may qualify a portion as non-Advisory and signal ERR_NONFATAL in addition
to ERR_COR. In this case, there's no way to determine which Non-Fatal
Error was Advisory. Assume none is to ensure that the Uncorrectable Error
code path is taken to recover from the errors.
Introduce aer_compute_anfe_status() to compute Advisory Non-Fatal Error
bits from AER registers, based on this policy. Use it for Firmware First
error handling in pci_print_aer(), which receives an AER register dump
from the platform (UEFI r2.11 sec N.2.7).
Introduce aer_get_anfe_status() to read AER registers from a device and
feed them to aer_compute_anfe_status(). Use it for native error handling
in aer_get_device_error_info(), which gathers registers from the device
and caches the computed Advisory Non-Fatal Error bits in a new anfe_status
field in struct aer_err_info.
Regardless whether error handling is native or Firmware First, the AER
driver needs to increment error counters, signal a trace event and log
each error. When Advisory Non-Fatal Errors occur, these steps must be
performed for Correctable Errors and for Uncorrectable Errors. Achieve
this through a recursive invocation of aer_print_error() (for native error
handling) and pci_print_aer() (for Firmware First error handling). The
recursive invocation reports the (Advisory) Uncorrectable Errors after
reporting the Correctable Errors.
Note that the First Error Pointer and TLP Prefix Log is only meaningful
for Uncorrectable Errors, but when Advisory Non-Fatal Errors occur,
aer_get_device_error_info() has to populate the first_error and
tlp_header_valid fields in struct aer_err_info for a Correctable Error.
Avoid incorrectly logging those fields for Correctable Errors by amending
__aer_print_error() and aer_print_error() with conditionals.
Sample log output for an Advisory Unsupported Request Error:
pcieport 0001:00:00.4: AER: Multiple Correctable Error messages received, first one from 0001:0e:00.0
idxd 0001:0e:00.0: PCIe Bus Error: severity=Correctable
idxd 0001:0e:00.0: device [8086:1216] error status/mask=00002000/00000000
idxd 0001:0e:00.0: [13] NonFatalErr | |
idxd 0001:0e:00.0: PCIe Bus Error: severity=Uncorrectable (Non-Fatal)
idxd 0001:0e:00.0: device [8086:1216] error status/mask=00100000/00000000
idxd 0001:0e:00.0: [20] UnsupReq | Receiver | Transaction Layer (First)
idxd 0001:0e:00.0: AER: TLP Header (Flit): 0x01000104 0x00000000 0x0000080e 0x0f800001
This commit takes inspiration (but differs significantly) from an earlier
submission by Zhenzhong Duan, which in turn was based on a submission by
Qingshun Wang:
https://lore.kernel.org/r/20240620025857.206647-1-zhenzhong.duan@intel.com/
Prior attempts at supporting Advisory Non-Fatal Errors were submitted by
Yicong Yang and Dio Sun:
https://lore.kernel.org/r/1614689994-10925-1-git-send-email-yangyicong@hisilicon.com/https://lore.kernel.org/r/BJXPR01MB0614C01A9523786117B1F1CBCEC8A@BJXPR01MB0614.CHNPR01.prod.partner.outlook.cn/
Signed-off-by: Lukas Wunner <lukas@wunner.de>
[bhelgaas: fold in https://lore.kernel.org/all/amdnMg_J6T3Sys45@wunner.de,
https://lore.kernel.org/all/120da0565eac0157ffd913423c7cfa66e985ff59.1786800931.git.lukas@wunner.de]
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://lore.kernel.org/r/20240620025857.206647-1-zhenzhong.duan@intel.com/
Link: https://lore.kernel.org/r/1614689994-10925-1-git-send-email-yangyicong@hisilicon.com/
Link: https://lore.kernel.org/r/BJXPR01MB0614C01A9523786117B1F1CBCEC8A@BJXPR01MB0614.CHNPR01.prod.partner.outlook.cn/
Link: https://patch.msgid.link/1b62915ffe06ee5b08e846531c42392e5f244337.1784905909.git.lukas@wunner.de
pci_quirk_enable_intel_rp_mpc_acs() reads a 32-bit DWORD from the MPC
register, sets bit 26 (INTEL_MPC_REG_IRBNCE), but it writes it back using
pci_write_config_word().
Because bit 26 resides in the upper 16 bits of the 32-bit register, a
16-bit write drops the newly set bit. The quirk logs that it is enabling
IRBNCE, but the hardware never actually receives the command.
Use pci_write_config_dword() to ensure the full 32-bit value is written
back to the hardware.
Fixes: d99321b63b ("PCI: Enable quirks for PCIe ACS on Intel PCH root ports")
Signed-off-by: Mohamad Raizudeen <raizudeen.kerneldev@gmail.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Reviewed-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260723171203.4892-1-raizudeen.kerneldev@gmail.com
Some platforms require selectively disabling specific ASPM states on a
given PCIe link to avoid link instability or functional failures caused by
board-level connectivity constraints such as PCB routing, connectors,
slots, or external cabling.
Devicetree supports disabling ASPM L0s, L1, and L1 PM Substates via the
'aspm-no-l0s', 'aspm-no-l1' [1], and 'aspm-no-l1ss' [2] properties.
However, the ASPM driver does not currently honor these properties when
initializing the default link state.
When firmware enables L1 PM Substates before the kernel takes over,
masking aspm_support alone is insufficient to disable them in hardware.
pcie_config_aspm_link() guards L1SS configuration behind a check on
aspm_capable, which is derived from aspm_support. Once aspm_support is
masked, pcie_config_aspm_l1ss() is never called, leaving
firmware-enabled L1SS substates active in hardware.
Fix this by introducing pcie_link_has_aspm_override() to check for DT
override properties on either endpoint of the link. In
pcie_aspm_override_default_link_state(), use it to:
- Mask aspm_support, aspm_default, and aspm_enabled for any disabled
state, so software's view of the link stays in sync with what is
actually programmed in hardware. Leaving aspm_enabled stale would make
pcie_aspm_enabled() and the aspm sysfs attributes report a state as
active even after it has been masked, and could cause
pcie_config_aspm_link()'s "already in requested state" check to skip
reprogramming hardware to match.
- Explicitly call pcie_config_aspm_l1ss(link, 0) before masking
aspm_support when firmware has L1SS active and DT requests disabling L1
or L1SS, since pcie_config_aspm_link() will no longer do so once
aspm_capable is derived from the masked aspm_support.
Move the aspm_default initialization and
pcie_aspm_override_default_link_state() call in pcie_aspm_cap_init() to
before the "Restore L0s/L1" block. pcie_aspm_cap_init() disables L1 in
hardware prior to aspm_l1ss_init() and re-enables it only in the restore
block. Calling pcie_config_aspm_l1ss() while L1 is already disabled
satisfies its precondition ("Caller must disable L1 first"), whereas the
previous placement after the restore violated it.
Since the restore block writes back the parent_lnkctl/child_lnkctl snapshot
taken from hardware before the DT override ran, mask the L0s and L1 enable
bits out of that snapshot for any state the override has just disabled in
aspm_support. Otherwise the restore step would unconditionally reprogram
the link back to firmware's original L0s/L1 configuration, defeating the
Devicetree override it is meant to enforce.
Move pcie_config_aspm_l1ss() earlier in the file so it can be called
from pcie_aspm_override_default_link_state().
Link [1]: https://github.com/devicetree-org/dt-schema/pull/188
Link [2]: https://github.com/devicetree-org/dt-schema/pull/190
Signed-off-by: Krishna Chaitanya Chundru <krishna.chundru@oss.qualcomm.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Reviewed-by: Manivannan Sadhasivam <mani@kernel.org>
Link: https://patch.msgid.link/20260727-aspm-v6-3-2ebb3ee7ef71@oss.qualcomm.com
pcie_aspm_cap_init() disables ASPM L0s/L1 before touching L1SS config, then
restores the pre-existing state afterward. Both steps only ever touched
link->downstream, i.e. function 0 of the downstream component, leaving
sibling functions (>0) on a multi-function device untouched.
This means the "disable" step does not actually disable ASPM link-wide on a
multi-function device: a sibling function can still have L1 enabled even
after this step runs. PCIe r7.0, sec 7.5.3.7, recommends programming the
same ASPM Control value for all functions of a multi-function device, and
pcie_config_aspm_link() already loops over every function on the bus for
exactly this reason.
Loop over every function on linkbus->devices for both the disable and
restore steps, keeping the existing sec 7.5.3.7 ordering (disable
downstream functions before upstream, restore upstream before downstream
functions). The masked pcie_capability_clear_and_set_word() accessor from
the previous commit makes this safe: it only ever touches the ASPM Control
bits, so function-specific bits elsewhere in LNKCTL (e.g. Read Completion
Boundary, CLKREQ Enable) on sibling functions are left untouched.
Fixes: 7447990137 ("PCI/ASPM: Disable L1 before disabling L1 PM Substates")
Closes: https://lore.kernel.org/all/20260721143945.86E7D1F000E9@smtp.kernel.org/
Signed-off-by: Krishna Chaitanya Chundru <krishna.chundru@oss.qualcomm.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Reviewed-by: Manivannan Sadhasivam <mani@kernel.org>
Link: https://patch.msgid.link/20260727-aspm-v6-2-2ebb3ee7ef71@oss.qualcomm.com
Writing a PCI Host Controller driver requires bringing up the Root Complex
hardware and registering it with the PCI core in a specific sequence. Add
a guide describing these steps to help developers write new drivers.
It covers the Root Complex topology and enumeration, and walks through the
driver flow, including resource setup, Configuration Space accessors,
address translation, interrupt handling, Link training, power management,
shutdown and removal, using standard guidelines/best practices.
Signed-off-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://patch.msgid.link/20260803-pci-doc-v1-1-2814f8672cad@oss.qualcomm.com
pcie_aspm_cap_init() disables ASPM L0s/L1 on both ends of the Link
before touching L1SS config, then later restores the LNKCTL state
that was in effect beforehand. Both steps use raw
pcie_capability_write_word() calls: the disable step computes the
new value by hand from a snapshot taken earlier in the function, and
the restore step writes that same snapshot straight back.
Switch both steps to pcie_capability_clear_and_set_word(), masked to
PCI_EXP_LNKCTL_ASPMC, matching the accessor pcie_config_aspm_dev()
already uses elsewhere in this file for the exact same register. This
does a live read-modify-write of just the ASPM Control bits instead of
relying on a stale snapshot for the rest of the word, and is
consistent with how the rest of the file already touches this
register. No functional change.
Fixes: 7447990137 ("PCI/ASPM: Disable L1 before disabling L1 PM Substates")
Closes: https://lore.kernel.org/all/20260721143945.86E7D1F000E9@smtp.kernel.org/
Signed-off-by: Krishna Chaitanya Chundru <krishna.chundru@oss.qualcomm.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Reviewed-by: Manivannan Sadhasivam <mani@kernel.org>
Link: https://patch.msgid.link/20260727-aspm-v6-1-2ebb3ee7ef71@oss.qualcomm.com
According to PCIe r7.0, sec 5.3.3.2, two link wakeup mechanisms are
defined: Beacon and WAKE#. Beacon is a hardware-only mechanism and is
invisible to software (sec 4.2.7.8.1). This change adds support for the
WAKE# mechanism in the PCI core.
According to the PCIe specification, multiple WAKE# signals can exist in a
system or several components in the hierarchy may share a single WAKE#
signal. In configurations involving a PCIe switch, each downstream port
(DSP) of the switch may be connected to a separate WAKE# line, allowing
each endpoint to signal WAKE# independently. From figure 5.4 in sec
5.3.3.2, WAKE# can also be terminated at the switch itself. Such topologies
are typically not described in Device Tree, therefore it is out of scope
for this series.
To support this, the WAKE# should be described in the device tree node of
the endpoint/bridge. If all endpoints share a single WAKE# line, then each
endpoint node shall describe the same WAKE# signal or a single WAKE# in the
Root Port node.
In pci_device_add(), PCI framework will search for the WAKE# in device
node. Once found, register for the wake IRQ through
dev_pm_set_dedicated_wake_irq() associates a wakeup IRQ with a device and
requests it, but the PM core keeps the IRQ disabled by default. The IRQ is
enabled by the PM core, only when the device is permitted to wake the
system, i.e. during system suspend and after runtime suspend, and only when
device wakeup is enabled.
If the same WAKE# GPIO is described in multiple device tree nodes, only the
first device that successfully registers the wake IRQ will succeed, while
subsequent registrations may fail. This limitation does not affect
functional correctness, since WAKE# is only used to bring the link to D0,
and endpoint-specific wakeup handling is resolved later through PME
detection (PME_EN is set in suspend path by PCI core by default).
When the wake IRQ fires, the wakeirq handler invokes pm_runtime_resume() to
bring the device back to an active power state, such as transitioning from
D3cold to D0. Once the device is active and the link is usable, the
endpoint may generate a PME, which is then handled by the PCI core through
PME polling or the PCIe PME service driver to complete the wakeup of the
endpoint.
WAKE# is added in dts schema and merged based on below links.
Link: https://lore.kernel.org/all/20250515090517.3506772-1-krishna.chundru@oss.qualcomm.com/
Link: https://github.com/devicetree-org/dt-schema/pull/170
Signed-off-by: Krishna Chaitanya Chundru <krishna.chundru@oss.qualcomm.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Reviewed-by: Linus Walleij <linus.walleij@linaro.org>
Reviewed-by: Bartosz Golaszewski <bartosz.golaszewski@oss.qualcomm.com>
Acked-by: Manivannan Sadhasivam <mani@kernel.org>
Link: https://patch.msgid.link/20260707-wakeirq_support-v12-1-b4453f5bcc97@oss.qualcomm.com
Commit eb3b5bf1a8 ("PCI: Whitelist native hotplug ports for runtime D3"),
prevented native hotplug-capable Root Ports from entering D3 citing issues
on old Intel SkyLake Xeon-SP platform.
But there is no reason to restrict D3 for native hotplug-capable Root Ports
on non-x86 platforms. We recently enabled D3 on non hotplug-capable Root
Ports on non-x86 platforms (specifically for DT platforms) in commit
a5fb3ff632 ("PCI: Allow PCI bridges to go to D3Hot on all non-x86"). So
do the same for native hotplug-capable Root Ports as well.
Signed-off-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://patch.msgid.link/20260729165005.896725-1-manivannan.sadhasivam@oss.qualcomm.com
Correct a few white-space issues, like double space after '=' or before
bracket '{' characters, which will be flagged by dt-check-style. No
functional changes.
Signed-off-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>
Signed-off-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com>
Reviewed-by: Marek Vasut <marek.vasut+renesas@mailbox.org>
Acked-by: Rob Herring (Arm) <robh@kernel.org>
Link: https://patch.msgid.link/20260801195517.235161-2-krzysztof.kozlowski@oss.qualcomm.com
The Realtek RTS525A PCIe card reader reports an AER Correctable Replay
Timer Timeout storm when ASPM L0s is enabled on its link. On an affected
HP ZBook Power 16 inch G11, the Root Port received tens of millions of AER
interrupts from the RTS525A even when the rtsx_pci driver was blacklisted
and the endpoint was not enabled by a driver.
For example:
pcieport 0000:00:1c.6: AER: Multiple Correctable error message received from 0000:58:00.0
rtsx_pci 0000:58:00.0: PCIe Bus Error: severity=Correctable, type=Data Link Layer, (Transmitter ID)
rtsx_pci 0000:58:00.0: device [10ec:525a] error status/mask=00001000/00006000
rtsx_pci 0000:58:00.0: [12] Timeout
pcieport 0000:00:1c.6: AER: Correctable error message received from 0000:58:00.0
Testing with OS-native AER control showed that disabling only L0s on the
RTS525A link stops new AER interrupt and counter growth while leaving L1
enabled. Disabling L1, L1 substates, or Clock PM alone did not stop the
storm.
Prevent the broken L0s configuration by removing L0s from the RTS525A
advertised ASPM capability. This avoids enabling the non-working ASPM
state instead of masking the resulting AER Replay Timer Timeout reports.
Signed-off-by: Max Lee <max.lee@canonical.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Reviewed-by: Lukas Wunner <lukas@wunner.de>
Reviewed-by: Manivannan Sadhasivam <mani@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260707021527.639611-1-max.lee@canonical.com
The cadence-plat driver has a single platform driver that can be built-in
or a loadable module, but it calls two separate backend drivers depending
on whether it is a host or endpoint.
If one of the mode is build as built-in and another as loadable module,
we end up with a situation where the built-in pcie-cadence-plat driver
tries to call the modular host or endpoint driver, which causes a link
failure:
ld: error: undefined symbol: cdns_pcie_ep_setup
>>> referenced by pcie-cadence-plat.c
>>> drivers/pci/controller/cadence/pcie-cadence-plat.o:(cdns_plat_pcie_probe) in archive vmlinux.a
ld: error: undefined symbol: cdns_pcie_host_setup
>>> referenced by pcie-cadence-plat.c
>>> drivers/pci/controller/cadence/pcie-cadence-plat.o:(cdns_plat_pcie_probe) in archive vmlinux.a
Fix this by moving the 'select' of PCIE_CADENCE_HOST and PCIE_CADENCE_EP
from the individual PLAT_HOST/PLAT_EP symbols into the common PCIE_CADENCE_PLAT
symbol, conditioned on which backends (modes) are enabled.
Fixes: 611627a4e5 ("PCI: cadence: Add module support for platform controller driver")
Reported-by: Randy Dunlap <rdunlap@infradead.org>
Closes: https://lore.kernel.org/linux-next/589ea512-93c6-4e1c-83d7-ba45a0b35843@infradead.org/
Signed-off-by: Aksh Garg <a-garg7@ti.com>
Signed-off-by: Manivannan Sadhasivam <manivannan.sadhasivam@oss.qualcomm.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Tested-by: Randy Dunlap <rdunlap@infradead.org>
Acked-by: Randy Dunlap <rdunlap@infradead.org>
Link: https://patch.msgid.link/20260805105014.3952686-1-a-garg7@ti.com
proc_bus_pci_read() decides how much of the config space is readable based
on capable(CAP_SYS_ADMIN), which checks the credentials of the task calling
read(), not the credentials of the process that opened the file.
The sysfs equivalent, pci_read_config(), has checked the credentials of the
opening process since commit de139a3393 ("pci: check caps from sysfs file
open to read device dependent config space"), so a privileged process can
open the config space file and pass the file descriptor to an unprivileged
process (for example, a process running a KVM guest with an assigned
device), which can then read the entire config space. The check was
subsequently routed through the LSM framework in commit 47970b1b2a ("pci:
use security_capable() when checking capablities during config space read")
and converted to the dedicated helper in commit ab0fa82b2d ("pci-sysfs:
use proper file capability helper function").
Thus, the two interfaces check the same capability against different
credentials. Checking the credentials of the task calling read() makes the
outcome depend on who reads rather than who opened, so the restriction is
bypassed whenever a more privileged process reads through the descriptor.
Checking the credentials recorded in file->f_cred settles the decision at
open() time and ties it to the file, where it cannot change with the
caller.
Use file_ns_capable() to check CAP_SYS_ADMIN against the credentials in
effect when the file was opened, bringing the procfs interface in line with
the sysfs behaviour.
As a result, a file descriptor opened by a privileged process and passed to
an unprivileged one now allows the entire config space to be read through
procfs, matching sysfs.
Signed-off-by: Krzysztof Wilczyński <kwilczynski@kernel.org>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260720204145.1500105-1-kwilczynski@kernel.org
Currently, a driver can claim a region of a device's config space as
exclusive using pci_request_config_region_exclusive(), after which a write
to that region originating from user space is expected to emit a warning
and taint the kernel. The check is advisory only, as the write itself is
still allowed to proceed.
Since commit 278294798a ("PCI: Allow drivers to request exclusive config
regions"), the sysfs config space attribute performs this check in
pci_write_config(), but the procfs interface was never updated. A write
performed through /proc/bus/pci/BB/DD.F therefore bypasses the detection
entirely, even though both interfaces offer the same level of access.
Add the same resource_is_exclusive() check to proc_bus_pci_write().
Signed-off-by: Krzysztof Wilczyński <kwilczynski@kernel.org>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260729075413.1215821-1-kwilczynski@kernel.org
Currently, proc_bus_pci_read() and proc_bus_pci_write() do not return early
for zero-length configuration space accesses at valid offsets.
Such an access invokes pci_config_pm_runtime_get() and
pci_config_pm_runtime_put() around transfer blocks that do nothing.
This is a problem because pci_config_pm_runtime_get() synchronously resumes
the upstream bridge through pm_runtime_get_sync(), and resumes the device
itself through pm_runtime_resume() when it is in D3cold, only for the
handler to return zero immediately afterwards. Such a spurious wakeup
wastes power and adds needless resume latency.
The sysfs core already returns early for in-range zero-length binary
attribute accesses before pci_read_config() or pci_write_config() is
invoked. In contrast, the VFS forwards zero-length requests to the procfs
callbacks, where they continue into runtime PM handling.
Return early from proc_bus_pci_read() and proc_bus_pci_write() when nbytes
is zero, before any runtime PM involvement.
The value returned to userspace at these offsets remains zero,
so the change is not visible to userspace.
Signed-off-by: Krzysztof Wilczyński <kwilczynski@kernel.org>
[bhelgaas: order tags]
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260729075909.1219906-1-kwilczynski@kernel.org
Currently, the static binary attributes for the PCI legacy I/O port
and ISA memory space files (legacy_io, legacy_io_sparse, legacy_mem and
legacy_mem_sparse) are open-coded, with each definition repeating the
same set of properties and callbacks.
Add two macros for declaring such attributes:
- pci_legacy_resource_io_attr(), for legacy I/O port space (read/write)
- pci_legacy_resource_mem_attr(), for legacy memory space (mmap)
Each macro takes the fixed attribute size as a parameter.
Then replace the open-coded definitions with the newly added macros.
No functional changes intended.
Signed-off-by: Krzysztof Wilczyński <kwilczynski@kernel.org>
[bhelgaas: shorten macros to fit in 80 columns]
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://patch.msgid.link/20260721020427.1541197-4-kwilczynski@kernel.org
Currently, the __pci_dev_resource_attr() helper macro takes the attribute
name suffix as its third parameter, even though the suffix is what
distinguishes the three attribute variants built on top of it.
Additionally, the pci_dev_resource_attr() wrapper passes an empty suffix,
and with the suffix placed in the middle of the parameter list its
invocation contains two consecutive commas, which checkpatch.pl highlights,
as follows:
ERROR: space required after that ',' (ctx:VxO)
Move the suffix to the front so that the variant selector comes first and
the empty argument follows the opening parenthesis, which checkpatch.pl
does not complain about. This also matches the parameter order used by the
PCI legacy I/O and memory attribute macros introduced in a subsequent
change.
No functional changes intended.
Signed-off-by: Krzysztof Wilczyński <kwilczynski@kernel.org>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://patch.msgid.link/20260721020427.1541197-3-kwilczynski@kernel.org
Currently, the static binary attributes for the PCI resource files are
generated with the names dev_resource<N>_io_attr, dev_resource<N>_uc_attr
and dev_resource<N>_wc_attr.
The macros that generate these attributes and the arrays that collect them
already carry the pci_ prefix, as do the sibling legacy I/O and memory
attributes, such as pci_legacy_io_attr. Only the generated variable names
lack it.
Rename the generated variables to pci_dev_resource<N>_io_attr,
pci_dev_resource<N>_uc_attr and pci_dev_resource<N>_wc_attr, and update the
attribute pointer arrays to match.
While at it, re-align the continuation backslashes in the resource
attribute macros to match.
No functional changes intended.
Signed-off-by: Krzysztof Wilczyński <kwilczynski@kernel.org>
[bhelgaas: shorten pci_dev_resource##_bar##_wc_attr to fit in 80 columns]
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://patch.msgid.link/20260721020427.1541197-2-kwilczynski@kernel.org
Currently, the legacy I/O and memory sysfs handlers do not check
security_locked_down(LOCKDOWN_PCI_ACCESS), leaving the legacy_io and
legacy_mem files unprotected when the kernel is locked down.
Commit eb627e1772 ("PCI: Lock down BAR access when the kernel is locked
down") added the check to pci_write_config(), pci_mmap_resource(), and
pci_write_resource_io() to prevent userspace from programming DMA-capable
hardware that could be used to modify kernel code, but did not cover the
legacy handlers.
As a result, root can still write arbitrary I/O ports and map the legacy
I/O and memory spaces while the kernel is locked down, which is the same
capability the lockdown is meant to remove.
Add the same check to pci_write_legacy_io(), pci_mmap_legacy_mem(), and
pci_mmap_legacy_io().
These generic handlers cover both architectures that define HAVE_PCI_LEGACY
(such as Alpha and PowerPC).
Fixes: eb627e1772 ("PCI: Lock down BAR access when the kernel is locked down")
Signed-off-by: Krzysztof Wilczyński <kwilczynski@kernel.org>
[bhelgaas: add Link]
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://patch.msgid.link/20260720211541.1509744-1-kwilczynski@kernel.org
Currently, mmap() of a resourceN file for an I/O BAR fails with -ENODEV on
architectures where arch_can_pci_mmap_io() is 0, such as x86, because the
attribute has no mmap callback there and the error comes from the generic
kernfs dispatch.
This is a side effect of commit e854d8b2a8 ("PCI: Add
arch_can_pci_mmap_io() on architectures which can mmap() I/O space"),
which removed the mmap callback from the I/O resource attribute on
these architectures. Previously the request reached the architecture
mmap code and failed with -EINVAL, and the same commit deliberately
kept -EINVAL for the identical operation on the procfs interface, so
the two PCI userspace interfaces have disagreed ever since.
Add a pci_mmap_resource_io_unsupported() callback that
returns -EINVAL and use it as the mmap handler of the I/O resource
attribute when arch_can_pci_mmap_io() is 0, so the failure is
produced deliberately by PCI code, consistent with the procfs
interface and with the behaviour before e854d8b2a8.
Architectures where arch_can_pci_mmap_io() is non-zero keep the real
pci_mmap_resource_uc() handler and are unaffected. The mmap() fails
either way. Only the reported error changes from -ENODEV to -EINVAL.
Fixes: e854d8b2a8 ("PCI: Add arch_can_pci_mmap_io() on architectures which can mmap() I/O space")
Signed-off-by: Krzysztof Wilczyński <kwilczynski@kernel.org>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Link: https://patch.msgid.link/20260720204624.1503794-1-kwilczynski@kernel.org