DRM Rust changes for v7.3-rc1

- I/O (shared from driver-core tree via signed tag rust-io-7.3-rc1):
 
   - Rework of I/O types: make I/O regions typed (with a
     dynamically-sized Region type for the existing untyped case), create
     view types representing subregions of a mapped I/O region, and add
     io_project!() for safely creating subviews.
 
   - Split Io into a base trait (IoBase) and an extension trait (Io) with
     a blanket implementation, preventing implementers from overriding
     provided methods that unsafe code relies on.
 
   - Add a SysMem backend for shared system memory with volatile access,
     and make Coherent implement Io via an I/O view type.  Add copying
     methods (memcpy_{from,to}io).
 
   - Replace dma_read!/dma_write! with io_read!/io_write!; drop the old
     macros.
 
 - DRM:
 
   - RegistrationGuard and RegistrationData:
     - Rework DeviceContext typestates: rename Uninit to Normal, add an
       Ioctl context, restrict AlwaysRefCounted to Normal for both Device
       and GEM Object, and establish a Deref chain from Registered to
       Normal.
 
     - Introduce RegistrationGuard, a guard representing a
       drm_dev_enter/exit SRCU critical section that proves the DRM
       device is registered, which implies the parent bus device is still
       bound.
 
     - Add RegistrationData as a GAT on drm::Driver. The data does not
       outlive driver unbind, so it can capture lifetime-annotated device
       resources and references. Accessible through the guard via a
       closure with HRTB lifetime.
 
     - Wrap ioctl dispatch in RegistrationGuard (returning ENODEV if
       unplugged) and pass registration data to handlers.
 
     - Add Driver::ParentDevice associated type.
 
   - Fix unbounded lifetimes in ioctl handler arguments.
 
   - Fix a race in drm_dev_register() where a partial failure allowed
     in-flight ioctls to proceed while the error path tore down
     resources.
 
   - GEM shmem: add DmaResvGuard helper, vmap functions, and sg_table()
     accessor.
 
   - GPUVM: require Send + Sync for the driver's associated data,
     implement Send and Sync for GpuVaAlloc and GpuVmBo, add SmContext
     lifetime bound, update DriverGpuVm for DeviceContext.
 
 - Nova:
 
   - nova-core / nova-drm cross-crate dependency:
     - Build nova-core and nova-drm from drivers/gpu/Makefile for build
       ordering, export nova-core Rust symbols for nova-drm. Workaround
       until the build system supports Rust cross-crate dependencies
       natively.
 
   - GSP boot process consolidation:
     - Introduce GspBootContext to bundle common boot parameters,
       replacing per-argument threading. Separate context and GPU
       lifetimes to support mutable borrows of GPU subdevices.
 
     - Turn FWSEC execution into a HAL method, make FWSEC bootloader
       usage a property of the TU102 HAL (GA102+ gets its own instance
       with it disabled). Move firmware file selection to the GSP HAL.
 
     - Store the Fsp instance in Gpu (lifetime tied to the GPU, not just
       a single boot invocation). Move GSP state and unload bundle into a
       pinned subobject for reliable teardown on partial init failure.
 
   - Boot GSP with vGPU enabled:
     - Add PRC (Product Reconfiguration Control) protocol to query device
       configuration from the FSP. Read vGPU mode, detect and store vGPU
       state.
     - Set RMSetSriovMode registry entry and reserve the larger WPR2 heap
       required when vGPU is enabled.
     - Build SetRegistry entries dynamically.
 
   - TLV firmware image format:
     - Add a TLV (type-length-value) parser for the new firmware image
       format. TLV files use unversioned filenames with a .tlv suffix,
       start with "NVFW" magic, and contain tagged blocks with 4-byte
       aligned payloads.
     - Transition all firmware loading (booter, gsp, gen_bootloader, fsp)
       to TLV images.
     - Note: this requires a development firmware not in linux-firmware
       [1]; this is temporary and serves the transition to r615.
 
   - Hopper/Blackwell fixes and cleanups:
     - Correct FRTS vidmem offset calculation, split FbLayout into FSP
       and non-FSP versions, fix Blackwell flush address composition, use
       absolute FBHUB0 flush registers on Blackwell, use correct sysmem
       flush registers on Hopper.
     - Harden FSP messaging: limit receive allocation size, catch bogus
       queue pointers, ensure DMA allocation lifetimes for FMC boot and
       LibOS, wait for RISC-V HALTED on unload.
 
   - I/O projection adoption:
     - Use io_project!() for PTE array, message queues, and Falcon DMA
       transfer bounds checking.
 
   - Misc:
     - Keep unloading if FWSEC-SB fails during Turing/Ampere GSP reset.
     - Don't declare booter firmware for FSP chipsets.
     - Fix packed registry table size.
     - Extract and display usable FB regions from GSP.
     - Store bar and dev directly in Falcon, simplifying the API.
     - Parse VBIOS structs via zerocopy.
     - Convert to kernel bitfield macro, remove local one.
     - Move register definitions into sub-modules.
     - Add FSP and PRC protocol documentation.
 
 - Tyr:
 
   - Firmware loading and MCU boot:
     - Add a generic slot manager for dynamically allocating limited
       hardware slots to software seats, with lazy eviction under
       contention.
     - Add MMU support wrapping the slot manager for address-space slot
       allocation, with MAIR-to-MEMATTR translation.
     - Add GPU virtual memory (VM) support using drm_gpuvm with ARM64
       LPAE Stage 1 page tables and 4KB/2MB page sizes.
     - Add a kernel buffer object type for internal driver allocations.
     - Add a parser for the Mali CSF firmware binary format.
     - Add MCU booting: load, parse, and map firmware sections into VM,
       then boot the MCU at probe().
 
 - Cross-subsystem:
 
   - Add faux::Device type with AsBusDevice support. Allow retrieving a
     bound Device from a Registration.
 
   - Add device lifetime to IoPageTable.
 
   - Add Vec::zeroed method.
 
   - Add firmware::request_into_buf() to load firmware into a
     caller-provided buffer.
 
   - Rename dma_handle to dma_address in the DMA abstraction.
 
   - Change pci_sriov_get_totalvfs() return type to unsigned int; add
     Rust helper.
 
 [1] https://github.com/ttabi/linux-firmware-nova
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQS2q/xV6QjXAdC7k+1FlHeO1qrKLgUCandOTAAKCRBFlHeO1qrK
 LgWSAP4wHxhEuOme55bgZkne0XMums8bLln69N/UR+Rim+agJQD7BmgsL0ANxGOu
 Csnsxej/tcktyraoy/QGHMjYX7pcGwM=
 =UXCM
 -----END PGP SIGNATURE-----

Merge tag 'drm-rust-next-2026-08-08' of https://gitlab.freedesktop.org/drm/rust/kernel into drm-next

DRM Rust changes for v7.3-rc1

- I/O (shared from driver-core tree via signed tag rust-io-7.3-rc1):

  - Rework of I/O types: make I/O regions typed (with a
    dynamically-sized Region type for the existing untyped case), create
    view types representing subregions of a mapped I/O region, and add
    io_project!() for safely creating subviews.

  - Split Io into a base trait (IoBase) and an extension trait (Io) with
    a blanket implementation, preventing implementers from overriding
    provided methods that unsafe code relies on.

  - Add a SysMem backend for shared system memory with volatile access,
    and make Coherent implement Io via an I/O view type.  Add copying
    methods (memcpy_{from,to}io).

  - Replace dma_read!/dma_write! with io_read!/io_write!; drop the old
    macros.

- DRM:

  - RegistrationGuard and RegistrationData:
    - Rework DeviceContext typestates: rename Uninit to Normal, add an
      Ioctl context, restrict AlwaysRefCounted to Normal for both Device
      and GEM Object, and establish a Deref chain from Registered to
      Normal.

    - Introduce RegistrationGuard, a guard representing a
      drm_dev_enter/exit SRCU critical section that proves the DRM
      device is registered, which implies the parent bus device is still
      bound.

    - Add RegistrationData as a GAT on drm::Driver. The data does not
      outlive driver unbind, so it can capture lifetime-annotated device
      resources and references. Accessible through the guard via a
      closure with HRTB lifetime.

    - Wrap ioctl dispatch in RegistrationGuard (returning ENODEV if
      unplugged) and pass registration data to handlers.

    - Add Driver::ParentDevice associated type.

  - Fix unbounded lifetimes in ioctl handler arguments.

  - Fix a race in drm_dev_register() where a partial failure allowed
    in-flight ioctls to proceed while the error path tore down
    resources.

  - GEM shmem: add DmaResvGuard helper, vmap functions, and sg_table()
    accessor.

  - GPUVM: require Send + Sync for the driver's associated data,
    implement Send and Sync for GpuVaAlloc and GpuVmBo, add SmContext
    lifetime bound, update DriverGpuVm for DeviceContext.

- Nova:

  - nova-core / nova-drm cross-crate dependency:
    - Build nova-core and nova-drm from drivers/gpu/Makefile for build
      ordering, export nova-core Rust symbols for nova-drm. Workaround
      until the build system supports Rust cross-crate dependencies
      natively.

  - GSP boot process consolidation:
    - Introduce GspBootContext to bundle common boot parameters,
      replacing per-argument threading. Separate context and GPU
      lifetimes to support mutable borrows of GPU subdevices.

    - Turn FWSEC execution into a HAL method, make FWSEC bootloader
      usage a property of the TU102 HAL (GA102+ gets its own instance
      with it disabled). Move firmware file selection to the GSP HAL.

    - Store the Fsp instance in Gpu (lifetime tied to the GPU, not just
      a single boot invocation). Move GSP state and unload bundle into a
      pinned subobject for reliable teardown on partial init failure.

  - Boot GSP with vGPU enabled:
    - Add PRC (Product Reconfiguration Control) protocol to query device
      configuration from the FSP. Read vGPU mode, detect and store vGPU
      state.
    - Set RMSetSriovMode registry entry and reserve the larger WPR2 heap
      required when vGPU is enabled.
    - Build SetRegistry entries dynamically.

  - TLV firmware image format:
    - Add a TLV (type-length-value) parser for the new firmware image
      format. TLV files use unversioned filenames with a .tlv suffix,
      start with "NVFW" magic, and contain tagged blocks with 4-byte
      aligned payloads.
    - Transition all firmware loading (booter, gsp, gen_bootloader, fsp)
      to TLV images.
    - Note: this requires a development firmware not in linux-firmware
      [1]; this is temporary and serves the transition to r615.

  - Hopper/Blackwell fixes and cleanups:
    - Correct FRTS vidmem offset calculation, split FbLayout into FSP
      and non-FSP versions, fix Blackwell flush address composition, use
      absolute FBHUB0 flush registers on Blackwell, use correct sysmem
      flush registers on Hopper.
    - Harden FSP messaging: limit receive allocation size, catch bogus
      queue pointers, ensure DMA allocation lifetimes for FMC boot and
      LibOS, wait for RISC-V HALTED on unload.

  - I/O projection adoption:
    - Use io_project!() for PTE array, message queues, and Falcon DMA
      transfer bounds checking.

  - Misc:
    - Keep unloading if FWSEC-SB fails during Turing/Ampere GSP reset.
    - Don't declare booter firmware for FSP chipsets.
    - Fix packed registry table size.
    - Extract and display usable FB regions from GSP.
    - Store bar and dev directly in Falcon, simplifying the API.
    - Parse VBIOS structs via zerocopy.
    - Convert to kernel bitfield macro, remove local one.
    - Move register definitions into sub-modules.
    - Add FSP and PRC protocol documentation.

- Tyr:

  - Firmware loading and MCU boot:
    - Add a generic slot manager for dynamically allocating limited
      hardware slots to software seats, with lazy eviction under
      contention.
    - Add MMU support wrapping the slot manager for address-space slot
      allocation, with MAIR-to-MEMATTR translation.
    - Add GPU virtual memory (VM) support using drm_gpuvm with ARM64
      LPAE Stage 1 page tables and 4KB/2MB page sizes.
    - Add a kernel buffer object type for internal driver allocations.
    - Add a parser for the Mali CSF firmware binary format.
    - Add MCU booting: load, parse, and map firmware sections into VM,
      then boot the MCU at probe().

- Cross-subsystem:

  - Add faux::Device type with AsBusDevice support. Allow retrieving a
    bound Device from a Registration.

  - Add device lifetime to IoPageTable.

  - Add Vec::zeroed method.

  - Add firmware::request_into_buf() to load firmware into a
    caller-provided buffer.

  - Rename dma_handle to dma_address in the DMA abstraction.

  - Change pci_sriov_get_totalvfs() return type to unsigned int; add
    Rust helper.

[1] https://github.com/ttabi/linux-firmware-nova

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: "Danilo Krummrich" <dakr@kernel.org>
Link: https://patch.msgid.link/DKJQQUOS0PVO.3JPR3MYK4PDVZ@kernel.org
This commit is contained in:
Dave Airlie 2026-08-10 14:47:20 +10:00
commit 29daacacf4
107 changed files with 8810 additions and 3491 deletions

View File

@ -0,0 +1,142 @@
.. SPDX-License-Identifier: GPL-2.0
===================================================
FSP (Foundation Security Processor) and Secure Boot
===================================================
This document describes the role of the FSP in the GPU boot sequence on
Hopper and Blackwell GPUs, and how it differs from the earlier Ampere boot
flow. It also provides a brief overview of the PRC (Product Reconfiguration
Control) protocol used to query device configuration through FSP. As with
other documents in this directory, the information is subject to change and
is intended to help developers understand the corresponding kernel code.
What is FSP?
============
The Foundation Security Processor (FSP) is the GPU's Internal Root of Trust
(IROT). It is a dedicated security processor that boots from immutable ROM
(Boot ROM) inside the GPU and is responsible for establishing the Chain of
Trust before any other firmware is allowed to run.
FSP runs independently of the host CPU and starts executing as soon as the
GPU is powered on. By the time the nova-core driver is loaded, FSP has
already completed its own secure boot and is ready to accept commands from
the driver.
Simplified boot flow (Hopper/Blackwell)
=======================================
Starting with Hopper, the boot flow is significantly simplified compared to
earlier GPU generations like Ampere.
On an **Ampere** GPU, the boot verification chain involves multiple Falcon
engines and multiple ucode stages (see falcon.rst for details)::
Hardware BROM (SEC2)
-> HS Booter (SEC2)
-> LS GSP-RM (GSP)
The driver must extract ucode from VBIOS, manage SEC2 and GSP, and
orchestrate the Booter to load GSP-RM. This involves FWSEC-FRTS, devinit,
and the Booter stages.
On **Hopper/Blackwell** GPUs, FSP replaces this multi-stage process with a
single message-driven interface::
FSP (hardware root of trust, boots from ROM)
-> FMC (Falcon Microcontroller, verified by FSP)
-> GSP-RM (verified and loaded by FMC)
The driver only needs to:
1. Wait for FSP to complete its own secure boot (polling a scratch register).
2. Send a Chain of Trust (COT) message to FSP with the FMC firmware location,
cryptographic signatures, and GSP boot parameters.
3. FSP authenticates the FMC firmware and boots it, FMC in turn loads GSP-RM.
There is no SEC2 involvement, no Booter ucode, and no FWSEC-FRTS stage. The
entire secure boot is driven by a single FSP message exchange.
Chain of Trust (COT) protocol
=============================
The Chain of Trust establishes a cryptographically enforced boot sequence,
ensuring the GPU reaches a known, trusted state.
The driver communicates with FSP using a message queue (Falcon MSGQ
interface). Each message consists of an MCTP (Management Component Transport
Protocol) transport header and an NVDM (NVIDIA Vendor Defined Message) header,
followed by a protocol-specific payload.
For Chain of Trust, the payload includes:
- The system memory address of the FMC firmware image.
- Cryptographic material: a SHA-384 hash, RSA-3K public key, and RSA-3K
signature extracted from the FMC ELF firmware.
- FRTS (Firmware Runtime Services) region information (vidmem offset and size).
- The system memory address of the GSP boot arguments structure.
FSP verifies the signature against the provided public key and hash, and if
verification succeeds, boots the FMC. The FMC then authenticates and launches
GSP-RM.
The message flow is::
nova-core FSP
| |
| 1. Poll scratch register |
| (wait for FSP boot complete) |
| |
| 2. COT message ------------> |
| (FMC addr, signatures, |
| boot params) |
| |
| |--- Verify FMC signature
| |--- Boot FMC
| |--- FMC loads GSP-RM
| |
| 3. COT response <------------ |
| (success/error) |
| |
FSP message format
==================
All FSP messages share a common header format consisting of two 32-bit words:
**MCTP header** (Management Component Transport Protocol):
- Bit 31: SOM (Start of Message)
- Bit 30: EOM (End of Message)
- Bits 29:28: Packet sequence number
- Bits 23:16: Source Endpoint ID
**NVDM header** (NVIDIA Vendor Defined Message):
- Bits 6:0: MCTP message type (0x7e = vendor-defined PCI)
- Bits 23:8: PCI vendor ID (0x10de = NVIDIA)
- Bits 31:24: NVDM type (0x14 = COT, 0x13 = PRC, 0x15 = FSP response)
PRC (Product Reconfiguration Control) protocol
===============================================
PRC is an API system exposed through FSP's Management Partition that allows
querying and modifying device configuration without firmware updates.
Configuration parameters are called "knobs". Each knob has a unique object
ID and controls a specific device behavior. Examples include vGPU mode, ECC
enable, confidential computing mode, and NVLINK configuration.
Each knob has two values:
- **Active**: the currently effective value for this boot cycle.
- **Persistent**: the value stored in InfoROM, applied on subsequent boots.
The nova-core driver uses PRC to read the vGPU mode knob (object ID 0x29)
during early boot, before firmware loading, to determine whether the GPU
should operate in vGPU mode.
The PRC message format follows the same MCTP/NVDM header structure as COT,
with NVDM type 0x13. The payload contains:
- A sub-command (e.g., 0x0c for read).
- Flags indicating which value to read (bit 0 = persistent, bit 1 = active).
- The knob object ID.
The response includes the common FSP response header (with error status)
followed by the knob's 16-bit state value.

View File

@ -0,0 +1,184 @@
.. SPDX-License-Identifier: (GPL-2.0+ OR MIT)
==================================
TLV Tags in Nova Firmware Images
==================================
Nova firmware images use a Type-Length-Value (TLV) format to encapsulate
firmware components and metadata. The TLV file begins with a 4-byte "magic"
header that contains the string "NVFW". Following the header is a sequence of
TLV blocks.
Each block consists of a 4-byte tag of ASCII characters, a 4-byte length
encoded as a little-endian unsigned integer, and a sequence of bytes, the size
of which is equal to the length rounded up to the next multiple of 4.
The driver code that reads the TLV and uses its contents is called the parser.
It is the responsibility of the parser to handle missing or malformed tags,
lengths, and values in the TLV.
::
+------+------+------+------+
| 'N' | 'V' | 'F' | 'W' | Magic header
+------+------+------+------+
| Tag (4 bytes, ASCII) | TLV block 0
+---------------------------+
| Length (4 bytes, LE) |
+---------------------------+
| |
| Value (length bytes, |
| padded to 4-byte align) |
| |
+---------------------------+
| Tag (4 bytes, ASCII) | TLV block 1
+---------------------------+
| Length (4 bytes, LE) |
+---------------------------+
| |
| Value (length bytes, |
| padded to 4-byte align) |
| |
+---------------------------+
| ... | More TLV blocks
+---------------------------+
Tags and Length
===============
TLV tags are always four-character words, with all letters being upper case.
Duplicate tags are not allowed.
A TLV file may contain additional tags not described in this document.
Values
======
Values are one of four types. The type is not encoded in the format; rather,
the parser expects a given tag to have a value of a given type.
1) Integers, encoded in 32-bit or 64-bit little-endian format.
2) Strings, encoded as-is and required to be only printable ASCII characters
and without a null terminator.
3) An array of bytes, for binary data.
4) Boolean, encoded as single byte, with a value of 0 for False or 1 for True.
Common Tags
===========
These tags are shared across firmware types and carry the same meaning
wherever they appear. Unlike the firmware-specific tags below, a common tag
is reserved: its meaning is fixed and may never be redefined for a particular
firmware type.
``VERS`` (string)
Human-readable firmware version string. Present in all TLV files.
A TLV image must contain either a single ``BLOB`` tag (firmware embedded
inline) or a ``SIZE``/``FILE`` pair (firmware stored in a separate file).
``BLOB`` (bytes)
If the firmware microcode binary is stored in the TLV, this tag contains
the actual firmware image bytes.
``FILE`` (string)
If the firmware binary is stored as a separate file, this tag contains the
name of that file, which is required to be in the same directory as the TLV,
so no paths are allowed in the filename. This tag is always paired with
``SIZE``, so as to allow the driver to pre-allocate the buffer before
loading the file.
``SIZE`` (u32)
Total size in bytes of the firmware image to be loaded from the companion
file named by ``FILE``. This tag is mandatory if ``FILE`` exists, so the
size of the firmware image must be known when the TLV is created. If the
firmware image is updated and its size changes, then the TLV must be
updated with it.
GSP Firmware Tags
=================
``SIGN`` (bytes)
Cryptographic signature for the GSP firmware.
``BLID`` (string)
The build ID, extracted from the ".note.gnu.build-id" section.
Booter Firmware Tags
====================
``DAOF`` (u32) - ``os_data_offset``
OS data section offset within the firmware image (absolute byte offset).
Maps to the DMEM load source.
``DASZ`` (u32) - ``os_data_size``
OS data section size in bytes.
``CDOF`` (u32) - ``os_code_offset``
OS code section offset within the firmware image (absolute byte offset).
Maps to the non-secure IMEM load source.
``CDSZ`` (u32) - ``os_code_size``
OS code section size in bytes.
``PLOC`` (u32) - ``patch_loc``
Signature patch location -- byte offset within the firmware image where the
selected signature should be written.
``FUSE`` (u32) - ``fuse_version``
Fuse version of the firmware, used with the hardware fuse register to
select the correct signature index.
``ENID`` (u32) - ``engine_id``
Engine ID mask identifying the falcon engine this firmware targets.
``UCID`` (u32) - ``ucode_id``
Microcode ID used together with the engine ID to query hardware signature
fuse registers.
``A0CO`` (u32) - ``app0_code_offset``
App0 code offset -- start of the secure code region within the firmware
image. Used as the IMEM secure section source.
``A0CS`` (u32) - ``app0_code_size``
App0 code size in bytes.
``NSIG`` (u32) - ``num_sigs``
Number of signatures included in the ``SIGN`` tag.
``SIGN`` (bytes)
Concatenated array of firmware signatures. The size of each signature is
the total length of the ``SIGN`` value divided by ``NSIG``. The correct
signature is selected using the fuse-version-derived index.
Generic Bootloader Tags
=======================
``CDSZ`` (u32) - ``code_size``
Size in bytes of the bootloader code to copy from the ``BLOB`` tag and
PIO-load into falcon IMEM.
``STRT`` (u32) - ``start_tag``
Start tag identifying the IMEM block where execution begins. The falcon
boot address is derived as ``start_tag << 8``.
GSP Bootloader Tags
===================
``CDOF`` (u32) - ``code_offset``
Offset within the firmware image at which the code section starts.
``DAOF`` (u32) - ``data_offset``
Offset within the firmware image at which the data section starts.
``MFOF`` (u32) - ``manifest_offset``
Offset within the firmware image at which the manifest starts.
``APPV`` (u32) - ``app_version``
Application version of the firmware.
FMC Firmware Tags
=================
``HASH`` (bytes)
SHA-384 hash of the FMC firmware, exactly 48 bytes long.
``PKEY`` (bytes)
Public key used to verify the FMC firmware. At most 384 bytes (RSA-3072),
but may be shorter.
``SIGN`` (bytes)
Signature of the FMC firmware. At most 384 bytes (RSA-3072), but may
be shorter.

View File

@ -30,5 +30,7 @@ vGPU manager VFIO driver and the nova-drm driver.
core/todo
core/vbios
core/devinit
core/fsp
core/fwsec
core/falcon
core/tlv

View File

@ -7,4 +7,60 @@ obj-$(CONFIG_GPU_BUDDY) += buddy.o
obj-y += host1x/ drm/ vga/ tests/
obj-$(CONFIG_IMX_IPUV3_CORE) += ipu-v3/
obj-$(CONFIG_TRACE_GPU_MEM) += trace/
obj-$(CONFIG_NOVA_CORE) += nova-core/
# nova-core and nova-drm are built from this Makefile so nova-drm's dependency
# on nova-core can be expressed as a plain Make prerequisite rather than a
# recursive sub-make. This is a temporary workaround until the Rust build
# system supports cross-crate dependencies natively.
obj-$(CONFIG_NOVA_CORE) += nova-core.o
nova-core-y := nova-core/nova_core.o nova-core/nova_core_exports.o
obj-$(CONFIG_DRM_NOVA) += nova-drm.o
nova-drm-y := drm/nova/nova.o
# Export Rust symbols from nova-core only if nova-drm actually references them.
nova-core-export-deps := $(if $(CONFIG_DRM_NOVA),$(obj)/drm/nova/nova.o)
rust_needed_exports = \
{ $(if $(strip $(2)),$(NM) -u $(2);,) echo "__DEFINED_RUST_SYMBOLS__"; \
$(NM) -p --defined-only $(1); } | \
awk -v fmt='$(3)' ' \
/^__DEFINED_RUST_SYMBOLS__$$/ { defs = 1; next } \
!defs { if ($$NF ~ /^_R/) needed[$$NF] = 1; next } \
defs && $$2 ~ /(T|R|D|B)/ && $$3 ~ /^_R/ && \
$$3 !~ /_(init|cleanup)_module$$/ && \
$$3 !~ /__(pfx|cfi|odr_asan)/ && \
$$3 in needed { printf fmt, $$3 } \
'
quiet_cmd_exports = EXPORTS $@
cmd_exports = \
$(call rust_needed_exports,$<,$(nova-core-export-deps),EXPORT_SYMBOL_RUST_GPL(%s);\n) > $@
$(obj)/nova-core/exports_nova_core_generated.h: $(obj)/nova-core/nova_core.o $(nova-core-export-deps) FORCE
$(call if_changed,exports)
targets += nova-core/exports_nova_core_generated.h
$(obj)/nova-core/nova_core_exports.o: $(obj)/nova-core/exports_nova_core_generated.h
CFLAGS_nova-core/nova_core_exports.o := -I $(objtree)/$(obj)/nova-core
ifdef CONFIG_MODVERSIONS
# The C export shim declares Rust symbols as `extern int`, so reuse its export
# list but generate symbol CRCs from the Rust object instead of the shim's DWARF.
$(obj)/nova-core/nova_core_exports.o: private cmd_gensymtypes_c = \
$(call getexportsymbols,\1) | \
$(objtree)/scripts/gendwarfksyms/gendwarfksyms \
$(if $(KBUILD_GENDWARFKSYMS_STABLE), --stable) \
$(if $(KBUILD_SYMTYPES), --symtypes $(@:.o=.symtypes),) \
$(obj)/nova-core/nova_core.o
endif
# Output nova-core's crate metadata for use by nova-drm at compile time.
RUSTFLAGS_nova-core/nova_core.o += \
--emit=metadata=$(objtree)/$(obj)/nova-core/libnova_core.rmeta
# Allow nova-drm to import nova-core's types.
$(obj)/drm/nova/nova.o: $(obj)/nova-core/nova_core.o
RUSTFLAGS_drm/nova/nova.o := -L $(objtree)/$(obj)/nova-core --extern nova_core

View File

@ -186,7 +186,7 @@ obj-$(CONFIG_DRM_VMWGFX)+= vmwgfx/
obj-$(CONFIG_DRM_VGEM) += vgem/
obj-$(CONFIG_DRM_VKMS) += vkms/
obj-$(CONFIG_DRM_NOUVEAU) +=nouveau/
obj-$(CONFIG_DRM_NOVA) += nova/
# nova-drm is built from drivers/gpu/Makefile together with nova-core.
obj-$(CONFIG_DRM_EXYNOS) +=exynos/
obj-$(CONFIG_DRM_ROCKCHIP) +=rockchip/
obj-$(CONFIG_DRM_GMA500) += gma500/

View File

@ -475,6 +475,22 @@ void drm_dev_exit(int idx)
}
EXPORT_SYMBOL(drm_dev_exit);
/*
* Mark the device as unplugged and wait for any in-flight drm_dev_enter()
* critical sections to complete.
*/
static void drm_dev_synchronize_unplug(struct drm_device *dev)
{
/*
* After synchronizing any critical read section is guaranteed to see
* the new value of ->unplugged, and any critical section which might
* still have seen the old value of ->unplugged is guaranteed to have
* finished.
*/
dev->unplugged = true;
synchronize_srcu(&drm_unplug_srcu);
}
/**
* drm_dev_unplug - unplug a DRM device
* @dev: DRM device
@ -487,15 +503,7 @@ EXPORT_SYMBOL(drm_dev_exit);
*/
void drm_dev_unplug(struct drm_device *dev)
{
/*
* After synchronizing any critical read section is guaranteed to see
* the new value of ->unplugged, and any critical section which might
* still have seen the old value of ->unplugged is guaranteed to have
* finished.
*/
dev->unplugged = true;
synchronize_srcu(&drm_unplug_srcu);
drm_dev_synchronize_unplug(dev);
drm_dev_unregister(dev);
/* Clear all CPU mappings pointing to this device */
@ -1095,6 +1103,7 @@ int drm_dev_register(struct drm_device *dev, unsigned long flags)
goto err_minors;
dev->registered = true;
dev->unplugged = false;
if (driver->load) {
ret = driver->load(dev, flags);
@ -1122,6 +1131,13 @@ int drm_dev_register(struct drm_device *dev, unsigned long flags)
if (dev->driver->unload)
dev->driver->unload(dev);
err_minors:
/*
* If a minor was registered before the failure, userspace could have
* opened it and entered a drm_dev_enter() critical section. Ensure all
* such sections complete before we clean up.
*/
drm_dev_synchronize_unplug(dev);
remove_compat_control_link(dev);
drm_minor_unregister(dev, DRM_MINOR_ACCEL);
drm_minor_unregister(dev, DRM_MINOR_PRIMARY);

View File

@ -1,4 +1,3 @@
# SPDX-License-Identifier: GPL-2.0
obj-$(CONFIG_DRM_NOVA) += nova-drm.o
nova-drm-y := nova.o
# nova-drm is built from drivers/gpu/Makefile.
# nova.o (rust-analyzer marker - DO NOT REMOVE).

View File

@ -2,7 +2,10 @@
use kernel::{
auxiliary,
device::Core,
device::{
Core,
DeviceContext, //
},
drm::{
self,
gem,
@ -17,18 +20,14 @@
pub(crate) struct NovaDriver;
pub(crate) struct Nova {
pub(crate) struct Nova<'bound> {
#[expect(unused)]
drm: ARef<drm::Device<NovaDriver>>,
_reg: drm::Registration<'bound, NovaDriver>,
}
/// Convienence type alias for the DRM device type for this driver
pub(crate) type NovaDevice<Ctx = drm::Registered> = drm::Device<NovaDriver, Ctx>;
#[pin_data]
pub(crate) struct NovaData {
pub(crate) adev: ARef<auxiliary::Device>,
}
pub(crate) type NovaDevice<Ctx = drm::Normal> = drm::Device<NovaDriver, Ctx>;
const INFO: drm::DriverInfo = drm::DriverInfo {
major: 0,
@ -53,27 +52,32 @@ pub(crate) struct NovaData {
impl auxiliary::Driver for NovaDriver {
type IdInfo = ();
type Data<'bound> = Nova;
type Data<'bound> = Nova<'bound>;
const ID_TABLE: auxiliary::IdTable<Self::IdInfo> = &AUX_TABLE;
fn probe<'bound>(
adev: &'bound auxiliary::Device<Core<'_>>,
_info: &'bound Self::IdInfo,
) -> impl PinInit<Self::Data<'bound>, Error> + 'bound {
let data = try_pin_init!(NovaData { adev: adev.into() });
let drm = drm::UnregisteredDevice::<Self>::new(adev, Ok(()))?;
// SAFETY: `reg` is stored in `Nova` and dropped when the driver is unbound; it is
// never forgotten.
let reg = unsafe { drm::Registration::new(adev.as_ref(), drm, (), 0)? };
let drm = drm::UnregisteredDevice::<Self>::new(adev.as_ref(), data)?;
let drm = drm::Registration::new_foreign_owned(drm, adev.as_ref(), 0)?;
Ok(Nova { drm: drm.into() })
Ok(Nova {
drm: reg.device().into(),
_reg: reg,
})
}
}
#[vtable]
impl drm::Driver for NovaDriver {
type Data = NovaData;
type Data = ();
type RegistrationData<'a> = ();
type File = File;
type Object<Ctx: drm::DeviceContext> = gem::Object<NovaObject, Ctx>;
type Object = gem::Object<NovaObject>;
type ParentDevice<Ctx: DeviceContext> = auxiliary::Device<Ctx>;
const INFO: drm::DriverInfo = INFO;

View File

@ -4,7 +4,13 @@
use crate::gem::NovaObject;
use kernel::{
alloc::flags::*,
drm::{self, gem::BaseObject},
auxiliary,
device::Bound,
drm::{
self,
gem::BaseObject,
Registered, //
},
pci,
prelude::*,
uapi,
@ -23,13 +29,13 @@ fn open(_dev: &NovaDevice) -> Result<Pin<KBox<Self>>> {
impl File {
/// IOCTL: get_param: Query GPU / driver metadata.
pub(crate) fn get_param(
dev: &NovaDevice,
dev: &NovaDevice<Registered>,
_reg_data: &(),
getparam: &mut uapi::drm_nova_getparam,
_file: &drm::File<File>,
) -> Result<u32> {
let adev = &dev.adev;
let parent = adev.parent();
let pdev: &pci::Device = parent.try_into()?;
let adev: &auxiliary::Device<Bound> = dev.as_ref();
let pdev: &pci::Device<Bound> = adev.parent().try_into()?;
let value = match getparam.param as u32 {
uapi::NOVA_GETPARAM_VRAM_BAR_SIZE => pdev.resource_len(1)?,
@ -43,7 +49,8 @@ pub(crate) fn get_param(
/// IOCTL: gem_create: Create a new DRM GEM object.
pub(crate) fn gem_create(
dev: &NovaDevice,
dev: &NovaDevice<Registered>,
_reg_data: &(),
req: &mut uapi::drm_nova_gem_create,
file: &drm::File<File>,
) -> Result<u32> {
@ -56,7 +63,8 @@ pub(crate) fn gem_create(
/// IOCTL: gem_info: Query GEM metadata.
pub(crate) fn gem_info(
_dev: &NovaDevice,
_dev: &NovaDevice<Registered>,
_reg_data: &(),
req: &mut uapi::drm_nova_gem_info,
file: &drm::File<File>,
) -> Result<u32> {

View File

@ -2,7 +2,10 @@
use kernel::{
drm,
drm::{gem, gem::BaseObject, DeviceContext},
drm::{
gem,
gem::BaseObject, //
},
page,
prelude::*,
sync::aref::ARef,
@ -21,27 +24,20 @@ impl gem::DriverObject for NovaObject {
type Driver = NovaDriver;
type Args = ();
fn new<Ctx: DeviceContext>(
_dev: &NovaDevice<Ctx>,
_size: usize,
_args: Self::Args,
) -> impl PinInit<Self, Error> {
fn new(_dev: &NovaDevice, _size: usize, _args: Self::Args) -> impl PinInit<Self, Error> {
try_pin_init!(NovaObject {})
}
}
impl NovaObject {
/// Create a new DRM GEM object.
pub(crate) fn new<Ctx: DeviceContext>(
dev: &NovaDevice<Ctx>,
size: usize,
) -> Result<ARef<gem::Object<Self, Ctx>>> {
pub(crate) fn new(dev: &NovaDevice, size: usize) -> Result<ARef<gem::Object<Self>>> {
if size == 0 {
return Err(EINVAL);
}
let aligned_size = page::page_align(size).ok_or(EINVAL)?;
gem::Object::<Self, Ctx>::new(dev, aligned_size, ())
gem::Object::<Self>::new(dev, aligned_size, ())
}
/// Look up a GEM object handle for a `File` and return an `ObjectRef` for it.

View File

@ -5,10 +5,15 @@ config DRM_TYR
depends on DRM=y
depends on RUST
depends on ARM || ARM64 || COMPILE_TEST
depends on MMU
depends on !GENERIC_ATOMIC64 # for IOMMU_IO_PGTABLE_LPAE
depends on COMMON_CLK
depends on IOMMU_SUPPORT
default n
select IOMMU_IO_PGTABLE_LPAE
select RUST_DRM_GEM_SHMEM_HELPER
select RUST_DRM_GPUVM
select RUST_FW_LOADER_ABSTRACTIONS
help
Rust DRM driver for ARM Mali CSF-based GPUs.

View File

@ -6,8 +6,10 @@
OptionalClk, //
},
device::{
Bound,
Core,
Device, //
Device,
DeviceContext, //
},
dma::{
Device as DmaDevice,
@ -27,7 +29,7 @@
regulator::Regulator,
sizes::SZ_2M,
sync::{
aref::ARef,
Arc,
Mutex, //
},
time, //
@ -35,9 +37,11 @@
use crate::{
file::TyrDrmFileData,
gem::BoData,
fw::Firmware,
gem::Bo,
gpu,
gpu::GpuInfo,
mmu::Mmu,
regs::gpu_control::*, //
};
@ -46,18 +50,26 @@
pub(crate) struct TyrDrmDriver;
/// Convenience type alias for the DRM device type for this driver.
pub(crate) type TyrDrmDevice<Ctx = drm::Registered> = drm::Device<TyrDrmDriver, Ctx>;
pub(crate) type TyrDrmDevice<Ctx = drm::Normal> = drm::Device<TyrDrmDriver, Ctx>;
pub(crate) struct TyrPlatformDriver;
#[pin_data(PinnedDrop)]
pub(crate) struct TyrPlatformDriverData {
_device: ARef<TyrDrmDevice>,
pub(crate) struct TyrPlatformDriverData<'bound> {
_reg: drm::Registration<'bound, TyrDrmDriver>,
}
/// Data owned by the DRM [`Registration`].
///
/// This data can have references tied to the parent platform device binding scope
/// and is accessible only while the DRM device is registered with userspace.
#[pin_data]
pub(crate) struct TyrDrmDeviceData {
pub(crate) pdev: ARef<platform::Device>,
pub(crate) struct TyrDrmRegistrationData<'drm> {
/// Parent platform device.
pub(crate) pdev: &'drm platform::Device<Bound>,
/// Firmware sections.
pub(crate) fw: Firmware<'drm>,
#[pin]
clks: Mutex<Clocks>,
@ -65,9 +77,10 @@ pub(crate) struct TyrDrmDeviceData {
#[pin]
regulators: Mutex<Regulators>,
/// Some information on the GPU.
///
/// This is mainly queried by userspace, i.e.: Mesa.
/// GPU MMIO register mapping.
pub(crate) iomem: Arc<IoMem<'drm>>,
/// GPU information read from hardware during probe.
pub(crate) gpu_info: GpuInfo,
}
@ -97,7 +110,7 @@ fn issue_soft_reset(dev: &Device, iomem: &IoMem<'_>) -> Result {
impl platform::Driver for TyrPlatformDriver {
type IdInfo = ();
type Data<'bound> = TyrPlatformDriverData;
type Data<'bound> = TyrPlatformDriverData<'bound>;
const OF_ID_TABLE: Option<of::IdTable<Self::IdInfo>> = Some(&OF_TABLE);
fn probe<'bound>(
@ -116,7 +129,8 @@ fn probe<'bound>(
let sram_regulator = Regulator::<regulator::Enabled>::get(pdev.as_ref(), c"sram")?;
let request = pdev.io_request_by_index(0).ok_or(ENODEV)?;
let iomem = request.iomap_sized::<SZ_2M>()?;
let iomem = Arc::new(request.iomap_sized::<SZ_2M>()?, GFP_KERNEL)?;
issue_soft_reset(pdev.as_ref(), &iomem)?;
gpu::l2_power_on(pdev.as_ref(), &iomem)?;
@ -132,10 +146,23 @@ fn probe<'bound>(
// other threads of execution.
unsafe { pdev.dma_set_mask_and_coherent(DmaMask::try_new(pa_bits)?)? };
let platform: ARef<platform::Device> = pdev.into();
let unreg_dev = drm::UnregisteredDevice::<TyrDrmDriver>::new(pdev, Ok(()))?;
let data = try_pin_init!(TyrDrmDeviceData {
pdev: platform.clone(),
let mmu = Mmu::new(pdev.as_ref(), iomem.as_arc_borrow(), &gpu_info)?;
let firmware = Firmware::new(
pdev.as_ref(),
iomem.clone(),
&unreg_dev,
mmu.as_arc_borrow(),
&gpu_info,
)?;
firmware.boot()?;
let reg_data = pin_init!(TyrDrmRegistrationData {
pdev,
fw: firmware,
clks <- new_mutex!(Clocks {
core: core_clk,
stacks: stacks_clk,
@ -145,25 +172,23 @@ fn probe<'bound>(
_mali: mali_regulator,
_sram: sram_regulator,
}),
iomem,
gpu_info,
});
let tdev = drm::UnregisteredDevice::<TyrDrmDriver>::new(pdev.as_ref(), data)?;
let tdev = drm::driver::Registration::new_foreign_owned(tdev, pdev.as_ref(), 0)?;
// SAFETY: `reg` is stored in `TyrPlatformDriverData` and dropped when the driver is
// unbound; it is never forgotten.
let reg = unsafe { drm::Registration::new(pdev.as_ref(), unreg_dev, reg_data, 0)? };
let driver = TyrPlatformDriverData {
_device: tdev.into(),
};
let driver = TyrPlatformDriverData { _reg: reg };
// We need this to be dev_info!() because dev_dbg!() does not work at
// all in Rust for now, and we need to see whether probe succeeded.
dev_info!(pdev, "Tyr initialized correctly.\n");
dev_dbg!(pdev, "Tyr initialized correctly.");
Ok(driver)
}
}
#[pinned_drop]
impl PinnedDrop for TyrPlatformDriverData {
impl PinnedDrop for TyrPlatformDriverData<'_> {
fn drop(self: Pin<&mut Self>) {}
}
@ -179,9 +204,11 @@ fn drop(self: Pin<&mut Self>) {}
#[vtable]
impl drm::Driver for TyrDrmDriver {
type Data = TyrDrmDeviceData;
type Data = ();
type RegistrationData<'drm> = TyrDrmRegistrationData<'drm>;
type File = TyrDrmFileData;
type Object<R: drm::DeviceContext> = drm::gem::shmem::Object<BoData, R>;
type Object = Bo;
type ParentDevice<Ctx: DeviceContext> = platform::Device<Ctx>;
const INFO: drm::DriverInfo = INFO;
const FEAT_RENDER: bool = true;

View File

@ -1,7 +1,10 @@
// SPDX-License-Identifier: GPL-2.0 or MIT
use kernel::{
drm,
drm::{
self,
Registered, //
},
prelude::*,
uaccess::UserSlice,
uapi, //
@ -9,7 +12,8 @@
use crate::driver::{
TyrDrmDevice,
TyrDrmDriver, //
TyrDrmDriver,
TyrDrmRegistrationData, //
};
#[pin_data]
@ -28,14 +32,15 @@ fn open(_dev: &drm::Device<Self::Driver>) -> Result<Pin<KBox<Self>>> {
impl TyrDrmFileData {
pub(crate) fn dev_query(
ddev: &TyrDrmDevice,
_ddev: &TyrDrmDevice<Registered>,
reg_data: &TyrDrmRegistrationData<'_>,
devquery: &mut uapi::drm_panthor_dev_query,
_file: &TyrDrmFile,
) -> Result<u32> {
if devquery.pointer == 0 {
match devquery.type_ {
uapi::drm_panthor_dev_query_type_DRM_PANTHOR_DEV_QUERY_GPU_INFO => {
devquery.size = core::mem::size_of_val(&ddev.gpu_info) as u32;
devquery.size = core::mem::size_of_val(&reg_data.gpu_info) as u32;
Ok(0)
}
_ => Err(EINVAL),
@ -49,7 +54,7 @@ pub(crate) fn dev_query(
)
.writer();
writer.write(&ddev.gpu_info)?;
writer.write(&reg_data.gpu_info)?;
Ok(0)
}

321
drivers/gpu/drm/tyr/fw.rs Normal file
View File

@ -0,0 +1,321 @@
// SPDX-License-Identifier: GPL-2.0 or MIT
//! Firmware loading and management for Mali CSF GPU.
//!
//! This module handles loading the Mali GPU firmware binary, parsing it into sections,
//! and mapping those sections into the MCU's virtual address space. Each firmware section
//! has specific properties (read/write/execute permissions, cache modes) and must be loaded
//! at specific virtual addresses expected by the MCU.
//!
//! See [`Firmware`] for the main firmware management interface and [`Section`] for
//! individual firmware sections.
//!
//! [`Firmware`]: crate::fw::Firmware
//! [`Section`]: crate::fw::Section
use kernel::{
device::{
Bound,
Device, //
},
drm::{
gem::BaseObject, //
},
io::{
poll,
Io, //
},
num::Bounded,
prelude::*,
register,
str::CString,
sync::{
Arc,
ArcBorrow, //
},
time, //
};
use crate::{
driver::{
IoMem,
TyrDrmDevice, //
},
fw::parser::{
FwParser,
ParsedSection, //
},
gem,
gem::{
KernelBo,
KernelBoVaAlloc, //
},
gpu::GpuInfo,
mmu::Mmu,
regs::{
gpu_control::{
McuControlMode,
McuStatus,
GPU_ID,
MCU_CONTROL,
MCU_STATUS, //
}, //
job_control::{
JOB_IRQ_CLEAR,
JOB_IRQ_RAWSTAT, //
}, //
},
vm::Vm, //
};
mod parser;
pub(super) const CSF_MCU_SHARED_REGION_START: u32 = 0x04000000;
#[derive(Copy, Clone, Debug, PartialEq, Eq)]
#[repr(u8)]
pub(super) enum CacheMode {
None = 0,
Cached = 1,
UncachedCoherent = 2,
CachedCoherent = 3,
}
impl From<Bounded<u32, 2>> for CacheMode {
fn from(value: Bounded<u32, 2>) -> Self {
match value.get() {
0 => Self::None,
1 => Self::Cached,
2 => Self::UncachedCoherent,
3 => Self::CachedCoherent,
_ => unreachable!(),
}
}
}
impl From<CacheMode> for Bounded<u32, 2> {
fn from(value: CacheMode) -> Self {
Bounded::try_new(value as u32).unwrap()
}
}
register! {
#[allow(non_upper_case_globals)]
pub(super) SectionFlags(u32) @ 0x0 {
0:0 read => bool;
1:1 write => bool;
2:2 exec => bool;
4:3 cache_mode => CacheMode;
5:5 prot => bool;
30:30 shared => bool;
31:31 zero => bool;
}
}
impl SectionFlags {
const VALID_MASK: u32 = Self::READ_MASK
| Self::WRITE_MASK
| Self::EXEC_MASK
| Self::CACHE_MODE_MASK
| Self::PROT_MASK
| Self::SHARED_MASK
| Self::ZERO_MASK;
fn try_from_fw(value: u32) -> Result<Self> {
if value & !Self::VALID_MASK != 0 {
Err(EINVAL)
} else {
Ok(Self::from_raw(value))
}
}
}
/// A parsed section of the firmware binary.
struct Section<'drm> {
// Raw firmware section data for reset purposes
#[expect(dead_code)]
data: KVec<u8>,
// Keep the BO backing this firmware section so that both the
// GPU mapping and CPU mapping remain valid until the Section is dropped.
#[expect(dead_code)]
mem: gem::KernelBo<'drm>,
}
/// Loaded firmware with sections mapped into MCU VM.
pub(crate) struct Firmware<'drm> {
/// Iomem need to access registers.
iomem: Arc<IoMem<'drm>>,
/// MCU VM.
vm: Arc<Vm<'drm>>,
/// List of firmware sections.
#[expect(dead_code)]
sections: KVec<Section<'drm>>,
}
impl<'drm> Drop for Firmware<'drm> {
fn drop(&mut self) {
// Stop the MCU before releasing its firmware mappings and memory.
let _ = self.stop();
// AS slots retain a VM ref, we need to kill the circular ref manually.
self.vm.kill();
}
}
impl<'drm> Firmware<'drm> {
fn init_section_mem(dev: &Device, mem: &mut KernelBo<'drm>, data: &KVec<u8>) -> Result {
if data.is_empty() {
return Ok(());
}
let vmap = mem.bo().vmap::<0>()?;
let size = mem.bo().size();
if data.len() > size {
dev_err!(dev, "fw section {} bigger than BO {}", data.len(), size);
return Err(EINVAL);
}
for (i, &byte) in data.iter().enumerate() {
vmap.try_write8(byte, i)?;
}
Ok(())
}
fn request(ddev: &TyrDrmDevice, gpu_info: &GpuInfo) -> Result<kernel::firmware::Firmware> {
let gpu_id = GPU_ID::from_raw(gpu_info.gpu_id);
let path = CString::try_from_fmt(fmt!(
"arm/mali/arch{}.{}/mali_csffw.bin",
gpu_id.arch_major().get(),
gpu_id.arch_minor().get()
))?;
kernel::firmware::Firmware::request(&path, ddev.as_ref().as_ref())
}
fn load(
dev: &Device,
ddev: &TyrDrmDevice,
gpu_info: &GpuInfo,
) -> Result<(kernel::firmware::Firmware, KVec<ParsedSection>)> {
let fw = Self::request(ddev, gpu_info)?;
let mut parser = FwParser::new(dev, fw.data());
let parsed_sections = parser.parse()?;
Ok((fw, parsed_sections))
}
/// Load firmware and map sections into MCU VM.
pub(crate) fn new(
dev: &'drm Device<Bound>,
iomem: Arc<IoMem<'drm>>,
ddev: &TyrDrmDevice,
mmu: ArcBorrow<'_, Mmu<'drm>>,
gpu_info: &GpuInfo,
) -> Result<Firmware<'drm>> {
let vm = Vm::new(dev, ddev, mmu, gpu_info)?;
vm.activate()?;
let result = (|| {
let (fw, parsed_sections) = Self::load(dev, ddev, gpu_info)?;
let mut sections = KVec::new();
for parsed in parsed_sections {
let size = u64::from(parsed.va.end.checked_sub(parsed.va.start).ok_or(EINVAL)?);
let va = u64::from(parsed.va.start);
let mut mem = KernelBo::new(
ddev,
vm.clone(),
size,
KernelBoVaAlloc::Explicit(va),
parsed.vm_map_flags,
)?;
let section_start = parsed.data_range.start as usize;
let section_end = parsed.data_range.end as usize;
let mut data = KVec::new();
// Ensure that the firmware slice is not out of bounds.
let fw_data = fw.data();
let bytes = fw_data.get(section_start..section_end).ok_or(EINVAL)?;
data.extend_from_slice(bytes, GFP_KERNEL)?;
Self::init_section_mem(dev, &mut mem, &data)?;
sections.push(Section { data, mem }, GFP_KERNEL)?;
}
Ok(Firmware {
iomem,
vm: vm.clone(),
sections,
})
})();
if result.is_err() {
vm.kill();
}
result
}
pub(crate) fn boot(&self) -> Result {
let io = &self.iomem;
// Discard any stale global interrupt.
io.write_reg(JOB_IRQ_CLEAR::zeroed().with_glb(true));
io.write_reg(MCU_CONTROL::zeroed().with_req(McuControlMode::Auto));
if let Err(e) = poll::read_poll_timeout(
|| Ok((io.read(MCU_STATUS), io.read(JOB_IRQ_RAWSTAT))),
|(mcu_status, irq_rawstat)| {
mcu_status.value() == McuStatus::Enabled && irq_rawstat.glb()
},
time::Delta::from_millis(1),
time::Delta::from_millis(100),
) {
let status = io.read(MCU_STATUS);
dev_err!(
self.vm.dev(),
"MCU failed to boot, status: {:?}",
status.value()
);
return Err(e);
}
io.write_reg(JOB_IRQ_CLEAR::zeroed().with_glb(true));
Ok(())
}
fn stop(&self) -> Result {
let io = &self.iomem;
io.write_reg(MCU_CONTROL::zeroed().with_req(McuControlMode::Disable));
if let Err(e) = poll::read_poll_timeout(
|| Ok(io.read(MCU_STATUS)),
|status| status.value() == McuStatus::Disabled,
time::Delta::from_micros(10),
time::Delta::from_millis(100),
) {
let status = io.read(MCU_STATUS);
dev_err!(
self.vm.dev(),
"MCU failed to stop, status: {:?}",
status.value()
);
return Err(e);
}
Ok(())
}
}

View File

@ -0,0 +1,588 @@
// SPDX-License-Identifier: GPL-2.0 or MIT
//! Firmware binary parser for Mali CSF (Command Stream Frontend) GPU.
//!
//! This module implements a parser for the Mali GPU firmware binary format. The firmware
//! file contains a header followed by a sequence of entries, each describing how to load
//! firmware sections into the MCU (Microcontroller Unit) memory. The parser extracts section
//! metadata including:
//! - Virtual address ranges where sections should be mapped
//! - Data ranges (byte offsets) within the firmware binary
//! - Section flags (permissions, cache modes)
use core::{
mem::size_of,
ops::Range, //
};
use kernel::{
bits::bit_u32,
device::Device,
prelude::*,
sizes::SZ_4K, //
};
use crate::{
fw::{
CacheMode,
SectionFlags,
CSF_MCU_SHARED_REGION_START, //
},
vm::{
VmFlag,
VmMapFlags, //
}, //
};
/// A parsed firmware section ready for loading into MCU memory.
///
/// Represents a single firmware section extracted from the firmware binary, containing
/// all information needed to map the section's data into the MCU's virtual address space.
pub(super) struct ParsedSection {
/// Byte offset range within the firmware binary where this section's data resides.
pub(super) data_range: Range<u32>,
/// MCU virtual address range where this section should be mapped.
pub(super) va: Range<u32>,
/// Memory protection and caching flags for the mapping.
pub(super) vm_map_flags: VmMapFlags,
}
/// A bare-bones `std::io::Cursor<[u8]>` clone to keep track of the current position in the
/// firmware binary.
///
/// Provides methods to sequentially read primitive types and byte arrays from the firmware
/// binary while maintaining the current read position.
struct Cursor<'a> {
dev: &'a Device,
data: &'a [u8],
pos: usize,
}
impl<'a> Cursor<'a> {
fn new(dev: &'a Device, data: &'a [u8]) -> Self {
Self { dev, data, pos: 0 }
}
fn len(&self) -> usize {
self.data.len()
}
fn pos(&self) -> usize {
self.pos
}
/// Returns a view into the cursor's data.
///
/// This spawns a new cursor, leaving the current cursor unchanged.
fn view(&self, range: Range<usize>) -> Result<Cursor<'_>> {
if range.start < self.pos || range.end > self.data.len() {
dev_err!(
self.dev,
"Invalid cursor range {:?} for data of length {}",
range,
self.data.len()
);
Err(EINVAL)
} else {
Ok(Self {
dev: self.dev,
data: &self.data[range],
pos: 0,
})
}
}
/// Reads a slice of bytes from the current position and advances the cursor.
///
/// Returns an error if the read would exceed the data bounds.
fn read(&mut self, nbytes: usize) -> Result<&[u8]> {
let start = self.pos;
let end = start + nbytes;
if end > self.data.len() {
dev_err!(
self.dev,
"Invalid firmware file: read of size {} at position {} is out of bounds",
nbytes,
start,
);
return Err(EINVAL);
}
self.pos += nbytes;
Ok(&self.data[start..end])
}
/// Reads a little-endian `u8` from the current position and advances the cursor.
fn read_u8(&mut self) -> Result<u8> {
let bytes = self.read(size_of::<u8>())?;
Ok(bytes[0])
}
/// Reads a little-endian `u16` from the current position and advances the cursor.
fn read_u16(&mut self) -> Result<u16> {
let bytes: [u8; 2] = self
.read(size_of::<u16>())?
.try_into()
.map_err(|_| EINVAL)?;
Ok(u16::from_le_bytes(bytes))
}
/// Reads a little-endian `u32` from the current position and advances the cursor.
fn read_u32(&mut self) -> Result<u32> {
let bytes: [u8; 4] = self
.read(size_of::<u32>())?
.try_into()
.map_err(|_| EINVAL)?;
Ok(u32::from_le_bytes(bytes))
}
/// Advances the cursor position by the specified number of bytes.
///
/// Returns an error if the advance would exceed the data bounds.
fn advance(&mut self, nbytes: usize) -> Result {
if self.pos + nbytes > self.data.len() {
dev_err!(
self.dev,
"Invalid firmware file: advance of size {} at position {} is out of bounds",
nbytes,
self.pos,
);
return Err(EINVAL);
}
self.pos += nbytes;
Ok(())
}
}
/// Parser for Mali CSF GPU firmware binaries.
///
/// Parses the firmware binary format, extracting section metadata including virtual
/// address ranges, data offsets, and memory protection flags needed to load firmware
/// into the MCU's memory.
pub(super) struct FwParser<'a> {
cursor: Cursor<'a>,
}
impl<'a> FwParser<'a> {
/// Creates a new firmware parser for the given firmware binary data.
pub(super) fn new(dev: &'a Device, data: &'a [u8]) -> Self {
Self {
cursor: Cursor::new(dev, data),
}
}
/// Parses the firmware binary and returns a collection of parsed sections.
///
/// This method validates the firmware header and iterates through all entries
/// in the binary, extracting section information needed for loading.
pub(super) fn parse(&mut self) -> Result<KVec<ParsedSection>> {
let fw_header = self.parse_fw_header()?;
let header_end = fw_header.size as usize;
let mut parsed_sections = KVec::new();
while self.cursor.pos() < header_end {
let entry_section = self.parse_entry(header_end)?;
if let Some(inner) = entry_section.inner {
parsed_sections.push(inner, GFP_KERNEL)?;
}
}
if parsed_sections.is_empty() {
dev_err!(self.cursor.dev, "Firmware contains no loadable sections");
return Err(EINVAL);
}
Ok(parsed_sections)
}
fn parse_fw_header(&mut self) -> Result<FirmwareHeader> {
let fw_header: FirmwareHeader = match FirmwareHeader::new(&mut self.cursor) {
Ok(fw_header) => fw_header,
Err(e) => {
dev_err!(self.cursor.dev, "Invalid firmware file: {}", e.to_errno());
return Err(e);
}
};
if fw_header.size as usize > self.cursor.len() {
dev_err!(self.cursor.dev, "Firmware image is truncated");
return Err(EINVAL);
}
Ok(fw_header)
}
fn parse_entry(&mut self, header_end: usize) -> Result<EntrySection> {
let entry_start = self.cursor.pos();
let entry_header_end = entry_start
.checked_add(size_of::<EntryHeader>())
.ok_or(EINVAL)?;
if entry_header_end > header_end {
dev_err!(
self.cursor.dev,
"Firmware entry header at {:#x} exceeds header region ending at {:#x}",
entry_start,
header_end
);
return Err(EINVAL);
}
let entry_section = EntrySection {
entry_hdr: EntryHeader(self.cursor.read_u32()?),
inner: None,
};
let firmware_size = self.cursor.len();
let entry_size = entry_section.entry_hdr.size() as usize;
if self.cursor.pos() % size_of::<u32>() != 0
|| entry_size % size_of::<u32>() != 0
|| entry_size < size_of::<EntryHeader>()
{
dev_err!(
self.cursor.dev,
"Firmware entry isn't 32 bit aligned, offset={:#x} size={:#x}",
self.cursor.pos() - size_of::<u32>(),
entry_size
);
return Err(EINVAL);
}
let entry_end = entry_start.checked_add(entry_size).ok_or(EINVAL)?;
if entry_end > header_end {
dev_err!(
self.cursor.dev,
"Firmware entry at {:#x} extends beyond header region ending at {:#x}",
entry_start,
header_end
);
return Err(EINVAL);
}
let section_hdr_size = entry_size - size_of::<EntryHeader>();
let entry_section = {
let mut entry_cursor = self.cursor.view(self.cursor.pos()..entry_end)?;
match entry_section.entry_hdr.entry_type() {
Ok(EntryType::Iface) => Ok(EntrySection {
entry_hdr: entry_section.entry_hdr,
inner: Self::parse_section_entry(&mut entry_cursor, firmware_size)?,
}),
Ok(
EntryType::Config
| EntryType::FutfTest
| EntryType::TraceBuffer
| EntryType::TimelineMetadata
| EntryType::BuildInfoMetadata,
) => Ok(entry_section),
Err(_) => {
if entry_section.entry_hdr.optional() {
Ok(entry_section)
} else {
dev_err!(
self.cursor.dev,
"Failed to handle firmware entry type: {}",
entry_section.entry_hdr.entry_type_raw()
);
Err(EINVAL)
}
}
}
};
if entry_section.is_ok() {
self.cursor.advance(section_hdr_size)?;
}
entry_section
}
fn parse_section_entry(
entry_cursor: &mut Cursor<'_>,
firmware_size: usize,
) -> Result<Option<ParsedSection>> {
let section_hdr: SectionHeader = SectionHeader::new(entry_cursor)?;
if section_hdr.data.end < section_hdr.data.start {
dev_err!(
entry_cursor.dev,
"Firmware corrupted, data.end < data.start (0x{:x} < 0x{:x})",
section_hdr.data.end,
section_hdr.data.start
);
return Err(EINVAL);
}
if section_hdr.data.end as usize > firmware_size {
dev_err!(
entry_cursor.dev,
"Firmware data range {:#x}..{:#x} exceeds firmware size {:#x}",
section_hdr.data.start,
section_hdr.data.end,
firmware_size,
);
return Err(EINVAL);
}
if section_hdr.va.start as usize % SZ_4K != 0 || section_hdr.va.end as usize % SZ_4K != 0 {
dev_err!(
entry_cursor.dev,
"Firmware virtual address range {:#x}..{:#x} is not page aligned",
section_hdr.va.start,
section_hdr.va.end
);
return Err(EINVAL);
}
if section_hdr.section_flags.prot() {
dev_dbg!(
entry_cursor.dev,
"Firmware protected mode entry not supported, ignoring"
);
return Ok(None);
}
if section_hdr.va.start == CSF_MCU_SHARED_REGION_START
&& !section_hdr.section_flags.shared()
{
dev_err!(
entry_cursor.dev,
"Interface at 0x{:x} must be shared",
CSF_MCU_SHARED_REGION_START
);
return Err(EINVAL);
}
if section_hdr.va.is_empty() {
return Ok(None);
}
let mut vm_map_flags = VmMapFlags::empty();
if !section_hdr.section_flags.write() {
vm_map_flags |= VmFlag::Readonly;
}
if !section_hdr.section_flags.exec() {
vm_map_flags |= VmFlag::Noexec;
}
// TODO: As in Panthor, map coherent firmware sections uncached until the VM
// supports a coherent mapping attribute.
if section_hdr.section_flags.cache_mode() != CacheMode::Cached {
vm_map_flags |= VmFlag::Uncached;
}
Ok(Some(ParsedSection {
data_range: section_hdr.data.clone(),
va: section_hdr.va,
vm_map_flags,
}))
}
}
/// Firmware binary header containing version and size information.
///
/// The header is located at the beginning of the firmware binary and contains
/// a magic value for validation, version information, and the total size of
/// all structured headers that follow.
#[expect(dead_code)]
struct FirmwareHeader {
/// Magic value to check binary validity.
magic: u32,
/// Minor firmware version.
minor: u8,
/// Major firmware version.
major: u8,
/// Padding. Must be set to zero.
_padding1: u16,
/// Firmware version hash.
version_hash: u32,
/// Padding. Must be set to zero.
_padding2: u32,
/// Total size of all the structured data headers at beginning of firmware binary.
size: u32,
}
impl FirmwareHeader {
const FW_BINARY_MAGIC: u32 = 0xc3f13a6e;
const FW_BINARY_MAJOR_MAX: u8 = 0;
/// Reads and validates a firmware header from the cursor.
///
/// Verifies the magic value, version compatibility, and padding fields.
fn new(cursor: &mut Cursor<'_>) -> Result<Self> {
let magic = cursor.read_u32()?;
if magic != Self::FW_BINARY_MAGIC {
dev_err!(cursor.dev, "Invalid firmware magic");
return Err(EINVAL);
}
let minor = cursor.read_u8()?;
let major = cursor.read_u8()?;
if major > Self::FW_BINARY_MAJOR_MAX {
dev_err!(
cursor.dev,
"Unsupported firmware binary header version {}.{} (expected {}.x)",
major,
minor,
Self::FW_BINARY_MAJOR_MAX
);
return Err(EINVAL);
}
let padding1 = cursor.read_u16()?;
let version_hash = cursor.read_u32()?;
let padding2 = cursor.read_u32()?;
let size = cursor.read_u32()?;
if padding1 != 0 || padding2 != 0 {
dev_err!(
cursor.dev,
"Invalid firmware file: header padding is not zero"
);
return Err(EINVAL);
}
let fw_header = Self {
magic,
minor,
major,
_padding1: padding1,
version_hash,
_padding2: padding2,
size,
};
Ok(fw_header)
}
}
/// Firmware section header for loading binary sections into MCU memory.
#[derive(Debug)]
struct SectionHeader {
section_flags: SectionFlags,
/// MCU virtual range to map this binary section to.
va: Range<u32>,
/// References the data in the FW binary.
data: Range<u32>,
}
impl SectionHeader {
/// Reads and validates a section header from the cursor.
///
/// Parses section flags, virtual address range, and data range from the firmware binary.
fn new(cursor: &mut Cursor<'_>) -> Result<Self> {
let section_flags = SectionFlags::try_from_fw(cursor.read_u32()?)?;
let va_start = cursor.read_u32()?;
let va_end = cursor.read_u32()?;
let va = va_start..va_end;
if va.end < va.start {
dev_err!(
cursor.dev,
"Invalid firmware file: VA end precedes start at pos {}",
cursor.pos(),
);
return Err(EINVAL);
}
let data_start = cursor.read_u32()?;
let data_end = cursor.read_u32()?;
let data = data_start..data_end;
Ok(Self {
section_flags,
va,
data,
})
}
}
/// A firmware entry containing a header and optional parsed section data.
///
/// Represents a single entry in the firmware binary, which may contain loadable
/// section data or metadata that doesn't require loading.
struct EntrySection {
entry_hdr: EntryHeader,
inner: Option<ParsedSection>,
}
/// Header for a firmware entry, packed into a single u32.
///
/// The entry header encodes the entry type, size, and optional flag in a
/// 32-bit value with the following layout:
/// - Bits 0-7: Entry type
/// - Bits 8-15: Size in bytes
/// - Bit 31: Optional flag
struct EntryHeader(u32);
impl EntryHeader {
fn entry_type_raw(&self) -> u8 {
(self.0 & 0xff) as u8
}
fn entry_type(&self) -> Result<EntryType> {
let v = self.entry_type_raw();
EntryType::try_from(v)
}
fn optional(&self) -> bool {
self.0 & bit_u32(31) != 0
}
fn size(&self) -> u32 {
self.0 >> 8 & 0xff
}
}
#[derive(Clone, Copy, Debug)]
#[repr(u8)]
enum EntryType {
/// Host <-> FW interface.
Iface = 0,
/// FW config.
Config = 1,
/// Unit tests.
FutfTest = 2,
/// Trace buffer interface.
TraceBuffer = 3,
/// Timeline metadata interface.
TimelineMetadata = 4,
/// Metadata about how the FW binary was built.
BuildInfoMetadata = 6,
}
impl TryFrom<u8> for EntryType {
type Error = Error;
fn try_from(value: u8) -> Result<Self, Self::Error> {
match value {
0 => Ok(EntryType::Iface),
1 => Ok(EntryType::Config),
2 => Ok(EntryType::FutfTest),
3 => Ok(EntryType::TraceBuffer),
4 => Ok(EntryType::TimelineMetadata),
6 => Ok(EntryType::BuildInfoMetadata),
_ => Err(EINVAL),
}
}
}

View File

@ -4,17 +4,29 @@
//! This module provides buffer object (BO) management functionality using
//! DRM's GEM subsystem with shmem backing.
use core::ops::Range;
use kernel::{
drm::{
gem,
DeviceContext, //
drm::gem::{
self,
shmem, //
},
prelude::*, //
prelude::*,
sync::{
aref::ARef,
Arc, //
}, //
};
use crate::driver::{
TyrDrmDevice,
TyrDrmDriver, //
use crate::{
driver::{
TyrDrmDevice,
TyrDrmDriver, //
},
vm::{
Vm,
VmMapFlags, //
},
};
/// Tyr's DriverObject type for GEM objects.
@ -33,11 +45,116 @@ impl gem::DriverObject for BoData {
type Driver = TyrDrmDriver;
type Args = BoCreateArgs;
fn new<Ctx: DeviceContext>(
_dev: &TyrDrmDevice<Ctx>,
_size: usize,
args: BoCreateArgs,
) -> impl PinInit<Self, Error> {
fn new(_dev: &TyrDrmDevice, _size: usize, args: BoCreateArgs) -> impl PinInit<Self, Error> {
try_pin_init!(Self { flags: args.flags })
}
}
/// Type alias for Tyr GEM buffer objects.
pub(crate) type Bo = gem::shmem::Object<BoData>;
/// Creates a dummy GEM object to serve as the root of a GPUVM.
pub(crate) fn new_dummy_object(ddev: &TyrDrmDevice) -> Result<ARef<Bo>> {
let bo = Bo::new(
ddev,
4096,
shmem::ObjectConfig {
map_wc: true,
parent_resv_obj: None,
},
BoCreateArgs { flags: 0 },
)?;
Ok(bo)
}
/// Specifies how to choose a GPU virtual address for a [`KernelBo`].
/// An automatic VA allocation strategy will be added in the future.
pub(crate) enum KernelBoVaAlloc {
/// Explicit VA address specified by the caller.
Explicit(u64),
}
/// A kernel-owned buffer object with automatic GPU virtual address mapping.
///
/// This structure represents a buffer object that is created and managed entirely
/// by the kernel driver, as opposed to userspace-created GEM objects. It combines
/// a GEM object with automatic GPU virtual address (VA) space mapping and cleanup.
///
/// When dropped, the buffer is automatically unmapped from the GPU VA space.
pub(crate) struct KernelBo<'drm> {
/// The underlying GEM buffer object.
bo: ARef<Bo>,
/// The GPU VM this buffer is mapped into.
vm: Arc<Vm<'drm>>,
/// The GPU VA range occupied by this buffer.
va_range: Range<u64>,
}
impl<'drm> KernelBo<'drm> {
/// Creates a new kernel-owned buffer object and maps it into GPU VA space.
///
/// This function allocates a new shmem-backed GEM object and immediately maps
/// it into the specified GPU virtual memory space. The mapping is automatically
/// cleaned up when the [`KernelBo`] is dropped.
pub(crate) fn new(
ddev: &TyrDrmDevice,
vm: Arc<Vm<'drm>>,
size: u64,
va_alloc: KernelBoVaAlloc,
flags: VmMapFlags,
) -> Result<Self> {
if size == 0 {
dev_err!(vm.dev(), "Cannot create KernelBo with size 0");
return Err(EINVAL);
}
let KernelBoVaAlloc::Explicit(va) = va_alloc;
let bo_size = usize::try_from(size).map_err(|_| EOVERFLOW)?;
let va_end = va.checked_add(size).ok_or(EINVAL)?;
let bo = Bo::new(
ddev,
bo_size,
shmem::ObjectConfig {
map_wc: true,
parent_resv_obj: None,
},
BoCreateArgs { flags: 0 },
)?;
vm.map_bo_range(&bo, 0, size, va, flags)?;
Ok(KernelBo {
bo,
vm,
va_range: va..va_end,
})
}
pub(crate) fn bo(&self) -> &Bo {
&self.bo
}
}
impl Drop for KernelBo<'_> {
fn drop(&mut self) {
let va = self.va_range.start;
let size = self.va_range.end - self.va_range.start;
if let Err(e) = self.vm.unmap_range(va, size) {
// If unmap_range fails, it is still safe to drop the
// KernelBo and its ARef to the GEM buffer object because
// GPUVM also holds a reference to the GEM buffer object.
// The physical pages won't be freed or reallocated.
dev_err!(
self.vm.dev(),
"Failed to unmap KernelBo range {:#x}..{:#x}: {:?}",
self.va_range.start,
self.va_range.end,
e
);
}
}
}

120
drivers/gpu/drm/tyr/mmu.rs Normal file
View File

@ -0,0 +1,120 @@
// SPDX-License-Identifier: GPL-2.0 or MIT
//! Memory Management Unit (MMU) module.
//!
//! The GPU MMU provides a limited number of memory address spaces for use by command streams.
//! The MMU translates virtual addresses to physical addresses and manages memory configuration
//! and access permissions.
//!
//! This MMU module is essentially a locked wrapper around a [`SlotManager`] instance.
//! The [`SlotManager`] manages the assignment of virtual address spaces to hardware address-space
//! (AS) slots. MMU commands such as updates and flushes are carried out by the
//! [`AddressSpaceManager`] which actually writes to the MMU registers.
use core::ops::Range;
use kernel::{
device::{
Bound,
Device, //
},
new_mutex,
prelude::*,
sync::{
Arc,
ArcBorrow,
Mutex, //
}, //
};
use crate::{
driver::IoMem,
gpu::GpuInfo,
mmu::address_space::{
AddressSpaceManager,
VmAsData, //
},
regs::{
gpu_control::AS_PRESENT,
MAX_AS, //
},
slot::SlotManager, //
};
pub(crate) mod address_space;
pub(crate) type AsSlotManager<'drm> = SlotManager<AddressSpaceManager<'drm>, MAX_AS>;
/// Locked wrapper for carrying out virtual memory (VM) operations on the MMU.
#[pin_data]
pub(crate) struct Mmu<'drm> {
/// Slot Manager instance used to allocate hardware slots and write to MMU registers.
#[pin]
pub(crate) as_manager: Mutex<AsSlotManager<'drm>>,
}
impl<'drm> Mmu<'drm> {
/// Create an MMU component for this device.
pub(crate) fn new(
dev: &'drm Device<Bound>,
iomem: ArcBorrow<'_, IoMem<'drm>>,
gpu_info: &GpuInfo,
) -> Result<Arc<Mmu<'drm>>> {
let present = AS_PRESENT::from_raw(gpu_info.as_present).present().get();
let slot_count = present.count_ones().try_into()?;
let address_space_manager = AddressSpaceManager::new(dev, iomem.into(), present)?;
let as_slot_manager =
SlotManager::new(address_space_manager, slot_count).inspect_err(|e| {
dev_err!(
dev,
"Failed to initialize MMU slot manager with {} slots: {:?}",
slot_count,
e
);
})?;
let mmu_init = try_pin_init!(Self{
as_manager <- new_mutex!(as_slot_manager),
});
Arc::pin_init(mmu_init, GFP_KERNEL)
}
/// Assign a VM to an AS slot, provide a translation table,
/// and update the MMU to make the VM resident.
pub(crate) fn activate_vm(&self, vm_as_data: ArcBorrow<'_, VmAsData<'drm>>) -> Result {
self.as_manager.lock().activate_vm(vm_as_data)
}
/// Evict a VM from its AS slot and flush the MMU.
pub(crate) fn deactivate_vm(&self, vm_as_data: &VmAsData<'drm>) -> Result {
self.as_manager.lock().deactivate_vm(vm_as_data)
}
/// Flush MMU translation caches after a VM update.
pub(crate) fn flush_vm(&self, vm_as_data: &VmAsData<'drm>) -> Result {
self.as_manager.lock().flush_vm(vm_as_data)
}
/// Flags the start of a VM update.
///
/// If the VM is resident, any GPU access on the memory range being
/// updated will be blocked until `Mmu::end_vm_update()` is called.
/// This guarantees the atomicity of a VM update.
/// If the VM is not resident, this is a NOP.
pub(crate) fn start_vm_update(
&self,
vm_as_data: &VmAsData<'drm>,
region: &Range<u64>,
) -> Result {
self.as_manager.lock().start_vm_update(vm_as_data, region)
}
/// Flags the end of a VM update.
///
/// If the VM is resident, this will let GPU accesses on the updated
/// range go through, in case any of them were blocked.
/// If the VM is not resident, this is a NOP.
pub(crate) fn end_vm_update(&self, vm_as_data: &VmAsData<'drm>) -> Result {
self.as_manager.lock().end_vm_update(vm_as_data)
}
}

View File

@ -0,0 +1,511 @@
// SPDX-License-Identifier: GPL-2.0 or MIT
//! Address space module.
//!
//! This module handles the hardware interaction for MMU operations through
//! MMIO register access.
//!
use core::ops::Range;
use kernel::{
device::{
Bound,
Device, //
}, //
error::Result,
io::{
poll,
register::Array,
Io, //
},
iommu::pgtable::{
Config,
IoPageTable,
ARM64LPAES1, //
},
num::Bounded,
prelude::*,
sizes::{
SZ_2M,
SZ_4K, //
},
sync::{
Arc,
ArcBorrow,
LockedBy, //
},
time::Delta, //
};
use crate::{
driver::IoMem,
mmu::{
AsSlotManager,
Mmu, //
},
regs::{
mmu_control::mmu_as_control,
mmu_control::mmu_as_control::*,
MAX_AS, //
},
slot::{
LockedSeat,
Seat,
SlotOperations, //
}, //
};
/// Address space configuration values to be written to MMU registers.
#[derive(Clone, Copy)]
struct AddressSpaceConfig {
/// Translation configuration. Configures how the MMU walks the page table for this
/// address space.
transcfg: u64,
/// Translation table base address. The address of the page table.
transtab: u64,
/// Memory attributes such as cacheability.
memattr: u64,
}
/// Virtual memory (VM) address space data for use in MMU operations.
#[pin_data]
pub(crate) struct VmAsData<'drm> {
/// This address-space seat tracks this VM's binding to a hardware address space slot.
/// It can only be accessed when holding the `Mmu::as_manager` lock.
as_seat: LockedSeat<AddressSpaceManager<'drm>, MAX_AS>,
/// Virtual address bits for this address space.
va_bits: u8,
/// The page table which maps GPU virtual addresses to physical addresses for this VM.
#[pin]
pub(crate) page_table: IoPageTable<'drm, ARM64LPAES1>,
}
impl<'drm> VmAsData<'drm> {
/// Creates VM address space data by initializing all of its fields.
pub(crate) fn new<'a>(
mmu: &'a Mmu<'drm>,
dev: &'drm Device<Bound>,
va_bits: u32,
pa_bits: u32,
) -> impl pin_init::PinInit<VmAsData<'drm>, Error> + 'a {
let pt_config = Config {
quirks: 0,
pgsize_bitmap: SZ_4K | SZ_2M,
ias: va_bits,
oas: pa_bits,
coherent_walk: false,
};
let page_table_init = IoPageTable::new(dev, pt_config);
try_pin_init!(Self {
as_seat: LockedBy::new(&mmu.as_manager, Seat::NoSeat),
va_bits: va_bits as u8,
page_table <- page_table_init,
}? Error)
}
/// Computes the hardware configuration for this address space.
fn as_config(&self) -> Result<AddressSpaceConfig> {
let pt = &self.page_table;
// The hardware computes the valid input address range as:
// INA_BITS_VALID = min(HW_INA_BITS, 55 - INA_BITS)
// To configure our desired va_bits, we solve for INA_BITS:
// INA_BITS = 55 - va_bits
// This assumes HW_INA_BITS (hardware capability) >= va_bits.
let field = 55u64.checked_sub(self.va_bits.into()).ok_or(EINVAL)?;
let ina_bits =
match mmu_as_control::InaBits::try_from(Bounded::try_new(field).ok_or(EINVAL)?)? {
mmu_as_control::InaBits::Reset => return Err(EINVAL),
bits => bits,
};
let transcfg = mmu_as_control::TRANSCFG::zeroed()
.with_ptw_memattr(mmu_as_control::PtwMemattr::WriteBack)
.with_r_allocate(true)
.with_mode(mmu_as_control::AddressSpaceMode::Aarch64_4K)
.with_ina_bits(ina_bits)
.into_raw();
Ok(AddressSpaceConfig {
transcfg,
// SAFETY: The SlotManager holds an `Arc<VmAsData>` as SlotData while this
// TTBR is programmed and stores that Arc in the active slot before
// returning. Eviction flushes and disables the slot before releasing
// the Arc; if eviction fails, the slot retains it. Therefore the page
// table cannot be dropped while the GPU is using it.
transtab: unsafe { pt.ttbr() },
memattr: MEMATTR::from_mair(pt.mair()).into_raw(),
})
}
}
/// Coordinates all hardware-level address space operations through MMIO register
/// operations including enabling, disabling, flushing, and updating address spaces.
pub(crate) struct AddressSpaceManager<'drm> {
/// Parent device used for logging.
dev: &'drm Device<Bound>,
/// Memory-mapped I/O region for GPU register access.
iomem: Arc<IoMem<'drm>>,
/// Bitmask of present address space slots from GPU_AS_PRESENT register.
as_present: u32,
}
impl<'drm> AddressSpaceManager<'drm> {
/// Creates a new address space manager.
///
/// Initializes the manager with references to the platform device and
/// I/O memory region, along with the bitmask of available AS slots.
pub(super) fn new(
dev: &'drm Device<Bound>,
iomem: Arc<IoMem<'drm>>,
as_present: u32,
) -> Result<AddressSpaceManager<'drm>> {
if as_present.trailing_ones() != as_present.count_ones() {
dev_err!(
dev,
"Sparse AS_PRESENT mask is unsupported: {:#x}",
as_present
);
return Err(EINVAL);
}
Ok(Self {
dev,
iomem,
as_present,
})
}
/// Validates that an AS slot number is within range and present in hardware.
///
/// Checks that the slot index is less than [`MAX_AS`] and that
/// the corresponding bit is set in the `as_present` mask read from the GPU.
///
/// Returns [`EINVAL`] if the slot is out of range or not present in hardware.
fn validate_as_slot(&self, as_nr: usize) -> Result {
if as_nr >= MAX_AS {
dev_err!(
self.dev,
"AS slot {} out of valid range (max {})",
as_nr,
MAX_AS
);
return Err(EINVAL);
}
if (self.as_present & (1 << as_nr)) == 0 {
dev_err!(
self.dev,
"AS slot {} not present in hardware (AS_PRESENT={:#x})",
as_nr,
self.as_present
);
return Err(EINVAL);
}
Ok(())
}
/// Waits for an AS slot to become ready (not active).
///
/// Returns an error if polling times out after 10ms or if register access fails.
fn as_wait_ready(&self, as_nr: usize) -> Result {
let io = &*self.iomem;
let op = || {
let status_reg = STATUS::try_at(as_nr).ok_or(EINVAL)?;
Ok(io.read(status_reg))
};
let cond = |status: &STATUS| -> bool { !status.active_ext() };
poll::read_poll_timeout(op, cond, Delta::from_micros(50), Delta::from_millis(10))?;
Ok(())
}
/// Sends a command to an AS slot.
///
/// Returns an error if waiting for ready times out or if register write fails.
fn as_send_cmd(&mut self, as_nr: usize, cmd: MmuCommand) -> Result {
self.as_wait_ready(as_nr)?;
let io = &*self.iomem;
let command_reg = COMMAND::try_at(as_nr).ok_or(EINVAL)?;
io.write(command_reg, COMMAND::zeroed().with_command(cmd));
Ok(())
}
/// Sends a command to an AS slot and waits for completion.
///
/// Returns an error if sending the command fails or if waiting for completion times out.
fn as_send_cmd_and_wait(&mut self, as_nr: usize, cmd: MmuCommand) -> Result {
self.as_send_cmd(as_nr, cmd)?;
self.as_wait_ready(as_nr)?;
Ok(())
}
/// Enables an AS slot with the provided configuration.
///
/// Returns an error if the slot is invalid or if register writes/commands fail.
fn as_enable(&mut self, as_nr: usize, as_config: &AddressSpaceConfig) -> Result {
self.validate_as_slot(as_nr)?;
let io = &*self.iomem;
let transtab = as_config.transtab;
io.write(
TRANSTAB_LO::try_at(as_nr).ok_or(EINVAL)?,
TRANSTAB_LO::from_raw(transtab as u32),
);
io.write(
TRANSTAB_HI::try_at(as_nr).ok_or(EINVAL)?,
TRANSTAB_HI::from_raw((transtab >> 32) as u32),
);
let transcfg = as_config.transcfg;
io.write(
TRANSCFG_LO::try_at(as_nr).ok_or(EINVAL)?,
TRANSCFG_LO::from_raw(transcfg as u32),
);
io.write(
TRANSCFG_HI::try_at(as_nr).ok_or(EINVAL)?,
TRANSCFG_HI::from_raw((transcfg >> 32) as u32),
);
let memattr = as_config.memattr;
io.write(
MEMATTR_LO::try_at(as_nr).ok_or(EINVAL)?,
MEMATTR_LO::from_raw(memattr as u32),
);
io.write(
MEMATTR_HI::try_at(as_nr).ok_or(EINVAL)?,
MEMATTR_HI::from_raw((memattr >> 32) as u32),
);
self.as_send_cmd_and_wait(as_nr, MmuCommand::Update)?;
Ok(())
}
/// Disables an AS slot and clears its configuration.
///
/// Returns an error if the slot is invalid or if register writes/commands fail.
fn as_disable(&mut self, as_nr: usize) -> Result {
self.validate_as_slot(as_nr)?;
// Flush AS before disabling
self.as_send_cmd_and_wait(as_nr, MmuCommand::FlushMem)?;
let io = &*self.iomem;
io.write(
TRANSTAB_LO::try_at(as_nr).ok_or(EINVAL)?,
TRANSTAB_LO::from_raw(0),
);
io.write(
TRANSTAB_HI::try_at(as_nr).ok_or(EINVAL)?,
TRANSTAB_HI::from_raw(0),
);
io.write(
MEMATTR_LO::try_at(as_nr).ok_or(EINVAL)?,
MEMATTR_LO::from_raw(0),
);
io.write(
MEMATTR_HI::try_at(as_nr).ok_or(EINVAL)?,
MEMATTR_HI::from_raw(0),
);
let transcfg = TRANSCFG::zeroed()
.with_mode(AddressSpaceMode::Unmapped)
.into_raw();
io.write(
TRANSCFG_LO::try_at(as_nr).ok_or(EINVAL)?,
TRANSCFG_LO::from_raw(transcfg as u32),
);
io.write(
TRANSCFG_HI::try_at(as_nr).ok_or(EINVAL)?,
TRANSCFG_HI::from_raw((transcfg >> 32) as u32),
);
self.as_send_cmd_and_wait(as_nr, MmuCommand::Update)?;
Ok(())
}
/// Locks a region of the translation tables for an atomic update.
///
/// Programs the MMU [`LOCKADDR`] register for the given address space and issues
/// the lock command. The hardware rounds the requested range up to a
/// power-of-two region aligned to its size.
///
/// Returns an error if the slot is invalid or if register writes/commands fail.
fn as_start_update(&mut self, as_nr: usize, region: &Range<u64>) -> Result {
self.validate_as_slot(as_nr)?;
// Avoid both an empty range and an inverted range.
if region.start >= region.end {
return Err(EINVAL);
}
// The lock operates on full 64-byte cache lines of translation table entries.
// Since each translation table entry (TTE) is 8 bytes, a cache line has 8 TTEs.
// Since each TTE maps one page, the minimum locked region size will be 8 pages.
//
// With 4KiB pages (Aarch64_4K mode), the minimum locked region is 32KiB.
let lock_region_min_size: u64 = 4096 * 8;
// Count the number of trailing zero bits (zeros at the right/least-significant
// end of the binary representation). For a power-of-two value, this equals the
// base-2 exponent (e.g., 32 KiB = 2^15 → 15).
let lock_region_min_size_log2 = lock_region_min_size.trailing_zeros() as u8;
// XOR the first and last addresses to identify which bits differ between them.
// The highest set bit in the result determines the exponent of the smallest
// power-of-two region that can contain both addresses.
//
// Example:
// addr_xor = 0x1000 ^ 0x2FFF = 0x3FFF
// highest set bit in 0x3FFF is bit 13
// minimum region size = 2^(13 + 1) = 16 KiB
let addr_xor = region.start ^ (region.end - 1);
let region_size_log2 = 64 - addr_xor.leading_zeros() as u8;
let lock_region_log2 = core::cmp::max(region_size_log2, lock_region_min_size_log2);
let lock_region_size = 1u64.checked_shl(lock_region_log2.into()).ok_or(EINVAL)?;
// Align the LOCKADDR base address down to the lock region size (1 << lock_region_log2).
//
// The MMU ignores the low lock_region_log2 bits of LOCKADDR base, so ensure
// they are cleared in software to avoid ambiguity.
//
// Example:
// lock_region_log2 = 14 (16 KiB)
// region.start = 0x1000
// lockaddr_base = 0x1000 & ~(0x3FFF) = 0x0000
let lockaddr_base = region.start & !(lock_region_size - 1);
// The LOCKADDR size field encodes the lock region size as log2(size) - 1,
// per the hardware definition. For example, a 32 KiB region is encoded as 14
// because log2(32 KiB) = 15.
let lockaddr_size = lock_region_log2 - 1;
let io = &*self.iomem;
// The LOCKADDR base field stores address bits 63:12, so remove the low 12 bits
// before passing this value to the register macro helper.
// These bits are guaranteed to be zero anyway because of the minimum
// size of the locked region.
let lockaddr_base_field = lockaddr_base >> 12;
let lockaddr_val = LOCKADDR::zeroed()
.try_with_size(lockaddr_size)?
.try_with_base(lockaddr_base_field)?
.into_raw();
io.write(
LOCKADDR_LO::try_at(as_nr).ok_or(EINVAL)?,
LOCKADDR_LO::from_raw(lockaddr_val as u32),
);
io.write(
LOCKADDR_HI::try_at(as_nr).ok_or(EINVAL)?,
LOCKADDR_HI::from_raw((lockaddr_val >> 32) as u32),
);
self.as_send_cmd_and_wait(as_nr, MmuCommand::Lock)
}
/// Completes an atomic translation table update.
///
/// Returns an error if the slot is invalid or if the flush command fails.
fn as_end_update(&mut self, as_nr: usize) -> Result {
self.validate_as_slot(as_nr)?;
self.as_send_cmd_and_wait(as_nr, MmuCommand::FlushPt)?;
Ok(())
}
/// Flushes the translation table cache for an AS slot.
///
/// Returns an error if the slot is invalid or if the flush command fails.
fn as_flush(&mut self, as_nr: usize) -> Result {
self.validate_as_slot(as_nr)?;
self.as_send_cmd_and_wait(as_nr, MmuCommand::FlushPt)
}
}
impl<'drm> SlotOperations<MAX_AS> for AddressSpaceManager<'drm> {
/// VM address space data associated with a hardware slot.
type SlotData = Arc<VmAsData<'drm>>;
fn seat(slot_data: &Self::SlotData) -> &LockedSeat<Self, MAX_AS> {
&slot_data.as_seat
}
/// Activates a VM in a hardware slot.
fn activate(&mut self, slot_idx: usize, slot_data: &Self::SlotData) -> Result {
let as_config = slot_data.as_config()?;
self.as_enable(slot_idx, &as_config)
}
/// Evicts a VM from a hardware slot.
fn evict(&mut self, slot_idx: usize, _slot_data: &Self::SlotData) -> Result {
self.as_flush(slot_idx)?;
self.as_disable(slot_idx)?;
Ok(())
}
}
impl<'drm> AsSlotManager<'drm> {
/// Locks a region for translation table updates if the VM has an active slot.
pub(super) fn start_vm_update(
&mut self,
vm_as_data: &VmAsData<'drm>,
region: &Range<u64>,
) -> Result {
let seat = vm_as_data.as_seat.access(self);
match seat.slot() {
Some(slot) => {
let as_nr = slot as usize;
self.as_start_update(as_nr, region)
}
_ => Ok(()),
}
}
/// Completes translation table updates and unlocks the region.
pub(super) fn end_vm_update(&mut self, vm_as_data: &VmAsData<'drm>) -> Result {
let seat = vm_as_data.as_seat.access(self);
match seat.slot() {
Some(slot) => {
let as_nr = slot as usize;
self.as_end_update(as_nr)
}
_ => Ok(()),
}
}
/// Flushes the translation table cache if the VM has an active slot.
pub(super) fn flush_vm(&mut self, vm_as_data: &VmAsData<'drm>) -> Result {
let seat = vm_as_data.as_seat.access(self);
match seat.slot() {
Some(slot) => {
let as_nr = slot as usize;
self.as_flush(as_nr)
}
_ => Ok(()),
}
}
/// Activates a VM by assigning it to a hardware slot.
pub(super) fn activate_vm(&mut self, vm_as_data: ArcBorrow<'_, VmAsData<'drm>>) -> Result {
self.activate(vm_as_data.into())
}
/// Deactivates a VM by evicting it from its hardware slot.
pub(super) fn deactivate_vm(&mut self, vm_as_data: &VmAsData<'drm>) -> Result {
self.evict(&vm_as_data.as_seat)
}
}

View File

@ -25,7 +25,7 @@
//
// Nevertheless, it is useful to have most of them defined, like the C driver
// does.
#![allow(dead_code)]
#![expect(dead_code)]
/// Combine two 32-bit values into a single 64-bit value.
pub(crate) fn join_u64(lo: u32, hi: u32) -> u64 {
@ -45,20 +45,17 @@ pub(crate) fn read_u64_no_tearing(lo_read: impl Fn() -> u32, hi_read: impl Fn()
}
}
pub(crate) use mmu_control::mmu_as_control::MAX_AS;
/// These registers correspond to the GPU_CONTROL register page.
/// They are involved in GPU configuration and control.
pub(crate) mod gpu_control {
use core::convert::TryFrom;
use kernel::{
error::{
code::EINVAL,
Error, //
},
num::Bounded,
prelude::*,
register,
uapi, //
};
use pin_init::Zeroable;
register! {
/// GPU identification register.
@ -964,17 +961,14 @@ pub(crate) mod mmu_control {
///
/// This array contains 16 instances of the MMU_AS_CONTROL register page.
pub(crate) mod mmu_as_control {
use core::convert::TryFrom;
use kernel::{
error::{
code::EINVAL,
Error, //
},
num::Bounded,
prelude::*,
register, //
};
use pin_init::Zeroable;
/// Maximum number of hardware address space slots.
/// The actual number of slots available is usually lower.
pub(crate) const MAX_AS: usize = 16;
@ -1168,7 +1162,136 @@ fn from(val: MMU_MEMATTR_STAGE1) -> Self {
pub(crate) MEMATTR_HI(u32)[MAX_AS, stride = STRIDE] @ 0x240c {
31:0 value;
}
}
impl MEMATTR {
/// Outer cache-policy nibble indicating device memory.
const ARM_MAIR_DEVICE_MEMORY: u8 = 0x0;
/// In the ARM Architecture Reference Manual, the MAIR encoding for Normal memory
/// uses the format `0bxxRW` where:
/// - `W` (bit 0) = Write-Allocate policy
/// - `R` (bit 1) = Read-Allocate policy
/// E.g., `0b0011` would allow both read and write allocation on a cache miss.
///
/// ARM MAIR Write-Allocate bit (bit 0 of a cache policy nibble).
const ARM_MAIR_WRITE_ALLOCATE: u8 = 0x1;
/// ARM MAIR Read-Allocate bit (bit 1 of a cache policy nibble).
const ARM_MAIR_READ_ALLOCATE: u8 = 0x2;
/// Write-back policy bit. For cacheable encodings, it is necessary but not
/// sufficient to set bit 2 of the cache policy nibble. Bit 2 does not
/// definitively determine write back because bit 2 is also set in `0b0100`
/// which encodes Normal non-cacheable memory.
const ARM_MAIR_WRITE_BACK_BIT: u8 = 0x4;
/// Complete cache-policy nibble encoding for Normal Non-cacheable memory.
const ARM_MAIR_NON_CACHEABLE: u8 = 0x4;
/// Mask for the inner cache policy nibble in MAIR attribute bytes.
const ARM_MAIR_INNER_MASK: u8 = 0x0f;
/// Check if a MAIR attribute byte represents device memory.
///
/// Device memory (memory-mapped I/O, registers) cannot be cached because
/// reading and writing to this memory may have side effects.
fn is_device_memory(mair_attr: u8) -> bool {
// In AArch64 MAIR, outer nibble only is 0 for device memory.
(mair_attr >> 4) == Self::ARM_MAIR_DEVICE_MEMORY
}
/// Check if normal memory is fully write-back cacheable.
///
/// ARM MAIR has two cache policy levels (outer [7:4] and inner [3:0]).
/// For memory to be truly write-back, BOTH levels must have the write-back bit set.
/// If only one level is write-back, treat it as non-cacheable for GPU purposes.
fn is_writeback_cacheable(mair_attr: u8) -> bool {
let outer = mair_attr >> 4;
let inner = mair_attr & Self::ARM_MAIR_INNER_MASK;
outer != Self::ARM_MAIR_NON_CACHEABLE
&& inner != Self::ARM_MAIR_NON_CACHEABLE
&& (outer & Self::ARM_MAIR_WRITE_BACK_BIT) != 0
&& (inner & Self::ARM_MAIR_WRITE_BACK_BIT) != 0
}
// Helper to encode a MEMATTR attribute from its individual fields.
fn encode_attribute(
alloc_w: bool,
alloc_r: bool,
alloc_sel: AllocPolicySelect,
coherency: Coherency,
memory_type: MemoryType,
) -> MMU_MEMATTR_STAGE1 {
MMU_MEMATTR_STAGE1::zeroed()
.with_alloc_w(alloc_w)
.with_alloc_r(alloc_r)
.with_alloc_sel(alloc_sel)
.with_coherency(coherency)
.with_memory_type(memory_type)
}
/// Convert one MAIR attribute byte into a MEMATTR attribute.
// TODO: Add a `coherent` parameter like panthor's mair_to_memattr().
// For now, assume a non-coherent system and always encode write-back
// memory with MidgardInnerDomain coherency.
fn attribute_from_mair(mair_attr: u8) -> MMU_MEMATTR_STAGE1 {
// Device memory or non-write-back normal memory
if Self::is_device_memory(mair_attr) || !Self::is_writeback_cacheable(mair_attr) {
return Self::encode_attribute(
false,
false,
AllocPolicySelect::Alloc,
Coherency::MidgardInnerDomain,
MemoryType::NonCacheable,
);
}
// Write-back cacheable normal memory
let inner: u8 = mair_attr & Self::ARM_MAIR_INNER_MASK;
Self::encode_attribute(
(inner & Self::ARM_MAIR_WRITE_ALLOCATE) != 0,
(inner & Self::ARM_MAIR_READ_ALLOCATE) != 0,
AllocPolicySelect::Alloc,
Coherency::MidgardInnerDomain,
MemoryType::WriteBack,
)
}
/// Write one converted MAIR attribute into a corresponding MEMATTR slot.
fn with_encoded_attribute(self, index: usize, attr: MMU_MEMATTR_STAGE1) -> Self {
debug_assert!(index < 8);
let shift = index * 8;
let mask = !(0xffu64 << shift);
let raw = (self.into_raw() & mask) | ((u64::from(attr.into_raw())) << shift);
Self::from_raw(raw)
}
/// Convert an AArch64 MAIR value into the GPU MEMATTR register encoding.
///
/// Both MAIR and MEMATTR are 64-bit values with eight 8-bit memory
/// attribute entries, but the bits do not map directly. The GPU MEMATTR encoding
/// is less detailed than the MAIR encoding, so MAIR is converted to MEMATTR
/// conservatively as follows:
///
/// 1. Device memory, or Normal Memory that is not write-back cacheable, is encoded
/// as GPU `NonCacheable`
///
/// 2. Normal memory that is write-back cacheable is encoded as GPU `WriteBack`,
/// and the inner allocation hints are preserved.
pub(crate) fn from_mair(mair: u64) -> Self {
mair.to_le_bytes()
.into_iter()
.enumerate()
.fold(Self::zeroed(), |acc, (i, attr)| {
acc.with_encoded_attribute(i, Self::attribute_from_mair(attr))
})
}
}
register! {
/// Lock region address for each address space.
pub(crate) LOCKADDR(u64)[MAX_AS, stride = STRIDE] @ 0x2410 {
/// Lock region size.

404
drivers/gpu/drm/tyr/slot.rs Normal file
View File

@ -0,0 +1,404 @@
// SPDX-License-Identifier: GPL-2.0 or MIT
//! Slot management abstraction for limited hardware resources.
//!
//! This module provides a generic [`SlotManager`] that assigns limited hardware
//! slots to logical "seats". A seat represents an entity (such as a virtual memory
//! (VM) address space) that needs access to a hardware slot.
//!
//! The [`SlotManager`] tracks slot allocation using sequence numbers (seqno) to detect
//! when a seat's binding has been invalidated. When a seat requests activation,
//! the manager will either reuse the seat's existing slot (if still valid),
//! allocate a free slot (if any are available), or evict the oldest idle slot if any
//! slots are idle.
//!
//! Hardware-specific behavior is customized by implementing the [`SlotOperations`]
//! trait, which allows callbacks when slots are activated or evicted.
//!
//! This is currently used for managing address space slots in the GPU, and it will
//! also be used to manage Command Stream Group (CSG) interface slots in the future.
//!
//! [SlotOperations]: crate::slot::SlotOperations
//! [SlotManager]: crate::slot::SlotManager
use core::{
mem,
ops::{
Deref,
DerefMut, //
}, //
};
use kernel::{
prelude::*,
sync::LockedBy, //
};
/// Seat information.
///
/// This can't be accessed directly by the element embedding a `Seat`,
/// but is used by the generic slot manager logic to control residency
/// of a certain object on a hardware slot.
pub(crate) struct SeatInfo {
/// Slot used by this seat.
///
/// This index is only valid if the slot pointed to by this index
/// has its `SlotInfo::seqno` match `SeatInfo::seqno`. Otherwise,
/// it means the object has been evicted from the hardware slot,
/// and a new slot needs to be acquired to make this object
/// resident again.
slot: u8,
/// Sequence number encoding the last time this seat was active.
/// We also use it to check if a slot is still bound to a seat.
seqno: u64,
}
/// Seat state.
///
/// This is meant to be embedded in the object that wants to acquire
/// hardware slots. It also starts in the `Seat::NoSeat` state, and
/// the slot manager will change the object value when an active/evict
/// request is issued.
#[derive(Default)]
pub(crate) enum Seat {
#[expect(clippy::enum_variant_names)]
/// Resource is not resident.
///
/// All objects start with a seat in the `Seat::NoSeat` state. The seat also
/// gets back to that state if the user requests eviction. It
/// can also end up in that state next time an operation is done
/// on a `Seat::Idle` seat and the slot manager finds out this
/// object has been evicted from the slot.
#[default]
NoSeat,
/// Resource is actively used and resident.
///
/// When a seat is in the `Seat::Active` state, it can't be evicted, and the
/// slot pointed to by `SeatInfo::slot` is guaranteed to be reserved
/// for this object as long as the seat stays active.
Active(SeatInfo),
/// Resource is idle and might or might not be resident.
///
/// When a seat is in the`Seat::Idle` state, we can't know for sure if the
/// object is resident or evicted until the next request we issue
/// to the slot manager. This tells the slot manager it can
/// reclaim the underlying slot if needed.
/// In order for the hardware to use this object again, the seat
/// needs to be turned into an `Seat::Active` state again
/// with a `SlotManager::activate()` call.
Idle(SeatInfo),
}
impl Seat {
/// Get the slot index this seat is pointing to.
///
/// If the seat is not `Seat::Active` we can't trust the
/// `SeatInfo`. In that case `None` is returned, otherwise
/// `Some(SeatInfo::slot)` is returned.
pub(crate) fn slot(&self) -> Option<u8> {
match self {
Self::Active(info) => Some(info.slot),
_ => None,
}
}
}
/// Information related to a slot.
struct SlotInfo<D> {
/// Type specific data attached to a slot.
slot_data: D,
/// Sequence number from when this slot was last activated.
seqno: u64,
}
/// Slot state.
#[derive(Default)]
enum Slot<D> {
/// Slot is free.
#[default]
Free,
/// Slot is active.
Active(SlotInfo<D>),
/// Slot is idle.
Idle(SlotInfo<D>),
}
pub(crate) type LockedSeat<T, const MAX_SLOTS: usize> = LockedBy<Seat, SlotManager<T, MAX_SLOTS>>;
/// Trait describing the slot-related operations.
pub(crate) trait SlotOperations<const MAX_SLOTS: usize>: Sized {
/// Implementation-specific data associated with each slot.
type SlotData;
/// Returns the seat belonging to this slot data.
fn seat(slot_data: &Self::SlotData) -> &LockedSeat<Self, MAX_SLOTS>;
/// Called when a slot is being activated for a seat.
fn activate(&mut self, _slot_idx: usize, _slot_data: &Self::SlotData) -> Result {
Ok(())
}
/// Called when a slot is being evicted and freed.
fn evict(&mut self, _slot_idx: usize, _slot_data: &Self::SlotData) -> Result {
Ok(())
}
}
/// A generic slot manager that provides access to a limited number of hardware slots.
pub(crate) struct SlotManager<T: SlotOperations<MAX_SLOTS>, const MAX_SLOTS: usize> {
/// A specific implementation of the generic slot manager.
manager: T,
/// Number of slots actually available.
slot_count: usize,
/// Slot array used to track the state of each slot.
slots: [Slot<T::SlotData>; MAX_SLOTS],
/// Sequence number incremented each time a Seat is successfully activated
use_seqno: u64,
}
impl<T: SlotOperations<MAX_SLOTS>, const MAX_SLOTS: usize> SlotManager<T, MAX_SLOTS> {
/// Creates a specific instance of a slot manager.
pub(crate) fn new(manager: T, slot_count: usize) -> Result<Self> {
if slot_count == 0 {
return Err(EINVAL);
}
if slot_count > MAX_SLOTS {
return Err(EINVAL);
}
// Since the slot index is stored in SeatInfo as a u8, the maximum number of slots is 256.
if slot_count > u8::MAX as usize + 1 {
return Err(EINVAL);
}
Ok(Self {
manager,
slot_count,
slots: [const { Slot::Free }; MAX_SLOTS],
use_seqno: 1,
})
}
/// Records a newly activated slot for the given seat.
/// The slot manager takes ownership of the hardware-specific slot data.
fn record_active_slot(&mut self, slot_idx: usize, slot_data: T::SlotData) {
let cur_seqno = self.use_seqno;
*T::seat(&slot_data).access_mut(self) = Seat::Active(SeatInfo {
slot: slot_idx as u8,
seqno: cur_seqno,
});
self.slots[slot_idx] = Slot::Active(SlotInfo {
slot_data,
seqno: cur_seqno,
});
self.use_seqno += 1;
}
/// Reactivates an active/idle slot for a given seat without reprogramming the hardware.
/// The SlotManager reuses the existing slot_data. This ensures that the hardware-specific
/// information is not changed between subsequent uses. It also ensures that resources
/// owned by the existing slot_data remain alive while the hardware is configured to use them.
fn reactivate_slot(&mut self, slot_idx: usize, slot_data: &T::SlotData) -> Result {
let cur_seqno = self.use_seqno;
let mut slot_info = match mem::take(&mut self.slots[slot_idx]) {
Slot::Active(slot_info) | Slot::Idle(slot_info) => slot_info,
Slot::Free => {
*T::seat(slot_data).access_mut(self) = Seat::NoSeat;
return Err(EINVAL);
}
};
*T::seat(slot_data).access_mut(self) = Seat::Active(SeatInfo {
slot: slot_idx as u8,
seqno: cur_seqno,
});
slot_info.seqno = cur_seqno;
self.slots[slot_idx] = Slot::Active(slot_info);
self.use_seqno += 1;
Ok(())
}
/// Activates a slot for the given seat.
fn activate_slot(&mut self, slot_idx: usize, slot_data: T::SlotData) -> Result {
self.manager.activate(slot_idx, &slot_data)?;
self.record_active_slot(slot_idx, slot_data);
Ok(())
}
/// Finds a slot for the given seat. A free slot is preferred, but if none
/// are available, the oldest idle slot is evicted and reused. Otherwise, if
/// there are no free or idle slots, return [`EBUSY`].
fn allocate_slot(&mut self, slot_data: T::SlotData) -> Result {
let slots = &self.slots[..self.slot_count];
let mut idle_slot_idx = None;
let mut idle_slot_seqno: u64 = 0;
for (slot_idx, slot) in slots.iter().enumerate() {
match slot {
Slot::Free => {
return self.activate_slot(slot_idx, slot_data);
}
Slot::Idle(slot_info) => {
if idle_slot_idx.is_none() || slot_info.seqno < idle_slot_seqno {
idle_slot_idx = Some(slot_idx);
idle_slot_seqno = slot_info.seqno;
}
}
Slot::Active(_) => (),
}
}
match idle_slot_idx {
Some(slot_idx) => {
// Lazily evict idle slot just before it is reused.
if let Slot::Idle(slot_info) = &self.slots[slot_idx] {
self.manager.evict(slot_idx, &slot_info.slot_data)?;
mem::take(&mut self.slots[slot_idx]);
}
self.activate_slot(slot_idx, slot_data)
}
None => Err(EBUSY),
}
}
/// Converts an active slot and its seat to idle state.
fn idle_slot(&mut self, slot_idx: usize, locked_seat: &LockedSeat<T, MAX_SLOTS>) -> Result {
let slot = mem::take(&mut self.slots[slot_idx]);
self.slots[slot_idx] = match slot {
// If the slot was active, make it idle.
Slot::Active(slot_info) => Slot::Idle(slot_info),
// Preserve an already-idle slot.
Slot::Idle(slot_info) => Slot::Idle(slot_info),
// A free slot remains free.
Slot::Free => Slot::Free,
};
// If the seat was active, make it idle, or keep it idle if it was already idle.
*locked_seat.access_mut(self) = match locked_seat.access(self) {
Seat::Active(seat_info) | Seat::Idle(seat_info) => Seat::Idle(SeatInfo {
slot: seat_info.slot,
seqno: seat_info.seqno,
}),
Seat::NoSeat => Seat::NoSeat,
};
Ok(())
}
/// Evicts an active or idle slot: calls the eviction callback and marks the slot as free
/// and the seat as NoSeat.
fn evict_slot(&mut self, slot_idx: usize, locked_seat: &LockedSeat<T, MAX_SLOTS>) -> Result {
match &self.slots[slot_idx] {
Slot::Active(slot_info) | Slot::Idle(slot_info) => {
// If hardware eviction fails (e.g. times out), the slot retains
// its SlotData so that any resources still referenced by the hardware
// will remain alive. This prevents use-after-free errors.
self.manager.evict(slot_idx, &slot_info.slot_data)?;
mem::take(&mut self.slots[slot_idx]);
}
_ => (),
}
*locked_seat.access_mut(self) = Seat::NoSeat;
Ok(())
}
/// Checks that the seat state matches the slot's state.
/// If they don't match, the seat is stale and is reset to `NoSeat`.
fn check_seat(&mut self, locked_seat: &LockedSeat<T, MAX_SLOTS>) {
let (slot_idx, seat_seqno, is_active) = match locked_seat.access(self) {
Seat::Active(seat_info) => (seat_info.slot as usize, seat_info.seqno, true),
Seat::Idle(seat_info) => (seat_info.slot as usize, seat_info.seqno, false),
_ => return,
};
let valid = if is_active {
!kernel::warn_on!(!matches!(
&self.slots[slot_idx],
Slot::Active(slot_info) if slot_info.seqno == seat_seqno
))
} else {
matches!(
&self.slots[slot_idx],
Slot::Idle(slot_info) if slot_info.seqno == seat_seqno
)
};
if !valid {
*locked_seat.access_mut(self) = Seat::NoSeat;
}
}
/// Activates a resource on any available/reclaimable slot.
pub(crate) fn activate(&mut self, slot_data: T::SlotData) -> Result {
self.check_seat(T::seat(&slot_data));
// Copy out only the slot index so the borrow of slot_data ends here.
let slot_idx = match T::seat(&slot_data).access(self) {
Seat::Active(seat_info) | Seat::Idle(seat_info) => Some(seat_info.slot as usize),
Seat::NoSeat => None,
};
match slot_idx {
Some(slot_idx) => self.reactivate_slot(slot_idx, &slot_data),
None => self.allocate_slot(slot_data),
}
}
/// Flag a resource as idle. This method will be used for user VM support.
#[expect(dead_code)]
pub(crate) fn idle(&mut self, locked_seat: &LockedSeat<T, MAX_SLOTS>) -> Result {
self.check_seat(locked_seat);
if let Seat::Active(seat_info) = locked_seat.access(self) {
self.idle_slot(seat_info.slot as usize, locked_seat)?;
}
Ok(())
}
/// Evict a resource from its slot.
pub(crate) fn evict(&mut self, locked_seat: &LockedSeat<T, MAX_SLOTS>) -> Result {
self.check_seat(locked_seat);
match locked_seat.access(self) {
Seat::Active(seat_info) | Seat::Idle(seat_info) => {
let slot_idx = seat_info.slot as usize;
self.evict_slot(slot_idx, locked_seat)?;
}
_ => (),
}
Ok(())
}
}
impl<T: SlotOperations<MAX_SLOTS>, const MAX_SLOTS: usize> Deref for SlotManager<T, MAX_SLOTS> {
type Target = T;
fn deref(&self) -> &Self::Target {
&self.manager
}
}
impl<T: SlotOperations<MAX_SLOTS>, const MAX_SLOTS: usize> DerefMut for SlotManager<T, MAX_SLOTS> {
fn deref_mut(&mut self) -> &mut Self::Target {
&mut self.manager
}
}

View File

@ -9,9 +9,13 @@
mod driver;
mod file;
mod fw;
mod gem;
mod gpu;
mod mmu;
mod regs;
mod slot;
mod vm;
kernel::module_platform_driver! {
type: TyrPlatformDriver,

950
drivers/gpu/drm/tyr/vm.rs Normal file
View File

@ -0,0 +1,950 @@
// SPDX-License-Identifier: GPL-2.0 or MIT
//! GPU virtual memory management using the DRM GPUVM framework.
//!
//! This module manages GPU virtual address spaces, providing memory isolation and
//! the illusion of owning the entire virtual address (VA) range, similar to CPU virtual memory.
//! Each virtual memory (VM) area is backed by ARM64 LPAE Stage 1 page tables and can be
//! mapped into hardware address space (AS) slots for GPU execution.
use core::marker::PhantomData;
use core::ops::Range;
use kernel::{
device::{
Bound,
Device, //
},
drm::{
gem::BaseObject,
gpuvm::{
DriverGpuVm,
GpuVaAlloc,
GpuVm,
GpuVmBo,
OpMap,
OpMapRequest,
OpMapped,
OpRemap,
OpRemapped,
OpUnmap,
OpUnmapped,
UniqueRefGpuVm, //
}, //
},
fmt,
impl_flags,
io::PhysAddr,
iommu::pgtable::{
prot,
IoPageTable,
ARM64LPAES1, //
},
new_mutex,
prelude::*,
sizes::{
SZ_1G,
SZ_2M,
SZ_4K, //
},
sync::{
aref::ARef,
Arc,
ArcBorrow,
Mutex, //
},
uapi, //
};
use crate::{
driver::{
TyrDrmDevice,
TyrDrmDriver, //
},
gem,
gem::Bo,
gpu::GpuInfo,
mmu::{
address_space::VmAsData,
Mmu, //
},
regs::gpu_control::MMU_FEATURES,
};
impl_flags!(
/// Flags controlling virtual memory mapping behavior.
///
/// These flags control access permissions and caching behavior for GPU virtual
/// memory mappings.
#[derive(Debug, Clone, Default, Copy, PartialEq, Eq)]
pub(crate) struct VmMapFlags(u32);
/// Individual flags that can be combined in [`VmMapFlags`].
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(crate) enum VmFlag {
/// Map as read-only.
Readonly = uapi::drm_panthor_vm_bind_op_flags_DRM_PANTHOR_VM_BIND_OP_MAP_READONLY as u32,
/// Map as non-executable.
Noexec = uapi::drm_panthor_vm_bind_op_flags_DRM_PANTHOR_VM_BIND_OP_MAP_NOEXEC as u32,
/// Map as uncached.
Uncached = uapi::drm_panthor_vm_bind_op_flags_DRM_PANTHOR_VM_BIND_OP_MAP_UNCACHED as u32,
}
);
impl VmMapFlags {
/// Convert the flags to `pgtable::prot`.
fn to_prot(self) -> u32 {
let mut prot = 0;
if self.contains(VmFlag::Readonly) {
prot |= prot::READ;
} else {
prot |= prot::READ | prot::WRITE;
}
if self.contains(VmFlag::Noexec) {
prot |= prot::NOEXEC;
}
if !self.contains(VmFlag::Uncached) {
prot |= prot::CACHE;
}
prot
}
}
impl fmt::Display for VmMapFlags {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
let mut first = true;
if self.contains(VmFlag::Readonly) {
write!(f, "READONLY")?;
first = false;
}
if self.contains(VmFlag::Noexec) {
if !first {
write!(f, " | ")?;
}
write!(f, "NOEXEC")?;
first = false;
}
if self.contains(VmFlag::Uncached) {
if !first {
write!(f, " | ")?;
}
write!(f, "UNCACHED")?;
}
Ok(())
}
}
impl TryFrom<u32> for VmMapFlags {
type Error = Error;
fn try_from(value: u32) -> Result<Self, Self::Error> {
let valid = VmFlag::Readonly as u32 | VmFlag::Noexec as u32 | VmFlag::Uncached as u32;
if value & !valid != 0 {
return Err(EINVAL);
}
Ok(Self(value))
}
}
/// Arguments for a virtual memory map operation.
struct VmMapArgs<'drm> {
/// Access permissions and caching behavior for the mapping.
flags: VmMapFlags,
/// GEM buffer object registered with the GPUVM framework.
vm_bo: ARef<GpuVmBo<GpuVmData<'drm>>>,
/// Offset in bytes from the start of the buffer object.
bo_offset: u64,
}
/// Type of virtual memory operation.
enum VmOpType<'drm> {
/// Map a GEM buffer object into the virtual address space.
Map(VmMapArgs<'drm>),
/// Unmap a region from the virtual address space.
Unmap,
}
/// Preallocated resources needed to execute a VM operation.
///
/// VM operations may require allocating new GPUVA objects to track mappings.
/// To avoid allocation failures during the operation, preallocate the
/// maximum number of GPUVAs that might be needed.
struct VmOpResources<'drm> {
/// Preallocated GPUVA objects for remap operations.
///
/// Partial unmap requests or map requests overlapping existing mappings
/// will trigger a remap call, which needs to register up to three VA
/// objects (one for the new mapping, and two for the previous and next
/// mappings).
preallocated_gpuvas: [Option<GpuVaAlloc<GpuVmData<'drm>>>; 3],
}
/// Request to execute a virtual memory operation.
struct VmOpRequest<'drm> {
/// Request type.
op_type: VmOpType<'drm>,
/// Region of the virtual address space covered by this request.
region: Range<u64>,
}
/// Arguments for a page table map operation.
struct PtMapArgs {
/// Memory protection flags describing allowed accesses for this mapping.
///
/// This is directly derived from [`VmMapFlags`] via [`VmMapFlags::to_prot`].
prot: u32,
}
/// Type of page table operation.
enum PtOpType {
/// Map pages into the page table.
Map(PtMapArgs),
/// Unmap pages from the page table.
Unmap,
}
/// Context for updating the GPU page table.
///
/// This context is created when beginning a page table update operation and
/// automatically flushes changes when dropped. It ensures that the
/// Memory Management Unit (MMU) state is properly managed and Translation
/// Lookaside Buffer (TLB) entries are flushed.
pub(crate) struct PtUpdateContext<'ctx, 'drm> {
/// Device used for DMA-mapping GEM shmem SG tables.
dev: &'ctx Device<Bound>,
/// Page table.
pt: &'ctx IoPageTable<'drm, ARM64LPAES1>,
/// MMU manager.
mmu: &'ctx Mmu<'drm>,
/// Reference to the address space data to pass to the MMU functions.
as_data: &'ctx VmAsData<'drm>,
/// Region of the virtual address space covered by this request.
region: Range<u64>,
/// Operation type.
op_type: PtOpType,
/// Preallocated resources that can be used when executing the request.
resources: &'ctx mut VmOpResources<'drm>,
}
impl<'ctx, 'drm> PtUpdateContext<'ctx, 'drm> {
/// Creates a new page table update context.
///
/// This prepares the MMU for a page table update.
/// The context will automatically flush the TLB and
/// complete the update when dropped.
fn new(
dev: &'ctx Device<Bound>,
pt: &'ctx IoPageTable<'drm, ARM64LPAES1>,
mmu: &'ctx Mmu<'drm>,
as_data: &'ctx VmAsData<'drm>,
region: Range<u64>,
op_type: PtOpType,
resources: &'ctx mut VmOpResources<'drm>,
) -> Result<PtUpdateContext<'ctx, 'drm>> {
mmu.start_vm_update(as_data, &region)?;
Ok(Self {
dev,
pt,
mmu,
as_data,
region,
op_type,
resources,
})
}
/// Finds one of our pre-allocated VAs.
fn preallocated_gpuva(&mut self) -> Result<GpuVaAlloc<GpuVmData<'drm>>> {
self.resources
.preallocated_gpuvas
.iter_mut()
.find_map(|f| f.take())
.ok_or(EINVAL)
}
/// Returns an unused GPUVA object to the preallocated pool.
/// If the pool is already full, the unused allocation is simply dropped.
fn return_preallocated_gpuva(&mut self, gpuva: GpuVaAlloc<GpuVmData<'drm>>) {
if let Some(slot) = self
.resources
.preallocated_gpuvas
.iter_mut()
.find(|slot| slot.is_none())
{
*slot = Some(gpuva);
}
}
}
impl Drop for PtUpdateContext<'_, '_> {
fn drop(&mut self) {
if let Err(e) = self.mmu.end_vm_update(self.as_data) {
dev_err!(self.dev, "Failed to end VM update {:?}", e);
}
if let Err(e) = self.mmu.flush_vm(self.as_data) {
dev_err!(self.dev, "Failed to flush VM {:?}", e);
}
}
}
/// Driver implementation for the GPUVM framework.
///
/// Implements [`DriverGpuVm`] to provide VM operation callbacks (map, unmap, remap)
/// and associated types for buffer objects, virtual addresses, and contexts.
pub(crate) struct GpuVmData<'drm> {
_phantom: PhantomData<&'drm ()>,
}
/// GPU virtual address space.
///
/// Each VM can be mapped into a hardware address space slot.
#[pin_data]
pub(crate) struct Vm<'drm> {
/// Data referenced by an AS when the VM is active
as_data: Arc<VmAsData<'drm>>,
/// MMU manager.
mmu: Arc<Mmu<'drm>>,
/// Parent device used for DMA mapping and page-table operations.
dev: &'drm Device<Bound>,
/// DRM GPUVM core for managing virtual address space.
#[pin]
gpuvm_unique: Mutex<UniqueRefGpuVm<GpuVmData<'drm>>>,
/// Non-core part of the GPUVM. Can be used for stuff that doesn't modify the
/// internal mapping tree, like GpuVm::obtain()
gpuvm: ARef<GpuVm<GpuVmData<'drm>>>,
/// VA range for this VM.
va_range: Range<u64>,
}
impl<'drm> Vm<'drm> {
/// Creates a new GPU virtual address space.
///
/// The VM is initialized with a page table configured according to the GPU's
/// address translation capabilities and registered with the GPUVM framework.
pub(crate) fn new(
dev: &'drm Device<Bound>,
ddev: &TyrDrmDevice,
mmu: ArcBorrow<'_, Mmu<'drm>>,
gpu_info: &GpuInfo,
) -> Result<Arc<Vm<'drm>>> {
let mmu_features = MMU_FEATURES::from_raw(gpu_info.mmu_features);
let va_bits = mmu_features.va_bits().get();
let pa_bits = mmu_features.pa_bits().get();
let range = 0..(1u64 << va_bits);
let reserve_range = 0..0u64;
// dummy_obj is used to initialize the GPUVM tree.
let dummy_obj = gem::new_dummy_object(ddev).inspect_err(|e| {
dev_err!(dev, "Failed to create dummy GEM object: {:?}", e);
})?;
let gpuvm_unique = GpuVm::new::<Error, _>(
c"Tyr::GpuVm",
ddev,
&*dummy_obj,
range.clone(),
reserve_range,
GpuVmData::<'drm> {
_phantom: PhantomData::<&()>,
},
)
.inspect_err(|e| {
dev_err!(dev, "Failed to create GpuVm: {:?}", e);
})?;
let gpuvm = ARef::from(&*gpuvm_unique);
let as_data = Arc::pin_init(VmAsData::new(&mmu, dev, va_bits, pa_bits), GFP_KERNEL)?;
let vm = Arc::pin_init(
pin_init!(Self{
as_data,
dev,
mmu: mmu.into(),
gpuvm,
gpuvm_unique <- new_mutex!(gpuvm_unique),
va_range: range,
}),
GFP_KERNEL,
)?;
Ok(vm)
}
/// Returns the parent device used by this VM for DMA mapping and page-table operations.
pub(crate) fn dev(&self) -> &'drm Device<Bound> {
self.dev
}
/// Activate the VM in a hardware address space slot.
pub(crate) fn activate(&self) -> Result {
self.mmu
.activate_vm(self.as_data.as_arc_borrow())
.inspect_err(|e| {
dev_err!(self.dev, "Failed to activate VM: {:?}", e);
})
}
/// Deactivate the VM by evicting it from its address space slot.
fn deactivate(&self) -> Result {
self.mmu.deactivate_vm(&self.as_data).inspect_err(|e| {
dev_err!(self.dev, "Failed to deactivate VM: {:?}", e);
})
}
/// Kills the VM by deactivating it and unmapping all regions.
pub(crate) fn kill(&self) {
// TODO: Turn the VM into a state where it can't be used.
let _ = self.deactivate();
let _ = self
.unmap_range(self.va_range.start, self.va_range.end - self.va_range.start)
.inspect_err(|e| {
dev_err!(self.dev, "Failed to unmap range during deactivate: {:?}", e);
});
}
/// Executes a virtual memory operation.
///
/// This handles both map and unmap operations by coordinating between the
/// GPUVM framework and the hardware page table.
fn exec_op<'a>(
&self,
gpuvm_unique: &mut UniqueRefGpuVm<GpuVmData<'drm>>,
req: VmOpRequest<'drm>,
resources: &'a mut VmOpResources<'drm>,
) -> Result {
let pt = &self.as_data.page_table;
match req.op_type {
VmOpType::Map(args) => {
let mut pt_upd = PtUpdateContext::new(
self.dev,
pt,
&self.mmu,
&self.as_data,
req.region,
PtOpType::Map(PtMapArgs {
prot: args.flags.to_prot(),
}),
resources,
)?;
gpuvm_unique.sm_map(OpMapRequest {
addr: pt_upd.region.start,
range: pt_upd.region.end - pt_upd.region.start,
gem_offset: args.bo_offset,
vm_bo: &args.vm_bo,
context: &mut pt_upd,
})
//PtUpdateContext drops here flushing the page table
}
VmOpType::Unmap => {
let mut pt_upd = PtUpdateContext::new(
self.dev,
pt,
&self.mmu,
&self.as_data,
req.region,
PtOpType::Unmap,
resources,
)?;
gpuvm_unique.sm_unmap(
pt_upd.region.start,
pt_upd.region.end - pt_upd.region.start,
&mut pt_upd,
)
//PtUpdateContext drops here flushing the page table
}
}
}
/// Maps a GEM buffer object range into the VM at the specified virtual address.
///
/// This creates a mapping from GPU virtual address `va` to the physical pages
/// backing the GEM object, starting at `bo_offset` bytes into the object and
/// spanning `map_size` bytes. The mapping respects the access permissions and
/// caching behavior specified in `flags`.
pub(crate) fn map_bo_range(
&self,
bo: &Bo,
bo_offset: u64,
map_size: u64,
va: u64,
flags: VmMapFlags,
) -> Result {
if map_size == 0
|| va % SZ_4K as u64 != 0
|| bo_offset % SZ_4K as u64 != 0
|| map_size % SZ_4K as u64 != 0
{
return Err(EINVAL);
}
let bo_size = u64::try_from(bo.size()).map_err(|_| EOVERFLOW)?;
let bo_end = bo_offset.checked_add(map_size).ok_or(EINVAL)?;
if bo_end > bo_size {
dev_err!(
self.dev,
"BO mapping range {:#x}..{:#x} exceeds BO size {:#x}",
bo_offset,
bo_end,
bo_size
);
return Err(EINVAL);
}
let va_end: u64 = va.checked_add(map_size).ok_or(EINVAL)?;
let req = VmOpRequest {
op_type: VmOpType::Map(VmMapArgs {
vm_bo: self.gpuvm.obtain(bo, ())?,
flags,
bo_offset,
}),
region: va..va_end,
};
let mut resources = VmOpResources {
preallocated_gpuvas: [
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
],
};
let result = {
let mut gpuvm_unique = self.gpuvm_unique.lock();
self.exec_op(gpuvm_unique.as_mut().get_mut(), req, &mut resources)
};
// We flush the defer cleanup list now. Things will be different in
// the asynchronous VM_BIND path, where we want the cleanup to
// happen outside the DMA signalling path.
self.gpuvm.deferred_cleanup();
result
}
/// Unmaps a virtual address range from the VM.
///
/// This removes any existing mappings in the specified range, freeing the
/// virtual address space for reuse.
pub(crate) fn unmap_range(&self, va: u64, size: u64) -> Result {
if size == 0 || va % SZ_4K as u64 != 0 || size % SZ_4K as u64 != 0 {
return Err(EINVAL);
}
let end = va.checked_add(size).ok_or(EINVAL)?;
if va < self.va_range.start || end > self.va_range.end {
dev_err!(
self.dev,
"Unmap range {:#x}..{:#x} exceeds VM range {:#x}..{:#x}",
va,
end,
self.va_range.start,
self.va_range.end
);
return Err(EINVAL);
}
let req = VmOpRequest {
op_type: VmOpType::Unmap,
region: va..end,
};
let full_vm = va == self.va_range.start && end == self.va_range.end;
let mut resources = VmOpResources {
preallocated_gpuvas: if full_vm {
// Unmapping the entire VM cannot split an existing mapping,
// so no GPUVA objects are needed for remap operations.
[None, None, None]
} else {
[
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
]
},
};
let result = {
let mut gpuvm_unique = self.gpuvm_unique.lock();
self.exec_op(gpuvm_unique.as_mut().get_mut(), req, &mut resources)
};
// We flush the defer cleanup list now. Things will be different in
// the asynchronous VM_BIND path, where we want the cleanup to
// happen outside the DMA signalling path.
self.gpuvm.deferred_cleanup();
result
}
}
impl<'drm> DriverGpuVm for GpuVmData<'drm> {
type Driver = TyrDrmDriver;
type Object = Bo;
type VmBoData = ();
type VaData = ();
type SmContext<'ctx>
= PtUpdateContext<'ctx, 'drm>
where
Self: 'ctx;
/// Create a new mapping.
fn sm_step_map<'op>(
&mut self,
op: OpMap<'op, Self>,
context: &mut Self::SmContext<'_>,
) -> Result<OpMapped<'op, Self>, Error> {
let start_iova = op.addr();
let mut iova = start_iova;
let mut bytes_left_to_map = op.length();
let mut gem_offset = op.gem_offset();
// Make sure that the end of the requested GEM range doesn't run past the
// end of the GEM buffer itself.
let gem_range_end = op.gem_offset().checked_add(op.length()).ok_or(EINVAL)?;
if gem_range_end > op.obj().size() as u64 {
dev_err!(
context.dev,
"Requested GEM range ends at {} which is beyond the GEM buffer size {}",
gem_range_end,
op.obj().size()
);
return Err(EINVAL);
}
let sgt = op.obj().sg_table(context.dev).inspect_err(|e| {
dev_err!(context.dev, "Failed to get sg_table: {:?}", e);
})?;
let prot = match &context.op_type {
PtOpType::Map(args) => args.prot,
_ => {
return Err(EINVAL);
}
};
for sgt_entry in sgt.iter() {
// Expressly convert to u64 to work with arm 32-bit builds.
#[allow(clippy::useless_conversion)]
let mut paddr = u64::from(sgt_entry.dma_address());
#[allow(clippy::useless_conversion)]
let mut sgt_entry_length = u64::from(sgt_entry.dma_len());
if bytes_left_to_map == 0 {
break;
}
if gem_offset > 0 {
// Skip the entire SGT entry if the gem_offset exceeds its length.
let skip = u64::min(sgt_entry_length, gem_offset);
paddr += skip;
sgt_entry_length -= skip;
gem_offset -= skip;
}
if sgt_entry_length == 0 {
continue;
}
let len = u64::min(sgt_entry_length, bytes_left_to_map);
let segment_mapped = match pt_map(context.dev, context.pt, iova, paddr, len, prot) {
Ok(segment_mapped) => segment_mapped,
Err(e) => {
// clean up any successful mappings from previous SGT entries.
let total_mapped = iova - start_iova;
if total_mapped > 0 {
let _ = pt_unmap(
context.dev,
context.pt,
start_iova..(start_iova + total_mapped),
);
}
return Err(e);
}
};
bytes_left_to_map -= segment_mapped;
iova += segment_mapped;
}
if bytes_left_to_map != 0 {
let total_mapped = iova - start_iova;
if total_mapped > 0 {
let _ = pt_unmap(context.dev, context.pt, start_iova..iova);
}
dev_err!(
context.dev,
"SG table is too small for requested mapping: {} bytes remain",
bytes_left_to_map
);
return Err(EINVAL);
}
let gpuva = context.preallocated_gpuva()?;
let op = op.insert(gpuva, pin_init::init_zeroed());
Ok(op)
}
/// Indicates that an existing mapping should be removed.
fn sm_step_unmap<'op>(
&mut self,
op: OpUnmap<'op, Self>,
context: &mut Self::SmContext<'_>,
) -> Result<OpUnmapped<'op, Self>, Error> {
let start_iova = op.va().addr();
let length = op.va().length();
let region = start_iova..(start_iova + length);
pt_unmap(context.dev, context.pt, region.clone()).inspect_err(|e| {
dev_err!(
context.dev,
"Failed to unmap region {:#x}..{:#x}: {:?}",
region.start,
region.end,
e
);
})?;
let (op_unmapped, _va_removed) = op.remove();
Ok(op_unmapped)
}
/// Split up an existing mapping.
fn sm_step_remap<'op>(
&mut self,
op: OpRemap<'op, Self>,
context: &mut Self::SmContext<'_>,
) -> Result<OpRemapped<'op, Self>, Error> {
let unmap_start = if let Some(prev) = op.prev() {
prev.addr() + prev.length()
} else {
op.va_to_unmap().addr()
};
let unmap_end = if let Some(next) = op.next() {
next.addr()
} else {
op.va_to_unmap().addr() + op.va_to_unmap().length()
};
let unmap_length = unmap_end - unmap_start;
if unmap_length > 0 {
let region = unmap_start..(unmap_start + unmap_length);
pt_unmap(context.dev, context.pt, region.clone()).inspect_err(|e| {
dev_err!(
context.dev,
"Failed to unmap remap region {:#x}..{:#x}: {:?}",
region.start,
region.end,
e
);
})?;
}
let prev_va = context.preallocated_gpuva()?;
let next_va = context.preallocated_gpuva()?;
let (op_remapped, remap_ret) = op.remap(
[prev_va, next_va],
pin_init::init_zeroed(),
pin_init::init_zeroed(),
);
if let Some(unused_va) = remap_ret.unused_va {
context.return_preallocated_gpuva(unused_va);
}
Ok(op_remapped)
}
}
/// This function selects the largest supported block size (currently 4KB or 2MB)
/// that can be used for a mapping at the given address and size, respecting alignment constraints.
///
/// We can map multiple pages at once but we can't exceed the size of the
/// table entry itself. So, if mapping 4KB pages, figure out how many pages
/// can be mapped before we hit the 2MB boundary. Or, if mapping 2MB pages,
/// figure out how many pages can be mapped before hitting the 1GB boundary
/// Returns the page size (4KB or 2MB) and the number of pages that can be mapped at that size.
fn get_pgsize(addr: u64, size: u64) -> (u64, u64) {
// Get the distance to the next boundary of 2MB block
let blk_offset_2m = addr.wrapping_neg() % (SZ_2M as u64);
// Use 4K blocks if the address is not 2MB aligned, or we have less than 2MB to map
if blk_offset_2m != 0 || size < SZ_2M as u64 {
let pgcount = if blk_offset_2m == 0 {
size / SZ_4K as u64
} else {
u64::min(blk_offset_2m, size) / SZ_4K as u64
};
return (SZ_4K as u64, pgcount);
}
let blk_offset_1g = addr.wrapping_neg() % (SZ_1G as u64);
let blk_offset = if blk_offset_1g == 0 {
SZ_1G as u64
} else {
blk_offset_1g
};
let pgcount = u64::min(blk_offset, size) / SZ_2M as u64;
(SZ_2M as u64, pgcount)
}
/// Maps a physical address range into the page table at the specified virtual address.
///
/// This function maps `len` bytes of physical memory starting at `paddr` to the
/// virtual address `iova`, using the protection flags specified in `prot`. It
/// automatically selects optimal page sizes to minimize page table overhead.
///
/// If the mapping fails partway through, all successfully mapped pages are
/// unmapped before returning an error.
///
/// Returns the number of bytes successfully mapped.
fn pt_map(
dev: &Device,
pt: &IoPageTable<'_, ARM64LPAES1>,
iova: u64,
paddr: u64,
len: u64,
prot: u32,
) -> Result<u64> {
let mut segment_mapped = 0u64;
while segment_mapped < len {
let remaining = len - segment_mapped;
let curr_iova = iova + segment_mapped;
let curr_paddr = paddr + segment_mapped;
let (pgsize, pgcount) = get_pgsize(curr_iova | curr_paddr, remaining);
// On 32-bit systems, usize is only 32 bits, so check that
// the iova can be converted without truncation.
let curr_iova = match usize::try_from(curr_iova) {
Ok(curr_iova) => curr_iova,
Err(_) => {
dev_err!(
dev,
"curr_iova {:#x} cannot be represented as usize (max {:#x})",
curr_iova,
usize::MAX
);
if segment_mapped > 0 {
let _ = pt_unmap(dev, pt, iova..(iova + segment_mapped));
}
return Err(EOVERFLOW);
}
};
// SAFETY:
// No other io-pgtable operation can currently access this range because Tyr holds
// the gpuvm_unique mutex for the entire sm_map() operation.
// The addresses being mapped won't overlap any existing mappings in this
// page table because drm_gpuvm_sm_map() checks each requested mapping and either unmaps
// or remaps any overlap before creating the new mapping.
let (mapped, result) = unsafe {
pt.map_pages(
curr_iova,
curr_paddr as PhysAddr,
pgsize as usize,
pgcount as usize,
prot,
GFP_KERNEL,
)
};
if let Err(e) = result {
// If map_pages fails, mapped will be zero because the ARM LPAE backend
// only updates the mapped value after the entire request succeeds.
dev_err!(dev, "pt.map_pages failed at iova {:#x}: {:?}", curr_iova, e);
if segment_mapped > 0 {
let _ = pt_unmap(dev, pt, iova..(iova + segment_mapped));
}
return Err(e);
}
if mapped == 0 {
dev_err!(dev, "Failed to map any pages at iova {:#x}", curr_iova);
if segment_mapped > 0 {
let _ = pt_unmap(dev, pt, iova..(iova + segment_mapped));
}
return Err(ENOMEM);
}
segment_mapped += mapped as u64;
}
Ok(segment_mapped)
}
/// Unmaps a virtual address range from the page table.
///
/// This function removes all page table entries in the specified range,
/// automatically handling different page sizes that may be present.
fn pt_unmap(dev: &Device, pt: &IoPageTable<'_, ARM64LPAES1>, range: Range<u64>) -> Result {
let mut iova = range.start;
let mut bytes_left_to_unmap = range.end - range.start;
while bytes_left_to_unmap > 0 {
// It is fine to use just the iova to determine the page size
// because if the actual mapping was represented with smaller page sizes,
// (e.g. because the physical address was not 2MiB aligned)
// the ARM LPAE backend will notice and handle the lower-level table correctly.
let (pgsize, pgcount) = get_pgsize(iova, bytes_left_to_unmap);
// On 32-bit systems, usize is only 32 bits, so check that
// the iova can be converted without truncation.
let iova_usize = usize::try_from(iova).map_err(|_| {
dev_err!(
dev,
"IOVA {:#x} cannot be represented as usize (max {:#x})",
iova,
usize::MAX
);
EOVERFLOW
})?;
// SAFETY:
// No other io-pgtable operation can currently access this range because Tyr holds
// the gpuvm_unique mutex for the entire sm_unmap() operation.
// We know that this page table has one or more consecutive mappings
// starting at `iova` with the total size of `pgcount * pgsize` because
// gpuvm callbacks provide exactly the range that was previously mapped.
let unmapped = unsafe { pt.unmap_pages(iova_usize, pgsize as usize, pgcount as usize) };
if unmapped == 0 {
dev_err!(dev, "Failed to unmap any bytes at iova {:#x}", iova_usize);
return Err(EINVAL);
}
bytes_left_to_unmap -= unmapped as u64;
iova += unmapped as u64;
}
Ok(())
}

1
drivers/gpu/nova-core/.gitignore vendored Normal file
View File

@ -0,0 +1 @@
exports_nova_core_generated.h

View File

@ -1,4 +1,3 @@
# SPDX-License-Identifier: GPL-2.0
obj-$(CONFIG_NOVA_CORE) += nova-core.o
nova-core-y := nova_core.o
# nova-core is built from drivers/gpu/Makefile.
# nova_core.o (rust-analyzer marker - DO NOT REMOVE).

View File

@ -1,329 +0,0 @@
// SPDX-License-Identifier: GPL-2.0
//! Bitfield library for Rust structures
//!
//! Support for defining bitfields in Rust structures. Also used by the [`register!`] macro.
/// Defines a struct with accessors to access bits within an inner unsigned integer.
///
/// # Syntax
///
/// ```rust
/// use nova_core::bitfield;
///
/// #[derive(Debug, Clone, Copy, Default)]
/// enum Mode {
/// #[default]
/// Low = 0,
/// High = 1,
/// Auto = 2,
/// }
///
/// impl TryFrom<u8> for Mode {
/// type Error = u8;
/// fn try_from(value: u8) -> Result<Self, Self::Error> {
/// match value {
/// 0 => Ok(Mode::Low),
/// 1 => Ok(Mode::High),
/// 2 => Ok(Mode::Auto),
/// _ => Err(value),
/// }
/// }
/// }
///
/// impl From<Mode> for u8 {
/// fn from(mode: Mode) -> u8 {
/// mode as u8
/// }
/// }
///
/// #[derive(Debug, Clone, Copy, Default)]
/// enum State {
/// #[default]
/// Inactive = 0,
/// Active = 1,
/// }
///
/// impl From<bool> for State {
/// fn from(value: bool) -> Self {
/// if value { State::Active } else { State::Inactive }
/// }
/// }
///
/// impl From<State> for bool {
/// fn from(state: State) -> bool {
/// match state {
/// State::Inactive => false,
/// State::Active => true,
/// }
/// }
/// }
///
/// bitfield! {
/// pub struct ControlReg(u32) {
/// 7:7 state as bool => State;
/// 3:0 mode as u8 ?=> Mode;
/// }
/// }
/// ```
///
/// This generates a struct with:
/// - Field accessors: `mode()`, `state()`, etc.
/// - Field setters: `set_mode()`, `set_state()`, etc. (supports chaining with builder pattern).
/// Note that the compiler will error out if the size of the setter's arg exceeds the
/// struct's storage size.
/// - Debug and Default implementations.
///
/// Note: Field accessors and setters inherit the same visibility as the struct itself.
/// In the example above, both `mode()` and `set_mode()` methods will be `pub`.
///
/// Fields are defined as follows:
///
/// - `as <type>` simply returns the field value casted to <type>, typically `u32`, `u16`, `u8` or
/// `bool`. Note that `bool` fields must have a range of 1 bit.
/// - `as <type> => <into_type>` calls `<into_type>`'s `From::<<type>>` implementation and returns
/// the result.
/// - `as <type> ?=> <try_into_type>` calls `<try_into_type>`'s `TryFrom::<<type>>` implementation
/// and returns the result. This is useful with fields for which not all values are valid.
macro_rules! bitfield {
// Main entry point - defines the bitfield struct with fields
($vis:vis struct $name:ident($storage:ty) $(, $comment:literal)? { $($fields:tt)* }) => {
bitfield!(@core $vis $name $storage $(, $comment)? { $($fields)* });
};
// All rules below are helpers.
// Defines the wrapper `$name` type, as well as its relevant implementations (`Debug`,
// `Default`, and conversion to the value type) and field accessor methods.
(@core $vis:vis $name:ident $storage:ty $(, $comment:literal)? { $($fields:tt)* }) => {
$(
#[doc=$comment]
)?
#[repr(transparent)]
#[derive(Clone, Copy)]
$vis struct $name($storage);
impl ::core::convert::From<$name> for $storage {
fn from(val: $name) -> $storage {
val.0
}
}
bitfield!(@fields_dispatcher $vis $name $storage { $($fields)* });
};
// Captures the fields and passes them to all the implementers that require field information.
//
// Used to simplify the matching rules for implementers, so they don't need to match the entire
// complex fields rule even though they only make use of part of it.
(@fields_dispatcher $vis:vis $name:ident $storage:ty {
$($hi:tt:$lo:tt $field:ident as $type:tt
$(?=> $try_into_type:ty)?
$(=> $into_type:ty)?
$(, $comment:literal)?
;
)*
}
) => {
bitfield!(@field_accessors $vis $name $storage {
$(
$hi:$lo $field as $type
$(?=> $try_into_type)?
$(=> $into_type)?
$(, $comment)?
;
)*
});
bitfield!(@debug $name { $($field;)* });
bitfield!(@default $name { $($field;)* });
};
// Defines all the field getter/setter methods for `$name`.
(
@field_accessors $vis:vis $name:ident $storage:ty {
$($hi:tt:$lo:tt $field:ident as $type:tt
$(?=> $try_into_type:ty)?
$(=> $into_type:ty)?
$(, $comment:literal)?
;
)*
}
) => {
$(
bitfield!(@check_field_bounds $hi:$lo $field as $type);
)*
#[allow(dead_code)]
impl $name {
$(
bitfield!(@field_accessor $vis $name $storage, $hi:$lo $field as $type
$(?=> $try_into_type)?
$(=> $into_type)?
$(, $comment)?
;
);
)*
}
};
// Boolean fields must have `$hi == $lo`.
(@check_field_bounds $hi:tt:$lo:tt $field:ident as bool) => {
#[allow(clippy::eq_op)]
const _: () = {
::kernel::build_assert::build_assert!(
$hi == $lo,
concat!("boolean field `", stringify!($field), "` covers more than one bit")
);
};
};
// Non-boolean fields must have `$hi >= $lo`.
(@check_field_bounds $hi:tt:$lo:tt $field:ident as $type:tt) => {
#[allow(clippy::eq_op)]
const _: () = {
::kernel::build_assert::build_assert!(
$hi >= $lo,
concat!("field `", stringify!($field), "`'s MSB is smaller than its LSB")
);
};
};
// Catches fields defined as `bool` and convert them into a boolean value.
(
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as bool
=> $into_type:ty $(, $comment:literal)?;
) => {
bitfield!(
@leaf_accessor $vis $name $storage, $hi:$lo $field
{ |f| <$into_type>::from(f != 0) }
bool $into_type => $into_type $(, $comment)?;
);
};
// Shortcut for fields defined as `bool` without the `=>` syntax.
(
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as bool
$(, $comment:literal)?;
) => {
bitfield!(
@field_accessor $vis $name $storage, $hi:$lo $field as bool => bool $(, $comment)?;
);
};
// Catches the `?=>` syntax for non-boolean fields.
(
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as $type:tt
?=> $try_into_type:ty $(, $comment:literal)?;
) => {
bitfield!(@leaf_accessor $vis $name $storage, $hi:$lo $field
{ |f| <$try_into_type>::try_from(f as $type) } $type $try_into_type =>
::core::result::Result<
$try_into_type,
<$try_into_type as ::core::convert::TryFrom<$type>>::Error
>
$(, $comment)?;);
};
// Catches the `=>` syntax for non-boolean fields.
(
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as $type:tt
=> $into_type:ty $(, $comment:literal)?;
) => {
bitfield!(@leaf_accessor $vis $name $storage, $hi:$lo $field
{ |f| <$into_type>::from(f as $type) } $type $into_type => $into_type $(, $comment)?;);
};
// Shortcut for non-boolean fields defined without the `=>` or `?=>` syntax.
(
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as $type:tt
$(, $comment:literal)?;
) => {
bitfield!(
@field_accessor $vis $name $storage, $hi:$lo $field as $type => $type $(, $comment)?;
);
};
// Generates the accessor methods for a single field.
(
@leaf_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident
{ $process:expr } $prim_type:tt $to_type:ty => $res_type:ty $(, $comment:literal)?;
) => {
::kernel::macros::paste!(
const [<$field:upper _RANGE>]: ::core::ops::RangeInclusive<u8> = $lo..=$hi;
const [<$field:upper _MASK>]: $storage = {
// Generate mask for shifting
match ::core::mem::size_of::<$storage>() {
1 => ::kernel::bits::genmask_u8($lo..=$hi) as $storage,
2 => ::kernel::bits::genmask_u16($lo..=$hi) as $storage,
4 => ::kernel::bits::genmask_u32($lo..=$hi) as $storage,
8 => ::kernel::bits::genmask_u64($lo..=$hi) as $storage,
_ => ::kernel::build_error!("Unsupported storage type size")
}
};
const [<$field:upper _SHIFT>]: u32 = $lo;
);
$(
#[doc="Returns the value of this field:"]
#[doc=$comment]
)?
#[inline(always)]
$vis fn $field(self) -> $res_type {
::kernel::macros::paste!(
const MASK: $storage = $name::[<$field:upper _MASK>];
const SHIFT: u32 = $name::[<$field:upper _SHIFT>];
);
let field = ((self.0 & MASK) >> SHIFT);
$process(field)
}
::kernel::macros::paste!(
$(
#[doc="Sets the value of this field:"]
#[doc=$comment]
)?
#[inline(always)]
$vis fn [<set_ $field>](mut self, value: $to_type) -> Self {
const MASK: $storage = $name::[<$field:upper _MASK>];
const SHIFT: u32 = $name::[<$field:upper _SHIFT>];
let value = ($storage::from($prim_type::from(value)) << SHIFT) & MASK;
self.0 = (self.0 & !MASK) | value;
self
}
);
};
// Generates the `Debug` implementation for `$name`.
(@debug $name:ident { $($field:ident;)* }) => {
impl ::kernel::fmt::Debug for $name {
fn fmt(&self, f: &mut ::kernel::fmt::Formatter<'_>) -> ::kernel::fmt::Result {
f.debug_struct(stringify!($name))
.field("<raw>", &::kernel::prelude::fmt!("{:#x}", &self.0))
$(
.field(stringify!($field), &self.$field())
)*
.finish()
}
}
};
// Generates the `Default` implementation for `$name`.
(@default $name:ident { $($field:ident;)* }) => {
/// Returns a value for the bitfield where all fields are set to their default value.
impl ::core::default::Default for $name {
fn default() -> Self {
let value = Self(Default::default());
::kernel::macros::paste!(
$(
let value = value.[<set_ $field>](Default::default());
)*
);
value
}
}
};
}

View File

@ -5,17 +5,14 @@
use hal::FalconHal;
use kernel::{
device::{
self,
Device, //
},
device,
dma::{
Coherent,
CoherentBox,
DmaAddress,
DmaMask, //
DmaAddress, //
},
io::{
io_project,
poll::read_poll_timeout,
register::{
RegisterBase,
@ -24,7 +21,6 @@
Io,
},
prelude::*,
sync::aref::ARef,
time::Delta,
};
@ -358,41 +354,47 @@ pub(crate) trait FalconFirmware {
}
/// Contains the base parameters common to all Falcon instances.
pub(crate) struct Falcon<E: FalconEngine> {
pub(crate) struct Falcon<'a, E: FalconEngine> {
hal: KBox<dyn FalconHal<E>>,
dev: ARef<device::Device>,
dev: &'a device::Device<device::Bound>,
bar: Bar0<'a>,
}
impl<E: FalconEngine + 'static> Falcon<E> {
impl<'a, E: FalconEngine + 'static> Falcon<'a, E> {
/// Create a new falcon instance.
pub(crate) fn new(dev: &device::Device, chipset: Chipset) -> Result<Self> {
pub(crate) fn new(
dev: &'a device::Device<device::Bound>,
chipset: Chipset,
bar: Bar0<'a>,
) -> Result<Self> {
Ok(Self {
hal: hal::falcon_hal(chipset)?,
dev: dev.into(),
dev,
bar,
})
}
/// Resets DMA-related registers.
pub(crate) fn dma_reset(&self, bar: Bar0<'_>) {
bar.update(regs::NV_PFALCON_FBIF_CTL::of::<E>(), |v| {
pub(crate) fn dma_reset(&self) {
self.bar.update(regs::NV_PFALCON_FBIF_CTL::of::<E>(), |v| {
v.with_allow_phys_no_ctx(true)
});
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_DMACTL::zeroed(),
);
}
/// Reset the controller, select the falcon core, and wait for memory scrubbing to complete.
pub(crate) fn reset(&self, bar: Bar0<'_>) -> Result {
self.hal.reset_eng(bar)?;
self.hal.select_core(self, bar)?;
self.hal.reset_wait_mem_scrubbing(bar)?;
pub(crate) fn reset(&self) -> Result {
self.hal.reset_eng(self)?;
self.hal.select_core(self)?;
self.hal.reset_wait_mem_scrubbing(self)?;
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_RM::from(bar.read(regs::NV_PMC_BOOT_0).into_raw()),
regs::NV_PFALCON_FALCON_RM::from(self.bar.read(regs::NV_PMC_BOOT_0).into_raw()),
);
Ok(())
@ -404,18 +406,14 @@ pub(crate) fn reset(&self, bar: Bar0<'_>) -> Result {
/// Write a slice to Falcon IMEM memory using programmed I/O (PIO).
///
/// Returns `EINVAL` if `img.len()` is not a multiple of 4.
fn pio_wr_imem_slice(
&self,
bar: Bar0<'_>,
load_offsets: FalconPioImemLoadTarget<'_>,
) -> Result {
fn pio_wr_imem_slice(&self, load_offsets: FalconPioImemLoadTarget<'_>) -> Result {
// Rejecting misaligned images here allows us to avoid checking
// inside the loops.
if load_offsets.data.len() % 4 != 0 {
return Err(EINVAL);
}
bar.write(
self.bar.write(
WithBase::of::<E>().at(Self::PIO_PORT),
regs::NV_PFALCON_FALCON_IMEMC::zeroed()
.with_secure(load_offsets.secure)
@ -426,13 +424,13 @@ fn pio_wr_imem_slice(
for (n, block) in load_offsets.data.chunks(MEM_BLOCK_ALIGNMENT).enumerate() {
let n = u16::try_from(n)?;
let tag: u16 = load_offsets.start_tag.checked_add(n).ok_or(ERANGE)?;
bar.write(
self.bar.write(
WithBase::of::<E>().at(Self::PIO_PORT),
regs::NV_PFALCON_FALCON_IMEMT::zeroed().with_tag(tag),
);
for word in block.chunks_exact(4) {
let w = [word[0], word[1], word[2], word[3]];
bar.write(
self.bar.write(
WithBase::of::<E>().at(Self::PIO_PORT),
regs::NV_PFALCON_FALCON_IMEMD::zeroed().with_data(u32::from_le_bytes(w)),
);
@ -445,18 +443,14 @@ fn pio_wr_imem_slice(
/// Write a slice to Falcon DMEM memory using programmed I/O (PIO).
///
/// Returns `EINVAL` if `img.len()` is not a multiple of 4.
fn pio_wr_dmem_slice(
&self,
bar: Bar0<'_>,
load_offsets: FalconPioDmemLoadTarget<'_>,
) -> Result {
fn pio_wr_dmem_slice(&self, load_offsets: FalconPioDmemLoadTarget<'_>) -> Result {
// Rejecting misaligned images here allows us to avoid checking
// inside the loops.
if load_offsets.data.len() % 4 != 0 {
return Err(EINVAL);
}
bar.write(
self.bar.write(
WithBase::of::<E>().at(Self::PIO_PORT),
regs::NV_PFALCON_FALCON_DMEMC::zeroed()
.with_aincw(true)
@ -465,7 +459,7 @@ fn pio_wr_dmem_slice(
for word in load_offsets.data.chunks_exact(4) {
let w = [word[0], word[1], word[2], word[3]];
bar.write(
self.bar.write(
WithBase::of::<E>().at(Self::PIO_PORT),
regs::NV_PFALCON_FALCON_DMEMD::zeroed().with_data(u32::from_le_bytes(w)),
);
@ -477,29 +471,28 @@ fn pio_wr_dmem_slice(
/// Perform a PIO copy into `IMEM` and `DMEM` of `fw`, and prepare the falcon to run it.
pub(crate) fn pio_load<F: FalconFirmware<Target = E> + FalconPioLoadable>(
&self,
bar: Bar0<'_>,
fw: &F,
) -> Result {
bar.update(regs::NV_PFALCON_FBIF_CTL::of::<E>(), |v| {
self.bar.update(regs::NV_PFALCON_FBIF_CTL::of::<E>(), |v| {
v.with_allow_phys_no_ctx(true)
});
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_DMACTL::zeroed(),
);
if let Some(imem_ns) = fw.imem_ns_load_params() {
self.pio_wr_imem_slice(bar, imem_ns)?;
self.pio_wr_imem_slice(imem_ns)?;
}
if let Some(imem_sec) = fw.imem_sec_load_params() {
self.pio_wr_imem_slice(bar, imem_sec)?;
self.pio_wr_imem_slice(imem_sec)?;
}
self.pio_wr_dmem_slice(bar, fw.dmem_load_params())?;
self.pio_wr_dmem_slice(fw.dmem_load_params())?;
self.hal.program_brom(self, bar, &fw.brom_params());
self.hal.program_brom(self, &fw.brom_params());
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_BOOTVEC::zeroed().with_value(fw.boot_addr()),
);
@ -507,33 +500,43 @@ pub(crate) fn pio_load<F: FalconFirmware<Target = E> + FalconPioLoadable>(
Ok(())
}
/// Perform a DMA write according to `load_offsets` from `dma_handle` into the falcon's
/// Perform a DMA write according to `load_offsets` from `dma_obj` into the falcon's
/// `target_mem`.
///
/// `sec` is set if the loaded firmware is expected to run in secure mode.
fn dma_wr(
&self,
bar: Bar0<'_>,
dma_obj: &Coherent<[u8]>,
target_mem: FalconMem,
load_offsets: FalconDmaLoadTarget,
) -> Result {
const DMA_LEN: u32 = num::usize_into_u32::<{ MEM_BLOCK_ALIGNMENT }>();
// DMA transfers can only be done in units of 256 bytes. Compute how many such transfers we
// need to perform.
let num_transfers = load_offsets.len.div_ceil(DMA_LEN);
// For IMEM, we want to use the start offset as a virtual address tag for each page, since
// code addresses in the firmware (and the boot vector) are virtual.
//
// For DMEM we can fold the start offset into the DMA handle.
// For DMEM, the start offset is folded into the DMA address.
let (src_start, dma_start) = match target_mem {
FalconMem::ImemSecure | FalconMem::ImemNonSecure => {
(load_offsets.src_start, dma_obj.dma_handle())
}
FalconMem::Dmem => (
0,
dma_obj.dma_handle() + DmaAddress::from(load_offsets.src_start),
),
FalconMem::ImemSecure | FalconMem::ImemNonSecure => (load_offsets.src_start, 0),
FalconMem::Dmem => (0, usize::from_safe_cast(load_offsets.src_start)),
};
if dma_start % DmaAddress::from(DMA_LEN) > 0 {
let dma_address = {
// Upper limit of transfer is `(num_transfers * DMA_LEN) + load_offsets.src_start`.
let dma_end = num_transfers
.checked_mul(DMA_LEN)
.and_then(|size| size.checked_add(load_offsets.src_start))
.map(usize::from_safe_cast)
.ok_or(EOVERFLOW)?;
io_project!(dma_obj, [try: dma_start..dma_end]).dma_address()
};
if dma_address % DmaAddress::from(DMA_LEN) > 0 {
dev_err!(
self.dev,
"DMA transfer start addresses must be a multiple of {}\n",
@ -542,46 +545,19 @@ fn dma_wr(
return Err(EINVAL);
}
// The DMATRFBASE/1 register pair only supports a 49-bit address.
if dma_start > DmaMask::new::<49>().value() {
dev_err!(self.dev, "DMA address {:#x} exceeds 49 bits\n", dma_start);
return Err(ERANGE);
}
// DMA transfers can only be done in units of 256 bytes. Compute how many such transfers we
// need to perform.
let num_transfers = load_offsets.len.div_ceil(DMA_LEN);
// Check that the area we are about to transfer is within the bounds of the DMA object.
// Upper limit of transfer is `(num_transfers * DMA_LEN) + load_offsets.src_start`.
match num_transfers
.checked_mul(DMA_LEN)
.and_then(|size| size.checked_add(load_offsets.src_start))
{
None => {
dev_err!(self.dev, "DMA transfer length overflow\n");
return Err(EOVERFLOW);
}
Some(upper_bound) if usize::from_safe_cast(upper_bound) > dma_obj.size() => {
dev_err!(self.dev, "DMA transfer goes beyond range of DMA object\n");
return Err(EINVAL);
}
Some(_) => (),
};
// Set up the base source DMA address.
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_DMATRFBASE::zeroed().with_base(
// CAST: `as u32` is used on purpose since we do want to strip the upper bits,
// which will be written to `NV_PFALCON_FALCON_DMATRFBASE1`.
(dma_start >> 8) as u32,
(dma_address >> 8) as u32,
),
);
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_DMATRFBASE1::zeroed().try_with_base(dma_start >> 40)?,
regs::NV_PFALCON_FALCON_DMATRFBASE1::zeroed().try_with_base(dma_address >> 40)?,
);
let cmd = regs::NV_PFALCON_FALCON_DMATRFCMD::zeroed()
@ -590,23 +566,23 @@ fn dma_wr(
for pos in (0..num_transfers).map(|i| i * DMA_LEN) {
// Perform a transfer of size `DMA_LEN`.
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_DMATRFMOFFS::zeroed()
.try_with_offs(load_offsets.dst_start + pos)?,
);
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_DMATRFFBOFFS::zeroed().with_offs(src_start + pos),
);
bar.write(WithBase::of::<E>(), cmd);
self.bar.write(WithBase::of::<E>(), cmd);
// Wait for the transfer to complete.
// TIMEOUT: arbitrarily large value, no DMA transfer to the falcon's small memories
// should ever take that long.
read_poll_timeout(
|| Ok(bar.read(regs::NV_PFALCON_FALCON_DMATRFCMD::of::<E>())),
|| Ok(self.bar.read(regs::NV_PFALCON_FALCON_DMATRFCMD::of::<E>())),
|r| r.idle(),
Delta::ZERO,
Delta::from_secs(2),
@ -617,12 +593,7 @@ fn dma_wr(
}
/// Perform a DMA load into `IMEM` and `DMEM` of `fw`, and prepare the falcon to run it.
fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
&self,
dev: &Device<device::Bound>,
bar: Bar0<'_>,
fw: &F,
) -> Result {
fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(&self, fw: &F) -> Result {
// DMA object with firmware content as the source of the DMA engine.
let dma_obj = {
let fw_slice = fw.as_slice();
@ -630,7 +601,7 @@ fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
// DMA copies are done in chunks of `MEM_BLOCK_ALIGNMENT`, so pad the length
// accordingly and fill with `0`.
let mut dma_obj = CoherentBox::zeroed_slice(
dev,
self.dev,
fw_slice.len().next_multiple_of(MEM_BLOCK_ALIGNMENT),
GFP_KERNEL,
)?;
@ -642,24 +613,20 @@ fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
dma_obj.into()
};
self.dma_reset(bar);
bar.update(regs::NV_PFALCON_FBIF_TRANSCFG::of::<E>().at(0), |v| {
v.with_target(FalconFbifTarget::CoherentSysmem)
.with_mem_type(FalconFbifMemType::Physical)
});
self.dma_reset();
self.bar
.update(regs::NV_PFALCON_FBIF_TRANSCFG::of::<E>().at(0), |v| {
v.with_target(FalconFbifTarget::CoherentSysmem)
.with_mem_type(FalconFbifMemType::Physical)
});
self.dma_wr(
bar,
&dma_obj,
FalconMem::ImemSecure,
fw.imem_sec_load_params(),
)?;
self.dma_wr(bar, &dma_obj, FalconMem::Dmem, fw.dmem_load_params())?;
self.dma_wr(&dma_obj, FalconMem::ImemSecure, fw.imem_sec_load_params())?;
self.dma_wr(&dma_obj, FalconMem::Dmem, fw.dmem_load_params())?;
self.hal.program_brom(self, bar, &fw.brom_params());
self.hal.program_brom(self, &fw.brom_params());
// Set `BootVec` to start of non-secure code.
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_BOOTVEC::zeroed().with_value(fw.boot_addr()),
);
@ -668,10 +635,10 @@ fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
}
/// Wait until the falcon CPU is halted.
pub(crate) fn wait_till_halted(&self, bar: Bar0<'_>) -> Result<()> {
pub(crate) fn wait_till_halted(&self) -> Result<()> {
// TIMEOUT: arbitrarily large value, firmwares should complete in less than 2 seconds.
read_poll_timeout(
|| Ok(bar.read(regs::NV_PFALCON_FALCON_CPUCTL::of::<E>())),
|| Ok(self.bar.read(regs::NV_PFALCON_FALCON_CPUCTL::of::<E>())),
|r| r.halted(),
Delta::ZERO,
Delta::from_secs(2),
@ -681,16 +648,17 @@ pub(crate) fn wait_till_halted(&self, bar: Bar0<'_>) -> Result<()> {
}
/// Start the falcon CPU.
pub(crate) fn start(&self, bar: Bar0<'_>) -> Result<()> {
match bar
pub(crate) fn start(&self) -> Result<()> {
match self
.bar
.read(regs::NV_PFALCON_FALCON_CPUCTL::of::<E>())
.alias_en()
{
true => bar.write(
true => self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_CPUCTL_ALIAS::zeroed().with_startcpu(true),
),
false => bar.write(
false => self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_CPUCTL::zeroed().with_startcpu(true),
),
@ -700,16 +668,16 @@ pub(crate) fn start(&self, bar: Bar0<'_>) -> Result<()> {
}
/// Writes values to the mailbox registers if provided.
pub(crate) fn write_mailboxes(&self, bar: Bar0<'_>, mbox0: Option<u32>, mbox1: Option<u32>) {
pub(crate) fn write_mailboxes(&self, mbox0: Option<u32>, mbox1: Option<u32>) {
if let Some(mbox0) = mbox0 {
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_MAILBOX0::zeroed().with_value(mbox0),
);
}
if let Some(mbox1) = mbox1 {
bar.write(
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_MAILBOX1::zeroed().with_value(mbox1),
);
@ -717,21 +685,23 @@ pub(crate) fn write_mailboxes(&self, bar: Bar0<'_>, mbox0: Option<u32>, mbox1: O
}
/// Reads the value from `mbox0` register.
pub(crate) fn read_mailbox0(&self, bar: Bar0<'_>) -> u32 {
bar.read(regs::NV_PFALCON_FALCON_MAILBOX0::of::<E>())
pub(crate) fn read_mailbox0(&self) -> u32 {
self.bar
.read(regs::NV_PFALCON_FALCON_MAILBOX0::of::<E>())
.value()
}
/// Reads the value from `mbox1` register.
pub(crate) fn read_mailbox1(&self, bar: Bar0<'_>) -> u32 {
bar.read(regs::NV_PFALCON_FALCON_MAILBOX1::of::<E>())
pub(crate) fn read_mailbox1(&self) -> u32 {
self.bar
.read(regs::NV_PFALCON_FALCON_MAILBOX1::of::<E>())
.value()
}
/// Reads values from both mailbox registers.
pub(crate) fn read_mailboxes(&self, bar: Bar0<'_>) -> (u32, u32) {
let mbox0 = self.read_mailbox0(bar);
let mbox1 = self.read_mailbox1(bar);
pub(crate) fn read_mailboxes(&self) -> (u32, u32) {
let mbox0 = self.read_mailbox0();
let mbox1 = self.read_mailbox1();
(mbox0, mbox1)
}
@ -743,54 +713,54 @@ pub(crate) fn read_mailboxes(&self, bar: Bar0<'_>) -> (u32, u32) {
///
/// Wait up to two seconds for the firmware to complete, and return its exit status read from
/// the `MBOX0` and `MBOX1` registers.
pub(crate) fn boot(
&self,
bar: Bar0<'_>,
mbox0: Option<u32>,
mbox1: Option<u32>,
) -> Result<(u32, u32)> {
self.write_mailboxes(bar, mbox0, mbox1);
self.start(bar)?;
self.wait_till_halted(bar)?;
Ok(self.read_mailboxes(bar))
pub(crate) fn boot(&self, mbox0: Option<u32>, mbox1: Option<u32>) -> Result<(u32, u32)> {
self.write_mailboxes(mbox0, mbox1);
self.start()?;
self.wait_till_halted()?;
Ok(self.read_mailboxes())
}
/// Returns the fused version of the signature to use in order to run a HS firmware on this
/// falcon instance. `engine_id_mask` and `ucode_id` are obtained from the firmware header.
pub(crate) fn signature_reg_fuse_version(
&self,
bar: Bar0<'_>,
engine_id_mask: u16,
ucode_id: u8,
) -> Result<u32> {
self.hal
.signature_reg_fuse_version(self, bar, engine_id_mask, ucode_id)
.signature_reg_fuse_version(self, engine_id_mask, ucode_id)
}
/// Check if the RISC-V core is active.
///
/// Note that this does not guarantee that the RISC-V core is halted if it returns `false`.
///
/// Returns `true` if the RISC-V core is active, `false` otherwise.
pub(crate) fn is_riscv_active(&self, bar: Bar0<'_>) -> bool {
self.hal.is_riscv_active(bar)
pub(crate) fn is_riscv_active(&self) -> bool {
self.hal.is_riscv_active(self)
}
/// Checks whether the RISC-V core is halted.
///
/// Note that this does not guarantee that the RISC-V core is active if it returns `false`.
///
/// Returns [`ENOTSUPP`] if the status is not available.
pub(crate) fn is_riscv_halted(&self) -> Result<bool> {
self.hal.is_riscv_halted(self)
}
/// Load a firmware image into Falcon memory, using the preferred method for the current
/// chipset.
pub(crate) fn load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
&self,
dev: &Device<device::Bound>,
bar: Bar0<'_>,
fw: &F,
) -> Result {
pub(crate) fn load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(&self, fw: &F) -> Result {
match self.hal.load_method() {
LoadMethod::Dma => self.dma_load(dev, bar, fw),
LoadMethod::Pio => self.pio_load(bar, &fw.try_as_pio_loadable()?),
LoadMethod::Dma => self.dma_load(fw),
LoadMethod::Pio => self.pio_load(&fw.try_as_pio_loadable()?),
}
}
/// Write the application version to the OS register.
pub(crate) fn write_os_version(&self, bar: Bar0<'_>, app_version: u32) {
bar.write(
pub(crate) fn write_os_version(&self, app_version: u32) {
self.bar.write(
WithBase::of::<E>(),
regs::NV_PFALCON_FALCON_OS::zeroed().with_value(app_version),
);

View File

@ -17,11 +17,11 @@
Io, //
},
prelude::*,
sizes::SZ_1K,
time::Delta,
};
use crate::{
driver::Bar0,
falcon::{
Falcon,
FalconEngine,
@ -35,6 +35,9 @@
/// FSP message timeout in milliseconds.
const FSP_MSG_TIMEOUT_MS: i64 = 2000;
/// Size of the FSP EMEM channel 0 that we can use.
const FSP_EMEM_CHANNEL_0_SIZE: usize = SZ_1K;
/// Type specifying the `Fsp` falcon engine. Cannot be instantiated.
pub(crate) struct Fsp(());
@ -48,18 +51,18 @@ impl RegisterBase<PFalcon2Base> for Fsp {
impl FalconEngine for Fsp {}
impl Falcon<Fsp> {
impl<'a> Falcon<'a, Fsp> {
/// Writes `data` to FSP external memory at offset `0`.
///
/// `data` is interpreted as little-endian 32-bit words. Returns `EINVAL`
/// if the `data` length is not 4-byte aligned.
fn write_emem(&mut self, bar: Bar0<'_>, data: &[u8]) -> Result {
fn write_emem(&mut self, data: &[u8]) -> Result {
if data.len() % 4 != 0 {
return Err(EINVAL);
}
// Begin a write burst at offset `0`, auto-incrementing on each write.
bar.write(
self.bar.write(
WithBase::of::<Fsp>(),
regs::NV_PFALCON_FALCON_EMEMC::zeroed().with_aincw(true),
);
@ -68,7 +71,7 @@ fn write_emem(&mut self, bar: Bar0<'_>, data: &[u8]) -> Result {
let value = u32::from_le_bytes([chunk[0], chunk[1], chunk[2], chunk[3]]);
// Write the next 32-bit `value`; hardware advances the offset.
bar.write(
self.bar.write(
WithBase::of::<Fsp>(),
regs::NV_PFALCON_FALCON_EMEMD::zeroed().with_data(value),
);
@ -81,20 +84,23 @@ fn write_emem(&mut self, bar: Bar0<'_>, data: &[u8]) -> Result {
///
/// `data` is stored as little-endian 32-bit words. Returns `EINVAL` if
/// the `data` length is not 4-byte aligned.
fn read_emem(&mut self, bar: Bar0<'_>, data: &mut [u8]) -> Result {
fn read_emem(&mut self, data: &mut [u8]) -> Result {
if data.len() % 4 != 0 {
return Err(EINVAL);
}
// Begin a read burst at offset `0`, auto-incrementing on each read.
bar.write(
self.bar.write(
WithBase::of::<Fsp>(),
regs::NV_PFALCON_FALCON_EMEMC::zeroed().with_aincr(true),
);
for chunk in data.chunks_exact_mut(4) {
// Read the next 32-bit word; hardware advances the offset.
let value = bar.read(regs::NV_PFALCON_FALCON_EMEMD::of::<Fsp>()).data();
let value = self
.bar
.read(regs::NV_PFALCON_FALCON_EMEMD::of::<Fsp>())
.data();
chunk.copy_from_slice(&value.to_le_bytes());
}
@ -105,37 +111,41 @@ fn read_emem(&mut self, bar: Bar0<'_>, data: &mut [u8]) -> Result {
///
/// Returns the size of available data in bytes, or 0 if no data is available.
///
/// Returns [`EIO`] if the queue pointers are bogus (`tail < head`).
///
/// The FSP message queue is not circular. Pointers are reset to 0 after each
/// message exchange, so `tail >= head` is always true when data is present.
fn poll_msgq(&self, bar: Bar0<'_>) -> u32 {
let head = bar.read(regs::NV_PFSP_MSGQ_HEAD::at(0)).val();
let tail = bar.read(regs::NV_PFSP_MSGQ_TAIL::at(0)).val();
fn poll_msgq(&self) -> Result<u32> {
let head = self.bar.read(regs::NV_PFSP_MSGQ_HEAD::at(0)).val();
let tail = self.bar.read(regs::NV_PFSP_MSGQ_TAIL::at(0)).val();
if head == tail {
return 0;
Ok(0)
} else {
// TAIL points at the last DWORD written, so the size is `tail - head + 4`.
tail.checked_sub(head)
.and_then(|delta| delta.checked_add(4))
.ok_or(EIO)
}
// TAIL points at last DWORD written, so add 4 to get total size.
tail.saturating_sub(head).saturating_add(4)
}
/// Writes `packet` to FSP EMEM and updates the queue pointers to notify FSP.
///
/// Returns `EINVAL` if `packet` is empty or its length is not 4-byte aligned.
pub(crate) fn send_msg(&mut self, bar: Bar0<'_>, packet: &[u8]) -> Result {
pub(crate) fn send_msg(&mut self, packet: &[u8]) -> Result {
if packet.is_empty() {
return Err(EINVAL);
}
self.write_emem(bar, packet)?;
self.write_emem(packet)?;
// Update queue pointers. TAIL points at the last DWORD written.
let tail_offset = u32::try_from(packet.len() - 4).map_err(|_| EINVAL)?;
bar.write(
self.bar.write(
Array::at(0),
regs::NV_PFSP_QUEUE_TAIL::zeroed().with_address(tail_offset),
);
bar.write(
self.bar.write(
Array::at(0),
regs::NV_PFSP_QUEUE_HEAD::zeroed().with_address(0),
);
@ -148,23 +158,30 @@ pub(crate) fn send_msg(&mut self, bar: Bar0<'_>, packet: &[u8]) -> Result {
///
/// Returns `ETIMEDOUT` if no message was available until timeout, or a regular error code if a
/// memory allocation error occurred.
pub(crate) fn recv_msg(&mut self, bar: Bar0<'_>) -> Result<KVec<u8>> {
pub(crate) fn recv_msg(&mut self) -> Result<KVec<u8>> {
let msg_size = read_poll_timeout(
|| Ok(self.poll_msgq(bar)),
|| self.poll_msgq(),
|&size| size > 0,
Delta::from_millis(10),
Delta::from_millis(FSP_MSG_TIMEOUT_MS),
)
.map(num::u32_as_usize)?;
// Don't blindly allocate more than the maximum we expect from FSP.
if msg_size > FSP_EMEM_CHANNEL_0_SIZE {
return Err(EMSGSIZE);
}
let mut buffer = KVec::<u8>::new();
buffer.resize(msg_size, 0, GFP_KERNEL)?;
self.read_emem(bar, &mut buffer)?;
self.read_emem(&mut buffer)?;
// Reset message queue pointers after reading.
bar.write(Array::at(0), regs::NV_PFSP_MSGQ_TAIL::zeroed().with_val(0));
bar.write(Array::at(0), regs::NV_PFSP_MSGQ_HEAD::zeroed().with_val(0));
self.bar
.write(Array::at(0), regs::NV_PFSP_MSGQ_TAIL::zeroed().with_val(0));
self.bar
.write(Array::at(0), regs::NV_PFSP_MSGQ_HEAD::zeroed().with_val(0));
Ok(buffer)
}

View File

@ -14,7 +14,6 @@
};
use crate::{
driver::Bar0,
falcon::{
Falcon,
FalconEngine,
@ -24,10 +23,6 @@
regs,
};
/// Pattern returned by GSP register reads while the PRIV target mask still blocks CPU access.
const GSP_TARGET_MASK_LOCKED_PATTERN: u32 = 0xbadf_4100;
const GSP_TARGET_MASK_LOCKED_MASK: u32 = 0xffff_ff00;
/// Type specifying the `Gsp` falcon engine. Cannot be instantiated.
pub(crate) struct Gsp(());
@ -41,20 +36,20 @@ impl RegisterBase<PFalcon2Base> for Gsp {
impl FalconEngine for Gsp {}
impl Falcon<Gsp> {
impl<'a> Falcon<'a, Gsp> {
/// Clears the SWGEN0 bit in the Falcon's IRQ status clear register to
/// allow GSP to signal CPU for processing new messages in message queue.
pub(crate) fn clear_swgen0_intr(&self, bar: Bar0<'_>) {
bar.write(
pub(crate) fn clear_swgen0_intr(&self) {
self.bar.write(
WithBase::of::<Gsp>(),
regs::NV_PFALCON_FALCON_IRQSCLR::zeroed().with_swgen0(true),
);
}
/// Checks if GSP reload/resume has completed during the boot process.
pub(crate) fn check_reload_completed(&self, bar: Bar0<'_>, timeout: Delta) -> Result<bool> {
pub(crate) fn check_reload_completed(&self, timeout: Delta) -> Result<bool> {
read_poll_timeout(
|| Ok(bar.read(regs::NV_PGC6_BSI_SECURE_SCRATCH_14)),
|| Ok(self.bar.read(regs::NV_PGC6_BSI_SECURE_SCRATCH_14)),
|val| val.boot_stage_3_handoff(),
Delta::ZERO,
timeout,
@ -63,17 +58,24 @@ pub(crate) fn check_reload_completed(&self, bar: Bar0<'_>, timeout: Delta) -> Re
}
/// Returns whether the RISC-V branch privilege lockdown bit is set.
pub(crate) fn riscv_branch_privilege_lockdown(&self, bar: Bar0<'_>) -> bool {
bar.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<Gsp>())
pub(crate) fn riscv_branch_privilege_lockdown(&self) -> bool {
self.bar
.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<Gsp>())
.riscv_br_priv_lockdown()
}
/// Returns whether GSP registers can be read by the CPU.
pub(crate) fn priv_target_mask_released(&self, bar: Bar0<'_>) -> bool {
let hwcfg2 = bar
pub(crate) fn priv_target_mask_released(&self) -> bool {
/// Pattern returned by GSP register reads while the PRIV target mask still blocks CPU
/// access. The low byte varies; the upper 24 bits are fixed.
const LOCKED_PATTERN: u32 = 0xbadf_4100;
const LOCKED_MASK: u32 = 0xffff_ff00;
let hwcfg2 = self
.bar
.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<Gsp>())
.into_raw();
hwcfg2 != 0 && (hwcfg2 & GSP_TARGET_MASK_LOCKED_MASK) != GSP_TARGET_MASK_LOCKED_PATTERN
hwcfg2 != 0 && (hwcfg2 & LOCKED_MASK) != LOCKED_PATTERN
}
}

View File

@ -3,7 +3,6 @@
use kernel::prelude::*;
use crate::{
driver::Bar0,
falcon::{
Falcon,
FalconBromParams,
@ -34,7 +33,7 @@ pub(crate) enum LoadMethod {
/// registers.
pub(crate) trait FalconHal<E: FalconEngine>: Send + Sync {
/// Activates the Falcon core if the engine is a risvc/falcon dual engine.
fn select_core(&self, _falcon: &Falcon<E>, _bar: Bar0<'_>) -> Result {
fn select_core(&self, _falcon: &Falcon<'_, E>) -> Result {
Ok(())
}
@ -42,24 +41,28 @@ fn select_core(&self, _falcon: &Falcon<E>, _bar: Bar0<'_>) -> Result {
/// falcon instance. `engine_id_mask` and `ucode_id` are obtained from the firmware header.
fn signature_reg_fuse_version(
&self,
falcon: &Falcon<E>,
bar: Bar0<'_>,
falcon: &Falcon<'_, E>,
engine_id_mask: u16,
ucode_id: u8,
) -> Result<u32>;
/// Program the boot ROM registers prior to starting a secure firmware.
fn program_brom(&self, falcon: &Falcon<E>, bar: Bar0<'_>, params: &FalconBromParams);
fn program_brom(&self, falcon: &Falcon<'_, E>, params: &FalconBromParams);
/// Check if the RISC-V core is active.
/// Returns `true` if the RISC-V core is active, `false` otherwise.
fn is_riscv_active(&self, bar: Bar0<'_>) -> bool;
fn is_riscv_active(&self, falcon: &Falcon<'_, E>) -> bool;
/// Checks whether the RISC-V core is halted.
///
/// Returns [`ENOTSUPP`] if the chipset does not expose RISC-V halt status.
fn is_riscv_halted(&self, falcon: &Falcon<'_, E>) -> Result<bool>;
/// Wait for memory scrubbing to complete.
fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result;
fn reset_wait_mem_scrubbing(&self, falcon: &Falcon<'_, E>) -> Result;
/// Reset the falcon engine.
fn reset_eng(&self, bar: Bar0<'_>) -> Result;
fn reset_eng(&self, falcon: &Falcon<'_, E>) -> Result;
/// Returns the method used to load data into the falcon's memory.
///

View File

@ -115,33 +115,41 @@ pub(super) fn new() -> Self {
}
impl<E: FalconEngine> FalconHal<E> for Ga102<E> {
fn select_core(&self, _falcon: &Falcon<E>, bar: Bar0<'_>) -> Result {
select_core_ga102::<E>(bar)
fn select_core(&self, falcon: &Falcon<'_, E>) -> Result {
select_core_ga102::<E>(falcon.bar)
}
fn signature_reg_fuse_version(
&self,
falcon: &Falcon<E>,
bar: Bar0<'_>,
falcon: &Falcon<'_, E>,
engine_id_mask: u16,
ucode_id: u8,
) -> Result<u32> {
signature_reg_fuse_version_ga102(&falcon.dev, bar, engine_id_mask, ucode_id)
signature_reg_fuse_version_ga102(falcon.dev, falcon.bar, engine_id_mask, ucode_id)
}
fn program_brom(&self, _falcon: &Falcon<E>, bar: Bar0<'_>, params: &FalconBromParams) {
program_brom_ga102::<E>(bar, params);
fn program_brom(&self, falcon: &Falcon<'_, E>, params: &FalconBromParams) {
program_brom_ga102::<E>(falcon.bar, params);
}
fn is_riscv_active(&self, bar: Bar0<'_>) -> bool {
bar.read(regs::NV_PRISCV_RISCV_CPUCTL::of::<E>())
fn is_riscv_active(&self, falcon: &Falcon<'_, E>) -> bool {
falcon
.bar
.read(regs::NV_PRISCV_RISCV_CPUCTL::of::<E>())
.active_stat()
}
fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result {
fn is_riscv_halted(&self, falcon: &Falcon<'_, E>) -> Result<bool> {
Ok(falcon
.bar
.read(regs::NV_PRISCV_RISCV_CPUCTL::of::<E>())
.halted())
}
fn reset_wait_mem_scrubbing(&self, falcon: &Falcon<'_, E>) -> Result {
// TIMEOUT: memory scrubbing should complete in less than 20ms.
read_poll_timeout(
|| Ok(bar.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<E>())),
|| Ok(falcon.bar.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<E>())),
|r| r.mem_scrubbing_done(),
Delta::ZERO,
Delta::from_millis(20),
@ -149,7 +157,9 @@ fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result {
.map(|_| ())
}
fn reset_eng(&self, bar: Bar0<'_>) -> Result {
fn reset_eng(&self, falcon: &Falcon<'_, E>) -> Result {
let bar = falcon.bar;
let _ = bar.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<E>());
// According to OpenRM's `kflcnPreResetWait_GA102` documentation, HW sometimes does not set
@ -162,7 +172,7 @@ fn reset_eng(&self, bar: Bar0<'_>) -> Result {
);
regs::NV_PFALCON_FALCON_ENGINE::reset_engine::<E>(bar);
self.reset_wait_mem_scrubbing(bar)?;
self.reset_wait_mem_scrubbing(falcon)?;
Ok(())
}

View File

@ -13,7 +13,6 @@
};
use crate::{
driver::Bar0,
falcon::{
hal::LoadMethod,
Falcon,
@ -34,31 +33,36 @@ pub(super) fn new() -> Self {
}
impl<E: FalconEngine> FalconHal<E> for Tu102<E> {
fn select_core(&self, _falcon: &Falcon<E>, _bar: Bar0<'_>) -> Result {
fn select_core(&self, _falcon: &Falcon<'_, E>) -> Result {
Ok(())
}
fn signature_reg_fuse_version(
&self,
_falcon: &Falcon<E>,
_bar: Bar0<'_>,
_falcon: &Falcon<'_, E>,
_engine_id_mask: u16,
_ucode_id: u8,
) -> Result<u32> {
Ok(0)
}
fn program_brom(&self, _falcon: &Falcon<E>, _bar: Bar0<'_>, _params: &FalconBromParams) {}
fn program_brom(&self, _falcon: &Falcon<'_, E>, _params: &FalconBromParams) {}
fn is_riscv_active(&self, bar: Bar0<'_>) -> bool {
bar.read(regs::NV_PRISCV_RISCV_CORE_SWITCH_RISCV_STATUS::of::<E>())
fn is_riscv_active(&self, falcon: &Falcon<'_, E>) -> bool {
falcon
.bar
.read(regs::NV_PRISCV_RISCV_CORE_SWITCH_RISCV_STATUS::of::<E>())
.active_stat()
}
fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result {
fn is_riscv_halted(&self, _falcon: &Falcon<'_, E>) -> Result<bool> {
Err(ENOTSUPP)
}
fn reset_wait_mem_scrubbing(&self, falcon: &Falcon<'_, E>) -> Result {
// TIMEOUT: memory scrubbing should complete in less than 10ms.
read_poll_timeout(
|| Ok(bar.read(regs::NV_PFALCON_FALCON_DMACTL::of::<E>())),
|| Ok(falcon.bar.read(regs::NV_PFALCON_FALCON_DMACTL::of::<E>())),
|r| r.mem_scrubbing_done(),
Delta::ZERO,
Delta::from_millis(10),
@ -66,9 +70,9 @@ fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result {
.map(|_| ())
}
fn reset_eng(&self, bar: Bar0<'_>) -> Result {
regs::NV_PFALCON_FALCON_ENGINE::reset_engine::<E>(bar);
self.reset_wait_mem_scrubbing(bar)?;
fn reset_eng(&self, falcon: &Falcon<'_, E>) -> Result {
regs::NV_PFALCON_FALCON_ENGINE::reset_engine::<E>(falcon.bar);
self.reset_wait_mem_scrubbing(falcon)?;
Ok(())
}

View File

@ -24,10 +24,11 @@
gpu::Chipset,
gsp,
num::FromSafeCast,
regs, //
vgpu::VgpuState, //
};
mod hal;
mod regs;
/// Type holding the sysmem flush memory page, a page of memory to be written into the
/// `NV_PFB_NISO_FLUSH_SYSMEM_ADDR*` registers and used to maintain memory coherency.
@ -60,7 +61,7 @@ pub(crate) fn register(
) -> Result<Self> {
let page = CoherentHandle::alloc(dev, kernel::page::PAGE_SIZE, GFP_KERNEL)?;
hal::fb_hal(chipset).write_sysmem_flush_page(bar, page.dma_handle())?;
hal::fb_hal(chipset).write_sysmem_flush_page(bar, page.dma_address())?;
Ok(Self {
chipset,
@ -75,7 +76,7 @@ impl Drop for SysmemFlush<'_> {
fn drop(&mut self) {
let hal = hal::fb_hal(self.chipset);
if hal.read_sysmem_flush_page(self.bar) == self.page.dma_handle() {
if hal.read_sysmem_flush_page(self.bar) == self.page.dma_address() {
let _ = hal.write_sysmem_flush_page(self.bar, 0).inspect_err(|e| {
dev_warn!(
&self.device,
@ -148,7 +149,7 @@ fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
///
/// Contains ranges of GPU memory reserved for a given purpose during the GSP boot process.
#[derive(Debug)]
pub(crate) struct FbLayout {
pub(crate) struct FbRanges {
/// Range of the framebuffer. Starts at `0`.
pub(crate) fb: FbRange,
/// VGA workspace, small area of reserved memory at the end of the framebuffer.
@ -163,15 +164,22 @@ pub(crate) struct FbLayout {
pub(crate) wpr2_heap: FbRange,
/// WPR2 region range, starting with an instance of `GspFwWprMeta`.
pub(crate) wpr2: FbRange,
pub(crate) heap: FbRange,
/// Non-WPR heap, located just below WPR2.
pub(crate) non_wpr_heap: FbRange,
/// Number of VF partitions.
pub(crate) vf_partition_count: u8,
/// PMU reserved memory size, in bytes.
pub(crate) pmu_reserved_size: u32,
}
impl FbLayout {
/// Computes the FB layout for `chipset` required to run the `gsp_fw` GSP firmware.
pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, gsp_fw: &GspFirmware) -> Result<Self> {
impl FbRanges {
/// Computes concrete framebuffer ranges required on non-FSP booting architectures.
pub(crate) fn new(
chipset: Chipset,
bar: Bar0<'_>,
gsp_fw: &GspFirmware,
vgpu_state: VgpuState,
) -> Result<Self> {
let hal = hal::fb_hal(chipset);
let fb = {
@ -234,11 +242,15 @@ pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, gsp_fw: &GspFirmware) -> Resu
FbRange(elf_addr..elf_addr + elf_size)
};
let (vf_partition_count, wpr2_heap_size) = wpr2_heap_params(chipset, vgpu_state, fb.end)?;
let wpr2_heap = {
const WPR2_HEAP_DOWN_ALIGN: Alignment = Alignment::new::<SZ_1M>();
let wpr2_heap_size =
gsp::LibosParams::from_chipset(chipset).wpr_heap_size(chipset, fb.end)?;
let wpr2_heap_addr = (elf.start - wpr2_heap_size).align_down(WPR2_HEAP_DOWN_ALIGN);
let wpr2_heap_addr = elf
.start
.checked_sub(wpr2_heap_size)
.ok_or(EOVERFLOW)?
.align_down(WPR2_HEAP_DOWN_ALIGN);
FbRange(wpr2_heap_addr..(elf.start).align_down(WPR2_HEAP_DOWN_ALIGN))
};
@ -251,9 +263,9 @@ pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, gsp_fw: &GspFirmware) -> Resu
FbRange(wpr2_addr..frts.end)
};
let heap = {
let heap_size = u64::from(hal.non_wpr_heap_size());
FbRange(wpr2.start - heap_size..wpr2.start)
let non_wpr_heap = {
let non_wpr_heap_size = hal.non_wpr_heap_size();
FbRange(wpr2.start - non_wpr_heap_size..wpr2.start)
};
Ok(Self {
@ -264,9 +276,69 @@ pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, gsp_fw: &GspFirmware) -> Resu
elf,
wpr2_heap,
wpr2,
heap,
vf_partition_count: 0,
non_wpr_heap,
vf_partition_count,
pmu_reserved_size: hal.pmu_reserved_size(),
})
}
}
/// Reads the WPR2 memory region registers and returns the range if set.
/// Returns `None` if the WPR2 region is not set.
pub(crate) fn wpr2_range(bar: Bar0<'_>) -> Option<Range<u64>> {
let wpr2_hi = bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI);
if !wpr2_hi.is_wpr2_set() {
return None;
}
let wpr2_lo = bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_LO);
Some(wpr2_lo.lower_bound()..wpr2_hi.higher_bound())
}
/// Computes the number of VF partitions and the WPR2 heap size from the vGPU state.
fn wpr2_heap_params(chipset: Chipset, vgpu_state: VgpuState, fb_size: u64) -> Result<(u8, u64)> {
Ok(match vgpu_state {
VgpuState::Disabled => (
0,
gsp::LibosParams::from_chipset(chipset).wpr_heap_size(chipset, fb_size)?,
),
VgpuState::Enabled { total_vfs } => (
u8::try_from(total_vfs.get()).map_err(|_| EINVAL)?,
gsp::LibosParams::vgpu_wpr_heap_size(),
),
})
}
/// Framebuffer region sizes needed for GSP-FMC boot.
#[derive(Debug)]
pub(crate) struct FbSizes {
/// FRTS size, in bytes.
pub(crate) frts_size: u64,
/// WPR2 heap size, in bytes.
pub(crate) wpr2_heap_size: u64,
/// Non-WPR heap size, in bytes.
pub(crate) non_wpr_heap_size: u64,
/// PMU reserved memory size, in bytes.
pub(crate) pmu_reserved_size: u32,
/// Number of VF partitions.
pub(crate) vf_partition_count: u8,
}
impl FbSizes {
/// Computes the framebuffer region sizes for GSP-FMC boot.
pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, vgpu_state: VgpuState) -> Result<Self> {
let hal = hal::fb_hal(chipset);
let fb_size = hal.vidmem_size(bar);
let (vf_partition_count, wpr2_heap_size) = wpr2_heap_params(chipset, vgpu_state, fb_size)?;
Ok(Self {
frts_size: hal.frts_size(),
wpr2_heap_size,
non_wpr_heap_size: hal.non_wpr_heap_size(),
pmu_reserved_size: hal.pmu_reserved_size(),
vf_partition_count,
})
}
}

View File

@ -37,7 +37,7 @@ pub(crate) trait FbHal {
fn pmu_reserved_size(&self) -> u32;
/// Returns the non-WPR heap size for this chipset, in bytes.
fn non_wpr_heap_size(&self) -> u32;
fn non_wpr_heap_size(&self) -> u64;
/// Returns the FRTS size, in bytes.
fn frts_size(&self) -> u64;

View File

@ -9,8 +9,10 @@
use crate::{
driver::Bar0,
fb::hal::FbHal,
regs, //
fb::{
hal::FbHal,
regs, //
},
};
use super::tu102::FLUSH_SYSMEM_ADDR_SHIFT;
@ -41,7 +43,7 @@ pub(super) fn write_sysmem_flush_page_ga100(bar: Bar0<'_>, addr: u64) {
}
pub(super) fn display_enabled_ga100(bar: Bar0<'_>) -> bool {
!bar.read(regs::ga100::NV_FUSE_STATUS_OPT_DISPLAY)
!bar.read(crate::regs::ga100::NV_FUSE_STATUS_OPT_DISPLAY)
.display_disabled()
}
@ -72,7 +74,7 @@ fn pmu_reserved_size(&self) -> u32 {
super::tu102::pmu_reserved_size_tu102()
}
fn non_wpr_heap_size(&self) -> u32 {
fn non_wpr_heap_size(&self) -> u64 {
super::tu102::non_wpr_heap_size_tu102()
}

View File

@ -41,7 +41,7 @@ fn pmu_reserved_size(&self) -> u32 {
super::tu102::pmu_reserved_size_tu102()
}
fn non_wpr_heap_size(&self) -> u32 {
fn non_wpr_heap_size(&self) -> u64 {
super::tu102::non_wpr_heap_size_tu102()
}

View File

@ -22,9 +22,11 @@
use crate::{
driver::Bar0,
fb::hal::FbHal,
fb::{
hal::FbHal,
regs, //
},
num::usize_into_u32,
regs, //
};
struct Gb100;
@ -78,6 +80,7 @@ fn write_sysmem_flush_page_gb100(bar: Bar0<'_>, addr: Bounded<u64, 52>) {
);
}
// This PMU reservation size is r570-specific.
pub(super) const fn pmu_reserved_size_gb100() -> u32 {
usize_into_u32::<{ const_align_up(SZ_8M + SZ_16M + SZ_4K, Alignment::new::<SZ_128K>()).unwrap() }>(
)
@ -108,9 +111,9 @@ fn pmu_reserved_size(&self) -> u32 {
pmu_reserved_size_gb100()
}
fn non_wpr_heap_size(&self) -> u32 {
fn non_wpr_heap_size(&self) -> u64 {
// Non-WPR heap for GB10x (see Open RM: kgspGetNonWprHeapSize, GB100/GB102).
u32::SZ_2M
u64::SZ_2M
}
fn frts_size(&self) -> u64 {

View File

@ -4,13 +4,7 @@
//! Blackwell GB20x framebuffer HAL.
use kernel::{
io::{
register::{
RegisterBase,
WithBase, //
},
Io, //
},
io::Io,
num::Bounded,
prelude::*,
sizes::SizeConstants, //
@ -18,41 +12,37 @@
use crate::{
driver::Bar0,
fb::hal::FbHal,
regs, //
fb::{
hal::FbHal,
regs, //
},
};
struct Gb202;
impl RegisterBase<regs::Fbhub0Base> for Gb202 {
const BASE: usize = 0x008a_0000;
}
fn read_sysmem_flush_page_gb202(bar: Bar0<'_>) -> u64 {
let lo = u64::from(
bar.read(regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO::of::<Gb202>())
bar.read(regs::NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_LO)
.adr(),
);
let hi = u64::from(
bar.read(regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI::of::<Gb202>())
bar.read(regs::NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_HI)
.adr(),
);
lo | (hi << 32)
(hi << 32) | lo
}
/// Write the sysmem flush page address through the GB20x FBHUB0 registers.
fn write_sysmem_flush_page_gb202(bar: Bar0<'_>, addr: Bounded<u64, 52>) {
// Write HI first. The hardware will trigger the flush on the LO write.
bar.write(
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI::of::<Gb202>(),
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI::zeroed()
bar.write_reg(
regs::NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_HI::zeroed()
.with_adr(addr.shr::<32, 20>().cast::<u32>()),
);
bar.write(
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO::of::<Gb202>(),
bar.write_reg(
// CAST: lower 32 bits. Hardware ignores bits 7:0.
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO::zeroed().with_adr(*addr as u32),
regs::NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_LO::zeroed().with_adr(*addr as u32),
);
}
@ -81,9 +71,10 @@ fn pmu_reserved_size(&self) -> u32 {
super::gb100::pmu_reserved_size_gb100()
}
fn non_wpr_heap_size(&self) -> u32 {
fn non_wpr_heap_size(&self) -> u64 {
// Non-WPR heap for GB20x (see Open RM: kgspGetNonWprHeapSize, GB202+).
u32::SZ_2M + u32::SZ_128K
// This size is r570-specific.
u64::SZ_2M + u64::SZ_128K
}
fn frts_size(&self) -> u64 {

View File

@ -2,24 +2,51 @@
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use kernel::{
io::Io,
num::Bounded,
prelude::*,
sizes::SizeConstants, //
};
use crate::{
driver::Bar0,
fb::hal::FbHal, //
fb::{
hal::FbHal,
regs, //
},
};
struct Gh100;
fn read_sysmem_flush_page_gh100(bar: Bar0<'_>) -> u64 {
let lo = u64::from(bar.read(regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO).adr());
let hi = u64::from(bar.read(regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI).adr());
(hi << 32) | lo
}
/// Write the sysmem flush page address through the Hopper FBHUB registers.
fn write_sysmem_flush_page_gh100(bar: Bar0<'_>, addr: Bounded<u64, 52>) {
// Write HI first. The hardware will trigger the flush on the LO write.
bar.write_reg(
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI::zeroed()
.with_adr(addr.shr::<32, 20>().cast::<u32>()),
);
bar.write_reg(
// CAST: lower 32 bits. Hardware ignores bits 7:0.
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO::zeroed().with_adr(*addr as u32),
);
}
impl FbHal for Gh100 {
fn read_sysmem_flush_page(&self, bar: Bar0<'_>) -> u64 {
super::ga100::read_sysmem_flush_page_ga100(bar)
read_sysmem_flush_page_gh100(bar)
}
fn write_sysmem_flush_page(&self, bar: Bar0<'_>, addr: u64) -> Result {
super::ga100::write_sysmem_flush_page_ga100(bar, addr);
let addr = Bounded::<u64, 52>::try_new(addr).ok_or(EINVAL)?;
write_sysmem_flush_page_gh100(bar, addr);
Ok(())
}
@ -36,9 +63,9 @@ fn pmu_reserved_size(&self) -> u32 {
super::tu102::pmu_reserved_size_tu102()
}
fn non_wpr_heap_size(&self) -> u32 {
fn non_wpr_heap_size(&self) -> u64 {
// Non-WPR heap for Hopper (see Open RM: kgspCalculateFbLayout_GH100).
u32::SZ_2M
u64::SZ_2M
}
fn frts_size(&self) -> u64 {

View File

@ -9,8 +9,10 @@
use crate::{
driver::Bar0,
fb::hal::FbHal,
regs, //
fb::{
hal::FbHal,
regs, //
},
};
/// Shift applied to the sysmem address before it is written into `NV_PFB_NISO_FLUSH_SYSMEM_ADDR`,
@ -31,7 +33,7 @@ pub(super) fn write_sysmem_flush_page_gm107(bar: Bar0<'_>, addr: u64) -> Result
}
pub(super) fn display_enabled_gm107(bar: Bar0<'_>) -> bool {
!bar.read(regs::gm107::NV_FUSE_STATUS_OPT_DISPLAY)
!bar.read(crate::regs::gm107::NV_FUSE_STATUS_OPT_DISPLAY)
.display_disabled()
}
@ -44,8 +46,8 @@ pub(super) const fn pmu_reserved_size_tu102() -> u32 {
0
}
pub(super) const fn non_wpr_heap_size_tu102() -> u32 {
u32::SZ_1M
pub(super) const fn non_wpr_heap_size_tu102() -> u64 {
u64::SZ_1M
}
pub(super) const fn frts_size_tu102() -> u64 {
@ -75,7 +77,7 @@ fn pmu_reserved_size(&self) -> u32 {
pmu_reserved_size_tu102()
}
fn non_wpr_heap_size(&self) -> u32 {
fn non_wpr_heap_size(&self) -> u64 {
non_wpr_heap_size_tu102()
}

View File

@ -0,0 +1,155 @@
// SPDX-License-Identifier: GPL-2.0
use kernel::{
io::register,
sizes::SizeConstants, //
};
// PDISP
register! {
pub(super) NV_PDISP_VGA_WORKSPACE_BASE(u32) @ 0x00625f04 {
/// VGA workspace base address divided by 0x10000.
31:8 addr;
/// Set if the `addr` field is valid.
3:3 status_valid => bool;
}
}
impl NV_PDISP_VGA_WORKSPACE_BASE {
/// Returns the base address of the VGA workspace, or `None` if none exists.
pub(super) fn vga_workspace_addr(self) -> Option<u64> {
if self.status_valid() {
Some(u64::from(self.addr()) << 16)
} else {
None
}
}
}
// PFB
register! {
/// Low bits of the physical system memory address used by the GPU to perform sysmembar
/// operations (see [`crate::fb::SysmemFlush`]).
pub(super) NV_PFB_NISO_FLUSH_SYSMEM_ADDR(u32) @ 0x00100c10 {
31:0 adr_39_08;
}
/// High bits of the physical system memory address used by the GPU to perform sysmembar
/// operations.
pub(super) NV_PFB_NISO_FLUSH_SYSMEM_ADDR_HI(u32) @ 0x00100c40 {
23:0 adr_63_40;
}
pub(super) NV_PFB_PRI_MMU_LOCAL_MEMORY_RANGE(u32) @ 0x00100ce0 {
30:30 ecc_mode_enabled => bool;
9:4 lower_mag;
3:0 lower_scale;
}
pub(super) NV_PFB_PRI_MMU_WPR2_ADDR_LO(u32) @ 0x001fa824 {
/// Bits 12..40 of the lower (inclusive) bound of the WPR2 region.
31:4 lo_val;
}
pub(super) NV_PFB_PRI_MMU_WPR2_ADDR_HI(u32) @ 0x001fa828 {
/// Bits 12..40 of the higher (exclusive) bound of the WPR2 region.
31:4 hi_val;
}
}
/// Base of the GB10x HSHUB0 register window (`NV_HSHUB0_PRIV_BASE` in Open RM).
///
/// The base is provided by the GB10x framebuffer HAL.
pub(super) struct Hshub0Base(());
register! {
// GB10x sysmem flush registers, relative to the HSHUB0 base. GB10x routes sysmembar
// through a primary and an EG (egress) pair that must both be programmed to the same
// address. Hardware ignores bits 7:0 of each LO register. The boot path uses a fixed
// HSHUB0 base, so the multiple runtime-discovered HSHUB bases are not needed here.
pub(super) NV_PFB_HSHUB_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Hshub0Base + 0x00000e50 {
31:0 adr => u32;
}
pub(super) NV_PFB_HSHUB_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Hshub0Base + 0x00000e54 {
19:0 adr;
}
pub(super) NV_PFB_HSHUB_EG_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Hshub0Base + 0x000006c0 {
31:0 adr => u32;
}
pub(super) NV_PFB_HSHUB_EG_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Hshub0Base + 0x000006c4 {
19:0 adr;
}
}
register! {
// GB20x FBHUB0 sysmem flush registers. Unlike the older
// NV_PFB_NISO_FLUSH_SYSMEM_ADDR registers, which encode the address with an
// 8-bit right-shift, these take the raw address split into lower and upper
// halves. Hardware ignores bits 7:0 of the LO register.
pub(super) NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ 0x008a1d58 {
31:0 adr => u32;
}
pub(super) NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ 0x008a1d5c {
19:0 adr;
}
}
register! {
/// Low bits of the physical system memory address used by the GPU to perform
/// sysmembar operations on Hopper.
///
/// Like the GB20x FBHUB0 registers, and unlike the Ampere
/// `NV_PFB_NISO_FLUSH_SYSMEM_ADDR` registers (which encode the address with an
/// 8-bit right-shift), these take the raw address split into lower and upper
/// halves. Hardware ignores bits 7:0 of the LO register.
pub(super) NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ 0x00100a34 {
31:0 adr => u32;
}
/// High bits of the physical system memory address used by the GPU to perform
/// sysmembar operations on Hopper.
pub(super) NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ 0x00100a38 {
19:0 adr;
}
}
impl NV_PFB_PRI_MMU_LOCAL_MEMORY_RANGE {
/// Returns the usable framebuffer size, in bytes.
pub(super) fn usable_fb_size(self) -> u64 {
let size = (u64::from(self.lower_mag()) << u64::from(self.lower_scale())) * u64::SZ_1M;
if self.ecc_mode_enabled() {
// Remove the amount of memory reserved for ECC (one per 16 units).
size / 16 * 15
} else {
size
}
}
}
impl NV_PFB_PRI_MMU_WPR2_ADDR_LO {
/// Returns the lower (inclusive) bound of the WPR2 region.
pub(super) fn lower_bound(self) -> u64 {
u64::from(self.lo_val()) << 12
}
}
impl NV_PFB_PRI_MMU_WPR2_ADDR_HI {
/// Returns the higher (exclusive) bound of the WPR2 region.
///
/// A value of zero means the WPR2 region is not set.
pub(super) fn higher_bound(self) -> u64 {
u64::from(self.hi_val()) << 12
}
/// Returns whether the WPR2 region is currently set.
pub(super) fn is_wpr2_set(self) -> bool {
self.hi_val() != 0
}
}

View File

@ -8,11 +8,8 @@
use core::ops::Deref;
use kernel::{
device,
firmware,
prelude::*,
str::CString,
transmute::FromBytes, //
prelude::*, //
};
use crate::{
@ -21,10 +18,8 @@
FalconFirmware, //
},
gpu,
num::{
FromSafeCast,
IntoSafeCast, //
},
gsp::boot_firmware_files,
num::IntoSafeCast, //
};
pub(crate) mod booter;
@ -32,21 +27,7 @@
pub(crate) mod fwsec;
pub(crate) mod gsp;
pub(crate) mod riscv;
pub(crate) const FIRMWARE_VERSION: &str = "570.144";
/// Requests the GPU firmware `name` suitable for `chipset`, with version `ver`.
fn request_firmware(
dev: &device::Device,
chipset: gpu::Chipset,
name: &str,
ver: &str,
) -> Result<firmware::Firmware> {
let chip_name = chipset.name();
CString::try_from_fmt(fmt!("nvidia/{chip_name}/gsp/{name}-{ver}.bin"))
.and_then(|path| firmware::Firmware::request(&path, dev))
}
pub(crate) mod tlv;
/// Structure used to describe some firmwares, notably FWSEC-FRTS.
#[repr(C)]
@ -88,7 +69,7 @@ pub(crate) struct FalconUCodeDescV2 {
/// Structure used to describe some firmwares, notably FWSEC-FRTS.
#[repr(C)]
#[derive(Debug, Clone)]
#[derive(Debug, Clone, FromBytes)]
pub(crate) struct FalconUCodeDescV3 {
/// Header defined by `NV_BIT_FALCON_UCODE_DESC_HEADER_VDESC*` in OpenRM.
hdr: u32,
@ -119,10 +100,6 @@ pub(crate) struct FalconUCodeDescV3 {
_reserved: u16,
}
// SAFETY: all bit patterns are valid for this type, and it doesn't use
// interior mutability.
unsafe impl FromBytes for FalconUCodeDescV3 {}
/// Enum wrapping the different versions of Falcon microcode descriptors.
///
/// This allows handling both V2 and V3 descriptor formats through a
@ -349,60 +326,6 @@ fn no_patch_signature(self) -> FirmwareObject<F, Signed> {
}
}
/// Header common to most firmware files.
#[repr(C)]
#[derive(Debug, Clone)]
struct BinHdr {
/// Magic number, must be `0x10de`.
bin_magic: u32,
/// Version of the header.
bin_ver: u32,
/// Size in bytes of the binary (to be ignored).
bin_size: u32,
/// Offset of the start of the application-specific header.
header_offset: u32,
/// Offset of the start of the data payload.
data_offset: u32,
/// Size in bytes of the data payload.
data_size: u32,
}
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for BinHdr {}
// A firmware blob starting with a `BinHdr`.
struct BinFirmware<'a> {
hdr: BinHdr,
fw: &'a [u8],
}
impl<'a> BinFirmware<'a> {
/// Interpret `fw` as a firmware image starting with a [`BinHdr`], and returns the
/// corresponding [`BinFirmware`] that can be used to extract its payload.
fn new(fw: &'a firmware::Firmware) -> Result<Self> {
const BIN_MAGIC: u32 = 0x10de;
let fw = fw.data();
fw.get(0..size_of::<BinHdr>())
// Extract header.
.and_then(BinHdr::from_bytes_copy)
// Validate header.
.filter(|hdr| hdr.bin_magic == BIN_MAGIC)
.map(|hdr| Self { hdr, fw })
.ok_or(EINVAL)
}
/// Returns the data payload of the firmware, or `None` if the data range is out of bounds of
/// the firmware image.
fn data(&self) -> Option<&[u8]> {
let fw_start = usize::from_safe_cast(self.hdr.data_offset);
let fw_size = usize::from_safe_cast(self.hdr.data_size);
let fw_end = fw_start.checked_add(fw_size)?;
self.fw.get(fw_start..fw_end)
}
}
pub(crate) struct ModInfoBuilder<const N: usize>(firmware::ModInfoBuilder<N>);
impl<const N: usize> ModInfoBuilder<N> {
@ -413,33 +336,28 @@ const fn make_entry_file(self, chipset: &str, fw: &str) -> Self {
.push("nvidia/")
.push(chipset)
.push("/gsp/")
.push(fw)
.push("-")
.push(FIRMWARE_VERSION)
.push(".bin"),
.push(fw),
)
}
const fn make_entry_chipset(self, chipset: gpu::Chipset) -> Self {
let name = chipset.name();
let this = self
.make_entry_file(name, "booter_load")
.make_entry_file(name, "booter_unload")
.make_entry_file(name, "bootloader")
.make_entry_file(name, "gsp");
// GSP firmware files are always present.
let mut this = self
.make_entry_file(name, "gsp_bootloader.tlv")
.make_entry_file(name, "gsp.tlv")
.make_entry_file(name, "gsp.bin");
let this = if chipset.needs_fwsec_bootloader() {
this.make_entry_file(name, "gen_bootloader")
} else {
this
};
if chipset.uses_fsp() {
this.make_entry_file(name, "fmc")
} else {
this
// Add the firmware files specific to the GSP boot method of `chipset`.
let boot_files = boot_firmware_files(chipset);
let mut i = 0;
while i < boot_files.len() {
this = this.make_entry_file(name, boot_files[i]);
i += 1;
}
this
}
pub(crate) const fn create(
@ -456,210 +374,3 @@ pub(crate) const fn create(
this.0
}
}
/// Ad-hoc and temporary module to extract sections from ELF images.
///
/// Some firmware images are currently packaged as ELF files, where sections names are used as keys
/// to specific and related bits of data. Future firmware versions are scheduled to move away from
/// that scheme before nova-core becomes stable, which means this module will eventually be
/// removed.
mod elf {
use core::mem::size_of;
use kernel::{
bindings,
str::CStr,
transmute::FromBytes, //
};
/// Trait to abstract over ELF header differences.
trait ElfHeader: FromBytes {
fn shnum(&self) -> u16;
fn shoff(&self) -> u64;
fn shstrndx(&self) -> u16;
}
/// Trait to abstract over ELF section-header differences.
trait ElfSectionHeader: FromBytes {
fn name(&self) -> u32;
fn offset(&self) -> u64;
fn size(&self) -> u64;
}
/// Trait describing a matching ELF header and section-header format.
trait ElfFormat {
type Header: ElfHeader;
type SectionHeader: ElfSectionHeader;
}
/// Newtype to provide a [`FromBytes`] implementation.
#[repr(transparent)]
struct Elf64Hdr(bindings::elf64_hdr);
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for Elf64Hdr {}
impl ElfHeader for Elf64Hdr {
fn shnum(&self) -> u16 {
self.0.e_shnum
}
fn shoff(&self) -> u64 {
self.0.e_shoff
}
fn shstrndx(&self) -> u16 {
self.0.e_shstrndx
}
}
#[repr(transparent)]
struct Elf64SHdr(bindings::elf64_shdr);
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for Elf64SHdr {}
impl ElfSectionHeader for Elf64SHdr {
fn name(&self) -> u32 {
self.0.sh_name
}
fn offset(&self) -> u64 {
self.0.sh_offset
}
fn size(&self) -> u64 {
self.0.sh_size
}
}
struct Elf64Format;
impl ElfFormat for Elf64Format {
type Header = Elf64Hdr;
type SectionHeader = Elf64SHdr;
}
/// Newtype to provide [`FromBytes`] and [`ElfHeader`] implementations for ELF32.
#[repr(transparent)]
struct Elf32Hdr(bindings::elf32_hdr);
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for Elf32Hdr {}
impl ElfHeader for Elf32Hdr {
fn shnum(&self) -> u16 {
self.0.e_shnum
}
fn shoff(&self) -> u64 {
u64::from(self.0.e_shoff)
}
fn shstrndx(&self) -> u16 {
self.0.e_shstrndx
}
}
/// Newtype to provide [`FromBytes`] and [`ElfSectionHeader`] implementations for ELF32.
#[repr(transparent)]
struct Elf32SHdr(bindings::elf32_shdr);
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for Elf32SHdr {}
impl ElfSectionHeader for Elf32SHdr {
fn name(&self) -> u32 {
self.0.sh_name
}
fn offset(&self) -> u64 {
u64::from(self.0.sh_offset)
}
fn size(&self) -> u64 {
u64::from(self.0.sh_size)
}
}
struct Elf32Format;
impl ElfFormat for Elf32Format {
type Header = Elf32Hdr;
type SectionHeader = Elf32SHdr;
}
/// Returns a NULL-terminated string from the ELF image at `offset`.
fn elf_str(elf: &[u8], offset: u64) -> Option<&str> {
let idx = usize::try_from(offset).ok()?;
let bytes = elf.get(idx..)?;
CStr::from_bytes_until_nul(bytes).ok()?.to_str().ok()
}
fn elf_section_generic<'a, F>(elf: &'a [u8], name: &str) -> Option<&'a [u8]>
where
F: ElfFormat,
{
let hdr = F::Header::from_bytes(elf.get(0..size_of::<F::Header>())?)?;
let shdr_num = usize::from(hdr.shnum());
let shdr_start = usize::try_from(hdr.shoff()).ok()?;
let shdr_end = shdr_num
.checked_mul(size_of::<F::SectionHeader>())
.and_then(|v| v.checked_add(shdr_start))?;
// Get all the section headers as an iterator over byte chunks.
let shdr_bytes = elf.get(shdr_start..shdr_end)?;
let mut shdr_iter = shdr_bytes.chunks_exact(size_of::<F::SectionHeader>());
// Get the strings table.
let strhdr = shdr_iter
.clone()
.nth(usize::from(hdr.shstrndx()))
.and_then(F::SectionHeader::from_bytes)?;
// Find the section which name matches `name` and return it.
shdr_iter.find_map(|sh_bytes| {
let sh = F::SectionHeader::from_bytes(sh_bytes)?;
let name_offset = strhdr.offset().checked_add(u64::from(sh.name()))?;
let section_name = elf_str(elf, name_offset)?;
if section_name != name {
return None;
}
let start = usize::try_from(sh.offset()).ok()?;
let end = usize::try_from(sh.size())
.ok()
.and_then(|sz| start.checked_add(sz))?;
elf.get(start..end)
})
}
/// Extract the section with name `name` from the ELF64 image `elf`.
fn elf64_section<'a>(elf: &'a [u8], name: &str) -> Option<&'a [u8]> {
elf_section_generic::<Elf64Format>(elf, name)
}
/// Extract the section with name `name` from the ELF32 image `elf`.
fn elf32_section<'a>(elf: &'a [u8], name: &str) -> Option<&'a [u8]> {
elf_section_generic::<Elf32Format>(elf, name)
}
/// Automatically detects ELF32 vs ELF64 based on the ELF header.
pub(super) fn elf_section<'a>(elf: &'a [u8], name: &str) -> Option<&'a [u8]> {
// ELF identification: a 4-byte magic followed by a class byte (32- vs 64-bit).
const ELFMAG: &[u8] = b"\x7fELF";
const SELFMAG: usize = ELFMAG.len();
const EI_CLASS: usize = 4;
const ELFCLASS32: u8 = 1;
const ELFCLASS64: u8 = 2;
if elf.get(0..SELFMAG) != Some(ELFMAG) {
return None;
}
match *elf.get(EI_CLASS)? {
ELFCLASS32 => elf32_section(elf, name),
ELFCLASS64 => elf64_section(elf, name),
_ => None,
}
}
}

View File

@ -10,12 +10,10 @@
use kernel::{
device,
dma::Coherent,
prelude::*,
transmute::FromBytes, //
prelude::*, //
};
use crate::{
driver::Bar0,
falcon::{
sec2::Sec2,
Falcon,
@ -25,224 +23,19 @@
FalconFirmware, //
},
firmware::{
BinFirmware,
tlv::{
request_tlv, //
Tlv,
},
FirmwareObject,
FirmwareSignature,
Signed,
Unsigned, //
},
gpu::Chipset,
num::{
FromSafeCast,
IntoSafeCast, //
},
num::IntoSafeCast,
};
/// Local convenience function to return a copy of `S` by reinterpreting the bytes starting at
/// `offset` in `slice`.
fn frombytes_at<S: FromBytes + Sized>(slice: &[u8], offset: usize) -> Result<S> {
let end = offset.checked_add(size_of::<S>()).ok_or(EINVAL)?;
slice
.get(offset..end)
.and_then(S::from_bytes_copy)
.ok_or(EINVAL)
}
/// Heavy-Secured firmware header.
///
/// Such firmwares have an application-specific payload that needs to be patched with a given
/// signature.
#[repr(C)]
#[derive(Debug, Clone)]
struct HsHeaderV2 {
/// Offset to the start of the signatures.
sig_prod_offset: u32,
/// Size in bytes of the signatures.
sig_prod_size: u32,
/// Offset to a `u32` containing the location at which to patch the signature in the microcode
/// image.
patch_loc_offset: u32,
/// Offset to a `u32` containing the index of the signature to patch.
patch_sig_offset: u32,
/// Start offset to the signature metadata.
meta_data_offset: u32,
/// Size in bytes of the signature metadata.
meta_data_size: u32,
/// Offset to a `u32` containing the number of signatures in the signatures section.
num_sig_offset: u32,
/// Offset of the application-specific header.
header_offset: u32,
/// Size in bytes of the application-specific header.
header_size: u32,
}
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for HsHeaderV2 {}
/// Heavy-Secured Firmware image container.
///
/// This provides convenient access to the fields of [`HsHeaderV2`] that are actually indices to
/// read from in the firmware data.
struct HsFirmwareV2<'a> {
hdr: HsHeaderV2,
fw: &'a [u8],
}
impl<'a> HsFirmwareV2<'a> {
/// Interprets the header of `bin_fw` as a [`HsHeaderV2`] and returns an instance of
/// `HsFirmwareV2` for further parsing.
///
/// Fails if the header pointed at by `bin_fw` is not within the bounds of the firmware image.
fn new(bin_fw: &BinFirmware<'a>) -> Result<Self> {
frombytes_at::<HsHeaderV2>(bin_fw.fw, bin_fw.hdr.header_offset.into_safe_cast())
.map(|hdr| Self { hdr, fw: bin_fw.fw })
}
/// Returns the location at which the signatures should be patched in the microcode image.
///
/// Fails if the offset of the patch location is outside the bounds of the firmware
/// image.
fn patch_location(&self) -> Result<u32> {
frombytes_at::<u32>(self.fw, self.hdr.patch_loc_offset.into_safe_cast())
}
/// Returns an iterator to the signatures of the firmware. The iterator can be empty if the
/// firmware is unsigned.
///
/// Fails if the pointed signatures are outside the bounds of the firmware image.
fn signatures_iter(&'a self) -> Result<impl Iterator<Item = BooterSignature<'a>>> {
let num_sig = frombytes_at::<u32>(self.fw, self.hdr.num_sig_offset.into_safe_cast())?;
let iter = match self.hdr.sig_prod_size.checked_div(num_sig) {
// If there are no signatures, return an iterator that will yield zero elements.
None => (&[] as &[u8]).chunks_exact(1),
Some(sig_size) => {
let patch_sig =
frombytes_at::<u32>(self.fw, self.hdr.patch_sig_offset.into_safe_cast())?;
let signatures_start = self
.hdr
.sig_prod_offset
.checked_add(patch_sig)
.map(usize::from_safe_cast)
.ok_or(EINVAL)?;
let signatures_end = signatures_start
.checked_add(usize::from_safe_cast(self.hdr.sig_prod_size))
.ok_or(EINVAL)?;
self.fw
// Get signatures range.
.get(signatures_start..signatures_end)
.ok_or(EINVAL)?
.chunks_exact(sig_size.into_safe_cast())
}
};
// Map the byte slices into signatures.
Ok(iter.map(BooterSignature))
}
}
/// Signature parameters, as defined in the firmware.
#[repr(C)]
struct HsSignatureParams {
/// Fuse version to use.
fuse_ver: u32,
/// Mask of engine IDs this firmware applies to.
engine_id_mask: u32,
/// ID of the microcode.
ucode_id: u32,
}
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for HsSignatureParams {}
impl HsSignatureParams {
/// Returns the signature parameters contained in `hs_fw`.
///
/// Fails if the meta data parameter of `hs_fw` is outside the bounds of the firmware image, or
/// if its size doesn't match that of [`HsSignatureParams`].
fn new(hs_fw: &HsFirmwareV2<'_>) -> Result<Self> {
let start = usize::from_safe_cast(hs_fw.hdr.meta_data_offset);
let end = start
.checked_add(hs_fw.hdr.meta_data_size.into_safe_cast())
.ok_or(EINVAL)?;
hs_fw
.fw
.get(start..end)
.and_then(Self::from_bytes_copy)
.ok_or(EINVAL)
}
}
/// Header for code and data load offsets.
#[repr(C)]
#[derive(Debug, Clone)]
struct HsLoadHeaderV2 {
// Offset at which the code starts.
os_code_offset: u32,
// Total size of the code, for all apps.
os_code_size: u32,
// Offset at which the data starts.
os_data_offset: u32,
// Size of the data.
os_data_size: u32,
// Number of apps following this header. Each app is described by a [`HsLoadHeaderV2App`].
num_apps: u32,
}
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for HsLoadHeaderV2 {}
impl HsLoadHeaderV2 {
/// Returns the load header contained in `hs_fw`.
///
/// Fails if the header pointed at by `hs_fw` is not within the bounds of the firmware image.
fn new(hs_fw: &HsFirmwareV2<'_>) -> Result<Self> {
frombytes_at::<Self>(hs_fw.fw, hs_fw.hdr.header_offset.into_safe_cast())
}
}
/// Header for app code loader.
#[repr(C)]
#[derive(Debug, Clone)]
struct HsLoadHeaderV2App {
/// Offset at which to load the app code.
offset: u32,
/// Length in bytes of the app code.
len: u32,
}
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for HsLoadHeaderV2App {}
impl HsLoadHeaderV2App {
/// Returns the [`HsLoadHeaderV2App`] for app `idx` of `hs_fw`.
///
/// Fails if `idx` is larger than the number of apps declared in `hs_fw`, or if the header is
/// not within the bounds of the firmware image.
fn new(hs_fw: &HsFirmwareV2<'_>, idx: u32) -> Result<Self> {
let load_hdr = HsLoadHeaderV2::new(hs_fw)?;
if idx >= load_hdr.num_apps {
Err(EINVAL)
} else {
frombytes_at::<Self>(
hs_fw.fw,
usize::from_safe_cast(hs_fw.hdr.header_offset)
// Skip the load header...
.checked_add(size_of::<HsLoadHeaderV2>())
// ... and jump to app header `idx`.
.and_then(|offset| {
offset
.checked_add(usize::from_safe_cast(idx).checked_mul(size_of::<Self>())?)
})
.ok_or(EINVAL)?,
)
}
}
}
/// Signature for Booter firmware. Their size is encoded into the header and not known a compile
/// time, so we just wrap a byte slices on which we can implement [`FirmwareSignature`].
struct BooterSignature<'a>(&'a [u8]);
@ -292,89 +85,76 @@ pub(crate) fn new(
dev: &device::Device<device::Bound>,
kind: BooterKind,
chipset: Chipset,
ver: &str,
falcon: &Falcon<<Self as FalconFirmware>::Target>,
bar: Bar0<'_>,
falcon: &Falcon<'_, <Self as FalconFirmware>::Target>,
) -> Result<Self> {
let fw_name = match kind {
BooterKind::Loader => "booter_load",
BooterKind::Unloader => "booter_unload",
};
let fw = super::request_firmware(dev, chipset, fw_name, ver)?;
let bin_fw = BinFirmware::new(&fw)?;
let fw = request_tlv(dev, chipset, fw_name)?;
let tlv = Tlv::new(fw.data())?;
dev_dbg!(
dev,
"loaded {} firmware v{}\n",
fw_name,
tlv.get_string(b"VERS")?
);
// The binary firmware embeds a Heavy-Secured firmware.
let hs_fw = HsFirmwareV2::new(&bin_fw)?;
let os_data_offset = tlv.get_u32(b"DAOF")?;
let os_data_size = tlv.get_u32(b"DASZ")?;
let os_code_offset = tlv.get_u32(b"CDOF")?;
let os_code_size = tlv.get_u32(b"CDSZ")?;
let patch_loc = tlv.get_u32(b"PLOC")?;
let fuse_version: usize = tlv.get_u32(b"FUSE")?.into_safe_cast();
let engine_id = tlv.get_u32(b"ENID")?;
let ucode_id = tlv.get_u32(b"UCID")?;
let app0_code_offset = tlv.get_u32(b"A0CO")?;
let app0_code_size = tlv.get_u32(b"A0CS")?;
// The Heavy-Secured firmware embeds a firmware load descriptor.
let load_hdr = HsLoadHeaderV2::new(&hs_fw)?;
// Offset in `ucode` where to patch the signature.
let patch_loc = hs_fw.patch_location()?;
let sig_params = HsSignatureParams::new(&hs_fw)?;
let brom_params = FalconBromParams {
// `load_hdr.os_data_offset` is an absolute index, but `pkc_data_offset` is from the
// `os_data_offset` is an absolute index, but `pkc_data_offset` is from the
// signature patch location.
pkc_data_offset: patch_loc
.checked_sub(load_hdr.os_data_offset)
.ok_or(EINVAL)?,
engine_id_mask: u16::try_from(sig_params.engine_id_mask).map_err(|_| EINVAL)?,
ucode_id: u8::try_from(sig_params.ucode_id).map_err(|_| EINVAL)?,
pkc_data_offset: patch_loc.checked_sub(os_data_offset).ok_or(EINVAL)?,
engine_id_mask: u16::try_from(engine_id).map_err(|_| EINVAL)?,
ucode_id: u8::try_from(ucode_id).map_err(|_| EINVAL)?,
};
let app0 = HsLoadHeaderV2App::new(&hs_fw, 0)?;
// Object containing the firmware microcode to be signature-patched.
let ucode = bin_fw
.data()
.ok_or(EINVAL)
let ucode = tlv
.get_bytes(b"BLOB")
.and_then(FirmwareObject::<Self, _>::new_booter)?;
let ucode_signed = {
let mut signatures = hs_fw.signatures_iter()?.peekable();
// Obtain the version from the fuse register, and extract the corresponding
// signature.
let reg_fuse_version: usize = falcon
.signature_reg_fuse_version(brom_params.engine_id_mask, brom_params.ucode_id)?
.into_safe_cast();
if signatures.peek().is_none() {
// If there are no signatures, then the firmware is unsigned.
ucode.no_patch_signature()
} else {
// Obtain the version from the fuse register, and extract the corresponding
// signature.
let reg_fuse_version = falcon.signature_reg_fuse_version(
bar,
brom_params.engine_id_mask,
brom_params.ucode_id,
)?;
const FUSE_VERSION_USE_LAST_SIG: usize = 0;
// `0` means the last signature should be used.
const FUSE_VERSION_USE_LAST_SIG: u32 = 0;
let signature = match reg_fuse_version {
FUSE_VERSION_USE_LAST_SIG => signatures.last(),
// Otherwise hardware fuse version needs to be subtracted to obtain the index.
reg_fuse_version => {
let Some(idx) = sig_params.fuse_ver.checked_sub(reg_fuse_version) else {
dev_err!(dev, "invalid fuse version for Booter firmware\n");
return Err(EINVAL);
};
signatures.nth(idx.into_safe_cast())
}
}
.ok_or(EINVAL)?;
ucode.patch_signature(&signature, patch_loc.into_safe_cast())?
}
let index = match reg_fuse_version {
// `0` means the last signature should be used.
FUSE_VERSION_USE_LAST_SIG => None,
// Otherwise, hardware fuse version needs to be subtracted to obtain the index.
_ => Some(fuse_version.checked_sub(reg_fuse_version).ok_or(EINVAL)?),
};
// Extract the nth signature. Booter is always signed.
let sig_chunk = tlv.get_signature(index)?;
let signature = BooterSignature(sig_chunk);
let ucode_signed = ucode.patch_signature(&signature, patch_loc.into_safe_cast())?;
// There are two versions of Booter, one for Turing/GA100, and another for
// GA102+. The extraction of the IMEM sections differs between the two
// versions. Unfortunately, the file names are the same, and the headers
// don't indicate the versions. The only way to differentiate is by the Chipset.
let (imem_sec_dst_start, imem_ns_load_target) = if chipset <= Chipset::GA100 {
(
app0.offset,
app0_code_offset,
Some(FalconDmaLoadTarget {
src_start: 0,
dst_start: load_hdr.os_code_offset,
len: load_hdr.os_code_size,
dst_start: os_code_offset,
len: os_code_size,
}),
)
} else {
@ -383,15 +163,15 @@ pub(crate) fn new(
Ok(Self {
imem_sec_load_target: FalconDmaLoadTarget {
src_start: app0.offset,
src_start: app0_code_offset,
dst_start: imem_sec_dst_start,
len: app0.len,
len: app0_code_size,
},
imem_ns_load_target,
dmem_load_target: FalconDmaLoadTarget {
src_start: load_hdr.os_data_offset,
src_start: os_data_offset,
dst_start: 0,
len: load_hdr.os_data_size,
len: os_data_size,
},
brom_params,
ucode: ucode_signed,
@ -405,17 +185,15 @@ pub(crate) fn new(
pub(crate) fn run<T>(
&self,
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
sec2_falcon: &Falcon<Sec2>,
sec2_falcon: &Falcon<'_, Sec2>,
wpr_meta: &Coherent<T>,
) -> Result {
sec2_falcon.reset(bar)?;
sec2_falcon.load(dev, bar, self)?;
let wpr_handle = wpr_meta.dma_handle();
sec2_falcon.reset()?;
sec2_falcon.load(self)?;
let wpr_dma_address = wpr_meta.dma_address();
let (mbox0, mbox1) = sec2_falcon.boot(
bar,
Some(wpr_handle as u32),
Some((wpr_handle >> 32) as u32),
Some(wpr_dma_address as u32),
Some((wpr_dma_address >> 32) as u32),
)?;
dev_dbg!(dev, "SEC2 MBOX0: {:#x}, MBOX1: {:#x}\n", mbox0, mbox1);

View File

@ -6,12 +6,14 @@
use kernel::{
device,
dma::Coherent,
firmware::Firmware,
prelude::*, //
};
use crate::{
firmware::elf,
firmware::tlv::{
request_tlv, //
Tlv,
},
gpu::Chipset, //
};
@ -19,11 +21,11 @@
const FSP_HASH_SIZE: usize = 48;
/// Maximum size of the FSP public key (RSA-3072), in bytes.
///
/// The FMC ELF `publickey` section may be shorter, so the remaining bytes are zero-padded.
/// The FMC `PKEY` tag may be shorter, so the remaining bytes are zero-padded.
const FSP_PKEY_SIZE: usize = 384;
/// Maximum size of the FSP signature (RSA-3072), in bytes.
///
/// The FMC ELF `signature` section may be shorter, so the remaining bytes are zero-padded.
/// The FMC `SIGN` tag may be shorter, so the remaining bytes are zero-padded.
const FSP_SIG_SIZE: usize = 384;
/// Structure to hold FMC signatures.
@ -38,64 +40,35 @@ pub(crate) struct FmcSignatures {
}
pub(crate) struct FspFirmware {
/// FMC firmware image data (only the "image" ELF section).
/// FMC firmware image data
pub(crate) fmc_image: Coherent<[u8]>,
/// FMC firmware signatures.
pub(crate) fmc_sigs: KBox<FmcSignatures>,
}
impl FspFirmware {
pub(crate) fn new(
dev: &device::Device<device::Bound>,
chipset: Chipset,
ver: &str,
) -> Result<Self> {
let fw = super::request_firmware(dev, chipset, "fmc", ver)?;
pub(crate) fn new(dev: &device::Device<device::Bound>, chipset: Chipset) -> Result<Self> {
let fw = request_tlv(dev, chipset, "fmc")?;
let tlv = Tlv::new(fw.data())?;
dev_dbg!(dev, "loaded fsp firmware v{}\n", tlv.get_string(b"VERS")?);
// FSP expects only the "image" section, not the entire ELF file.
let fmc_image_data = elf::elf_section(fw.data(), "image").ok_or_else(|| {
dev_err!(dev, "FMC ELF file missing 'image' section\n");
EINVAL
})?;
let fmc_image_data = tlv.get_bytes(b"BLOB")?;
let fmc_image = Coherent::from_slice(dev, fmc_image_data, GFP_KERNEL)?;
Ok(Self {
fmc_image,
fmc_sigs: Self::extract_fmc_signatures(&fw, dev)?,
fmc_sigs: Self::extract_fmc_signatures(&tlv, dev)?,
})
}
/// Extract FMC firmware signatures for Chain of Trust verification.
///
/// Extracts real cryptographic signatures from FMC ELF32 firmware sections.
/// Extracts real cryptographic signatures from FMC TLV firmware tags.
/// Returns signatures in a heap-allocated structure to prevent stack overflow.
fn extract_fmc_signatures(
fmc_fw: &Firmware,
dev: &device::Device,
) -> Result<KBox<FmcSignatures>> {
let get_section = |name: &str, max_len: usize| {
elf::elf_section(fmc_fw.data(), name)
.ok_or(EINVAL)
.inspect_err(|_| dev_err!(dev, "FMC firmware missing '{}' section\n", name))
.and_then(|section| {
if section.len() > max_len {
dev_err!(
dev,
"FMC {} section size {} > maximum {}\n",
name,
section.len(),
max_len
);
Err(EINVAL)
} else {
Ok(section)
}
})
};
let hash_section = get_section("hash", FSP_HASH_SIZE)?;
let pkey_section = get_section("publickey", FSP_PKEY_SIZE)?;
let sig_section = get_section("signature", FSP_SIG_SIZE)?;
fn extract_fmc_signatures(tlv: &Tlv<'_>, dev: &device::Device) -> Result<KBox<FmcSignatures>> {
let hash_section = tlv.get_bytes(b"HASH")?;
let pkey_section = tlv.get_bytes(b"PKEY")?;
let sig_section = tlv.get_bytes(b"SIGN")?;
// The hash section is a SHA-384 output: it must be exactly FSP_HASH_SIZE bytes.
if hash_section.len() != FSP_HASH_SIZE {
@ -108,15 +81,36 @@ fn extract_fmc_signatures(
return Err(EINVAL);
}
// The key and signature sections are zero-padded to a fixed maximum, so they may be
// shorter, but must not exceed the destination buffers.
if pkey_section.len() > FSP_PKEY_SIZE {
dev_err!(
dev,
"FMC public key section size {} > maximum {}\n",
pkey_section.len(),
FSP_PKEY_SIZE
);
return Err(EINVAL);
}
if sig_section.len() > FSP_SIG_SIZE {
dev_err!(
dev,
"FMC signature section size {} > maximum {}\n",
sig_section.len(),
FSP_SIG_SIZE
);
return Err(EINVAL);
}
// Initialize the signatures in place to avoid building the large `FmcSignatures` on the
// stack, then fill each section from the firmware.
let signatures = KBox::init(
pin_init::init_zeroed::<FmcSignatures>().chain(|sigs| {
// PANIC: src and dst lengths are both FSP_HASH_SIZE (verified above).
sigs.hash384.copy_from_slice(hash_section);
// PANIC: dst is sliced to src.len(); src.len() <= FSP_PKEY_SIZE per `get_section`.
// PANIC: dst is sliced to src.len(); src.len() <= FSP_PKEY_SIZE (verified above).
sigs.public_key[..pkey_section.len()].copy_from_slice(pkey_section);
// PANIC: dst is sliced to src.len(); src.len() <= FSP_SIG_SIZE per `get_section`.
// PANIC: dst is sliced to src.len(); src.len() <= FSP_SIG_SIZE (verified above).
sigs.signature[..sig_section.len()].copy_from_slice(sig_section);
Ok(())
}),

View File

@ -27,7 +27,6 @@
};
use crate::{
driver::Bar0,
falcon::{
gsp::Gsp,
Falcon,
@ -320,8 +319,7 @@ impl FwsecFirmware {
/// command.
pub(crate) fn new(
dev: &Device<device::Bound>,
falcon: &Falcon<Gsp>,
bar: Bar0<'_>,
falcon: &Falcon<'_, Gsp>,
bios: &Vbios,
cmd: FwsecCommand,
) -> Result<Self> {
@ -337,7 +335,7 @@ pub(crate) fn new(
.ok_or(EINVAL)?;
let desc_sig_versions = u32::from(desc.signature_versions());
let reg_fuse_version =
falcon.signature_reg_fuse_version(bar, desc.engine_id_mask(), desc.ucode_id())?;
falcon.signature_reg_fuse_version(desc.engine_id_mask(), desc.ucode_id())?;
dev_dbg!(
dev,
"desc_sig_versions: {:#x}, reg_fuse_version: {}\n",
@ -387,24 +385,18 @@ pub(crate) fn new(
/// Loads the FWSEC firmware into `falcon` and execute it.
///
/// This must only be called on chipsets that do not need the FWSEC bootloader (i.e., where
/// [`Chipset::needs_fwsec_bootloader()`](crate::gpu::Chipset::needs_fwsec_bootloader) returns
/// `false`). On chipsets that do, use [`bootloader::FwsecFirmwareWithBl`] instead.
pub(crate) fn run(
&self,
dev: &Device<device::Bound>,
falcon: &Falcon<Gsp>,
bar: Bar0<'_>,
) -> Result<()> {
/// This must only be called on chipsets that do not need the FWSEC bootloader. On chipsets
/// where the bootloader is required, use [`bootloader::FwsecFirmwareWithBl`] instead.
pub(crate) fn run(&self, dev: &Device<device::Bound>, falcon: &Falcon<'_, Gsp>) -> Result<()> {
// Reset falcon, load the firmware, and run it.
falcon
.reset(bar)
.reset()
.inspect_err(|e| dev_err!(dev, "Failed to reset GSP falcon: {:?}\n", e))?;
falcon
.load(dev, bar, self)
.load(self)
.inspect_err(|e| dev_err!(dev, "Failed to load FWSEC firmware: {:?}\n", e))?;
let (mbox0, _) = falcon
.boot(bar, Some(0), None)
.boot(Some(0), None)
.inspect_err(|e| dev_err!(dev, "Failed to boot FWSEC firmware: {:?}\n", e))?;
if mbox0 != 0 {
dev_err!(dev, "FWSEC firmware returned error {}\n", mbox0);

View File

@ -7,26 +7,19 @@
//! be loaded using PIO.
use kernel::{
alloc::KVec,
device::{
self,
Device, //
},
dma::Coherent,
io::{
register::WithBase, //
Io,
},
io::{register::WithBase, Io},
prelude::*,
ptr::{
Alignable,
Alignment, //
},
sizes,
transmute::{
AsBytes,
FromBytes, //
},
transmute::AsBytes,
};
use crate::{
@ -46,38 +39,16 @@
},
firmware::{
fwsec::FwsecFirmware,
request_firmware,
BinHdr,
FIRMWARE_VERSION, //
tlv::{
request_tlv, //
Tlv,
},
},
gpu::Chipset,
num::FromSafeCast,
num::FromSafeCast, //
regs,
};
/// Descriptor used by RM to figure out the requirements of the boot loader.
///
/// Most of its fields appear to be legacy and carry incorrect values, so they are left unused.
#[repr(C)]
#[derive(Debug, Clone)]
struct BootloaderDesc {
/// Starting tag of bootloader.
start_tag: u32,
/// DMEM load offset - unused here as we always load at offset `0`.
_dmem_load_off: u32,
/// Offset of code section in the image. Unused as there is only one section in the bootloader
/// binary.
_code_off: u32,
/// Size of code section in the image.
code_size: u32,
/// Offset of data section in the image. Unused as we build the data section ourselves.
_data_off: u32,
/// Size of data section in the image. Unused as we build the data section ourselves.
_data_size: u32,
}
// SAFETY: any byte sequence is valid for this struct.
unsafe impl FromBytes for BootloaderDesc {}
/// Structure used by the boot-loader to load the rest of the code.
///
/// This has to be filled by the GPU driver and copied into DMEM at offset
@ -150,38 +121,24 @@ pub(crate) fn new(
dev: &Device<device::Bound>,
chipset: Chipset,
) -> Result<Self> {
let fw = request_firmware(dev, chipset, "gen_bootloader", FIRMWARE_VERSION)?;
let hdr = fw
.data()
.get(0..size_of::<BinHdr>())
.and_then(BinHdr::from_bytes_copy)
.ok_or(EINVAL)?;
let desc = {
let desc_offset = usize::from_safe_cast(hdr.header_offset);
fw.data()
.get(desc_offset..)
.and_then(BootloaderDesc::from_bytes_copy_prefix)
.ok_or(EINVAL)?
.0
};
let fw = request_tlv(dev, chipset, "gen_bootloader")?;
let tlv = Tlv::new(fw.data())?;
dev_dbg!(
dev,
"loaded generic bootloader firmware v{}\n",
tlv.get_string(b"VERS")?
);
let ucode = {
let ucode_start = usize::from_safe_cast(hdr.data_offset);
let code_size = usize::from_safe_cast(desc.code_size);
// Align to falcon block size (256 bytes).
let blob = tlv.get_bytes(b"BLOB")?;
let code_size = usize::from_safe_cast(tlv.get_u32(b"CDSZ")?);
let code = blob.get(..code_size).ok_or(EINVAL)?;
let aligned_code_size = code_size
.align_up(Alignment::new::<{ falcon::MEM_BLOCK_ALIGNMENT }>())
.ok_or(EINVAL)?;
let mut ucode = KVec::with_capacity(aligned_code_size, GFP_KERNEL)?;
ucode.extend_from_slice(
fw.data()
.get(ucode_start..ucode_start + code_size)
.ok_or(EINVAL)?,
GFP_KERNEL,
)?;
ucode.extend_from_slice(code, GFP_KERNEL)?;
ucode.resize(aligned_code_size, 0, GFP_KERNEL)?;
ucode
@ -234,7 +191,7 @@ pub(crate) fn new(
reserved: [0; 4],
signature: [0; 4],
ctx_dma: FALCON_DMAIDX_PHYS_SYS_NCOH,
code_dma_base: firmware_dma.dma_handle(),
code_dma_base: firmware_dma.dma_address(),
// `dst_start` is also valid as the source offset since the firmware DMA object is
// a mirror image of the target IMEM layout.
non_sec_code_off: imem_ns.dst_start,
@ -246,7 +203,7 @@ pub(crate) fn new(
code_entry_point: 0,
// Start of data section is the added padding + the DMEM `src_start` field.
data_dma_base: firmware_dma
.dma_handle()
.dma_address()
.checked_add(u64::from_safe_cast(align_padding))
.and_then(|offset| offset.checked_add(dmem.src_start.into()))
.ok_or(EOVERFLOW)?,
@ -262,13 +219,15 @@ pub(crate) fn new(
.checked_sub(ucode.len())
.ok_or(EOVERFLOW)?;
let start_tag = u16::try_from(tlv.get_u32(b"STRT")?)?;
Ok(Self {
_firmware_dma: firmware_dma,
ucode,
dmem_desc,
brom_params: firmware.brom_params(),
imem_dst_start: u16::try_from(imem_dst_start)?,
start_tag: u16::try_from(desc.start_tag)?,
start_tag,
})
}
@ -279,15 +238,15 @@ pub(crate) fn new(
pub(crate) fn run(
&self,
dev: &Device<device::Bound>,
falcon: &Falcon<Gsp>,
falcon: &Falcon<'_, Gsp>,
bar: Bar0<'_>,
) -> Result<()> {
// Reset falcon, load the firmware, and run it.
falcon
.reset(bar)
.reset()
.inspect_err(|e| dev_err!(dev, "Failed to reset GSP falcon: {:?}\n", e))?;
falcon
.pio_load(bar, self)
.pio_load(self)
.inspect_err(|e| dev_err!(dev, "Failed to load FWSEC firmware: {:?}\n", e))?;
// Configure DMA index for the bootloader to fetch the FWSEC firmware from system memory.
@ -302,7 +261,7 @@ pub(crate) fn run(
);
let (mbox0, _) = falcon
.boot(bar, Some(0), None)
.boot(Some(0), None)
.inspect_err(|e| dev_err!(dev, "Failed to boot FWSEC firmware: {:?}\n", e))?;
if mbox0 != 0 {
dev_err!(dev, "FWSEC firmware returned error {}\n", mbox0);

View File

@ -8,22 +8,24 @@
DataDirection,
DmaAddress, //
},
firmware,
prelude::*,
scatterlist::{
Owned,
SGTable, //
},
str::CString,
};
use crate::{
firmware::{
elf,
riscv::RiscvFirmware, //
tlv::{
request_tlv, //
Tlv,
},
},
gpu::{
Architecture,
Chipset, //
},
gpu::Chipset,
gsp::GSP_PAGE_SIZE,
num::FromSafeCast,
};
@ -63,43 +65,26 @@ pub(crate) struct GspFirmware {
}
impl GspFirmware {
fn find_gsp_sigs_section(chipset: Chipset) -> &'static str {
match chipset.arch() {
Architecture::Turing if matches!(chipset, Chipset::TU116 | Chipset::TU117) => {
".fwsignature_tu11x"
}
Architecture::Turing => ".fwsignature_tu10x",
Architecture::Ampere if chipset == Chipset::GA100 => ".fwsignature_ga100",
Architecture::Ampere => ".fwsignature_ga10x",
Architecture::Ada => ".fwsignature_ad10x",
Architecture::Hopper => ".fwsignature_gh10x",
Architecture::BlackwellGB10x => ".fwsignature_gb10x",
Architecture::BlackwellGB20x => ".fwsignature_gb20x",
}
}
/// Loads the GSP firmware binaries, map them into `dev`'s address-space, and creates the page
/// tables expected by the GSP bootloader to load it.
pub(crate) fn new<'a>(
dev: &'a device::Device<device::Bound>,
chipset: Chipset,
ver: &'a str,
) -> impl PinInit<Self, Error> + 'a {
pin_init::pin_init_scope(move || {
let firmware = super::request_firmware(dev, chipset, "gsp", ver)?;
let firmware = request_tlv(dev, chipset, "gsp")?;
let tlv = Tlv::new(firmware.data())?;
dev_dbg!(dev, "loaded gsp firmware v{}\n", tlv.get_string(b"VERS")?);
let fw_section = elf::elf_section(firmware.data(), ".fwimage").ok_or(EINVAL)?;
let size = usize::from_safe_cast(tlv.get_u32(b"SIZE")?);
let mut fw_vvec = VVec::zeroed(size, GFP_KERNEL).map_err(|_| ENOMEM)?;
let size = fw_section.len();
let chip_name = chipset.name();
let file = tlv.get_string(b"FILE")?;
let filename = CString::try_from_fmt(fmt!("nvidia/{chip_name}/gsp/{file}"))?;
firmware::request_into_buf(&filename, dev, fw_vvec.as_mut_slice())?;
// Move the firmware into a vmalloc'd vector and map it into the device address
// space.
let fw_vvec = VVec::with_capacity(fw_section.len(), GFP_KERNEL)
.and_then(|mut v| {
v.extend_from_slice(fw_section, GFP_KERNEL)?;
Ok(v)
})
.map_err(|_| ENOMEM)?;
let signatures = Coherent::from_slice(dev, tlv.get_bytes(b"SIGN")?, GFP_KERNEL)?;
Ok(try_pin_init!(Self {
fw <- SGTable::new(dev, fw_vvec, DataDirection::ToDevice, GFP_KERNEL),
@ -145,15 +130,9 @@ pub(crate) fn new<'a>(
level0.into()
},
size,
signatures: {
let sigs_section = Self::find_gsp_sigs_section(chipset);
elf::elf_section(firmware.data(), sigs_section)
.ok_or(EINVAL)
.and_then(|data| Coherent::from_slice(dev, data, GFP_KERNEL))?
},
signatures,
bootloader: {
let bl = super::request_firmware(dev, chipset, "bootloader", ver)?;
let bl = request_tlv(dev, chipset, "gsp_bootloader")?;
RiscvFirmware::new(dev, &bl)?
},
@ -161,9 +140,9 @@ pub(crate) fn new<'a>(
})
}
/// Returns the DMA handle of the radix3 level 0 page table.
pub(crate) fn radix3_dma_handle(&self) -> DmaAddress {
self.level0.dma_handle()
/// Returns the DMA address of the radix3 level 0 page table.
pub(crate) fn radix3_dma_address(&self) -> DmaAddress {
self.level0.dma_address()
}
}

View File

@ -7,53 +7,10 @@
device,
dma::Coherent,
firmware::Firmware,
prelude::*,
transmute::FromBytes, //
prelude::*, //
};
use crate::{
firmware::BinFirmware,
num::FromSafeCast, //
};
/// Descriptor for microcode running on a RISC-V core.
#[repr(C)]
#[derive(Debug)]
struct RmRiscvUCodeDesc {
version: u32,
bootloader_offset: u32,
bootloader_size: u32,
bootloader_param_offset: u32,
bootloader_param_size: u32,
riscv_elf_offset: u32,
riscv_elf_size: u32,
app_version: u32,
manifest_offset: u32,
manifest_size: u32,
monitor_data_offset: u32,
monitor_data_size: u32,
monitor_code_offset: u32,
monitor_code_size: u32,
}
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
unsafe impl FromBytes for RmRiscvUCodeDesc {}
impl RmRiscvUCodeDesc {
/// Interprets the header of `bin_fw` as a [`RmRiscvUCodeDesc`] and returns it.
///
/// Fails if the header pointed at by `bin_fw` is not within the bounds of the firmware image.
fn new(bin_fw: &BinFirmware<'_>) -> Result<Self> {
let offset = usize::from_safe_cast(bin_fw.hdr.header_offset);
let end = offset.checked_add(size_of::<Self>()).ok_or(EINVAL)?;
bin_fw
.fw
.get(offset..end)
.and_then(Self::from_bytes_copy)
.ok_or(EINVAL)
}
}
use crate::firmware::tlv::Tlv;
/// A parsed firmware for a RISC-V core, ready to be loaded and run.
pub(crate) struct RiscvFirmware {
@ -72,24 +29,26 @@ pub(crate) struct RiscvFirmware {
impl RiscvFirmware {
/// Parses the RISC-V firmware image contained in `fw`.
pub(crate) fn new(dev: &device::Device<device::Bound>, fw: &Firmware) -> Result<Self> {
let bin_fw = BinFirmware::new(fw)?;
let tlv = Tlv::new(fw.data())?;
dev_dbg!(
dev,
"loaded gsp bootloader firmware v{}\n",
tlv.get_string(b"VERS")?
);
let riscv_desc = RmRiscvUCodeDesc::new(&bin_fw)?;
let code_offset = tlv.get_u32(b"CDOF")?;
let data_offset = tlv.get_u32(b"DAOF")?;
let manifest_offset = tlv.get_u32(b"MFOF")?;
let app_version = tlv.get_u32(b"APPV")?;
let ucode = {
let start = usize::from_safe_cast(bin_fw.hdr.data_offset);
let len = usize::from_safe_cast(bin_fw.hdr.data_size);
let end = start.checked_add(len).ok_or(EINVAL)?;
Coherent::from_slice(dev, fw.data().get(start..end).ok_or(EINVAL)?, GFP_KERNEL)?
};
let ucode = Coherent::from_slice(dev, tlv.get_bytes(b"BLOB")?, GFP_KERNEL)?;
Ok(Self {
ucode,
code_offset: riscv_desc.monitor_code_offset,
data_offset: riscv_desc.monitor_data_offset,
manifest_offset: riscv_desc.manifest_offset,
app_version: riscv_desc.app_version,
code_offset,
data_offset,
manifest_offset,
app_version,
})
}
}

View File

@ -0,0 +1,262 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use kernel::{
device,
firmware,
prelude::*,
str::CString, //
};
use crate::{
gpu,
num::*, //
};
/// Requests the GPU firmware TLV `name` suitable for `chipset`.
pub(crate) fn request_tlv(
dev: &device::Device,
chipset: gpu::Chipset,
name: &str,
) -> Result<firmware::Firmware> {
let chip_name = chipset.name();
let filename = CString::try_from_fmt(fmt!("nvidia/{chip_name}/gsp/{name}.tlv"))?;
dev_dbg!(dev, "loading firmware image {:?}\n", &filename);
firmware::Firmware::request(&filename, dev)
}
struct TlvBlock<'a> {
tag: [u8; 4],
value: &'a [u8],
}
/// On-wire TLV block header: 4-byte ASCII tag + little-endian payload length (bytes, excluding
/// padding to a 4-byte boundary).
struct TlvBlockHeader {
tag: [u8; 4],
length: usize,
}
impl TlvBlockHeader {
const SIZE: usize = size_of::<[u8; 4]>() + size_of::<u32>();
/// Parses the first [`Self::SIZE`] bytes of `hdr` (caller may pass a longer slice).
fn parse(hdr: &[u8]) -> Option<Self> {
let hdr = hdr.get(..Self::SIZE)?;
let tag = <[u8; 4]>::try_from(hdr.get(..4)?).ok()?;
if !tag.is_ascii() {
return None;
}
let len_arr = <[u8; 4]>::try_from(hdr.get(4..Self::SIZE)?).ok()?;
let length = u32_as_usize(u32::from_le_bytes(len_arr));
Some(Self { tag, length })
}
}
/// Iterator over the [`TlvBlock`]s of a [`Tlv`].
///
/// # Invariants
///
/// `pos` is a byte offset into `tlv.data` that always lies on a block boundary (in the sense
/// of the [`Tlv`] invariant): it is either the start of a well-formed block, or equal to
/// `tlv.data.len()` (end of iteration).
struct TlvIter<'tlv, 'a> {
tlv: &'tlv Tlv<'a>,
pos: usize,
}
impl<'tlv, 'a> Iterator for TlvIter<'tlv, 'a> {
type Item = TlvBlock<'a>;
/// Returns the block starting at `self.pos` and advances the cursor past it, or [`None`]
/// once the cursor reaches the end of the data or encounters an error.
///
/// Note that errors cannot actually occur because the data is validated in the constructor.
fn next(&mut self) -> Option<Self::Item> {
if self.pos >= self.tlv.data.len() {
return None;
}
let tail = self.tlv.data.get(self.pos..)?;
let hdr = tail.get(..TlvBlockHeader::SIZE)?;
let header = TlvBlockHeader::parse(hdr)?;
let stored_size = header.length.checked_next_multiple_of(4)?;
let advance = TlvBlockHeader::SIZE.checked_add(stored_size)?;
let payload_end = TlvBlockHeader::SIZE.checked_add(header.length)?;
let value = tail
.get(..advance)?
.get(TlvBlockHeader::SIZE..payload_end)?;
// INVARIANT: by the `Tlv` invariant the block at `self.pos` occupies exactly `advance`
// bytes, so `self.pos + advance` is the next block boundary (or `data.len()`).
self.pos = self.pos.checked_add(advance)?;
Some(TlvBlock {
tag: header.tag,
value,
})
}
}
/// The post-header part of a validated TLV (type, length, value) firmware image.
///
/// TLV firmware images start with a 4-byte "NVFW" magic header, followed by a sequence of
/// blocks. Each block has a 4-byte type tag, a 4-byte length field, and a data payload
/// (value) whose stored size is the length rounded up to the nearest multiple of 4.
///
/// [`Self::new`] checks the magic header and walks every block: tags must be ASCII,
/// lengths and padding must fit without overflow, and the byte stream after `NVFW` must
/// be exactly partitionable into blocks (no trailing partial header or slack). After
/// that, [`TlvIter`] only signals end-of-stream via [`None`], not parse failure.
///
/// Although the spec forbids duplicate tags, neither the constructor nor the iterator
/// enforces this restriction. Instead, duplicate tags are simply ignored.
///
/// # Invariants
///
/// `data` is a validated TLV payload (the bytes *after* the `NVFW` magic): it is the exact
/// concatenation of zero or more well-formed blocks, with no trailing partial header or slack.
/// Consequently, any offset `o` into `data` that is a block boundary and satisfies
/// `o < data.len()` is the start of a complete block whose header parses and whose stored
/// extent (`TlvBlockHeader::SIZE + header.length.next_multiple_of(4)` bytes) lies within
/// `data`. `data.len()` is itself a boundary.
pub(crate) struct Tlv<'a> {
data: &'a [u8],
}
impl<'a> Tlv<'a> {
const MAGIC: &'static [u8; 4] = b"NVFW";
/// Parses `data` as a TLV firmware image, returning [`EINVAL`] if the image is malformed.
pub(crate) fn new(data: &'a [u8]) -> Result<Self> {
// Verify that the magic bytes exist and are the correct value
let magic_len = Self::MAGIC.len();
if data
.get(..magic_len)
.is_none_or(|magic| magic != Self::MAGIC)
{
return Err(EINVAL);
}
// The payload is the contiguous sequence of TLV blocks after the magic.
let payload = data.get(magic_len..).ok_or(EINVAL)?;
// The spec says every TLV must have a VERS tag.
let mut has_vers = false;
let mut rest = payload;
while !rest.is_empty() {
// Validate and extract the header (type, length).
let Some(header): Option<TlvBlockHeader> = rest
.get(..TlvBlockHeader::SIZE)
.and_then(TlvBlockHeader::parse)
else {
return Err(EINVAL);
};
has_vers |= header.tag == *b"VERS";
// The `length` field of a TLV block contains the actual byte length of the
// value, but each TLV block is aligned to a 4-byte boundary.
let Some(stored_size) = header.length.checked_next_multiple_of(4) else {
return Err(EINVAL);
};
let length = TlvBlockHeader::SIZE
.checked_add(stored_size)
.ok_or(EINVAL)?;
rest = rest.split_at_checked(length).ok_or(EINVAL)?.1;
}
if !has_vers {
return Err(EINVAL);
}
// INVARIANT: the loop above walked `payload` block-by-block. For each block, the
// header is parsed (`TlvBlockHeader::parse` rejects non-ASCII tags), and the
// stored extent (`SIZE + length.next_multiple_of(4)`) is computed without
// overflow and split off `rest` only when it fits. The loop ends only when `rest`
// is empty, so the byte stream is an exact concatenation of blocks with no
// trailing partial header or slack.
Ok(Self { data: payload })
}
fn iter(&self) -> TlvIter<'_, 'a> {
// INVARIANT: 0 is a block boundary, either the start of the first block,
// or `data.len()` when `data` is empty.
TlvIter { tlv: self, pos: 0 }
}
fn find(&self, tag: &[u8; 4]) -> Result<TlvBlock<'a>> {
self.iter().find(|b| b.tag == *tag).ok_or(EINVAL)
}
/// Return a slice of bytes.
///
/// Returns `EINVAL` if the value is empty.
pub(crate) fn get_bytes(&self, tag: &[u8; 4]) -> Result<&'a [u8]> {
let tlv = self.find(tag)?;
// Treat empty value as an error, to avoid trying to parse nothing.
if tlv.value.is_empty() {
return Err(EINVAL); // TODO: Use ENODATA once available.
}
Ok(tlv.value)
}
/// Return a little-endian u32.
pub(crate) fn get_u32(&self, tag: &[u8; 4]) -> Result<u32> {
let tlv = self.find(tag)?;
tlv.value
.try_into()
.ok()
.map(u32::from_le_bytes)
.ok_or(EINVAL)
}
/// Return a string value.
pub(crate) fn get_string(&self, tag: &[u8; 4]) -> Result<&'a str> {
let tlv = self.find(tag)?;
let bytes = tlv.value;
// Strings can only contain printable ASCII characters.
if bytes.iter().any(|&b| !(32..127).contains(&b)) {
return Err(EINVAL);
}
core::str::from_utf8(bytes).map_err(|_| EINVAL)
}
/// Obtain the nth signature from a SIGN tag. If `index` is None,
/// then return the last signature.
pub(crate) fn get_signature(&self, index: Option<usize>) -> Result<&'a [u8]> {
let num_sigs: usize = match self.get_u32(b"NSIG")? {
0 => return Err(EINVAL),
n => n.into_safe_cast(),
};
let sig_bytes = self.get_bytes(b"SIGN")?;
// Ensure that sig_bytes can be divided evenly into chunks.
if sig_bytes.len() % num_sigs != 0 {
return Err(EINVAL);
}
// num_sigs cannot be 0, and sig_bytes cannot be empty, so this cannot panic.
let sig_size = sig_bytes.len() / num_sigs;
let index = index.unwrap_or(num_sigs - 1);
sig_bytes.chunks_exact(sig_size).nth(index).ok_or(EINVAL)
}
}

View File

@ -11,6 +11,7 @@
device,
dma::Coherent,
io::poll::read_poll_timeout,
num::TryIntoBounded,
prelude::*,
ptr::{
Alignable,
@ -30,13 +31,17 @@
fsp::Fsp as FspEngine,
Falcon, //
},
fb::FbLayout,
fb::FbSizes,
firmware::fsp::{
FmcSignatures,
FspFirmware, //
},
gpu::Chipset,
gsp::GspFmcBootParams,
gsp::{
GspFmcBootParams,
GspFwWprMeta,
LibosMemoryRegionInitArgument, //
},
mctp::{
MctpHeader,
NvdmHeader,
@ -48,6 +53,56 @@
mod hal;
/// PRC message sub-command.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
#[repr(u8)]
enum PrcMessageSubcmd {
/// Read a PRC knob value.
Read = 0x0c,
}
impl From<PrcMessageSubcmd> for u8 {
fn from(value: PrcMessageSubcmd) -> Self {
value as u8
}
}
/// PRC object identifier.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
#[repr(u8)]
enum PrcObjectId {
/// vGPU mode configuration knob.
VgpuMode = 0x29,
}
impl From<PrcObjectId> for u8 {
fn from(value: PrcObjectId) -> Self {
value as u8
}
}
kernel::impl_flags!(
/// PRC request flags.
#[derive(Clone, Copy, Default, PartialEq, Eq)]
struct PrcFlags(u8);
/// Individual PRC request flag.
#[derive(Clone, Copy, PartialEq, Eq)]
enum PrcFlag {
/// Request the active knob value for the current boot.
Active = 1 << 1,
}
);
/// vGPU operating mode as reported by FSP via the PRC protocol.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(crate) enum VgpuMode {
/// vGPU support is disabled on this GPU.
Disabled,
/// vGPU support is enabled on this GPU.
Enabled,
}
/// FSP command response payload (`NVDM_PAYLOAD_COMMAND_RESPONSE`).
#[repr(C, packed)]
#[derive(Clone, Copy)]
@ -57,17 +112,107 @@ struct NvdmPayloadCommandResponse {
error_code: u32,
}
/// Complete FSP response structure with MCTP and NVDM headers.
/// PRC message payload.
///
/// Sent to FSP to query or modify a device configuration knob.
#[repr(C, packed)]
#[derive(Clone, Copy)]
struct FspResponse {
struct NvdmPayloadPrc {
sub_message_id: u8,
flags: u8,
object_id: u8,
reserved: u8,
}
impl NvdmPayloadPrc {
/// Constructs a PRC payload from typed protocol fields.
fn new(subcmd: PrcMessageSubcmd, object_id: PrcObjectId, flags: PrcFlags) -> Self {
Self {
sub_message_id: subcmd.into(),
flags: flags.into(),
object_id: object_id.into(),
reserved: 0,
}
}
}
// SAFETY: NvdmPayloadPrc is a packed C struct with only integral fields.
unsafe impl AsBytes for NvdmPayloadPrc {}
/// PRC response payload containing the knob state value.
#[repr(C, packed)]
#[derive(Clone, Copy)]
struct NvdmPayloadPrcResponse {
value_low: u8,
value_high: u8,
reserved1: u8,
reserved2: u8,
}
impl NvdmPayloadPrcResponse {
/// Returns the PRC knob value as a little-endian 16-bit integer.
fn value(self) -> u16 {
u16::from(self.value_low) | (u16::from(self.value_high) << 8)
}
}
impl TryFrom<NvdmPayloadPrcResponse> for VgpuMode {
type Error = kernel::error::Error;
fn try_from(value: NvdmPayloadPrcResponse) -> Result<Self> {
match value.value() {
0 => Ok(VgpuMode::Disabled),
1 => Ok(VgpuMode::Enabled),
_ => Err(EINVAL),
}
}
}
/// Common MCTP and NVDM headers shared by all FSP messages.
#[repr(C, packed)]
#[derive(Clone, Copy)]
struct FspMessageHeader {
mctp_header: MctpHeader,
nvdm_header: NvdmHeader,
}
// SAFETY: FspMessageHeader is a packed C struct with only integral fields.
unsafe impl AsBytes for FspMessageHeader {}
// SAFETY: FspMessageHeader is a packed C struct with only integral fields.
unsafe impl FromBytes for FspMessageHeader {}
impl FspMessageHeader {
/// Construct a standard FSP message header for the given NVDM type.
fn new(nvdm_type: NvdmType) -> Self {
Self {
mctp_header: MctpHeader::single_packet(),
nvdm_header: NvdmHeader::new(nvdm_type),
}
}
}
/// Common FSP response header with MCTP, NVDM and command response payloads.
#[repr(C, packed)]
#[derive(Clone, Copy)]
struct FspResponseHeader {
header: FspMessageHeader,
response: NvdmPayloadCommandResponse,
}
// SAFETY: FspResponse is a packed C struct with only integral fields.
unsafe impl FromBytes for FspResponse {}
// SAFETY: FspResponseHeader is a packed C struct with only integral fields.
unsafe impl FromBytes for FspResponseHeader {}
/// Complete FSP PRC response including the knob state payload.
#[repr(C, packed)]
#[derive(Clone, Copy)]
struct FspPrcResponse {
header: FspResponseHeader,
prc_data: NvdmPayloadPrcResponse,
}
// SAFETY: FspPrcResponse is a packed C struct with only integral fields.
unsafe impl FromBytes for FspPrcResponse {}
/// Trait implemented by types representing a message to send to FSP.
///
@ -94,45 +239,56 @@ struct NvdmPayloadCot {
gsp_boot_args_sysmem_offset: u64,
}
/// Complete FSP message structure with MCTP and NVDM headers.
/// Complete FSP COT (Chain of Trust) message structure.
#[repr(C)]
#[derive(Clone, Copy)]
struct FspMessage {
mctp_header: MctpHeader,
nvdm_header: NvdmHeader,
struct FspCotMessage {
header: FspMessageHeader,
cot: NvdmPayloadCot,
}
impl FspMessage {
/// Returns an in-place initializer for [`FspMessage`].
fn new<'a>(
fb_layout: &FbLayout,
fsp_fw: &'a FspFirmware,
args: &'a FmcBootArgs,
) -> Result<impl Init<Self> + 'a> {
// frts_offset is relative to FB end: FRTS_location = FB_END - frts_offset
let frts_vidmem_offset = if !args.resume {
let frts_reserved_size = fb_layout.heap.len() + u64::from(fb_layout.pmu_reserved_size);
impl FspCotMessage {
/// Computes the FRTS vidmem offset for the Chain-of-Trust message. It is measured backwards
/// from the end of the framebuffer.
fn frts_vidmem_offset(hal: &dyn hal::FspHal, fb_info: &FbSizes) -> Result<u64> {
let mut offset = hal.fb_end_reserved_size();
frts_reserved_size
// As per OpenRM's `kfspPrepareBootCommands_GH100`.
if fb_info.pmu_reserved_size != 0 {
offset = (offset + u64::from(fb_info.pmu_reserved_size))
// The 2 MiB alignment is r570-specific.
.align_up(Alignment::new::<SZ_2M>())
.ok_or(EINVAL)?
.ok_or(EINVAL)?;
}
Ok(offset)
}
/// Returns an in-place initializer for [`FspCotMessage`].
fn new<'a>(
fb_info: &FbSizes,
fsp_fw: &'a FspFirmware,
args: &'a FmcBootArgs<'_>,
) -> Result<impl Init<Self> + 'a> {
let hal = hal::fsp_hal(args.chipset).ok_or(ENOTSUPP)?;
let frts_vidmem_offset = if !args.resume {
Self::frts_vidmem_offset(hal, fb_info)?
} else {
0
};
let frts_size: u32 = if !args.resume {
fb_layout.frts.len().try_into()?
fb_info.frts_size.try_into()?
} else {
0
};
let version = hal::fsp_hal(args.chipset).ok_or(ENOTSUPP)?.cot_version();
let version = hal.cot_version();
let size = num::usize_into_u16::<{ core::mem::size_of::<NvdmPayloadCot>() }>();
Ok(init!(Self {
mctp_header: MctpHeader::single_packet(),
nvdm_header: NvdmHeader::new(NvdmType::Cot),
header: FspMessageHeader::new(NvdmType::Cot),
// The payload is packed, so we cannot use `init!`. Initialize it member-by-member using
// `chain`.
cot <- pin_init::init_zeroed(),
@ -140,12 +296,12 @@ fn new<'a>(
.chain(move |msg| {
msg.cot.version = version;
msg.cot.size = size;
msg.cot.gsp_fmc_sysmem_offset = fsp_fw.fmc_image.dma_handle();
msg.cot.gsp_fmc_sysmem_offset = fsp_fw.fmc_image.dma_address();
msg.cot.frts_vidmem_offset = frts_vidmem_offset;
msg.cot.frts_vidmem_size = frts_size;
// frts_sysmem_* intentionally left at zero for now, but will be needed for e.g.
// systems without VRAM.
msg.cot.gsp_boot_args_sysmem_offset = args.fmc_boot_params.dma_handle();
// frts_sysmem_* are left at zero because this path places FRTS in vidmem. The sysmem
// fields point to an FRTS buffer in sysmem instead, for systems without VRAM.
msg.cot.gsp_boot_args_sysmem_offset = args.fmc_boot_params.dma_address();
msg.cot.sigs = *fsp_fw.fmc_sigs;
Ok(())
@ -153,44 +309,73 @@ fn new<'a>(
}
}
// SAFETY: `FspMessage` is `#[repr(C)]` with no padding, so all of its
// SAFETY: `FspCotMessage` is `#[repr(C)]` with no padding, so all of its
// bytes are initialized.
unsafe impl AsBytes for FspMessage {}
unsafe impl AsBytes for FspCotMessage {}
impl MessageToFsp for FspMessage {
/// Complete FSP PRC message.
#[repr(C, packed)]
#[derive(Clone, Copy)]
struct FspPrcMessage {
header: FspMessageHeader,
prc: NvdmPayloadPrc,
}
impl FspPrcMessage {
/// Constructs a PRC message.
fn new(subcmd: PrcMessageSubcmd, object_id: PrcObjectId, flags: PrcFlags) -> Self {
Self {
header: FspMessageHeader::new(NvdmType::Prc),
prc: NvdmPayloadPrc::new(subcmd, object_id, flags),
}
}
}
// SAFETY: FspPrcMessage is a packed C struct with only integral fields.
unsafe impl AsBytes for FspPrcMessage {}
impl MessageToFsp for FspCotMessage {
const NVDM_TYPE: NvdmType = NvdmType::Cot;
}
impl MessageToFsp for FspPrcMessage {
const NVDM_TYPE: NvdmType = NvdmType::Prc;
}
/// Bundled arguments for FMC boot via FSP Chain of Trust.
pub(crate) struct FmcBootArgs {
pub(crate) struct FmcBootArgs<'a> {
chipset: Chipset,
fmc_boot_params: Coherent<GspFmcBootParams>,
resume: bool,
// Additional dependencies required to be kept alive for FMC boot.
_wpr_meta: Coherent<GspFwWprMeta>,
_libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
}
impl FmcBootArgs {
impl<'a> FmcBootArgs<'a> {
/// Builds FMC boot arguments, allocating the DMA-coherent boot parameter
/// structure that FSP will read.
pub(crate) fn new(
dev: &device::Device<device::Bound>,
chipset: Chipset,
wpr_meta_addr: u64,
libos_addr: u64,
wpr_meta: Coherent<GspFwWprMeta>,
libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
resume: bool,
) -> Result<Self> {
let init = GspFmcBootParams::new(wpr_meta_addr, libos_addr);
let init = GspFmcBootParams::new(wpr_meta.dma_address(), libos.dma_address());
Ok(Self {
chipset,
fmc_boot_params: Coherent::<GspFmcBootParams>::init(dev, GFP_KERNEL, init)?,
resume,
_wpr_meta: wpr_meta,
_libos: libos,
})
}
/// DMA address of the FMC boot parameters, needed after boot for lockdown
/// release polling.
pub(crate) fn boot_params_dma_handle(&self) -> u64 {
self.fmc_boot_params.dma_handle()
/// Returns the FMC boot parameters allocation.
pub(crate) fn boot_params(&self) -> &Coherent<GspFmcBootParams> {
&self.fmc_boot_params
}
}
@ -199,28 +384,45 @@ pub(crate) fn boot_params_dma_handle(&self) -> u64 {
/// An `Fsp` is produced by [`Fsp::wait_secure_boot`], which only returns once FSP secure boot
/// has completed. It owns the FSP falcon and the FMC firmware, which are used for the subsequent
/// Chain of Trust boot.
pub(crate) struct Fsp {
falcon: Falcon<FspEngine>,
pub(crate) struct Fsp<'a> {
falcon: Falcon<'a, FspEngine>,
fsp_fw: FspFirmware,
}
impl Fsp {
impl<'a> Fsp<'a> {
/// Attempts to create a `Fsp` instance.
///
/// This can involve waiting for FSP secure boot completion, but should be instantaneous in
/// practice.
///
/// If `chipset` doesn't support FSP, `Ok(None)` is returned.
pub(crate) fn try_new(
dev: &'a device::Device<device::Bound>,
bar: Bar0<'a>,
chipset: Chipset,
) -> Result<Option<Self>> {
match hal::fsp_hal(chipset) {
None => Ok(None),
Some(hal) => Self::wait_secure_boot(dev, bar, chipset, hal).map(Option::Some),
}
}
/// Waits for FSP secure boot completion, then returns the [`Fsp`] interface.
///
/// Polls the thermal scratch register until FSP signals boot completion or the timeout
/// elapses. Returning an [`Fsp`] only on success guarantees, at the API level, that the
/// interface is not used before secure boot has completed.
pub(crate) fn wait_secure_boot(
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
fn wait_secure_boot(
dev: &'a device::Device<device::Bound>,
bar: Bar0<'a>,
chipset: Chipset,
fsp_fw: FspFirmware,
) -> Result<Fsp> {
hal: &'static dyn hal::FspHal,
) -> Result<Fsp<'a>> {
/// FSP secure boot completion timeout in milliseconds.
const FSP_SECURE_BOOT_TIMEOUT_MS: i64 = 5000;
let hal = hal::fsp_hal(chipset).ok_or(ENOTSUPP)?;
let falcon = Falcon::<FspEngine>::new(dev, chipset)?;
let falcon = Falcon::<FspEngine>::new(dev, chipset, bar)?;
let fsp_fw = FspFirmware::new(dev, chipset)?;
read_poll_timeout(
|| Ok(hal.fsp_boot_status(bar)),
@ -236,23 +438,25 @@ pub(crate) fn wait_secure_boot(
}
/// Sends a message to FSP and waits for the response.
fn send_sync_fsp<M>(&mut self, dev: &device::Device, bar: Bar0<'_>, msg: &M) -> Result
/// Returns the full response buffer on success.
fn send_sync_fsp<M>(&mut self, dev: &device::Device, msg: &M) -> Result<KVec<u8>>
where
M: MessageToFsp,
{
self.falcon.send_msg(bar, msg.as_bytes())?;
self.falcon.send_msg(msg.as_bytes())?;
let response_buf = self.falcon.recv_msg(bar).inspect_err(|e| {
let response_buf = self.falcon.recv_msg().inspect_err(|e| {
dev_err!(dev, "FSP response error: {:?}\n", e);
})?;
let (response, _) = FspResponse::from_bytes_prefix(&response_buf[..]).ok_or_else(|| {
dev_err!(dev, "FSP response too small: {}\n", response_buf.len());
EIO
})?;
let (response, _) =
FspResponseHeader::from_bytes_prefix(&response_buf[..]).ok_or_else(|| {
dev_err!(dev, "FSP response too small: {}\n", response_buf.len());
EIO
})?;
let mctp_header = response.mctp_header;
let nvdm_header = response.nvdm_header;
let mctp_header = response.header.mctp_header;
let nvdm_header = response.header.nvdm_header;
let command_nvdm_type = response.response.command_nvdm_type;
let error_code = response.response.error_code;
@ -274,7 +478,7 @@ fn send_sync_fsp<M>(&mut self, dev: &device::Device, bar: Bar0<'_>, msg: &M) ->
return Err(EIO);
}
if command_nvdm_type != u8::from(M::NVDM_TYPE).into() {
if command_nvdm_type.try_into_bounded() != Some(M::NVDM_TYPE.into()) {
dev_err!(
dev,
"Expected NVDM type {:?} in reply, got {:#x}\n",
@ -294,7 +498,34 @@ fn send_sync_fsp<M>(&mut self, dev: &device::Device, bar: Bar0<'_>, msg: &M) ->
return Err(EIO);
}
Ok(())
Ok(response_buf)
}
/// Reads the active vGPU mode from FSP using the PRC protocol.
///
/// Queries FSP's Management Partition for the active vGPU mode knob value.
pub(crate) fn read_vgpu_mode(
&mut self,
dev: &device::Device<device::Bound>,
) -> Result<VgpuMode> {
let msg = FspPrcMessage::new(
PrcMessageSubcmd::Read,
PrcObjectId::VgpuMode,
PrcFlags::from(PrcFlag::Active),
);
let response_buf = self.send_sync_fsp(dev, &msg)?;
let (prc_response, _) =
FspPrcResponse::from_bytes_prefix(&response_buf[..]).ok_or_else(|| {
dev_err!(dev, "PRC response too small: {}\n", response_buf.len());
EIO
})?;
let prc_data = prc_response.prc_data;
VgpuMode::try_from(prc_data).inspect_err(|_| {
dev_err!(dev, "Unexpected vGPU mode value: {:#x}\n", prc_data.value());
})
}
/// Boots GSP FMC via FSP Chain of Trust.
@ -304,15 +535,14 @@ fn send_sync_fsp<M>(&mut self, dev: &device::Device, bar: Bar0<'_>, msg: &M) ->
pub(crate) fn boot_fmc(
&mut self,
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
fb_layout: &FbLayout,
args: &FmcBootArgs,
fb_info: &FbSizes,
args: &FmcBootArgs<'_>,
) -> Result {
dev_dbg!(dev, "Starting FSP boot sequence for {}\n", args.chipset);
let msg = KBox::init(FspMessage::new(fb_layout, &self.fsp_fw, args)?, GFP_KERNEL)?;
let msg = KBox::init(FspCotMessage::new(fb_info, &self.fsp_fw, args)?, GFP_KERNEL)?;
self.send_sync_fsp(dev, bar, &*msg)?;
let _response_buf = self.send_sync_fsp(dev, &*msg)?;
dev_dbg!(dev, "FSP Chain of Trust completed successfully\n");
Ok(())

View File

@ -19,6 +19,10 @@ pub(super) trait FspHal {
/// Returns the FSP Chain of Trust protocol version this chipset advertises.
fn cot_version(&self) -> u16;
// TODO: consider moving this into the TLV firmware metadata when ready
/// Returns the size reserved at the end of the framebuffer, in bytes.
fn fb_end_reserved_size(&self) -> u64;
}
/// Returns the FSP HAL, or `None` if the architecture doesn't support FSP.

View File

@ -1,6 +1,8 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use kernel::sizes::SizeConstants;
use crate::{
driver::Bar0,
fsp::hal::FspHal, //
@ -17,6 +19,10 @@ fn fsp_boot_status(&self, bar: Bar0<'_>) -> u32 {
fn cot_version(&self) -> u16 {
2
}
fn fb_end_reserved_size(&self) -> u64 {
u64::SZ_2M + u64::SZ_128K
}
}
const GB100: Gb100 = Gb100;

View File

@ -1,7 +1,10 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use kernel::io::Io;
use kernel::{
io::Io,
sizes::SizeConstants, //
};
use crate::{
driver::Bar0,
@ -21,6 +24,10 @@ fn fsp_boot_status(&self, bar: Bar0<'_>) -> u32 {
fn cot_version(&self) -> u16 {
2
}
fn fb_end_reserved_size(&self) -> u64 {
u64::SZ_2M + u64::SZ_128K
}
}
const GB202: Gb202 = Gb202;

View File

@ -1,7 +1,10 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use kernel::io::Io;
use kernel::{
io::Io,
sizes::SizeConstants, //
};
use crate::{
driver::Bar0,
@ -26,6 +29,10 @@ fn fsp_boot_status(&self, bar: Bar0<'_>) -> u32 {
fn cot_version(&self) -> u16 {
1
}
fn fb_end_reserved_size(&self) -> u64 {
u64::SZ_2M
}
}
const GH100: Gh100 = Gh100;

View File

@ -9,7 +9,8 @@
io::Io,
num::Bounded,
pci,
prelude::*, //
prelude::*,
sizes::SizeConstants, //
};
use crate::{
@ -21,11 +22,15 @@
Falcon, //
},
fb::SysmemFlush,
fsp::Fsp,
gsp::{
self,
Gsp, //
commands::GetGspStaticInfoReply,
Gsp,
GspBootContext, //
},
regs,
vgpu::VgpuManager, //
};
mod hal;
@ -130,22 +135,6 @@ pub(crate) const fn arch(self) -> Architecture {
}
}
/// Returns `true` if this chipset requires the PIO-loaded bootloader in order to boot FWSEC.
///
/// This includes all chipsets < GA102.
pub(crate) const fn needs_fwsec_bootloader(self) -> bool {
matches!(self.arch(), Architecture::Turing) || matches!(self, Self::GA100)
}
/// Returns `true` if this chipset boots via FSP (Hopper and later), which requires the FMC
/// firmware image.
pub(crate) const fn uses_fsp(self) -> bool {
matches!(
self.arch(),
Architecture::Hopper | Architecture::BlackwellGB10x | Architecture::BlackwellGB20x
)
}
/// Returns the address range of the PCI config mirror space.
pub(crate) fn pci_config_mirror_range(self) -> Range<u32> {
hal::gpu_hal(self).pci_config_mirror_range()
@ -262,37 +251,87 @@ fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
}
}
/// Structure holding the resources required to operate the GPU.
/// Self-contained resources to operate and drop the GSP.
#[pin_data(PinnedDrop)]
pub(crate) struct Gpu<'gpu> {
struct GspResources<'gpu> {
/// Device owning the GPU.
device: &'gpu device::Device<device::Bound>,
device: &'gpu pci::Device<device::Bound>,
/// Details about the chipset.
spec: Spec,
/// MMIO mapping of PCI BAR 0.
bar: Bar0<'gpu>,
/// System memory page required for flushing all pending GPU-side memory writes done through
/// PCIE into system memory, via sysmembar (A GPU-initiated HW memory-barrier operation).
sysmem_flush: SysmemFlush<'gpu>,
/// GSP falcon instance, used for GSP boot up and cleanup.
gsp_falcon: Falcon<GspFalcon>,
gsp_falcon: Falcon<'gpu, GspFalcon>,
/// SEC2 falcon instance, used for GSP boot up and cleanup.
sec2_falcon: Falcon<Sec2Falcon>,
/// GSP runtime data. Temporarily an empty placeholder.
sec2_falcon: Falcon<'gpu, Sec2Falcon>,
/// FSP instance, if on an arch that supports it.
// TODO: use different resource types for each boot method, and make the relevant Gsp methods
// generic against them.
fsp: Option<Fsp<'gpu>>,
/// vGPU state detected before GSP boot.
vgpu: VgpuManager,
/// GSP runtime data.
#[pin]
gsp: Gsp,
/// GSP unload firmware bundle, if any.
unload_bundle: Option<gsp::UnloadBundle>,
}
/// Structure holding the resources required to operate the GPU.
#[pin_data]
pub(crate) struct Gpu<'gpu> {
spec: Spec,
/// Static GPU information as provided by the GSP.
gsp_static_info: GetGspStaticInfoReply,
/// GSP and its resources.
#[pin]
gsp_resources: GspResources<'gpu>,
/// System memory page required for flushing all pending GPU-side memory writes done through
/// PCIE into system memory, via sysmembar (A GPU-initiated HW memory-barrier operation).
///
/// Must be kept declared *after* `gsp_resources`, as the latter's `PinnedDrop` implementation
/// requires the sysmem flush page to be in place.
sysmem_flush: SysmemFlush<'gpu>,
}
#[pinned_drop]
impl PinnedDrop for GspResources<'_> {
fn drop(self: Pin<&mut Self>) {
let this = self.project();
let device = *this.device;
let bar = *this.bar;
let bundle = this.unload_bundle.take();
let _ = this
.gsp
.as_ref()
.get_ref()
.unload(
GspBootContext {
pdev: device,
bar,
chipset: this.spec.chipset,
gsp_falcon: &*this.gsp_falcon,
sec2_falcon: &*this.sec2_falcon,
fsp: this.fsp.as_mut(),
vgpu: &*this.vgpu,
},
bundle,
)
.inspect_err(|e| dev_err!(device, "failed to unload GSP: {:?}\n", e));
}
}
impl<'gpu> Gpu<'gpu> {
pub(crate) fn new(
pdev: &'gpu pci::Device<device::Core<'_>>,
bar: Bar0<'gpu>,
) -> impl PinInit<Self, Error> + 'gpu {
let dev = pdev.as_ref();
try_pin_init!(Self {
device: pdev.as_ref(),
spec: Spec::new(pdev.as_ref(), bar).inspect(|spec| {
dev_info!(pdev,"NVIDIA ({})\n", spec);
spec: Spec::new(dev, bar).inspect(|spec| {
dev_info!(dev,"NVIDIA ({})\n", spec);
})?,
// We must wait for GFW_BOOT completion before doing any significant setup on the GPU.
@ -305,43 +344,73 @@ pub(crate) fn new(
unsafe { pdev.dma_set_mask_and_coherent(dma_mask)? };
hal.wait_gfw_boot_completion(bar)
.inspect_err(|_| dev_err!(pdev, "GFW boot did not complete\n"))?;
.inspect_err(|_| dev_err!(dev, "GFW boot did not complete\n"))?;
},
sysmem_flush: SysmemFlush::register(pdev.as_ref(), bar, spec.chipset)?,
// Initialize this early because `gsp_resources` depends on it.
sysmem_flush: SysmemFlush::register(dev, bar, spec.chipset)?,
gsp_falcon: Falcon::new(
pdev.as_ref(),
spec.chipset,
)
.inspect(|falcon| falcon.clear_swgen0_intr(bar))?,
gsp_resources <- try_pin_init!(GspResources {
device: pdev,
sec2_falcon: Falcon::new(pdev.as_ref(), spec.chipset)?,
spec: *spec,
gsp <- Gsp::new(pdev),
bar,
// This member must be initialized last, so the `UnloadBundle` can never be dropped from
// outside of the constructed `Gpu`, ensuring that the unload sequence is properly run
// in case of failure.
unload_bundle: gsp.boot(pdev, bar, spec.chipset, gsp_falcon, sec2_falcon)?,
bar,
gsp_falcon: Falcon::new(
dev,
spec.chipset,
bar
)
.inspect(|falcon| falcon.clear_swgen0_intr())?,
sec2_falcon: Falcon::new(dev, spec.chipset, bar)?,
fsp: Fsp::try_new(dev, bar, spec.chipset)?,
vgpu: VgpuManager::new(pdev, spec.chipset, fsp.as_mut()),
gsp <- Gsp::new(pdev),
// This member must be initialized last, so the `UnloadBundle` can never be dropped
// from outside of the constructed `GspResources`, ensuring that the unload sequence
// is properly run in case of failure.
unload_bundle: gsp.boot(GspBootContext {
pdev,
bar,
chipset: spec.chipset,
gsp_falcon,
sec2_falcon,
fsp: fsp.as_mut(),
vgpu,
})?,
}),
gsp_static_info: {
// Obtain and display basic GPU information.
let info = gsp_resources.gsp.get_static_info(bar)?;
match info.gpu_name() {
Ok(name) => dev_info!(dev, "GPU name: {}\n", name),
Err(e) => dev_warn!(dev, "GPU name unavailable: {:?}\n", e),
}
if !info.usable_fb_regions.is_empty() {
dev_dbg!(dev, "Usable FB regions:\n");
for region in &info.usable_fb_regions {
dev_dbg!(dev, " - {:#x?}\n", region);
}
dev_dbg!(
dev,
"Total usable VRAM: {} MiB\n",
info.usable_fb_regions.iter().fold(0u64, |res, region| res
.saturating_add(region.end - region.start))
/ u64::SZ_1M
);
}
info
}
})
}
}
#[pinned_drop]
impl PinnedDrop for Gpu<'_> {
fn drop(self: Pin<&mut Self>) {
let this = self.project();
let device = *this.device;
let bar = *this.bar;
let bundle = this.unload_bundle.take();
let _ = this
.gsp
.as_ref()
.get_ref()
.unload(device, bar, &*this.gsp_falcon, &*this.sec2_falcon, bundle)
.inspect_err(|e| dev_err!(device, "failed to unload GSP: {:?}\n", e));
}
}

View File

@ -9,60 +9,95 @@
dma::{
Coherent,
CoherentBox,
CoherentView,
DmaAddress, //
},
io::{
io_project,
io_write,
Io, //
},
pci,
prelude::*,
transmute::{
AsBytes,
FromBytes, //
}, //
prelude::*, //
};
pub(crate) mod cmdq;
pub(crate) mod commands;
mod fw;
mod regs;
mod sequencer;
pub(crate) use fw::{
GspFmcBootParams,
GspFwWprMeta,
LibosMemoryRegionInitArgument,
LibosParams, //
};
pub(crate) use hal::boot_firmware_files;
use crate::{
gsp::cmdq::Cmdq,
gsp::fw::{
GspArgumentsPadded,
LibosMemoryRegionInitArgument, //
driver::Bar0,
falcon::{
gsp::Gsp as GspFalcon,
sec2::Sec2 as Sec2Falcon,
Falcon, //
},
fsp::Fsp,
gpu::Chipset,
gsp::{
cmdq::Cmdq,
fw::GspArgumentsPadded, //
},
num,
vgpu::VgpuManager, //
};
pub(crate) const GSP_PAGE_SHIFT: usize = 12;
pub(crate) const GSP_PAGE_SIZE: usize = 1 << GSP_PAGE_SHIFT;
/// Common context for the GSP boot process.
///
/// It carries two distinct lifetimes:
///
/// - `'gpu` is the lifetime of the bound GPU device, as captured by the GPU subdevices.
/// - `'ctx` is a shorter lifetime during which this context borrows those subdevices.
pub(crate) struct GspBootContext<'ctx, 'gpu> {
pub(crate) pdev: &'gpu pci::Device<device::Bound>,
pub(crate) bar: Bar0<'gpu>,
pub(crate) chipset: Chipset,
pub(crate) gsp_falcon: &'ctx Falcon<'gpu, GspFalcon>,
pub(crate) sec2_falcon: &'ctx Falcon<'gpu, Sec2Falcon>,
pub(crate) fsp: Option<&'ctx mut Fsp<'gpu>>,
pub(crate) vgpu: &'ctx VgpuManager,
}
impl<'ctx, 'gpu> GspBootContext<'ctx, 'gpu> {
pub(crate) fn dev(&self) -> &'gpu device::Device<device::Bound> {
self.pdev.as_ref()
}
}
/// Number of GSP pages to use in a RM log buffer.
const RM_LOG_BUFFER_NUM_PAGES: usize = 0x10;
const LOG_BUFFER_SIZE: usize = RM_LOG_BUFFER_NUM_PAGES * GSP_PAGE_SIZE;
/// Array of page table entries, as understood by the GSP bootloader.
#[repr(C)]
#[derive(FromBytes, IntoBytes)]
struct PteArray<const NUM_ENTRIES: usize>([u64; NUM_ENTRIES]);
/// SAFETY: arrays of `u64` implement `FromBytes` and we are but a wrapper around one.
unsafe impl<const NUM_ENTRIES: usize> FromBytes for PteArray<NUM_ENTRIES> {}
/// SAFETY: arrays of `u64` implement `AsBytes` and we are but a wrapper around one.
unsafe impl<const NUM_ENTRIES: usize> AsBytes for PteArray<NUM_ENTRIES> {}
impl<const NUM_PAGES: usize> PteArray<NUM_PAGES> {
/// Returns the page table entry for `index`, for a mapping starting at `start`.
// TODO: Replace with `IoView` projection once available.
fn entry(start: DmaAddress, index: usize) -> Result<u64> {
start
.checked_add(num::usize_as_u64(index) << GSP_PAGE_SHIFT)
.ok_or(EOVERFLOW)
/// Initialize a new page table array mapping `NUM_PAGES` GSP pages starting at address `start`.
fn init(view: CoherentView<'_, Self>, start: DmaAddress) -> Result<()> {
for i in 0..NUM_PAGES {
io_write!(view, .0[build: i],
start
.checked_add(num::usize_as_u64(i) << GSP_PAGE_SHIFT)
.ok_or(EOVERFLOW)?
);
}
Ok(())
}
}
@ -87,19 +122,14 @@ impl LogBuffer {
fn new(dev: &device::Device<device::Bound>) -> Result<Self> {
let obj = Self(Coherent::zeroed(dev, GFP_KERNEL)?);
let start_addr = obj.0.dma_handle();
let start_addr = obj.0.dma_address();
// SAFETY: `obj` has just been created and we are its sole user.
let pte_region = unsafe {
&mut obj.0.as_mut()[size_of::<u64>()..][..RM_LOG_BUFFER_NUM_PAGES * size_of::<u64>()]
};
// Write values one by one to avoid an on-stack instance of `PteArray`.
for (i, chunk) in pte_region.chunks_exact_mut(size_of::<u64>()).enumerate() {
let pte_value = PteArray::<0>::entry(start_addr, i)?;
chunk.copy_from_slice(&pte_value.to_ne_bytes());
}
let pte_view = io_project!(
obj.0,
[build: size_of::<u64>()..][build: ..RM_LOG_BUFFER_NUM_PAGES * size_of::<u64>()]
)
.try_cast::<PteArray<RM_LOG_BUFFER_NUM_PAGES>>()?;
PteArray::init(pte_view, start_addr)?;
Ok(obj)
}
@ -185,6 +215,11 @@ pub(crate) fn new(pdev: &pci::Device<device::Bound>) -> impl PinInit<Self, Error
}))
})
}
/// Query the GSP for the static GPU information.
pub(crate) fn get_static_info(&self, bar: Bar0<'_>) -> Result<commands::GetGspStaticInfoReply> {
self.cmdq.send_command(bar, commands::GetGspStaticInfo)
}
}
/// Opaque bundle required to unload the GSP. Created by [`Gsp::boot`], consumed by [`Gsp::unload`].

View File

@ -3,10 +3,7 @@
use kernel::{
bits,
device,
dma::Coherent,
io::poll::read_poll_timeout,
pci,
prelude::*,
time::Delta,
types::ScopeGuard, //
@ -16,82 +13,15 @@
driver::Bar0,
falcon::{
gsp::Gsp,
sec2::Sec2,
Falcon, //
},
fb::FbLayout,
firmware::{
gsp::GspFirmware,
FIRMWARE_VERSION, //
},
gpu::Chipset,
firmware::gsp::GspFirmware,
gsp::{
cmdq::Cmdq,
commands,
GspFwWprMeta, //
commands, //
},
};
/// Arguments required to call [`Gsp::unload`](super::Gsp::unload).
///
/// Stored as their own type to avoid repeating a long and tedious list in [`BootUnloadGuard`].
pub(super) struct BootUnloadArgs<'a> {
gsp: &'a super::Gsp,
dev: &'a device::Device<device::Bound>,
bar: Bar0<'a>,
gsp_falcon: &'a Falcon<Gsp>,
sec2_falcon: &'a Falcon<Sec2>,
unload_bundle: Option<super::UnloadBundle>,
}
/// Guard that calls [`Gsp::unload`](super::Gsp::unload) with a
/// [`UnloadBundle`](super::UnloadBundle) when dropped.
///
/// Used to ensure the `UnloadBundle` is run during failure paths.
pub(super) struct BootUnloadGuard<'a> {
guard: ScopeGuard<BootUnloadArgs<'a>, fn(BootUnloadArgs<'a>)>,
}
impl<'a> BootUnloadGuard<'a> {
/// Wraps `unload_bundle` into a guard that executes it when dropped.
pub(super) fn new(
gsp: &'a super::Gsp,
dev: &'a device::Device<device::Bound>,
bar: Bar0<'a>,
gsp_falcon: &'a Falcon<Gsp>,
sec2_falcon: &'a Falcon<Sec2>,
unload_bundle: Option<super::UnloadBundle>,
) -> Self {
Self {
guard: ScopeGuard::new_with_data(
BootUnloadArgs {
gsp,
dev,
bar,
gsp_falcon,
sec2_falcon,
unload_bundle,
},
|args| {
let _ = super::Gsp::unload(
args.gsp,
args.dev,
args.bar,
args.gsp_falcon,
args.sec2_falcon,
args.unload_bundle,
);
},
),
}
}
/// Disarms the guard and returns the [`UnloadBundle`](super::UnloadBundle) it contains.
pub(super) fn dismiss(self) -> Option<super::UnloadBundle> {
self.guard.dismiss().unload_bundle
}
}
impl super::Gsp {
/// Attempt to boot the GSP.
///
@ -103,71 +33,64 @@ impl super::Gsp {
/// [`Self::unload`]) returned.
pub(crate) fn boot(
self: Pin<&mut Self>,
pdev: &pci::Device<device::Bound>,
bar: Bar0<'_>,
chipset: Chipset,
gsp_falcon: &Falcon<Gsp>,
sec2_falcon: &Falcon<Sec2>,
mut ctx: super::GspBootContext<'_, '_>,
) -> Result<Option<super::UnloadBundle>> {
let pdev = ctx.pdev;
let bar = ctx.bar;
let chipset = ctx.chipset;
let gsp_falcon = ctx.gsp_falcon;
let dev = pdev.as_ref();
let hal = super::hal::gsp_hal(chipset);
let gsp_fw = KBox::pin_init(GspFirmware::new(dev, chipset, FIRMWARE_VERSION), GFP_KERNEL)?;
let fb_layout = FbLayout::new(chipset, bar, &gsp_fw)?;
dev_dbg!(dev, "{:#x?}\n", fb_layout);
let wpr_meta = Coherent::init(dev, GFP_KERNEL, GspFwWprMeta::new(&gsp_fw, &fb_layout))?;
let gsp_fw = KBox::pin_init(GspFirmware::new(dev, chipset), GFP_KERNEL)?;
// Perform the chipset-specific boot sequence, and retrieve the unload bundle.
let unload_guard = hal.boot(
&self,
dev,
bar,
chipset,
&fb_layout,
&wpr_meta,
gsp_falcon,
sec2_falcon,
)?;
let unload_bundle = hal.boot(&self, &mut ctx, &gsp_fw)?.or_else(|| {
dev_warn!(dev, "The GSP won't be able to unload properly on unbind.\n");
dev_warn!(
dev,
"The GPU will need to be reset before the driver can bind again.\n"
);
gsp_falcon.write_os_version(bar, gsp_fw.bootloader.app_version);
None
});
let mut unload_guard =
ScopeGuard::new_with_data((ctx, unload_bundle), |(ctx, unload_bundle)| {
let _ = self.unload(ctx, unload_bundle);
});
let ctx = &mut unload_guard.0;
gsp_falcon.write_os_version(gsp_fw.bootloader.app_version);
// Poll for RISC-V to become active before continuing.
read_poll_timeout(
|| Ok(gsp_falcon.is_riscv_active(bar)),
|| Ok(gsp_falcon.is_riscv_active()),
|val: &bool| *val,
Delta::from_millis(10),
Delta::from_secs(5),
)?;
dev_dbg!(pdev, "RISC-V active? {}\n", gsp_falcon.is_riscv_active(bar),);
dev_dbg!(pdev, "RISC-V active? {}\n", gsp_falcon.is_riscv_active(),);
self.cmdq
.send_command_no_wait(bar, commands::SetSystemInfo::new(pdev, chipset))?;
self.cmdq
.send_command_no_wait(bar, commands::SetRegistry::new())?;
.send_command_no_wait(bar, commands::SetRegistry::new(ctx.vgpu.state())?)?;
hal.post_boot(&self, dev, bar, &gsp_fw, gsp_falcon, sec2_falcon)?;
hal.post_boot(&self, ctx, &gsp_fw)?;
// Wait until GSP is fully initialized.
commands::wait_gsp_init_done(&self.cmdq)?;
// Obtain and display basic GPU information.
let info = self.cmdq.send_command(bar, commands::GetGspStaticInfo)?;
match info.gpu_name() {
Ok(name) => dev_info!(pdev, "GPU name: {}\n", name),
Err(e) => dev_warn!(pdev, "GPU name unavailable: {:?}\n", e),
}
Ok(unload_guard.dismiss())
Ok(unload_guard.dismiss().1)
}
/// Shut down the GSP and wait until it is offline.
fn shutdown_gsp(
cmdq: &Cmdq,
bar: Bar0<'_>,
gsp_falcon: &Falcon<Gsp>,
gsp_falcon: &Falcon<'_, Gsp>,
mode: commands::PowerStateLevel,
) -> Result {
// Command to shut the GSP down.
@ -176,7 +99,7 @@ fn shutdown_gsp(
// Wait until GSP signals it is suspended.
const LIBOS_INTERRUPT_PROCESSOR_SUSPENDED: u32 = bits::bit_u32(31);
read_poll_timeout(
|| Ok(gsp_falcon.read_mailbox0(bar)),
|| Ok(gsp_falcon.read_mailbox0()),
|&mb0| mb0 & LIBOS_INTERRUPT_PROCESSOR_SUSPENDED != 0,
Delta::from_millis(10),
Delta::from_secs(5),
@ -189,17 +112,16 @@ fn shutdown_gsp(
/// This stops all activity on the GSP.
pub(crate) fn unload(
&self,
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
gsp_falcon: &Falcon<Gsp>,
sec2_falcon: &Falcon<Sec2>,
mut ctx: super::GspBootContext<'_, '_>,
unload_bundle: Option<super::UnloadBundle>,
) -> Result {
let dev = ctx.dev();
// Shut down the GSP. Keep going even in case of error.
let mut res = Self::shutdown_gsp(
&self.cmdq,
bar,
gsp_falcon,
ctx.bar,
ctx.gsp_falcon,
commands::PowerStateLevel::Level0,
)
.inspect_err(|e| dev_err!(dev, "GSP shutdown failed: {:?}\n", e));
@ -209,7 +131,7 @@ pub(crate) fn unload(
res = res.and(
unload_bundle
.0
.run(dev, bar, gsp_falcon, sec2_falcon)
.run(&mut ctx)
.inspect_err(|e| dev_err!(dev, "Unload bundle failed: {:?}\n", e)),
);
} else {

View File

@ -2,16 +2,23 @@
mod continuation;
use core::mem;
use core::{
mem,
sync::atomic::{
fence,
Ordering, //
},
};
use kernel::{
device,
dma::{
Coherent,
CoherentBox,
DmaAddress, //
},
dma_write,
io::{
io_project,
poll::read_poll_timeout,
Io, //
},
@ -51,10 +58,11 @@
GSP_PAGE_SIZE, //
},
num,
regs,
sbuffer::SBufferIter, //
};
use super::regs;
/// Marker type representing the absence of a reply for a command. Commands using this as their
/// reply type are sent using [`Cmdq::send_command_no_wait`].
pub(crate) struct NoReply;
@ -171,20 +179,18 @@ struct MsgqData {
#[repr(C)]
// There is no struct defined for this in the open-gpu-kernel-source headers.
// Instead it is defined by code in `GspMsgQueuesInit()`.
// TODO: Revert to private once `IoView` projections replace the `gsp_mem` module.
pub(super) struct Msgq {
struct Msgq {
/// Header for sending messages, including the write pointer.
pub(super) tx: MsgqTxHeader,
tx: MsgqTxHeader,
/// Header for receiving messages, including the read pointer.
pub(super) rx: MsgqRxHeader,
rx: MsgqRxHeader,
/// The message queue proper.
msgq: MsgqData,
}
/// Structure shared between the driver and the GSP and containing the command and message queues.
#[repr(C)]
// TODO: Revert to private once `IoView` projections replace the `gsp_mem` module.
pub(super) struct GspMem {
struct GspMem {
/// Self-mapping page table entries.
ptes: PteArray<{ Self::PTE_ARRAY_SIZE }>,
/// CPU queue: the driver writes commands here, and the GSP reads them. It also contains the
@ -192,13 +198,13 @@ pub(super) struct GspMem {
/// index into the GSP queue.
///
/// This member is read-only for the GSP.
pub(super) cpuq: Msgq,
cpuq: Msgq,
/// GSP queue: the GSP writes messages here, and the driver reads them. It also contains the
/// write and read pointers that the GSP updates. This means that the read pointer here is an
/// index into the CPU queue.
///
/// This member is read-only for the driver.
pub(super) gspq: Msgq,
gspq: Msgq,
}
impl GspMem {
@ -232,20 +238,12 @@ fn new(dev: &device::Device<device::Bound>) -> Result<Self> {
const MSGQ_SIZE: u32 = num::usize_into_u32::<{ size_of::<Msgq>() }>();
const RX_HDR_OFF: u32 = num::usize_into_u32::<{ mem::offset_of!(Msgq, rx) }>();
let gsp_mem = Coherent::<GspMem>::zeroed(dev, GFP_KERNEL)?;
let mut gsp_mem = CoherentBox::<GspMem>::zeroed(dev, GFP_KERNEL)?;
gsp_mem.cpuq.tx = MsgqTxHeader::new(MSGQ_SIZE, RX_HDR_OFF, MSGQ_NUM_PAGES);
gsp_mem.cpuq.rx = MsgqRxHeader::new();
let start = gsp_mem.dma_handle();
// Write values one by one to avoid an on-stack instance of `PteArray`.
for i in 0..GspMem::PTE_ARRAY_SIZE {
dma_write!(gsp_mem, .ptes.0[build: i], PteArray::<0>::entry(start, i)?);
}
dma_write!(
gsp_mem,
.cpuq.tx,
MsgqTxHeader::new(MSGQ_SIZE, RX_HDR_OFF, MSGQ_NUM_PAGES)
);
dma_write!(gsp_mem, .cpuq.rx, MsgqRxHeader::new());
let gsp_mem: Coherent<_> = gsp_mem.into();
PteArray::init(io_project!(gsp_mem, .ptes), gsp_mem.dma_address())?;
Ok(Self(gsp_mem))
}
@ -406,7 +404,7 @@ fn allocate_command(&mut self, size: usize, timeout: Delta) -> Result<GspCommand
//
// - The returned value is within `0..MSGQ_NUM_PAGES`.
fn gsp_write_ptr(&self) -> u32 {
super::fw::gsp_mem::gsp_write_ptr(&self.0)
MsgqTxHeader::write_ptr(io_project!(self.0, .gspq.tx)) % MSGQ_NUM_PAGES
}
// Returns the index of the memory page the GSP will read the next command from.
@ -415,7 +413,7 @@ fn gsp_write_ptr(&self) -> u32 {
//
// - The returned value is within `0..MSGQ_NUM_PAGES`.
fn gsp_read_ptr(&self) -> u32 {
super::fw::gsp_mem::gsp_read_ptr(&self.0)
MsgqRxHeader::read_ptr(io_project!(self.0, .gspq.rx)) % MSGQ_NUM_PAGES
}
// Returns the index of the memory page the CPU can read the next message from.
@ -424,12 +422,18 @@ fn gsp_read_ptr(&self) -> u32 {
//
// - The returned value is within `0..MSGQ_NUM_PAGES`.
fn cpu_read_ptr(&self) -> u32 {
super::fw::gsp_mem::cpu_read_ptr(&self.0)
MsgqRxHeader::read_ptr(io_project!(self.0, .cpuq.rx)) % MSGQ_NUM_PAGES
}
// Informs the GSP that it can send `elem_count` new pages into the message queue.
fn advance_cpu_read_ptr(&mut self, elem_count: u32) {
super::fw::gsp_mem::advance_cpu_read_ptr(&self.0, elem_count)
let rx = io_project!(self.0, .cpuq.rx);
let rptr = MsgqRxHeader::read_ptr(rx).wrapping_add(elem_count) % MSGQ_NUM_PAGES;
// Ensure read pointer is properly ordered.
fence(Ordering::SeqCst);
MsgqRxHeader::set_read_ptr(rx, rptr)
}
// Returns the index of the memory page the CPU can write the next command to.
@ -438,12 +442,17 @@ fn advance_cpu_read_ptr(&mut self, elem_count: u32) {
//
// - The returned value is within `0..MSGQ_NUM_PAGES`.
fn cpu_write_ptr(&self) -> u32 {
super::fw::gsp_mem::cpu_write_ptr(&self.0)
MsgqTxHeader::write_ptr(io_project!(self.0, .cpuq.tx)) % MSGQ_NUM_PAGES
}
// Informs the GSP that it can process `elem_count` new pages from the command queue.
fn advance_cpu_write_ptr(&mut self, elem_count: u32) {
super::fw::gsp_mem::advance_cpu_write_ptr(&self.0, elem_count)
let tx = io_project!(self.0, .cpuq.tx);
let wptr = MsgqTxHeader::write_ptr(tx).wrapping_add(elem_count) % MSGQ_NUM_PAGES;
MsgqTxHeader::set_write_ptr(tx, wptr);
// Ensure all command data is visible before triggering the GSP read.
fence(Ordering::SeqCst);
}
}
@ -478,8 +487,8 @@ pub(crate) struct Cmdq {
/// Inner mutex-protected state.
#[pin]
inner: Mutex<CmdqInner>,
/// DMA handle of the command queue's shared memory region.
pub(super) dma_handle: DmaAddress,
/// DMA address of the command queue's shared memory region.
pub(super) dma_addr: DmaAddress,
}
impl Cmdq {
@ -508,7 +517,7 @@ pub(crate) fn new(dev: &device::Device<device::Bound>) -> impl PinInit<Self, Err
let gsp_mem = DmaGspMem::new(dev)?;
Ok(try_pin_init!(Self {
dma_handle: gsp_mem.0.dma_handle(),
dma_addr: gsp_mem.0.dma_address(),
inner <- new_mutex!(CmdqInner {
dev: dev.into(),
gsp_mem,

View File

@ -5,6 +5,7 @@
array,
convert::Infallible,
ffi::FromBytesUntilNulError,
ops::Range,
str::Utf8Error, //
};
@ -33,6 +34,7 @@
},
},
sbuffer::SBufferIter,
vgpu::VgpuState, //
};
/// The `GspSetSystemInfo` command.
@ -66,37 +68,55 @@ struct RegistryEntry {
/// The `SetRegistry` command.
pub(crate) struct SetRegistry {
entries: [RegistryEntry; Self::NUM_ENTRIES],
entries: KVec<RegistryEntry>,
}
impl SetRegistry {
// For now we hard-code the registry entries. Future work will allow others to
// be added as module parameters.
const NUM_ENTRIES: usize = 3;
/// Creates a new `SetRegistry` command, using a set of hardcoded entries.
pub(crate) fn new() -> Self {
Self {
entries: [
// RMSecBusResetEnable - enables PCI secondary bus reset
pub(crate) fn new(vgpu_state: VgpuState) -> Result<Self> {
let mut entries = KVec::new();
// RMSecBusResetEnable - enables PCI secondary bus reset
entries.push(
RegistryEntry {
key: "RMSecBusResetEnable",
value: 1,
},
GFP_KERNEL,
)?;
// RMForcePcieConfigSave - forces GSP-RM to preserve PCI configuration registers on
// any PCI reset.
entries.push(
RegistryEntry {
key: "RMForcePcieConfigSave",
value: 1,
},
GFP_KERNEL,
)?;
// RMDevidCheckIgnore - allows GSP-RM to boot even if the PCI dev ID is not found
// in the internal product name database.
entries.push(
RegistryEntry {
key: "RMDevidCheckIgnore",
value: 1,
},
GFP_KERNEL,
)?;
if matches!(vgpu_state, VgpuState::Enabled { .. }) {
// RMSetSriovMode - required when vGPU is enabled.
entries.push(
RegistryEntry {
key: "RMSecBusResetEnable",
key: "RMSetSriovMode",
value: 1,
},
// RMForcePcieConfigSave - forces GSP-RM to preserve PCI configuration registers on
// any PCI reset.
RegistryEntry {
key: "RMForcePcieConfigSave",
value: 1,
},
// RMDevidCheckIgnore - allows GSP-RM to boot even if the PCI dev ID is not found
// in the internal product name database.
RegistryEntry {
key: "RMDevidCheckIgnore",
value: 1,
},
],
GFP_KERNEL,
)?;
}
Ok(Self { entries })
}
}
@ -107,15 +127,15 @@ impl CommandToGsp for SetRegistry {
type InitError = Infallible;
fn init(&self) -> impl Init<Self::Command, Self::InitError> {
Self::Command::init(Self::NUM_ENTRIES as u32, self.variable_payload_len() as u32)
Self::Command::init(self.entries.len() as u32, self.size() as u32)
}
fn variable_payload_len(&self) -> usize {
let mut key_size = 0;
for i in 0..Self::NUM_ENTRIES {
key_size += self.entries[i].key.len() + 1; // +1 for NULL terminator
for entry in self.entries.iter() {
key_size += entry.key.len() + 1; // +1 for NULL terminator
}
Self::NUM_ENTRIES * size_of::<fw::commands::PackedRegistryEntry>() + key_size
self.entries.len() * size_of::<fw::commands::PackedRegistryEntry>() + key_size
}
fn init_variable_payload(
@ -123,12 +143,12 @@ fn init_variable_payload(
dst: &mut SBufferIter<core::array::IntoIter<&mut [u8], 2>>,
) -> Result {
let string_data_start_offset = size_of::<Self::Command>()
+ Self::NUM_ENTRIES * size_of::<fw::commands::PackedRegistryEntry>();
+ self.entries.len() * size_of::<fw::commands::PackedRegistryEntry>();
// Array for string data.
let mut string_data = KVec::new();
for entry in self.entries.iter().take(Self::NUM_ENTRIES) {
for entry in self.entries.iter() {
dst.write_all(
fw::commands::PackedRegistryEntry::new(
(string_data_start_offset + string_data.len()) as u32,
@ -191,22 +211,30 @@ fn init(&self) -> impl Init<Self::Command, Self::InitError> {
}
}
/// The reply from the GSP to the [`GetGspInfo`] command.
/// The reply from the GSP to the [`GetGspStaticInfo`] command.
pub(crate) struct GetGspStaticInfoReply {
gpu_name: [u8; 64],
/// Usable FB (VRAM) regions for driver memory allocation.
pub(crate) usable_fb_regions: KVec<Range<u64>>,
}
impl MessageFromGsp for GetGspStaticInfoReply {
const FUNCTION: MsgFunction = MsgFunction::GetGspStaticInfo;
type Message = fw::commands::GspStaticConfigInfo;
type InitError = Infallible;
type InitError = Error;
fn read(
msg: &Self::Message,
_sbuffer: &mut SBufferIter<array::IntoIter<&[u8], 2>>,
) -> Result<Self, Self::InitError> {
let mut usable_fb_regions = KVec::new();
for region in msg.usable_fb_regions() {
usable_fb_regions.push(region, GFP_KERNEL)?;
}
Ok(GetGspStaticInfoReply {
gpu_name: msg.gpu_name_str(),
usable_fb_regions,
})
}
}

View File

@ -10,7 +10,15 @@
use core::ops::Range;
use kernel::{
dma::Coherent,
bitfield,
dma::{
Coherent,
CoherentView, //
},
io::{
io_read,
io_write, //
},
prelude::*,
ptr::{
Alignable,
@ -28,7 +36,10 @@
};
use crate::{
fb::FbLayout,
fb::{
FbRanges,
FbSizes, //
},
firmware::gsp::GspFirmware,
gpu::{
Architecture,
@ -44,59 +55,6 @@
},
};
// TODO: Replace with `IoView` projections once available.
pub(super) mod gsp_mem {
use core::sync::atomic::{
fence,
Ordering, //
};
use kernel::{
dma::Coherent,
dma_read,
dma_write, //
};
use crate::gsp::cmdq::{
GspMem,
MSGQ_NUM_PAGES, //
};
pub(in crate::gsp) fn gsp_write_ptr(qs: &Coherent<GspMem>) -> u32 {
dma_read!(qs, .gspq.tx.0.writePtr) % MSGQ_NUM_PAGES
}
pub(in crate::gsp) fn gsp_read_ptr(qs: &Coherent<GspMem>) -> u32 {
dma_read!(qs, .gspq.rx.0.readPtr) % MSGQ_NUM_PAGES
}
pub(in crate::gsp) fn cpu_read_ptr(qs: &Coherent<GspMem>) -> u32 {
dma_read!(qs, .cpuq.rx.0.readPtr) % MSGQ_NUM_PAGES
}
pub(in crate::gsp) fn advance_cpu_read_ptr(qs: &Coherent<GspMem>, count: u32) {
let rptr = cpu_read_ptr(qs).wrapping_add(count) % MSGQ_NUM_PAGES;
// Ensure read pointer is properly ordered.
fence(Ordering::SeqCst);
dma_write!(qs, .cpuq.rx.0.readPtr, rptr);
}
pub(in crate::gsp) fn cpu_write_ptr(qs: &Coherent<GspMem>) -> u32 {
dma_read!(qs, .cpuq.tx.0.writePtr) % MSGQ_NUM_PAGES
}
pub(in crate::gsp) fn advance_cpu_write_ptr(qs: &Coherent<GspMem>, count: u32) {
let wptr = cpu_write_ptr(qs).wrapping_add(count) % MSGQ_NUM_PAGES;
dma_write!(qs, .cpuq.tx.0.writePtr, wptr);
// Ensure all command data is visible before triggering the GSP read.
fence(Ordering::SeqCst);
}
}
/// Maximum size of a single GSP message queue element in bytes.
pub(crate) const GSP_MSG_QUEUE_ELEMENT_SIZE_MAX: usize =
num::u32_as_usize(bindings::GSP_MSG_QUEUE_ELEMENT_SIZE_MAX);
@ -177,6 +135,11 @@ pub(crate) fn from_chipset(chipset: Chipset) -> &'static LibosParams {
}
}
/// Returns the WPR heap size to reserve when vGPU is enabled.
pub(crate) fn vgpu_wpr_heap_size() -> u64 {
u64::from(bindings::GSP_FW_HEAP_SIZE_VGPU_DEFAULT)
}
/// Returns the amount of memory (in bytes) to allocate for the WPR heap for a framebuffer size
/// of `fb_size` (in bytes) for `chipset`.
pub(crate) fn wpr_heap_size(&self, chipset: Chipset, fb_size: u64) -> Result<u64> {
@ -214,48 +177,89 @@ unsafe impl FromBytes for GspFwWprMeta {}
impl GspFwWprMeta {
/// Returns an initializer for a `GspFwWprMeta` suitable for booting `gsp_firmware` using the
/// `fb_layout` layout.
pub(crate) fn new<'a>(
/// framebuffer ranges `ranges`.
pub(crate) fn from_ranges<'a>(
gsp_firmware: &'a GspFirmware,
fb_layout: &'a FbLayout,
ranges: &'a FbRanges,
) -> impl Init<Self> + 'a {
#[allow(non_snake_case)]
let init_inner = init!(bindings::GspFwWprMeta {
// CAST: we want to store the bits of `GSP_FW_WPR_META_MAGIC` unmodified.
magic: bindings::GSP_FW_WPR_META_MAGIC as u64,
revision: u64::from(bindings::GSP_FW_WPR_META_REVISION),
sysmemAddrOfRadix3Elf: gsp_firmware.radix3_dma_handle(),
sysmemAddrOfRadix3Elf: gsp_firmware.radix3_dma_address(),
sizeOfRadix3Elf: u64::from_safe_cast(gsp_firmware.size),
sysmemAddrOfBootloader: gsp_firmware.bootloader.ucode.dma_handle(),
sysmemAddrOfBootloader: gsp_firmware.bootloader.ucode.dma_address(),
sizeOfBootloader: u64::from_safe_cast(gsp_firmware.bootloader.ucode.size()),
bootloaderCodeOffset: u64::from(gsp_firmware.bootloader.code_offset),
bootloaderDataOffset: u64::from(gsp_firmware.bootloader.data_offset),
bootloaderManifestOffset: u64::from(gsp_firmware.bootloader.manifest_offset),
__bindgen_anon_1: GspFwWprMetaBootResumeInfo {
__bindgen_anon_1: GspFwWprMetaBootInfo {
sysmemAddrOfSignature: gsp_firmware.signatures.dma_handle(),
sysmemAddrOfSignature: gsp_firmware.signatures.dma_address(),
sizeOfSignature: u64::from_safe_cast(gsp_firmware.signatures.size()),
},
},
gspFwRsvdStart: fb_layout.heap.start,
nonWprHeapOffset: fb_layout.heap.start,
nonWprHeapSize: fb_layout.heap.end - fb_layout.heap.start,
gspFwWprStart: fb_layout.wpr2.start,
gspFwHeapOffset: fb_layout.wpr2_heap.start,
gspFwHeapSize: fb_layout.wpr2_heap.end - fb_layout.wpr2_heap.start,
gspFwOffset: fb_layout.elf.start,
bootBinOffset: fb_layout.boot.start,
frtsOffset: fb_layout.frts.start,
frtsSize: fb_layout.frts.end - fb_layout.frts.start,
gspFwWprEnd: fb_layout
gspFwRsvdStart: ranges.non_wpr_heap.start,
nonWprHeapOffset: ranges.non_wpr_heap.start,
nonWprHeapSize: ranges.non_wpr_heap.len(),
gspFwWprStart: ranges.wpr2.start,
gspFwHeapOffset: ranges.wpr2_heap.start,
gspFwHeapSize: ranges.wpr2_heap.len(),
gspFwOffset: ranges.elf.start,
bootBinOffset: ranges.boot.start,
frtsOffset: ranges.frts.start,
frtsSize: ranges.frts.len(),
gspFwWprEnd: ranges
.vga_workspace
.start
.align_down(Alignment::new::<SZ_128K>()),
gspFwHeapVfPartitionCount: fb_layout.vf_partition_count,
fbSize: fb_layout.fb.end - fb_layout.fb.start,
vgaWorkspaceOffset: fb_layout.vga_workspace.start,
vgaWorkspaceSize: fb_layout.vga_workspace.end - fb_layout.vga_workspace.start,
pmuReservedSize: fb_layout.pmu_reserved_size,
gspFwHeapVfPartitionCount: ranges.vf_partition_count,
fbSize: ranges.fb.len(),
vgaWorkspaceOffset: ranges.vga_workspace.start,
vgaWorkspaceSize: ranges.vga_workspace.len(),
pmuReservedSize: ranges.pmu_reserved_size,
..Zeroable::init_zeroed()
});
init!(GspFwWprMeta {
inner <- init_inner,
})
}
/// Returns an initializer for a `GspFwWprMeta` suitable for booting `gsp_firmware` using the
/// framebuffer region sizes `sizes`.
///
/// The region offsets are left at zero: the ACR ucode computes them when it sets up WPR2.
pub(crate) fn from_sizes<'a>(
gsp_firmware: &'a GspFirmware,
sizes: &'a FbSizes,
) -> impl Init<Self> + 'a {
/// VGA workspace size to reserve at the end of the framebuffer, in bytes.
const VGA_WORKSPACE_SIZE: u64 = u64::SZ_128K;
let init_inner = init!(bindings::GspFwWprMeta {
// CAST: we want to store the bits of `GSP_FW_WPR_META_MAGIC` unmodified.
magic: bindings::GSP_FW_WPR_META_MAGIC as u64,
revision: u64::from(bindings::GSP_FW_WPR_META_REVISION),
sysmemAddrOfRadix3Elf: gsp_firmware.radix3_dma_address(),
sizeOfRadix3Elf: u64::from_safe_cast(gsp_firmware.size),
sysmemAddrOfBootloader: gsp_firmware.bootloader.ucode.dma_address(),
sizeOfBootloader: u64::from_safe_cast(gsp_firmware.bootloader.ucode.size()),
bootloaderCodeOffset: u64::from(gsp_firmware.bootloader.code_offset),
bootloaderDataOffset: u64::from(gsp_firmware.bootloader.data_offset),
bootloaderManifestOffset: u64::from(gsp_firmware.bootloader.manifest_offset),
__bindgen_anon_1: GspFwWprMetaBootResumeInfo {
__bindgen_anon_1: GspFwWprMetaBootInfo {
sysmemAddrOfSignature: gsp_firmware.signatures.dma_address(),
sizeOfSignature: u64::from_safe_cast(gsp_firmware.signatures.size()),
},
},
nonWprHeapSize: sizes.non_wpr_heap_size,
gspFwHeapSize: sizes.wpr2_heap_size,
frtsSize: sizes.frts_size,
gspFwHeapVfPartitionCount: sizes.vf_partition_count,
vgaWorkspaceSize: VGA_WORKSPACE_SIZE,
pmuReservedSize: sizes.pmu_reserved_size,
..Zeroable::init_zeroed()
});
@ -674,10 +678,9 @@ fn id8(name: &str) -> u64 {
u64::from_ne_bytes(bytes)
}
#[allow(non_snake_case)]
let init_inner = init!(bindings::LibosMemoryRegionInitArgument {
id8: id8(name),
pa: obj.dma_handle(),
pa: obj.dma_address(),
size: num::usize_as_u64(obj.size()),
kind: num::u32_into_u8::<
{ bindings::LibosMemoryRegionKind_LIBOS_MEMORY_REGION_CONTIGUOUS },
@ -720,6 +723,16 @@ pub(crate) fn new(msgq_size: u32, rx_hdr_offset: u32, msg_count: u32) -> Self {
entryOff: num::usize_into_u32::<GSP_PAGE_SIZE>(),
})
}
/// Returns the value of the write pointer for this queue.
pub(crate) fn write_ptr(this: CoherentView<'_, Self>) -> u32 {
io_read!(this, .0.writePtr)
}
/// Sets the value of the write pointer for this queue.
pub(crate) fn set_write_ptr(this: CoherentView<'_, Self>, val: u32) {
io_write!(this, .0.writePtr, val)
}
}
// SAFETY: Padding is explicit and does not contain uninitialized data.
@ -735,6 +748,16 @@ impl MsgqRxHeader {
pub(crate) fn new() -> Self {
Self(Default::default())
}
/// Returns the value of the read pointer for this queue.
pub(crate) fn read_ptr(this: CoherentView<'_, Self>) -> u32 {
io_read!(this, .0.readPtr)
}
/// Sets the value of the read pointer for this queue.
pub(crate) fn set_read_ptr(this: CoherentView<'_, Self>, val: u32) {
io_write!(this, .0.readPtr, val)
}
}
// SAFETY: Padding is explicit and does not contain uninitialized data.
@ -742,8 +765,8 @@ unsafe impl AsBytes for MsgqRxHeader {}
bitfield! {
struct MsgHeaderVersion(u32) {
31:24 major as u8;
23:16 minor as u8;
31:24 major;
23:16 minor;
}
}
@ -752,9 +775,9 @@ impl MsgHeaderVersion {
const MINOR_TOT: u8 = 0;
fn new() -> Self {
Self::default()
.set_major(Self::MAJOR_TOT)
.set_minor(Self::MINOR_TOT)
Self::zeroed()
.with_major(Self::MAJOR_TOT)
.with_minor(Self::MINOR_TOT)
}
}
@ -793,7 +816,6 @@ impl GspMsgElement {
/// * `sequence` - Sequence number of the message.
/// * `cmd_size` - Size of the command (not including the message element), in bytes.
/// * `function` - Function of the message.
#[allow(non_snake_case)]
pub(crate) fn init(
sequence: u32,
cmd_size: usize,
@ -876,7 +898,6 @@ pub(crate) struct GspArgumentsCached {
impl GspArgumentsCached {
/// Creates the arguments for starting the GSP up using `cmdq` as its command queue.
pub(crate) fn new(cmdq: &Cmdq) -> impl Init<Self> + '_ {
#[allow(non_snake_case)]
let init_inner = init!(bindings::GSP_ARGUMENTS_CACHED {
messageQueueInitArguments <- MessageQueueInitArguments::new(cmdq),
bDmemStack: 1,
@ -923,10 +944,9 @@ unsafe impl FromBytes for GspArgumentsPadded {}
impl MessageQueueInitArguments {
/// Creates a new init arguments structure for `cmdq`.
#[allow(non_snake_case)]
fn new(cmdq: &Cmdq) -> impl Init<Self> + '_ {
init!(MessageQueueInitArguments {
sharedMemPhysAddr: cmdq.dma_handle,
sharedMemPhysAddr: cmdq.dma_addr,
pageTableEntryCount: num::usize_into_u32::<{ Cmdq::NUM_PTES }>(),
cmdQueueOffset: num::usize_as_u64(Cmdq::CMDQ_OFFSET),
statQueueOffset: num::usize_as_u64(Cmdq::STATQ_OFFSET),
@ -947,7 +967,6 @@ pub(crate) enum GspDmaTarget {
impl GspAcrBootGspRmParams {
fn new(target: GspDmaTarget, wpr_meta_addr: u64) -> impl Init<Self> {
#[allow(non_snake_case)]
let params = init!(Self {
target: target as u32,
gspRmDescSize: num::usize_into_u32::<{ size_of::<GspFwWprMeta>() }>(),
@ -966,7 +985,6 @@ fn new(target: GspDmaTarget, wpr_meta_addr: u64) -> impl Init<Self> {
impl GspRmParams {
fn new(target: GspDmaTarget, libos_addr: u64) -> impl Init<Self> {
#[allow(non_snake_case)]
let params = init!(Self {
target: target as u32,
bootArgsOffset: libos_addr,
@ -986,7 +1004,6 @@ unsafe impl FromBytes for GspFmcBootParams {}
impl GspFmcBootParams {
pub(crate) fn new(wpr_meta_addr: u64, libos_addr: u64) -> impl Init<Self> {
#[allow(non_snake_case)]
let init = init!(Self {
// Blackwell FSP obtains WPR info from other sources, so
// wprCarveoutOffset and wprCarveoutSize are left zero.

View File

@ -1,6 +1,8 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use core::ops::Range;
use kernel::{
device,
pci,
@ -13,7 +15,8 @@
use crate::{
gpu::Chipset,
gsp::GSP_PAGE_SIZE, //
gsp::GSP_PAGE_SIZE,
num::IntoSafeCast, //
};
use super::bindings;
@ -27,7 +30,6 @@ pub(crate) struct GspSetSystemInfo {
impl GspSetSystemInfo {
/// Returns an in-place initializer for the `GspSetSystemInfo` command.
#[allow(non_snake_case)]
pub(crate) fn init<'a>(
dev: &'a pci::Device<device::Bound>,
chipset: Chipset,
@ -99,7 +101,6 @@ pub(crate) struct PackedRegistryTable {
}
impl PackedRegistryTable {
#[allow(non_snake_case)]
pub(crate) fn init(num_entries: u32, size: u32) -> impl Init<Self> {
type InnerPackedRegistryTable = bindings::PACKED_REGISTRY_TABLE;
let init_inner = init!(InnerPackedRegistryTable {
@ -129,6 +130,41 @@ impl GspStaticConfigInfo {
pub(crate) fn gpu_name_str(&self) -> [u8; 64] {
self.0.gpuNameString
}
/// Returns an iterator over valid FB regions from GSP firmware data.
fn fb_regions(
&self,
) -> impl Iterator<Item = &bindings::NV2080_CTRL_CMD_FB_GET_FB_REGION_FB_REGION_INFO> {
let fb_info = &self.0.fbRegionInfoParams;
fb_info
.fbRegion
.iter()
.take(fb_info.numFBRegions.into_safe_cast())
.filter(|reg| reg.limit >= reg.base)
}
/// Iterates over usable FB regions from GSP firmware data.
///
/// Each yielded region is a [`Range<u64>`] suitable for driver memory allocation.
/// Usable regions are those that satisfy all the following properties:
/// - Are not reserved for firmware internal use.
/// - Are not protected (hardware-enforced access restrictions).
/// - Support compression (can use GPU memory compression for bandwidth).
/// - Support ISO (isochronous memory for display requiring guaranteed bandwidth).
pub(crate) fn usable_fb_regions(&self) -> impl Iterator<Item = Range<u64>> + '_ {
self.fb_regions().filter_map(|reg| {
// Filter: not reserved, not protected, supports compression and ISO.
if reg.reserved == 0
&& reg.bProtected == 0
&& reg.supportCompressed != 0
&& reg.supportISO != 0
{
reg.limit.checked_add(1).map(|end| reg.base..end)
} else {
None
}
})
}
}
// SAFETY: Padding is explicit and will not contain uninitialized data.

View File

@ -40,6 +40,7 @@ fn fmt(&self, fmt: &mut ::core::fmt::Formatter<'_>) -> ::core::fmt::Result {
pub const GSP_FW_HEAP_PARAM_BASE_RM_SIZE_GH100: u32 = 14680064;
pub const GSP_FW_HEAP_PARAM_SIZE_PER_GB_FB: u32 = 98304;
pub const GSP_FW_HEAP_PARAM_CLIENT_ALLOC_SIZE: u32 = 100663296;
pub const GSP_FW_HEAP_SIZE_VGPU_DEFAULT: u32 = 609222656;
pub const GSP_FW_HEAP_SIZE_OVERRIDE_LIBOS2_MIN_MB: u32 = 64;
pub const GSP_FW_HEAP_SIZE_OVERRIDE_LIBOS2_MAX_MB: u32 = 256;
pub const GSP_FW_HEAP_SIZE_OVERRIDE_LIBOS3_BAREMETAL_MIN_MB: u32 = 88;

View File

@ -1,33 +1,21 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
mod ga102;
mod gh100;
mod tu102;
use kernel::prelude::*;
use kernel::{
device,
dma::Coherent, //
};
use crate::{
driver::Bar0,
falcon::{
gsp::Gsp as GspEngine,
sec2::Sec2,
Falcon, //
},
fb::FbLayout,
firmware::gsp::GspFirmware,
gpu::{
Architecture,
Chipset, //
},
gsp::{
boot::BootUnloadGuard,
Gsp,
GspFwWprMeta, //
GspBootContext, //
},
};
@ -38,33 +26,21 @@
/// required for unloading is prepared at load time, and stored here until it needs to be run.
pub(super) trait UnloadBundle: Send {
/// Performs the steps required to properly reset the GSP after it has been stopped.
fn run(
&self,
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
gsp_falcon: &Falcon<GspEngine>,
sec2_falcon: &Falcon<Sec2>,
) -> Result;
fn run(&self, ctx: &mut GspBootContext<'_, '_>) -> Result;
}
/// Trait implemented by GSP HALs.
pub(super) trait GspHal: Send {
/// Performs the GSP boot process, loading and running the required firmwares as needed.
///
/// Upon success, returns a guard that runs the GSP unload sequence if GSP boot does not
/// complete.
#[allow(clippy::too_many_arguments)]
fn boot<'a>(
/// Upon success, returns the [`crate::gsp::UnloadBundle`] to use with [`Gsp::unload`], if one
/// could be created.
fn boot(
&self,
gsp: &'a Gsp,
dev: &'a device::Device<device::Bound>,
bar: Bar0<'a>,
chipset: Chipset,
fb_layout: &FbLayout,
wpr_meta: &Coherent<GspFwWprMeta>,
gsp_falcon: &'a Falcon<GspEngine>,
sec2_falcon: &'a Falcon<Sec2>,
) -> Result<BootUnloadGuard<'a>>;
gsp: &Gsp,
ctx: &mut GspBootContext<'_, '_>,
gsp_fw: &GspFirmware,
) -> Result<Option<crate::gsp::UnloadBundle>>;
/// Performs HAL-specific post-GSP boot tasks.
///
@ -73,20 +49,38 @@ fn boot<'a>(
fn post_boot(
&self,
_gsp: &Gsp,
_dev: &device::Device<device::Bound>,
_bar: Bar0<'_>,
_ctx: &mut GspBootContext<'_, '_>,
_gsp_fw: &GspFirmware,
_gsp_falcon: &Falcon<GspEngine>,
_sec2_falcon: &Falcon<Sec2>,
) -> Result {
Ok(())
}
}
/// Returns the names of the firmware files required to boot the GSP of `chipset`, in addition to
/// the "bootloader" and "gsp" images required by all chipsets.
pub(crate) const fn boot_firmware_files(chipset: Chipset) -> &'static [&'static str] {
match chipset.arch() {
// Turing chipsets boot the GSP via the SEC2 Booter, and require the FWSEC bootloader.
Architecture::Turing => &["booter_load.tlv", "booter_unload.tlv", "gen_bootloader.tlv"],
// GA100 also requires the FWSEC bootloader.
Architecture::Ampere if matches!(chipset, Chipset::GA100) => {
&["booter_load.tlv", "booter_unload.tlv", "gen_bootloader.tlv"]
}
// Other Ampere chipsets, as well as Ada chipsets, run FWSEC directly.
Architecture::Ampere | Architecture::Ada => &["booter_load.tlv", "booter_unload.tlv"],
// Hopper and later chipsets boot the GSP via the FMC image loaded by FSP.
Architecture::Hopper | Architecture::BlackwellGB10x | Architecture::BlackwellGB20x => {
&["fmc.tlv"]
}
}
}
/// Returns the GSP HAL to be used for `chipset`.
pub(super) fn gsp_hal(chipset: Chipset) -> &'static dyn GspHal {
match chipset.arch() {
Architecture::Turing | Architecture::Ampere | Architecture::Ada => tu102::TU102_HAL,
Architecture::Turing => tu102::TU102_HAL,
Architecture::Ampere if matches!(chipset, Chipset::GA100) => tu102::TU102_HAL,
Architecture::Ampere | Architecture::Ada => ga102::GA102_HAL,
Architecture::Hopper | Architecture::BlackwellGB10x | Architecture::BlackwellGB20x => {
gh100::GH100_HAL
}

View File

@ -0,0 +1,14 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use crate::gsp::hal::{
tu102::Tu102,
GspHal, //
};
/// The GA102 HAL is like the TU102 one, except it doesn't use the bootloader.
const GA102: Tu102 = Tu102 {
needs_fwsec_bootloader: false,
};
pub(super) const GA102_HAL: &dyn GspHal = &GA102;

View File

@ -7,33 +7,26 @@
device,
dma::Coherent,
io::poll::read_poll_timeout,
time::Delta, //
time::Delta,
types::ScopeGuard, //
};
use crate::{
driver::Bar0,
falcon::{
gsp::Gsp as GspEngine,
sec2::Sec2,
Falcon, //
},
fb::FbLayout,
firmware::{
fsp::FspFirmware,
FIRMWARE_VERSION, //
},
fsp::{
FmcBootArgs,
Fsp, //
},
gpu::Chipset,
fb::FbSizes,
firmware::gsp::GspFirmware,
fsp::FmcBootArgs,
gsp::{
boot::BootUnloadGuard,
hal::{
GspHal,
UnloadBundle, //
},
Gsp,
GspBootContext,
GspFmcBootParams,
GspFwWprMeta, //
},
};
@ -46,10 +39,10 @@ struct GspMbox {
impl GspMbox {
/// Reads both mailboxes from the GSP falcon.
fn read(gsp_falcon: &Falcon<GspEngine>, bar: Bar0<'_>) -> Self {
fn read(gsp_falcon: &Falcon<'_, GspEngine>) -> Self {
Self {
mbox0: gsp_falcon.read_mailbox0(bar),
mbox1: gsp_falcon.read_mailbox1(bar),
mbox0: gsp_falcon.read_mailbox0(),
mbox1: gsp_falcon.read_mailbox1(),
}
}
@ -64,27 +57,25 @@ fn combined_addr(&self) -> u64 {
/// either condition should stop the poll loop.
fn lockdown_released_or_error(
&self,
gsp_falcon: &Falcon<GspEngine>,
bar: Bar0<'_>,
fmc_boot_params_addr: u64,
gsp_falcon: &Falcon<'_, GspEngine>,
fmc_boot_params: &Coherent<GspFmcBootParams>,
) -> bool {
// GSP-FMC normally clears the boot parameters address from the mailboxes early during
// boot. If the address is still there, keep polling rather than treating it as an error.
// Any other non-zero mailbox0 value is a GSP-FMC error code.
if self.mbox0 != 0 {
return self.combined_addr() != fmc_boot_params_addr;
return self.combined_addr() != fmc_boot_params.dma_address();
}
!gsp_falcon.riscv_branch_privilege_lockdown(bar)
!gsp_falcon.riscv_branch_privilege_lockdown()
}
}
/// Waits for GSP lockdown to be released after FSP Chain of Trust.
fn wait_for_gsp_lockdown_release(
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
gsp_falcon: &Falcon<GspEngine>,
fmc_boot_params_addr: u64,
gsp_falcon: &Falcon<'_, GspEngine>,
fmc_boot_params: &Coherent<GspFmcBootParams>,
) -> Result {
dev_dbg!(dev, "Waiting for GSP lockdown release\n");
@ -92,14 +83,14 @@ fn wait_for_gsp_lockdown_release(
|| {
// While the PRIV target mask is still locked to FSP, GSP register and mailbox reads
// are not meaningful. Wait until HWCFG2 says the CPU can read them.
Ok(match gsp_falcon.priv_target_mask_released(bar) {
Ok(match gsp_falcon.priv_target_mask_released() {
false => None,
true => Some(GspMbox::read(gsp_falcon, bar)),
true => Some(GspMbox::read(gsp_falcon)),
})
},
|mbox| match mbox {
None => false,
Some(mbox) => mbox.lockdown_released_or_error(gsp_falcon, bar, fmc_boot_params_addr),
Some(mbox) => mbox.lockdown_released_or_error(gsp_falcon, fmc_boot_params),
},
Delta::from_millis(10),
Delta::from_secs(30),
@ -123,22 +114,23 @@ fn wait_for_gsp_lockdown_release(
struct FspUnloadBundle;
impl UnloadBundle for FspUnloadBundle {
fn run(
&self,
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
gsp_falcon: &Falcon<GspEngine>,
_sec2_falcon: &Falcon<Sec2>,
) -> Result {
fn run(&self, ctx: &mut GspBootContext<'_, '_>) -> Result {
// GSP falcon does most of the work of resetting, so just wait for it to finish.
read_poll_timeout(
|| Ok(gsp_falcon.is_riscv_active(bar)),
|&active| !active,
|| {
// GSP register reads are not meaningful until the PRIV target mask is released.
if !ctx.gsp_falcon.priv_target_mask_released() {
return Ok(false);
}
ctx.gsp_falcon.is_riscv_halted()
},
|&halted| halted,
Delta::from_millis(10),
Delta::from_secs(5),
)
.map(|_| ())
.inspect_err(|_| dev_err!(dev, "GSP falcon failed to halt\n"))
.inspect_err(|_| dev_err!(ctx.dev(), "GSP falcon failed to halt\n"))
}
}
@ -149,42 +141,44 @@ impl GspHal for Gh100 {
///
/// This path uses FSP to establish a chain of trust and boot GSP-FMC. FSP handles
/// the GSP boot internally - no manual GSP reset/boot is needed.
fn boot<'a>(
fn boot(
&self,
gsp: &'a Gsp,
dev: &'a device::Device<device::Bound>,
bar: Bar0<'a>,
chipset: Chipset,
fb_layout: &FbLayout,
wpr_meta: &Coherent<GspFwWprMeta>,
gsp_falcon: &'a Falcon<GspEngine>,
sec2_falcon: &'a Falcon<Sec2>,
) -> Result<BootUnloadGuard<'a>> {
let fsp_fw = FspFirmware::new(dev, chipset, FIRMWARE_VERSION)?;
gsp: &Gsp,
ctx: &mut GspBootContext<'_, '_>,
gsp_fw: &GspFirmware,
) -> Result<Option<crate::gsp::UnloadBundle>> {
let dev = ctx.dev();
let chipset = ctx.chipset;
let gsp_falcon = ctx.gsp_falcon;
let fb_sizes = FbSizes::new(chipset, ctx.bar, ctx.vgpu.state())?;
dev_dbg!(dev, "{:#x?}\n", fb_sizes);
let wpr_meta =
Coherent::init(dev, GFP_KERNEL, GspFwWprMeta::from_sizes(gsp_fw, &fb_sizes))?;
let args = FmcBootArgs::new(dev, chipset, wpr_meta, &gsp.libos, false)?;
let unload_bundle = crate::gsp::UnloadBundle(
KBox::new(FspUnloadBundle, GFP_KERNEL)? as KBox<dyn UnloadBundle>
);
// Wrap the unload bundle into a drop guard so it is automatically run upon failure.
let unload_guard =
BootUnloadGuard::new(gsp, dev, bar, gsp_falcon, sec2_falcon, Some(unload_bundle));
// Wait for the GSP RISC-V core to halt in case of error. We create this guard after `args`
// to make sure that the boot args and the WPR metadata they own are kept alive until halt,
// in case they are still being accessed.
let mut unload_guard =
ScopeGuard::new_with_data((unload_bundle, ctx), |(unload_bundle, ctx)| {
let _ = unload_bundle.0.run(ctx);
});
let mut fsp = Fsp::wait_secure_boot(dev, bar, chipset, fsp_fw)?;
let fsp = unload_guard.1.fsp.as_mut().ok_or(ENODEV)?;
let args = FmcBootArgs::new(
dev,
chipset,
wpr_meta.dma_handle(),
gsp.libos.dma_handle(),
false,
)?;
fsp.boot_fmc(dev, &fb_sizes, &args)?;
fsp.boot_fmc(dev, bar, fb_layout, &args)?;
// Wait for GSP-FMC to release the GSP lockdown, indicating that `args` is not accessed
// anymore.
wait_for_gsp_lockdown_release(dev, gsp_falcon, args.boot_params())?;
wait_for_gsp_lockdown_release(dev, bar, gsp_falcon, args.boot_params_dma_handle())?;
Ok(unload_guard)
Ok(Some(unload_guard.dismiss().0))
}
}

View File

@ -6,7 +6,8 @@
use kernel::{
device,
dma::Coherent,
io::Io, //
io::Io,
types::ScopeGuard, //
};
use crate::{
@ -16,7 +17,10 @@
sec2::Sec2,
Falcon, //
},
fb::FbLayout,
fb::{
wpr2_range,
FbRanges, //
},
firmware::{
booter::{
BooterFirmware,
@ -27,24 +31,20 @@
FwsecCommand,
FwsecFirmware, //
},
gsp::GspFirmware,
FIRMWARE_VERSION, //
gsp::GspFirmware, //
},
gpu::Chipset,
gsp::{
boot::BootUnloadGuard,
hal::{
GspHal,
UnloadBundle, //
},
sequencer::{
GspSequencer,
GspSequencerParams, //
},
regs,
sequencer::GspSequencer,
Gsp,
GspBootContext,
GspFwWprMeta, //
},
regs,
vbios::Vbios, //
};
@ -58,32 +58,15 @@ enum FwsecUnloadFirmware {
}
impl FwsecUnloadFirmware {
/// Loads the FWSEC SB firmware, as well as its bootloader if `chipset` requires it.
fn new(
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
chipset: Chipset,
bios: &Vbios,
gsp_falcon: &Falcon<GspEngine>,
) -> Result<Self> {
let fwsec_sb = FwsecFirmware::new(dev, gsp_falcon, bar, bios, FwsecCommand::Sb)?;
Ok(if chipset.needs_fwsec_bootloader() {
Self::WithBl(FwsecFirmwareWithBl::new(fwsec_sb, dev, chipset)?)
} else {
Self::WithoutBl(fwsec_sb)
})
}
/// Runs the FWSEC SB firmware.
fn run(
&self,
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
gsp_falcon: &Falcon<GspEngine>,
gsp_falcon: &Falcon<'_, GspEngine>,
) -> Result {
match self {
Self::WithoutBl(fw) => fw.run(dev, gsp_falcon, bar),
Self::WithoutBl(fw) => fw.run(dev, gsp_falcon),
Self::WithBl(fw) => fw.run(dev, gsp_falcon, bar),
}
}
@ -96,210 +79,225 @@ struct Sec2UnloadBundle {
booter_unloader: BooterFirmware,
}
impl Sec2UnloadBundle {
/// Load and prepare the resources required to properly reset the GSP after it has been stopped.
fn build(
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
chipset: Chipset,
bios: &Vbios,
gsp_falcon: &Falcon<GspEngine>,
sec2_falcon: &Falcon<Sec2>,
) -> Result<KBox<dyn UnloadBundle>> {
KBox::new(
Self {
fwsec_sb: FwsecUnloadFirmware::new(dev, bar, chipset, bios, gsp_falcon)?,
booter_unloader: BooterFirmware::new(
dev,
BooterKind::Unloader,
chipset,
FIRMWARE_VERSION,
sec2_falcon,
bar,
)?,
},
GFP_KERNEL,
)
.map(|b| b as KBox<dyn UnloadBundle>)
.map_err(Into::into)
}
}
impl UnloadBundle for Sec2UnloadBundle {
fn run(
&self,
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
gsp_falcon: &Falcon<GspEngine>,
sec2_falcon: &Falcon<Sec2>,
) -> Result {
fn run(&self, ctx: &mut GspBootContext<'_, '_>) -> Result {
let dev = ctx.dev();
let bar = ctx.bar;
// Run FWSEC-SB to reset the GSP falcon to its pre-libos state.
self.fwsec_sb.run(dev, bar, gsp_falcon)?;
// Log errors but keep going if it fails.
let fwsec_sb_res = self
.fwsec_sb
.run(dev, bar, ctx.gsp_falcon)
.inspect_err(|e| dev_err!(dev, "FWSEC-SB failed to run: {:?}\n", e));
// Remove WPR2 region if set.
let wpr2_hi = bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI);
if wpr2_hi.is_wpr2_set() {
sec2_falcon.reset(bar)?;
sec2_falcon.load(dev, bar, &self.booter_unloader)?;
let booter_unloader_res = (|| {
if wpr2_range(bar).is_none() {
return Ok(());
}
ctx.sec2_falcon.reset()?;
ctx.sec2_falcon.load(&self.booter_unloader)?;
// Sentinel value to confirm that Booter Unloader has run.
const MAILBOX_SENTINEL: u32 = 0xff;
let (mbox0, _) =
sec2_falcon.boot(bar, Some(MAILBOX_SENTINEL), Some(MAILBOX_SENTINEL))?;
let (mbox0, _) = ctx
.sec2_falcon
.boot(Some(MAILBOX_SENTINEL), Some(MAILBOX_SENTINEL))?;
if mbox0 != 0 {
dev_err!(dev, "Booter Unloader returned error 0x{:x}\n", mbox0);
return Err(EINVAL);
}
// Confirm that the WPR2 region has been removed.
let wpr2_hi = bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI);
if wpr2_hi.is_wpr2_set() {
if wpr2_range(bar).is_some() {
dev_err!(
dev,
"WPR2 region still set after Booter Unloader returned\n"
);
return Err(EBUSY);
}
}
Ok(())
Ok(())
})()
.inspect_err(|e| dev_err!(dev, "Booter Unloader failed to run: {:?}\n", e));
fwsec_sb_res.and(booter_unloader_res)
}
}
/// Helper function to load and run the FWSEC-FRTS firmware and confirm that it has properly
/// created the WPR2 region.
fn run_fwsec_frts(
dev: &device::Device<device::Bound>,
chipset: Chipset,
falcon: &Falcon<GspEngine>,
bar: Bar0<'_>,
bios: &Vbios,
fb_layout: &FbLayout,
) -> Result {
// Check that the WPR2 region does not already exist - if it does, we cannot run
// FWSEC-FRTS until the GPU is reset.
if bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI).higher_bound() != 0 {
dev_err!(
pub(super) struct Tu102 {
/// If `true`, then the FWSEC-FRTS bootloader will be used to load the actual firmware.
pub(super) needs_fwsec_bootloader: bool,
}
impl Tu102 {
/// Helper method to load and run the FWSEC-FRTS firmware and confirm that it has properly
/// created the WPR2 region.
fn run_fwsec_frts(
&self,
dev: &device::Device<device::Bound>,
chipset: Chipset,
falcon: &Falcon<'_, GspEngine>,
bar: Bar0<'_>,
bios: &Vbios,
fb_ranges: &FbRanges,
) -> Result {
// Check that the WPR2 region does not already exist - if it does, we cannot run
// FWSEC-FRTS until the GPU is reset.
if wpr2_range(bar).is_some() {
dev_err!(
dev,
"WPR2 region already exists - GPU needs to be reset to proceed\n"
);
return Err(EBUSY);
}
// FWSEC-FRTS will create the WPR2 region.
let fwsec_frts = FwsecFirmware::new(
dev,
"WPR2 region already exists - GPU needs to be reset to proceed\n"
);
return Err(EBUSY);
}
falcon,
bios,
FwsecCommand::Frts {
frts_addr: fb_ranges.frts.start,
frts_size: fb_ranges.frts.len(),
},
)?;
// FWSEC-FRTS will create the WPR2 region.
let fwsec_frts = FwsecFirmware::new(
dev,
falcon,
bar,
bios,
FwsecCommand::Frts {
frts_addr: fb_layout.frts.start,
frts_size: fb_layout.frts.len(),
},
)?;
if self.needs_fwsec_bootloader {
let fwsec_frts_bl = FwsecFirmwareWithBl::new(fwsec_frts, dev, chipset)?;
// Load and run the bootloader, which will load FWSEC-FRTS and run it.
fwsec_frts_bl.run(dev, falcon, bar)?;
} else {
// Load and run FWSEC-FRTS directly.
fwsec_frts.run(dev, falcon)?;
}
if chipset.needs_fwsec_bootloader() {
let fwsec_frts_bl = FwsecFirmwareWithBl::new(fwsec_frts, dev, chipset)?;
// Load and run the bootloader, which will load FWSEC-FRTS and run it.
fwsec_frts_bl.run(dev, falcon, bar)?;
} else {
// Load and run FWSEC-FRTS directly.
fwsec_frts.run(dev, falcon, bar)?;
}
// SCRATCH_E contains the error code for FWSEC-FRTS.
let frts_status = bar
.read(regs::NV_PBUS_SW_SCRATCH_0E_FRTS_ERR)
.frts_err_code();
if frts_status != 0 {
dev_err!(
dev,
"FWSEC-FRTS returned with error code {:#x}\n",
frts_status
);
// SCRATCH_E contains the error code for FWSEC-FRTS.
let frts_status = bar
.read(regs::NV_PBUS_SW_SCRATCH_0E_FRTS_ERR)
.frts_err_code();
if frts_status != 0 {
dev_err!(
dev,
"FWSEC-FRTS returned with error code {:#x}\n",
frts_status
);
return Err(EIO);
}
return Err(EIO);
}
// Check that the WPR2 region has been created as we requested.
let (wpr2_lo, wpr2_hi) = (
bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_LO).lower_bound(),
bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI).higher_bound(),
);
match (wpr2_lo, wpr2_hi) {
(_, 0) => {
// Check that the WPR2 region has been created as we requested.
let Some(wpr2_range) = wpr2_range(bar) else {
dev_err!(dev, "WPR2 region not created after running FWSEC-FRTS\n");
Err(EIO)
}
(wpr2_lo, _) if wpr2_lo != fb_layout.frts.start => {
return Err(EIO);
};
if wpr2_range.start != fb_ranges.frts.start {
dev_err!(
dev,
"WPR2 region created at unexpected address {:#x}; expected {:#x}\n",
wpr2_lo,
fb_layout.frts.start,
wpr2_range.start,
fb_ranges.frts.start,
);
Err(EIO)
return Err(EIO);
}
(wpr2_lo, wpr2_hi) => {
dev_dbg!(dev, "WPR2: {:#x}-{:#x}\n", wpr2_lo, wpr2_hi);
dev_dbg!(dev, "GPU instance built\n");
Ok(())
}
dev_dbg!(dev, "WPR2: {:#x}-{:#x}\n", wpr2_range.start, wpr2_range.end);
dev_dbg!(dev, "GPU instance built\n");
Ok(())
}
/// Load and prepare the resources required to properly reset the GSP after it has been stopped.
fn build_unload_bundle(
&self,
dev: &device::Device<device::Bound>,
chipset: Chipset,
bios: &Vbios,
gsp_falcon: &Falcon<'_, GspEngine>,
sec2_falcon: &Falcon<'_, Sec2>,
) -> Result<crate::gsp::UnloadBundle> {
// Load the FWSEC SB firmware, as well as its bootloader if required.
let fwsec_sb = FwsecFirmware::new(dev, gsp_falcon, bios, FwsecCommand::Sb)?;
let fwsec_sb = if self.needs_fwsec_bootloader {
FwsecUnloadFirmware::WithBl(FwsecFirmwareWithBl::new(fwsec_sb, dev, chipset)?)
} else {
FwsecUnloadFirmware::WithoutBl(fwsec_sb)
};
KBox::new(
Sec2UnloadBundle {
fwsec_sb,
booter_unloader: BooterFirmware::new(
dev,
BooterKind::Unloader,
chipset,
sec2_falcon,
)?,
},
GFP_KERNEL,
)
.map(|b| crate::gsp::UnloadBundle(b))
.map_err(Into::into)
}
}
struct Tu102;
impl GspHal for Tu102 {
fn boot<'a>(
fn boot(
&self,
gsp: &'a Gsp,
dev: &'a device::Device<device::Bound>,
bar: Bar0<'a>,
chipset: Chipset,
fb_layout: &FbLayout,
wpr_meta: &Coherent<GspFwWprMeta>,
gsp_falcon: &'a Falcon<GspEngine>,
sec2_falcon: &'a Falcon<Sec2>,
) -> Result<BootUnloadGuard<'a>> {
gsp: &Gsp,
ctx: &mut GspBootContext<'_, '_>,
gsp_fw: &GspFirmware,
) -> Result<Option<crate::gsp::UnloadBundle>> {
let dev = ctx.dev();
let bar = ctx.bar;
let chipset = ctx.chipset;
let gsp_falcon = ctx.gsp_falcon;
let sec2_falcon = ctx.sec2_falcon;
let fb_ranges = FbRanges::new(chipset, bar, gsp_fw, ctx.vgpu.state())?;
dev_dbg!(dev, "{:#x?}\n", fb_ranges);
// Declared before the unload guard so that if Booter fails while running, SEC2 is reset
// by the guard before this allocation is freed.
let wpr_meta = Coherent::init(
dev,
GFP_KERNEL,
GspFwWprMeta::from_ranges(gsp_fw, &fb_ranges),
)?;
let bios = Vbios::new(dev, bar)?;
// Try and prepare the unload bundle.
//
// If the unload bundle creation fails, the GPU will need to be reset before the driver can
// be probed again.
let unload_bundle =
Sec2UnloadBundle::build(dev, bar, chipset, &bios, gsp_falcon, sec2_falcon)
.inspect_err(|e| {
dev_warn!(dev, "Failed to prepare unload firmware: {:?}\n", e);
dev_warn!(dev, "The GSP won't be able to unload properly on unbind.\n");
dev_warn!(
dev,
"The GPU will need to be reset before the driver can bind again.\n"
);
})
.ok()
.map(crate::gsp::UnloadBundle);
let unload_bundle = self
.build_unload_bundle(dev, chipset, &bios, gsp_falcon, sec2_falcon)
.inspect_err(|e| dev_warn!(dev, "Failed to prepare unload firmware: {:?}\n", e))
.ok();
// Wrap the unload bundle into a drop guard so it is automatically run upon failure.
let unload_guard =
BootUnloadGuard::new(gsp, dev, bar, gsp_falcon, sec2_falcon, unload_bundle);
// Run the unload bundle to try and recover the GSP if an error occurs.
let unload_guard = ScopeGuard::new_with_data(unload_bundle, |unload_bundle| {
if let Some(unload_bundle) = unload_bundle {
let _ = unload_bundle.0.run(ctx);
}
});
// FWSEC-FRTS is not executed on chips where the FRTS region size is 0 (e.g. GA100).
if !fb_layout.frts.is_empty() {
run_fwsec_frts(dev, chipset, gsp_falcon, bar, &bios, fb_layout)?;
if !fb_ranges.frts.is_empty() {
self.run_fwsec_frts(dev, chipset, gsp_falcon, bar, &bios, &fb_ranges)?;
}
gsp_falcon.reset(bar)?;
let libos_handle = gsp.libos.dma_handle();
gsp_falcon.reset()?;
let libos_dma_address = gsp.libos.dma_address();
let (mbox0, mbox1) = gsp_falcon.boot(
bar,
Some(libos_handle as u32),
Some((libos_handle >> 32) as u32),
Some(libos_dma_address as u32),
Some((libos_dma_address >> 32) as u32),
)?;
dev_dbg!(dev, "GSP MBOX0: {:#x}, MBOX1: {:#x}\n", mbox0, mbox1);
@ -308,42 +306,30 @@ fn boot<'a>(
"Using SEC2 to load and run the booter_load firmware...\n"
);
BooterFirmware::new(
BooterFirmware::new(dev, BooterKind::Loader, chipset, sec2_falcon)?.run(
dev,
BooterKind::Loader,
chipset,
FIRMWARE_VERSION,
sec2_falcon,
bar,
)?
.run(dev, bar, sec2_falcon, wpr_meta)?;
&wpr_meta,
)?;
Ok(unload_guard)
Ok(unload_guard.dismiss())
}
fn post_boot(
&self,
gsp: &Gsp,
dev: &device::Device<device::Bound>,
bar: Bar0<'_>,
ctx: &mut GspBootContext<'_, '_>,
gsp_fw: &GspFirmware,
gsp_falcon: &Falcon<GspEngine>,
sec2_falcon: &Falcon<Sec2>,
) -> Result {
// Create and run the GSP sequencer.
let seq_params = GspSequencerParams {
bootloader_app_version: gsp_fw.bootloader.app_version,
libos_dma_handle: gsp.libos.dma_handle(),
gsp_falcon,
sec2_falcon,
dev,
bar,
};
GspSequencer::run(&gsp.cmdq, seq_params)?;
GspSequencer::run(&gsp.cmdq, ctx, &gsp.libos, gsp_fw.bootloader.app_version)?;
Ok(())
}
}
const TU102: Tu102 = Tu102;
/// The TU102 HAL requires the use of the FWSEC bootloader.
const TU102: Tu102 = Tu102 {
needs_fwsec_bootloader: true,
};
pub(super) const TU102_HAL: &dyn GspHal = &TU102;

View File

@ -0,0 +1,22 @@
// SPDX-License-Identifier: GPL-2.0
use kernel::io::register;
use crate::regs::NV_PBUS_SW_SCRATCH;
// PGSP
register! {
pub(super) NV_PGSP_QUEUE_HEAD(u32) @ 0x00110c00 {
31:0 address;
}
}
// PBUS
register! {
/// Scratch register 0xe used as FRTS firmware error code.
pub(super) NV_PBUS_SW_SCRATCH_0E_FRTS_ERR(u32) => NV_PBUS_SW_SCRATCH[0xe] {
31:16 frts_err_code;
}
}

View File

@ -6,6 +6,7 @@
use kernel::{
device,
dma::Coherent,
io::{
poll::read_poll_timeout,
Io, //
@ -31,6 +32,8 @@
MessageFromGsp, //
},
fw,
GspBootContext,
LibosMemoryRegionInitArgument, //
},
num::FromSafeCast,
sbuffer::SBufferIter,
@ -128,16 +131,14 @@ pub(crate) fn new(data: &[u8], dev: &device::Device) -> Result<(Self, usize)> {
/// GSP Sequencer for executing firmware commands during boot.
pub(crate) struct GspSequencer<'a> {
/// Sequencer information with command data.
seq_info: GspSequence,
/// `Bar0` for register access.
bar: Bar0<'a>,
/// SEC2 falcon for core operations.
sec2_falcon: &'a Falcon<Sec2>,
sec2_falcon: &'a Falcon<'a, Sec2>,
/// GSP falcon for core operations.
gsp_falcon: &'a Falcon<Gsp>,
/// LibOS DMA handle address.
libos_dma_handle: u64,
gsp_falcon: &'a Falcon<'a, Gsp>,
/// LibOS memory region init arguments.
libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
/// Bootloader application version.
bootloader_app_version: u32,
/// Device for logging.
@ -213,16 +214,16 @@ fn run(&self, seq: &GspSequencer<'_>) -> Result {
GspSeqCmd::DelayUs(cmd) => cmd.run(seq),
GspSeqCmd::RegStore(cmd) => cmd.run(seq),
GspSeqCmd::CoreReset => {
seq.gsp_falcon.reset(seq.bar)?;
seq.gsp_falcon.dma_reset(seq.bar);
seq.gsp_falcon.reset()?;
seq.gsp_falcon.dma_reset();
Ok(())
}
GspSeqCmd::CoreStart => {
seq.gsp_falcon.start(seq.bar)?;
seq.gsp_falcon.start()?;
Ok(())
}
GspSeqCmd::CoreWaitForHalt => {
seq.gsp_falcon.wait_till_halted(seq.bar)?;
seq.gsp_falcon.wait_till_halted()?;
Ok(())
}
GspSeqCmd::CoreResume => {
@ -231,35 +232,34 @@ fn run(&self, seq: &GspSequencer<'_>) -> Result {
// sequencer will start both.
// Reset the GSP to prepare it for resuming.
seq.gsp_falcon.reset(seq.bar)?;
seq.gsp_falcon.reset()?;
// Write the libOS DMA handle to GSP mailboxes.
let libos_dma_address = seq.libos.dma_address();
// Write the libOS DMA address to GSP mailboxes.
seq.gsp_falcon.write_mailboxes(
seq.bar,
Some(seq.libos_dma_handle as u32),
Some((seq.libos_dma_handle >> 32) as u32),
Some(libos_dma_address as u32),
Some((libos_dma_address >> 32) as u32),
);
// Start the SEC2 falcon which will trigger GSP-RM to resume on the GSP.
seq.sec2_falcon.start(seq.bar)?;
seq.sec2_falcon.start()?;
// Poll until GSP-RM reload/resume has completed (up to 2 seconds).
seq.gsp_falcon
.check_reload_completed(seq.bar, Delta::from_secs(2))?;
seq.gsp_falcon.check_reload_completed(Delta::from_secs(2))?;
// Verify SEC2 completed successfully by checking its mailbox for errors.
let mbox0 = seq.sec2_falcon.read_mailbox0(seq.bar);
let mbox0 = seq.sec2_falcon.read_mailbox0();
if mbox0 != 0 {
dev_err!(seq.dev, "Sequencer: sec2 errors: {:?}\n", mbox0);
return Err(EIO);
}
// Configure GSP with the bootloader version.
seq.gsp_falcon
.write_os_version(seq.bar, seq.bootloader_app_version);
seq.gsp_falcon.write_os_version(seq.bootloader_app_version);
// Verify the GSP's RISC-V core is active indicating successful GSP boot.
if !seq.gsp_falcon.is_riscv_active(seq.bar) {
if !seq.gsp_falcon.is_riscv_active() {
dev_err!(seq.dev, "Sequencer: RISC-V core is not active\n");
return Err(EIO);
}
@ -270,7 +270,7 @@ fn run(&self, seq: &GspSequencer<'_>) -> Result {
}
/// Iterator over GSP sequencer commands.
pub(crate) struct GspSeqIter<'a> {
struct GspSeqIter<'a> {
/// Command data buffer.
cmd_data: &'a [u8],
/// Current position in the buffer.
@ -283,6 +283,18 @@ pub(crate) struct GspSeqIter<'a> {
dev: &'a device::Device,
}
impl<'a> GspSeqIter<'a> {
fn new(seq: &'a GspSequence, dev: &'a device::Device) -> Self {
Self {
cmd_data: &seq.cmd_data,
current_offset: 0,
total_cmds: seq.cmd_index,
cmds_processed: 0,
dev,
}
}
}
impl<'a> Iterator for GspSeqIter<'a> {
type Item = Result<GspSeqCmd>;
@ -325,37 +337,12 @@ fn next(&mut self) -> Option<Self::Item> {
}
impl<'a> GspSequencer<'a> {
fn iter(&self) -> GspSeqIter<'_> {
let cmd_data = &self.seq_info.cmd_data[..];
GspSeqIter {
cmd_data,
current_offset: 0,
total_cmds: self.seq_info.cmd_index,
cmds_processed: 0,
dev: self.dev,
}
}
}
/// Parameters for running the GSP sequencer.
pub(crate) struct GspSequencerParams<'a> {
/// Bootloader application version.
pub(crate) bootloader_app_version: u32,
/// LibOS DMA handle address.
pub(crate) libos_dma_handle: u64,
/// GSP falcon for core operations.
pub(crate) gsp_falcon: &'a Falcon<Gsp>,
/// SEC2 falcon for core operations.
pub(crate) sec2_falcon: &'a Falcon<Sec2>,
/// Device for logging.
pub(crate) dev: &'a device::Device,
/// BAR0 for register access.
pub(crate) bar: Bar0<'a>,
}
impl<'a> GspSequencer<'a> {
pub(crate) fn run(cmdq: &Cmdq, params: GspSequencerParams<'a>) -> Result {
pub(crate) fn run(
cmdq: &Cmdq,
ctx: &'a GspBootContext<'_, '_>,
libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
bootloader_app_version: u32,
) -> Result {
let seq_info = loop {
match cmdq.receive_msg::<GspSequence>(Cmdq::RECEIVE_TIMEOUT) {
Ok(seq_info) => break seq_info,
@ -365,25 +352,24 @@ pub(crate) fn run(cmdq: &Cmdq, params: GspSequencerParams<'a>) -> Result {
};
let sequencer = GspSequencer {
seq_info,
bar: params.bar,
sec2_falcon: params.sec2_falcon,
gsp_falcon: params.gsp_falcon,
libos_dma_handle: params.libos_dma_handle,
bootloader_app_version: params.bootloader_app_version,
dev: params.dev,
bar: ctx.bar,
sec2_falcon: ctx.sec2_falcon,
gsp_falcon: ctx.gsp_falcon,
libos,
bootloader_app_version,
dev: ctx.dev(),
};
dev_dbg!(sequencer.dev, "Running CPU Sequencer commands\n");
for cmd_result in sequencer.iter() {
for cmd_result in GspSeqIter::new(&seq_info, sequencer.dev) {
match cmd_result {
Ok(cmd) => cmd.run(&sequencer)?,
Err(e) => {
dev_err!(
sequencer.dev,
"Error running command at index {}\n",
sequencer.seq_info.cmd_index
seq_info.cmd_index
);
return Err(e);
}

View File

@ -7,55 +7,53 @@
//! Data Model) messages between the kernel driver and GPU firmware processors
//! such as FSP and GSP.
use kernel::pci::Vendor;
use kernel::{
bitfield,
pci::Vendor,
prelude::*, //
};
/// NVDM message type identifiers carried over MCTP.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
#[repr(u8)]
pub(crate) enum NvdmType {
#[default]
/// Chain of Trust boot message.
Cot = 0x14,
/// FSP command response.
FspResponse = 0x15,
}
use crate::{
bounded_enum,
num, //
};
impl TryFrom<u8> for NvdmType {
type Error = u8;
fn try_from(value: u8) -> Result<Self, Self::Error> {
match value {
x if x == u8::from(Self::Cot) => Ok(Self::Cot),
x if x == u8::from(Self::FspResponse) => Ok(Self::FspResponse),
_ => Err(value),
}
}
}
impl From<NvdmType> for u8 {
fn from(value: NvdmType) -> Self {
value as u8
bounded_enum! {
/// NVDM message type identifiers carried over MCTP.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(crate) enum NvdmType with TryFrom<Bounded<u32, 8>> {
/// PRC (Product Reconfiguration Control) message.
Prc = 0x13,
/// Chain of Trust boot message.
Cot = 0x14,
/// FSP command response.
FspResponse = 0x15,
}
}
bitfield! {
pub(crate) struct MctpHeader(u32), "MCTP transport header for NVIDIA firmware messages." {
31:31 som as bool, "Start-of-message bit.";
30:30 eom as bool, "End-of-message bit.";
29:28 seq as u8, "Packet sequence number.";
23:16 seid as u8, "Source endpoint ID.";
/// MCTP transport header for NVIDIA firmware messages.
pub(crate) struct MctpHeader(u32) {
/// Start-of-message bit.
31:31 som;
/// End-of-message bit.
30:30 eom;
/// Packet sequence number.
29:28 seq;
/// Source endpoint ID.
23:16 seid;
}
}
impl MctpHeader {
/// Builds a single-packet MCTP header (`SOM=1`, `EOM=1`, `SEQ=0`, `SEID=0`).
pub(crate) fn single_packet() -> Self {
Self::default().set_som(true).set_eom(true)
Self::zeroed().with_som(true).with_eom(true)
}
/// Returns whether this is a complete single-packet message (`SOM=1` and `EOM=1`).
pub(crate) fn is_single_packet(self) -> bool {
self.som() && self.eom()
self.som().into_bool() && self.eom().into_bool()
}
}
@ -63,26 +61,30 @@ pub(crate) fn is_single_packet(self) -> bool {
const MSG_TYPE_VENDOR_PCI: u8 = 0x7e;
bitfield! {
pub(crate) struct NvdmHeader(u32), "NVIDIA Vendor-Defined Message header over MCTP." {
31:24 nvdm_type as u8 ?=> NvdmType, "NVDM message type.";
23:8 vendor_id as u16, "PCI vendor ID.";
6:0 msg_type as u8, "MCTP vendor-defined message type.";
/// NVIDIA Vendor-Defined Message header over MCTP.
pub(crate) struct NvdmHeader(u32) {
/// NVDM message type.
31:24 nvdm_type ?=> NvdmType;
/// PCI vendor ID.
23:8 vendor_id;
/// MCTP vendor-defined message type.
6:0 msg_type;
}
}
impl NvdmHeader {
/// Builds an NVDM header for the given message type.
pub(crate) fn new(nvdm_type: NvdmType) -> Self {
Self::default()
.set_msg_type(MSG_TYPE_VENDOR_PCI)
.set_vendor_id(Vendor::NVIDIA.as_raw())
.set_nvdm_type(nvdm_type)
Self::zeroed()
.with_const_msg_type::<{ num::u8_as_u32(MSG_TYPE_VENDOR_PCI) }>()
.with_vendor_id(Vendor::NVIDIA.as_raw())
.with_nvdm_type(nvdm_type)
}
/// Validates this header against the expected NVIDIA NVDM format and type.
pub(crate) fn validate(self, expected_type: NvdmType) -> bool {
self.msg_type() == MSG_TYPE_VENDOR_PCI
&& self.vendor_id() == Vendor::NVIDIA.as_raw()
u8::from(self.msg_type()) == MSG_TYPE_VENDOR_PCI
&& u16::from(self.vendor_id()) == Vendor::NVIDIA.as_raw()
&& matches!(self.nvdm_type(), Ok(nvdm_type) if nvdm_type == expected_type)
}
}

View File

@ -10,9 +10,6 @@
InPlaceModule, //
};
#[macro_use]
mod bitfield;
mod driver;
mod falcon;
mod fb;
@ -26,6 +23,7 @@
mod regs;
mod sbuffer;
mod vbios;
mod vgpu;
pub(crate) const MODULE_NAME: &core::ffi::CStr = <LocalModule as kernel::ModuleMetadata>::NAME;
@ -54,7 +52,7 @@ struct NovaCoreModule {
impl InPlaceModule for NovaCoreModule {
fn init(module: &'static kernel::ThisModule) -> impl PinInit<Self, Error> {
let dir = debugfs::Dir::new(kernel::c_str!("nova-core"));
let dir = debugfs::Dir::new(c"nova-core");
// SAFETY: We are the only driver code running during init, so there
// cannot be any concurrent access to `DEBUGFS_ROOT`.

View File

@ -0,0 +1,15 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
/*
* Exports Rust symbols from the `nova_core` crate for use by dependent modules.
*
* This is a workaround until the build system supports Rust cross-module
* dependencies natively.
*/
#include <linux/export.h>
#define EXPORT_SYMBOL_RUST_GPL(sym) extern int sym; EXPORT_SYMBOL_GPL(sym)
#include "exports_nova_core_generated.h"

View File

@ -109,130 +109,6 @@ fn fmt(&self, f: &mut kernel::fmt::Formatter<'_>) -> kernel::fmt::Result {
register! {
pub(crate) NV_PBUS_SW_SCRATCH(u32)[64] @ 0x00001400 {}
/// Scratch register 0xe used as FRTS firmware error code.
pub(crate) NV_PBUS_SW_SCRATCH_0E_FRTS_ERR(u32) => NV_PBUS_SW_SCRATCH[0xe] {
31:16 frts_err_code;
}
}
// PFB
register! {
/// Low bits of the physical system memory address used by the GPU to perform sysmembar
/// operations (see [`crate::fb::SysmemFlush`]).
pub(crate) NV_PFB_NISO_FLUSH_SYSMEM_ADDR(u32) @ 0x00100c10 {
31:0 adr_39_08;
}
/// High bits of the physical system memory address used by the GPU to perform sysmembar
/// operations (see [`crate::fb::SysmemFlush`]).
pub(crate) NV_PFB_NISO_FLUSH_SYSMEM_ADDR_HI(u32) @ 0x00100c40 {
23:0 adr_63_40;
}
pub(crate) NV_PFB_PRI_MMU_LOCAL_MEMORY_RANGE(u32) @ 0x00100ce0 {
30:30 ecc_mode_enabled => bool;
9:4 lower_mag;
3:0 lower_scale;
}
pub(crate) NV_PFB_PRI_MMU_WPR2_ADDR_LO(u32) @ 0x001fa824 {
/// Bits 12..40 of the lower (inclusive) bound of the WPR2 region.
31:4 lo_val;
}
pub(crate) NV_PFB_PRI_MMU_WPR2_ADDR_HI(u32) @ 0x001fa828 {
/// Bits 12..40 of the higher (exclusive) bound of the WPR2 region.
31:4 hi_val;
}
}
/// Base of the GB10x HSHUB0 register window (`NV_HSHUB0_PRIV_BASE` in Open RM).
///
/// The base is provided by the GB10x framebuffer HAL.
pub(crate) struct Hshub0Base(());
/// Base of the GB20x FBHUB0 register window (`NV_FBHUB0_PRI_BASE` in Open RM).
///
/// The base is provided by the GB20x framebuffer HAL.
pub(crate) struct Fbhub0Base(());
register! {
// GB10x sysmem flush registers, relative to the HSHUB0 base. GB10x routes sysmembar
// through a primary and an EG (egress) pair that must both be programmed to the same
// address. Hardware ignores bits 7:0 of each LO register. The boot path uses a fixed
// HSHUB0 base, so the multiple runtime-discovered HSHUB bases are not needed here.
pub(crate) NV_PFB_HSHUB_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Hshub0Base + 0x00000e50 {
31:0 adr => u32;
}
pub(crate) NV_PFB_HSHUB_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Hshub0Base + 0x00000e54 {
19:0 adr;
}
pub(crate) NV_PFB_HSHUB_EG_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Hshub0Base + 0x000006c0 {
31:0 adr => u32;
}
pub(crate) NV_PFB_HSHUB_EG_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Hshub0Base + 0x000006c4 {
19:0 adr;
}
// GB20x sysmem flush registers, relative to the FBHUB0 base. Unlike the older
// NV_PFB_NISO_FLUSH_SYSMEM_ADDR registers which encode the address with an 8-bit
// right-shift, these take the raw address split into lower and upper halves. Hardware
// ignores bits 7:0 of the LO register.
pub(crate) NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Fbhub0Base + 0x00001d58 {
31:0 adr => u32;
}
pub(crate) NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Fbhub0Base + 0x00001d5c {
19:0 adr;
}
}
impl NV_PFB_PRI_MMU_LOCAL_MEMORY_RANGE {
/// Returns the usable framebuffer size, in bytes.
pub(crate) fn usable_fb_size(self) -> u64 {
let size = (u64::from(self.lower_mag()) << u64::from(self.lower_scale())) * u64::SZ_1M;
if self.ecc_mode_enabled() {
// Remove the amount of memory reserved for ECC (one per 16 units).
size / 16 * 15
} else {
size
}
}
}
impl NV_PFB_PRI_MMU_WPR2_ADDR_LO {
/// Returns the lower (inclusive) bound of the WPR2 region.
pub(crate) fn lower_bound(self) -> u64 {
u64::from(self.lo_val()) << 12
}
}
impl NV_PFB_PRI_MMU_WPR2_ADDR_HI {
/// Returns the higher (exclusive) bound of the WPR2 region.
///
/// A value of zero means the WPR2 region is not set.
pub(crate) fn higher_bound(self) -> u64 {
u64::from(self.hi_val()) << 12
}
/// Returns whether the WPR2 region is currently set.
pub(crate) fn is_wpr2_set(self) -> bool {
self.hi_val() != 0
}
}
// PGSP
register! {
pub(crate) NV_PGSP_QUEUE_HEAD(u32) @ 0x00110c00 {
31:0 address;
}
}
// PGC6 register space.
@ -294,28 +170,6 @@ pub(crate) fn usable_fb_size(self) -> u64 {
}
}
// PDISP
register! {
pub(crate) NV_PDISP_VGA_WORKSPACE_BASE(u32) @ 0x00625f04 {
/// VGA workspace base address divided by 0x10000.
31:8 addr;
/// Set if the `addr` field is valid.
3:3 status_valid => bool;
}
}
impl NV_PDISP_VGA_WORKSPACE_BASE {
/// Returns the base address of the VGA workspace, or `None` if none exists.
pub(crate) fn vga_workspace_addr(self) -> Option<u64> {
if self.status_valid() {
Some(u64::from(self.addr()) << 16)
} else {
None
}
}
}
// FUSE
pub(crate) const NV_FUSE_OPT_FPF_SIZE: usize = 16;
@ -570,7 +424,7 @@ pub(crate) fn mem_scrubbing_done(self) -> bool {
/// GA102 and later.
pub(crate) NV_PRISCV_RISCV_CPUCTL(u32) @ PFalcon2Base + 0x00000388 {
7:7 active_stat => bool;
0:0 halted => bool;
4:4 halted => bool;
}
/// GA102 and later.

View File

@ -13,11 +13,8 @@
register,
sizes::SZ_4K,
sync::aref::ARef,
transmute::FromBytes,
};
use zerocopy::FromBytes as _;
use crate::{
driver::Bar0,
firmware::{
@ -359,7 +356,7 @@ pub(crate) fn fwsec_image(&self) -> &FwSecBiosImage {
}
/// PCI Data Structure as defined in PCI Firmware Specification
#[derive(Debug, Clone)]
#[derive(Debug, Clone, FromBytes)]
#[repr(C)]
struct PcirStruct {
/// PCI Data Structure signature ("PCIR" or "NPDS")
@ -388,15 +385,12 @@ struct PcirStruct {
max_runtime_image_len: u16,
}
// SAFETY: all bit patterns are valid for `PcirStruct`.
unsafe impl FromBytes for PcirStruct {}
impl PcirStruct {
/// The bit in `last_image` that indicates the last image.
const LAST_IMAGE_BIT_MASK: u8 = 0x80;
fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
let (pcir, _) = PcirStruct::from_bytes_copy_prefix(data).ok_or(EINVAL)?;
let (pcir, _) = PcirStruct::read_from_prefix(data).map_err(|_| EINVAL)?;
// Signature should be "PCIR" (0x52494350) or "NPDS" (0x5344504e).
if &pcir.signature != b"PCIR" && &pcir.signature != b"NPDS" {
@ -432,7 +426,7 @@ fn image_size_bytes(&self) -> usize {
/// This is the head of the BIT table, that is used to locate the Falcon data. The BIT table (with
/// its header) is in the [`PciAtBiosImage`] and the falcon data it is pointing to is in the
/// [`FwSecBiosImage`].
#[derive(Debug, Clone, Copy)]
#[derive(Debug, Clone, Copy, FromBytes)]
#[repr(C)]
struct BitHeader {
/// 0h: BIT Header Identifier (BMP=0x7FFF/BIT=0xB8FF)
@ -451,12 +445,9 @@ struct BitHeader {
checksum: u8,
}
// SAFETY: all bit patterns are valid for `BitHeader`.
unsafe impl FromBytes for BitHeader {}
impl BitHeader {
fn new(data: &[u8]) -> Result<Self> {
let (header, _) = BitHeader::from_bytes_copy_prefix(data).ok_or(EINVAL)?;
let (header, _) = BitHeader::read_from_prefix(data).map_err(|_| EINVAL)?;
// Check header ID and signature
if header.id != 0xB8FF || &header.signature != b"BIT\0" {
@ -468,7 +459,7 @@ fn new(data: &[u8]) -> Result<Self> {
}
/// BIT Token Entry: Records in the BIT table followed by the BIT header.
#[derive(Debug, Clone, Copy)]
#[derive(Debug, Clone, Copy, FromBytes)]
#[repr(C)]
struct BitToken {
/// 00h: Token identifier
@ -481,9 +472,6 @@ struct BitToken {
data_offset: u16,
}
// SAFETY: all bit patterns are valid for `BitToken`.
unsafe impl FromBytes for BitToken {}
impl BitToken {
/// BIT token ID for Falcon data.
const ID_FALCON_DATA: u8 = 0x70;
@ -508,7 +496,7 @@ fn from_id(image: &PciAtBiosImage, token_id: u8) -> Result<Self> {
.and_then(|data| data.get(..entry_size))
.ok_or(EINVAL)?;
let (token, _) = BitToken::from_bytes_copy_prefix(entry).ok_or(EINVAL)?;
let (token, _) = BitToken::read_from_prefix(entry).map_err(|_| EINVAL)?;
// Check if this token has the requested ID
if token.id == token_id {
@ -525,7 +513,7 @@ fn from_id(image: &PciAtBiosImage, token_id: u8) -> Result<Self> {
///
/// This header is at the beginning of every image in the set of images in the ROM. It contains a
/// pointer to the PCI Data Structure which describes the image.
#[derive(Debug, Clone, Copy)]
#[derive(Debug, Clone, Copy, FromBytes)]
#[repr(C)]
struct PciRomHeader {
/// 00h: Signature (0xAA55)
@ -536,13 +524,10 @@ struct PciRomHeader {
pci_data_struct_offset: u16,
}
// SAFETY: all bit patterns are valid for `PciRomHeader`.
unsafe impl FromBytes for PciRomHeader {}
impl PciRomHeader {
fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
let (rom_header, _) = PciRomHeader::from_bytes_copy_prefix(data)
.ok_or(EINVAL)
let (rom_header, _) = PciRomHeader::read_from_prefix(data)
.map_err(|_| EINVAL)
.inspect_err(|_| dev_err!(dev, "Not enough data for ROM header\n"))?;
// Check for valid ROM signatures.
@ -564,7 +549,7 @@ fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
/// PCI Data Structure. It contains some fields that are redundant with the PCI Data Structure, but
/// are needed for traversing the BIOS images. It is expected to be present in all BIOS images
/// except for NBSI images.
#[derive(Debug, Clone)]
#[derive(Debug, Clone, FromBytes)]
#[repr(C)]
struct NpdeStruct {
/// 00h: Signature ("NPDE")
@ -579,15 +564,12 @@ struct NpdeStruct {
last_image: u8,
}
// SAFETY: all bit patterns are valid for `NpdeStruct`.
unsafe impl FromBytes for NpdeStruct {}
impl NpdeStruct {
/// The bit in `last_image` that indicates the last image.
const LAST_IMAGE_BIT_MASK: u8 = 0x80;
fn new(dev: &device::Device, data: &[u8]) -> Option<Self> {
let (npde, _) = NpdeStruct::from_bytes_copy_prefix(data)?;
let (npde, _) = NpdeStruct::read_from_prefix(data).ok()?;
// Signature should be "NPDE" (0x4544504E).
if &npde.signature != b"NPDE" {
@ -784,7 +766,7 @@ fn falcon_data_offset(&self, dev: &device::Device) -> Result<usize> {
let data = &self.base.data;
let (ptr, _) = data
.get(offset..)
.and_then(u32::from_bytes_copy_prefix)
.and_then(|p| u32::read_from_prefix(p).ok())
.ok_or(EINVAL)?;
usize::from_safe_cast(ptr)
@ -814,6 +796,7 @@ fn try_from(base: BiosImage) -> Result<Self> {
/// The [`PmuLookupTableEntry`] structure is a single entry in the [`PmuLookupTable`].
///
/// See the [`PmuLookupTable`] description for more information.
#[derive(FromBytes)]
#[repr(C, packed)]
struct PmuLookupTableEntry {
application_id: u8,
@ -821,9 +804,6 @@ struct PmuLookupTableEntry {
data: u32,
}
// SAFETY: all bit patterns are valid for `PmuLookupTableEntry`.
unsafe impl FromBytes for PmuLookupTableEntry {}
impl PmuLookupTableEntry {
/// PMU lookup table application ID for firmware security license ucode.
#[expect(dead_code)]
@ -836,6 +816,7 @@ impl PmuLookupTableEntry {
}
#[repr(C)]
#[derive(FromBytes)]
struct PmuLookupTableHeader {
version: u8,
header_len: u8,
@ -843,9 +824,6 @@ struct PmuLookupTableHeader {
entry_count: u8,
}
// SAFETY: all bit patterns are valid for `PmuLookupTableHeader`.
unsafe impl FromBytes for PmuLookupTableHeader {}
/// The [`PmuLookupTableEntry`] structure is used to find the [`PmuLookupTableEntry`] for a given
/// application ID.
///
@ -857,7 +835,7 @@ struct PmuLookupTable {
impl PmuLookupTable {
fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
let (header, _) = PmuLookupTableHeader::from_bytes_copy_prefix(data).ok_or(EINVAL)?;
let (header, _) = PmuLookupTableHeader::read_from_prefix(data).map_err(|_| EINVAL)?;
let header_len = usize::from(header.header_len);
let entry_len = usize::from(header.entry_len);
@ -872,8 +850,8 @@ fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
let mut entries = KVVec::with_capacity(entry_count, GFP_KERNEL)?;
for i in 0..entry_count {
let (entry, _) = PmuLookupTableEntry::from_bytes_copy_prefix(&data[i * entry_len..])
.ok_or(EINVAL)?;
let (entry, _) = PmuLookupTableEntry::read_from_prefix(&data[i * entry_len..])
.map_err(|_| EINVAL)?;
entries.push(entry, GFP_KERNEL)?;
}
@ -929,15 +907,11 @@ pub(crate) fn header(&self) -> Result<FalconUCodeDesc> {
let ver = data.get(1).copied().ok_or(EINVAL)?;
match ver {
2 => {
let v2 = FalconUCodeDescV2::read_from_prefix(data)
.map_err(|_| EINVAL)?
.0;
let (v2, _) = FalconUCodeDescV2::read_from_prefix(data).map_err(|_| EINVAL)?;
Ok(FalconUCodeDesc::V2(v2))
}
3 => {
let v3 = FalconUCodeDescV3::from_bytes_copy_prefix(data)
.ok_or(EINVAL)?
.0;
let (v3, _) = FalconUCodeDescV3::read_from_prefix(data).map_err(|_| EINVAL)?;
Ok(FalconUCodeDesc::V3(v3))
}
_ => {

View File

@ -0,0 +1,91 @@
// SPDX-License-Identifier: GPL-2.0
use core::num::NonZero;
use kernel::{
device,
pci,
prelude::*, //
};
use crate::{
fsp::{
Fsp,
VgpuMode, //
},
gpu::Chipset, //
};
mod hal;
/// vGPU state detected during GPU construction.
#[derive(Debug, Clone, Copy)]
pub(crate) enum VgpuState {
/// vGPU mode is not enabled for this boot.
Disabled,
/// vGPU mode is enabled for this boot.
Enabled {
/// Total number of SR-IOV VFs supported by this device.
total_vfs: NonZero<u16>,
},
}
/// vGPU state manager.
pub(crate) struct VgpuManager {
state: VgpuState,
}
impl VgpuManager {
/// Creates a vGPU manager by querying SR-IOV and the FSP PRC vGPU knob.
pub(crate) fn new(
pdev: &pci::Device<device::Core<'_>>,
chipset: Chipset,
fsp: Option<&mut Fsp<'_>>,
) -> Self {
let state = Self::detect_state(pdev, chipset, fsp).unwrap_or_else(|e| {
dev_warn!(
pdev,
"vGPU state detection failed: {:?}; disabling vGPU\n",
e
);
VgpuState::Disabled
});
dev_dbg!(pdev, "vGPU state: {:?}\n", state);
Self { state }
}
/// Detects the vGPU state from the chipset, SR-IOV capability and FSP PRC knob.
fn detect_state(
pdev: &pci::Device<device::Core<'_>>,
chipset: Chipset,
fsp: Option<&mut Fsp<'_>>,
) -> Result<VgpuState> {
if !hal::vgpu_hal(chipset).supports_vgpu() {
return Ok(VgpuState::Disabled);
}
let Some(total_vfs) = pdev.sriov_get_totalvfs() else {
return Ok(VgpuState::Disabled);
};
if total_vfs.get() < 2 {
// The current vGPU path does not support single-VF SR-IOV devices yet.
// Treat one total VF as vGPU-disabled for now; single-VF support can relax
// this gate once the manager handles that topology.
return Ok(VgpuState::Disabled);
}
let fsp = fsp.ok_or(ENODEV)?;
match fsp.read_vgpu_mode(pdev.as_ref())? {
VgpuMode::Enabled => Ok(VgpuState::Enabled { total_vfs }),
VgpuMode::Disabled => Ok(VgpuState::Disabled),
}
}
/// Returns the detected vGPU state for this boot.
pub(crate) fn state(&self) -> VgpuState {
self.state
}
}

View File

@ -0,0 +1,25 @@
// SPDX-License-Identifier: GPL-2.0
use crate::gpu::{
Architecture,
Chipset, //
};
mod gb202;
mod tu102;
pub(super) trait VgpuHal {
/// Returns whether this chipset can support vGPU.
fn supports_vgpu(&self) -> bool;
}
pub(super) fn vgpu_hal(chipset: Chipset) -> &'static dyn VgpuHal {
match chipset.arch() {
Architecture::BlackwellGB20x => gb202::GB202_HAL,
Architecture::Turing
| Architecture::Ampere
| Architecture::Hopper
| Architecture::Ada
| Architecture::BlackwellGB10x => tu102::TU102_HAL,
}
}

View File

@ -0,0 +1,15 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use crate::vgpu::hal::VgpuHal;
struct Gb202;
impl VgpuHal for Gb202 {
fn supports_vgpu(&self) -> bool {
true
}
}
const GB202: Gb202 = Gb202;
pub(super) const GB202_HAL: &dyn VgpuHal = &GB202;

View File

@ -0,0 +1,15 @@
// SPDX-License-Identifier: GPL-2.0
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
use crate::vgpu::hal::VgpuHal;
struct Tu102;
impl VgpuHal for Tu102 {
fn supports_vgpu(&self) -> bool {
false
}
}
const TU102: Tu102 = Tu102;
pub(super) const TU102_HAL: &dyn VgpuHal = &TU102;

View File

@ -1283,7 +1283,7 @@ EXPORT_SYMBOL_GPL(pci_sriov_set_totalvfs);
* SRIOV capability value of TotalVFs or the value of driver_max_VFs
* if the driver reduced it. Otherwise 0.
*/
int pci_sriov_get_totalvfs(struct pci_dev *dev)
unsigned int pci_sriov_get_totalvfs(struct pci_dev *dev)
{
if (!dev->is_physfn)
return 0;

View File

@ -20,7 +20,6 @@
//! this method is not used in this driver.
//!
use core::ops::Deref;
use kernel::{
clk::Clk,
device::{Bound, Core, Device},
@ -213,8 +212,7 @@ fn read_waveform(
) -> Result<Self::WfHw> {
let data = chip.drvdata();
let hwpwm = pwm.hwpwm();
let iomem_accessor = data.iomem.access(parent_dev)?;
let iomap = iomem_accessor.deref();
let iomap = data.iomem.access(parent_dev)?;
let ctrl = iomap.try_read32(th1520_pwm_ctrl(hwpwm))?;
let period_cycles = iomap.try_read32(th1520_pwm_per(hwpwm))?;
@ -248,8 +246,7 @@ fn write_waveform(
) -> Result {
let data = chip.drvdata();
let hwpwm = pwm.hwpwm();
let iomem_accessor = data.iomem.access(parent_dev)?;
let iomap = iomem_accessor.deref();
let iomap = data.iomem.access(parent_dev)?;
let duty_cycles = iomap.try_read32(th1520_pwm_fp(hwpwm))?;
let was_enabled = duty_cycles != 0;

View File

@ -2569,7 +2569,7 @@ void pci_iov_remove_virtfn(struct pci_dev *dev, int id);
int pci_num_vf(struct pci_dev *dev);
int pci_vfs_assigned(struct pci_dev *dev);
int pci_sriov_set_totalvfs(struct pci_dev *dev, u16 numvfs);
int pci_sriov_get_totalvfs(struct pci_dev *dev);
unsigned int pci_sriov_get_totalvfs(struct pci_dev *dev);
int pci_sriov_configure_simple(struct pci_dev *dev, int nr_virtfn);
resource_size_t pci_iov_resource_size(const struct pci_dev *dev, int resno);
int pci_iov_vf_bar_set_size(struct pci_dev *dev, int resno, int size);
@ -2622,7 +2622,7 @@ static inline int pci_vfs_assigned(struct pci_dev *dev)
{ return 0; }
static inline int pci_sriov_set_totalvfs(struct pci_dev *dev, u16 numvfs)
{ return 0; }
static inline int pci_sriov_get_totalvfs(struct pci_dev *dev)
static inline unsigned int pci_sriov_get_totalvfs(struct pci_dev *dev)
{ return 0; }
#define pci_sriov_configure_simple NULL
static inline resource_size_t pci_iov_resource_size(const struct pci_dev *dev,

View File

@ -19,6 +19,19 @@ __rust_helper void rust_helper_iounmap(void __iomem *addr)
iounmap(addr);
}
__rust_helper void rust_helper_memcpy_fromio(void *dst,
const volatile void __iomem *src,
size_t count)
{
memcpy_fromio(dst, src, count);
}
__rust_helper void rust_helper_memcpy_toio(volatile void __iomem *dst,
const void *src, size_t count)
{
memcpy_toio(dst, src, count);
}
__rust_helper u8 rust_helper_readb(const void __iomem *addr)
{
return readb(addr);

View File

@ -24,6 +24,14 @@ __rust_helper bool rust_helper_dev_is_pci(const struct device *dev)
return dev_is_pci(dev);
}
#ifndef CONFIG_PCI_IOV
__rust_helper unsigned int
rust_helper_pci_sriov_get_totalvfs(struct pci_dev *pdev)
{
return pci_sriov_get_totalvfs(pdev);
}
#endif
#ifndef CONFIG_PCI_MSI
__rust_helper int rust_helper_pci_alloc_irq_vectors(struct pci_dev *dev,
unsigned int min_vecs,

View File

@ -9,6 +9,7 @@
Vmalloc,
VmallocPageIter, //
},
flags::__GFP_ZERO,
layout::ArrayLayout,
AllocError,
Allocator,
@ -51,6 +52,8 @@
}, //
};
use pin_init::Zeroable;
mod errors;
pub use self::errors::{InsertError, PushError, RemoveError};
@ -532,6 +535,30 @@ pub fn with_capacity(capacity: usize, flags: Flags) -> Result<Self, AllocError>
Ok(v)
}
/// Creates a new [`Vec`] with `n` zero-initialized elements.
///
/// # Examples
///
/// ```
/// let v = KVec::<u32>::zeroed(20, GFP_KERNEL)?;
///
/// assert!(v.iter().all(|&x| x == 0));
/// # Ok::<(), Error>(())
/// ```
pub fn zeroed(n: usize, flags: Flags) -> Result<Self, AllocError>
where
T: Zeroable,
{
let mut v = Self::with_capacity(n, flags | __GFP_ZERO)?;
// SAFETY:
// - `n <= capacity - len`: `with_capacity(n)` guarantees capacity >= n, len is 0.
// - All elements in `[0, n)` are initialized: `__GFP_ZERO` zeroes the allocation,
// and `T: Zeroable` guarantees all-zeroes is a valid bit pattern.
unsafe { v.inc_len(n) };
Ok(v)
}
/// Creates a `Vec<T, A>` from a pointer, a length and a capacity using the allocator `A`.
///
/// # Examples

View File

@ -68,17 +68,19 @@ struct Inner<T> {
/// devres::Devres,
/// io::{
/// Io,
/// IoKnownSize,
/// IoBase,
/// Mmio,
/// MmioRaw,
/// PhysAddr, //
/// MmioBackend,
/// PhysAddr,
/// Region, //
/// },
/// prelude::*,
/// };
/// use core::ops::Deref;
///
/// // See also [`pci::Bar`] for a real example.
/// struct IoMem<const SIZE: usize>(MmioRaw<SIZE>);
/// struct IoMem<const SIZE: usize>(MmioRaw<Region<SIZE>>);
///
/// impl<const SIZE: usize> IoMem<SIZE> {
/// /// # Safety
@ -93,7 +95,7 @@ struct Inner<T> {
/// return Err(ENOMEM);
/// }
///
/// Ok(IoMem(MmioRaw::new(addr as usize, SIZE)?))
/// Ok(IoMem(MmioRaw::new_region(addr as usize, SIZE)?))
/// }
/// }
///
@ -104,12 +106,13 @@ struct Inner<T> {
/// }
/// }
///
/// impl<const SIZE: usize> Deref for IoMem<SIZE> {
/// type Target = Mmio<SIZE>;
/// impl<'a, const SIZE: usize> IoBase<'a> for &'a IoMem<SIZE> {
/// type Backend = MmioBackend;
/// type Target = Region<SIZE>;
///
/// fn deref(&self) -> &Self::Target {
/// fn as_view(self) -> Mmio<'a, Region<SIZE>> {
/// // SAFETY: The memory range stored in `self` has been properly mapped in `Self::new`.
/// unsafe { Mmio::from_raw(&self.0) }
/// unsafe { Mmio::from_raw(self.0) }
/// }
/// }
/// # fn no_run(dev: &Device<Bound>) -> Result<(), Error> {
@ -297,10 +300,7 @@ pub fn device(&self) -> &Device {
/// use kernel::{
/// device::Core,
/// devres::Devres,
/// io::{
/// Io,
/// IoKnownSize, //
/// },
/// io::Io,
/// pci, //
/// };
///

View File

@ -14,14 +14,22 @@
},
error::to_result,
fs::file,
io::{
IoBackend,
IoBase,
IoCapable,
IoCopyable,
SysMem,
SysMemBackend, //
},
prelude::*,
ptr::KnownSize,
sync::aref::ARef,
transmute::{
AsBytes,
FromBytes, //
}, //
uaccess::UserSliceWriter,
},
uaccess::UserSliceWriter, //
};
use core::{
ops::{
@ -577,7 +585,7 @@ fn from(value: CoherentBox<T>) -> Self {
/// # Invariants
///
/// - For the lifetime of an instance of [`Coherent`], the `cpu_addr` is a valid pointer
/// to an allocated region of coherent memory and `dma_handle` is the DMA address base of the
/// to an allocated region of coherent memory and `dma_addr` is the DMA address base of the
/// region.
/// - The size in bytes of the allocation is equal to size information via pointer.
// TODO
@ -594,7 +602,7 @@ fn from(value: CoherentBox<T>) -> Self {
// entire `Coherent` including the allocated memory itself.
pub struct Coherent<T: KnownSize + ?Sized> {
dev: ARef<device::Device>,
dma_handle: DmaAddress,
dma_addr: DmaAddress,
cpu_addr: NonNull<T>,
dma_attrs: Attrs,
}
@ -619,11 +627,10 @@ pub fn as_mut_ptr(&self) -> *mut T {
self.cpu_addr.as_ptr()
}
/// Returns a DMA handle which may be given to the device as the DMA address base of
/// the region.
/// Returns a DMA address which may be given to the device as the base of the region.
#[inline]
pub fn dma_handle(&self) -> DmaAddress {
self.dma_handle
pub fn dma_address(&self) -> DmaAddress {
self.dma_addr
}
/// Returns a reference to the data in the region.
@ -654,52 +661,6 @@ pub unsafe fn as_mut(&self) -> &mut T {
// SAFETY: per safety requirement.
unsafe { &mut *self.as_mut_ptr() }
}
/// Reads the value of `field` and ensures that its type is [`FromBytes`].
///
/// # Safety
///
/// This must be called from the [`dma_read`] macro which ensures that the `field` pointer is
/// validated beforehand.
///
/// Public but hidden since it should only be used from [`dma_read`] macro.
#[doc(hidden)]
pub unsafe fn field_read<F: FromBytes>(&self, field: *const F) -> F {
// SAFETY:
// - By the safety requirements field is valid.
// - Using read_volatile() here is not sound as per the usual rules, the usage here is
// a special exception with the following notes in place. When dealing with a potential
// race from a hardware or code outside kernel (e.g. user-space program), we need that
// read on a valid memory is not UB. Currently read_volatile() is used for this, and the
// rationale behind is that it should generate the same code as READ_ONCE() which the
// kernel already relies on to avoid UB on data races. Note that the usage of
// read_volatile() is limited to this particular case, it cannot be used to prevent
// the UB caused by racing between two kernel functions nor do they provide atomicity.
unsafe { field.read_volatile() }
}
/// Writes a value to `field` and ensures that its type is [`AsBytes`].
///
/// # Safety
///
/// This must be called from the [`dma_write`] macro which ensures that the `field` pointer is
/// validated beforehand.
///
/// Public but hidden since it should only be used from [`dma_write`] macro.
#[doc(hidden)]
pub unsafe fn field_write<F: AsBytes>(&self, field: *mut F, val: F) {
// SAFETY:
// - By the safety requirements field is valid.
// - Using write_volatile() here is not sound as per the usual rules, the usage here is
// a special exception with the following notes in place. When dealing with a potential
// race from a hardware or code outside kernel (e.g. user-space program), we need that
// write on a valid memory is not UB. Currently write_volatile() is used for this, and the
// rationale behind is that it should generate the same code as WRITE_ONCE() which the
// kernel already relies on to avoid UB on data races. Note that the usage of
// write_volatile() is limited to this particular case, it cannot be used to prevent
// the UB caused by racing between two kernel functions nor do they provide atomicity.
unsafe { field.write_volatile(val) }
}
}
impl<T: AsBytes + FromBytes> Coherent<T> {
@ -716,13 +677,13 @@ fn alloc_with_attrs(
);
}
let mut dma_handle = 0;
let mut dma_addr = 0;
// SAFETY: Device pointer is guaranteed as valid by the type invariant on `Device`.
let addr = unsafe {
bindings::dma_alloc_attrs(
dev.as_raw(),
core::mem::size_of::<T>(),
&mut dma_handle,
&mut dma_addr,
gfp_flags.as_raw(),
dma_attrs.as_raw(),
)
@ -734,7 +695,7 @@ fn alloc_with_attrs(
// - We also hold a refcounted reference to the device.
Ok(Self {
dev: dev.into(),
dma_handle,
dma_addr,
cpu_addr,
dma_attrs,
})
@ -833,13 +794,13 @@ fn alloc_slice_with_attrs(
}
let size = core::mem::size_of::<T>().checked_mul(len).ok_or(ENOMEM)?;
let mut dma_handle = 0;
let mut dma_addr = 0;
// SAFETY: Device pointer is guaranteed as valid by the type invariant on `Device`.
let addr = unsafe {
bindings::dma_alloc_attrs(
dev.as_raw(),
size,
&mut dma_handle,
&mut dma_addr,
gfp_flags.as_raw(),
dma_attrs.as_raw(),
)
@ -851,7 +812,7 @@ fn alloc_slice_with_attrs(
// - We also hold a refcounted reference to the device.
Ok(Coherent {
dev: dev.into(),
dma_handle,
dma_addr,
cpu_addr,
dma_attrs,
})
@ -965,14 +926,14 @@ impl<T: KnownSize + ?Sized> Drop for Coherent<T> {
fn drop(&mut self) {
let size = T::size(self.cpu_addr.as_ptr());
// SAFETY: Device pointer is guaranteed as valid by the type invariant on `Device`.
// The cpu address, and the dma handle are valid due to the type invariants on
// The cpu address, and the dma address are valid due to the type invariants on
// `Coherent`.
unsafe {
bindings::dma_free_attrs(
self.dev.as_raw(),
size,
self.cpu_addr.as_ptr().cast(),
self.dma_handle,
self.dma_addr,
self.dma_attrs.as_raw(),
)
}
@ -1027,13 +988,13 @@ fn write_to_slice(
///
/// - `cpu_handle` holds the opaque handle returned by `dma_alloc_attrs` with
/// `DMA_ATTR_NO_KERNEL_MAPPING` set, and is only valid for passing back to `dma_free_attrs`.
/// - `dma_handle` is the corresponding bus address for device DMA.
/// - `dma_addr` is the corresponding bus address for device DMA.
/// - `size` is the allocation size in bytes as passed to `dma_alloc_attrs`.
/// - `dma_attrs` contains the attributes used for the allocation, always including
/// `DMA_ATTR_NO_KERNEL_MAPPING`.
pub struct CoherentHandle {
dev: ARef<device::Device>,
dma_handle: DmaAddress,
dma_addr: DmaAddress,
cpu_handle: NonNull<c_void>,
size: usize,
dma_attrs: Attrs,
@ -1057,13 +1018,13 @@ pub fn alloc_with_attrs(
}
let dma_attrs = dma_attrs | Attrs(bindings::DMA_ATTR_NO_KERNEL_MAPPING);
let mut dma_handle = 0;
let mut dma_addr = 0;
// SAFETY: `dev.as_raw()` is valid by the type invariant on `device::Device`.
let cpu_handle = unsafe {
bindings::dma_alloc_attrs(
dev.as_raw(),
size,
&mut dma_handle,
&mut dma_addr,
gfp_flags.as_raw(),
dma_attrs.as_raw(),
)
@ -1072,11 +1033,11 @@ pub fn alloc_with_attrs(
let cpu_handle = NonNull::new(cpu_handle).ok_or(ENOMEM)?;
// INVARIANT: `cpu_handle` is the opaque handle from a successful `dma_alloc_attrs` call
// with `DMA_ATTR_NO_KERNEL_MAPPING`, `dma_handle` is the corresponding DMA address,
// with `DMA_ATTR_NO_KERNEL_MAPPING`, `dma_addr` is the corresponding DMA address,
// and we hold a refcounted reference to the device.
Ok(Self {
dev: dev.into(),
dma_handle,
dma_addr,
cpu_handle,
size,
dma_attrs,
@ -1093,12 +1054,12 @@ pub fn alloc(
Self::alloc_with_attrs(dev, size, gfp_flags, Attrs(0))
}
/// Returns the DMA handle for this allocation.
/// Returns the DMA address for this allocation.
///
/// This address can be programmed into device hardware for DMA access.
#[inline]
pub fn dma_handle(&self) -> DmaAddress {
self.dma_handle
pub fn dma_address(&self) -> DmaAddress {
self.dma_addr
}
/// Returns the size in bytes of this allocation.
@ -1117,100 +1078,170 @@ fn drop(&mut self) {
self.dev.as_raw(),
self.size,
self.cpu_handle.as_ptr(),
self.dma_handle,
self.dma_addr,
self.dma_attrs.as_raw(),
)
}
}
}
// SAFETY: `CoherentHandle` only holds a device reference, a DMA handle, an opaque CPU handle,
// SAFETY: `CoherentHandle` only holds a device reference, a DMA address, an opaque CPU handle,
// and a size. None of these are tied to a specific thread.
unsafe impl Send for CoherentHandle {}
// SAFETY: `CoherentHandle` provides no CPU access to the underlying allocation. The only
// operations on `&CoherentHandle` are reading the DMA handle and size, both of which are
// operations on `&CoherentHandle` are reading the DMA address and size, both of which are
// plain `Copy` values.
unsafe impl Sync for CoherentHandle {}
/// Reads a field of an item from an allocated region of structs.
/// View type for `Coherent`.
///
/// The syntax is of the form `kernel::dma_read!(dma, proj)` where `dma` is an expression evaluating
/// to a [`Coherent`] and `proj` is a [projection specification](kernel::ptr::project!).
///
/// # Examples
///
/// ```
/// use kernel::device::Device;
/// use kernel::dma::{attrs::*, Coherent};
///
/// struct MyStruct { field: u32, }
///
/// // SAFETY: All bit patterns are acceptable values for `MyStruct`.
/// unsafe impl kernel::transmute::FromBytes for MyStruct{};
/// // SAFETY: Instances of `MyStruct` have no uninitialized portions.
/// unsafe impl kernel::transmute::AsBytes for MyStruct{};
///
/// # fn test(alloc: &kernel::dma::Coherent<[MyStruct]>) -> Result {
/// let whole = kernel::dma_read!(alloc, [try: 2]);
/// let field = kernel::dma_read!(alloc, [panic: 1].field);
/// # Ok::<(), Error>(()) }
/// ```
#[macro_export]
macro_rules! dma_read {
($dma:expr, $($proj:tt)*) => {{
let dma = &$dma;
let ptr = $crate::ptr::project!(
$crate::dma::Coherent::as_ptr(dma), $($proj)*
);
// SAFETY: The pointer created by the projection is within the DMA region.
unsafe { $crate::dma::Coherent::field_read(dma, ptr) }
}};
/// This is same as [`SysMem`] but with additional information that allows handing out a DMA
/// address.
pub struct CoherentView<'a, T: ?Sized> {
cpu_addr: SysMem<'a, T>,
dma_addr: DmaAddress,
}
/// Writes to a field of an item from an allocated region of structs.
///
/// The syntax is of the form `kernel::dma_write!(dma, proj, val)` where `dma` is an expression
/// evaluating to a [`Coherent`], `proj` is a
/// [projection specification](kernel::ptr::project!), and `val` is the value to be written to the
/// projected location.
///
/// # Examples
///
/// ```
/// use kernel::device::Device;
/// use kernel::dma::{attrs::*, Coherent};
///
/// struct MyStruct { member: u32, }
///
/// // SAFETY: All bit patterns are acceptable values for `MyStruct`.
/// unsafe impl kernel::transmute::FromBytes for MyStruct{};
/// // SAFETY: Instances of `MyStruct` have no uninitialized portions.
/// unsafe impl kernel::transmute::AsBytes for MyStruct{};
///
/// # fn test(alloc: &kernel::dma::Coherent<[MyStruct]>) -> Result {
/// kernel::dma_write!(alloc, [try: 2].member, 0xf);
/// kernel::dma_write!(alloc, [panic: 1], MyStruct { member: 0xf });
/// # Ok::<(), Error>(()) }
/// ```
#[macro_export]
macro_rules! dma_write {
(@parse [$dma:expr] [$($proj:tt)*] [, $val:expr]) => {{
let dma = &$dma;
let ptr = $crate::ptr::project!(
mut $crate::dma::Coherent::as_mut_ptr(dma), $($proj)*
);
let val = $val;
// SAFETY: The pointer created by the projection is within the DMA region.
unsafe { $crate::dma::Coherent::field_write(dma, ptr, val) }
}};
(@parse [$dma:expr] [$($proj:tt)*] [.$field:tt $($rest:tt)*]) => {
$crate::dma_write!(@parse [$dma] [$($proj)* .$field] [$($rest)*])
};
(@parse [$dma:expr] [$($proj:tt)*] [[$flavor:ident: $index:expr] $($rest:tt)*]) => {
$crate::dma_write!(@parse [$dma] [$($proj)* [$flavor: $index]] [$($rest)*])
};
($dma:expr, $($rest:tt)*) => {
$crate::dma_write!(@parse [$dma] [] [$($rest)*])
};
impl<T: ?Sized> Copy for CoherentView<'_, T> {}
impl<T: ?Sized> Clone for CoherentView<'_, T> {
#[inline]
fn clone(&self) -> Self {
*self
}
}
impl<'a, T: ?Sized> CoherentView<'a, T> {
/// Erase the DMA address information and obtain a [`SysMem`] view of the same memory region.
#[inline]
pub fn as_sys_mem(self) -> SysMem<'a, T> {
self.cpu_addr
}
/// Returns the DMA address which may be given to the device as base of the region.
#[inline]
pub fn dma_address(self) -> DmaAddress {
self.dma_addr
}
/// Returns a reference to the data in the region.
///
/// # Safety
///
/// * Callers must ensure that the device does not read/write to/from memory while the returned
/// reference is live.
/// * Callers must ensure that this call does not race with a write (including call to `as_mut`)
/// to the same region while the returned reference is live.
#[inline]
pub unsafe fn as_ref(self) -> &'a T {
// SAFETY: pointer is aligned and valid per type invariant. Aliasing rule is satisfied per
// safety requirement.
unsafe { &*self.cpu_addr.as_ptr() }
}
/// Returns a mutable reference to the data in the region.
///
/// # Safety
///
/// * Callers must ensure that the device does not read/write to/from memory while the returned
/// reference is live.
/// * Callers must ensure that this call does not race with a read (including call to `as_ref`)
/// or write (including call to `as_mut`) to the same region while the returned reference is
/// live.
#[inline]
pub unsafe fn as_mut(self) -> &'a mut T {
// SAFETY: pointer is aligned and valid per type invariant. Aliasing rule is satisfied per
// safety requirement.
unsafe { &mut *self.cpu_addr.as_ptr() }
}
}
/// `IoBackend` implementation for `Coherent`.
pub struct CoherentIoBackend;
impl IoBackend for CoherentIoBackend {
type View<'a, T: ?Sized + KnownSize> = CoherentView<'a, T>;
#[inline]
fn as_ptr<'a, T: ?Sized + KnownSize>(view: Self::View<'a, T>) -> *mut T {
SysMemBackend::as_ptr(view.cpu_addr)
}
#[inline]
unsafe fn project_view<'a, T: ?Sized + KnownSize, U: ?Sized + KnownSize>(
view: Self::View<'a, T>,
ptr: *mut U,
) -> Self::View<'a, U> {
let offset = ptr.addr() - view.cpu_addr.as_ptr().addr();
// CAST: The offset DMA address can never overflow.
let dma_addr = view.dma_addr + offset as DmaAddress;
CoherentView {
dma_addr,
// SAFETY: Per safety requirement.
cpu_addr: unsafe { SysMemBackend::project_view(view.cpu_addr, ptr) },
}
}
}
impl<T> IoCapable<T> for CoherentIoBackend
where
SysMemBackend: IoCapable<T>,
{
#[inline]
fn io_read<'a>(view: Self::View<'a, T>) -> T {
SysMemBackend::io_read(view.cpu_addr)
}
#[inline]
fn io_write<'a>(view: Self::View<'a, T>, value: T) {
SysMemBackend::io_write(view.cpu_addr, value)
}
}
impl IoCopyable for CoherentIoBackend {
#[inline]
unsafe fn copy_from_io(view: Self::View<'_, [u8]>, buffer: *mut u8) {
// SAFETY: Per safety requirement.
unsafe { SysMemBackend::copy_from_io(view.cpu_addr, buffer) }
}
#[inline]
unsafe fn copy_to_io(view: Self::View<'_, [u8]>, buffer: *const u8) {
// SAFETY: Per safety requirement.
unsafe { SysMemBackend::copy_to_io(view.cpu_addr, buffer) }
}
#[inline]
fn copy_read<T: zerocopy::FromBytes>(view: Self::View<'_, T>) -> T {
SysMemBackend::copy_read(view.cpu_addr)
}
#[inline]
fn copy_write<T: zerocopy::IntoBytes>(view: Self::View<'_, T>, value: T) {
SysMemBackend::copy_write(view.cpu_addr, value)
}
}
impl<'a, T: ?Sized + KnownSize> IoBase<'a> for CoherentView<'a, T> {
type Backend = CoherentIoBackend;
type Target = T;
#[inline]
fn as_view(self) -> CoherentView<'a, Self::Target> {
self
}
}
impl<'a, T: ?Sized + KnownSize> IoBase<'a> for &'a Coherent<T> {
type Backend = CoherentIoBackend;
type Target = T;
#[inline]
fn as_view(self) -> CoherentView<'a, Self::Target> {
CoherentView {
// SAFETY: `cpu_addr` is valid and aligned kernel accessible memory.
cpu_addr: unsafe { SysMem::new(self.cpu_addr.as_ptr()) },
dma_addr: self.dma_addr,
}
}
}

View File

@ -32,6 +32,7 @@
};
use core::{
alloc::Layout,
cell::UnsafeCell,
marker::PhantomData,
mem,
ops::Deref,
@ -74,66 +75,59 @@ macro_rules! drm_legacy_fields {
/// A trait implemented by all possible contexts a [`Device`] can be used in.
///
/// Setting up a new [`Device`] is a multi-stage process. Each step of the process that a user
/// interacts with in Rust has a respective [`DeviceContext`] typestate. For example,
/// `Device<T, Registered>` would be a [`Device`] that reached the [`Registered`] [`DeviceContext`].
/// A [`Device`] can be in one of the following contexts:
///
/// Each stage of this process is described below:
/// - [`Normal`]: The general-purpose, reference-counted context. A [`Device`] in this context may
/// or may not be registered with userspace.
/// - [`Ioctl`]: The device has been registered with userspace at some point; used in ioctl
/// dispatch context.
/// - [`Registered`]: The device is currently registered with userspace and the parent bus device
/// is bound.
///
/// ```text
/// 1 2 3
/// +--------------+ +------------------+ +-----------------------+
/// |Device created| → |Device initialized| → |Registered w/ userspace|
/// +--------------+ +------------------+ +-----------------------+
/// (Uninit) (Registered)
/// ```
///
/// 1. The [`Device`] is in the [`Uninit`] context and is not guaranteed to be initialized or
/// registered with userspace. Only a limited subset of DRM core functionality is available.
/// 2. The [`Device`] is guaranteed to be fully initialized, but is not guaranteed to be registered
/// with userspace. All DRM core functionality which doesn't interact with userspace is
/// available. We currently don't have a context for representing this.
/// 3. The [`Device`] is guaranteed to be fully initialized, and is guaranteed to have been
/// registered with userspace at some point - thus putting it in the [`Registered`] context.
///
/// An important caveat of [`DeviceContext`] which must be kept in mind: when used as a typestate
/// for a reference type, it can only guarantee that a [`Device`] reached a particular stage in the
/// initialization process _at the time the reference was taken_. No guarantee is made in regards to
/// what stage of the process the [`Device`] is currently in. This means for instance that a
/// `&Device<T, Uninit>` may actually be registered with userspace, it just wasn't known to be
/// registered at the time the reference was taken.
pub trait DeviceContext: Sealed + Send + Sync {}
/// Both `Device<T, Ioctl>` and `Device<T, Registered>` dereference to `Device<T>` ([`Normal`]),
/// so any method available on a [`Normal`] device is also available in the other contexts.
pub trait DeviceContext: Sealed + Send + Sync + 'static {}
/// The [`DeviceContext`] of a [`Device`] that was registered with userspace at some point.
/// The general-purpose, reference-counted [`DeviceContext`].
///
/// This represents a [`Device`] which is guaranteed to have been registered with userspace at
/// some point in time. Such a DRM device is guaranteed to have been fully-initialized.
/// A [`Device`] in this context may or may not be registered with userspace. This context is used
/// for reference-counted device handles and during device setup via [`UnregisteredDevice`].
///
/// Note: A device in this context is not guaranteed to remain registered with userspace for its
/// entire lifetime, as this is impossible to guarantee at compile-time.
/// [`AlwaysRefCounted`] is only implemented for `Device<T, Normal>`, making this the required
/// context for [`ARef`]-based device handles.
pub struct Normal;
impl Sealed for Normal {}
impl DeviceContext for Normal {}
/// The [`DeviceContext`] of a [`Device`] that is currently registered with userspace.
///
/// A [`Device`] in this context is guaranteed to be registered and its parent bus device is
/// guaranteed to be bound. This is enforced at runtime by [`RegistrationGuard`], which holds a
/// `drm_dev_enter()` / `drm_dev_exit()` SRCU critical section.
///
/// # Invariants
///
/// A [`Device`] in this [`DeviceContext`] is guaranteed to have been registered with userspace
/// at some point in time.
/// The parent bus device is bound for the duration of any reference to a `Device<T, Registered>`.
pub struct Registered;
impl Sealed for Registered {}
impl DeviceContext for Registered {}
/// The [`DeviceContext`] of a [`Device`] that may be unregistered and partly uninitialized.
/// The [`DeviceContext`] of a [`Device`] that has been registered with userspace previously.
///
/// A [`Device`] in this context is only guaranteed to be partly initialized, and may or may not
/// be registered with userspace. Thus operations which depend on the [`Device`] being fully
/// initialized, or which depend on the [`Device`] being registered with userspace are not
/// available through this [`DeviceContext`].
/// A [`Device`] in this context has been registered at some point, but may be concurrently
/// unregistering or already unregistered. `drm_dev_enter()` can guard against this, ensuring the
/// device remains registered for the duration of the critical section.
///
/// A [`Device`] in this context can be used to create a
/// [`Registration`](drm::driver::Registration).
pub struct Uninit;
/// # Invariants
///
/// A [`Device`] in this context has been registered with userspace via `drm_dev_register()` at
/// some point.
pub struct Ioctl;
impl Sealed for Uninit {}
impl DeviceContext for Uninit {}
impl Sealed for Ioctl {}
impl DeviceContext for Ioctl {}
/// A [`Device`] which is known at compile-time to be unregistered with userspace.
///
@ -147,10 +141,10 @@ impl DeviceContext for Uninit {}
///
/// The device in `self.0` is guaranteed to be a newly created [`Device`] that has not yet been
/// registered with userspace until this type is dropped.
pub struct UnregisteredDevice<T: drm::Driver>(ARef<Device<T, Uninit>>, NotThreadSafe);
pub struct UnregisteredDevice<T: drm::Driver>(ARef<Device<T, Normal>>, NotThreadSafe);
impl<T: drm::Driver> Deref for UnregisteredDevice<T> {
type Target = Device<T, Uninit>;
type Target = Device<T, Normal>;
fn deref(&self) -> &Self::Target {
&self.0
@ -178,15 +172,13 @@ const fn compute_features() -> u32 {
master_drop: None,
debugfs_init: None,
// Ignore the Uninit DeviceContext below. It is only provided because it is required by the
// compiler, and it is not actually used by these functions.
gem_create_object: T::Object::<Uninit>::ALLOC_OPS.gem_create_object,
prime_handle_to_fd: T::Object::<Uninit>::ALLOC_OPS.prime_handle_to_fd,
prime_fd_to_handle: T::Object::<Uninit>::ALLOC_OPS.prime_fd_to_handle,
gem_prime_import: T::Object::<Uninit>::ALLOC_OPS.gem_prime_import,
gem_prime_import_sg_table: T::Object::<Uninit>::ALLOC_OPS.gem_prime_import_sg_table,
dumb_create: T::Object::<Uninit>::ALLOC_OPS.dumb_create,
dumb_map_offset: T::Object::<Uninit>::ALLOC_OPS.dumb_map_offset,
gem_create_object: T::Object::ALLOC_OPS.gem_create_object,
prime_handle_to_fd: T::Object::ALLOC_OPS.prime_handle_to_fd,
prime_fd_to_handle: T::Object::ALLOC_OPS.prime_fd_to_handle,
gem_prime_import: T::Object::ALLOC_OPS.gem_prime_import,
gem_prime_import_sg_table: T::Object::ALLOC_OPS.gem_prime_import_sg_table,
dumb_create: T::Object::ALLOC_OPS.dumb_create,
dumb_map_offset: T::Object::ALLOC_OPS.dumb_map_offset,
show_fdinfo: None,
fbdev_probe: None,
@ -208,10 +200,13 @@ const fn compute_features() -> u32 {
/// Create a new `UnregisteredDevice` for a `drm::Driver`.
///
/// This can be used to create a [`Registration`](kernel::drm::Registration).
pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<Self> {
pub fn new(
dev: &T::ParentDevice<device::Bound>,
data: impl PinInit<T::Data, Error>,
) -> Result<Self> {
// `__drm_dev_alloc` uses `kmalloc()` to allocate memory, hence ensure a `kmalloc()`
// compatible `Layout`.
let layout = Kmalloc::aligned_layout(Layout::new::<Device<T, Uninit>>());
let layout = Kmalloc::aligned_layout(Layout::new::<Device<T, Normal>>());
// Use a temporary vtable without a `release` callback until `data` is initialized, so
// init failure can release the DRM device without dropping uninitialized fields.
@ -223,12 +218,12 @@ pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<S
// SAFETY:
// - `alloc_vtable` reference remains valid until no longer used,
// - `dev` is valid by its type invarants,
let raw_drm: *mut Device<T, Uninit> = unsafe {
let raw_drm: *mut Device<T, Normal> = unsafe {
bindings::__drm_dev_alloc(
dev.as_raw(),
dev.as_ref().as_raw(),
&alloc_vtable,
layout.size(),
mem::offset_of!(Device<T, Uninit>, dev),
mem::offset_of!(Device<T, Normal>, dev),
)
}
.cast();
@ -253,6 +248,9 @@ pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<S
// SAFETY: `drm_dev` is still private to this function.
unsafe { (*drm_dev).driver = const { &Self::VTABLE } };
// SAFETY: `raw_drm` is valid; no concurrent access before registration.
unsafe { (*raw_drm.as_ptr()).registration_data = UnsafeCell::new(NonNull::dangling()) };
// SAFETY: The reference count is one, and now we take ownership of that reference as a
// `drm::Device`.
// INVARIANT: We just created the device above, but have yet to call `drm_dev_register`.
@ -264,16 +262,8 @@ pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<S
/// A typed DRM device with a specific [`drm::Driver`] implementation and [`DeviceContext`].
///
/// Since DRM devices can be used before being fully initialized and registered with userspace, `C`
/// represents the furthest [`DeviceContext`] we can guarantee that this [`Device`] has reached.
///
/// Keep in mind: this means that an unregistered device can still have the registration state
/// [`Registered`] as long as it was registered with userspace once in the past, and that the
/// behavior of such a device is still well-defined. Additionally, a device with the registration
/// state [`Uninit`] simply does not have a guaranteed registration state at compile time, and could
/// be either registered or unregistered. Since there is no way to guarantee a long-lived reference
/// to an unregistered device would remain unregistered, we do not provide a [`DeviceContext`] for
/// this.
/// A device in the [`Registered`] context is currently registered with userspace and its parent
/// bus device is bound. The [`Normal`] context is the general-purpose, reference-counted context.
///
/// # Invariants
///
@ -281,9 +271,10 @@ pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<S
/// * The data layout of `Self` remains the same across all implementations of `C`.
/// * Any invariants for `C` also apply.
#[repr(C)]
pub struct Device<T: drm::Driver, C: DeviceContext = Registered> {
pub struct Device<T: drm::Driver, C: DeviceContext = Normal> {
dev: Opaque<bindings::drm_device>,
data: T::Data,
pub(super) registration_data: UnsafeCell<NonNull<T::RegistrationData<'static>>>,
_ctx: PhantomData<C>,
}
@ -352,7 +343,111 @@ pub(crate) unsafe fn assume_ctx<NewCtx: DeviceContext>(&self) -> &Device<T, NewC
}
}
impl<T: drm::Driver, C: DeviceContext> Deref for Device<T, C> {
impl<T: drm::Driver> Device<T, Ioctl> {
/// Guard against the parent bus device being unbound.
///
/// Returns a [`RegistrationGuard`] if the device has not been unplugged, [`None`] otherwise.
///
/// While [`RegistrationGuard`] is held the parent device is guaranteed to be bound.
#[must_use]
pub fn registration_guard(&self) -> Option<RegistrationGuard<'_, T>> {
let mut idx: i32 = 0;
// SAFETY: `self.as_raw()` is a valid pointer to a `struct drm_device`.
if unsafe { bindings::drm_dev_enter(self.as_raw(), &mut idx) } {
// INVARIANT:
// - `idx` is the SRCU index from the successful `drm_dev_enter()` above.
// - The parent bus device is bound: `drm_dev_enter()` succeeded, meaning
// `drm_dev_unplug()` has not completed; since it is only called from
// `Registration::drop()` during parent unbind, the parent is still bound.
Some(RegistrationGuard {
// SAFETY: See INVARIANT above; the `Registered` context invariant holds.
dev: unsafe { self.assume_ctx() },
idx,
_not_send: NotThreadSafe,
})
} else {
None
}
}
}
/// A guard proving the DRM device is registered and the parent bus device is bound.
///
/// The guard dereferences to [`Device<T, Registered>`], providing access to the DRM device with
/// the guarantee that the parent bus device is bound for the entire duration of the critical
/// section.
///
/// Internally this is backed by a `drm_dev_enter()` / `drm_dev_exit()` SRCU critical section.
///
/// # Invariants
///
/// - `idx` is the SRCU read lock index returned by a successful `drm_dev_enter()` call.
/// - The parent bus device of `dev` is bound for the lifetime of this guard.
#[must_use]
pub struct RegistrationGuard<'a, T: drm::Driver> {
dev: &'a Device<T, Registered>,
idx: i32,
_not_send: NotThreadSafe,
}
impl<T: drm::Driver> Device<T, Registered> {
/// Returns a reference to the registration data with lifetime shortened from `'static`.
///
/// # Safety
///
/// The returned reference must not be exposed to code that can choose a concrete lifetime for
/// it, as that would be unsound for types that are invariant over their lifetime parameter
/// (e.g. it must be passed through an HRTB-bounded closure).
#[inline]
unsafe fn registration_data_unchecked(&self) -> &T::RegistrationData<'_> {
// SAFETY:
// - `Registered` guarantees the parent bus device is bound, hence the pointer is valid.
// - The pointer cast from `Of<'static>` to `Of<'_>` is layout-compatible since lifetimes
// are erased at runtime.
// - Caller guarantees the reference is only used behind an HRTB, making the lifetime
// shortening sound regardless of variance.
unsafe { (*self.registration_data.get()).cast::<_>().as_ref() }
}
/// Access the registration data through a closure, with the lifetime tied to the closure
/// scope.
///
/// The data is owned by [`Registration`](drm::Registration) and is guaranteed to remain valid
/// as long as the device is registered, since [`Registration`](drm::Registration)'s `drop`
/// calls `drm_dev_unplug()` which waits for all `drm_dev_enter()` critical sections to
/// complete.
#[inline]
pub fn registration_data_with<R, F>(&self, f: F) -> R
where
F: for<'a> FnOnce(&'a T::RegistrationData<'a>) -> R,
{
// SAFETY: `Registered` guarantees the device is registered and the parent bus device is
// bound. The closure's HRTB `for<'a>` prevents the caller from smuggling in references
// with a concrete short lifetime, satisfying the lifetime requirement of
// `registration_data_unchecked`.
f(unsafe { self.registration_data_unchecked() })
}
}
impl<T: drm::Driver> Deref for RegistrationGuard<'_, T> {
type Target = Device<T, Registered>;
#[inline]
fn deref(&self) -> &Self::Target {
self.dev
}
}
impl<T: drm::Driver> Drop for RegistrationGuard<'_, T> {
#[inline]
fn drop(&mut self) {
// SAFETY: `self.idx` was returned by a successful `drm_dev_enter()` call, as guaranteed
// by the type invariants of `RegistrationGuard`.
unsafe { bindings::drm_dev_exit(self.idx) };
}
}
impl<T: drm::Driver> Deref for Device<T> {
type Target = T::Data;
fn deref(&self) -> &Self::Target {
@ -360,9 +455,31 @@ fn deref(&self) -> &Self::Target {
}
}
impl<T: drm::Driver> Deref for Device<T, Registered> {
type Target = Device<T>;
#[inline]
fn deref(&self) -> &Self::Target {
// SAFETY: The caller holds a `Device<T, Registered>`, which guarantees all invariants
// of the weaker `Normal` context.
unsafe { self.assume_ctx() }
}
}
impl<T: drm::Driver> Deref for Device<T, Ioctl> {
type Target = Device<T>;
#[inline]
fn deref(&self) -> &Self::Target {
// SAFETY: The caller holds a `Device<T, Ioctl>`, which guarantees all invariants
// of the weaker `Normal` context.
unsafe { self.assume_ctx() }
}
}
// SAFETY: DRM device objects are always reference counted and the get/put functions
// satisfy the requirements.
unsafe impl<T: drm::Driver, C: DeviceContext> AlwaysRefCounted for Device<T, C> {
unsafe impl<T: drm::Driver> AlwaysRefCounted for Device<T> {
fn inc_ref(&self) {
// SAFETY: The existence of a shared reference guarantees that the refcount is non-zero.
unsafe { bindings::drm_dev_get(self.as_raw()) };
@ -377,11 +494,29 @@ unsafe fn dec_ref(obj: NonNull<Self>) {
}
}
impl<T: drm::Driver, C: DeviceContext> AsRef<device::Device> for Device<T, C> {
fn as_ref(&self) -> &device::Device {
impl<T: drm::Driver> AsRef<T::ParentDevice<device::Normal>> for Device<T> {
fn as_ref(&self) -> &T::ParentDevice<device::Normal> {
// SAFETY: `bindings::drm_device::dev` is valid as long as the DRM device itself is valid,
// which is guaranteed by the type invariant.
unsafe { device::Device::from_raw((*self.as_raw()).dev) }
let dev = unsafe { device::Device::from_raw((*self.as_raw()).dev) };
// SAFETY: The DRM device was constructed in `UnregisteredDevice::new()` with a parent
// device of type `T::ParentDevice`, hence `dev` is contained in a `T::ParentDevice`.
unsafe { device::AsBusDevice::from_device(dev) }
}
}
impl<T: drm::Driver> AsRef<T::ParentDevice<device::Bound>> for Device<T, Registered> {
#[inline]
fn as_ref(&self) -> &T::ParentDevice<device::Bound> {
let dev = (**self).as_ref().as_ref();
// SAFETY: A `Device<T, Registered>` guarantees that the parent device is bound.
let dev = unsafe { dev.as_bound() };
// SAFETY: The DRM device was constructed in `UnregisteredDevice::new()` with a parent
// device of type `T::ParentDevice`, hence `dev` is contained in a `T::ParentDevice`.
unsafe { device::AsBusDevice::from_device(dev) }
}
}
@ -392,12 +527,10 @@ unsafe impl<T: drm::Driver, C: DeviceContext> Send for Device<T, C> {}
// by the synchronization in `struct drm_device`.
unsafe impl<T: drm::Driver, C: DeviceContext> Sync for Device<T, C> {}
impl<T, C, const ID: u64> WorkItem<ID> for Device<T, C>
impl<T: drm::Driver, const ID: u64> WorkItem<ID> for Device<T>
where
T: drm::Driver,
T::Data: WorkItem<ID, Pointer = ARef<Self>>,
T::Data: HasWork<Self, ID>,
C: DeviceContext,
{
type Pointer = ARef<Self>;

View File

@ -7,16 +7,12 @@
use crate::{
bindings,
device,
devres,
drm,
error::to_result,
prelude::*,
sync::aref::ARef, //
};
use core::{
mem,
ptr::NonNull, //
};
use core::ptr::NonNull;
/// Driver use the GEM memory manager. This should be set for all modern drivers.
pub(crate) const FEAT_GEM: u32 = bindings::drm_driver_feature_DRIVER_GEM;
@ -110,12 +106,23 @@ pub trait Driver {
/// Context data associated with the DRM driver
type Data: Sync + Send;
/// Data owned by the [`Registration`] and accessible within a
/// [`RegistrationGuard`](drm::RegistrationGuard) critical section via
/// [`Device::registration_data_with()`](drm::Device::registration_data_with).
///
/// The lifetime parameter is tied to the [`Registration`] scope, which is enclosed in the
/// parent bus device binding scope but may be shorter.
type RegistrationData<'a>: Send + Sync + 'a;
/// The type used to manage memory for this driver.
type Object<Ctx: drm::DeviceContext>: AllocImpl;
type Object: AllocImpl;
/// The type used to represent a DRM File (client)
type File: drm::file::DriverFile;
/// The bus device type of the parent device that the DRM device is associated with.
type ParentDevice<Ctx: device::DeviceContext>: device::AsBusDevice<Ctx>;
/// Driver metadata
const INFO: DriverInfo;
@ -125,7 +132,7 @@ pub trait Driver {
/// Sets the `DRIVER_RENDER` feature for this driver.
///
/// When enabled, the driver exposes `/dev/dri/renderDXX` render nodes to
/// userspace. The render node is an alternate low-priviledge way to access
/// userspace. The render node is an alternate low-privilege way to access
/// the driver, which is enforced on a per-ioctl level. Userspace processes
/// that open the render node can only invoke ioctls explicitly listed as
/// usable from the render node (i.e. marked DRM_RENDER_ALLOW), whereas
@ -136,68 +143,84 @@ pub trait Driver {
/// The registration type of a `drm::Device`.
///
/// Once the `Registration` structure is dropped, the device is unregistered.
pub struct Registration<T: Driver>(ARef<drm::Device<T>>);
pub struct Registration<'a, T: Driver> {
drm: ARef<drm::Device<T>>,
_reg_data: Pin<KBox<T::RegistrationData<'a>>>,
}
impl<T: Driver> Registration<T> {
fn new(drm: drm::UnregisteredDevice<T>, flags: usize) -> Result<Self> {
// SAFETY: `drm.as_raw()` is valid by the invariants of `drm::Device`.
to_result(unsafe { bindings::drm_dev_register(drm.as_raw(), flags) })?;
// SAFETY: We just called `drm_dev_register` above
let new = NonNull::from(unsafe { drm.assume_ctx() });
// Leak the ARef from UnregisteredDevice in preparation for transferring its ownership.
mem::forget(drm);
// SAFETY: `drm`'s `Drop` constructor was never called, ensuring that there remains at least
// one reference to the device - which we take ownership over here.
let new = unsafe { ARef::from_raw(new) };
Ok(Self(new))
}
/// Registers a new [`UnregisteredDevice`](drm::UnregisteredDevice) with userspace.
impl<'a, T: Driver> Registration<'a, T> {
/// Register a new [`UnregisteredDevice`](drm::UnregisteredDevice) with userspace.
///
/// Ownership of the [`Registration`] object is passed to [`devres::register`].
pub fn new_foreign_owned<'a>(
drm: drm::UnregisteredDevice<T>,
/// # Safety
///
/// The caller must not `mem::forget()` the returned [`Registration`] or otherwise prevent its
/// [`Drop`] implementation from running, since the registration data may contain borrowed
/// references that become invalid after `'a` ends.
pub unsafe fn new<E>(
dev: &'a device::Device<device::Bound>,
drm: drm::UnregisteredDevice<T>,
reg_data: impl PinInit<T::RegistrationData<'a>, E>,
flags: usize,
) -> Result<&'a drm::Device<T>>
) -> Result<Self>
where
T: 'static,
Error: From<E>,
{
if drm.as_ref().as_raw() != dev.as_raw() {
let parent = drm.as_ref();
if parent.as_ref().as_raw() != dev.as_raw() {
return Err(EINVAL);
}
let reg = Registration::<T>::new(drm, flags)?;
let drm = NonNull::from(reg.device());
let reg_data: Pin<KBox<T::RegistrationData<'a>>> = KBox::pin_init(reg_data, GFP_KERNEL)?;
devres::register(dev, reg, GFP_KERNEL)?;
// Store the registration data pointer in the device before registration, so that it is
// visible once ioctls can be called.
let ptr: NonNull<T::RegistrationData<'static>> =
NonNull::from(Pin::get_ref(reg_data.as_ref())).cast();
// SAFETY: Since `reg` was passed to devres::register(), the device now owns the lifetime
// of the DRM registration - ensuring that this references lives for at least as long as 'a.
Ok(unsafe { drm.as_ref() })
// SAFETY: No concurrent access; the device is not yet registered.
unsafe { *drm.registration_data.get() = ptr };
// SAFETY: `drm` is a valid, initialized but not yet registered DRM device.
let ret = unsafe { bindings::drm_dev_register(drm.as_raw(), flags) };
if let Err(e) = to_result(ret) {
// SAFETY: `drm_dev_register()` synchronizes SRCU on failure, so no concurrent
// access to `registration_data` is possible at this point.
unsafe { *drm.registration_data.get() = NonNull::dangling() };
return Err(e);
}
Ok(Self {
drm: (&*drm).into(),
_reg_data: reg_data,
})
}
/// Returns a reference to the `Device` instance for this registration.
pub fn device(&self) -> &drm::Device<T> {
&self.0
&self.drm
}
}
// SAFETY: `Registration` doesn't offer any methods or access to fields when shared between
// threads, hence it's safe to share it.
unsafe impl<T: Driver> Sync for Registration<T> {}
unsafe impl<T: Driver> Sync for Registration<'_, T> {}
// SAFETY: Registration with and unregistration from the DRM subsystem can happen from any thread.
unsafe impl<T: Driver> Send for Registration<T> {}
unsafe impl<T: Driver> Send for Registration<'_, T> {}
impl<T: Driver> Drop for Registration<T> {
impl<T: Driver> Drop for Registration<'_, T> {
fn drop(&mut self) {
// Use `drm_dev_unplug` rather than `drm_dev_unregister` to ensure that existing
// `drm_dev_enter()` critical sections complete before unregistration proceeds. This
// is required for the safety of `RegistrationGuard`, which relies on the SRCU barrier in
// `drm_dev_unplug()` to guarantee that the parent device is still bound within the
// critical section.
//
// SAFETY: Safe by the invariant of `ARef<drm::Device<T>>`. The existence of this
// `Registration` also guarantees the this `drm::Device` is actually registered.
unsafe { bindings::drm_dev_unregister(self.0.as_raw()) };
// `Registration` also guarantees that this `drm::Device` is actually registered.
unsafe { bindings::drm_dev_unplug(self.drm.as_raw()) };
// After drm_dev_unplug(), the SRCU barrier guarantees that all RegistrationGuard critical
// sections have completed, so no one holds a reference to reg_data anymore.
// reg_data is dropped here automatically.
}
}

View File

@ -10,7 +10,7 @@
self,
device::{
DeviceContext,
Registered, //
Normal, //
},
driver::{
AllocImpl,
@ -81,11 +81,10 @@ unsafe fn dec_ref(obj: core::ptr::NonNull<Self>) {
/// A type alias for retrieving the current [`AllocImpl`] for a given [`DriverObject`].
///
/// [`Driver`]: drm::Driver
pub type DriverAllocImpl<T, Ctx = Registered> =
<<T as DriverObject>::Driver as drm::Driver>::Object<Ctx>;
pub type DriverAllocImpl<T> = <<T as DriverObject>::Driver as drm::Driver>::Object;
/// GEM object functions, which must be implemented by drivers.
pub trait DriverObject: Sync + Send + Sized {
pub trait DriverObject: Sync + Send + Sized + 'static {
/// Parent `Driver` for this object.
type Driver: drm::Driver;
@ -93,8 +92,8 @@ pub trait DriverObject: Sync + Send + Sized {
type Args;
/// Create a new driver data object for a GEM object of a given size.
fn new<Ctx: DeviceContext>(
dev: &drm::Device<Self::Driver, Ctx>,
fn new(
dev: &drm::Device<Self::Driver>,
size: usize,
args: Self::Args,
) -> impl PinInit<Self, Error>;
@ -109,7 +108,7 @@ fn close(_obj: &DriverAllocImpl<Self>, _file: &DriverFile<Self>) {}
}
/// Trait that represents a GEM object subtype
pub trait IntoGEMObject: Sized + super::private::Sealed + AlwaysRefCounted {
pub trait IntoGEMObject: Sized + super::private::Sealed {
/// Returns a reference to the raw `drm_gem_object` structure, which must be valid as long as
/// this owning object is valid.
fn as_raw(&self) -> *mut bindings::drm_gem_object;
@ -118,7 +117,8 @@ pub trait IntoGEMObject: Sized + super::private::Sealed + AlwaysRefCounted {
///
/// # Safety
///
/// - `self_ptr` must be a valid pointer to `Self`.
/// - `self_ptr` must be a valid pointer to the `struct drm_gem_object` embedded in a
/// valid instance of `Self`.
/// - The caller promises that holding the immutable reference returned by this function does
/// not violate rust's data aliasing rules and remains valid throughout the lifetime of `'a`.
unsafe fn from_raw<'a>(self_ptr: *mut bindings::drm_gem_object) -> &'a Self;
@ -183,7 +183,7 @@ fn size(&self) -> usize {
fn create_handle<D, F>(&self, file: &drm::File<F>) -> Result<u32>
where
Self: AllocImpl<Driver = D>,
D: drm::Driver<Object<Registered> = Self, File = F>,
D: drm::Driver<Object = Self, File = F>,
F: drm::file::DriverFile<Driver = D>,
{
let mut handle: u32 = 0;
@ -197,8 +197,8 @@ fn create_handle<D, F>(&self, file: &drm::File<F>) -> Result<u32>
/// Looks up an object by its handle for a given `File`.
fn lookup_handle<D, F>(file: &drm::File<F>, handle: u32) -> Result<ARef<Self>>
where
Self: AllocImpl<Driver = D>,
D: drm::Driver<Object<Registered> = Self, File = F>,
Self: AllocImpl<Driver = D> + AlwaysRefCounted,
D: drm::Driver<Object = Self, File = F>,
F: drm::file::DriverFile<Driver = D>,
{
// SAFETY: The arguments are all valid per the type invariants.
@ -254,7 +254,7 @@ impl<T: IntoGEMObject> BaseObjectPrivate for T {}
/// * Any type invariants of `Ctx` apply to the parent DRM device for this GEM object.
#[repr(C)]
#[pin_data]
pub struct Object<T: DriverObject + Send + Sync, Ctx: DeviceContext = Registered> {
pub struct Object<T: DriverObject + Send + Sync, Ctx: DeviceContext = Normal> {
obj: Opaque<bindings::drm_gem_object>,
#[pin]
data: T,
@ -280,48 +280,6 @@ impl<T: DriverObject, Ctx: DeviceContext> Object<T, Ctx> {
rss: None,
};
/// Create a new GEM object.
pub fn new(
dev: &drm::Device<T::Driver, Ctx>,
size: usize,
args: T::Args,
) -> Result<ARef<Self>> {
let obj: Pin<KBox<Self>> = KBox::pin_init(
try_pin_init!(Self {
obj: Opaque::new(bindings::drm_gem_object::default()),
data <- T::new(dev, size, args),
_ctx: PhantomData,
}),
GFP_KERNEL,
)?;
// SAFETY: `obj.as_raw()` is guaranteed to be valid by the initialization above.
unsafe { (*obj.as_raw()).funcs = &Self::OBJECT_FUNCS };
// INVARIANT: `dev` and the GEM object are in the same state at the moment, and upgrading
// the typestate in `dev` will not carry over to the GEM object.
if let Err(err) =
// SAFETY: The arguments are all valid per the type invariants.
to_result(unsafe {
bindings::drm_gem_object_init(dev.as_raw(), obj.obj.get(), size)
})
{
// SAFETY: `drm_gem_object_init()` initializes the private GEM object state before
// failing, so `drm_gem_private_object_fini()` is the matching cleanup.
unsafe { bindings::drm_gem_private_object_fini(obj.obj.get()) };
return Err(err);
}
// SAFETY: We will never move out of `Self` as `ARef<Self>` is always treated as pinned.
let ptr = KBox::into_raw(unsafe { Pin::into_inner_unchecked(obj) });
// SAFETY: `ptr` comes from `KBox::into_raw` and hence can't be NULL.
let ptr = unsafe { NonNull::new_unchecked(ptr) };
// SAFETY: We take over the initial reference count from `drm_gem_object_init()`.
Ok(unsafe { ARef::from_raw(ptr) })
}
/// Returns the `Device` that owns this GEM object.
pub fn dev(&self) -> &drm::Device<T::Driver, Ctx> {
// SAFETY:
@ -356,11 +314,50 @@ extern "C" fn free_callback(obj: *mut bindings::drm_gem_object) {
}
}
impl<T: DriverObject> Object<T> {
/// Create a new GEM object.
pub fn new(dev: &drm::Device<T::Driver>, size: usize, args: T::Args) -> Result<ARef<Self>> {
let obj: Pin<KBox<Self>> = KBox::pin_init(
try_pin_init!(Self {
obj: Opaque::new(bindings::drm_gem_object::default()),
data <- T::new(dev, size, args),
_ctx: PhantomData,
}),
GFP_KERNEL,
)?;
// SAFETY: `obj.as_raw()` is guaranteed to be valid by the initialization above.
unsafe { (*obj.as_raw()).funcs = &Self::OBJECT_FUNCS };
// INVARIANT: `dev` and the GEM object are in the same state at the moment, and upgrading
// the typestate in `dev` will not carry over to the GEM object.
if let Err(err) =
// SAFETY: The arguments are all valid per the type invariants.
to_result(unsafe {
bindings::drm_gem_object_init(dev.as_raw(), obj.obj.get(), size)
})
{
// SAFETY: `drm_gem_object_init()` initializes the private GEM object state before
// failing, so `drm_gem_private_object_fini()` is the matching cleanup.
unsafe { bindings::drm_gem_private_object_fini(obj.obj.get()) };
return Err(err);
}
// SAFETY: We will never move out of `Self` as `ARef<Self>` is always treated as pinned.
let ptr = KBox::into_raw(unsafe { Pin::into_inner_unchecked(obj) });
// SAFETY: `ptr` comes from `KBox::into_raw` and hence can't be NULL.
let ptr = unsafe { NonNull::new_unchecked(ptr) };
// SAFETY: We take over the initial reference count from `drm_gem_object_init()`.
Ok(unsafe { ARef::from_raw(ptr) })
}
}
impl_aref_for_gem_obj! {
impl<T, C> for Object<T, C>
impl<T> for Object<T>
where
T: DriverObject,
C: DeviceContext
T: DriverObject
}
impl<T: DriverObject, Ctx: DeviceContext> super::private::Sealed for Object<T, Ctx> {}

View File

@ -11,28 +11,57 @@
use crate::{
container_of,
device::{
self,
Bound, //
},
devres::*,
drm::{
driver,
gem,
private::Sealed,
Device,
DeviceContext,
Registered, //
Device, //
},
error::{
from_err_ptr,
to_result, //
},
io::{
IoBase,
Region,
SysMem,
SysMemBackend, //
},
error::to_result,
prelude::*,
sync::aref::ARef,
types::Opaque, //
scatterlist,
sync::{
aref::ARef,
new_mutex,
Mutex,
SetOnce, //
},
types::{
NotThreadSafe,
Opaque, //
},
};
use core::{
marker::PhantomData,
ffi::c_void,
mem::{
ManuallyDrop,
MaybeUninit, //
},
ops::{
Deref,
DerefMut, //
},
ptr::NonNull, //
ptr::{
self,
NonNull, //
},
};
use gem::{
BaseObject,
BaseObjectPrivate,
DriverObject,
IntoGEMObject, //
@ -42,15 +71,24 @@
///
/// This is used with [`Object::new()`] to control various properties that can only be set when
/// initially creating a shmem-backed GEM object.
#[derive(Default)]
pub struct ObjectConfig<'a, T: DriverObject, C: DeviceContext = Registered> {
pub struct ObjectConfig<'a, T: DriverObject> {
/// Whether to set the write-combine map flag.
pub map_wc: bool,
/// Reuse the DMA reservation from another GEM object.
///
/// The newly created [`Object`] will hold an owned refcount to `parent_resv_obj` if specified.
pub parent_resv_obj: Option<&'a Object<T, C>>,
pub parent_resv_obj: Option<&'a Object<T>>,
}
impl<'a, T: DriverObject> Default for ObjectConfig<'a, T> {
#[inline(always)]
fn default() -> Self {
Self {
map_wc: false,
parent_resv_obj: None,
}
}
}
/// A shmem-backed GEM object.
@ -59,33 +97,35 @@ pub struct ObjectConfig<'a, T: DriverObject, C: DeviceContext = Registered> {
///
/// - `obj` contains a valid initialized `struct drm_gem_shmem_object` for the lifetime of this
/// object.
/// - Any type invariants of `C` apply to the parent DRM device for this GEM object.
#[repr(C)]
#[pin_data]
pub struct Object<T: DriverObject, C: DeviceContext = Registered> {
pub struct Object<T: DriverObject> {
#[pin]
obj: Opaque<bindings::drm_gem_shmem_object>,
/// Parent object that owns this object's DMA reservation object.
parent_resv_obj: Option<ARef<Object<T, C>>>,
parent_resv_obj: Option<ARef<Object<T>>>,
/// Devres object for unmapping any SGTable on driver-unbind.
sgt_res: ManuallyDrop<SetOnce<Devres<SGTableMap<T>>>>,
#[pin]
/// Lock for protecting initialization of `sgt_res`.
sgt_lock: Mutex<()>,
#[pin]
inner: T,
_ctx: PhantomData<C>,
}
super::impl_aref_for_gem_obj! {
impl<T, C> for Object<T, C>
impl<T> for Object<T>
where
T: DriverObject,
C: DeviceContext
T: DriverObject
}
// SAFETY: All GEM objects are thread-safe.
unsafe impl<T: DriverObject, C: DeviceContext> Send for Object<T, C> {}
unsafe impl<T: DriverObject> Send for Object<T> {}
// SAFETY: All GEM objects are thread-safe.
unsafe impl<T: DriverObject, C: DeviceContext> Sync for Object<T, C> {}
unsafe impl<T: DriverObject> Sync for Object<T> {}
impl<T: DriverObject, C: DeviceContext> Object<T, C> {
impl<T: DriverObject> Object<T> {
/// `drm_gem_object_funcs` vtable suitable for GEM shmem objects.
const VTABLE: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs {
free: Some(Self::free_callback),
@ -112,21 +152,166 @@ fn as_raw_shmem(&self) -> *mut bindings::drm_gem_shmem_object {
self.obj.get()
}
/// Returns the `Device` that owns this GEM object.
pub fn dev(&self) -> &Device<T::Driver> {
// SAFETY: `dev` will have been initialized in `Self::new()` by `drm_gem_shmem_init()`.
unsafe { Device::from_raw((*self.as_raw()).dev) }
}
extern "C" fn free_callback(obj: *mut bindings::drm_gem_object) {
// SAFETY:
// - DRM always passes a valid gem object here
// - We used drm_gem_shmem_create() in our create_gem_object callback, so we know that
// `obj` is contained within a drm_gem_shmem_object
let base = unsafe { container_of!(obj, bindings::drm_gem_shmem_object, base) };
// SAFETY:
// - We verified above that `obj` is valid, which makes `this` valid
// - This function is set in AllocOps, so we know that `this` is contained within an
// `Object<T>`
let this = unsafe { container_of!(Opaque::cast_from(base), Self, obj) }.cast_mut();
// We need to drop `sgt_res` first, since doing so requires that the GEM object is still
// alive.
// SAFETY:
// - We verified above that `this` is valid.
// - We are in free_callback, guaranteeing we have exclusive access to `this` and that
// `sgt_res` will not be used after dropping it here.
unsafe { ManuallyDrop::drop(&mut (*this).sgt_res) };
// SAFETY:
// - We're in free_callback - so this function is safe to call.
// - We won't be using the gem resources on `this` after this call.
unsafe { bindings::drm_gem_shmem_release(base) };
// SAFETY: We're recovering the Kbox<> we created in gem_create_object()
let _ = unsafe { KBox::from_raw(this) };
}
/// Attempt to create a vmap from the gem object, and confirm the size of said vmap.
fn make_vmap<'a, R, const SIZE: usize>(&'a self) -> Result<VMap<T, R, SIZE>>
where
R: Deref<Target = Self> + From<&'a Self>,
{
// INVARIANT: We check here that the gem object is at least as large as `SIZE`.
if self.size() < SIZE {
return Err(ENOSPC);
}
let mut map: MaybeUninit<bindings::iosys_map> = MaybeUninit::uninit();
let guard = DmaResvGuard::new(self);
// SAFETY: `drm_gem_shmem_vmap()` can be called with the DMA reservation lock held.
to_result(unsafe {
bindings::drm_gem_shmem_vmap_locked(self.as_raw_shmem(), map.as_mut_ptr())
})?;
// Drop the guard explicitly here, since we may need to call `raw_vunmap()` (which
// re-acquires the lock).
drop(guard);
// SAFETY: The call to `drm_gem_shmem_vmap_locked()` succeeded above, so we are guaranteed
// that map is properly initialized.
let map = unsafe { map.assume_init() };
// XXX: We don't currently support iomem allocations
if map.is_iomem {
// SAFETY: The vmap operation above succeeded, guaranteeing that `map` points to a valid
// memory mapping.
unsafe { self.raw_vunmap(map) };
Err(ENOTSUPP)
} else {
Ok(VMap {
// INVARIANT: `addr` remains valid for as long as `owner` does, which extends to the
// lifetime of `VMap` itself.
// SAFETY: We checked that this is not an iomem allocation, making it safe to read
// vaddr.
addr: unsafe { map.__bindgen_anon_1.vaddr },
owner: self.into(),
})
}
}
/// Unmap a vmap from the gem object.
///
/// # Safety
///
/// - The caller promises that `map` is a valid vmap on this gem object.
/// - The caller promises that the memory pointed to by map will no longer be accesed through
/// this instance.
unsafe fn raw_vunmap(&self, mut map: bindings::iosys_map) {
let _guard = DmaResvGuard::new(self);
// SAFETY:
// - This function is safe to call with the DMA reservation lock held.
// - The caller promises that `map` is a valid vmap on this gem object.
unsafe { bindings::drm_gem_shmem_vunmap_locked(self.as_raw_shmem(), &mut map) };
}
/// Creates and returns a virtual kernel memory mapping for this object.
#[inline]
pub fn vmap<const SIZE: usize>(&self) -> Result<VMapRef<'_, T, SIZE>> {
self.make_vmap()
}
/// Creates (if necessary) and returns an immutable reference to a scatter-gather table of DMA
/// pages for this object.
///
/// This will pin the object in memory. It is expected that `dev` should be a pointer to the
/// same [`device::Device`] which `self` belongs to, otherwise this function will return
/// `Err(EINVAL)`.
pub fn sg_table<'a>(
&'a self,
dev: &'a device::Device<Bound>,
) -> Result<&'a scatterlist::SGTable> {
let parent = self.dev().as_ref();
if dev.as_raw() != parent.as_ref().as_raw() {
return Err(EINVAL);
}
let sgt_res = 'out: {
// Fast path: sgt_res is already initialized
if let Some(sgt_res) = self.sgt_res.as_ref() {
break 'out sgt_res;
}
// Slow path: Grab the lock and see if we need to initialize sgt_res.
let _guard = self.sgt_lock.lock();
// If someone initialized it while we were waiting, we can exit early.
if let Some(sgt_res) = self.sgt_res.as_ref() {
break 'out sgt_res;
}
// If not, finish initializing and return. `populate()` cannot return false, as
// `sgt_res` must be unpopulated, and we must hold `sgt_lock` to reach this point.
self.sgt_res
.populate(Devres::new(dev, SGTableMap::new(self))?);
// SAFETY: We just populated sgt_res above.
unsafe { self.sgt_res.as_ref().unwrap_unchecked() }
};
Ok(sgt_res.access(dev)?)
}
/// Create a new shmem-backed DRM object of the given size.
///
/// Additional config options can be specified using `config`.
pub fn new(
dev: &Device<T::Driver, C>,
dev: &Device<T::Driver>,
size: usize,
config: ObjectConfig<'_, T, C>,
config: ObjectConfig<'_, T>,
args: T::Args,
) -> Result<ARef<Self>> {
let new: Pin<KBox<Self>> = KBox::try_pin_init(
try_pin_init!(Self {
obj <- Opaque::init_zeroed(),
parent_resv_obj: config.parent_resv_obj.map(|p| p.into()),
sgt_res: ManuallyDrop::new(SetOnce::new()),
sgt_lock <- new_mutex!(()),
inner <- T::new(dev, size, args),
_ctx: PhantomData::<C>,
}),
GFP_KERNEL,
)?;
@ -158,36 +343,14 @@ pub fn new(
Ok(obj)
}
/// Returns the `Device` that owns this GEM object.
pub fn dev(&self) -> &Device<T::Driver, C> {
// SAFETY: `dev` will have been initialized in `Self::new()` by `drm_gem_shmem_init()`.
unsafe { Device::from_raw((*self.as_raw()).dev) }
}
extern "C" fn free_callback(obj: *mut bindings::drm_gem_object) {
// SAFETY:
// - DRM always passes a valid gem object here
// - We used drm_gem_shmem_create() in our create_gem_object callback, so we know that
// `obj` is contained within a drm_gem_shmem_object
let this = unsafe { container_of!(obj, bindings::drm_gem_shmem_object, base) };
// SAFETY:
// - We're in free_callback - so this function is safe to call.
// - We won't be using the gem resources on `this` after this call.
unsafe { bindings::drm_gem_shmem_release(this) };
// SAFETY:
// - We verified above that `obj` is valid, which makes `this` valid
// - This function is set in AllocOps, so we know that `this` is contained within a
// `Object<T, C>`
let this = unsafe { container_of!(Opaque::cast_from(this), Self, obj) }.cast_mut();
// SAFETY: We're recovering the Kbox<> we created in gem_create_object()
let _ = unsafe { KBox::from_raw(this) };
/// Creates and returns an owned reference to a virtual kernel memory mapping for this object.
#[inline]
pub fn owned_vmap<const SIZE: usize>(&self) -> Result<VMapOwned<T, SIZE>> {
self.make_vmap()
}
}
impl<T: DriverObject, C: DeviceContext> Deref for Object<T, C> {
impl<T: DriverObject> Deref for Object<T> {
type Target = T;
fn deref(&self) -> &Self::Target {
@ -195,15 +358,15 @@ fn deref(&self) -> &Self::Target {
}
}
impl<T: DriverObject, C: DeviceContext> DerefMut for Object<T, C> {
impl<T: DriverObject> DerefMut for Object<T> {
fn deref_mut(&mut self) -> &mut Self::Target {
&mut self.inner
}
}
impl<T: DriverObject, C: DeviceContext> Sealed for Object<T, C> {}
impl<T: DriverObject> Sealed for Object<T> {}
impl<T: DriverObject, C: DeviceContext> gem::IntoGEMObject for Object<T, C> {
impl<T: DriverObject> gem::IntoGEMObject for Object<T> {
fn as_raw(&self) -> *mut bindings::drm_gem_object {
// SAFETY:
// - Our immutable reference is proof that this is safe to dereference.
@ -222,7 +385,7 @@ unsafe fn from_raw<'a>(obj: *mut bindings::drm_gem_object) -> &'a Self {
}
}
impl<T: DriverObject, C: DeviceContext> driver::AllocImpl for Object<T, C> {
impl<T: DriverObject> driver::AllocImpl for Object<T> {
type Driver = T::Driver;
const ALLOC_OPS: driver::AllocOps = driver::AllocOps {
@ -235,3 +398,324 @@ impl<T: DriverObject, C: DeviceContext> driver::AllocImpl for Object<T, C> {
dumb_map_offset: None,
};
}
/// Private helper-type for holding the `dma_resv` object for a GEM shmem object.
///
/// When this is dropped, the `dma_resv` lock is dropped as well.
///
// TODO: This should be replace with a WwMutex equivalent once we have such bindings in the kernel.
struct DmaResvGuard<'a, T: DriverObject>(&'a Object<T>, NotThreadSafe);
impl<'a, T: DriverObject> DmaResvGuard<'a, T> {
#[inline]
fn new(obj: &'a Object<T>) -> Self {
// SAFETY: This lock is initialized throughout the lifetime of `object`.
unsafe { bindings::dma_resv_lock(obj.raw_dma_resv(), ptr::null_mut()) };
Self(obj, NotThreadSafe)
}
}
impl<'a, T: DriverObject> Drop for DmaResvGuard<'a, T> {
#[inline]
fn drop(&mut self) {
// SAFETY: We are releasing the lock grabbed during the creation of this object.
unsafe { bindings::dma_resv_unlock(self.0.raw_dma_resv()) };
}
}
/// A reference to a virtual mapping for an shmem-based GEM object in kernel address space.
///
/// # Invariants
///
/// - The size of `owner` is >= SIZE.
/// - The memory pointed to by `addr` remains valid at least until this object is dropped.
pub struct VMap<D, R, const SIZE: usize = 0>
where
D: DriverObject,
R: Deref<Target = Object<D>>,
{
addr: *mut c_void,
owner: R,
}
/// An alias type for a reference to a shmem-based GEM object's VMap.
pub type VMapRef<'a, D, const SIZE: usize = 0> = VMap<D, &'a Object<D>, SIZE>;
/// An alias type for an owned reference to a shmem-based GEM object's VMap.
pub type VMapOwned<D, const SIZE: usize = 0> = VMap<D, ARef<Object<D>>, SIZE>;
impl<D, R, const SIZE: usize> VMap<D, R, SIZE>
where
D: DriverObject,
R: Deref<Target = Object<D>>,
{
/// Borrows a reference to the object that owns this virtual mapping.
#[inline]
pub fn owner(&self) -> &Object<D> {
&self.owner
}
}
impl<'a, D, R, const SIZE: usize> IoBase<'a> for &'a VMap<D, R, SIZE>
where
D: DriverObject,
R: Deref<Target = Object<D>>,
{
type Backend = SysMemBackend;
type Target = Region<SIZE>;
#[inline]
fn as_view(self) -> SysMem<'a, Region<SIZE>> {
let ptr = Region::ptr_from_raw_parts_mut(self.addr.cast(), self.owner.size());
// SAFETY: Per type invariants of `VMap`:
// - `addr .. addr + owner.size()` is a valid kernel accessible memory region.
// - `addr` is page-aligned, which satisfies `Region`'s 4-byte alignment requirement.
// - The memory remains valid until this `VMap` is dropped; since `self` is `&'a VMap`,
// the borrow prevents the `VMap` from being dropped for the lifetime `'a`.
unsafe { SysMem::new(ptr) }
}
}
impl<D, R, const SIZE: usize> Drop for VMap<D, R, SIZE>
where
D: DriverObject,
R: Deref<Target = Object<D>>,
{
#[inline]
fn drop(&mut self) {
// SAFETY:
// - Our existence is proof that this map was previously created using self.owner.
// - Since we are in Drop, we are guaranteed that no one will access the memory
// through this mapping after calling this.
unsafe {
self.owner.raw_vunmap(bindings::iosys_map {
is_iomem: false,
__bindgen_anon_1: bindings::iosys_map__bindgen_ty_1 { vaddr: self.addr },
})
};
}
}
// SAFETY: `addr` points to a valid memory address for as long as `owner` exists, meaning that so
// long as `owner` is `Send` so is `VMap`.
unsafe impl<D, R, const SIZE: usize> Send for VMap<D, R, SIZE>
where
D: DriverObject,
R: Deref<Target = Object<D>> + Send,
{
}
// SAFETY: `addr` points to a valid memory address for as long as `owner` exists, meaning that so
// long as `owner` is `Sync` so is `VMap`.
unsafe impl<D, R, const SIZE: usize> Sync for VMap<D, R, SIZE>
where
D: DriverObject,
R: Deref<Target = Object<D>> + Sync,
{
}
/// A reference to a GEM object that is known to have a mapped [`SGTable`].
///
/// This is used by the Rust bindings with [`Devres`] in order to ensure that mappings for SGTables
/// on GEM shmem objects are revoked on driver-unbind.
///
/// # Invariants
///
/// - `self.obj` always points to a valid GEM object.
/// - This object is proof that `self.obj.owner.sgt_res` has an initialized and valid pointer to an
/// [`SGTable`].
///
/// [`SGTable`]: scatterlist::SGTable
pub struct SGTableMap<T: DriverObject> {
obj: NonNull<Object<T>>,
}
impl<T: DriverObject> Deref for SGTableMap<T> {
type Target = scatterlist::SGTable;
fn deref(&self) -> &Self::Target {
// SAFETY:
// - The NonNull is guaranteed to be valid via our type invariants.
// - The sgt field is guaranteed to be initialized and valid via our type invariants.
unsafe { scatterlist::SGTable::from_raw((*self.obj.as_ref().as_raw_shmem()).sgt) }
}
}
impl<T: DriverObject> Drop for SGTableMap<T> {
fn drop(&mut self) {
// SAFETY: `obj` is always valid via our type invariants
let obj = unsafe { self.obj.as_ref() };
let _lock = DmaResvGuard::new(obj);
// SAFETY: We acquired the lock needed for calling this function above
unsafe { bindings::__drm_gem_shmem_free_sgt_locked(obj.as_raw_shmem()) };
}
}
impl<T: DriverObject> SGTableMap<T> {
fn new(obj: &Object<T>) -> impl Init<Self, Error> {
// INVARIANT:
// - We call drm_gem_shmem_get_pages_sgt below and check whether or not it succeeds,
// fulfilling the invariant of SGTableMap that the object's `sgt` field is initialized.
// SAFETY:
// - `obj` is fully initialized, making this function safe to call.
from_err_ptr(unsafe { bindings::drm_gem_shmem_get_pages_sgt(obj.as_raw_shmem()) })?;
Ok(Self { obj: obj.into() })
}
}
// SAFETY: The NonNull in SGTableMap is guaranteed valid by our type invariants, and the GEM object
// it points to is guaranteed to be thread-safe.
unsafe impl<T: DriverObject> Send for SGTableMap<T> {}
// SAFETY: The NonNull in SGTableMap is guaranteed valid by our type invariants, and the GEM object
// it points to is guaranteed to be thread-safe.
unsafe impl<T: DriverObject> Sync for SGTableMap<T> {}
#[kunit_tests(rust_drm_gem_shmem)]
mod tests {
use super::*;
use crate::{
drm::{
self,
UnregisteredDevice, //
},
faux,
io::Io,
page::PAGE_SIZE, //
};
// The bare minimum needed to create a fake drm driver for kunit
#[pin_data]
struct KunitData {}
struct KunitDriver;
struct KunitFile;
#[pin_data]
struct KunitObject {}
const INFO: drm::DriverInfo = drm::DriverInfo {
major: 0,
minor: 0,
patchlevel: 0,
name: c"kunit",
desc: c"Kunit",
};
impl drm::file::DriverFile for KunitFile {
type Driver = KunitDriver;
fn open(_dev: &drm::Device<KunitDriver>) -> Result<Pin<KBox<Self>>> {
Ok(KBox::new(Self, GFP_KERNEL)?.into())
}
}
impl gem::DriverObject for KunitObject {
type Driver = KunitDriver;
type Args = ();
fn new(
_dev: &drm::Device<KunitDriver>,
_size: usize,
_args: Self::Args,
) -> impl PinInit<Self, Error> {
try_pin_init!(KunitObject {})
}
}
#[vtable]
impl drm::Driver for KunitDriver {
type Data = KunitData;
type RegistrationData<'a> = ();
type File = KunitFile;
type Object = Object<KunitObject>;
type ParentDevice<Ctx: device::DeviceContext> = faux::Device<Ctx>;
const INFO: drm::DriverInfo = INFO;
const IOCTLS: &'static [drm::ioctl::DrmIoctlDescriptor] = &[];
}
fn create_drm_dev() -> Result<(faux::Registration, UnregisteredDevice<KunitDriver>)> {
// Create a faux DRM device so we can test gem object creation.
let data = try_pin_init!(KunitData {});
let reg = faux::Registration::new(c"Kunit", None)?;
let fdev = reg.as_ref();
let drm = UnregisteredDevice::new(fdev, data)?;
Ok((reg, drm))
}
#[test]
fn compile_time_vmap_sizes() -> Result {
let (_dev, drm) = create_drm_dev()?;
let obj = Object::<KunitObject>::new(&drm, PAGE_SIZE, ObjectConfig::default(), ())?;
// Try creating a normal vmap
obj.vmap::<PAGE_SIZE>()?;
// Try creating a vmap that's smaller then the size we specified
let vmap = obj.vmap::<{ PAGE_SIZE - 100 }>()?;
// Verify the owner matches
assert!(ptr::eq(vmap.owner(), obj.deref()));
// Verify the size matches the actual object size
assert_eq!(vmap.size(), PAGE_SIZE);
// Make sure creating a vmap that's too large fails
assert!(obj.vmap::<{ PAGE_SIZE + 200 }>().is_err());
Ok(())
}
#[test]
fn vmap_io() -> Result {
let (_dev, drm) = create_drm_dev()?;
let obj = Object::<KunitObject>::new(&drm, PAGE_SIZE, ObjectConfig::default(), ())?;
let vmap = obj.vmap::<PAGE_SIZE>()?;
vmap.write8(0xDE, 0x0);
assert_eq!(vmap.read8(0x0), 0xDE);
vmap.write32(0xFEDCBA98, 0x20);
assert_eq!(vmap.read32(0x20), 0xFEDCBA98);
// Ensure the ordering in memory is correct
let expected = 0xFEDCBA98_u32.to_ne_bytes().into_iter();
for (offset, expected) in (0x20..=0x23).zip(expected) {
assert_eq!(vmap.try_read8(offset).unwrap(), expected);
}
Ok(())
}
// TODO: I would love to actually test the success paths of sg_table(), but that would require
// also implementing dummy dma_ops so that trying to create a mapping doesn't explode. So, leave
// that for someone else.
// Ensures that passing the wrong device to sg_table() fails as we expect, and also ensure it
// skips initializing `sgt_res` since we could otherwise create `sgt_res` with the wrong device
// bound to it.
#[test]
fn fail_sg_table_on_wrong_dev() -> Result {
let (_dev, drm) = create_drm_dev()?;
let reg = faux::Registration::new(c"EvilKunit", None)?;
let wrong_dev = reg.as_ref();
let obj = Object::<KunitObject>::new(&drm, PAGE_SIZE, ObjectConfig::default(), ())?;
assert_eq!(obj.sg_table(wrong_dev.as_ref()).err().unwrap(), EINVAL);
// If sgt_res was not initialized mistakenly with the wrong device, this should still fail.
assert_eq!(obj.sg_table(wrong_dev.as_ref()).err().unwrap(), EINVAL);
// TODO: Someday, we should test that creating an sg_table here still succeeds.
Ok(())
}
}

View File

@ -72,10 +72,12 @@ pub struct GpuVm<T: DriverGpuVm> {
data: UnsafeCell<T>,
}
// SAFETY: The GPUVM api does not assume that it is tied to a specific thread. The destructor will
// drop the `data` field, which is okay because it is guaranteed `Send` by the `DriverGpuVm` trait.
// SAFETY: It is safe to send a `GpuVm<T>` to another thread: all data reachable through it
// (`T`, `T::VmBoData`, and the GEM `T::Object`) is `Send` by the `DriverGpuVm` bounds.
unsafe impl<T: DriverGpuVm> Send for GpuVm<T> {}
// SAFETY: The GPUVM api is designed to allow &self methods to be called in parallel.
// SAFETY: It is safe to share a `&GpuVm<T>` between threads: `&self` methods only alias data
// that is `Sync` by the `DriverGpuVm` bounds, and any thread may drop that data, or upgrade the
// reference and ultimately drop `T`, which the same bounds make `Send`.
unsafe impl<T: DriverGpuVm> Sync for GpuVm<T> {}
// SAFETY: By type invariants, the allocation is managed by the refcount in `self.vm`.
@ -116,9 +118,9 @@ const fn vtable() -> &'static bindings::drm_gpuvm_ops {
/// Creates a GPUVM instance.
#[expect(clippy::new_ret_no_self)]
pub fn new<E>(
pub fn new<E, Ctx: drm::DeviceContext>(
name: &'static CStr,
dev: &drm::Device<T::Driver>,
dev: &drm::Device<T::Driver, Ctx>,
r_obj: &T::Object,
range: Range<u64>,
reserve_range: Range<u64>,
@ -250,21 +252,27 @@ fn raw_resv(&self) -> *mut bindings::dma_resv {
}
/// The manager for a GPUVM.
pub trait DriverGpuVm: Sized + Send {
pub trait DriverGpuVm: Sized + Send + Sync {
/// Parent `Driver` for this object.
type Driver: drm::Driver<Object = Self::Object>;
type Driver: drm::Driver;
/// The kind of GEM object stored in this GPUVM.
type Object: IntoGEMObject;
type Object: drm::driver::AllocImpl<Driver = Self::Driver> + Send + Sync;
/// Data stored with each [`struct drm_gpuva`](struct@GpuVa).
type VaData;
///
/// Only `Send` is required: the data has a single owner at all times, moving
/// between threads by value (handed back as a [`GpuVaRemoved`]) but never
/// accessed by two threads concurrently.
type VaData: Send;
/// Data stored with each [`struct drm_gpuvm_bo`](struct@GpuVmBo).
type VmBoData;
type VmBoData: Send + Sync;
/// The private data passed to callbacks.
type SmContext<'ctx>;
type SmContext<'ctx>
where
Self: 'ctx;
/// Indicates that a new mapping should be created.
fn sm_step_map<'op, 'ctx>(
@ -296,12 +304,10 @@ fn sm_step_remap<'op, 'ctx>(
/// # Invariants
///
/// Each `GpuVm` instance has at most one `UniqueRefGpuVm` reference.
// `Send`/`Sync` derive from `ARef<GpuVm<T>>`; the trait bounds make them correct for the unique
// handle's `&mut T` access.
pub struct UniqueRefGpuVm<T: DriverGpuVm>(ARef<GpuVm<T>>);
// SAFETY: The GPUVM api is designed to allow &self methods to be called in parallel, and
// concurrent access to `data` is safe due to the `T: Sync` requirement.
unsafe impl<T: DriverGpuVm + Sync> Sync for UniqueRefGpuVm<T> {}
impl<T: DriverGpuVm> UniqueRefGpuVm<T> {
/// Access the data owned by this `UniqueRefGpuVm` immutably.
#[inline]

View File

@ -3,7 +3,7 @@
use super::*;
/// The actual data that gets threaded through the callbacks.
struct SmData<'a, 'ctx, T: DriverGpuVm> {
struct SmData<'a, 'ctx, T: DriverGpuVm + 'ctx> {
gpuvm: &'a mut UniqueRefGpuVm<T>,
user_context: &'a mut T::SmContext<'ctx>,
}
@ -20,7 +20,7 @@ struct SmMapData<'a, 'ctx, T: DriverGpuVm> {
}
/// The argument for [`UniqueRefGpuVm::sm_map`].
pub struct OpMapRequest<'a, 'ctx, T: DriverGpuVm> {
pub struct OpMapRequest<'a, 'ctx, T: DriverGpuVm + 'ctx> {
/// Address in GPU virtual address space.
pub addr: u64,
/// Length of mapping to create.

View File

@ -104,6 +104,14 @@ pub fn vm_bo(&self) -> &GpuVmBo<T> {
/// The memory is zeroed.
pub struct GpuVaAlloc<T: DriverGpuVm>(KBox<MaybeUninit<GpuVa<T>>>);
// SAFETY: A `GpuVaAlloc` is an owned, uninitialised allocation with no live `T::VaData` and no
// thread-bound state.
unsafe impl<T: DriverGpuVm> Send for GpuVaAlloc<T> {}
// SAFETY: A `GpuVaAlloc` has no `&self` method that reaches its contents, so a shared
// `&GpuVaAlloc` cannot access the allocation.
unsafe impl<T: DriverGpuVm> Sync for GpuVaAlloc<T> {}
impl<T: DriverGpuVm> GpuVaAlloc<T> {
/// Pre-allocate a [`GpuVa`] object.
pub fn new(flags: AllocFlags) -> Result<GpuVaAlloc<T>, AllocError> {

View File

@ -19,6 +19,15 @@ pub struct GpuVmBo<T: DriverGpuVm> {
data: T::VmBoData,
}
// SAFETY: It is safe to send a `GpuVmBo<T>` to another thread: dropping it there drops
// `T::VmBoData` and the GEM `T::Object`, both `Send` by the `DriverGpuVm` bounds.
unsafe impl<T: DriverGpuVm> Send for GpuVmBo<T> {}
// SAFETY: It is safe to share a `&GpuVmBo<T>` between threads: it effectively shares
// `&T::VmBoData` and the GEM `&T::Object` (both `Sync`), and any thread may upgrade to an
// `ARef` and ultimately drop them (both `Send`), per the `DriverGpuVm` bounds.
unsafe impl<T: DriverGpuVm> Sync for GpuVmBo<T> {}
// SAFETY: By type invariants, the allocation is managed by the refcount in `self.inner`.
unsafe impl<T: DriverGpuVm> AlwaysRefCounted for GpuVmBo<T> {
fn inc_ref(&self) {

View File

@ -70,6 +70,18 @@ pub mod internal {
pub use bindings::drm_device;
pub use bindings::drm_file;
pub use bindings::drm_ioctl_desc;
/// Cast an [`Ioctl`] DRM device pointer to [`Registered`], preserving the driver type
/// parameter `T`.
///
/// Used by [`declare_drm_ioctls!`] to anchor type inference.
#[doc(hidden)]
#[inline]
pub const fn __dev_ctx_cast<T: crate::drm::Driver>(
ptr: *const crate::drm::Device<T, crate::drm::Ioctl>,
) -> *const crate::drm::Device<T, crate::drm::Registered> {
ptr.cast()
}
}
/// Declare the DRM ioctls for a driver.
@ -82,7 +94,8 @@ pub mod internal {
/// `user_callback` should have the following prototype:
///
/// ```ignore
/// fn foo(device: &kernel::drm::Device<Self>,
/// fn foo(device: &kernel::drm::Device<Self, kernel::drm::Registered>,
/// reg_data: &Self::RegistrationData<'_>,
/// data: &mut uapi::argument_type,
/// file: &kernel::drm::File<Self::File>,
/// ) -> Result<u32>
@ -131,10 +144,45 @@ macro_rules! declare_drm_ioctls {
// - The DRM device must have been registered when we're called through
// an IOCTL.
//
// INVARIANT: The `Ioctl` context requires that the device has been
// registered via `drm_dev_register()` at some point; the DRM core
// guarantees this for ioctl dispatch callbacks.
//
// FIXME: Currently there is nothing enforcing that the types of the
// dev/file match the current driver these ioctls are being declared
// for, and it's not clear how to enforce this within the type system.
let dev = $crate::drm::device::Device::from_raw(raw_dev);
let dev: &$crate::drm::device::Device<_, $crate::drm::Ioctl> =
$crate::drm::device::Device::from_raw(raw_dev);
// Type-inference anchor: the closure is never called but ties `dev`'s
// type to `$func`'s first parameter, which the compiler cannot infer
// through method resolution and associated-type projections alone.
#[allow(unreachable_code)]
let _ = || {
let __ptr = $crate::drm::ioctl::internal::__dev_ctx_cast(
::core::ptr::from_ref(dev),
);
$func(
// SAFETY: This closure is never executed; the dereference
// exists purely to unify the type parameter with `$func`.
// The pointer is valid regardless.
unsafe { &*__ptr },
unreachable!(),
unreachable!(),
unreachable!(),
)
};
// Enforce that the handler accepts higher-ranked
// lifetimes, preventing it from requiring 'static
// references that could escape this scope.
let _: for<'a> fn(&'a _, &'a _, &'a mut _, &'a _) -> _ = $func;
let Some(guard) = dev.registration_guard() else {
return $crate::error::code::ENODEV.to_errno();
};
// SAFETY: The ioctl argument has size `_IOC_SIZE(cmd)`, which we
// asserted above matches the size of this type, and all bit patterns of
// UAPI structs must be valid.
@ -147,7 +195,9 @@ macro_rules! declare_drm_ioctls {
// SAFETY: This is just the DRM file structure
let file = unsafe { $crate::drm::File::from_raw(raw_file) };
match $func(dev, data, file) {
match guard.registration_data_with(|reg_data| {
$func(&*guard, reg_data, data, file)
}) {
Err(e) => e.to_errno(),
Ok(i) => i.try_into()
.unwrap_or($crate::error::code::ERANGE.to_errno()),

View File

@ -11,8 +11,10 @@
pub use self::device::Device;
pub use self::device::DeviceContext;
pub use self::device::Ioctl;
pub use self::device::Normal;
pub use self::device::Registered;
pub use self::device::Uninit;
pub use self::device::RegistrationGuard;
pub use self::device::UnregisteredDevice;
pub use self::driver::Driver;
pub use self::driver::DriverInfo;

View File

@ -9,15 +9,63 @@
use crate::{
bindings,
device,
prelude::*, //
prelude::*,
types::Opaque, //
};
use core::ptr::{
addr_of_mut,
null,
null_mut,
NonNull, //
use core::{
marker::PhantomData,
ptr::{
null,
null_mut,
NonNull, //
},
};
/// A faux device.
///
/// A faux device is a virtual device backed by the faux bus, primarily used for scenarios where a
/// real hardware device is not available or for testing.
///
/// # Invariants
///
/// The underlying `struct faux_device` is valid.
#[repr(transparent)]
pub struct Device<Ctx: device::DeviceContext = device::Normal>(
Opaque<bindings::faux_device>,
PhantomData<Ctx>,
);
impl<Ctx: device::DeviceContext> Device<Ctx> {
#[inline]
fn as_raw(&self) -> *mut bindings::faux_device {
self.0.get()
}
/// # Safety
///
/// `ptr` must be a valid pointer to a `struct faux_device`.
#[inline]
unsafe fn from_raw<'a>(ptr: *mut bindings::faux_device) -> &'a Self {
// SAFETY: `Device` is a transparent wrapper of `Opaque<bindings::faux_device>`.
unsafe { &*ptr.cast() }
}
}
impl<Ctx: device::DeviceContext> AsRef<device::Device<Ctx>> for Device<Ctx> {
#[inline]
fn as_ref(&self) -> &device::Device<Ctx> {
// SAFETY: By the type invariant of `Self`, `self.as_raw()` is a pointer to a valid
// `struct faux_device`. `dev` points to a valid `struct device`.
unsafe { device::Device::from_raw(&raw mut (*self.as_raw()).dev) }
}
}
// SAFETY: `faux::Device` is a transparent wrapper of `struct faux_device`.
// The offset is guaranteed to point to a valid device field inside `faux::Device`.
unsafe impl<Ctx: device::DeviceContext> device::AsBusDevice<Ctx> for Device<Ctx> {
const OFFSET: usize = core::mem::offset_of!(bindings::faux_device, dev);
}
/// The registration of a faux device.
///
/// This type represents the registration of a [`struct faux_device`]. When an instance of this type
@ -25,7 +73,8 @@
///
/// # Invariants
///
/// `self.0` always holds a valid pointer to an initialized and registered [`struct faux_device`].
/// - `self.0` always holds a valid pointer to an initialized and registered [`struct faux_device`].
/// - This object is proof that the object described by this `Registration` is bound to a device.
///
/// [`struct faux_device`]: srctree/include/linux/device/faux.h
pub struct Registration(NonNull<bindings::faux_device>);
@ -59,11 +108,19 @@ fn as_raw(&self) -> *mut bindings::faux_device {
}
}
impl AsRef<device::Device> for Registration {
fn as_ref(&self) -> &device::Device {
// SAFETY: The underlying `device` in `faux_device` is guaranteed by the C API to be
// a valid initialized `device`.
unsafe { device::Device::from_raw(addr_of_mut!((*self.as_raw()).dev)) }
impl AsRef<Device<device::Bound>> for Registration {
#[inline]
fn as_ref(&self) -> &Device<device::Bound> {
// SAFETY:
// - The underlying `struct faux_device` is guaranteed by the C API to be a valid
// initialized `device`.
// - `faux_match()` always returns 1, and probe runs synchronously
// (PROBE_FORCE_SYNCHRONOUS).
// - `suppress_bind_attrs = true` on faux_driver prevents userspace-triggered unbind via
// sysfs.
// - `mem::forget(Registration)` is not a problem; if the `Registration` is leaked, the faux
// device stays bound forever.
unsafe { Device::from_raw(self.as_raw()) }
}
}

View File

@ -7,9 +7,9 @@
use crate::{
bindings,
device::Device,
error::Error,
error::Result,
error::to_result,
ffi,
prelude::*,
str::{CStr, CStrExt as _},
};
use core::ptr::NonNull;
@ -120,6 +120,48 @@ fn drop(&mut self) {
}
}
/// Load firmware directly into the caller-provided `buf`.
///
/// On success the firmware image has been copied into `buf`; the caller accesses the data
/// through `buf` itself.
///
/// This is intentionally a stand-alone function rather than a `Firmware` constructor. For
/// the `into_buf` path, the firmware data lives in the caller's `buf`, not in a
/// kernel-owned buffer, so returning a `Firmware` would expose `Firmware::data()` as a
/// second handle aliasing `buf` (and `release_firmware()` does not free `buf` anyway).
pub fn request_into_buf(name: &CStr, dev: &Device, buf: &mut [u8]) -> Result {
// `as_mut_ptr()` on an empty slice returns a non-NULL pointer to
// memory which the loader does not own. Passing that pointer with `size == 0`
// makes the loader believe that it is buffer it allocated itself, so when
// `release_firmware()` is called, it will vfree the pointer and trigger a
// bug. Reject empty slices to avoid this situation.
if buf.is_empty() {
return Err(EINVAL);
}
let mut fw: *const bindings::firmware = core::ptr::null();
// SAFETY: `&raw mut fw` is a valid pointer to a NULL initialized `bindings::firmware` pointer.
// `name` and `dev` are valid as by their type invariants. `buf` is a valid writable
// buffer of `buf.len()` bytes.
to_result(unsafe {
bindings::request_firmware_into_buf(
&raw mut fw,
name.as_char_ptr(),
dev.as_raw(),
buf.as_mut_ptr().cast(),
buf.len(),
)
})?;
// The firmware bytes are now in `buf`, which the caller owns, so we don't need
// the kernel to hang on to it any more.
// SAFETY: `fw` is a valid pointer returned by `request_firmware_into_buf`.
unsafe { bindings::release_firmware(fw) };
Ok(())
}
// SAFETY: `Firmware` only holds a pointer to a C `struct firmware`, which is safe to be used from
// any thread.
unsafe impl Send for Firmware {}

File diff suppressed because it is too large Load Diff

View File

@ -2,8 +2,6 @@
//! Generic memory-mapped IO.
use core::ops::Deref;
use crate::{
device::{
Bound,
@ -16,7 +14,9 @@
Region,
Resource, //
},
IoBase,
Mmio,
MmioBackend,
MmioRaw, //
},
prelude::*,
@ -210,11 +210,13 @@ pub fn into_devres(self) -> Result<Devres<ExclusiveIoMem<'static, SIZE>>> {
}
}
impl<const SIZE: usize> Deref for ExclusiveIoMem<'_, SIZE> {
type Target = Mmio<SIZE>;
impl<'a, const SIZE: usize> IoBase<'a> for &'a ExclusiveIoMem<'_, SIZE> {
type Backend = MmioBackend;
type Target = super::Region<SIZE>;
fn deref(&self) -> &Self::Target {
&self.iomem
#[inline]
fn as_view(self) -> Mmio<'a, Self::Target> {
self.iomem.as_view()
}
}
@ -229,7 +231,7 @@ fn deref(&self) -> &Self::Target {
/// start of the I/O memory mapped region.
pub struct IoMem<'a, const SIZE: usize = 0> {
dev: &'a Device<Bound>,
io: MmioRaw<SIZE>,
io: MmioRaw<super::Region<SIZE>>,
}
impl<'a, const SIZE: usize> IoMem<'a, SIZE> {
@ -264,8 +266,7 @@ fn ioremap(dev: &'a Device<Bound>, resource: &Resource) -> Result<Self> {
return Err(ENOMEM);
}
let io = MmioRaw::new(addr as usize, size)?;
let io = MmioRaw::new_region(addr as usize, size)?;
Ok(IoMem { dev, io })
}
@ -291,11 +292,13 @@ fn drop(&mut self) {
}
}
impl<const SIZE: usize> Deref for IoMem<'_, SIZE> {
type Target = Mmio<SIZE>;
impl<'a, const SIZE: usize> IoBase<'a> for &'a IoMem<'_, SIZE> {
type Backend = MmioBackend;
type Target = super::Region<SIZE>;
fn deref(&self) -> &Self::Target {
#[inline]
fn as_view(self) -> Mmio<'a, Self::Target> {
// SAFETY: Safe as by the invariant of `IoMem`.
unsafe { Mmio::from_raw(&self.io) }
unsafe { Mmio::from_raw(self.io) }
}
}

View File

@ -48,13 +48,14 @@
/// use kernel::io::{
/// Io,
/// Mmio,
/// Region,
/// poll::read_poll_timeout, //
/// };
/// use kernel::time::Delta;
///
/// const HW_READY: u16 = 0x01;
///
/// fn wait_for_hardware<const SIZE: usize>(io: &Mmio<SIZE>) -> Result {
/// fn wait_for_hardware<const SIZE: usize>(io: Mmio<'_, Region<SIZE>>) -> Result {
/// read_poll_timeout(
/// // The `op` closure reads the value of a specific status register.
/// || io.try_read16(0x1000),
@ -135,13 +136,14 @@ pub fn read_poll_timeout<Op, Cond, T>(
/// use kernel::io::{
/// Io,
/// Mmio,
/// Region,
/// poll::read_poll_timeout_atomic, //
/// };
/// use kernel::time::Delta;
///
/// const HW_READY: u16 = 0x01;
///
/// fn wait_for_hardware<const SIZE: usize>(io: &Mmio<SIZE>) -> Result {
/// fn wait_for_hardware<const SIZE: usize>(io: Mmio<'_, Region<SIZE>>) -> Result {
/// read_poll_timeout_atomic(
/// // The `op` closure reads the value of a specific status register.
/// || io.try_read16(0x1000),

View File

@ -58,7 +58,7 @@
//! },
//! num::Bounded,
//! };
//! # use kernel::io::Mmio;
//! # use kernel::io::{Mmio, Region};
//! # register! {
//! # pub BOOT_0(u32) @ 0x00000100 {
//! # 15:8 vendor_id;
@ -66,7 +66,7 @@
//! # 3:0 minor_revision;
//! # }
//! # }
//! # fn test(io: &Mmio<0x1000>) {
//! # fn test(io: Mmio<'_, Region<0x1000>>) {
//! # fn obtain_vendor_id() -> u8 { 0xff }
//!
//! // Read from the register's defined offset (0x100).
@ -113,6 +113,8 @@
io::IoLoc, //
};
use super::Region;
/// Trait implemented by all registers.
pub trait Register: Sized {
/// Backing primitive type of the register.
@ -129,7 +131,7 @@ pub trait FixedRegister: Register {}
/// Allows `()` to be used as the `location` parameter of [`Io::write`](super::Io::write) when
/// passing a [`FixedRegister`] value.
impl<T> IoLoc<T> for ()
impl<const SIZE: usize, T> IoLoc<Region<SIZE>, T> for ()
where
T: FixedRegister,
{
@ -143,7 +145,7 @@ fn offset(self) -> usize {
/// A [`FixedRegister`] carries its location in its type. Thus `FixedRegister` values can be used
/// as an [`IoLoc`].
impl<T> IoLoc<T> for T
impl<const SIZE: usize, T> IoLoc<Region<SIZE>, T> for T
where
T: FixedRegister,
{
@ -168,7 +170,7 @@ pub const fn new() -> Self {
}
}
impl<T> IoLoc<T> for FixedRegisterLoc<T>
impl<const SIZE: usize, T> IoLoc<Region<SIZE>, T> for FixedRegisterLoc<T>
where
T: FixedRegister,
{
@ -239,7 +241,7 @@ const fn offset(self) -> usize {
}
}
impl<T, B> IoLoc<T> for RelativeRegisterLoc<T, B>
impl<const SIZE: usize, T, B> IoLoc<Region<SIZE>, T> for RelativeRegisterLoc<T, B>
where
T: RelativeRegister,
B: RegisterBase<T::BaseFamily> + ?Sized,
@ -283,7 +285,7 @@ pub fn try_new(idx: usize) -> Option<Self> {
}
}
impl<T> IoLoc<T> for RegisterArrayLoc<T>
impl<const SIZE: usize, T> IoLoc<Region<SIZE>, T> for RegisterArrayLoc<T>
where
T: RegisterArray,
{
@ -370,7 +372,7 @@ pub fn try_at(self, idx: usize) -> Option<RelativeRegisterArrayLoc<T, B>> {
}
}
impl<T, B> IoLoc<T> for RelativeRegisterArrayLoc<T, B>
impl<const SIZE: usize, T, B> IoLoc<Region<SIZE>, T> for RelativeRegisterArrayLoc<T, B>
where
T: RelativeRegisterArray,
B: RegisterBase<T::BaseFamily> + ?Sized,
@ -387,18 +389,18 @@ fn offset(self) -> usize {
/// which to write it.
///
/// Implementors can be used with [`Io::write_reg`](super::Io::write_reg).
pub trait LocatedRegister {
pub trait LocatedRegister<Base: ?Sized> {
/// Register value to write.
type Value: Register;
/// Full location information at which to write the value.
type Location: IoLoc<Self::Value>;
type Location: IoLoc<Base, Self::Value>;
/// Consumes `self` and returns a `(location, value)` tuple describing a valid I/O write
/// operation.
fn into_io_op(self) -> (Self::Location, Self::Value);
}
impl<T> LocatedRegister for T
impl<const SIZE: usize, T> LocatedRegister<Region<SIZE>> for T
where
T: FixedRegister,
{
@ -444,7 +446,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// Io,
/// },
/// };
/// # use kernel::io::Mmio;
/// # use kernel::io::{Mmio, Region};
///
/// register! {
/// FIXED_REG(u32) @ 0x100 {
@ -453,7 +455,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// }
/// }
///
/// # fn test(io: &Mmio<0x1000>) {
/// # fn test(io: Mmio<'_, Region<0x1000>>) {
/// let val = io.read(FIXED_REG);
///
/// // Write from an already-existing value.
@ -557,7 +559,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// Io,
/// },
/// };
/// # use kernel::io::Mmio;
/// # use kernel::io::{Mmio, Region};
///
/// // Type used to identify the base.
/// pub struct CpuCtlBase;
@ -582,7 +584,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// }
/// }
///
/// # fn test(io: Mmio<0x1000>) {
/// # fn test(io: Mmio<'_, Region<0x1000>>) {
/// // Read the status of `Cpu0`.
/// let cpu0_started = io.read(CPU_CTL::of::<Cpu0>());
///
@ -599,7 +601,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// }
/// }
///
/// # fn test2(io: Mmio<0x1000>) {
/// # fn test2(io: Mmio<'_, Region<0x1000>>) {
/// // Start the aliased `CPU0`, leaving its other fields untouched.
/// io.update(CPU_CTL_ALIAS::of::<Cpu0>(), |r| r.with_alias_start(true));
/// # }
@ -636,7 +638,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// Io,
/// },
/// };
/// # use kernel::io::Mmio;
/// # use kernel::io::{Mmio, Region};
/// # fn get_scratch_idx() -> usize {
/// # 0x15
/// # }
@ -649,7 +651,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// }
/// }
///
/// # fn test(io: &Mmio<0x1000>)
/// # fn test(io: Mmio<'_, Region<0x1000>>)
/// # -> Result<(), Error>{
/// // Read scratch register 0, i.e. I/O address `0x80`.
/// let scratch_0 = io.read(SCRATCH::at(0)).value();
@ -722,7 +724,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// Io,
/// },
/// };
/// # use kernel::io::Mmio;
/// # use kernel::io::{Mmio, Region};
/// # fn get_scratch_idx() -> usize {
/// # 0x15
/// # }
@ -750,7 +752,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// }
/// }
///
/// # fn test(io: &Mmio<0x1000>) -> Result<(), Error> {
/// # fn test(io: Mmio<'_, Region<0x1000>>) -> Result<(), Error> {
/// // Read scratch register 0 of CPU0.
/// let scratch = io.read(CPU_SCRATCH::of::<Cpu0>().at(0));
///
@ -792,7 +794,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
/// }
/// }
///
/// # fn test2(io: &Mmio<0x1000>) -> Result<(), Error> {
/// # fn test2(io: Mmio<'_, Region<0x1000>>) -> Result<(), Error> {
/// let cpu0_status = io.read(CPU_FIRMWARE_STATUS::of::<Cpu0>()).status();
/// # Ok(())
/// # }

Some files were not shown because too many files have changed in this diff Show More