mirror of
https://github.com/torvalds/linux.git
synced 2026-09-13 15:40:03 +02:00
DRM Rust changes for v7.3-rc1
- I/O (shared from driver-core tree via signed tag rust-io-7.3-rc1):
- Rework of I/O types: make I/O regions typed (with a
dynamically-sized Region type for the existing untyped case), create
view types representing subregions of a mapped I/O region, and add
io_project!() for safely creating subviews.
- Split Io into a base trait (IoBase) and an extension trait (Io) with
a blanket implementation, preventing implementers from overriding
provided methods that unsafe code relies on.
- Add a SysMem backend for shared system memory with volatile access,
and make Coherent implement Io via an I/O view type. Add copying
methods (memcpy_{from,to}io).
- Replace dma_read!/dma_write! with io_read!/io_write!; drop the old
macros.
- DRM:
- RegistrationGuard and RegistrationData:
- Rework DeviceContext typestates: rename Uninit to Normal, add an
Ioctl context, restrict AlwaysRefCounted to Normal for both Device
and GEM Object, and establish a Deref chain from Registered to
Normal.
- Introduce RegistrationGuard, a guard representing a
drm_dev_enter/exit SRCU critical section that proves the DRM
device is registered, which implies the parent bus device is still
bound.
- Add RegistrationData as a GAT on drm::Driver. The data does not
outlive driver unbind, so it can capture lifetime-annotated device
resources and references. Accessible through the guard via a
closure with HRTB lifetime.
- Wrap ioctl dispatch in RegistrationGuard (returning ENODEV if
unplugged) and pass registration data to handlers.
- Add Driver::ParentDevice associated type.
- Fix unbounded lifetimes in ioctl handler arguments.
- Fix a race in drm_dev_register() where a partial failure allowed
in-flight ioctls to proceed while the error path tore down
resources.
- GEM shmem: add DmaResvGuard helper, vmap functions, and sg_table()
accessor.
- GPUVM: require Send + Sync for the driver's associated data,
implement Send and Sync for GpuVaAlloc and GpuVmBo, add SmContext
lifetime bound, update DriverGpuVm for DeviceContext.
- Nova:
- nova-core / nova-drm cross-crate dependency:
- Build nova-core and nova-drm from drivers/gpu/Makefile for build
ordering, export nova-core Rust symbols for nova-drm. Workaround
until the build system supports Rust cross-crate dependencies
natively.
- GSP boot process consolidation:
- Introduce GspBootContext to bundle common boot parameters,
replacing per-argument threading. Separate context and GPU
lifetimes to support mutable borrows of GPU subdevices.
- Turn FWSEC execution into a HAL method, make FWSEC bootloader
usage a property of the TU102 HAL (GA102+ gets its own instance
with it disabled). Move firmware file selection to the GSP HAL.
- Store the Fsp instance in Gpu (lifetime tied to the GPU, not just
a single boot invocation). Move GSP state and unload bundle into a
pinned subobject for reliable teardown on partial init failure.
- Boot GSP with vGPU enabled:
- Add PRC (Product Reconfiguration Control) protocol to query device
configuration from the FSP. Read vGPU mode, detect and store vGPU
state.
- Set RMSetSriovMode registry entry and reserve the larger WPR2 heap
required when vGPU is enabled.
- Build SetRegistry entries dynamically.
- TLV firmware image format:
- Add a TLV (type-length-value) parser for the new firmware image
format. TLV files use unversioned filenames with a .tlv suffix,
start with "NVFW" magic, and contain tagged blocks with 4-byte
aligned payloads.
- Transition all firmware loading (booter, gsp, gen_bootloader, fsp)
to TLV images.
- Note: this requires a development firmware not in linux-firmware
[1]; this is temporary and serves the transition to r615.
- Hopper/Blackwell fixes and cleanups:
- Correct FRTS vidmem offset calculation, split FbLayout into FSP
and non-FSP versions, fix Blackwell flush address composition, use
absolute FBHUB0 flush registers on Blackwell, use correct sysmem
flush registers on Hopper.
- Harden FSP messaging: limit receive allocation size, catch bogus
queue pointers, ensure DMA allocation lifetimes for FMC boot and
LibOS, wait for RISC-V HALTED on unload.
- I/O projection adoption:
- Use io_project!() for PTE array, message queues, and Falcon DMA
transfer bounds checking.
- Misc:
- Keep unloading if FWSEC-SB fails during Turing/Ampere GSP reset.
- Don't declare booter firmware for FSP chipsets.
- Fix packed registry table size.
- Extract and display usable FB regions from GSP.
- Store bar and dev directly in Falcon, simplifying the API.
- Parse VBIOS structs via zerocopy.
- Convert to kernel bitfield macro, remove local one.
- Move register definitions into sub-modules.
- Add FSP and PRC protocol documentation.
- Tyr:
- Firmware loading and MCU boot:
- Add a generic slot manager for dynamically allocating limited
hardware slots to software seats, with lazy eviction under
contention.
- Add MMU support wrapping the slot manager for address-space slot
allocation, with MAIR-to-MEMATTR translation.
- Add GPU virtual memory (VM) support using drm_gpuvm with ARM64
LPAE Stage 1 page tables and 4KB/2MB page sizes.
- Add a kernel buffer object type for internal driver allocations.
- Add a parser for the Mali CSF firmware binary format.
- Add MCU booting: load, parse, and map firmware sections into VM,
then boot the MCU at probe().
- Cross-subsystem:
- Add faux::Device type with AsBusDevice support. Allow retrieving a
bound Device from a Registration.
- Add device lifetime to IoPageTable.
- Add Vec::zeroed method.
- Add firmware::request_into_buf() to load firmware into a
caller-provided buffer.
- Rename dma_handle to dma_address in the DMA abstraction.
- Change pci_sriov_get_totalvfs() return type to unsigned int; add
Rust helper.
[1] https://github.com/ttabi/linux-firmware-nova
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQS2q/xV6QjXAdC7k+1FlHeO1qrKLgUCandOTAAKCRBFlHeO1qrK
LgWSAP4wHxhEuOme55bgZkne0XMums8bLln69N/UR+Rim+agJQD7BmgsL0ANxGOu
Csnsxej/tcktyraoy/QGHMjYX7pcGwM=
=UXCM
-----END PGP SIGNATURE-----
Merge tag 'drm-rust-next-2026-08-08' of https://gitlab.freedesktop.org/drm/rust/kernel into drm-next
DRM Rust changes for v7.3-rc1
- I/O (shared from driver-core tree via signed tag rust-io-7.3-rc1):
- Rework of I/O types: make I/O regions typed (with a
dynamically-sized Region type for the existing untyped case), create
view types representing subregions of a mapped I/O region, and add
io_project!() for safely creating subviews.
- Split Io into a base trait (IoBase) and an extension trait (Io) with
a blanket implementation, preventing implementers from overriding
provided methods that unsafe code relies on.
- Add a SysMem backend for shared system memory with volatile access,
and make Coherent implement Io via an I/O view type. Add copying
methods (memcpy_{from,to}io).
- Replace dma_read!/dma_write! with io_read!/io_write!; drop the old
macros.
- DRM:
- RegistrationGuard and RegistrationData:
- Rework DeviceContext typestates: rename Uninit to Normal, add an
Ioctl context, restrict AlwaysRefCounted to Normal for both Device
and GEM Object, and establish a Deref chain from Registered to
Normal.
- Introduce RegistrationGuard, a guard representing a
drm_dev_enter/exit SRCU critical section that proves the DRM
device is registered, which implies the parent bus device is still
bound.
- Add RegistrationData as a GAT on drm::Driver. The data does not
outlive driver unbind, so it can capture lifetime-annotated device
resources and references. Accessible through the guard via a
closure with HRTB lifetime.
- Wrap ioctl dispatch in RegistrationGuard (returning ENODEV if
unplugged) and pass registration data to handlers.
- Add Driver::ParentDevice associated type.
- Fix unbounded lifetimes in ioctl handler arguments.
- Fix a race in drm_dev_register() where a partial failure allowed
in-flight ioctls to proceed while the error path tore down
resources.
- GEM shmem: add DmaResvGuard helper, vmap functions, and sg_table()
accessor.
- GPUVM: require Send + Sync for the driver's associated data,
implement Send and Sync for GpuVaAlloc and GpuVmBo, add SmContext
lifetime bound, update DriverGpuVm for DeviceContext.
- Nova:
- nova-core / nova-drm cross-crate dependency:
- Build nova-core and nova-drm from drivers/gpu/Makefile for build
ordering, export nova-core Rust symbols for nova-drm. Workaround
until the build system supports Rust cross-crate dependencies
natively.
- GSP boot process consolidation:
- Introduce GspBootContext to bundle common boot parameters,
replacing per-argument threading. Separate context and GPU
lifetimes to support mutable borrows of GPU subdevices.
- Turn FWSEC execution into a HAL method, make FWSEC bootloader
usage a property of the TU102 HAL (GA102+ gets its own instance
with it disabled). Move firmware file selection to the GSP HAL.
- Store the Fsp instance in Gpu (lifetime tied to the GPU, not just
a single boot invocation). Move GSP state and unload bundle into a
pinned subobject for reliable teardown on partial init failure.
- Boot GSP with vGPU enabled:
- Add PRC (Product Reconfiguration Control) protocol to query device
configuration from the FSP. Read vGPU mode, detect and store vGPU
state.
- Set RMSetSriovMode registry entry and reserve the larger WPR2 heap
required when vGPU is enabled.
- Build SetRegistry entries dynamically.
- TLV firmware image format:
- Add a TLV (type-length-value) parser for the new firmware image
format. TLV files use unversioned filenames with a .tlv suffix,
start with "NVFW" magic, and contain tagged blocks with 4-byte
aligned payloads.
- Transition all firmware loading (booter, gsp, gen_bootloader, fsp)
to TLV images.
- Note: this requires a development firmware not in linux-firmware
[1]; this is temporary and serves the transition to r615.
- Hopper/Blackwell fixes and cleanups:
- Correct FRTS vidmem offset calculation, split FbLayout into FSP
and non-FSP versions, fix Blackwell flush address composition, use
absolute FBHUB0 flush registers on Blackwell, use correct sysmem
flush registers on Hopper.
- Harden FSP messaging: limit receive allocation size, catch bogus
queue pointers, ensure DMA allocation lifetimes for FMC boot and
LibOS, wait for RISC-V HALTED on unload.
- I/O projection adoption:
- Use io_project!() for PTE array, message queues, and Falcon DMA
transfer bounds checking.
- Misc:
- Keep unloading if FWSEC-SB fails during Turing/Ampere GSP reset.
- Don't declare booter firmware for FSP chipsets.
- Fix packed registry table size.
- Extract and display usable FB regions from GSP.
- Store bar and dev directly in Falcon, simplifying the API.
- Parse VBIOS structs via zerocopy.
- Convert to kernel bitfield macro, remove local one.
- Move register definitions into sub-modules.
- Add FSP and PRC protocol documentation.
- Tyr:
- Firmware loading and MCU boot:
- Add a generic slot manager for dynamically allocating limited
hardware slots to software seats, with lazy eviction under
contention.
- Add MMU support wrapping the slot manager for address-space slot
allocation, with MAIR-to-MEMATTR translation.
- Add GPU virtual memory (VM) support using drm_gpuvm with ARM64
LPAE Stage 1 page tables and 4KB/2MB page sizes.
- Add a kernel buffer object type for internal driver allocations.
- Add a parser for the Mali CSF firmware binary format.
- Add MCU booting: load, parse, and map firmware sections into VM,
then boot the MCU at probe().
- Cross-subsystem:
- Add faux::Device type with AsBusDevice support. Allow retrieving a
bound Device from a Registration.
- Add device lifetime to IoPageTable.
- Add Vec::zeroed method.
- Add firmware::request_into_buf() to load firmware into a
caller-provided buffer.
- Rename dma_handle to dma_address in the DMA abstraction.
- Change pci_sriov_get_totalvfs() return type to unsigned int; add
Rust helper.
[1] https://github.com/ttabi/linux-firmware-nova
Signed-off-by: Dave Airlie <airlied@redhat.com>
From: "Danilo Krummrich" <dakr@kernel.org>
Link: https://patch.msgid.link/DKJQQUOS0PVO.3JPR3MYK4PDVZ@kernel.org
This commit is contained in:
commit
29daacacf4
142
Documentation/gpu/nova/core/fsp.rst
Normal file
142
Documentation/gpu/nova/core/fsp.rst
Normal file
|
|
@ -0,0 +1,142 @@
|
|||
.. SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
===================================================
|
||||
FSP (Foundation Security Processor) and Secure Boot
|
||||
===================================================
|
||||
This document describes the role of the FSP in the GPU boot sequence on
|
||||
Hopper and Blackwell GPUs, and how it differs from the earlier Ampere boot
|
||||
flow. It also provides a brief overview of the PRC (Product Reconfiguration
|
||||
Control) protocol used to query device configuration through FSP. As with
|
||||
other documents in this directory, the information is subject to change and
|
||||
is intended to help developers understand the corresponding kernel code.
|
||||
|
||||
What is FSP?
|
||||
============
|
||||
The Foundation Security Processor (FSP) is the GPU's Internal Root of Trust
|
||||
(IROT). It is a dedicated security processor that boots from immutable ROM
|
||||
(Boot ROM) inside the GPU and is responsible for establishing the Chain of
|
||||
Trust before any other firmware is allowed to run.
|
||||
|
||||
FSP runs independently of the host CPU and starts executing as soon as the
|
||||
GPU is powered on. By the time the nova-core driver is loaded, FSP has
|
||||
already completed its own secure boot and is ready to accept commands from
|
||||
the driver.
|
||||
|
||||
Simplified boot flow (Hopper/Blackwell)
|
||||
=======================================
|
||||
Starting with Hopper, the boot flow is significantly simplified compared to
|
||||
earlier GPU generations like Ampere.
|
||||
|
||||
On an **Ampere** GPU, the boot verification chain involves multiple Falcon
|
||||
engines and multiple ucode stages (see falcon.rst for details)::
|
||||
|
||||
Hardware BROM (SEC2)
|
||||
-> HS Booter (SEC2)
|
||||
-> LS GSP-RM (GSP)
|
||||
|
||||
The driver must extract ucode from VBIOS, manage SEC2 and GSP, and
|
||||
orchestrate the Booter to load GSP-RM. This involves FWSEC-FRTS, devinit,
|
||||
and the Booter stages.
|
||||
|
||||
On **Hopper/Blackwell** GPUs, FSP replaces this multi-stage process with a
|
||||
single message-driven interface::
|
||||
|
||||
FSP (hardware root of trust, boots from ROM)
|
||||
-> FMC (Falcon Microcontroller, verified by FSP)
|
||||
-> GSP-RM (verified and loaded by FMC)
|
||||
|
||||
The driver only needs to:
|
||||
|
||||
1. Wait for FSP to complete its own secure boot (polling a scratch register).
|
||||
2. Send a Chain of Trust (COT) message to FSP with the FMC firmware location,
|
||||
cryptographic signatures, and GSP boot parameters.
|
||||
3. FSP authenticates the FMC firmware and boots it, FMC in turn loads GSP-RM.
|
||||
|
||||
There is no SEC2 involvement, no Booter ucode, and no FWSEC-FRTS stage. The
|
||||
entire secure boot is driven by a single FSP message exchange.
|
||||
|
||||
Chain of Trust (COT) protocol
|
||||
=============================
|
||||
The Chain of Trust establishes a cryptographically enforced boot sequence,
|
||||
ensuring the GPU reaches a known, trusted state.
|
||||
|
||||
The driver communicates with FSP using a message queue (Falcon MSGQ
|
||||
interface). Each message consists of an MCTP (Management Component Transport
|
||||
Protocol) transport header and an NVDM (NVIDIA Vendor Defined Message) header,
|
||||
followed by a protocol-specific payload.
|
||||
|
||||
For Chain of Trust, the payload includes:
|
||||
|
||||
- The system memory address of the FMC firmware image.
|
||||
- Cryptographic material: a SHA-384 hash, RSA-3K public key, and RSA-3K
|
||||
signature extracted from the FMC ELF firmware.
|
||||
- FRTS (Firmware Runtime Services) region information (vidmem offset and size).
|
||||
- The system memory address of the GSP boot arguments structure.
|
||||
|
||||
FSP verifies the signature against the provided public key and hash, and if
|
||||
verification succeeds, boots the FMC. The FMC then authenticates and launches
|
||||
GSP-RM.
|
||||
|
||||
The message flow is::
|
||||
|
||||
nova-core FSP
|
||||
| |
|
||||
| 1. Poll scratch register |
|
||||
| (wait for FSP boot complete) |
|
||||
| |
|
||||
| 2. COT message ------------> |
|
||||
| (FMC addr, signatures, |
|
||||
| boot params) |
|
||||
| |
|
||||
| |--- Verify FMC signature
|
||||
| |--- Boot FMC
|
||||
| |--- FMC loads GSP-RM
|
||||
| |
|
||||
| 3. COT response <------------ |
|
||||
| (success/error) |
|
||||
| |
|
||||
|
||||
FSP message format
|
||||
==================
|
||||
All FSP messages share a common header format consisting of two 32-bit words:
|
||||
|
||||
**MCTP header** (Management Component Transport Protocol):
|
||||
|
||||
- Bit 31: SOM (Start of Message)
|
||||
- Bit 30: EOM (End of Message)
|
||||
- Bits 29:28: Packet sequence number
|
||||
- Bits 23:16: Source Endpoint ID
|
||||
|
||||
**NVDM header** (NVIDIA Vendor Defined Message):
|
||||
|
||||
- Bits 6:0: MCTP message type (0x7e = vendor-defined PCI)
|
||||
- Bits 23:8: PCI vendor ID (0x10de = NVIDIA)
|
||||
- Bits 31:24: NVDM type (0x14 = COT, 0x13 = PRC, 0x15 = FSP response)
|
||||
|
||||
PRC (Product Reconfiguration Control) protocol
|
||||
===============================================
|
||||
PRC is an API system exposed through FSP's Management Partition that allows
|
||||
querying and modifying device configuration without firmware updates.
|
||||
|
||||
Configuration parameters are called "knobs". Each knob has a unique object
|
||||
ID and controls a specific device behavior. Examples include vGPU mode, ECC
|
||||
enable, confidential computing mode, and NVLINK configuration.
|
||||
|
||||
Each knob has two values:
|
||||
|
||||
- **Active**: the currently effective value for this boot cycle.
|
||||
- **Persistent**: the value stored in InfoROM, applied on subsequent boots.
|
||||
|
||||
The nova-core driver uses PRC to read the vGPU mode knob (object ID 0x29)
|
||||
during early boot, before firmware loading, to determine whether the GPU
|
||||
should operate in vGPU mode.
|
||||
|
||||
The PRC message format follows the same MCTP/NVDM header structure as COT,
|
||||
with NVDM type 0x13. The payload contains:
|
||||
|
||||
- A sub-command (e.g., 0x0c for read).
|
||||
- Flags indicating which value to read (bit 0 = persistent, bit 1 = active).
|
||||
- The knob object ID.
|
||||
|
||||
The response includes the common FSP response header (with error status)
|
||||
followed by the knob's 16-bit state value.
|
||||
184
Documentation/gpu/nova/core/tlv.rst
Normal file
184
Documentation/gpu/nova/core/tlv.rst
Normal file
|
|
@ -0,0 +1,184 @@
|
|||
.. SPDX-License-Identifier: (GPL-2.0+ OR MIT)
|
||||
|
||||
==================================
|
||||
TLV Tags in Nova Firmware Images
|
||||
==================================
|
||||
|
||||
Nova firmware images use a Type-Length-Value (TLV) format to encapsulate
|
||||
firmware components and metadata. The TLV file begins with a 4-byte "magic"
|
||||
header that contains the string "NVFW". Following the header is a sequence of
|
||||
TLV blocks.
|
||||
|
||||
Each block consists of a 4-byte tag of ASCII characters, a 4-byte length
|
||||
encoded as a little-endian unsigned integer, and a sequence of bytes, the size
|
||||
of which is equal to the length rounded up to the next multiple of 4.
|
||||
|
||||
The driver code that reads the TLV and uses its contents is called the parser.
|
||||
It is the responsibility of the parser to handle missing or malformed tags,
|
||||
lengths, and values in the TLV.
|
||||
|
||||
::
|
||||
|
||||
+------+------+------+------+
|
||||
| 'N' | 'V' | 'F' | 'W' | Magic header
|
||||
+------+------+------+------+
|
||||
| Tag (4 bytes, ASCII) | TLV block 0
|
||||
+---------------------------+
|
||||
| Length (4 bytes, LE) |
|
||||
+---------------------------+
|
||||
| |
|
||||
| Value (length bytes, |
|
||||
| padded to 4-byte align) |
|
||||
| |
|
||||
+---------------------------+
|
||||
| Tag (4 bytes, ASCII) | TLV block 1
|
||||
+---------------------------+
|
||||
| Length (4 bytes, LE) |
|
||||
+---------------------------+
|
||||
| |
|
||||
| Value (length bytes, |
|
||||
| padded to 4-byte align) |
|
||||
| |
|
||||
+---------------------------+
|
||||
| ... | More TLV blocks
|
||||
+---------------------------+
|
||||
|
||||
Tags and Length
|
||||
===============
|
||||
TLV tags are always four-character words, with all letters being upper case.
|
||||
Duplicate tags are not allowed.
|
||||
|
||||
A TLV file may contain additional tags not described in this document.
|
||||
|
||||
Values
|
||||
======
|
||||
Values are one of four types. The type is not encoded in the format; rather,
|
||||
the parser expects a given tag to have a value of a given type.
|
||||
|
||||
1) Integers, encoded in 32-bit or 64-bit little-endian format.
|
||||
2) Strings, encoded as-is and required to be only printable ASCII characters
|
||||
and without a null terminator.
|
||||
3) An array of bytes, for binary data.
|
||||
4) Boolean, encoded as single byte, with a value of 0 for False or 1 for True.
|
||||
|
||||
Common Tags
|
||||
===========
|
||||
These tags are shared across firmware types and carry the same meaning
|
||||
wherever they appear. Unlike the firmware-specific tags below, a common tag
|
||||
is reserved: its meaning is fixed and may never be redefined for a particular
|
||||
firmware type.
|
||||
|
||||
``VERS`` (string)
|
||||
Human-readable firmware version string. Present in all TLV files.
|
||||
|
||||
A TLV image must contain either a single ``BLOB`` tag (firmware embedded
|
||||
inline) or a ``SIZE``/``FILE`` pair (firmware stored in a separate file).
|
||||
|
||||
``BLOB`` (bytes)
|
||||
If the firmware microcode binary is stored in the TLV, this tag contains
|
||||
the actual firmware image bytes.
|
||||
|
||||
``FILE`` (string)
|
||||
If the firmware binary is stored as a separate file, this tag contains the
|
||||
name of that file, which is required to be in the same directory as the TLV,
|
||||
so no paths are allowed in the filename. This tag is always paired with
|
||||
``SIZE``, so as to allow the driver to pre-allocate the buffer before
|
||||
loading the file.
|
||||
|
||||
``SIZE`` (u32)
|
||||
Total size in bytes of the firmware image to be loaded from the companion
|
||||
file named by ``FILE``. This tag is mandatory if ``FILE`` exists, so the
|
||||
size of the firmware image must be known when the TLV is created. If the
|
||||
firmware image is updated and its size changes, then the TLV must be
|
||||
updated with it.
|
||||
|
||||
GSP Firmware Tags
|
||||
=================
|
||||
``SIGN`` (bytes)
|
||||
Cryptographic signature for the GSP firmware.
|
||||
|
||||
``BLID`` (string)
|
||||
The build ID, extracted from the ".note.gnu.build-id" section.
|
||||
|
||||
Booter Firmware Tags
|
||||
====================
|
||||
``DAOF`` (u32) - ``os_data_offset``
|
||||
OS data section offset within the firmware image (absolute byte offset).
|
||||
Maps to the DMEM load source.
|
||||
|
||||
``DASZ`` (u32) - ``os_data_size``
|
||||
OS data section size in bytes.
|
||||
|
||||
``CDOF`` (u32) - ``os_code_offset``
|
||||
OS code section offset within the firmware image (absolute byte offset).
|
||||
Maps to the non-secure IMEM load source.
|
||||
|
||||
``CDSZ`` (u32) - ``os_code_size``
|
||||
OS code section size in bytes.
|
||||
|
||||
``PLOC`` (u32) - ``patch_loc``
|
||||
Signature patch location -- byte offset within the firmware image where the
|
||||
selected signature should be written.
|
||||
|
||||
``FUSE`` (u32) - ``fuse_version``
|
||||
Fuse version of the firmware, used with the hardware fuse register to
|
||||
select the correct signature index.
|
||||
|
||||
``ENID`` (u32) - ``engine_id``
|
||||
Engine ID mask identifying the falcon engine this firmware targets.
|
||||
|
||||
``UCID`` (u32) - ``ucode_id``
|
||||
Microcode ID used together with the engine ID to query hardware signature
|
||||
fuse registers.
|
||||
|
||||
``A0CO`` (u32) - ``app0_code_offset``
|
||||
App0 code offset -- start of the secure code region within the firmware
|
||||
image. Used as the IMEM secure section source.
|
||||
|
||||
``A0CS`` (u32) - ``app0_code_size``
|
||||
App0 code size in bytes.
|
||||
|
||||
``NSIG`` (u32) - ``num_sigs``
|
||||
Number of signatures included in the ``SIGN`` tag.
|
||||
|
||||
``SIGN`` (bytes)
|
||||
Concatenated array of firmware signatures. The size of each signature is
|
||||
the total length of the ``SIGN`` value divided by ``NSIG``. The correct
|
||||
signature is selected using the fuse-version-derived index.
|
||||
|
||||
Generic Bootloader Tags
|
||||
=======================
|
||||
``CDSZ`` (u32) - ``code_size``
|
||||
Size in bytes of the bootloader code to copy from the ``BLOB`` tag and
|
||||
PIO-load into falcon IMEM.
|
||||
|
||||
``STRT`` (u32) - ``start_tag``
|
||||
Start tag identifying the IMEM block where execution begins. The falcon
|
||||
boot address is derived as ``start_tag << 8``.
|
||||
|
||||
GSP Bootloader Tags
|
||||
===================
|
||||
``CDOF`` (u32) - ``code_offset``
|
||||
Offset within the firmware image at which the code section starts.
|
||||
|
||||
``DAOF`` (u32) - ``data_offset``
|
||||
Offset within the firmware image at which the data section starts.
|
||||
|
||||
``MFOF`` (u32) - ``manifest_offset``
|
||||
Offset within the firmware image at which the manifest starts.
|
||||
|
||||
``APPV`` (u32) - ``app_version``
|
||||
Application version of the firmware.
|
||||
|
||||
FMC Firmware Tags
|
||||
=================
|
||||
``HASH`` (bytes)
|
||||
SHA-384 hash of the FMC firmware, exactly 48 bytes long.
|
||||
|
||||
``PKEY`` (bytes)
|
||||
Public key used to verify the FMC firmware. At most 384 bytes (RSA-3072),
|
||||
but may be shorter.
|
||||
|
||||
``SIGN`` (bytes)
|
||||
Signature of the FMC firmware. At most 384 bytes (RSA-3072), but may
|
||||
be shorter.
|
||||
|
|
@ -30,5 +30,7 @@ vGPU manager VFIO driver and the nova-drm driver.
|
|||
core/todo
|
||||
core/vbios
|
||||
core/devinit
|
||||
core/fsp
|
||||
core/fwsec
|
||||
core/falcon
|
||||
core/tlv
|
||||
|
|
|
|||
|
|
@ -7,4 +7,60 @@ obj-$(CONFIG_GPU_BUDDY) += buddy.o
|
|||
obj-y += host1x/ drm/ vga/ tests/
|
||||
obj-$(CONFIG_IMX_IPUV3_CORE) += ipu-v3/
|
||||
obj-$(CONFIG_TRACE_GPU_MEM) += trace/
|
||||
obj-$(CONFIG_NOVA_CORE) += nova-core/
|
||||
|
||||
# nova-core and nova-drm are built from this Makefile so nova-drm's dependency
|
||||
# on nova-core can be expressed as a plain Make prerequisite rather than a
|
||||
# recursive sub-make. This is a temporary workaround until the Rust build
|
||||
# system supports cross-crate dependencies natively.
|
||||
|
||||
obj-$(CONFIG_NOVA_CORE) += nova-core.o
|
||||
nova-core-y := nova-core/nova_core.o nova-core/nova_core_exports.o
|
||||
|
||||
obj-$(CONFIG_DRM_NOVA) += nova-drm.o
|
||||
nova-drm-y := drm/nova/nova.o
|
||||
|
||||
# Export Rust symbols from nova-core only if nova-drm actually references them.
|
||||
nova-core-export-deps := $(if $(CONFIG_DRM_NOVA),$(obj)/drm/nova/nova.o)
|
||||
|
||||
rust_needed_exports = \
|
||||
{ $(if $(strip $(2)),$(NM) -u $(2);,) echo "__DEFINED_RUST_SYMBOLS__"; \
|
||||
$(NM) -p --defined-only $(1); } | \
|
||||
awk -v fmt='$(3)' ' \
|
||||
/^__DEFINED_RUST_SYMBOLS__$$/ { defs = 1; next } \
|
||||
!defs { if ($$NF ~ /^_R/) needed[$$NF] = 1; next } \
|
||||
defs && $$2 ~ /(T|R|D|B)/ && $$3 ~ /^_R/ && \
|
||||
$$3 !~ /_(init|cleanup)_module$$/ && \
|
||||
$$3 !~ /__(pfx|cfi|odr_asan)/ && \
|
||||
$$3 in needed { printf fmt, $$3 } \
|
||||
'
|
||||
|
||||
quiet_cmd_exports = EXPORTS $@
|
||||
cmd_exports = \
|
||||
$(call rust_needed_exports,$<,$(nova-core-export-deps),EXPORT_SYMBOL_RUST_GPL(%s);\n) > $@
|
||||
|
||||
$(obj)/nova-core/exports_nova_core_generated.h: $(obj)/nova-core/nova_core.o $(nova-core-export-deps) FORCE
|
||||
$(call if_changed,exports)
|
||||
|
||||
targets += nova-core/exports_nova_core_generated.h
|
||||
|
||||
$(obj)/nova-core/nova_core_exports.o: $(obj)/nova-core/exports_nova_core_generated.h
|
||||
CFLAGS_nova-core/nova_core_exports.o := -I $(objtree)/$(obj)/nova-core
|
||||
|
||||
ifdef CONFIG_MODVERSIONS
|
||||
# The C export shim declares Rust symbols as `extern int`, so reuse its export
|
||||
# list but generate symbol CRCs from the Rust object instead of the shim's DWARF.
|
||||
$(obj)/nova-core/nova_core_exports.o: private cmd_gensymtypes_c = \
|
||||
$(call getexportsymbols,\1) | \
|
||||
$(objtree)/scripts/gendwarfksyms/gendwarfksyms \
|
||||
$(if $(KBUILD_GENDWARFKSYMS_STABLE), --stable) \
|
||||
$(if $(KBUILD_SYMTYPES), --symtypes $(@:.o=.symtypes),) \
|
||||
$(obj)/nova-core/nova_core.o
|
||||
endif
|
||||
|
||||
# Output nova-core's crate metadata for use by nova-drm at compile time.
|
||||
RUSTFLAGS_nova-core/nova_core.o += \
|
||||
--emit=metadata=$(objtree)/$(obj)/nova-core/libnova_core.rmeta
|
||||
|
||||
# Allow nova-drm to import nova-core's types.
|
||||
$(obj)/drm/nova/nova.o: $(obj)/nova-core/nova_core.o
|
||||
RUSTFLAGS_drm/nova/nova.o := -L $(objtree)/$(obj)/nova-core --extern nova_core
|
||||
|
|
|
|||
|
|
@ -186,7 +186,7 @@ obj-$(CONFIG_DRM_VMWGFX)+= vmwgfx/
|
|||
obj-$(CONFIG_DRM_VGEM) += vgem/
|
||||
obj-$(CONFIG_DRM_VKMS) += vkms/
|
||||
obj-$(CONFIG_DRM_NOUVEAU) +=nouveau/
|
||||
obj-$(CONFIG_DRM_NOVA) += nova/
|
||||
# nova-drm is built from drivers/gpu/Makefile together with nova-core.
|
||||
obj-$(CONFIG_DRM_EXYNOS) +=exynos/
|
||||
obj-$(CONFIG_DRM_ROCKCHIP) +=rockchip/
|
||||
obj-$(CONFIG_DRM_GMA500) += gma500/
|
||||
|
|
|
|||
|
|
@ -475,6 +475,22 @@ void drm_dev_exit(int idx)
|
|||
}
|
||||
EXPORT_SYMBOL(drm_dev_exit);
|
||||
|
||||
/*
|
||||
* Mark the device as unplugged and wait for any in-flight drm_dev_enter()
|
||||
* critical sections to complete.
|
||||
*/
|
||||
static void drm_dev_synchronize_unplug(struct drm_device *dev)
|
||||
{
|
||||
/*
|
||||
* After synchronizing any critical read section is guaranteed to see
|
||||
* the new value of ->unplugged, and any critical section which might
|
||||
* still have seen the old value of ->unplugged is guaranteed to have
|
||||
* finished.
|
||||
*/
|
||||
dev->unplugged = true;
|
||||
synchronize_srcu(&drm_unplug_srcu);
|
||||
}
|
||||
|
||||
/**
|
||||
* drm_dev_unplug - unplug a DRM device
|
||||
* @dev: DRM device
|
||||
|
|
@ -487,15 +503,7 @@ EXPORT_SYMBOL(drm_dev_exit);
|
|||
*/
|
||||
void drm_dev_unplug(struct drm_device *dev)
|
||||
{
|
||||
/*
|
||||
* After synchronizing any critical read section is guaranteed to see
|
||||
* the new value of ->unplugged, and any critical section which might
|
||||
* still have seen the old value of ->unplugged is guaranteed to have
|
||||
* finished.
|
||||
*/
|
||||
dev->unplugged = true;
|
||||
synchronize_srcu(&drm_unplug_srcu);
|
||||
|
||||
drm_dev_synchronize_unplug(dev);
|
||||
drm_dev_unregister(dev);
|
||||
|
||||
/* Clear all CPU mappings pointing to this device */
|
||||
|
|
@ -1095,6 +1103,7 @@ int drm_dev_register(struct drm_device *dev, unsigned long flags)
|
|||
goto err_minors;
|
||||
|
||||
dev->registered = true;
|
||||
dev->unplugged = false;
|
||||
|
||||
if (driver->load) {
|
||||
ret = driver->load(dev, flags);
|
||||
|
|
@ -1122,6 +1131,13 @@ int drm_dev_register(struct drm_device *dev, unsigned long flags)
|
|||
if (dev->driver->unload)
|
||||
dev->driver->unload(dev);
|
||||
err_minors:
|
||||
/*
|
||||
* If a minor was registered before the failure, userspace could have
|
||||
* opened it and entered a drm_dev_enter() critical section. Ensure all
|
||||
* such sections complete before we clean up.
|
||||
*/
|
||||
drm_dev_synchronize_unplug(dev);
|
||||
|
||||
remove_compat_control_link(dev);
|
||||
drm_minor_unregister(dev, DRM_MINOR_ACCEL);
|
||||
drm_minor_unregister(dev, DRM_MINOR_PRIMARY);
|
||||
|
|
|
|||
|
|
@ -1,4 +1,3 @@
|
|||
# SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
obj-$(CONFIG_DRM_NOVA) += nova-drm.o
|
||||
nova-drm-y := nova.o
|
||||
# nova-drm is built from drivers/gpu/Makefile.
|
||||
# nova.o (rust-analyzer marker - DO NOT REMOVE).
|
||||
|
|
|
|||
|
|
@ -2,7 +2,10 @@
|
|||
|
||||
use kernel::{
|
||||
auxiliary,
|
||||
device::Core,
|
||||
device::{
|
||||
Core,
|
||||
DeviceContext, //
|
||||
},
|
||||
drm::{
|
||||
self,
|
||||
gem,
|
||||
|
|
@ -17,18 +20,14 @@
|
|||
|
||||
pub(crate) struct NovaDriver;
|
||||
|
||||
pub(crate) struct Nova {
|
||||
pub(crate) struct Nova<'bound> {
|
||||
#[expect(unused)]
|
||||
drm: ARef<drm::Device<NovaDriver>>,
|
||||
_reg: drm::Registration<'bound, NovaDriver>,
|
||||
}
|
||||
|
||||
/// Convienence type alias for the DRM device type for this driver
|
||||
pub(crate) type NovaDevice<Ctx = drm::Registered> = drm::Device<NovaDriver, Ctx>;
|
||||
|
||||
#[pin_data]
|
||||
pub(crate) struct NovaData {
|
||||
pub(crate) adev: ARef<auxiliary::Device>,
|
||||
}
|
||||
pub(crate) type NovaDevice<Ctx = drm::Normal> = drm::Device<NovaDriver, Ctx>;
|
||||
|
||||
const INFO: drm::DriverInfo = drm::DriverInfo {
|
||||
major: 0,
|
||||
|
|
@ -53,27 +52,32 @@ pub(crate) struct NovaData {
|
|||
|
||||
impl auxiliary::Driver for NovaDriver {
|
||||
type IdInfo = ();
|
||||
type Data<'bound> = Nova;
|
||||
type Data<'bound> = Nova<'bound>;
|
||||
const ID_TABLE: auxiliary::IdTable<Self::IdInfo> = &AUX_TABLE;
|
||||
|
||||
fn probe<'bound>(
|
||||
adev: &'bound auxiliary::Device<Core<'_>>,
|
||||
_info: &'bound Self::IdInfo,
|
||||
) -> impl PinInit<Self::Data<'bound>, Error> + 'bound {
|
||||
let data = try_pin_init!(NovaData { adev: adev.into() });
|
||||
let drm = drm::UnregisteredDevice::<Self>::new(adev, Ok(()))?;
|
||||
// SAFETY: `reg` is stored in `Nova` and dropped when the driver is unbound; it is
|
||||
// never forgotten.
|
||||
let reg = unsafe { drm::Registration::new(adev.as_ref(), drm, (), 0)? };
|
||||
|
||||
let drm = drm::UnregisteredDevice::<Self>::new(adev.as_ref(), data)?;
|
||||
let drm = drm::Registration::new_foreign_owned(drm, adev.as_ref(), 0)?;
|
||||
|
||||
Ok(Nova { drm: drm.into() })
|
||||
Ok(Nova {
|
||||
drm: reg.device().into(),
|
||||
_reg: reg,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
#[vtable]
|
||||
impl drm::Driver for NovaDriver {
|
||||
type Data = NovaData;
|
||||
type Data = ();
|
||||
type RegistrationData<'a> = ();
|
||||
type File = File;
|
||||
type Object<Ctx: drm::DeviceContext> = gem::Object<NovaObject, Ctx>;
|
||||
type Object = gem::Object<NovaObject>;
|
||||
type ParentDevice<Ctx: DeviceContext> = auxiliary::Device<Ctx>;
|
||||
|
||||
const INFO: drm::DriverInfo = INFO;
|
||||
|
||||
|
|
|
|||
|
|
@ -4,7 +4,13 @@
|
|||
use crate::gem::NovaObject;
|
||||
use kernel::{
|
||||
alloc::flags::*,
|
||||
drm::{self, gem::BaseObject},
|
||||
auxiliary,
|
||||
device::Bound,
|
||||
drm::{
|
||||
self,
|
||||
gem::BaseObject,
|
||||
Registered, //
|
||||
},
|
||||
pci,
|
||||
prelude::*,
|
||||
uapi,
|
||||
|
|
@ -23,13 +29,13 @@ fn open(_dev: &NovaDevice) -> Result<Pin<KBox<Self>>> {
|
|||
impl File {
|
||||
/// IOCTL: get_param: Query GPU / driver metadata.
|
||||
pub(crate) fn get_param(
|
||||
dev: &NovaDevice,
|
||||
dev: &NovaDevice<Registered>,
|
||||
_reg_data: &(),
|
||||
getparam: &mut uapi::drm_nova_getparam,
|
||||
_file: &drm::File<File>,
|
||||
) -> Result<u32> {
|
||||
let adev = &dev.adev;
|
||||
let parent = adev.parent();
|
||||
let pdev: &pci::Device = parent.try_into()?;
|
||||
let adev: &auxiliary::Device<Bound> = dev.as_ref();
|
||||
let pdev: &pci::Device<Bound> = adev.parent().try_into()?;
|
||||
|
||||
let value = match getparam.param as u32 {
|
||||
uapi::NOVA_GETPARAM_VRAM_BAR_SIZE => pdev.resource_len(1)?,
|
||||
|
|
@ -43,7 +49,8 @@ pub(crate) fn get_param(
|
|||
|
||||
/// IOCTL: gem_create: Create a new DRM GEM object.
|
||||
pub(crate) fn gem_create(
|
||||
dev: &NovaDevice,
|
||||
dev: &NovaDevice<Registered>,
|
||||
_reg_data: &(),
|
||||
req: &mut uapi::drm_nova_gem_create,
|
||||
file: &drm::File<File>,
|
||||
) -> Result<u32> {
|
||||
|
|
@ -56,7 +63,8 @@ pub(crate) fn gem_create(
|
|||
|
||||
/// IOCTL: gem_info: Query GEM metadata.
|
||||
pub(crate) fn gem_info(
|
||||
_dev: &NovaDevice,
|
||||
_dev: &NovaDevice<Registered>,
|
||||
_reg_data: &(),
|
||||
req: &mut uapi::drm_nova_gem_info,
|
||||
file: &drm::File<File>,
|
||||
) -> Result<u32> {
|
||||
|
|
|
|||
|
|
@ -2,7 +2,10 @@
|
|||
|
||||
use kernel::{
|
||||
drm,
|
||||
drm::{gem, gem::BaseObject, DeviceContext},
|
||||
drm::{
|
||||
gem,
|
||||
gem::BaseObject, //
|
||||
},
|
||||
page,
|
||||
prelude::*,
|
||||
sync::aref::ARef,
|
||||
|
|
@ -21,27 +24,20 @@ impl gem::DriverObject for NovaObject {
|
|||
type Driver = NovaDriver;
|
||||
type Args = ();
|
||||
|
||||
fn new<Ctx: DeviceContext>(
|
||||
_dev: &NovaDevice<Ctx>,
|
||||
_size: usize,
|
||||
_args: Self::Args,
|
||||
) -> impl PinInit<Self, Error> {
|
||||
fn new(_dev: &NovaDevice, _size: usize, _args: Self::Args) -> impl PinInit<Self, Error> {
|
||||
try_pin_init!(NovaObject {})
|
||||
}
|
||||
}
|
||||
|
||||
impl NovaObject {
|
||||
/// Create a new DRM GEM object.
|
||||
pub(crate) fn new<Ctx: DeviceContext>(
|
||||
dev: &NovaDevice<Ctx>,
|
||||
size: usize,
|
||||
) -> Result<ARef<gem::Object<Self, Ctx>>> {
|
||||
pub(crate) fn new(dev: &NovaDevice, size: usize) -> Result<ARef<gem::Object<Self>>> {
|
||||
if size == 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
let aligned_size = page::page_align(size).ok_or(EINVAL)?;
|
||||
|
||||
gem::Object::<Self, Ctx>::new(dev, aligned_size, ())
|
||||
gem::Object::<Self>::new(dev, aligned_size, ())
|
||||
}
|
||||
|
||||
/// Look up a GEM object handle for a `File` and return an `ObjectRef` for it.
|
||||
|
|
|
|||
|
|
@ -5,10 +5,15 @@ config DRM_TYR
|
|||
depends on DRM=y
|
||||
depends on RUST
|
||||
depends on ARM || ARM64 || COMPILE_TEST
|
||||
depends on MMU
|
||||
depends on !GENERIC_ATOMIC64 # for IOMMU_IO_PGTABLE_LPAE
|
||||
depends on COMMON_CLK
|
||||
depends on IOMMU_SUPPORT
|
||||
default n
|
||||
select IOMMU_IO_PGTABLE_LPAE
|
||||
select RUST_DRM_GEM_SHMEM_HELPER
|
||||
select RUST_DRM_GPUVM
|
||||
select RUST_FW_LOADER_ABSTRACTIONS
|
||||
help
|
||||
Rust DRM driver for ARM Mali CSF-based GPUs.
|
||||
|
||||
|
|
|
|||
|
|
@ -6,8 +6,10 @@
|
|||
OptionalClk, //
|
||||
},
|
||||
device::{
|
||||
Bound,
|
||||
Core,
|
||||
Device, //
|
||||
Device,
|
||||
DeviceContext, //
|
||||
},
|
||||
dma::{
|
||||
Device as DmaDevice,
|
||||
|
|
@ -27,7 +29,7 @@
|
|||
regulator::Regulator,
|
||||
sizes::SZ_2M,
|
||||
sync::{
|
||||
aref::ARef,
|
||||
Arc,
|
||||
Mutex, //
|
||||
},
|
||||
time, //
|
||||
|
|
@ -35,9 +37,11 @@
|
|||
|
||||
use crate::{
|
||||
file::TyrDrmFileData,
|
||||
gem::BoData,
|
||||
fw::Firmware,
|
||||
gem::Bo,
|
||||
gpu,
|
||||
gpu::GpuInfo,
|
||||
mmu::Mmu,
|
||||
regs::gpu_control::*, //
|
||||
};
|
||||
|
||||
|
|
@ -46,18 +50,26 @@
|
|||
pub(crate) struct TyrDrmDriver;
|
||||
|
||||
/// Convenience type alias for the DRM device type for this driver.
|
||||
pub(crate) type TyrDrmDevice<Ctx = drm::Registered> = drm::Device<TyrDrmDriver, Ctx>;
|
||||
pub(crate) type TyrDrmDevice<Ctx = drm::Normal> = drm::Device<TyrDrmDriver, Ctx>;
|
||||
|
||||
pub(crate) struct TyrPlatformDriver;
|
||||
|
||||
#[pin_data(PinnedDrop)]
|
||||
pub(crate) struct TyrPlatformDriverData {
|
||||
_device: ARef<TyrDrmDevice>,
|
||||
pub(crate) struct TyrPlatformDriverData<'bound> {
|
||||
_reg: drm::Registration<'bound, TyrDrmDriver>,
|
||||
}
|
||||
|
||||
/// Data owned by the DRM [`Registration`].
|
||||
///
|
||||
/// This data can have references tied to the parent platform device binding scope
|
||||
/// and is accessible only while the DRM device is registered with userspace.
|
||||
#[pin_data]
|
||||
pub(crate) struct TyrDrmDeviceData {
|
||||
pub(crate) pdev: ARef<platform::Device>,
|
||||
pub(crate) struct TyrDrmRegistrationData<'drm> {
|
||||
/// Parent platform device.
|
||||
pub(crate) pdev: &'drm platform::Device<Bound>,
|
||||
|
||||
/// Firmware sections.
|
||||
pub(crate) fw: Firmware<'drm>,
|
||||
|
||||
#[pin]
|
||||
clks: Mutex<Clocks>,
|
||||
|
|
@ -65,9 +77,10 @@ pub(crate) struct TyrDrmDeviceData {
|
|||
#[pin]
|
||||
regulators: Mutex<Regulators>,
|
||||
|
||||
/// Some information on the GPU.
|
||||
///
|
||||
/// This is mainly queried by userspace, i.e.: Mesa.
|
||||
/// GPU MMIO register mapping.
|
||||
pub(crate) iomem: Arc<IoMem<'drm>>,
|
||||
|
||||
/// GPU information read from hardware during probe.
|
||||
pub(crate) gpu_info: GpuInfo,
|
||||
}
|
||||
|
||||
|
|
@ -97,7 +110,7 @@ fn issue_soft_reset(dev: &Device, iomem: &IoMem<'_>) -> Result {
|
|||
|
||||
impl platform::Driver for TyrPlatformDriver {
|
||||
type IdInfo = ();
|
||||
type Data<'bound> = TyrPlatformDriverData;
|
||||
type Data<'bound> = TyrPlatformDriverData<'bound>;
|
||||
const OF_ID_TABLE: Option<of::IdTable<Self::IdInfo>> = Some(&OF_TABLE);
|
||||
|
||||
fn probe<'bound>(
|
||||
|
|
@ -116,7 +129,8 @@ fn probe<'bound>(
|
|||
let sram_regulator = Regulator::<regulator::Enabled>::get(pdev.as_ref(), c"sram")?;
|
||||
|
||||
let request = pdev.io_request_by_index(0).ok_or(ENODEV)?;
|
||||
let iomem = request.iomap_sized::<SZ_2M>()?;
|
||||
|
||||
let iomem = Arc::new(request.iomap_sized::<SZ_2M>()?, GFP_KERNEL)?;
|
||||
|
||||
issue_soft_reset(pdev.as_ref(), &iomem)?;
|
||||
gpu::l2_power_on(pdev.as_ref(), &iomem)?;
|
||||
|
|
@ -132,10 +146,23 @@ fn probe<'bound>(
|
|||
// other threads of execution.
|
||||
unsafe { pdev.dma_set_mask_and_coherent(DmaMask::try_new(pa_bits)?)? };
|
||||
|
||||
let platform: ARef<platform::Device> = pdev.into();
|
||||
let unreg_dev = drm::UnregisteredDevice::<TyrDrmDriver>::new(pdev, Ok(()))?;
|
||||
|
||||
let data = try_pin_init!(TyrDrmDeviceData {
|
||||
pdev: platform.clone(),
|
||||
let mmu = Mmu::new(pdev.as_ref(), iomem.as_arc_borrow(), &gpu_info)?;
|
||||
|
||||
let firmware = Firmware::new(
|
||||
pdev.as_ref(),
|
||||
iomem.clone(),
|
||||
&unreg_dev,
|
||||
mmu.as_arc_borrow(),
|
||||
&gpu_info,
|
||||
)?;
|
||||
|
||||
firmware.boot()?;
|
||||
|
||||
let reg_data = pin_init!(TyrDrmRegistrationData {
|
||||
pdev,
|
||||
fw: firmware,
|
||||
clks <- new_mutex!(Clocks {
|
||||
core: core_clk,
|
||||
stacks: stacks_clk,
|
||||
|
|
@ -145,25 +172,23 @@ fn probe<'bound>(
|
|||
_mali: mali_regulator,
|
||||
_sram: sram_regulator,
|
||||
}),
|
||||
iomem,
|
||||
gpu_info,
|
||||
});
|
||||
|
||||
let tdev = drm::UnregisteredDevice::<TyrDrmDriver>::new(pdev.as_ref(), data)?;
|
||||
let tdev = drm::driver::Registration::new_foreign_owned(tdev, pdev.as_ref(), 0)?;
|
||||
// SAFETY: `reg` is stored in `TyrPlatformDriverData` and dropped when the driver is
|
||||
// unbound; it is never forgotten.
|
||||
let reg = unsafe { drm::Registration::new(pdev.as_ref(), unreg_dev, reg_data, 0)? };
|
||||
|
||||
let driver = TyrPlatformDriverData {
|
||||
_device: tdev.into(),
|
||||
};
|
||||
let driver = TyrPlatformDriverData { _reg: reg };
|
||||
|
||||
// We need this to be dev_info!() because dev_dbg!() does not work at
|
||||
// all in Rust for now, and we need to see whether probe succeeded.
|
||||
dev_info!(pdev, "Tyr initialized correctly.\n");
|
||||
dev_dbg!(pdev, "Tyr initialized correctly.");
|
||||
Ok(driver)
|
||||
}
|
||||
}
|
||||
|
||||
#[pinned_drop]
|
||||
impl PinnedDrop for TyrPlatformDriverData {
|
||||
impl PinnedDrop for TyrPlatformDriverData<'_> {
|
||||
fn drop(self: Pin<&mut Self>) {}
|
||||
}
|
||||
|
||||
|
|
@ -179,9 +204,11 @@ fn drop(self: Pin<&mut Self>) {}
|
|||
|
||||
#[vtable]
|
||||
impl drm::Driver for TyrDrmDriver {
|
||||
type Data = TyrDrmDeviceData;
|
||||
type Data = ();
|
||||
type RegistrationData<'drm> = TyrDrmRegistrationData<'drm>;
|
||||
type File = TyrDrmFileData;
|
||||
type Object<R: drm::DeviceContext> = drm::gem::shmem::Object<BoData, R>;
|
||||
type Object = Bo;
|
||||
type ParentDevice<Ctx: DeviceContext> = platform::Device<Ctx>;
|
||||
|
||||
const INFO: drm::DriverInfo = INFO;
|
||||
const FEAT_RENDER: bool = true;
|
||||
|
|
|
|||
|
|
@ -1,7 +1,10 @@
|
|||
// SPDX-License-Identifier: GPL-2.0 or MIT
|
||||
|
||||
use kernel::{
|
||||
drm,
|
||||
drm::{
|
||||
self,
|
||||
Registered, //
|
||||
},
|
||||
prelude::*,
|
||||
uaccess::UserSlice,
|
||||
uapi, //
|
||||
|
|
@ -9,7 +12,8 @@
|
|||
|
||||
use crate::driver::{
|
||||
TyrDrmDevice,
|
||||
TyrDrmDriver, //
|
||||
TyrDrmDriver,
|
||||
TyrDrmRegistrationData, //
|
||||
};
|
||||
|
||||
#[pin_data]
|
||||
|
|
@ -28,14 +32,15 @@ fn open(_dev: &drm::Device<Self::Driver>) -> Result<Pin<KBox<Self>>> {
|
|||
|
||||
impl TyrDrmFileData {
|
||||
pub(crate) fn dev_query(
|
||||
ddev: &TyrDrmDevice,
|
||||
_ddev: &TyrDrmDevice<Registered>,
|
||||
reg_data: &TyrDrmRegistrationData<'_>,
|
||||
devquery: &mut uapi::drm_panthor_dev_query,
|
||||
_file: &TyrDrmFile,
|
||||
) -> Result<u32> {
|
||||
if devquery.pointer == 0 {
|
||||
match devquery.type_ {
|
||||
uapi::drm_panthor_dev_query_type_DRM_PANTHOR_DEV_QUERY_GPU_INFO => {
|
||||
devquery.size = core::mem::size_of_val(&ddev.gpu_info) as u32;
|
||||
devquery.size = core::mem::size_of_val(®_data.gpu_info) as u32;
|
||||
Ok(0)
|
||||
}
|
||||
_ => Err(EINVAL),
|
||||
|
|
@ -49,7 +54,7 @@ pub(crate) fn dev_query(
|
|||
)
|
||||
.writer();
|
||||
|
||||
writer.write(&ddev.gpu_info)?;
|
||||
writer.write(®_data.gpu_info)?;
|
||||
|
||||
Ok(0)
|
||||
}
|
||||
|
|
|
|||
321
drivers/gpu/drm/tyr/fw.rs
Normal file
321
drivers/gpu/drm/tyr/fw.rs
Normal file
|
|
@ -0,0 +1,321 @@
|
|||
// SPDX-License-Identifier: GPL-2.0 or MIT
|
||||
|
||||
//! Firmware loading and management for Mali CSF GPU.
|
||||
//!
|
||||
//! This module handles loading the Mali GPU firmware binary, parsing it into sections,
|
||||
//! and mapping those sections into the MCU's virtual address space. Each firmware section
|
||||
//! has specific properties (read/write/execute permissions, cache modes) and must be loaded
|
||||
//! at specific virtual addresses expected by the MCU.
|
||||
//!
|
||||
//! See [`Firmware`] for the main firmware management interface and [`Section`] for
|
||||
//! individual firmware sections.
|
||||
//!
|
||||
//! [`Firmware`]: crate::fw::Firmware
|
||||
//! [`Section`]: crate::fw::Section
|
||||
|
||||
use kernel::{
|
||||
device::{
|
||||
Bound,
|
||||
Device, //
|
||||
},
|
||||
drm::{
|
||||
gem::BaseObject, //
|
||||
},
|
||||
io::{
|
||||
poll,
|
||||
Io, //
|
||||
},
|
||||
num::Bounded,
|
||||
prelude::*,
|
||||
register,
|
||||
str::CString,
|
||||
sync::{
|
||||
Arc,
|
||||
ArcBorrow, //
|
||||
},
|
||||
time, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::{
|
||||
IoMem,
|
||||
TyrDrmDevice, //
|
||||
},
|
||||
fw::parser::{
|
||||
FwParser,
|
||||
ParsedSection, //
|
||||
},
|
||||
gem,
|
||||
gem::{
|
||||
KernelBo,
|
||||
KernelBoVaAlloc, //
|
||||
},
|
||||
gpu::GpuInfo,
|
||||
|
||||
mmu::Mmu,
|
||||
regs::{
|
||||
gpu_control::{
|
||||
McuControlMode,
|
||||
McuStatus,
|
||||
GPU_ID,
|
||||
MCU_CONTROL,
|
||||
MCU_STATUS, //
|
||||
}, //
|
||||
job_control::{
|
||||
JOB_IRQ_CLEAR,
|
||||
JOB_IRQ_RAWSTAT, //
|
||||
}, //
|
||||
},
|
||||
vm::Vm, //
|
||||
};
|
||||
|
||||
mod parser;
|
||||
|
||||
pub(super) const CSF_MCU_SHARED_REGION_START: u32 = 0x04000000;
|
||||
|
||||
#[derive(Copy, Clone, Debug, PartialEq, Eq)]
|
||||
#[repr(u8)]
|
||||
pub(super) enum CacheMode {
|
||||
None = 0,
|
||||
Cached = 1,
|
||||
UncachedCoherent = 2,
|
||||
CachedCoherent = 3,
|
||||
}
|
||||
|
||||
impl From<Bounded<u32, 2>> for CacheMode {
|
||||
fn from(value: Bounded<u32, 2>) -> Self {
|
||||
match value.get() {
|
||||
0 => Self::None,
|
||||
1 => Self::Cached,
|
||||
2 => Self::UncachedCoherent,
|
||||
3 => Self::CachedCoherent,
|
||||
_ => unreachable!(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl From<CacheMode> for Bounded<u32, 2> {
|
||||
fn from(value: CacheMode) -> Self {
|
||||
Bounded::try_new(value as u32).unwrap()
|
||||
}
|
||||
}
|
||||
|
||||
register! {
|
||||
#[allow(non_upper_case_globals)]
|
||||
pub(super) SectionFlags(u32) @ 0x0 {
|
||||
0:0 read => bool;
|
||||
1:1 write => bool;
|
||||
2:2 exec => bool;
|
||||
4:3 cache_mode => CacheMode;
|
||||
5:5 prot => bool;
|
||||
30:30 shared => bool;
|
||||
31:31 zero => bool;
|
||||
}
|
||||
}
|
||||
|
||||
impl SectionFlags {
|
||||
const VALID_MASK: u32 = Self::READ_MASK
|
||||
| Self::WRITE_MASK
|
||||
| Self::EXEC_MASK
|
||||
| Self::CACHE_MODE_MASK
|
||||
| Self::PROT_MASK
|
||||
| Self::SHARED_MASK
|
||||
| Self::ZERO_MASK;
|
||||
|
||||
fn try_from_fw(value: u32) -> Result<Self> {
|
||||
if value & !Self::VALID_MASK != 0 {
|
||||
Err(EINVAL)
|
||||
} else {
|
||||
Ok(Self::from_raw(value))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A parsed section of the firmware binary.
|
||||
struct Section<'drm> {
|
||||
// Raw firmware section data for reset purposes
|
||||
#[expect(dead_code)]
|
||||
data: KVec<u8>,
|
||||
|
||||
// Keep the BO backing this firmware section so that both the
|
||||
// GPU mapping and CPU mapping remain valid until the Section is dropped.
|
||||
#[expect(dead_code)]
|
||||
mem: gem::KernelBo<'drm>,
|
||||
}
|
||||
|
||||
/// Loaded firmware with sections mapped into MCU VM.
|
||||
pub(crate) struct Firmware<'drm> {
|
||||
/// Iomem need to access registers.
|
||||
iomem: Arc<IoMem<'drm>>,
|
||||
|
||||
/// MCU VM.
|
||||
vm: Arc<Vm<'drm>>,
|
||||
|
||||
/// List of firmware sections.
|
||||
#[expect(dead_code)]
|
||||
sections: KVec<Section<'drm>>,
|
||||
}
|
||||
|
||||
impl<'drm> Drop for Firmware<'drm> {
|
||||
fn drop(&mut self) {
|
||||
// Stop the MCU before releasing its firmware mappings and memory.
|
||||
let _ = self.stop();
|
||||
|
||||
// AS slots retain a VM ref, we need to kill the circular ref manually.
|
||||
self.vm.kill();
|
||||
}
|
||||
}
|
||||
|
||||
impl<'drm> Firmware<'drm> {
|
||||
fn init_section_mem(dev: &Device, mem: &mut KernelBo<'drm>, data: &KVec<u8>) -> Result {
|
||||
if data.is_empty() {
|
||||
return Ok(());
|
||||
}
|
||||
|
||||
let vmap = mem.bo().vmap::<0>()?;
|
||||
let size = mem.bo().size();
|
||||
|
||||
if data.len() > size {
|
||||
dev_err!(dev, "fw section {} bigger than BO {}", data.len(), size);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
for (i, &byte) in data.iter().enumerate() {
|
||||
vmap.try_write8(byte, i)?;
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn request(ddev: &TyrDrmDevice, gpu_info: &GpuInfo) -> Result<kernel::firmware::Firmware> {
|
||||
let gpu_id = GPU_ID::from_raw(gpu_info.gpu_id);
|
||||
|
||||
let path = CString::try_from_fmt(fmt!(
|
||||
"arm/mali/arch{}.{}/mali_csffw.bin",
|
||||
gpu_id.arch_major().get(),
|
||||
gpu_id.arch_minor().get()
|
||||
))?;
|
||||
|
||||
kernel::firmware::Firmware::request(&path, ddev.as_ref().as_ref())
|
||||
}
|
||||
|
||||
fn load(
|
||||
dev: &Device,
|
||||
ddev: &TyrDrmDevice,
|
||||
gpu_info: &GpuInfo,
|
||||
) -> Result<(kernel::firmware::Firmware, KVec<ParsedSection>)> {
|
||||
let fw = Self::request(ddev, gpu_info)?;
|
||||
let mut parser = FwParser::new(dev, fw.data());
|
||||
|
||||
let parsed_sections = parser.parse()?;
|
||||
|
||||
Ok((fw, parsed_sections))
|
||||
}
|
||||
|
||||
/// Load firmware and map sections into MCU VM.
|
||||
pub(crate) fn new(
|
||||
dev: &'drm Device<Bound>,
|
||||
iomem: Arc<IoMem<'drm>>,
|
||||
ddev: &TyrDrmDevice,
|
||||
mmu: ArcBorrow<'_, Mmu<'drm>>,
|
||||
gpu_info: &GpuInfo,
|
||||
) -> Result<Firmware<'drm>> {
|
||||
let vm = Vm::new(dev, ddev, mmu, gpu_info)?;
|
||||
vm.activate()?;
|
||||
|
||||
let result = (|| {
|
||||
let (fw, parsed_sections) = Self::load(dev, ddev, gpu_info)?;
|
||||
let mut sections = KVec::new();
|
||||
for parsed in parsed_sections {
|
||||
let size = u64::from(parsed.va.end.checked_sub(parsed.va.start).ok_or(EINVAL)?);
|
||||
|
||||
let va = u64::from(parsed.va.start);
|
||||
|
||||
let mut mem = KernelBo::new(
|
||||
ddev,
|
||||
vm.clone(),
|
||||
size,
|
||||
KernelBoVaAlloc::Explicit(va),
|
||||
parsed.vm_map_flags,
|
||||
)?;
|
||||
|
||||
let section_start = parsed.data_range.start as usize;
|
||||
let section_end = parsed.data_range.end as usize;
|
||||
let mut data = KVec::new();
|
||||
|
||||
// Ensure that the firmware slice is not out of bounds.
|
||||
let fw_data = fw.data();
|
||||
let bytes = fw_data.get(section_start..section_end).ok_or(EINVAL)?;
|
||||
data.extend_from_slice(bytes, GFP_KERNEL)?;
|
||||
|
||||
Self::init_section_mem(dev, &mut mem, &data)?;
|
||||
|
||||
sections.push(Section { data, mem }, GFP_KERNEL)?;
|
||||
}
|
||||
|
||||
Ok(Firmware {
|
||||
iomem,
|
||||
vm: vm.clone(),
|
||||
sections,
|
||||
})
|
||||
})();
|
||||
|
||||
if result.is_err() {
|
||||
vm.kill();
|
||||
}
|
||||
|
||||
result
|
||||
}
|
||||
|
||||
pub(crate) fn boot(&self) -> Result {
|
||||
let io = &self.iomem;
|
||||
|
||||
// Discard any stale global interrupt.
|
||||
io.write_reg(JOB_IRQ_CLEAR::zeroed().with_glb(true));
|
||||
|
||||
io.write_reg(MCU_CONTROL::zeroed().with_req(McuControlMode::Auto));
|
||||
|
||||
if let Err(e) = poll::read_poll_timeout(
|
||||
|| Ok((io.read(MCU_STATUS), io.read(JOB_IRQ_RAWSTAT))),
|
||||
|(mcu_status, irq_rawstat)| {
|
||||
mcu_status.value() == McuStatus::Enabled && irq_rawstat.glb()
|
||||
},
|
||||
time::Delta::from_millis(1),
|
||||
time::Delta::from_millis(100),
|
||||
) {
|
||||
let status = io.read(MCU_STATUS);
|
||||
dev_err!(
|
||||
self.vm.dev(),
|
||||
"MCU failed to boot, status: {:?}",
|
||||
status.value()
|
||||
);
|
||||
return Err(e);
|
||||
}
|
||||
|
||||
io.write_reg(JOB_IRQ_CLEAR::zeroed().with_glb(true));
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn stop(&self) -> Result {
|
||||
let io = &self.iomem;
|
||||
io.write_reg(MCU_CONTROL::zeroed().with_req(McuControlMode::Disable));
|
||||
|
||||
if let Err(e) = poll::read_poll_timeout(
|
||||
|| Ok(io.read(MCU_STATUS)),
|
||||
|status| status.value() == McuStatus::Disabled,
|
||||
time::Delta::from_micros(10),
|
||||
time::Delta::from_millis(100),
|
||||
) {
|
||||
let status = io.read(MCU_STATUS);
|
||||
dev_err!(
|
||||
self.vm.dev(),
|
||||
"MCU failed to stop, status: {:?}",
|
||||
status.value()
|
||||
);
|
||||
return Err(e);
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
588
drivers/gpu/drm/tyr/fw/parser.rs
Normal file
588
drivers/gpu/drm/tyr/fw/parser.rs
Normal file
|
|
@ -0,0 +1,588 @@
|
|||
// SPDX-License-Identifier: GPL-2.0 or MIT
|
||||
|
||||
//! Firmware binary parser for Mali CSF (Command Stream Frontend) GPU.
|
||||
//!
|
||||
//! This module implements a parser for the Mali GPU firmware binary format. The firmware
|
||||
//! file contains a header followed by a sequence of entries, each describing how to load
|
||||
//! firmware sections into the MCU (Microcontroller Unit) memory. The parser extracts section
|
||||
//! metadata including:
|
||||
//! - Virtual address ranges where sections should be mapped
|
||||
//! - Data ranges (byte offsets) within the firmware binary
|
||||
//! - Section flags (permissions, cache modes)
|
||||
|
||||
use core::{
|
||||
mem::size_of,
|
||||
ops::Range, //
|
||||
};
|
||||
|
||||
use kernel::{
|
||||
bits::bit_u32,
|
||||
device::Device,
|
||||
prelude::*,
|
||||
sizes::SZ_4K, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
fw::{
|
||||
CacheMode,
|
||||
SectionFlags,
|
||||
CSF_MCU_SHARED_REGION_START, //
|
||||
},
|
||||
vm::{
|
||||
VmFlag,
|
||||
VmMapFlags, //
|
||||
}, //
|
||||
};
|
||||
|
||||
/// A parsed firmware section ready for loading into MCU memory.
|
||||
///
|
||||
/// Represents a single firmware section extracted from the firmware binary, containing
|
||||
/// all information needed to map the section's data into the MCU's virtual address space.
|
||||
pub(super) struct ParsedSection {
|
||||
/// Byte offset range within the firmware binary where this section's data resides.
|
||||
pub(super) data_range: Range<u32>,
|
||||
/// MCU virtual address range where this section should be mapped.
|
||||
pub(super) va: Range<u32>,
|
||||
/// Memory protection and caching flags for the mapping.
|
||||
pub(super) vm_map_flags: VmMapFlags,
|
||||
}
|
||||
|
||||
/// A bare-bones `std::io::Cursor<[u8]>` clone to keep track of the current position in the
|
||||
/// firmware binary.
|
||||
///
|
||||
/// Provides methods to sequentially read primitive types and byte arrays from the firmware
|
||||
/// binary while maintaining the current read position.
|
||||
struct Cursor<'a> {
|
||||
dev: &'a Device,
|
||||
data: &'a [u8],
|
||||
pos: usize,
|
||||
}
|
||||
|
||||
impl<'a> Cursor<'a> {
|
||||
fn new(dev: &'a Device, data: &'a [u8]) -> Self {
|
||||
Self { dev, data, pos: 0 }
|
||||
}
|
||||
|
||||
fn len(&self) -> usize {
|
||||
self.data.len()
|
||||
}
|
||||
|
||||
fn pos(&self) -> usize {
|
||||
self.pos
|
||||
}
|
||||
|
||||
/// Returns a view into the cursor's data.
|
||||
///
|
||||
/// This spawns a new cursor, leaving the current cursor unchanged.
|
||||
fn view(&self, range: Range<usize>) -> Result<Cursor<'_>> {
|
||||
if range.start < self.pos || range.end > self.data.len() {
|
||||
dev_err!(
|
||||
self.dev,
|
||||
"Invalid cursor range {:?} for data of length {}",
|
||||
range,
|
||||
self.data.len()
|
||||
);
|
||||
|
||||
Err(EINVAL)
|
||||
} else {
|
||||
Ok(Self {
|
||||
dev: self.dev,
|
||||
data: &self.data[range],
|
||||
pos: 0,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// Reads a slice of bytes from the current position and advances the cursor.
|
||||
///
|
||||
/// Returns an error if the read would exceed the data bounds.
|
||||
fn read(&mut self, nbytes: usize) -> Result<&[u8]> {
|
||||
let start = self.pos;
|
||||
let end = start + nbytes;
|
||||
|
||||
if end > self.data.len() {
|
||||
dev_err!(
|
||||
self.dev,
|
||||
"Invalid firmware file: read of size {} at position {} is out of bounds",
|
||||
nbytes,
|
||||
start,
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
self.pos += nbytes;
|
||||
Ok(&self.data[start..end])
|
||||
}
|
||||
|
||||
/// Reads a little-endian `u8` from the current position and advances the cursor.
|
||||
fn read_u8(&mut self) -> Result<u8> {
|
||||
let bytes = self.read(size_of::<u8>())?;
|
||||
Ok(bytes[0])
|
||||
}
|
||||
|
||||
/// Reads a little-endian `u16` from the current position and advances the cursor.
|
||||
fn read_u16(&mut self) -> Result<u16> {
|
||||
let bytes: [u8; 2] = self
|
||||
.read(size_of::<u16>())?
|
||||
.try_into()
|
||||
.map_err(|_| EINVAL)?;
|
||||
|
||||
Ok(u16::from_le_bytes(bytes))
|
||||
}
|
||||
|
||||
/// Reads a little-endian `u32` from the current position and advances the cursor.
|
||||
fn read_u32(&mut self) -> Result<u32> {
|
||||
let bytes: [u8; 4] = self
|
||||
.read(size_of::<u32>())?
|
||||
.try_into()
|
||||
.map_err(|_| EINVAL)?;
|
||||
|
||||
Ok(u32::from_le_bytes(bytes))
|
||||
}
|
||||
|
||||
/// Advances the cursor position by the specified number of bytes.
|
||||
///
|
||||
/// Returns an error if the advance would exceed the data bounds.
|
||||
fn advance(&mut self, nbytes: usize) -> Result {
|
||||
if self.pos + nbytes > self.data.len() {
|
||||
dev_err!(
|
||||
self.dev,
|
||||
"Invalid firmware file: advance of size {} at position {} is out of bounds",
|
||||
nbytes,
|
||||
self.pos,
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
self.pos += nbytes;
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
/// Parser for Mali CSF GPU firmware binaries.
|
||||
///
|
||||
/// Parses the firmware binary format, extracting section metadata including virtual
|
||||
/// address ranges, data offsets, and memory protection flags needed to load firmware
|
||||
/// into the MCU's memory.
|
||||
pub(super) struct FwParser<'a> {
|
||||
cursor: Cursor<'a>,
|
||||
}
|
||||
|
||||
impl<'a> FwParser<'a> {
|
||||
/// Creates a new firmware parser for the given firmware binary data.
|
||||
pub(super) fn new(dev: &'a Device, data: &'a [u8]) -> Self {
|
||||
Self {
|
||||
cursor: Cursor::new(dev, data),
|
||||
}
|
||||
}
|
||||
|
||||
/// Parses the firmware binary and returns a collection of parsed sections.
|
||||
///
|
||||
/// This method validates the firmware header and iterates through all entries
|
||||
/// in the binary, extracting section information needed for loading.
|
||||
pub(super) fn parse(&mut self) -> Result<KVec<ParsedSection>> {
|
||||
let fw_header = self.parse_fw_header()?;
|
||||
let header_end = fw_header.size as usize;
|
||||
|
||||
let mut parsed_sections = KVec::new();
|
||||
while self.cursor.pos() < header_end {
|
||||
let entry_section = self.parse_entry(header_end)?;
|
||||
|
||||
if let Some(inner) = entry_section.inner {
|
||||
parsed_sections.push(inner, GFP_KERNEL)?;
|
||||
}
|
||||
}
|
||||
|
||||
if parsed_sections.is_empty() {
|
||||
dev_err!(self.cursor.dev, "Firmware contains no loadable sections");
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
Ok(parsed_sections)
|
||||
}
|
||||
|
||||
fn parse_fw_header(&mut self) -> Result<FirmwareHeader> {
|
||||
let fw_header: FirmwareHeader = match FirmwareHeader::new(&mut self.cursor) {
|
||||
Ok(fw_header) => fw_header,
|
||||
Err(e) => {
|
||||
dev_err!(self.cursor.dev, "Invalid firmware file: {}", e.to_errno());
|
||||
return Err(e);
|
||||
}
|
||||
};
|
||||
|
||||
if fw_header.size as usize > self.cursor.len() {
|
||||
dev_err!(self.cursor.dev, "Firmware image is truncated");
|
||||
return Err(EINVAL);
|
||||
}
|
||||
Ok(fw_header)
|
||||
}
|
||||
|
||||
fn parse_entry(&mut self, header_end: usize) -> Result<EntrySection> {
|
||||
let entry_start = self.cursor.pos();
|
||||
|
||||
let entry_header_end = entry_start
|
||||
.checked_add(size_of::<EntryHeader>())
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
if entry_header_end > header_end {
|
||||
dev_err!(
|
||||
self.cursor.dev,
|
||||
"Firmware entry header at {:#x} exceeds header region ending at {:#x}",
|
||||
entry_start,
|
||||
header_end
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let entry_section = EntrySection {
|
||||
entry_hdr: EntryHeader(self.cursor.read_u32()?),
|
||||
inner: None,
|
||||
};
|
||||
|
||||
let firmware_size = self.cursor.len();
|
||||
let entry_size = entry_section.entry_hdr.size() as usize;
|
||||
|
||||
if self.cursor.pos() % size_of::<u32>() != 0
|
||||
|| entry_size % size_of::<u32>() != 0
|
||||
|| entry_size < size_of::<EntryHeader>()
|
||||
{
|
||||
dev_err!(
|
||||
self.cursor.dev,
|
||||
"Firmware entry isn't 32 bit aligned, offset={:#x} size={:#x}",
|
||||
self.cursor.pos() - size_of::<u32>(),
|
||||
entry_size
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let entry_end = entry_start.checked_add(entry_size).ok_or(EINVAL)?;
|
||||
|
||||
if entry_end > header_end {
|
||||
dev_err!(
|
||||
self.cursor.dev,
|
||||
"Firmware entry at {:#x} extends beyond header region ending at {:#x}",
|
||||
entry_start,
|
||||
header_end
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let section_hdr_size = entry_size - size_of::<EntryHeader>();
|
||||
|
||||
let entry_section = {
|
||||
let mut entry_cursor = self.cursor.view(self.cursor.pos()..entry_end)?;
|
||||
|
||||
match entry_section.entry_hdr.entry_type() {
|
||||
Ok(EntryType::Iface) => Ok(EntrySection {
|
||||
entry_hdr: entry_section.entry_hdr,
|
||||
inner: Self::parse_section_entry(&mut entry_cursor, firmware_size)?,
|
||||
}),
|
||||
Ok(
|
||||
EntryType::Config
|
||||
| EntryType::FutfTest
|
||||
| EntryType::TraceBuffer
|
||||
| EntryType::TimelineMetadata
|
||||
| EntryType::BuildInfoMetadata,
|
||||
) => Ok(entry_section),
|
||||
|
||||
Err(_) => {
|
||||
if entry_section.entry_hdr.optional() {
|
||||
Ok(entry_section)
|
||||
} else {
|
||||
dev_err!(
|
||||
self.cursor.dev,
|
||||
"Failed to handle firmware entry type: {}",
|
||||
entry_section.entry_hdr.entry_type_raw()
|
||||
);
|
||||
Err(EINVAL)
|
||||
}
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
if entry_section.is_ok() {
|
||||
self.cursor.advance(section_hdr_size)?;
|
||||
}
|
||||
|
||||
entry_section
|
||||
}
|
||||
|
||||
fn parse_section_entry(
|
||||
entry_cursor: &mut Cursor<'_>,
|
||||
firmware_size: usize,
|
||||
) -> Result<Option<ParsedSection>> {
|
||||
let section_hdr: SectionHeader = SectionHeader::new(entry_cursor)?;
|
||||
|
||||
if section_hdr.data.end < section_hdr.data.start {
|
||||
dev_err!(
|
||||
entry_cursor.dev,
|
||||
"Firmware corrupted, data.end < data.start (0x{:x} < 0x{:x})",
|
||||
section_hdr.data.end,
|
||||
section_hdr.data.start
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
if section_hdr.data.end as usize > firmware_size {
|
||||
dev_err!(
|
||||
entry_cursor.dev,
|
||||
"Firmware data range {:#x}..{:#x} exceeds firmware size {:#x}",
|
||||
section_hdr.data.start,
|
||||
section_hdr.data.end,
|
||||
firmware_size,
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
if section_hdr.va.start as usize % SZ_4K != 0 || section_hdr.va.end as usize % SZ_4K != 0 {
|
||||
dev_err!(
|
||||
entry_cursor.dev,
|
||||
"Firmware virtual address range {:#x}..{:#x} is not page aligned",
|
||||
section_hdr.va.start,
|
||||
section_hdr.va.end
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
if section_hdr.section_flags.prot() {
|
||||
dev_dbg!(
|
||||
entry_cursor.dev,
|
||||
"Firmware protected mode entry not supported, ignoring"
|
||||
);
|
||||
return Ok(None);
|
||||
}
|
||||
|
||||
if section_hdr.va.start == CSF_MCU_SHARED_REGION_START
|
||||
&& !section_hdr.section_flags.shared()
|
||||
{
|
||||
dev_err!(
|
||||
entry_cursor.dev,
|
||||
"Interface at 0x{:x} must be shared",
|
||||
CSF_MCU_SHARED_REGION_START
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
if section_hdr.va.is_empty() {
|
||||
return Ok(None);
|
||||
}
|
||||
|
||||
let mut vm_map_flags = VmMapFlags::empty();
|
||||
|
||||
if !section_hdr.section_flags.write() {
|
||||
vm_map_flags |= VmFlag::Readonly;
|
||||
}
|
||||
|
||||
if !section_hdr.section_flags.exec() {
|
||||
vm_map_flags |= VmFlag::Noexec;
|
||||
}
|
||||
|
||||
// TODO: As in Panthor, map coherent firmware sections uncached until the VM
|
||||
// supports a coherent mapping attribute.
|
||||
if section_hdr.section_flags.cache_mode() != CacheMode::Cached {
|
||||
vm_map_flags |= VmFlag::Uncached;
|
||||
}
|
||||
|
||||
Ok(Some(ParsedSection {
|
||||
data_range: section_hdr.data.clone(),
|
||||
va: section_hdr.va,
|
||||
vm_map_flags,
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
/// Firmware binary header containing version and size information.
|
||||
///
|
||||
/// The header is located at the beginning of the firmware binary and contains
|
||||
/// a magic value for validation, version information, and the total size of
|
||||
/// all structured headers that follow.
|
||||
#[expect(dead_code)]
|
||||
struct FirmwareHeader {
|
||||
/// Magic value to check binary validity.
|
||||
magic: u32,
|
||||
|
||||
/// Minor firmware version.
|
||||
minor: u8,
|
||||
|
||||
/// Major firmware version.
|
||||
major: u8,
|
||||
|
||||
/// Padding. Must be set to zero.
|
||||
_padding1: u16,
|
||||
|
||||
/// Firmware version hash.
|
||||
version_hash: u32,
|
||||
|
||||
/// Padding. Must be set to zero.
|
||||
_padding2: u32,
|
||||
|
||||
/// Total size of all the structured data headers at beginning of firmware binary.
|
||||
size: u32,
|
||||
}
|
||||
|
||||
impl FirmwareHeader {
|
||||
const FW_BINARY_MAGIC: u32 = 0xc3f13a6e;
|
||||
const FW_BINARY_MAJOR_MAX: u8 = 0;
|
||||
|
||||
/// Reads and validates a firmware header from the cursor.
|
||||
///
|
||||
/// Verifies the magic value, version compatibility, and padding fields.
|
||||
fn new(cursor: &mut Cursor<'_>) -> Result<Self> {
|
||||
let magic = cursor.read_u32()?;
|
||||
if magic != Self::FW_BINARY_MAGIC {
|
||||
dev_err!(cursor.dev, "Invalid firmware magic");
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let minor = cursor.read_u8()?;
|
||||
let major = cursor.read_u8()?;
|
||||
|
||||
if major > Self::FW_BINARY_MAJOR_MAX {
|
||||
dev_err!(
|
||||
cursor.dev,
|
||||
"Unsupported firmware binary header version {}.{} (expected {}.x)",
|
||||
major,
|
||||
minor,
|
||||
Self::FW_BINARY_MAJOR_MAX
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let padding1 = cursor.read_u16()?;
|
||||
let version_hash = cursor.read_u32()?;
|
||||
let padding2 = cursor.read_u32()?;
|
||||
let size = cursor.read_u32()?;
|
||||
|
||||
if padding1 != 0 || padding2 != 0 {
|
||||
dev_err!(
|
||||
cursor.dev,
|
||||
"Invalid firmware file: header padding is not zero"
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let fw_header = Self {
|
||||
magic,
|
||||
minor,
|
||||
major,
|
||||
_padding1: padding1,
|
||||
version_hash,
|
||||
_padding2: padding2,
|
||||
size,
|
||||
};
|
||||
|
||||
Ok(fw_header)
|
||||
}
|
||||
}
|
||||
|
||||
/// Firmware section header for loading binary sections into MCU memory.
|
||||
#[derive(Debug)]
|
||||
struct SectionHeader {
|
||||
section_flags: SectionFlags,
|
||||
/// MCU virtual range to map this binary section to.
|
||||
va: Range<u32>,
|
||||
/// References the data in the FW binary.
|
||||
data: Range<u32>,
|
||||
}
|
||||
|
||||
impl SectionHeader {
|
||||
/// Reads and validates a section header from the cursor.
|
||||
///
|
||||
/// Parses section flags, virtual address range, and data range from the firmware binary.
|
||||
fn new(cursor: &mut Cursor<'_>) -> Result<Self> {
|
||||
let section_flags = SectionFlags::try_from_fw(cursor.read_u32()?)?;
|
||||
|
||||
let va_start = cursor.read_u32()?;
|
||||
let va_end = cursor.read_u32()?;
|
||||
|
||||
let va = va_start..va_end;
|
||||
|
||||
if va.end < va.start {
|
||||
dev_err!(
|
||||
cursor.dev,
|
||||
"Invalid firmware file: VA end precedes start at pos {}",
|
||||
cursor.pos(),
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let data_start = cursor.read_u32()?;
|
||||
let data_end = cursor.read_u32()?;
|
||||
let data = data_start..data_end;
|
||||
|
||||
Ok(Self {
|
||||
section_flags,
|
||||
va,
|
||||
data,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// A firmware entry containing a header and optional parsed section data.
|
||||
///
|
||||
/// Represents a single entry in the firmware binary, which may contain loadable
|
||||
/// section data or metadata that doesn't require loading.
|
||||
struct EntrySection {
|
||||
entry_hdr: EntryHeader,
|
||||
inner: Option<ParsedSection>,
|
||||
}
|
||||
|
||||
/// Header for a firmware entry, packed into a single u32.
|
||||
///
|
||||
/// The entry header encodes the entry type, size, and optional flag in a
|
||||
/// 32-bit value with the following layout:
|
||||
/// - Bits 0-7: Entry type
|
||||
/// - Bits 8-15: Size in bytes
|
||||
/// - Bit 31: Optional flag
|
||||
struct EntryHeader(u32);
|
||||
|
||||
impl EntryHeader {
|
||||
fn entry_type_raw(&self) -> u8 {
|
||||
(self.0 & 0xff) as u8
|
||||
}
|
||||
|
||||
fn entry_type(&self) -> Result<EntryType> {
|
||||
let v = self.entry_type_raw();
|
||||
EntryType::try_from(v)
|
||||
}
|
||||
|
||||
fn optional(&self) -> bool {
|
||||
self.0 & bit_u32(31) != 0
|
||||
}
|
||||
|
||||
fn size(&self) -> u32 {
|
||||
self.0 >> 8 & 0xff
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Clone, Copy, Debug)]
|
||||
#[repr(u8)]
|
||||
enum EntryType {
|
||||
/// Host <-> FW interface.
|
||||
Iface = 0,
|
||||
/// FW config.
|
||||
Config = 1,
|
||||
/// Unit tests.
|
||||
FutfTest = 2,
|
||||
/// Trace buffer interface.
|
||||
TraceBuffer = 3,
|
||||
/// Timeline metadata interface.
|
||||
TimelineMetadata = 4,
|
||||
/// Metadata about how the FW binary was built.
|
||||
BuildInfoMetadata = 6,
|
||||
}
|
||||
|
||||
impl TryFrom<u8> for EntryType {
|
||||
type Error = Error;
|
||||
|
||||
fn try_from(value: u8) -> Result<Self, Self::Error> {
|
||||
match value {
|
||||
0 => Ok(EntryType::Iface),
|
||||
1 => Ok(EntryType::Config),
|
||||
2 => Ok(EntryType::FutfTest),
|
||||
3 => Ok(EntryType::TraceBuffer),
|
||||
4 => Ok(EntryType::TimelineMetadata),
|
||||
6 => Ok(EntryType::BuildInfoMetadata),
|
||||
_ => Err(EINVAL),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -4,17 +4,29 @@
|
|||
//! This module provides buffer object (BO) management functionality using
|
||||
//! DRM's GEM subsystem with shmem backing.
|
||||
|
||||
use core::ops::Range;
|
||||
|
||||
use kernel::{
|
||||
drm::{
|
||||
gem,
|
||||
DeviceContext, //
|
||||
drm::gem::{
|
||||
self,
|
||||
shmem, //
|
||||
},
|
||||
prelude::*, //
|
||||
prelude::*,
|
||||
sync::{
|
||||
aref::ARef,
|
||||
Arc, //
|
||||
}, //
|
||||
};
|
||||
|
||||
use crate::driver::{
|
||||
TyrDrmDevice,
|
||||
TyrDrmDriver, //
|
||||
use crate::{
|
||||
driver::{
|
||||
TyrDrmDevice,
|
||||
TyrDrmDriver, //
|
||||
},
|
||||
vm::{
|
||||
Vm,
|
||||
VmMapFlags, //
|
||||
},
|
||||
};
|
||||
|
||||
/// Tyr's DriverObject type for GEM objects.
|
||||
|
|
@ -33,11 +45,116 @@ impl gem::DriverObject for BoData {
|
|||
type Driver = TyrDrmDriver;
|
||||
type Args = BoCreateArgs;
|
||||
|
||||
fn new<Ctx: DeviceContext>(
|
||||
_dev: &TyrDrmDevice<Ctx>,
|
||||
_size: usize,
|
||||
args: BoCreateArgs,
|
||||
) -> impl PinInit<Self, Error> {
|
||||
fn new(_dev: &TyrDrmDevice, _size: usize, args: BoCreateArgs) -> impl PinInit<Self, Error> {
|
||||
try_pin_init!(Self { flags: args.flags })
|
||||
}
|
||||
}
|
||||
|
||||
/// Type alias for Tyr GEM buffer objects.
|
||||
pub(crate) type Bo = gem::shmem::Object<BoData>;
|
||||
|
||||
/// Creates a dummy GEM object to serve as the root of a GPUVM.
|
||||
pub(crate) fn new_dummy_object(ddev: &TyrDrmDevice) -> Result<ARef<Bo>> {
|
||||
let bo = Bo::new(
|
||||
ddev,
|
||||
4096,
|
||||
shmem::ObjectConfig {
|
||||
map_wc: true,
|
||||
parent_resv_obj: None,
|
||||
},
|
||||
BoCreateArgs { flags: 0 },
|
||||
)?;
|
||||
|
||||
Ok(bo)
|
||||
}
|
||||
|
||||
/// Specifies how to choose a GPU virtual address for a [`KernelBo`].
|
||||
/// An automatic VA allocation strategy will be added in the future.
|
||||
pub(crate) enum KernelBoVaAlloc {
|
||||
/// Explicit VA address specified by the caller.
|
||||
Explicit(u64),
|
||||
}
|
||||
|
||||
/// A kernel-owned buffer object with automatic GPU virtual address mapping.
|
||||
///
|
||||
/// This structure represents a buffer object that is created and managed entirely
|
||||
/// by the kernel driver, as opposed to userspace-created GEM objects. It combines
|
||||
/// a GEM object with automatic GPU virtual address (VA) space mapping and cleanup.
|
||||
///
|
||||
/// When dropped, the buffer is automatically unmapped from the GPU VA space.
|
||||
pub(crate) struct KernelBo<'drm> {
|
||||
/// The underlying GEM buffer object.
|
||||
bo: ARef<Bo>,
|
||||
/// The GPU VM this buffer is mapped into.
|
||||
vm: Arc<Vm<'drm>>,
|
||||
/// The GPU VA range occupied by this buffer.
|
||||
va_range: Range<u64>,
|
||||
}
|
||||
|
||||
impl<'drm> KernelBo<'drm> {
|
||||
/// Creates a new kernel-owned buffer object and maps it into GPU VA space.
|
||||
///
|
||||
/// This function allocates a new shmem-backed GEM object and immediately maps
|
||||
/// it into the specified GPU virtual memory space. The mapping is automatically
|
||||
/// cleaned up when the [`KernelBo`] is dropped.
|
||||
pub(crate) fn new(
|
||||
ddev: &TyrDrmDevice,
|
||||
vm: Arc<Vm<'drm>>,
|
||||
size: u64,
|
||||
va_alloc: KernelBoVaAlloc,
|
||||
flags: VmMapFlags,
|
||||
) -> Result<Self> {
|
||||
if size == 0 {
|
||||
dev_err!(vm.dev(), "Cannot create KernelBo with size 0");
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let KernelBoVaAlloc::Explicit(va) = va_alloc;
|
||||
|
||||
let bo_size = usize::try_from(size).map_err(|_| EOVERFLOW)?;
|
||||
let va_end = va.checked_add(size).ok_or(EINVAL)?;
|
||||
|
||||
let bo = Bo::new(
|
||||
ddev,
|
||||
bo_size,
|
||||
shmem::ObjectConfig {
|
||||
map_wc: true,
|
||||
parent_resv_obj: None,
|
||||
},
|
||||
BoCreateArgs { flags: 0 },
|
||||
)?;
|
||||
|
||||
vm.map_bo_range(&bo, 0, size, va, flags)?;
|
||||
|
||||
Ok(KernelBo {
|
||||
bo,
|
||||
vm,
|
||||
va_range: va..va_end,
|
||||
})
|
||||
}
|
||||
|
||||
pub(crate) fn bo(&self) -> &Bo {
|
||||
&self.bo
|
||||
}
|
||||
}
|
||||
|
||||
impl Drop for KernelBo<'_> {
|
||||
fn drop(&mut self) {
|
||||
let va = self.va_range.start;
|
||||
let size = self.va_range.end - self.va_range.start;
|
||||
|
||||
if let Err(e) = self.vm.unmap_range(va, size) {
|
||||
// If unmap_range fails, it is still safe to drop the
|
||||
// KernelBo and its ARef to the GEM buffer object because
|
||||
// GPUVM also holds a reference to the GEM buffer object.
|
||||
// The physical pages won't be freed or reallocated.
|
||||
dev_err!(
|
||||
self.vm.dev(),
|
||||
"Failed to unmap KernelBo range {:#x}..{:#x}: {:?}",
|
||||
self.va_range.start,
|
||||
self.va_range.end,
|
||||
e
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
120
drivers/gpu/drm/tyr/mmu.rs
Normal file
120
drivers/gpu/drm/tyr/mmu.rs
Normal file
|
|
@ -0,0 +1,120 @@
|
|||
// SPDX-License-Identifier: GPL-2.0 or MIT
|
||||
|
||||
//! Memory Management Unit (MMU) module.
|
||||
//!
|
||||
//! The GPU MMU provides a limited number of memory address spaces for use by command streams.
|
||||
//! The MMU translates virtual addresses to physical addresses and manages memory configuration
|
||||
//! and access permissions.
|
||||
//!
|
||||
//! This MMU module is essentially a locked wrapper around a [`SlotManager`] instance.
|
||||
//! The [`SlotManager`] manages the assignment of virtual address spaces to hardware address-space
|
||||
//! (AS) slots. MMU commands such as updates and flushes are carried out by the
|
||||
//! [`AddressSpaceManager`] which actually writes to the MMU registers.
|
||||
|
||||
use core::ops::Range;
|
||||
|
||||
use kernel::{
|
||||
device::{
|
||||
Bound,
|
||||
Device, //
|
||||
},
|
||||
new_mutex,
|
||||
prelude::*,
|
||||
sync::{
|
||||
Arc,
|
||||
ArcBorrow,
|
||||
Mutex, //
|
||||
}, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::IoMem,
|
||||
gpu::GpuInfo,
|
||||
mmu::address_space::{
|
||||
AddressSpaceManager,
|
||||
VmAsData, //
|
||||
},
|
||||
regs::{
|
||||
gpu_control::AS_PRESENT,
|
||||
MAX_AS, //
|
||||
},
|
||||
slot::SlotManager, //
|
||||
};
|
||||
|
||||
pub(crate) mod address_space;
|
||||
|
||||
pub(crate) type AsSlotManager<'drm> = SlotManager<AddressSpaceManager<'drm>, MAX_AS>;
|
||||
|
||||
/// Locked wrapper for carrying out virtual memory (VM) operations on the MMU.
|
||||
#[pin_data]
|
||||
pub(crate) struct Mmu<'drm> {
|
||||
/// Slot Manager instance used to allocate hardware slots and write to MMU registers.
|
||||
#[pin]
|
||||
pub(crate) as_manager: Mutex<AsSlotManager<'drm>>,
|
||||
}
|
||||
|
||||
impl<'drm> Mmu<'drm> {
|
||||
/// Create an MMU component for this device.
|
||||
pub(crate) fn new(
|
||||
dev: &'drm Device<Bound>,
|
||||
iomem: ArcBorrow<'_, IoMem<'drm>>,
|
||||
gpu_info: &GpuInfo,
|
||||
) -> Result<Arc<Mmu<'drm>>> {
|
||||
let present = AS_PRESENT::from_raw(gpu_info.as_present).present().get();
|
||||
let slot_count = present.count_ones().try_into()?;
|
||||
|
||||
let address_space_manager = AddressSpaceManager::new(dev, iomem.into(), present)?;
|
||||
let as_slot_manager =
|
||||
SlotManager::new(address_space_manager, slot_count).inspect_err(|e| {
|
||||
dev_err!(
|
||||
dev,
|
||||
"Failed to initialize MMU slot manager with {} slots: {:?}",
|
||||
slot_count,
|
||||
e
|
||||
);
|
||||
})?;
|
||||
let mmu_init = try_pin_init!(Self{
|
||||
as_manager <- new_mutex!(as_slot_manager),
|
||||
});
|
||||
Arc::pin_init(mmu_init, GFP_KERNEL)
|
||||
}
|
||||
|
||||
/// Assign a VM to an AS slot, provide a translation table,
|
||||
/// and update the MMU to make the VM resident.
|
||||
pub(crate) fn activate_vm(&self, vm_as_data: ArcBorrow<'_, VmAsData<'drm>>) -> Result {
|
||||
self.as_manager.lock().activate_vm(vm_as_data)
|
||||
}
|
||||
|
||||
/// Evict a VM from its AS slot and flush the MMU.
|
||||
pub(crate) fn deactivate_vm(&self, vm_as_data: &VmAsData<'drm>) -> Result {
|
||||
self.as_manager.lock().deactivate_vm(vm_as_data)
|
||||
}
|
||||
|
||||
/// Flush MMU translation caches after a VM update.
|
||||
pub(crate) fn flush_vm(&self, vm_as_data: &VmAsData<'drm>) -> Result {
|
||||
self.as_manager.lock().flush_vm(vm_as_data)
|
||||
}
|
||||
|
||||
/// Flags the start of a VM update.
|
||||
///
|
||||
/// If the VM is resident, any GPU access on the memory range being
|
||||
/// updated will be blocked until `Mmu::end_vm_update()` is called.
|
||||
/// This guarantees the atomicity of a VM update.
|
||||
/// If the VM is not resident, this is a NOP.
|
||||
pub(crate) fn start_vm_update(
|
||||
&self,
|
||||
vm_as_data: &VmAsData<'drm>,
|
||||
region: &Range<u64>,
|
||||
) -> Result {
|
||||
self.as_manager.lock().start_vm_update(vm_as_data, region)
|
||||
}
|
||||
|
||||
/// Flags the end of a VM update.
|
||||
///
|
||||
/// If the VM is resident, this will let GPU accesses on the updated
|
||||
/// range go through, in case any of them were blocked.
|
||||
/// If the VM is not resident, this is a NOP.
|
||||
pub(crate) fn end_vm_update(&self, vm_as_data: &VmAsData<'drm>) -> Result {
|
||||
self.as_manager.lock().end_vm_update(vm_as_data)
|
||||
}
|
||||
}
|
||||
511
drivers/gpu/drm/tyr/mmu/address_space.rs
Normal file
511
drivers/gpu/drm/tyr/mmu/address_space.rs
Normal file
|
|
@ -0,0 +1,511 @@
|
|||
// SPDX-License-Identifier: GPL-2.0 or MIT
|
||||
|
||||
//! Address space module.
|
||||
//!
|
||||
//! This module handles the hardware interaction for MMU operations through
|
||||
//! MMIO register access.
|
||||
//!
|
||||
|
||||
use core::ops::Range;
|
||||
|
||||
use kernel::{
|
||||
device::{
|
||||
Bound,
|
||||
Device, //
|
||||
}, //
|
||||
error::Result,
|
||||
io::{
|
||||
poll,
|
||||
register::Array,
|
||||
Io, //
|
||||
},
|
||||
iommu::pgtable::{
|
||||
Config,
|
||||
IoPageTable,
|
||||
ARM64LPAES1, //
|
||||
},
|
||||
num::Bounded,
|
||||
prelude::*,
|
||||
sizes::{
|
||||
SZ_2M,
|
||||
SZ_4K, //
|
||||
},
|
||||
sync::{
|
||||
Arc,
|
||||
ArcBorrow,
|
||||
LockedBy, //
|
||||
},
|
||||
time::Delta, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::IoMem,
|
||||
mmu::{
|
||||
AsSlotManager,
|
||||
Mmu, //
|
||||
},
|
||||
regs::{
|
||||
mmu_control::mmu_as_control,
|
||||
mmu_control::mmu_as_control::*,
|
||||
MAX_AS, //
|
||||
},
|
||||
slot::{
|
||||
LockedSeat,
|
||||
Seat,
|
||||
SlotOperations, //
|
||||
}, //
|
||||
};
|
||||
|
||||
/// Address space configuration values to be written to MMU registers.
|
||||
#[derive(Clone, Copy)]
|
||||
struct AddressSpaceConfig {
|
||||
/// Translation configuration. Configures how the MMU walks the page table for this
|
||||
/// address space.
|
||||
transcfg: u64,
|
||||
|
||||
/// Translation table base address. The address of the page table.
|
||||
transtab: u64,
|
||||
|
||||
/// Memory attributes such as cacheability.
|
||||
memattr: u64,
|
||||
}
|
||||
|
||||
/// Virtual memory (VM) address space data for use in MMU operations.
|
||||
#[pin_data]
|
||||
pub(crate) struct VmAsData<'drm> {
|
||||
/// This address-space seat tracks this VM's binding to a hardware address space slot.
|
||||
/// It can only be accessed when holding the `Mmu::as_manager` lock.
|
||||
as_seat: LockedSeat<AddressSpaceManager<'drm>, MAX_AS>,
|
||||
|
||||
/// Virtual address bits for this address space.
|
||||
va_bits: u8,
|
||||
|
||||
/// The page table which maps GPU virtual addresses to physical addresses for this VM.
|
||||
#[pin]
|
||||
pub(crate) page_table: IoPageTable<'drm, ARM64LPAES1>,
|
||||
}
|
||||
|
||||
impl<'drm> VmAsData<'drm> {
|
||||
/// Creates VM address space data by initializing all of its fields.
|
||||
pub(crate) fn new<'a>(
|
||||
mmu: &'a Mmu<'drm>,
|
||||
dev: &'drm Device<Bound>,
|
||||
va_bits: u32,
|
||||
pa_bits: u32,
|
||||
) -> impl pin_init::PinInit<VmAsData<'drm>, Error> + 'a {
|
||||
let pt_config = Config {
|
||||
quirks: 0,
|
||||
pgsize_bitmap: SZ_4K | SZ_2M,
|
||||
ias: va_bits,
|
||||
oas: pa_bits,
|
||||
coherent_walk: false,
|
||||
};
|
||||
|
||||
let page_table_init = IoPageTable::new(dev, pt_config);
|
||||
|
||||
try_pin_init!(Self {
|
||||
as_seat: LockedBy::new(&mmu.as_manager, Seat::NoSeat),
|
||||
va_bits: va_bits as u8,
|
||||
page_table <- page_table_init,
|
||||
}? Error)
|
||||
}
|
||||
|
||||
/// Computes the hardware configuration for this address space.
|
||||
fn as_config(&self) -> Result<AddressSpaceConfig> {
|
||||
let pt = &self.page_table;
|
||||
// The hardware computes the valid input address range as:
|
||||
// INA_BITS_VALID = min(HW_INA_BITS, 55 - INA_BITS)
|
||||
// To configure our desired va_bits, we solve for INA_BITS:
|
||||
// INA_BITS = 55 - va_bits
|
||||
// This assumes HW_INA_BITS (hardware capability) >= va_bits.
|
||||
let field = 55u64.checked_sub(self.va_bits.into()).ok_or(EINVAL)?;
|
||||
let ina_bits =
|
||||
match mmu_as_control::InaBits::try_from(Bounded::try_new(field).ok_or(EINVAL)?)? {
|
||||
mmu_as_control::InaBits::Reset => return Err(EINVAL),
|
||||
bits => bits,
|
||||
};
|
||||
|
||||
let transcfg = mmu_as_control::TRANSCFG::zeroed()
|
||||
.with_ptw_memattr(mmu_as_control::PtwMemattr::WriteBack)
|
||||
.with_r_allocate(true)
|
||||
.with_mode(mmu_as_control::AddressSpaceMode::Aarch64_4K)
|
||||
.with_ina_bits(ina_bits)
|
||||
.into_raw();
|
||||
|
||||
Ok(AddressSpaceConfig {
|
||||
transcfg,
|
||||
// SAFETY: The SlotManager holds an `Arc<VmAsData>` as SlotData while this
|
||||
// TTBR is programmed and stores that Arc in the active slot before
|
||||
// returning. Eviction flushes and disables the slot before releasing
|
||||
// the Arc; if eviction fails, the slot retains it. Therefore the page
|
||||
// table cannot be dropped while the GPU is using it.
|
||||
transtab: unsafe { pt.ttbr() },
|
||||
memattr: MEMATTR::from_mair(pt.mair()).into_raw(),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// Coordinates all hardware-level address space operations through MMIO register
|
||||
/// operations including enabling, disabling, flushing, and updating address spaces.
|
||||
pub(crate) struct AddressSpaceManager<'drm> {
|
||||
/// Parent device used for logging.
|
||||
dev: &'drm Device<Bound>,
|
||||
|
||||
/// Memory-mapped I/O region for GPU register access.
|
||||
iomem: Arc<IoMem<'drm>>,
|
||||
|
||||
/// Bitmask of present address space slots from GPU_AS_PRESENT register.
|
||||
as_present: u32,
|
||||
}
|
||||
|
||||
impl<'drm> AddressSpaceManager<'drm> {
|
||||
/// Creates a new address space manager.
|
||||
///
|
||||
/// Initializes the manager with references to the platform device and
|
||||
/// I/O memory region, along with the bitmask of available AS slots.
|
||||
pub(super) fn new(
|
||||
dev: &'drm Device<Bound>,
|
||||
iomem: Arc<IoMem<'drm>>,
|
||||
as_present: u32,
|
||||
) -> Result<AddressSpaceManager<'drm>> {
|
||||
if as_present.trailing_ones() != as_present.count_ones() {
|
||||
dev_err!(
|
||||
dev,
|
||||
"Sparse AS_PRESENT mask is unsupported: {:#x}",
|
||||
as_present
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
Ok(Self {
|
||||
dev,
|
||||
iomem,
|
||||
as_present,
|
||||
})
|
||||
}
|
||||
|
||||
/// Validates that an AS slot number is within range and present in hardware.
|
||||
///
|
||||
/// Checks that the slot index is less than [`MAX_AS`] and that
|
||||
/// the corresponding bit is set in the `as_present` mask read from the GPU.
|
||||
///
|
||||
/// Returns [`EINVAL`] if the slot is out of range or not present in hardware.
|
||||
fn validate_as_slot(&self, as_nr: usize) -> Result {
|
||||
if as_nr >= MAX_AS {
|
||||
dev_err!(
|
||||
self.dev,
|
||||
"AS slot {} out of valid range (max {})",
|
||||
as_nr,
|
||||
MAX_AS
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
if (self.as_present & (1 << as_nr)) == 0 {
|
||||
dev_err!(
|
||||
self.dev,
|
||||
"AS slot {} not present in hardware (AS_PRESENT={:#x})",
|
||||
as_nr,
|
||||
self.as_present
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Waits for an AS slot to become ready (not active).
|
||||
///
|
||||
/// Returns an error if polling times out after 10ms or if register access fails.
|
||||
fn as_wait_ready(&self, as_nr: usize) -> Result {
|
||||
let io = &*self.iomem;
|
||||
let op = || {
|
||||
let status_reg = STATUS::try_at(as_nr).ok_or(EINVAL)?;
|
||||
Ok(io.read(status_reg))
|
||||
};
|
||||
let cond = |status: &STATUS| -> bool { !status.active_ext() };
|
||||
poll::read_poll_timeout(op, cond, Delta::from_micros(50), Delta::from_millis(10))?;
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Sends a command to an AS slot.
|
||||
///
|
||||
/// Returns an error if waiting for ready times out or if register write fails.
|
||||
fn as_send_cmd(&mut self, as_nr: usize, cmd: MmuCommand) -> Result {
|
||||
self.as_wait_ready(as_nr)?;
|
||||
let io = &*self.iomem;
|
||||
let command_reg = COMMAND::try_at(as_nr).ok_or(EINVAL)?;
|
||||
io.write(command_reg, COMMAND::zeroed().with_command(cmd));
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Sends a command to an AS slot and waits for completion.
|
||||
///
|
||||
/// Returns an error if sending the command fails or if waiting for completion times out.
|
||||
fn as_send_cmd_and_wait(&mut self, as_nr: usize, cmd: MmuCommand) -> Result {
|
||||
self.as_send_cmd(as_nr, cmd)?;
|
||||
self.as_wait_ready(as_nr)?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Enables an AS slot with the provided configuration.
|
||||
///
|
||||
/// Returns an error if the slot is invalid or if register writes/commands fail.
|
||||
fn as_enable(&mut self, as_nr: usize, as_config: &AddressSpaceConfig) -> Result {
|
||||
self.validate_as_slot(as_nr)?;
|
||||
|
||||
let io = &*self.iomem;
|
||||
|
||||
let transtab = as_config.transtab;
|
||||
io.write(
|
||||
TRANSTAB_LO::try_at(as_nr).ok_or(EINVAL)?,
|
||||
TRANSTAB_LO::from_raw(transtab as u32),
|
||||
);
|
||||
io.write(
|
||||
TRANSTAB_HI::try_at(as_nr).ok_or(EINVAL)?,
|
||||
TRANSTAB_HI::from_raw((transtab >> 32) as u32),
|
||||
);
|
||||
|
||||
let transcfg = as_config.transcfg;
|
||||
io.write(
|
||||
TRANSCFG_LO::try_at(as_nr).ok_or(EINVAL)?,
|
||||
TRANSCFG_LO::from_raw(transcfg as u32),
|
||||
);
|
||||
io.write(
|
||||
TRANSCFG_HI::try_at(as_nr).ok_or(EINVAL)?,
|
||||
TRANSCFG_HI::from_raw((transcfg >> 32) as u32),
|
||||
);
|
||||
|
||||
let memattr = as_config.memattr;
|
||||
io.write(
|
||||
MEMATTR_LO::try_at(as_nr).ok_or(EINVAL)?,
|
||||
MEMATTR_LO::from_raw(memattr as u32),
|
||||
);
|
||||
io.write(
|
||||
MEMATTR_HI::try_at(as_nr).ok_or(EINVAL)?,
|
||||
MEMATTR_HI::from_raw((memattr >> 32) as u32),
|
||||
);
|
||||
|
||||
self.as_send_cmd_and_wait(as_nr, MmuCommand::Update)?;
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Disables an AS slot and clears its configuration.
|
||||
///
|
||||
/// Returns an error if the slot is invalid or if register writes/commands fail.
|
||||
fn as_disable(&mut self, as_nr: usize) -> Result {
|
||||
self.validate_as_slot(as_nr)?;
|
||||
|
||||
// Flush AS before disabling
|
||||
self.as_send_cmd_and_wait(as_nr, MmuCommand::FlushMem)?;
|
||||
|
||||
let io = &*self.iomem;
|
||||
|
||||
io.write(
|
||||
TRANSTAB_LO::try_at(as_nr).ok_or(EINVAL)?,
|
||||
TRANSTAB_LO::from_raw(0),
|
||||
);
|
||||
io.write(
|
||||
TRANSTAB_HI::try_at(as_nr).ok_or(EINVAL)?,
|
||||
TRANSTAB_HI::from_raw(0),
|
||||
);
|
||||
|
||||
io.write(
|
||||
MEMATTR_LO::try_at(as_nr).ok_or(EINVAL)?,
|
||||
MEMATTR_LO::from_raw(0),
|
||||
);
|
||||
io.write(
|
||||
MEMATTR_HI::try_at(as_nr).ok_or(EINVAL)?,
|
||||
MEMATTR_HI::from_raw(0),
|
||||
);
|
||||
|
||||
let transcfg = TRANSCFG::zeroed()
|
||||
.with_mode(AddressSpaceMode::Unmapped)
|
||||
.into_raw();
|
||||
|
||||
io.write(
|
||||
TRANSCFG_LO::try_at(as_nr).ok_or(EINVAL)?,
|
||||
TRANSCFG_LO::from_raw(transcfg as u32),
|
||||
);
|
||||
io.write(
|
||||
TRANSCFG_HI::try_at(as_nr).ok_or(EINVAL)?,
|
||||
TRANSCFG_HI::from_raw((transcfg >> 32) as u32),
|
||||
);
|
||||
|
||||
self.as_send_cmd_and_wait(as_nr, MmuCommand::Update)?;
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Locks a region of the translation tables for an atomic update.
|
||||
///
|
||||
/// Programs the MMU [`LOCKADDR`] register for the given address space and issues
|
||||
/// the lock command. The hardware rounds the requested range up to a
|
||||
/// power-of-two region aligned to its size.
|
||||
///
|
||||
/// Returns an error if the slot is invalid or if register writes/commands fail.
|
||||
fn as_start_update(&mut self, as_nr: usize, region: &Range<u64>) -> Result {
|
||||
self.validate_as_slot(as_nr)?;
|
||||
|
||||
// Avoid both an empty range and an inverted range.
|
||||
if region.start >= region.end {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// The lock operates on full 64-byte cache lines of translation table entries.
|
||||
// Since each translation table entry (TTE) is 8 bytes, a cache line has 8 TTEs.
|
||||
// Since each TTE maps one page, the minimum locked region size will be 8 pages.
|
||||
//
|
||||
// With 4KiB pages (Aarch64_4K mode), the minimum locked region is 32KiB.
|
||||
let lock_region_min_size: u64 = 4096 * 8;
|
||||
|
||||
// Count the number of trailing zero bits (zeros at the right/least-significant
|
||||
// end of the binary representation). For a power-of-two value, this equals the
|
||||
// base-2 exponent (e.g., 32 KiB = 2^15 → 15).
|
||||
let lock_region_min_size_log2 = lock_region_min_size.trailing_zeros() as u8;
|
||||
|
||||
// XOR the first and last addresses to identify which bits differ between them.
|
||||
// The highest set bit in the result determines the exponent of the smallest
|
||||
// power-of-two region that can contain both addresses.
|
||||
//
|
||||
// Example:
|
||||
// addr_xor = 0x1000 ^ 0x2FFF = 0x3FFF
|
||||
// highest set bit in 0x3FFF is bit 13
|
||||
// minimum region size = 2^(13 + 1) = 16 KiB
|
||||
let addr_xor = region.start ^ (region.end - 1);
|
||||
let region_size_log2 = 64 - addr_xor.leading_zeros() as u8;
|
||||
|
||||
let lock_region_log2 = core::cmp::max(region_size_log2, lock_region_min_size_log2);
|
||||
|
||||
let lock_region_size = 1u64.checked_shl(lock_region_log2.into()).ok_or(EINVAL)?;
|
||||
// Align the LOCKADDR base address down to the lock region size (1 << lock_region_log2).
|
||||
//
|
||||
// The MMU ignores the low lock_region_log2 bits of LOCKADDR base, so ensure
|
||||
// they are cleared in software to avoid ambiguity.
|
||||
//
|
||||
// Example:
|
||||
// lock_region_log2 = 14 (16 KiB)
|
||||
// region.start = 0x1000
|
||||
// lockaddr_base = 0x1000 & ~(0x3FFF) = 0x0000
|
||||
let lockaddr_base = region.start & !(lock_region_size - 1);
|
||||
|
||||
// The LOCKADDR size field encodes the lock region size as log2(size) - 1,
|
||||
// per the hardware definition. For example, a 32 KiB region is encoded as 14
|
||||
// because log2(32 KiB) = 15.
|
||||
let lockaddr_size = lock_region_log2 - 1;
|
||||
|
||||
let io = &*self.iomem;
|
||||
|
||||
// The LOCKADDR base field stores address bits 63:12, so remove the low 12 bits
|
||||
// before passing this value to the register macro helper.
|
||||
// These bits are guaranteed to be zero anyway because of the minimum
|
||||
// size of the locked region.
|
||||
let lockaddr_base_field = lockaddr_base >> 12;
|
||||
let lockaddr_val = LOCKADDR::zeroed()
|
||||
.try_with_size(lockaddr_size)?
|
||||
.try_with_base(lockaddr_base_field)?
|
||||
.into_raw();
|
||||
|
||||
io.write(
|
||||
LOCKADDR_LO::try_at(as_nr).ok_or(EINVAL)?,
|
||||
LOCKADDR_LO::from_raw(lockaddr_val as u32),
|
||||
);
|
||||
io.write(
|
||||
LOCKADDR_HI::try_at(as_nr).ok_or(EINVAL)?,
|
||||
LOCKADDR_HI::from_raw((lockaddr_val >> 32) as u32),
|
||||
);
|
||||
|
||||
self.as_send_cmd_and_wait(as_nr, MmuCommand::Lock)
|
||||
}
|
||||
|
||||
/// Completes an atomic translation table update.
|
||||
///
|
||||
/// Returns an error if the slot is invalid or if the flush command fails.
|
||||
fn as_end_update(&mut self, as_nr: usize) -> Result {
|
||||
self.validate_as_slot(as_nr)?;
|
||||
self.as_send_cmd_and_wait(as_nr, MmuCommand::FlushPt)?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Flushes the translation table cache for an AS slot.
|
||||
///
|
||||
/// Returns an error if the slot is invalid or if the flush command fails.
|
||||
fn as_flush(&mut self, as_nr: usize) -> Result {
|
||||
self.validate_as_slot(as_nr)?;
|
||||
self.as_send_cmd_and_wait(as_nr, MmuCommand::FlushPt)
|
||||
}
|
||||
}
|
||||
|
||||
impl<'drm> SlotOperations<MAX_AS> for AddressSpaceManager<'drm> {
|
||||
/// VM address space data associated with a hardware slot.
|
||||
type SlotData = Arc<VmAsData<'drm>>;
|
||||
|
||||
fn seat(slot_data: &Self::SlotData) -> &LockedSeat<Self, MAX_AS> {
|
||||
&slot_data.as_seat
|
||||
}
|
||||
|
||||
/// Activates a VM in a hardware slot.
|
||||
fn activate(&mut self, slot_idx: usize, slot_data: &Self::SlotData) -> Result {
|
||||
let as_config = slot_data.as_config()?;
|
||||
self.as_enable(slot_idx, &as_config)
|
||||
}
|
||||
|
||||
/// Evicts a VM from a hardware slot.
|
||||
fn evict(&mut self, slot_idx: usize, _slot_data: &Self::SlotData) -> Result {
|
||||
self.as_flush(slot_idx)?;
|
||||
self.as_disable(slot_idx)?;
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
impl<'drm> AsSlotManager<'drm> {
|
||||
/// Locks a region for translation table updates if the VM has an active slot.
|
||||
pub(super) fn start_vm_update(
|
||||
&mut self,
|
||||
vm_as_data: &VmAsData<'drm>,
|
||||
region: &Range<u64>,
|
||||
) -> Result {
|
||||
let seat = vm_as_data.as_seat.access(self);
|
||||
match seat.slot() {
|
||||
Some(slot) => {
|
||||
let as_nr = slot as usize;
|
||||
self.as_start_update(as_nr, region)
|
||||
}
|
||||
_ => Ok(()),
|
||||
}
|
||||
}
|
||||
|
||||
/// Completes translation table updates and unlocks the region.
|
||||
pub(super) fn end_vm_update(&mut self, vm_as_data: &VmAsData<'drm>) -> Result {
|
||||
let seat = vm_as_data.as_seat.access(self);
|
||||
match seat.slot() {
|
||||
Some(slot) => {
|
||||
let as_nr = slot as usize;
|
||||
self.as_end_update(as_nr)
|
||||
}
|
||||
_ => Ok(()),
|
||||
}
|
||||
}
|
||||
|
||||
/// Flushes the translation table cache if the VM has an active slot.
|
||||
pub(super) fn flush_vm(&mut self, vm_as_data: &VmAsData<'drm>) -> Result {
|
||||
let seat = vm_as_data.as_seat.access(self);
|
||||
match seat.slot() {
|
||||
Some(slot) => {
|
||||
let as_nr = slot as usize;
|
||||
self.as_flush(as_nr)
|
||||
}
|
||||
_ => Ok(()),
|
||||
}
|
||||
}
|
||||
|
||||
/// Activates a VM by assigning it to a hardware slot.
|
||||
pub(super) fn activate_vm(&mut self, vm_as_data: ArcBorrow<'_, VmAsData<'drm>>) -> Result {
|
||||
self.activate(vm_as_data.into())
|
||||
}
|
||||
|
||||
/// Deactivates a VM by evicting it from its hardware slot.
|
||||
pub(super) fn deactivate_vm(&mut self, vm_as_data: &VmAsData<'drm>) -> Result {
|
||||
self.evict(&vm_as_data.as_seat)
|
||||
}
|
||||
}
|
||||
|
|
@ -25,7 +25,7 @@
|
|||
//
|
||||
// Nevertheless, it is useful to have most of them defined, like the C driver
|
||||
// does.
|
||||
#![allow(dead_code)]
|
||||
#![expect(dead_code)]
|
||||
|
||||
/// Combine two 32-bit values into a single 64-bit value.
|
||||
pub(crate) fn join_u64(lo: u32, hi: u32) -> u64 {
|
||||
|
|
@ -45,20 +45,17 @@ pub(crate) fn read_u64_no_tearing(lo_read: impl Fn() -> u32, hi_read: impl Fn()
|
|||
}
|
||||
}
|
||||
|
||||
pub(crate) use mmu_control::mmu_as_control::MAX_AS;
|
||||
|
||||
/// These registers correspond to the GPU_CONTROL register page.
|
||||
/// They are involved in GPU configuration and control.
|
||||
pub(crate) mod gpu_control {
|
||||
use core::convert::TryFrom;
|
||||
use kernel::{
|
||||
error::{
|
||||
code::EINVAL,
|
||||
Error, //
|
||||
},
|
||||
num::Bounded,
|
||||
prelude::*,
|
||||
register,
|
||||
uapi, //
|
||||
};
|
||||
use pin_init::Zeroable;
|
||||
|
||||
register! {
|
||||
/// GPU identification register.
|
||||
|
|
@ -964,17 +961,14 @@ pub(crate) mod mmu_control {
|
|||
///
|
||||
/// This array contains 16 instances of the MMU_AS_CONTROL register page.
|
||||
pub(crate) mod mmu_as_control {
|
||||
use core::convert::TryFrom;
|
||||
|
||||
use kernel::{
|
||||
error::{
|
||||
code::EINVAL,
|
||||
Error, //
|
||||
},
|
||||
num::Bounded,
|
||||
prelude::*,
|
||||
register, //
|
||||
};
|
||||
|
||||
use pin_init::Zeroable;
|
||||
|
||||
/// Maximum number of hardware address space slots.
|
||||
/// The actual number of slots available is usually lower.
|
||||
pub(crate) const MAX_AS: usize = 16;
|
||||
|
|
@ -1168,7 +1162,136 @@ fn from(val: MMU_MEMATTR_STAGE1) -> Self {
|
|||
pub(crate) MEMATTR_HI(u32)[MAX_AS, stride = STRIDE] @ 0x240c {
|
||||
31:0 value;
|
||||
}
|
||||
}
|
||||
|
||||
impl MEMATTR {
|
||||
/// Outer cache-policy nibble indicating device memory.
|
||||
const ARM_MAIR_DEVICE_MEMORY: u8 = 0x0;
|
||||
|
||||
/// In the ARM Architecture Reference Manual, the MAIR encoding for Normal memory
|
||||
/// uses the format `0bxxRW` where:
|
||||
/// - `W` (bit 0) = Write-Allocate policy
|
||||
/// - `R` (bit 1) = Read-Allocate policy
|
||||
/// E.g., `0b0011` would allow both read and write allocation on a cache miss.
|
||||
///
|
||||
/// ARM MAIR Write-Allocate bit (bit 0 of a cache policy nibble).
|
||||
const ARM_MAIR_WRITE_ALLOCATE: u8 = 0x1;
|
||||
/// ARM MAIR Read-Allocate bit (bit 1 of a cache policy nibble).
|
||||
const ARM_MAIR_READ_ALLOCATE: u8 = 0x2;
|
||||
|
||||
/// Write-back policy bit. For cacheable encodings, it is necessary but not
|
||||
/// sufficient to set bit 2 of the cache policy nibble. Bit 2 does not
|
||||
/// definitively determine write back because bit 2 is also set in `0b0100`
|
||||
/// which encodes Normal non-cacheable memory.
|
||||
const ARM_MAIR_WRITE_BACK_BIT: u8 = 0x4;
|
||||
|
||||
/// Complete cache-policy nibble encoding for Normal Non-cacheable memory.
|
||||
const ARM_MAIR_NON_CACHEABLE: u8 = 0x4;
|
||||
|
||||
/// Mask for the inner cache policy nibble in MAIR attribute bytes.
|
||||
const ARM_MAIR_INNER_MASK: u8 = 0x0f;
|
||||
|
||||
/// Check if a MAIR attribute byte represents device memory.
|
||||
///
|
||||
/// Device memory (memory-mapped I/O, registers) cannot be cached because
|
||||
/// reading and writing to this memory may have side effects.
|
||||
fn is_device_memory(mair_attr: u8) -> bool {
|
||||
// In AArch64 MAIR, outer nibble only is 0 for device memory.
|
||||
(mair_attr >> 4) == Self::ARM_MAIR_DEVICE_MEMORY
|
||||
}
|
||||
|
||||
/// Check if normal memory is fully write-back cacheable.
|
||||
///
|
||||
/// ARM MAIR has two cache policy levels (outer [7:4] and inner [3:0]).
|
||||
/// For memory to be truly write-back, BOTH levels must have the write-back bit set.
|
||||
/// If only one level is write-back, treat it as non-cacheable for GPU purposes.
|
||||
fn is_writeback_cacheable(mair_attr: u8) -> bool {
|
||||
let outer = mair_attr >> 4;
|
||||
let inner = mair_attr & Self::ARM_MAIR_INNER_MASK;
|
||||
|
||||
outer != Self::ARM_MAIR_NON_CACHEABLE
|
||||
&& inner != Self::ARM_MAIR_NON_CACHEABLE
|
||||
&& (outer & Self::ARM_MAIR_WRITE_BACK_BIT) != 0
|
||||
&& (inner & Self::ARM_MAIR_WRITE_BACK_BIT) != 0
|
||||
}
|
||||
|
||||
// Helper to encode a MEMATTR attribute from its individual fields.
|
||||
fn encode_attribute(
|
||||
alloc_w: bool,
|
||||
alloc_r: bool,
|
||||
alloc_sel: AllocPolicySelect,
|
||||
coherency: Coherency,
|
||||
memory_type: MemoryType,
|
||||
) -> MMU_MEMATTR_STAGE1 {
|
||||
MMU_MEMATTR_STAGE1::zeroed()
|
||||
.with_alloc_w(alloc_w)
|
||||
.with_alloc_r(alloc_r)
|
||||
.with_alloc_sel(alloc_sel)
|
||||
.with_coherency(coherency)
|
||||
.with_memory_type(memory_type)
|
||||
}
|
||||
|
||||
/// Convert one MAIR attribute byte into a MEMATTR attribute.
|
||||
// TODO: Add a `coherent` parameter like panthor's mair_to_memattr().
|
||||
// For now, assume a non-coherent system and always encode write-back
|
||||
// memory with MidgardInnerDomain coherency.
|
||||
fn attribute_from_mair(mair_attr: u8) -> MMU_MEMATTR_STAGE1 {
|
||||
// Device memory or non-write-back normal memory
|
||||
if Self::is_device_memory(mair_attr) || !Self::is_writeback_cacheable(mair_attr) {
|
||||
return Self::encode_attribute(
|
||||
false,
|
||||
false,
|
||||
AllocPolicySelect::Alloc,
|
||||
Coherency::MidgardInnerDomain,
|
||||
MemoryType::NonCacheable,
|
||||
);
|
||||
}
|
||||
|
||||
// Write-back cacheable normal memory
|
||||
let inner: u8 = mair_attr & Self::ARM_MAIR_INNER_MASK;
|
||||
Self::encode_attribute(
|
||||
(inner & Self::ARM_MAIR_WRITE_ALLOCATE) != 0,
|
||||
(inner & Self::ARM_MAIR_READ_ALLOCATE) != 0,
|
||||
AllocPolicySelect::Alloc,
|
||||
Coherency::MidgardInnerDomain,
|
||||
MemoryType::WriteBack,
|
||||
)
|
||||
}
|
||||
|
||||
/// Write one converted MAIR attribute into a corresponding MEMATTR slot.
|
||||
fn with_encoded_attribute(self, index: usize, attr: MMU_MEMATTR_STAGE1) -> Self {
|
||||
debug_assert!(index < 8);
|
||||
|
||||
let shift = index * 8;
|
||||
let mask = !(0xffu64 << shift);
|
||||
let raw = (self.into_raw() & mask) | ((u64::from(attr.into_raw())) << shift);
|
||||
|
||||
Self::from_raw(raw)
|
||||
}
|
||||
|
||||
/// Convert an AArch64 MAIR value into the GPU MEMATTR register encoding.
|
||||
///
|
||||
/// Both MAIR and MEMATTR are 64-bit values with eight 8-bit memory
|
||||
/// attribute entries, but the bits do not map directly. The GPU MEMATTR encoding
|
||||
/// is less detailed than the MAIR encoding, so MAIR is converted to MEMATTR
|
||||
/// conservatively as follows:
|
||||
///
|
||||
/// 1. Device memory, or Normal Memory that is not write-back cacheable, is encoded
|
||||
/// as GPU `NonCacheable`
|
||||
///
|
||||
/// 2. Normal memory that is write-back cacheable is encoded as GPU `WriteBack`,
|
||||
/// and the inner allocation hints are preserved.
|
||||
pub(crate) fn from_mair(mair: u64) -> Self {
|
||||
mair.to_le_bytes()
|
||||
.into_iter()
|
||||
.enumerate()
|
||||
.fold(Self::zeroed(), |acc, (i, attr)| {
|
||||
acc.with_encoded_attribute(i, Self::attribute_from_mair(attr))
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
register! {
|
||||
/// Lock region address for each address space.
|
||||
pub(crate) LOCKADDR(u64)[MAX_AS, stride = STRIDE] @ 0x2410 {
|
||||
/// Lock region size.
|
||||
|
|
|
|||
404
drivers/gpu/drm/tyr/slot.rs
Normal file
404
drivers/gpu/drm/tyr/slot.rs
Normal file
|
|
@ -0,0 +1,404 @@
|
|||
// SPDX-License-Identifier: GPL-2.0 or MIT
|
||||
|
||||
//! Slot management abstraction for limited hardware resources.
|
||||
//!
|
||||
//! This module provides a generic [`SlotManager`] that assigns limited hardware
|
||||
//! slots to logical "seats". A seat represents an entity (such as a virtual memory
|
||||
//! (VM) address space) that needs access to a hardware slot.
|
||||
//!
|
||||
//! The [`SlotManager`] tracks slot allocation using sequence numbers (seqno) to detect
|
||||
//! when a seat's binding has been invalidated. When a seat requests activation,
|
||||
//! the manager will either reuse the seat's existing slot (if still valid),
|
||||
//! allocate a free slot (if any are available), or evict the oldest idle slot if any
|
||||
//! slots are idle.
|
||||
//!
|
||||
//! Hardware-specific behavior is customized by implementing the [`SlotOperations`]
|
||||
//! trait, which allows callbacks when slots are activated or evicted.
|
||||
//!
|
||||
//! This is currently used for managing address space slots in the GPU, and it will
|
||||
//! also be used to manage Command Stream Group (CSG) interface slots in the future.
|
||||
//!
|
||||
//! [SlotOperations]: crate::slot::SlotOperations
|
||||
//! [SlotManager]: crate::slot::SlotManager
|
||||
|
||||
use core::{
|
||||
mem,
|
||||
ops::{
|
||||
Deref,
|
||||
DerefMut, //
|
||||
}, //
|
||||
};
|
||||
|
||||
use kernel::{
|
||||
prelude::*,
|
||||
sync::LockedBy, //
|
||||
};
|
||||
|
||||
/// Seat information.
|
||||
///
|
||||
/// This can't be accessed directly by the element embedding a `Seat`,
|
||||
/// but is used by the generic slot manager logic to control residency
|
||||
/// of a certain object on a hardware slot.
|
||||
pub(crate) struct SeatInfo {
|
||||
/// Slot used by this seat.
|
||||
///
|
||||
/// This index is only valid if the slot pointed to by this index
|
||||
/// has its `SlotInfo::seqno` match `SeatInfo::seqno`. Otherwise,
|
||||
/// it means the object has been evicted from the hardware slot,
|
||||
/// and a new slot needs to be acquired to make this object
|
||||
/// resident again.
|
||||
slot: u8,
|
||||
|
||||
/// Sequence number encoding the last time this seat was active.
|
||||
/// We also use it to check if a slot is still bound to a seat.
|
||||
seqno: u64,
|
||||
}
|
||||
|
||||
/// Seat state.
|
||||
///
|
||||
/// This is meant to be embedded in the object that wants to acquire
|
||||
/// hardware slots. It also starts in the `Seat::NoSeat` state, and
|
||||
/// the slot manager will change the object value when an active/evict
|
||||
/// request is issued.
|
||||
#[derive(Default)]
|
||||
pub(crate) enum Seat {
|
||||
#[expect(clippy::enum_variant_names)]
|
||||
/// Resource is not resident.
|
||||
///
|
||||
/// All objects start with a seat in the `Seat::NoSeat` state. The seat also
|
||||
/// gets back to that state if the user requests eviction. It
|
||||
/// can also end up in that state next time an operation is done
|
||||
/// on a `Seat::Idle` seat and the slot manager finds out this
|
||||
/// object has been evicted from the slot.
|
||||
#[default]
|
||||
NoSeat,
|
||||
|
||||
/// Resource is actively used and resident.
|
||||
///
|
||||
/// When a seat is in the `Seat::Active` state, it can't be evicted, and the
|
||||
/// slot pointed to by `SeatInfo::slot` is guaranteed to be reserved
|
||||
/// for this object as long as the seat stays active.
|
||||
Active(SeatInfo),
|
||||
|
||||
/// Resource is idle and might or might not be resident.
|
||||
///
|
||||
/// When a seat is in the`Seat::Idle` state, we can't know for sure if the
|
||||
/// object is resident or evicted until the next request we issue
|
||||
/// to the slot manager. This tells the slot manager it can
|
||||
/// reclaim the underlying slot if needed.
|
||||
/// In order for the hardware to use this object again, the seat
|
||||
/// needs to be turned into an `Seat::Active` state again
|
||||
/// with a `SlotManager::activate()` call.
|
||||
Idle(SeatInfo),
|
||||
}
|
||||
|
||||
impl Seat {
|
||||
/// Get the slot index this seat is pointing to.
|
||||
///
|
||||
/// If the seat is not `Seat::Active` we can't trust the
|
||||
/// `SeatInfo`. In that case `None` is returned, otherwise
|
||||
/// `Some(SeatInfo::slot)` is returned.
|
||||
pub(crate) fn slot(&self) -> Option<u8> {
|
||||
match self {
|
||||
Self::Active(info) => Some(info.slot),
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Information related to a slot.
|
||||
struct SlotInfo<D> {
|
||||
/// Type specific data attached to a slot.
|
||||
slot_data: D,
|
||||
|
||||
/// Sequence number from when this slot was last activated.
|
||||
seqno: u64,
|
||||
}
|
||||
|
||||
/// Slot state.
|
||||
#[derive(Default)]
|
||||
enum Slot<D> {
|
||||
/// Slot is free.
|
||||
#[default]
|
||||
Free,
|
||||
|
||||
/// Slot is active.
|
||||
Active(SlotInfo<D>),
|
||||
|
||||
/// Slot is idle.
|
||||
Idle(SlotInfo<D>),
|
||||
}
|
||||
|
||||
pub(crate) type LockedSeat<T, const MAX_SLOTS: usize> = LockedBy<Seat, SlotManager<T, MAX_SLOTS>>;
|
||||
|
||||
/// Trait describing the slot-related operations.
|
||||
pub(crate) trait SlotOperations<const MAX_SLOTS: usize>: Sized {
|
||||
/// Implementation-specific data associated with each slot.
|
||||
type SlotData;
|
||||
|
||||
/// Returns the seat belonging to this slot data.
|
||||
fn seat(slot_data: &Self::SlotData) -> &LockedSeat<Self, MAX_SLOTS>;
|
||||
|
||||
/// Called when a slot is being activated for a seat.
|
||||
fn activate(&mut self, _slot_idx: usize, _slot_data: &Self::SlotData) -> Result {
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Called when a slot is being evicted and freed.
|
||||
fn evict(&mut self, _slot_idx: usize, _slot_data: &Self::SlotData) -> Result {
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
/// A generic slot manager that provides access to a limited number of hardware slots.
|
||||
pub(crate) struct SlotManager<T: SlotOperations<MAX_SLOTS>, const MAX_SLOTS: usize> {
|
||||
/// A specific implementation of the generic slot manager.
|
||||
manager: T,
|
||||
|
||||
/// Number of slots actually available.
|
||||
slot_count: usize,
|
||||
|
||||
/// Slot array used to track the state of each slot.
|
||||
slots: [Slot<T::SlotData>; MAX_SLOTS],
|
||||
|
||||
/// Sequence number incremented each time a Seat is successfully activated
|
||||
use_seqno: u64,
|
||||
}
|
||||
|
||||
impl<T: SlotOperations<MAX_SLOTS>, const MAX_SLOTS: usize> SlotManager<T, MAX_SLOTS> {
|
||||
/// Creates a specific instance of a slot manager.
|
||||
pub(crate) fn new(manager: T, slot_count: usize) -> Result<Self> {
|
||||
if slot_count == 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
if slot_count > MAX_SLOTS {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
// Since the slot index is stored in SeatInfo as a u8, the maximum number of slots is 256.
|
||||
if slot_count > u8::MAX as usize + 1 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
Ok(Self {
|
||||
manager,
|
||||
slot_count,
|
||||
slots: [const { Slot::Free }; MAX_SLOTS],
|
||||
use_seqno: 1,
|
||||
})
|
||||
}
|
||||
|
||||
/// Records a newly activated slot for the given seat.
|
||||
/// The slot manager takes ownership of the hardware-specific slot data.
|
||||
fn record_active_slot(&mut self, slot_idx: usize, slot_data: T::SlotData) {
|
||||
let cur_seqno = self.use_seqno;
|
||||
|
||||
*T::seat(&slot_data).access_mut(self) = Seat::Active(SeatInfo {
|
||||
slot: slot_idx as u8,
|
||||
seqno: cur_seqno,
|
||||
});
|
||||
|
||||
self.slots[slot_idx] = Slot::Active(SlotInfo {
|
||||
slot_data,
|
||||
seqno: cur_seqno,
|
||||
});
|
||||
|
||||
self.use_seqno += 1;
|
||||
}
|
||||
|
||||
/// Reactivates an active/idle slot for a given seat without reprogramming the hardware.
|
||||
/// The SlotManager reuses the existing slot_data. This ensures that the hardware-specific
|
||||
/// information is not changed between subsequent uses. It also ensures that resources
|
||||
/// owned by the existing slot_data remain alive while the hardware is configured to use them.
|
||||
fn reactivate_slot(&mut self, slot_idx: usize, slot_data: &T::SlotData) -> Result {
|
||||
let cur_seqno = self.use_seqno;
|
||||
|
||||
let mut slot_info = match mem::take(&mut self.slots[slot_idx]) {
|
||||
Slot::Active(slot_info) | Slot::Idle(slot_info) => slot_info,
|
||||
Slot::Free => {
|
||||
*T::seat(slot_data).access_mut(self) = Seat::NoSeat;
|
||||
return Err(EINVAL);
|
||||
}
|
||||
};
|
||||
|
||||
*T::seat(slot_data).access_mut(self) = Seat::Active(SeatInfo {
|
||||
slot: slot_idx as u8,
|
||||
seqno: cur_seqno,
|
||||
});
|
||||
|
||||
slot_info.seqno = cur_seqno;
|
||||
self.slots[slot_idx] = Slot::Active(slot_info);
|
||||
|
||||
self.use_seqno += 1;
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Activates a slot for the given seat.
|
||||
fn activate_slot(&mut self, slot_idx: usize, slot_data: T::SlotData) -> Result {
|
||||
self.manager.activate(slot_idx, &slot_data)?;
|
||||
self.record_active_slot(slot_idx, slot_data);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Finds a slot for the given seat. A free slot is preferred, but if none
|
||||
/// are available, the oldest idle slot is evicted and reused. Otherwise, if
|
||||
/// there are no free or idle slots, return [`EBUSY`].
|
||||
fn allocate_slot(&mut self, slot_data: T::SlotData) -> Result {
|
||||
let slots = &self.slots[..self.slot_count];
|
||||
|
||||
let mut idle_slot_idx = None;
|
||||
let mut idle_slot_seqno: u64 = 0;
|
||||
|
||||
for (slot_idx, slot) in slots.iter().enumerate() {
|
||||
match slot {
|
||||
Slot::Free => {
|
||||
return self.activate_slot(slot_idx, slot_data);
|
||||
}
|
||||
Slot::Idle(slot_info) => {
|
||||
if idle_slot_idx.is_none() || slot_info.seqno < idle_slot_seqno {
|
||||
idle_slot_idx = Some(slot_idx);
|
||||
idle_slot_seqno = slot_info.seqno;
|
||||
}
|
||||
}
|
||||
Slot::Active(_) => (),
|
||||
}
|
||||
}
|
||||
|
||||
match idle_slot_idx {
|
||||
Some(slot_idx) => {
|
||||
// Lazily evict idle slot just before it is reused.
|
||||
if let Slot::Idle(slot_info) = &self.slots[slot_idx] {
|
||||
self.manager.evict(slot_idx, &slot_info.slot_data)?;
|
||||
mem::take(&mut self.slots[slot_idx]);
|
||||
}
|
||||
self.activate_slot(slot_idx, slot_data)
|
||||
}
|
||||
None => Err(EBUSY),
|
||||
}
|
||||
}
|
||||
|
||||
/// Converts an active slot and its seat to idle state.
|
||||
fn idle_slot(&mut self, slot_idx: usize, locked_seat: &LockedSeat<T, MAX_SLOTS>) -> Result {
|
||||
let slot = mem::take(&mut self.slots[slot_idx]);
|
||||
|
||||
self.slots[slot_idx] = match slot {
|
||||
// If the slot was active, make it idle.
|
||||
Slot::Active(slot_info) => Slot::Idle(slot_info),
|
||||
|
||||
// Preserve an already-idle slot.
|
||||
Slot::Idle(slot_info) => Slot::Idle(slot_info),
|
||||
|
||||
// A free slot remains free.
|
||||
Slot::Free => Slot::Free,
|
||||
};
|
||||
|
||||
// If the seat was active, make it idle, or keep it idle if it was already idle.
|
||||
*locked_seat.access_mut(self) = match locked_seat.access(self) {
|
||||
Seat::Active(seat_info) | Seat::Idle(seat_info) => Seat::Idle(SeatInfo {
|
||||
slot: seat_info.slot,
|
||||
seqno: seat_info.seqno,
|
||||
}),
|
||||
Seat::NoSeat => Seat::NoSeat,
|
||||
};
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Evicts an active or idle slot: calls the eviction callback and marks the slot as free
|
||||
/// and the seat as NoSeat.
|
||||
fn evict_slot(&mut self, slot_idx: usize, locked_seat: &LockedSeat<T, MAX_SLOTS>) -> Result {
|
||||
match &self.slots[slot_idx] {
|
||||
Slot::Active(slot_info) | Slot::Idle(slot_info) => {
|
||||
// If hardware eviction fails (e.g. times out), the slot retains
|
||||
// its SlotData so that any resources still referenced by the hardware
|
||||
// will remain alive. This prevents use-after-free errors.
|
||||
self.manager.evict(slot_idx, &slot_info.slot_data)?;
|
||||
mem::take(&mut self.slots[slot_idx]);
|
||||
}
|
||||
_ => (),
|
||||
}
|
||||
|
||||
*locked_seat.access_mut(self) = Seat::NoSeat;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Checks that the seat state matches the slot's state.
|
||||
/// If they don't match, the seat is stale and is reset to `NoSeat`.
|
||||
fn check_seat(&mut self, locked_seat: &LockedSeat<T, MAX_SLOTS>) {
|
||||
let (slot_idx, seat_seqno, is_active) = match locked_seat.access(self) {
|
||||
Seat::Active(seat_info) => (seat_info.slot as usize, seat_info.seqno, true),
|
||||
Seat::Idle(seat_info) => (seat_info.slot as usize, seat_info.seqno, false),
|
||||
_ => return,
|
||||
};
|
||||
|
||||
let valid = if is_active {
|
||||
!kernel::warn_on!(!matches!(
|
||||
&self.slots[slot_idx],
|
||||
Slot::Active(slot_info) if slot_info.seqno == seat_seqno
|
||||
))
|
||||
} else {
|
||||
matches!(
|
||||
&self.slots[slot_idx],
|
||||
Slot::Idle(slot_info) if slot_info.seqno == seat_seqno
|
||||
)
|
||||
};
|
||||
|
||||
if !valid {
|
||||
*locked_seat.access_mut(self) = Seat::NoSeat;
|
||||
}
|
||||
}
|
||||
|
||||
/// Activates a resource on any available/reclaimable slot.
|
||||
pub(crate) fn activate(&mut self, slot_data: T::SlotData) -> Result {
|
||||
self.check_seat(T::seat(&slot_data));
|
||||
|
||||
// Copy out only the slot index so the borrow of slot_data ends here.
|
||||
let slot_idx = match T::seat(&slot_data).access(self) {
|
||||
Seat::Active(seat_info) | Seat::Idle(seat_info) => Some(seat_info.slot as usize),
|
||||
Seat::NoSeat => None,
|
||||
};
|
||||
|
||||
match slot_idx {
|
||||
Some(slot_idx) => self.reactivate_slot(slot_idx, &slot_data),
|
||||
None => self.allocate_slot(slot_data),
|
||||
}
|
||||
}
|
||||
|
||||
/// Flag a resource as idle. This method will be used for user VM support.
|
||||
#[expect(dead_code)]
|
||||
pub(crate) fn idle(&mut self, locked_seat: &LockedSeat<T, MAX_SLOTS>) -> Result {
|
||||
self.check_seat(locked_seat);
|
||||
if let Seat::Active(seat_info) = locked_seat.access(self) {
|
||||
self.idle_slot(seat_info.slot as usize, locked_seat)?;
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Evict a resource from its slot.
|
||||
pub(crate) fn evict(&mut self, locked_seat: &LockedSeat<T, MAX_SLOTS>) -> Result {
|
||||
self.check_seat(locked_seat);
|
||||
|
||||
match locked_seat.access(self) {
|
||||
Seat::Active(seat_info) | Seat::Idle(seat_info) => {
|
||||
let slot_idx = seat_info.slot as usize;
|
||||
self.evict_slot(slot_idx, locked_seat)?;
|
||||
}
|
||||
_ => (),
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: SlotOperations<MAX_SLOTS>, const MAX_SLOTS: usize> Deref for SlotManager<T, MAX_SLOTS> {
|
||||
type Target = T;
|
||||
|
||||
fn deref(&self) -> &Self::Target {
|
||||
&self.manager
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: SlotOperations<MAX_SLOTS>, const MAX_SLOTS: usize> DerefMut for SlotManager<T, MAX_SLOTS> {
|
||||
fn deref_mut(&mut self) -> &mut Self::Target {
|
||||
&mut self.manager
|
||||
}
|
||||
}
|
||||
|
|
@ -9,9 +9,13 @@
|
|||
|
||||
mod driver;
|
||||
mod file;
|
||||
mod fw;
|
||||
mod gem;
|
||||
mod gpu;
|
||||
mod mmu;
|
||||
mod regs;
|
||||
mod slot;
|
||||
mod vm;
|
||||
|
||||
kernel::module_platform_driver! {
|
||||
type: TyrPlatformDriver,
|
||||
|
|
|
|||
950
drivers/gpu/drm/tyr/vm.rs
Normal file
950
drivers/gpu/drm/tyr/vm.rs
Normal file
|
|
@ -0,0 +1,950 @@
|
|||
// SPDX-License-Identifier: GPL-2.0 or MIT
|
||||
|
||||
//! GPU virtual memory management using the DRM GPUVM framework.
|
||||
//!
|
||||
//! This module manages GPU virtual address spaces, providing memory isolation and
|
||||
//! the illusion of owning the entire virtual address (VA) range, similar to CPU virtual memory.
|
||||
//! Each virtual memory (VM) area is backed by ARM64 LPAE Stage 1 page tables and can be
|
||||
//! mapped into hardware address space (AS) slots for GPU execution.
|
||||
|
||||
use core::marker::PhantomData;
|
||||
use core::ops::Range;
|
||||
|
||||
use kernel::{
|
||||
device::{
|
||||
Bound,
|
||||
Device, //
|
||||
},
|
||||
drm::{
|
||||
gem::BaseObject,
|
||||
gpuvm::{
|
||||
DriverGpuVm,
|
||||
GpuVaAlloc,
|
||||
GpuVm,
|
||||
GpuVmBo,
|
||||
OpMap,
|
||||
OpMapRequest,
|
||||
OpMapped,
|
||||
OpRemap,
|
||||
OpRemapped,
|
||||
OpUnmap,
|
||||
OpUnmapped,
|
||||
UniqueRefGpuVm, //
|
||||
}, //
|
||||
},
|
||||
fmt,
|
||||
impl_flags,
|
||||
io::PhysAddr,
|
||||
iommu::pgtable::{
|
||||
prot,
|
||||
IoPageTable,
|
||||
ARM64LPAES1, //
|
||||
},
|
||||
new_mutex,
|
||||
prelude::*,
|
||||
sizes::{
|
||||
SZ_1G,
|
||||
SZ_2M,
|
||||
SZ_4K, //
|
||||
},
|
||||
sync::{
|
||||
aref::ARef,
|
||||
Arc,
|
||||
ArcBorrow,
|
||||
Mutex, //
|
||||
},
|
||||
uapi, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::{
|
||||
TyrDrmDevice,
|
||||
TyrDrmDriver, //
|
||||
},
|
||||
gem,
|
||||
gem::Bo,
|
||||
gpu::GpuInfo,
|
||||
mmu::{
|
||||
address_space::VmAsData,
|
||||
Mmu, //
|
||||
},
|
||||
regs::gpu_control::MMU_FEATURES,
|
||||
};
|
||||
|
||||
impl_flags!(
|
||||
/// Flags controlling virtual memory mapping behavior.
|
||||
///
|
||||
/// These flags control access permissions and caching behavior for GPU virtual
|
||||
/// memory mappings.
|
||||
#[derive(Debug, Clone, Default, Copy, PartialEq, Eq)]
|
||||
pub(crate) struct VmMapFlags(u32);
|
||||
|
||||
/// Individual flags that can be combined in [`VmMapFlags`].
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub(crate) enum VmFlag {
|
||||
/// Map as read-only.
|
||||
Readonly = uapi::drm_panthor_vm_bind_op_flags_DRM_PANTHOR_VM_BIND_OP_MAP_READONLY as u32,
|
||||
/// Map as non-executable.
|
||||
Noexec = uapi::drm_panthor_vm_bind_op_flags_DRM_PANTHOR_VM_BIND_OP_MAP_NOEXEC as u32,
|
||||
/// Map as uncached.
|
||||
Uncached = uapi::drm_panthor_vm_bind_op_flags_DRM_PANTHOR_VM_BIND_OP_MAP_UNCACHED as u32,
|
||||
}
|
||||
);
|
||||
|
||||
impl VmMapFlags {
|
||||
/// Convert the flags to `pgtable::prot`.
|
||||
fn to_prot(self) -> u32 {
|
||||
let mut prot = 0;
|
||||
|
||||
if self.contains(VmFlag::Readonly) {
|
||||
prot |= prot::READ;
|
||||
} else {
|
||||
prot |= prot::READ | prot::WRITE;
|
||||
}
|
||||
|
||||
if self.contains(VmFlag::Noexec) {
|
||||
prot |= prot::NOEXEC;
|
||||
}
|
||||
|
||||
if !self.contains(VmFlag::Uncached) {
|
||||
prot |= prot::CACHE;
|
||||
}
|
||||
|
||||
prot
|
||||
}
|
||||
}
|
||||
|
||||
impl fmt::Display for VmMapFlags {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
let mut first = true;
|
||||
|
||||
if self.contains(VmFlag::Readonly) {
|
||||
write!(f, "READONLY")?;
|
||||
first = false;
|
||||
}
|
||||
if self.contains(VmFlag::Noexec) {
|
||||
if !first {
|
||||
write!(f, " | ")?;
|
||||
}
|
||||
write!(f, "NOEXEC")?;
|
||||
first = false;
|
||||
}
|
||||
|
||||
if self.contains(VmFlag::Uncached) {
|
||||
if !first {
|
||||
write!(f, " | ")?;
|
||||
}
|
||||
write!(f, "UNCACHED")?;
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
impl TryFrom<u32> for VmMapFlags {
|
||||
type Error = Error;
|
||||
|
||||
fn try_from(value: u32) -> Result<Self, Self::Error> {
|
||||
let valid = VmFlag::Readonly as u32 | VmFlag::Noexec as u32 | VmFlag::Uncached as u32;
|
||||
|
||||
if value & !valid != 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
Ok(Self(value))
|
||||
}
|
||||
}
|
||||
|
||||
/// Arguments for a virtual memory map operation.
|
||||
struct VmMapArgs<'drm> {
|
||||
/// Access permissions and caching behavior for the mapping.
|
||||
flags: VmMapFlags,
|
||||
/// GEM buffer object registered with the GPUVM framework.
|
||||
vm_bo: ARef<GpuVmBo<GpuVmData<'drm>>>,
|
||||
/// Offset in bytes from the start of the buffer object.
|
||||
bo_offset: u64,
|
||||
}
|
||||
|
||||
/// Type of virtual memory operation.
|
||||
enum VmOpType<'drm> {
|
||||
/// Map a GEM buffer object into the virtual address space.
|
||||
Map(VmMapArgs<'drm>),
|
||||
/// Unmap a region from the virtual address space.
|
||||
Unmap,
|
||||
}
|
||||
|
||||
/// Preallocated resources needed to execute a VM operation.
|
||||
///
|
||||
/// VM operations may require allocating new GPUVA objects to track mappings.
|
||||
/// To avoid allocation failures during the operation, preallocate the
|
||||
/// maximum number of GPUVAs that might be needed.
|
||||
struct VmOpResources<'drm> {
|
||||
/// Preallocated GPUVA objects for remap operations.
|
||||
///
|
||||
/// Partial unmap requests or map requests overlapping existing mappings
|
||||
/// will trigger a remap call, which needs to register up to three VA
|
||||
/// objects (one for the new mapping, and two for the previous and next
|
||||
/// mappings).
|
||||
preallocated_gpuvas: [Option<GpuVaAlloc<GpuVmData<'drm>>>; 3],
|
||||
}
|
||||
|
||||
/// Request to execute a virtual memory operation.
|
||||
struct VmOpRequest<'drm> {
|
||||
/// Request type.
|
||||
op_type: VmOpType<'drm>,
|
||||
|
||||
/// Region of the virtual address space covered by this request.
|
||||
region: Range<u64>,
|
||||
}
|
||||
|
||||
/// Arguments for a page table map operation.
|
||||
struct PtMapArgs {
|
||||
/// Memory protection flags describing allowed accesses for this mapping.
|
||||
///
|
||||
/// This is directly derived from [`VmMapFlags`] via [`VmMapFlags::to_prot`].
|
||||
prot: u32,
|
||||
}
|
||||
|
||||
/// Type of page table operation.
|
||||
enum PtOpType {
|
||||
/// Map pages into the page table.
|
||||
Map(PtMapArgs),
|
||||
/// Unmap pages from the page table.
|
||||
Unmap,
|
||||
}
|
||||
|
||||
/// Context for updating the GPU page table.
|
||||
///
|
||||
/// This context is created when beginning a page table update operation and
|
||||
/// automatically flushes changes when dropped. It ensures that the
|
||||
/// Memory Management Unit (MMU) state is properly managed and Translation
|
||||
/// Lookaside Buffer (TLB) entries are flushed.
|
||||
pub(crate) struct PtUpdateContext<'ctx, 'drm> {
|
||||
/// Device used for DMA-mapping GEM shmem SG tables.
|
||||
dev: &'ctx Device<Bound>,
|
||||
|
||||
/// Page table.
|
||||
pt: &'ctx IoPageTable<'drm, ARM64LPAES1>,
|
||||
|
||||
/// MMU manager.
|
||||
mmu: &'ctx Mmu<'drm>,
|
||||
|
||||
/// Reference to the address space data to pass to the MMU functions.
|
||||
as_data: &'ctx VmAsData<'drm>,
|
||||
|
||||
/// Region of the virtual address space covered by this request.
|
||||
region: Range<u64>,
|
||||
|
||||
/// Operation type.
|
||||
op_type: PtOpType,
|
||||
|
||||
/// Preallocated resources that can be used when executing the request.
|
||||
resources: &'ctx mut VmOpResources<'drm>,
|
||||
}
|
||||
|
||||
impl<'ctx, 'drm> PtUpdateContext<'ctx, 'drm> {
|
||||
/// Creates a new page table update context.
|
||||
///
|
||||
/// This prepares the MMU for a page table update.
|
||||
/// The context will automatically flush the TLB and
|
||||
/// complete the update when dropped.
|
||||
fn new(
|
||||
dev: &'ctx Device<Bound>,
|
||||
pt: &'ctx IoPageTable<'drm, ARM64LPAES1>,
|
||||
mmu: &'ctx Mmu<'drm>,
|
||||
as_data: &'ctx VmAsData<'drm>,
|
||||
region: Range<u64>,
|
||||
op_type: PtOpType,
|
||||
resources: &'ctx mut VmOpResources<'drm>,
|
||||
) -> Result<PtUpdateContext<'ctx, 'drm>> {
|
||||
mmu.start_vm_update(as_data, ®ion)?;
|
||||
|
||||
Ok(Self {
|
||||
dev,
|
||||
pt,
|
||||
mmu,
|
||||
as_data,
|
||||
region,
|
||||
op_type,
|
||||
resources,
|
||||
})
|
||||
}
|
||||
|
||||
/// Finds one of our pre-allocated VAs.
|
||||
fn preallocated_gpuva(&mut self) -> Result<GpuVaAlloc<GpuVmData<'drm>>> {
|
||||
self.resources
|
||||
.preallocated_gpuvas
|
||||
.iter_mut()
|
||||
.find_map(|f| f.take())
|
||||
.ok_or(EINVAL)
|
||||
}
|
||||
|
||||
/// Returns an unused GPUVA object to the preallocated pool.
|
||||
/// If the pool is already full, the unused allocation is simply dropped.
|
||||
fn return_preallocated_gpuva(&mut self, gpuva: GpuVaAlloc<GpuVmData<'drm>>) {
|
||||
if let Some(slot) = self
|
||||
.resources
|
||||
.preallocated_gpuvas
|
||||
.iter_mut()
|
||||
.find(|slot| slot.is_none())
|
||||
{
|
||||
*slot = Some(gpuva);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl Drop for PtUpdateContext<'_, '_> {
|
||||
fn drop(&mut self) {
|
||||
if let Err(e) = self.mmu.end_vm_update(self.as_data) {
|
||||
dev_err!(self.dev, "Failed to end VM update {:?}", e);
|
||||
}
|
||||
|
||||
if let Err(e) = self.mmu.flush_vm(self.as_data) {
|
||||
dev_err!(self.dev, "Failed to flush VM {:?}", e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Driver implementation for the GPUVM framework.
|
||||
///
|
||||
/// Implements [`DriverGpuVm`] to provide VM operation callbacks (map, unmap, remap)
|
||||
/// and associated types for buffer objects, virtual addresses, and contexts.
|
||||
pub(crate) struct GpuVmData<'drm> {
|
||||
_phantom: PhantomData<&'drm ()>,
|
||||
}
|
||||
|
||||
/// GPU virtual address space.
|
||||
///
|
||||
/// Each VM can be mapped into a hardware address space slot.
|
||||
#[pin_data]
|
||||
pub(crate) struct Vm<'drm> {
|
||||
/// Data referenced by an AS when the VM is active
|
||||
as_data: Arc<VmAsData<'drm>>,
|
||||
/// MMU manager.
|
||||
mmu: Arc<Mmu<'drm>>,
|
||||
/// Parent device used for DMA mapping and page-table operations.
|
||||
dev: &'drm Device<Bound>,
|
||||
/// DRM GPUVM core for managing virtual address space.
|
||||
#[pin]
|
||||
gpuvm_unique: Mutex<UniqueRefGpuVm<GpuVmData<'drm>>>,
|
||||
/// Non-core part of the GPUVM. Can be used for stuff that doesn't modify the
|
||||
/// internal mapping tree, like GpuVm::obtain()
|
||||
gpuvm: ARef<GpuVm<GpuVmData<'drm>>>,
|
||||
/// VA range for this VM.
|
||||
va_range: Range<u64>,
|
||||
}
|
||||
|
||||
impl<'drm> Vm<'drm> {
|
||||
/// Creates a new GPU virtual address space.
|
||||
///
|
||||
/// The VM is initialized with a page table configured according to the GPU's
|
||||
/// address translation capabilities and registered with the GPUVM framework.
|
||||
pub(crate) fn new(
|
||||
dev: &'drm Device<Bound>,
|
||||
ddev: &TyrDrmDevice,
|
||||
mmu: ArcBorrow<'_, Mmu<'drm>>,
|
||||
gpu_info: &GpuInfo,
|
||||
) -> Result<Arc<Vm<'drm>>> {
|
||||
let mmu_features = MMU_FEATURES::from_raw(gpu_info.mmu_features);
|
||||
let va_bits = mmu_features.va_bits().get();
|
||||
let pa_bits = mmu_features.pa_bits().get();
|
||||
|
||||
let range = 0..(1u64 << va_bits);
|
||||
let reserve_range = 0..0u64;
|
||||
|
||||
// dummy_obj is used to initialize the GPUVM tree.
|
||||
let dummy_obj = gem::new_dummy_object(ddev).inspect_err(|e| {
|
||||
dev_err!(dev, "Failed to create dummy GEM object: {:?}", e);
|
||||
})?;
|
||||
|
||||
let gpuvm_unique = GpuVm::new::<Error, _>(
|
||||
c"Tyr::GpuVm",
|
||||
ddev,
|
||||
&*dummy_obj,
|
||||
range.clone(),
|
||||
reserve_range,
|
||||
GpuVmData::<'drm> {
|
||||
_phantom: PhantomData::<&()>,
|
||||
},
|
||||
)
|
||||
.inspect_err(|e| {
|
||||
dev_err!(dev, "Failed to create GpuVm: {:?}", e);
|
||||
})?;
|
||||
let gpuvm = ARef::from(&*gpuvm_unique);
|
||||
|
||||
let as_data = Arc::pin_init(VmAsData::new(&mmu, dev, va_bits, pa_bits), GFP_KERNEL)?;
|
||||
|
||||
let vm = Arc::pin_init(
|
||||
pin_init!(Self{
|
||||
as_data,
|
||||
dev,
|
||||
mmu: mmu.into(),
|
||||
gpuvm,
|
||||
gpuvm_unique <- new_mutex!(gpuvm_unique),
|
||||
va_range: range,
|
||||
}),
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
|
||||
Ok(vm)
|
||||
}
|
||||
|
||||
/// Returns the parent device used by this VM for DMA mapping and page-table operations.
|
||||
pub(crate) fn dev(&self) -> &'drm Device<Bound> {
|
||||
self.dev
|
||||
}
|
||||
|
||||
/// Activate the VM in a hardware address space slot.
|
||||
pub(crate) fn activate(&self) -> Result {
|
||||
self.mmu
|
||||
.activate_vm(self.as_data.as_arc_borrow())
|
||||
.inspect_err(|e| {
|
||||
dev_err!(self.dev, "Failed to activate VM: {:?}", e);
|
||||
})
|
||||
}
|
||||
|
||||
/// Deactivate the VM by evicting it from its address space slot.
|
||||
fn deactivate(&self) -> Result {
|
||||
self.mmu.deactivate_vm(&self.as_data).inspect_err(|e| {
|
||||
dev_err!(self.dev, "Failed to deactivate VM: {:?}", e);
|
||||
})
|
||||
}
|
||||
|
||||
/// Kills the VM by deactivating it and unmapping all regions.
|
||||
pub(crate) fn kill(&self) {
|
||||
// TODO: Turn the VM into a state where it can't be used.
|
||||
let _ = self.deactivate();
|
||||
let _ = self
|
||||
.unmap_range(self.va_range.start, self.va_range.end - self.va_range.start)
|
||||
.inspect_err(|e| {
|
||||
dev_err!(self.dev, "Failed to unmap range during deactivate: {:?}", e);
|
||||
});
|
||||
}
|
||||
|
||||
/// Executes a virtual memory operation.
|
||||
///
|
||||
/// This handles both map and unmap operations by coordinating between the
|
||||
/// GPUVM framework and the hardware page table.
|
||||
fn exec_op<'a>(
|
||||
&self,
|
||||
gpuvm_unique: &mut UniqueRefGpuVm<GpuVmData<'drm>>,
|
||||
req: VmOpRequest<'drm>,
|
||||
resources: &'a mut VmOpResources<'drm>,
|
||||
) -> Result {
|
||||
let pt = &self.as_data.page_table;
|
||||
|
||||
match req.op_type {
|
||||
VmOpType::Map(args) => {
|
||||
let mut pt_upd = PtUpdateContext::new(
|
||||
self.dev,
|
||||
pt,
|
||||
&self.mmu,
|
||||
&self.as_data,
|
||||
req.region,
|
||||
PtOpType::Map(PtMapArgs {
|
||||
prot: args.flags.to_prot(),
|
||||
}),
|
||||
resources,
|
||||
)?;
|
||||
|
||||
gpuvm_unique.sm_map(OpMapRequest {
|
||||
addr: pt_upd.region.start,
|
||||
range: pt_upd.region.end - pt_upd.region.start,
|
||||
gem_offset: args.bo_offset,
|
||||
vm_bo: &args.vm_bo,
|
||||
context: &mut pt_upd,
|
||||
})
|
||||
//PtUpdateContext drops here flushing the page table
|
||||
}
|
||||
VmOpType::Unmap => {
|
||||
let mut pt_upd = PtUpdateContext::new(
|
||||
self.dev,
|
||||
pt,
|
||||
&self.mmu,
|
||||
&self.as_data,
|
||||
req.region,
|
||||
PtOpType::Unmap,
|
||||
resources,
|
||||
)?;
|
||||
|
||||
gpuvm_unique.sm_unmap(
|
||||
pt_upd.region.start,
|
||||
pt_upd.region.end - pt_upd.region.start,
|
||||
&mut pt_upd,
|
||||
)
|
||||
//PtUpdateContext drops here flushing the page table
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Maps a GEM buffer object range into the VM at the specified virtual address.
|
||||
///
|
||||
/// This creates a mapping from GPU virtual address `va` to the physical pages
|
||||
/// backing the GEM object, starting at `bo_offset` bytes into the object and
|
||||
/// spanning `map_size` bytes. The mapping respects the access permissions and
|
||||
/// caching behavior specified in `flags`.
|
||||
pub(crate) fn map_bo_range(
|
||||
&self,
|
||||
bo: &Bo,
|
||||
bo_offset: u64,
|
||||
map_size: u64,
|
||||
va: u64,
|
||||
flags: VmMapFlags,
|
||||
) -> Result {
|
||||
if map_size == 0
|
||||
|| va % SZ_4K as u64 != 0
|
||||
|| bo_offset % SZ_4K as u64 != 0
|
||||
|| map_size % SZ_4K as u64 != 0
|
||||
{
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let bo_size = u64::try_from(bo.size()).map_err(|_| EOVERFLOW)?;
|
||||
let bo_end = bo_offset.checked_add(map_size).ok_or(EINVAL)?;
|
||||
|
||||
if bo_end > bo_size {
|
||||
dev_err!(
|
||||
self.dev,
|
||||
"BO mapping range {:#x}..{:#x} exceeds BO size {:#x}",
|
||||
bo_offset,
|
||||
bo_end,
|
||||
bo_size
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let va_end: u64 = va.checked_add(map_size).ok_or(EINVAL)?;
|
||||
|
||||
let req = VmOpRequest {
|
||||
op_type: VmOpType::Map(VmMapArgs {
|
||||
vm_bo: self.gpuvm.obtain(bo, ())?,
|
||||
flags,
|
||||
bo_offset,
|
||||
}),
|
||||
region: va..va_end,
|
||||
};
|
||||
let mut resources = VmOpResources {
|
||||
preallocated_gpuvas: [
|
||||
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
|
||||
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
|
||||
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
|
||||
],
|
||||
};
|
||||
let result = {
|
||||
let mut gpuvm_unique = self.gpuvm_unique.lock();
|
||||
self.exec_op(gpuvm_unique.as_mut().get_mut(), req, &mut resources)
|
||||
};
|
||||
// We flush the defer cleanup list now. Things will be different in
|
||||
// the asynchronous VM_BIND path, where we want the cleanup to
|
||||
// happen outside the DMA signalling path.
|
||||
self.gpuvm.deferred_cleanup();
|
||||
result
|
||||
}
|
||||
|
||||
/// Unmaps a virtual address range from the VM.
|
||||
///
|
||||
/// This removes any existing mappings in the specified range, freeing the
|
||||
/// virtual address space for reuse.
|
||||
pub(crate) fn unmap_range(&self, va: u64, size: u64) -> Result {
|
||||
if size == 0 || va % SZ_4K as u64 != 0 || size % SZ_4K as u64 != 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let end = va.checked_add(size).ok_or(EINVAL)?;
|
||||
|
||||
if va < self.va_range.start || end > self.va_range.end {
|
||||
dev_err!(
|
||||
self.dev,
|
||||
"Unmap range {:#x}..{:#x} exceeds VM range {:#x}..{:#x}",
|
||||
va,
|
||||
end,
|
||||
self.va_range.start,
|
||||
self.va_range.end
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let req = VmOpRequest {
|
||||
op_type: VmOpType::Unmap,
|
||||
region: va..end,
|
||||
};
|
||||
|
||||
let full_vm = va == self.va_range.start && end == self.va_range.end;
|
||||
|
||||
let mut resources = VmOpResources {
|
||||
preallocated_gpuvas: if full_vm {
|
||||
// Unmapping the entire VM cannot split an existing mapping,
|
||||
// so no GPUVA objects are needed for remap operations.
|
||||
[None, None, None]
|
||||
} else {
|
||||
[
|
||||
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
|
||||
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
|
||||
Some(GpuVaAlloc::<GpuVmData<'drm>>::new(GFP_KERNEL)?),
|
||||
]
|
||||
},
|
||||
};
|
||||
let result = {
|
||||
let mut gpuvm_unique = self.gpuvm_unique.lock();
|
||||
self.exec_op(gpuvm_unique.as_mut().get_mut(), req, &mut resources)
|
||||
};
|
||||
// We flush the defer cleanup list now. Things will be different in
|
||||
// the asynchronous VM_BIND path, where we want the cleanup to
|
||||
// happen outside the DMA signalling path.
|
||||
self.gpuvm.deferred_cleanup();
|
||||
result
|
||||
}
|
||||
}
|
||||
|
||||
impl<'drm> DriverGpuVm for GpuVmData<'drm> {
|
||||
type Driver = TyrDrmDriver;
|
||||
type Object = Bo;
|
||||
type VmBoData = ();
|
||||
type VaData = ();
|
||||
type SmContext<'ctx>
|
||||
= PtUpdateContext<'ctx, 'drm>
|
||||
where
|
||||
Self: 'ctx;
|
||||
|
||||
/// Create a new mapping.
|
||||
fn sm_step_map<'op>(
|
||||
&mut self,
|
||||
op: OpMap<'op, Self>,
|
||||
context: &mut Self::SmContext<'_>,
|
||||
) -> Result<OpMapped<'op, Self>, Error> {
|
||||
let start_iova = op.addr();
|
||||
let mut iova = start_iova;
|
||||
let mut bytes_left_to_map = op.length();
|
||||
let mut gem_offset = op.gem_offset();
|
||||
|
||||
// Make sure that the end of the requested GEM range doesn't run past the
|
||||
// end of the GEM buffer itself.
|
||||
let gem_range_end = op.gem_offset().checked_add(op.length()).ok_or(EINVAL)?;
|
||||
|
||||
if gem_range_end > op.obj().size() as u64 {
|
||||
dev_err!(
|
||||
context.dev,
|
||||
"Requested GEM range ends at {} which is beyond the GEM buffer size {}",
|
||||
gem_range_end,
|
||||
op.obj().size()
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let sgt = op.obj().sg_table(context.dev).inspect_err(|e| {
|
||||
dev_err!(context.dev, "Failed to get sg_table: {:?}", e);
|
||||
})?;
|
||||
let prot = match &context.op_type {
|
||||
PtOpType::Map(args) => args.prot,
|
||||
_ => {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
};
|
||||
|
||||
for sgt_entry in sgt.iter() {
|
||||
// Expressly convert to u64 to work with arm 32-bit builds.
|
||||
#[allow(clippy::useless_conversion)]
|
||||
let mut paddr = u64::from(sgt_entry.dma_address());
|
||||
#[allow(clippy::useless_conversion)]
|
||||
let mut sgt_entry_length = u64::from(sgt_entry.dma_len());
|
||||
|
||||
if bytes_left_to_map == 0 {
|
||||
break;
|
||||
}
|
||||
|
||||
if gem_offset > 0 {
|
||||
// Skip the entire SGT entry if the gem_offset exceeds its length.
|
||||
let skip = u64::min(sgt_entry_length, gem_offset);
|
||||
paddr += skip;
|
||||
sgt_entry_length -= skip;
|
||||
gem_offset -= skip;
|
||||
}
|
||||
|
||||
if sgt_entry_length == 0 {
|
||||
continue;
|
||||
}
|
||||
|
||||
let len = u64::min(sgt_entry_length, bytes_left_to_map);
|
||||
|
||||
let segment_mapped = match pt_map(context.dev, context.pt, iova, paddr, len, prot) {
|
||||
Ok(segment_mapped) => segment_mapped,
|
||||
Err(e) => {
|
||||
// clean up any successful mappings from previous SGT entries.
|
||||
let total_mapped = iova - start_iova;
|
||||
if total_mapped > 0 {
|
||||
let _ = pt_unmap(
|
||||
context.dev,
|
||||
context.pt,
|
||||
start_iova..(start_iova + total_mapped),
|
||||
);
|
||||
}
|
||||
return Err(e);
|
||||
}
|
||||
};
|
||||
|
||||
bytes_left_to_map -= segment_mapped;
|
||||
iova += segment_mapped;
|
||||
}
|
||||
|
||||
if bytes_left_to_map != 0 {
|
||||
let total_mapped = iova - start_iova;
|
||||
|
||||
if total_mapped > 0 {
|
||||
let _ = pt_unmap(context.dev, context.pt, start_iova..iova);
|
||||
}
|
||||
|
||||
dev_err!(
|
||||
context.dev,
|
||||
"SG table is too small for requested mapping: {} bytes remain",
|
||||
bytes_left_to_map
|
||||
);
|
||||
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let gpuva = context.preallocated_gpuva()?;
|
||||
let op = op.insert(gpuva, pin_init::init_zeroed());
|
||||
|
||||
Ok(op)
|
||||
}
|
||||
|
||||
/// Indicates that an existing mapping should be removed.
|
||||
fn sm_step_unmap<'op>(
|
||||
&mut self,
|
||||
op: OpUnmap<'op, Self>,
|
||||
context: &mut Self::SmContext<'_>,
|
||||
) -> Result<OpUnmapped<'op, Self>, Error> {
|
||||
let start_iova = op.va().addr();
|
||||
let length = op.va().length();
|
||||
|
||||
let region = start_iova..(start_iova + length);
|
||||
pt_unmap(context.dev, context.pt, region.clone()).inspect_err(|e| {
|
||||
dev_err!(
|
||||
context.dev,
|
||||
"Failed to unmap region {:#x}..{:#x}: {:?}",
|
||||
region.start,
|
||||
region.end,
|
||||
e
|
||||
);
|
||||
})?;
|
||||
|
||||
let (op_unmapped, _va_removed) = op.remove();
|
||||
|
||||
Ok(op_unmapped)
|
||||
}
|
||||
|
||||
/// Split up an existing mapping.
|
||||
fn sm_step_remap<'op>(
|
||||
&mut self,
|
||||
op: OpRemap<'op, Self>,
|
||||
context: &mut Self::SmContext<'_>,
|
||||
) -> Result<OpRemapped<'op, Self>, Error> {
|
||||
let unmap_start = if let Some(prev) = op.prev() {
|
||||
prev.addr() + prev.length()
|
||||
} else {
|
||||
op.va_to_unmap().addr()
|
||||
};
|
||||
|
||||
let unmap_end = if let Some(next) = op.next() {
|
||||
next.addr()
|
||||
} else {
|
||||
op.va_to_unmap().addr() + op.va_to_unmap().length()
|
||||
};
|
||||
|
||||
let unmap_length = unmap_end - unmap_start;
|
||||
|
||||
if unmap_length > 0 {
|
||||
let region = unmap_start..(unmap_start + unmap_length);
|
||||
pt_unmap(context.dev, context.pt, region.clone()).inspect_err(|e| {
|
||||
dev_err!(
|
||||
context.dev,
|
||||
"Failed to unmap remap region {:#x}..{:#x}: {:?}",
|
||||
region.start,
|
||||
region.end,
|
||||
e
|
||||
);
|
||||
})?;
|
||||
}
|
||||
|
||||
let prev_va = context.preallocated_gpuva()?;
|
||||
let next_va = context.preallocated_gpuva()?;
|
||||
|
||||
let (op_remapped, remap_ret) = op.remap(
|
||||
[prev_va, next_va],
|
||||
pin_init::init_zeroed(),
|
||||
pin_init::init_zeroed(),
|
||||
);
|
||||
|
||||
if let Some(unused_va) = remap_ret.unused_va {
|
||||
context.return_preallocated_gpuva(unused_va);
|
||||
}
|
||||
|
||||
Ok(op_remapped)
|
||||
}
|
||||
}
|
||||
|
||||
/// This function selects the largest supported block size (currently 4KB or 2MB)
|
||||
/// that can be used for a mapping at the given address and size, respecting alignment constraints.
|
||||
///
|
||||
/// We can map multiple pages at once but we can't exceed the size of the
|
||||
/// table entry itself. So, if mapping 4KB pages, figure out how many pages
|
||||
/// can be mapped before we hit the 2MB boundary. Or, if mapping 2MB pages,
|
||||
/// figure out how many pages can be mapped before hitting the 1GB boundary
|
||||
/// Returns the page size (4KB or 2MB) and the number of pages that can be mapped at that size.
|
||||
fn get_pgsize(addr: u64, size: u64) -> (u64, u64) {
|
||||
// Get the distance to the next boundary of 2MB block
|
||||
let blk_offset_2m = addr.wrapping_neg() % (SZ_2M as u64);
|
||||
|
||||
// Use 4K blocks if the address is not 2MB aligned, or we have less than 2MB to map
|
||||
if blk_offset_2m != 0 || size < SZ_2M as u64 {
|
||||
let pgcount = if blk_offset_2m == 0 {
|
||||
size / SZ_4K as u64
|
||||
} else {
|
||||
u64::min(blk_offset_2m, size) / SZ_4K as u64
|
||||
};
|
||||
return (SZ_4K as u64, pgcount);
|
||||
}
|
||||
|
||||
let blk_offset_1g = addr.wrapping_neg() % (SZ_1G as u64);
|
||||
let blk_offset = if blk_offset_1g == 0 {
|
||||
SZ_1G as u64
|
||||
} else {
|
||||
blk_offset_1g
|
||||
};
|
||||
let pgcount = u64::min(blk_offset, size) / SZ_2M as u64;
|
||||
|
||||
(SZ_2M as u64, pgcount)
|
||||
}
|
||||
|
||||
/// Maps a physical address range into the page table at the specified virtual address.
|
||||
///
|
||||
/// This function maps `len` bytes of physical memory starting at `paddr` to the
|
||||
/// virtual address `iova`, using the protection flags specified in `prot`. It
|
||||
/// automatically selects optimal page sizes to minimize page table overhead.
|
||||
///
|
||||
/// If the mapping fails partway through, all successfully mapped pages are
|
||||
/// unmapped before returning an error.
|
||||
///
|
||||
/// Returns the number of bytes successfully mapped.
|
||||
fn pt_map(
|
||||
dev: &Device,
|
||||
pt: &IoPageTable<'_, ARM64LPAES1>,
|
||||
iova: u64,
|
||||
paddr: u64,
|
||||
len: u64,
|
||||
prot: u32,
|
||||
) -> Result<u64> {
|
||||
let mut segment_mapped = 0u64;
|
||||
while segment_mapped < len {
|
||||
let remaining = len - segment_mapped;
|
||||
let curr_iova = iova + segment_mapped;
|
||||
let curr_paddr = paddr + segment_mapped;
|
||||
|
||||
let (pgsize, pgcount) = get_pgsize(curr_iova | curr_paddr, remaining);
|
||||
|
||||
// On 32-bit systems, usize is only 32 bits, so check that
|
||||
// the iova can be converted without truncation.
|
||||
let curr_iova = match usize::try_from(curr_iova) {
|
||||
Ok(curr_iova) => curr_iova,
|
||||
Err(_) => {
|
||||
dev_err!(
|
||||
dev,
|
||||
"curr_iova {:#x} cannot be represented as usize (max {:#x})",
|
||||
curr_iova,
|
||||
usize::MAX
|
||||
);
|
||||
|
||||
if segment_mapped > 0 {
|
||||
let _ = pt_unmap(dev, pt, iova..(iova + segment_mapped));
|
||||
}
|
||||
|
||||
return Err(EOVERFLOW);
|
||||
}
|
||||
};
|
||||
|
||||
// SAFETY:
|
||||
// No other io-pgtable operation can currently access this range because Tyr holds
|
||||
// the gpuvm_unique mutex for the entire sm_map() operation.
|
||||
// The addresses being mapped won't overlap any existing mappings in this
|
||||
// page table because drm_gpuvm_sm_map() checks each requested mapping and either unmaps
|
||||
// or remaps any overlap before creating the new mapping.
|
||||
let (mapped, result) = unsafe {
|
||||
pt.map_pages(
|
||||
curr_iova,
|
||||
curr_paddr as PhysAddr,
|
||||
pgsize as usize,
|
||||
pgcount as usize,
|
||||
prot,
|
||||
GFP_KERNEL,
|
||||
)
|
||||
};
|
||||
|
||||
if let Err(e) = result {
|
||||
// If map_pages fails, mapped will be zero because the ARM LPAE backend
|
||||
// only updates the mapped value after the entire request succeeds.
|
||||
dev_err!(dev, "pt.map_pages failed at iova {:#x}: {:?}", curr_iova, e);
|
||||
if segment_mapped > 0 {
|
||||
let _ = pt_unmap(dev, pt, iova..(iova + segment_mapped));
|
||||
}
|
||||
return Err(e);
|
||||
}
|
||||
|
||||
if mapped == 0 {
|
||||
dev_err!(dev, "Failed to map any pages at iova {:#x}", curr_iova);
|
||||
if segment_mapped > 0 {
|
||||
let _ = pt_unmap(dev, pt, iova..(iova + segment_mapped));
|
||||
}
|
||||
return Err(ENOMEM);
|
||||
}
|
||||
|
||||
segment_mapped += mapped as u64;
|
||||
}
|
||||
|
||||
Ok(segment_mapped)
|
||||
}
|
||||
|
||||
/// Unmaps a virtual address range from the page table.
|
||||
///
|
||||
/// This function removes all page table entries in the specified range,
|
||||
/// automatically handling different page sizes that may be present.
|
||||
fn pt_unmap(dev: &Device, pt: &IoPageTable<'_, ARM64LPAES1>, range: Range<u64>) -> Result {
|
||||
let mut iova = range.start;
|
||||
let mut bytes_left_to_unmap = range.end - range.start;
|
||||
|
||||
while bytes_left_to_unmap > 0 {
|
||||
// It is fine to use just the iova to determine the page size
|
||||
// because if the actual mapping was represented with smaller page sizes,
|
||||
// (e.g. because the physical address was not 2MiB aligned)
|
||||
// the ARM LPAE backend will notice and handle the lower-level table correctly.
|
||||
let (pgsize, pgcount) = get_pgsize(iova, bytes_left_to_unmap);
|
||||
|
||||
// On 32-bit systems, usize is only 32 bits, so check that
|
||||
// the iova can be converted without truncation.
|
||||
let iova_usize = usize::try_from(iova).map_err(|_| {
|
||||
dev_err!(
|
||||
dev,
|
||||
"IOVA {:#x} cannot be represented as usize (max {:#x})",
|
||||
iova,
|
||||
usize::MAX
|
||||
);
|
||||
EOVERFLOW
|
||||
})?;
|
||||
|
||||
// SAFETY:
|
||||
// No other io-pgtable operation can currently access this range because Tyr holds
|
||||
// the gpuvm_unique mutex for the entire sm_unmap() operation.
|
||||
// We know that this page table has one or more consecutive mappings
|
||||
// starting at `iova` with the total size of `pgcount * pgsize` because
|
||||
// gpuvm callbacks provide exactly the range that was previously mapped.
|
||||
let unmapped = unsafe { pt.unmap_pages(iova_usize, pgsize as usize, pgcount as usize) };
|
||||
|
||||
if unmapped == 0 {
|
||||
dev_err!(dev, "Failed to unmap any bytes at iova {:#x}", iova_usize);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
bytes_left_to_unmap -= unmapped as u64;
|
||||
iova += unmapped as u64;
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
1
drivers/gpu/nova-core/.gitignore
vendored
Normal file
1
drivers/gpu/nova-core/.gitignore
vendored
Normal file
|
|
@ -0,0 +1 @@
|
|||
exports_nova_core_generated.h
|
||||
|
|
@ -1,4 +1,3 @@
|
|||
# SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
obj-$(CONFIG_NOVA_CORE) += nova-core.o
|
||||
nova-core-y := nova_core.o
|
||||
# nova-core is built from drivers/gpu/Makefile.
|
||||
# nova_core.o (rust-analyzer marker - DO NOT REMOVE).
|
||||
|
|
|
|||
|
|
@ -1,329 +0,0 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
//! Bitfield library for Rust structures
|
||||
//!
|
||||
//! Support for defining bitfields in Rust structures. Also used by the [`register!`] macro.
|
||||
|
||||
/// Defines a struct with accessors to access bits within an inner unsigned integer.
|
||||
///
|
||||
/// # Syntax
|
||||
///
|
||||
/// ```rust
|
||||
/// use nova_core::bitfield;
|
||||
///
|
||||
/// #[derive(Debug, Clone, Copy, Default)]
|
||||
/// enum Mode {
|
||||
/// #[default]
|
||||
/// Low = 0,
|
||||
/// High = 1,
|
||||
/// Auto = 2,
|
||||
/// }
|
||||
///
|
||||
/// impl TryFrom<u8> for Mode {
|
||||
/// type Error = u8;
|
||||
/// fn try_from(value: u8) -> Result<Self, Self::Error> {
|
||||
/// match value {
|
||||
/// 0 => Ok(Mode::Low),
|
||||
/// 1 => Ok(Mode::High),
|
||||
/// 2 => Ok(Mode::Auto),
|
||||
/// _ => Err(value),
|
||||
/// }
|
||||
/// }
|
||||
/// }
|
||||
///
|
||||
/// impl From<Mode> for u8 {
|
||||
/// fn from(mode: Mode) -> u8 {
|
||||
/// mode as u8
|
||||
/// }
|
||||
/// }
|
||||
///
|
||||
/// #[derive(Debug, Clone, Copy, Default)]
|
||||
/// enum State {
|
||||
/// #[default]
|
||||
/// Inactive = 0,
|
||||
/// Active = 1,
|
||||
/// }
|
||||
///
|
||||
/// impl From<bool> for State {
|
||||
/// fn from(value: bool) -> Self {
|
||||
/// if value { State::Active } else { State::Inactive }
|
||||
/// }
|
||||
/// }
|
||||
///
|
||||
/// impl From<State> for bool {
|
||||
/// fn from(state: State) -> bool {
|
||||
/// match state {
|
||||
/// State::Inactive => false,
|
||||
/// State::Active => true,
|
||||
/// }
|
||||
/// }
|
||||
/// }
|
||||
///
|
||||
/// bitfield! {
|
||||
/// pub struct ControlReg(u32) {
|
||||
/// 7:7 state as bool => State;
|
||||
/// 3:0 mode as u8 ?=> Mode;
|
||||
/// }
|
||||
/// }
|
||||
/// ```
|
||||
///
|
||||
/// This generates a struct with:
|
||||
/// - Field accessors: `mode()`, `state()`, etc.
|
||||
/// - Field setters: `set_mode()`, `set_state()`, etc. (supports chaining with builder pattern).
|
||||
/// Note that the compiler will error out if the size of the setter's arg exceeds the
|
||||
/// struct's storage size.
|
||||
/// - Debug and Default implementations.
|
||||
///
|
||||
/// Note: Field accessors and setters inherit the same visibility as the struct itself.
|
||||
/// In the example above, both `mode()` and `set_mode()` methods will be `pub`.
|
||||
///
|
||||
/// Fields are defined as follows:
|
||||
///
|
||||
/// - `as <type>` simply returns the field value casted to <type>, typically `u32`, `u16`, `u8` or
|
||||
/// `bool`. Note that `bool` fields must have a range of 1 bit.
|
||||
/// - `as <type> => <into_type>` calls `<into_type>`'s `From::<<type>>` implementation and returns
|
||||
/// the result.
|
||||
/// - `as <type> ?=> <try_into_type>` calls `<try_into_type>`'s `TryFrom::<<type>>` implementation
|
||||
/// and returns the result. This is useful with fields for which not all values are valid.
|
||||
macro_rules! bitfield {
|
||||
// Main entry point - defines the bitfield struct with fields
|
||||
($vis:vis struct $name:ident($storage:ty) $(, $comment:literal)? { $($fields:tt)* }) => {
|
||||
bitfield!(@core $vis $name $storage $(, $comment)? { $($fields)* });
|
||||
};
|
||||
|
||||
// All rules below are helpers.
|
||||
|
||||
// Defines the wrapper `$name` type, as well as its relevant implementations (`Debug`,
|
||||
// `Default`, and conversion to the value type) and field accessor methods.
|
||||
(@core $vis:vis $name:ident $storage:ty $(, $comment:literal)? { $($fields:tt)* }) => {
|
||||
$(
|
||||
#[doc=$comment]
|
||||
)?
|
||||
#[repr(transparent)]
|
||||
#[derive(Clone, Copy)]
|
||||
$vis struct $name($storage);
|
||||
|
||||
impl ::core::convert::From<$name> for $storage {
|
||||
fn from(val: $name) -> $storage {
|
||||
val.0
|
||||
}
|
||||
}
|
||||
|
||||
bitfield!(@fields_dispatcher $vis $name $storage { $($fields)* });
|
||||
};
|
||||
|
||||
// Captures the fields and passes them to all the implementers that require field information.
|
||||
//
|
||||
// Used to simplify the matching rules for implementers, so they don't need to match the entire
|
||||
// complex fields rule even though they only make use of part of it.
|
||||
(@fields_dispatcher $vis:vis $name:ident $storage:ty {
|
||||
$($hi:tt:$lo:tt $field:ident as $type:tt
|
||||
$(?=> $try_into_type:ty)?
|
||||
$(=> $into_type:ty)?
|
||||
$(, $comment:literal)?
|
||||
;
|
||||
)*
|
||||
}
|
||||
) => {
|
||||
bitfield!(@field_accessors $vis $name $storage {
|
||||
$(
|
||||
$hi:$lo $field as $type
|
||||
$(?=> $try_into_type)?
|
||||
$(=> $into_type)?
|
||||
$(, $comment)?
|
||||
;
|
||||
)*
|
||||
});
|
||||
bitfield!(@debug $name { $($field;)* });
|
||||
bitfield!(@default $name { $($field;)* });
|
||||
};
|
||||
|
||||
// Defines all the field getter/setter methods for `$name`.
|
||||
(
|
||||
@field_accessors $vis:vis $name:ident $storage:ty {
|
||||
$($hi:tt:$lo:tt $field:ident as $type:tt
|
||||
$(?=> $try_into_type:ty)?
|
||||
$(=> $into_type:ty)?
|
||||
$(, $comment:literal)?
|
||||
;
|
||||
)*
|
||||
}
|
||||
) => {
|
||||
$(
|
||||
bitfield!(@check_field_bounds $hi:$lo $field as $type);
|
||||
)*
|
||||
|
||||
#[allow(dead_code)]
|
||||
impl $name {
|
||||
$(
|
||||
bitfield!(@field_accessor $vis $name $storage, $hi:$lo $field as $type
|
||||
$(?=> $try_into_type)?
|
||||
$(=> $into_type)?
|
||||
$(, $comment)?
|
||||
;
|
||||
);
|
||||
)*
|
||||
}
|
||||
};
|
||||
|
||||
// Boolean fields must have `$hi == $lo`.
|
||||
(@check_field_bounds $hi:tt:$lo:tt $field:ident as bool) => {
|
||||
#[allow(clippy::eq_op)]
|
||||
const _: () = {
|
||||
::kernel::build_assert::build_assert!(
|
||||
$hi == $lo,
|
||||
concat!("boolean field `", stringify!($field), "` covers more than one bit")
|
||||
);
|
||||
};
|
||||
};
|
||||
|
||||
// Non-boolean fields must have `$hi >= $lo`.
|
||||
(@check_field_bounds $hi:tt:$lo:tt $field:ident as $type:tt) => {
|
||||
#[allow(clippy::eq_op)]
|
||||
const _: () = {
|
||||
::kernel::build_assert::build_assert!(
|
||||
$hi >= $lo,
|
||||
concat!("field `", stringify!($field), "`'s MSB is smaller than its LSB")
|
||||
);
|
||||
};
|
||||
};
|
||||
|
||||
// Catches fields defined as `bool` and convert them into a boolean value.
|
||||
(
|
||||
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as bool
|
||||
=> $into_type:ty $(, $comment:literal)?;
|
||||
) => {
|
||||
bitfield!(
|
||||
@leaf_accessor $vis $name $storage, $hi:$lo $field
|
||||
{ |f| <$into_type>::from(f != 0) }
|
||||
bool $into_type => $into_type $(, $comment)?;
|
||||
);
|
||||
};
|
||||
|
||||
// Shortcut for fields defined as `bool` without the `=>` syntax.
|
||||
(
|
||||
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as bool
|
||||
$(, $comment:literal)?;
|
||||
) => {
|
||||
bitfield!(
|
||||
@field_accessor $vis $name $storage, $hi:$lo $field as bool => bool $(, $comment)?;
|
||||
);
|
||||
};
|
||||
|
||||
// Catches the `?=>` syntax for non-boolean fields.
|
||||
(
|
||||
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as $type:tt
|
||||
?=> $try_into_type:ty $(, $comment:literal)?;
|
||||
) => {
|
||||
bitfield!(@leaf_accessor $vis $name $storage, $hi:$lo $field
|
||||
{ |f| <$try_into_type>::try_from(f as $type) } $type $try_into_type =>
|
||||
::core::result::Result<
|
||||
$try_into_type,
|
||||
<$try_into_type as ::core::convert::TryFrom<$type>>::Error
|
||||
>
|
||||
$(, $comment)?;);
|
||||
};
|
||||
|
||||
// Catches the `=>` syntax for non-boolean fields.
|
||||
(
|
||||
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as $type:tt
|
||||
=> $into_type:ty $(, $comment:literal)?;
|
||||
) => {
|
||||
bitfield!(@leaf_accessor $vis $name $storage, $hi:$lo $field
|
||||
{ |f| <$into_type>::from(f as $type) } $type $into_type => $into_type $(, $comment)?;);
|
||||
};
|
||||
|
||||
// Shortcut for non-boolean fields defined without the `=>` or `?=>` syntax.
|
||||
(
|
||||
@field_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident as $type:tt
|
||||
$(, $comment:literal)?;
|
||||
) => {
|
||||
bitfield!(
|
||||
@field_accessor $vis $name $storage, $hi:$lo $field as $type => $type $(, $comment)?;
|
||||
);
|
||||
};
|
||||
|
||||
// Generates the accessor methods for a single field.
|
||||
(
|
||||
@leaf_accessor $vis:vis $name:ident $storage:ty, $hi:tt:$lo:tt $field:ident
|
||||
{ $process:expr } $prim_type:tt $to_type:ty => $res_type:ty $(, $comment:literal)?;
|
||||
) => {
|
||||
::kernel::macros::paste!(
|
||||
const [<$field:upper _RANGE>]: ::core::ops::RangeInclusive<u8> = $lo..=$hi;
|
||||
const [<$field:upper _MASK>]: $storage = {
|
||||
// Generate mask for shifting
|
||||
match ::core::mem::size_of::<$storage>() {
|
||||
1 => ::kernel::bits::genmask_u8($lo..=$hi) as $storage,
|
||||
2 => ::kernel::bits::genmask_u16($lo..=$hi) as $storage,
|
||||
4 => ::kernel::bits::genmask_u32($lo..=$hi) as $storage,
|
||||
8 => ::kernel::bits::genmask_u64($lo..=$hi) as $storage,
|
||||
_ => ::kernel::build_error!("Unsupported storage type size")
|
||||
}
|
||||
};
|
||||
const [<$field:upper _SHIFT>]: u32 = $lo;
|
||||
);
|
||||
|
||||
$(
|
||||
#[doc="Returns the value of this field:"]
|
||||
#[doc=$comment]
|
||||
)?
|
||||
#[inline(always)]
|
||||
$vis fn $field(self) -> $res_type {
|
||||
::kernel::macros::paste!(
|
||||
const MASK: $storage = $name::[<$field:upper _MASK>];
|
||||
const SHIFT: u32 = $name::[<$field:upper _SHIFT>];
|
||||
);
|
||||
let field = ((self.0 & MASK) >> SHIFT);
|
||||
|
||||
$process(field)
|
||||
}
|
||||
|
||||
::kernel::macros::paste!(
|
||||
$(
|
||||
#[doc="Sets the value of this field:"]
|
||||
#[doc=$comment]
|
||||
)?
|
||||
#[inline(always)]
|
||||
$vis fn [<set_ $field>](mut self, value: $to_type) -> Self {
|
||||
const MASK: $storage = $name::[<$field:upper _MASK>];
|
||||
const SHIFT: u32 = $name::[<$field:upper _SHIFT>];
|
||||
let value = ($storage::from($prim_type::from(value)) << SHIFT) & MASK;
|
||||
self.0 = (self.0 & !MASK) | value;
|
||||
|
||||
self
|
||||
}
|
||||
);
|
||||
};
|
||||
|
||||
// Generates the `Debug` implementation for `$name`.
|
||||
(@debug $name:ident { $($field:ident;)* }) => {
|
||||
impl ::kernel::fmt::Debug for $name {
|
||||
fn fmt(&self, f: &mut ::kernel::fmt::Formatter<'_>) -> ::kernel::fmt::Result {
|
||||
f.debug_struct(stringify!($name))
|
||||
.field("<raw>", &::kernel::prelude::fmt!("{:#x}", &self.0))
|
||||
$(
|
||||
.field(stringify!($field), &self.$field())
|
||||
)*
|
||||
.finish()
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
// Generates the `Default` implementation for `$name`.
|
||||
(@default $name:ident { $($field:ident;)* }) => {
|
||||
/// Returns a value for the bitfield where all fields are set to their default value.
|
||||
impl ::core::default::Default for $name {
|
||||
fn default() -> Self {
|
||||
let value = Self(Default::default());
|
||||
|
||||
::kernel::macros::paste!(
|
||||
$(
|
||||
let value = value.[<set_ $field>](Default::default());
|
||||
)*
|
||||
);
|
||||
|
||||
value
|
||||
}
|
||||
}
|
||||
};
|
||||
}
|
||||
|
|
@ -5,17 +5,14 @@
|
|||
use hal::FalconHal;
|
||||
|
||||
use kernel::{
|
||||
device::{
|
||||
self,
|
||||
Device, //
|
||||
},
|
||||
device,
|
||||
dma::{
|
||||
Coherent,
|
||||
CoherentBox,
|
||||
DmaAddress,
|
||||
DmaMask, //
|
||||
DmaAddress, //
|
||||
},
|
||||
io::{
|
||||
io_project,
|
||||
poll::read_poll_timeout,
|
||||
register::{
|
||||
RegisterBase,
|
||||
|
|
@ -24,7 +21,6 @@
|
|||
Io,
|
||||
},
|
||||
prelude::*,
|
||||
sync::aref::ARef,
|
||||
time::Delta,
|
||||
};
|
||||
|
||||
|
|
@ -358,41 +354,47 @@ pub(crate) trait FalconFirmware {
|
|||
}
|
||||
|
||||
/// Contains the base parameters common to all Falcon instances.
|
||||
pub(crate) struct Falcon<E: FalconEngine> {
|
||||
pub(crate) struct Falcon<'a, E: FalconEngine> {
|
||||
hal: KBox<dyn FalconHal<E>>,
|
||||
dev: ARef<device::Device>,
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
bar: Bar0<'a>,
|
||||
}
|
||||
|
||||
impl<E: FalconEngine + 'static> Falcon<E> {
|
||||
impl<'a, E: FalconEngine + 'static> Falcon<'a, E> {
|
||||
/// Create a new falcon instance.
|
||||
pub(crate) fn new(dev: &device::Device, chipset: Chipset) -> Result<Self> {
|
||||
pub(crate) fn new(
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
bar: Bar0<'a>,
|
||||
) -> Result<Self> {
|
||||
Ok(Self {
|
||||
hal: hal::falcon_hal(chipset)?,
|
||||
dev: dev.into(),
|
||||
dev,
|
||||
bar,
|
||||
})
|
||||
}
|
||||
|
||||
/// Resets DMA-related registers.
|
||||
pub(crate) fn dma_reset(&self, bar: Bar0<'_>) {
|
||||
bar.update(regs::NV_PFALCON_FBIF_CTL::of::<E>(), |v| {
|
||||
pub(crate) fn dma_reset(&self) {
|
||||
self.bar.update(regs::NV_PFALCON_FBIF_CTL::of::<E>(), |v| {
|
||||
v.with_allow_phys_no_ctx(true)
|
||||
});
|
||||
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_DMACTL::zeroed(),
|
||||
);
|
||||
}
|
||||
|
||||
/// Reset the controller, select the falcon core, and wait for memory scrubbing to complete.
|
||||
pub(crate) fn reset(&self, bar: Bar0<'_>) -> Result {
|
||||
self.hal.reset_eng(bar)?;
|
||||
self.hal.select_core(self, bar)?;
|
||||
self.hal.reset_wait_mem_scrubbing(bar)?;
|
||||
pub(crate) fn reset(&self) -> Result {
|
||||
self.hal.reset_eng(self)?;
|
||||
self.hal.select_core(self)?;
|
||||
self.hal.reset_wait_mem_scrubbing(self)?;
|
||||
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_RM::from(bar.read(regs::NV_PMC_BOOT_0).into_raw()),
|
||||
regs::NV_PFALCON_FALCON_RM::from(self.bar.read(regs::NV_PMC_BOOT_0).into_raw()),
|
||||
);
|
||||
|
||||
Ok(())
|
||||
|
|
@ -404,18 +406,14 @@ pub(crate) fn reset(&self, bar: Bar0<'_>) -> Result {
|
|||
/// Write a slice to Falcon IMEM memory using programmed I/O (PIO).
|
||||
///
|
||||
/// Returns `EINVAL` if `img.len()` is not a multiple of 4.
|
||||
fn pio_wr_imem_slice(
|
||||
&self,
|
||||
bar: Bar0<'_>,
|
||||
load_offsets: FalconPioImemLoadTarget<'_>,
|
||||
) -> Result {
|
||||
fn pio_wr_imem_slice(&self, load_offsets: FalconPioImemLoadTarget<'_>) -> Result {
|
||||
// Rejecting misaligned images here allows us to avoid checking
|
||||
// inside the loops.
|
||||
if load_offsets.data.len() % 4 != 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>().at(Self::PIO_PORT),
|
||||
regs::NV_PFALCON_FALCON_IMEMC::zeroed()
|
||||
.with_secure(load_offsets.secure)
|
||||
|
|
@ -426,13 +424,13 @@ fn pio_wr_imem_slice(
|
|||
for (n, block) in load_offsets.data.chunks(MEM_BLOCK_ALIGNMENT).enumerate() {
|
||||
let n = u16::try_from(n)?;
|
||||
let tag: u16 = load_offsets.start_tag.checked_add(n).ok_or(ERANGE)?;
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>().at(Self::PIO_PORT),
|
||||
regs::NV_PFALCON_FALCON_IMEMT::zeroed().with_tag(tag),
|
||||
);
|
||||
for word in block.chunks_exact(4) {
|
||||
let w = [word[0], word[1], word[2], word[3]];
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>().at(Self::PIO_PORT),
|
||||
regs::NV_PFALCON_FALCON_IMEMD::zeroed().with_data(u32::from_le_bytes(w)),
|
||||
);
|
||||
|
|
@ -445,18 +443,14 @@ fn pio_wr_imem_slice(
|
|||
/// Write a slice to Falcon DMEM memory using programmed I/O (PIO).
|
||||
///
|
||||
/// Returns `EINVAL` if `img.len()` is not a multiple of 4.
|
||||
fn pio_wr_dmem_slice(
|
||||
&self,
|
||||
bar: Bar0<'_>,
|
||||
load_offsets: FalconPioDmemLoadTarget<'_>,
|
||||
) -> Result {
|
||||
fn pio_wr_dmem_slice(&self, load_offsets: FalconPioDmemLoadTarget<'_>) -> Result {
|
||||
// Rejecting misaligned images here allows us to avoid checking
|
||||
// inside the loops.
|
||||
if load_offsets.data.len() % 4 != 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>().at(Self::PIO_PORT),
|
||||
regs::NV_PFALCON_FALCON_DMEMC::zeroed()
|
||||
.with_aincw(true)
|
||||
|
|
@ -465,7 +459,7 @@ fn pio_wr_dmem_slice(
|
|||
|
||||
for word in load_offsets.data.chunks_exact(4) {
|
||||
let w = [word[0], word[1], word[2], word[3]];
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>().at(Self::PIO_PORT),
|
||||
regs::NV_PFALCON_FALCON_DMEMD::zeroed().with_data(u32::from_le_bytes(w)),
|
||||
);
|
||||
|
|
@ -477,29 +471,28 @@ fn pio_wr_dmem_slice(
|
|||
/// Perform a PIO copy into `IMEM` and `DMEM` of `fw`, and prepare the falcon to run it.
|
||||
pub(crate) fn pio_load<F: FalconFirmware<Target = E> + FalconPioLoadable>(
|
||||
&self,
|
||||
bar: Bar0<'_>,
|
||||
fw: &F,
|
||||
) -> Result {
|
||||
bar.update(regs::NV_PFALCON_FBIF_CTL::of::<E>(), |v| {
|
||||
self.bar.update(regs::NV_PFALCON_FBIF_CTL::of::<E>(), |v| {
|
||||
v.with_allow_phys_no_ctx(true)
|
||||
});
|
||||
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_DMACTL::zeroed(),
|
||||
);
|
||||
|
||||
if let Some(imem_ns) = fw.imem_ns_load_params() {
|
||||
self.pio_wr_imem_slice(bar, imem_ns)?;
|
||||
self.pio_wr_imem_slice(imem_ns)?;
|
||||
}
|
||||
if let Some(imem_sec) = fw.imem_sec_load_params() {
|
||||
self.pio_wr_imem_slice(bar, imem_sec)?;
|
||||
self.pio_wr_imem_slice(imem_sec)?;
|
||||
}
|
||||
self.pio_wr_dmem_slice(bar, fw.dmem_load_params())?;
|
||||
self.pio_wr_dmem_slice(fw.dmem_load_params())?;
|
||||
|
||||
self.hal.program_brom(self, bar, &fw.brom_params());
|
||||
self.hal.program_brom(self, &fw.brom_params());
|
||||
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_BOOTVEC::zeroed().with_value(fw.boot_addr()),
|
||||
);
|
||||
|
|
@ -507,33 +500,43 @@ pub(crate) fn pio_load<F: FalconFirmware<Target = E> + FalconPioLoadable>(
|
|||
Ok(())
|
||||
}
|
||||
|
||||
/// Perform a DMA write according to `load_offsets` from `dma_handle` into the falcon's
|
||||
/// Perform a DMA write according to `load_offsets` from `dma_obj` into the falcon's
|
||||
/// `target_mem`.
|
||||
///
|
||||
/// `sec` is set if the loaded firmware is expected to run in secure mode.
|
||||
fn dma_wr(
|
||||
&self,
|
||||
bar: Bar0<'_>,
|
||||
dma_obj: &Coherent<[u8]>,
|
||||
target_mem: FalconMem,
|
||||
load_offsets: FalconDmaLoadTarget,
|
||||
) -> Result {
|
||||
const DMA_LEN: u32 = num::usize_into_u32::<{ MEM_BLOCK_ALIGNMENT }>();
|
||||
|
||||
// DMA transfers can only be done in units of 256 bytes. Compute how many such transfers we
|
||||
// need to perform.
|
||||
let num_transfers = load_offsets.len.div_ceil(DMA_LEN);
|
||||
|
||||
// For IMEM, we want to use the start offset as a virtual address tag for each page, since
|
||||
// code addresses in the firmware (and the boot vector) are virtual.
|
||||
//
|
||||
// For DMEM we can fold the start offset into the DMA handle.
|
||||
// For DMEM, the start offset is folded into the DMA address.
|
||||
let (src_start, dma_start) = match target_mem {
|
||||
FalconMem::ImemSecure | FalconMem::ImemNonSecure => {
|
||||
(load_offsets.src_start, dma_obj.dma_handle())
|
||||
}
|
||||
FalconMem::Dmem => (
|
||||
0,
|
||||
dma_obj.dma_handle() + DmaAddress::from(load_offsets.src_start),
|
||||
),
|
||||
FalconMem::ImemSecure | FalconMem::ImemNonSecure => (load_offsets.src_start, 0),
|
||||
FalconMem::Dmem => (0, usize::from_safe_cast(load_offsets.src_start)),
|
||||
};
|
||||
if dma_start % DmaAddress::from(DMA_LEN) > 0 {
|
||||
|
||||
let dma_address = {
|
||||
// Upper limit of transfer is `(num_transfers * DMA_LEN) + load_offsets.src_start`.
|
||||
let dma_end = num_transfers
|
||||
.checked_mul(DMA_LEN)
|
||||
.and_then(|size| size.checked_add(load_offsets.src_start))
|
||||
.map(usize::from_safe_cast)
|
||||
.ok_or(EOVERFLOW)?;
|
||||
|
||||
io_project!(dma_obj, [try: dma_start..dma_end]).dma_address()
|
||||
};
|
||||
|
||||
if dma_address % DmaAddress::from(DMA_LEN) > 0 {
|
||||
dev_err!(
|
||||
self.dev,
|
||||
"DMA transfer start addresses must be a multiple of {}\n",
|
||||
|
|
@ -542,46 +545,19 @@ fn dma_wr(
|
|||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// The DMATRFBASE/1 register pair only supports a 49-bit address.
|
||||
if dma_start > DmaMask::new::<49>().value() {
|
||||
dev_err!(self.dev, "DMA address {:#x} exceeds 49 bits\n", dma_start);
|
||||
return Err(ERANGE);
|
||||
}
|
||||
|
||||
// DMA transfers can only be done in units of 256 bytes. Compute how many such transfers we
|
||||
// need to perform.
|
||||
let num_transfers = load_offsets.len.div_ceil(DMA_LEN);
|
||||
|
||||
// Check that the area we are about to transfer is within the bounds of the DMA object.
|
||||
// Upper limit of transfer is `(num_transfers * DMA_LEN) + load_offsets.src_start`.
|
||||
match num_transfers
|
||||
.checked_mul(DMA_LEN)
|
||||
.and_then(|size| size.checked_add(load_offsets.src_start))
|
||||
{
|
||||
None => {
|
||||
dev_err!(self.dev, "DMA transfer length overflow\n");
|
||||
return Err(EOVERFLOW);
|
||||
}
|
||||
Some(upper_bound) if usize::from_safe_cast(upper_bound) > dma_obj.size() => {
|
||||
dev_err!(self.dev, "DMA transfer goes beyond range of DMA object\n");
|
||||
return Err(EINVAL);
|
||||
}
|
||||
Some(_) => (),
|
||||
};
|
||||
|
||||
// Set up the base source DMA address.
|
||||
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_DMATRFBASE::zeroed().with_base(
|
||||
// CAST: `as u32` is used on purpose since we do want to strip the upper bits,
|
||||
// which will be written to `NV_PFALCON_FALCON_DMATRFBASE1`.
|
||||
(dma_start >> 8) as u32,
|
||||
(dma_address >> 8) as u32,
|
||||
),
|
||||
);
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_DMATRFBASE1::zeroed().try_with_base(dma_start >> 40)?,
|
||||
regs::NV_PFALCON_FALCON_DMATRFBASE1::zeroed().try_with_base(dma_address >> 40)?,
|
||||
);
|
||||
|
||||
let cmd = regs::NV_PFALCON_FALCON_DMATRFCMD::zeroed()
|
||||
|
|
@ -590,23 +566,23 @@ fn dma_wr(
|
|||
|
||||
for pos in (0..num_transfers).map(|i| i * DMA_LEN) {
|
||||
// Perform a transfer of size `DMA_LEN`.
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_DMATRFMOFFS::zeroed()
|
||||
.try_with_offs(load_offsets.dst_start + pos)?,
|
||||
);
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_DMATRFFBOFFS::zeroed().with_offs(src_start + pos),
|
||||
);
|
||||
|
||||
bar.write(WithBase::of::<E>(), cmd);
|
||||
self.bar.write(WithBase::of::<E>(), cmd);
|
||||
|
||||
// Wait for the transfer to complete.
|
||||
// TIMEOUT: arbitrarily large value, no DMA transfer to the falcon's small memories
|
||||
// should ever take that long.
|
||||
read_poll_timeout(
|
||||
|| Ok(bar.read(regs::NV_PFALCON_FALCON_DMATRFCMD::of::<E>())),
|
||||
|| Ok(self.bar.read(regs::NV_PFALCON_FALCON_DMATRFCMD::of::<E>())),
|
||||
|r| r.idle(),
|
||||
Delta::ZERO,
|
||||
Delta::from_secs(2),
|
||||
|
|
@ -617,12 +593,7 @@ fn dma_wr(
|
|||
}
|
||||
|
||||
/// Perform a DMA load into `IMEM` and `DMEM` of `fw`, and prepare the falcon to run it.
|
||||
fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
|
||||
&self,
|
||||
dev: &Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
fw: &F,
|
||||
) -> Result {
|
||||
fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(&self, fw: &F) -> Result {
|
||||
// DMA object with firmware content as the source of the DMA engine.
|
||||
let dma_obj = {
|
||||
let fw_slice = fw.as_slice();
|
||||
|
|
@ -630,7 +601,7 @@ fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
|
|||
// DMA copies are done in chunks of `MEM_BLOCK_ALIGNMENT`, so pad the length
|
||||
// accordingly and fill with `0`.
|
||||
let mut dma_obj = CoherentBox::zeroed_slice(
|
||||
dev,
|
||||
self.dev,
|
||||
fw_slice.len().next_multiple_of(MEM_BLOCK_ALIGNMENT),
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
|
|
@ -642,24 +613,20 @@ fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
|
|||
dma_obj.into()
|
||||
};
|
||||
|
||||
self.dma_reset(bar);
|
||||
bar.update(regs::NV_PFALCON_FBIF_TRANSCFG::of::<E>().at(0), |v| {
|
||||
v.with_target(FalconFbifTarget::CoherentSysmem)
|
||||
.with_mem_type(FalconFbifMemType::Physical)
|
||||
});
|
||||
self.dma_reset();
|
||||
self.bar
|
||||
.update(regs::NV_PFALCON_FBIF_TRANSCFG::of::<E>().at(0), |v| {
|
||||
v.with_target(FalconFbifTarget::CoherentSysmem)
|
||||
.with_mem_type(FalconFbifMemType::Physical)
|
||||
});
|
||||
|
||||
self.dma_wr(
|
||||
bar,
|
||||
&dma_obj,
|
||||
FalconMem::ImemSecure,
|
||||
fw.imem_sec_load_params(),
|
||||
)?;
|
||||
self.dma_wr(bar, &dma_obj, FalconMem::Dmem, fw.dmem_load_params())?;
|
||||
self.dma_wr(&dma_obj, FalconMem::ImemSecure, fw.imem_sec_load_params())?;
|
||||
self.dma_wr(&dma_obj, FalconMem::Dmem, fw.dmem_load_params())?;
|
||||
|
||||
self.hal.program_brom(self, bar, &fw.brom_params());
|
||||
self.hal.program_brom(self, &fw.brom_params());
|
||||
|
||||
// Set `BootVec` to start of non-secure code.
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_BOOTVEC::zeroed().with_value(fw.boot_addr()),
|
||||
);
|
||||
|
|
@ -668,10 +635,10 @@ fn dma_load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
|
|||
}
|
||||
|
||||
/// Wait until the falcon CPU is halted.
|
||||
pub(crate) fn wait_till_halted(&self, bar: Bar0<'_>) -> Result<()> {
|
||||
pub(crate) fn wait_till_halted(&self) -> Result<()> {
|
||||
// TIMEOUT: arbitrarily large value, firmwares should complete in less than 2 seconds.
|
||||
read_poll_timeout(
|
||||
|| Ok(bar.read(regs::NV_PFALCON_FALCON_CPUCTL::of::<E>())),
|
||||
|| Ok(self.bar.read(regs::NV_PFALCON_FALCON_CPUCTL::of::<E>())),
|
||||
|r| r.halted(),
|
||||
Delta::ZERO,
|
||||
Delta::from_secs(2),
|
||||
|
|
@ -681,16 +648,17 @@ pub(crate) fn wait_till_halted(&self, bar: Bar0<'_>) -> Result<()> {
|
|||
}
|
||||
|
||||
/// Start the falcon CPU.
|
||||
pub(crate) fn start(&self, bar: Bar0<'_>) -> Result<()> {
|
||||
match bar
|
||||
pub(crate) fn start(&self) -> Result<()> {
|
||||
match self
|
||||
.bar
|
||||
.read(regs::NV_PFALCON_FALCON_CPUCTL::of::<E>())
|
||||
.alias_en()
|
||||
{
|
||||
true => bar.write(
|
||||
true => self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_CPUCTL_ALIAS::zeroed().with_startcpu(true),
|
||||
),
|
||||
false => bar.write(
|
||||
false => self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_CPUCTL::zeroed().with_startcpu(true),
|
||||
),
|
||||
|
|
@ -700,16 +668,16 @@ pub(crate) fn start(&self, bar: Bar0<'_>) -> Result<()> {
|
|||
}
|
||||
|
||||
/// Writes values to the mailbox registers if provided.
|
||||
pub(crate) fn write_mailboxes(&self, bar: Bar0<'_>, mbox0: Option<u32>, mbox1: Option<u32>) {
|
||||
pub(crate) fn write_mailboxes(&self, mbox0: Option<u32>, mbox1: Option<u32>) {
|
||||
if let Some(mbox0) = mbox0 {
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_MAILBOX0::zeroed().with_value(mbox0),
|
||||
);
|
||||
}
|
||||
|
||||
if let Some(mbox1) = mbox1 {
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_MAILBOX1::zeroed().with_value(mbox1),
|
||||
);
|
||||
|
|
@ -717,21 +685,23 @@ pub(crate) fn write_mailboxes(&self, bar: Bar0<'_>, mbox0: Option<u32>, mbox1: O
|
|||
}
|
||||
|
||||
/// Reads the value from `mbox0` register.
|
||||
pub(crate) fn read_mailbox0(&self, bar: Bar0<'_>) -> u32 {
|
||||
bar.read(regs::NV_PFALCON_FALCON_MAILBOX0::of::<E>())
|
||||
pub(crate) fn read_mailbox0(&self) -> u32 {
|
||||
self.bar
|
||||
.read(regs::NV_PFALCON_FALCON_MAILBOX0::of::<E>())
|
||||
.value()
|
||||
}
|
||||
|
||||
/// Reads the value from `mbox1` register.
|
||||
pub(crate) fn read_mailbox1(&self, bar: Bar0<'_>) -> u32 {
|
||||
bar.read(regs::NV_PFALCON_FALCON_MAILBOX1::of::<E>())
|
||||
pub(crate) fn read_mailbox1(&self) -> u32 {
|
||||
self.bar
|
||||
.read(regs::NV_PFALCON_FALCON_MAILBOX1::of::<E>())
|
||||
.value()
|
||||
}
|
||||
|
||||
/// Reads values from both mailbox registers.
|
||||
pub(crate) fn read_mailboxes(&self, bar: Bar0<'_>) -> (u32, u32) {
|
||||
let mbox0 = self.read_mailbox0(bar);
|
||||
let mbox1 = self.read_mailbox1(bar);
|
||||
pub(crate) fn read_mailboxes(&self) -> (u32, u32) {
|
||||
let mbox0 = self.read_mailbox0();
|
||||
let mbox1 = self.read_mailbox1();
|
||||
|
||||
(mbox0, mbox1)
|
||||
}
|
||||
|
|
@ -743,54 +713,54 @@ pub(crate) fn read_mailboxes(&self, bar: Bar0<'_>) -> (u32, u32) {
|
|||
///
|
||||
/// Wait up to two seconds for the firmware to complete, and return its exit status read from
|
||||
/// the `MBOX0` and `MBOX1` registers.
|
||||
pub(crate) fn boot(
|
||||
&self,
|
||||
bar: Bar0<'_>,
|
||||
mbox0: Option<u32>,
|
||||
mbox1: Option<u32>,
|
||||
) -> Result<(u32, u32)> {
|
||||
self.write_mailboxes(bar, mbox0, mbox1);
|
||||
self.start(bar)?;
|
||||
self.wait_till_halted(bar)?;
|
||||
Ok(self.read_mailboxes(bar))
|
||||
pub(crate) fn boot(&self, mbox0: Option<u32>, mbox1: Option<u32>) -> Result<(u32, u32)> {
|
||||
self.write_mailboxes(mbox0, mbox1);
|
||||
self.start()?;
|
||||
self.wait_till_halted()?;
|
||||
Ok(self.read_mailboxes())
|
||||
}
|
||||
|
||||
/// Returns the fused version of the signature to use in order to run a HS firmware on this
|
||||
/// falcon instance. `engine_id_mask` and `ucode_id` are obtained from the firmware header.
|
||||
pub(crate) fn signature_reg_fuse_version(
|
||||
&self,
|
||||
bar: Bar0<'_>,
|
||||
engine_id_mask: u16,
|
||||
ucode_id: u8,
|
||||
) -> Result<u32> {
|
||||
self.hal
|
||||
.signature_reg_fuse_version(self, bar, engine_id_mask, ucode_id)
|
||||
.signature_reg_fuse_version(self, engine_id_mask, ucode_id)
|
||||
}
|
||||
|
||||
/// Check if the RISC-V core is active.
|
||||
///
|
||||
/// Note that this does not guarantee that the RISC-V core is halted if it returns `false`.
|
||||
///
|
||||
/// Returns `true` if the RISC-V core is active, `false` otherwise.
|
||||
pub(crate) fn is_riscv_active(&self, bar: Bar0<'_>) -> bool {
|
||||
self.hal.is_riscv_active(bar)
|
||||
pub(crate) fn is_riscv_active(&self) -> bool {
|
||||
self.hal.is_riscv_active(self)
|
||||
}
|
||||
|
||||
/// Checks whether the RISC-V core is halted.
|
||||
///
|
||||
/// Note that this does not guarantee that the RISC-V core is active if it returns `false`.
|
||||
///
|
||||
/// Returns [`ENOTSUPP`] if the status is not available.
|
||||
pub(crate) fn is_riscv_halted(&self) -> Result<bool> {
|
||||
self.hal.is_riscv_halted(self)
|
||||
}
|
||||
|
||||
/// Load a firmware image into Falcon memory, using the preferred method for the current
|
||||
/// chipset.
|
||||
pub(crate) fn load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(
|
||||
&self,
|
||||
dev: &Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
fw: &F,
|
||||
) -> Result {
|
||||
pub(crate) fn load<F: FalconFirmware<Target = E> + FalconDmaLoadable>(&self, fw: &F) -> Result {
|
||||
match self.hal.load_method() {
|
||||
LoadMethod::Dma => self.dma_load(dev, bar, fw),
|
||||
LoadMethod::Pio => self.pio_load(bar, &fw.try_as_pio_loadable()?),
|
||||
LoadMethod::Dma => self.dma_load(fw),
|
||||
LoadMethod::Pio => self.pio_load(&fw.try_as_pio_loadable()?),
|
||||
}
|
||||
}
|
||||
|
||||
/// Write the application version to the OS register.
|
||||
pub(crate) fn write_os_version(&self, bar: Bar0<'_>, app_version: u32) {
|
||||
bar.write(
|
||||
pub(crate) fn write_os_version(&self, app_version: u32) {
|
||||
self.bar.write(
|
||||
WithBase::of::<E>(),
|
||||
regs::NV_PFALCON_FALCON_OS::zeroed().with_value(app_version),
|
||||
);
|
||||
|
|
|
|||
|
|
@ -17,11 +17,11 @@
|
|||
Io, //
|
||||
},
|
||||
prelude::*,
|
||||
sizes::SZ_1K,
|
||||
time::Delta,
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
Falcon,
|
||||
FalconEngine,
|
||||
|
|
@ -35,6 +35,9 @@
|
|||
/// FSP message timeout in milliseconds.
|
||||
const FSP_MSG_TIMEOUT_MS: i64 = 2000;
|
||||
|
||||
/// Size of the FSP EMEM channel 0 that we can use.
|
||||
const FSP_EMEM_CHANNEL_0_SIZE: usize = SZ_1K;
|
||||
|
||||
/// Type specifying the `Fsp` falcon engine. Cannot be instantiated.
|
||||
pub(crate) struct Fsp(());
|
||||
|
||||
|
|
@ -48,18 +51,18 @@ impl RegisterBase<PFalcon2Base> for Fsp {
|
|||
|
||||
impl FalconEngine for Fsp {}
|
||||
|
||||
impl Falcon<Fsp> {
|
||||
impl<'a> Falcon<'a, Fsp> {
|
||||
/// Writes `data` to FSP external memory at offset `0`.
|
||||
///
|
||||
/// `data` is interpreted as little-endian 32-bit words. Returns `EINVAL`
|
||||
/// if the `data` length is not 4-byte aligned.
|
||||
fn write_emem(&mut self, bar: Bar0<'_>, data: &[u8]) -> Result {
|
||||
fn write_emem(&mut self, data: &[u8]) -> Result {
|
||||
if data.len() % 4 != 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// Begin a write burst at offset `0`, auto-incrementing on each write.
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<Fsp>(),
|
||||
regs::NV_PFALCON_FALCON_EMEMC::zeroed().with_aincw(true),
|
||||
);
|
||||
|
|
@ -68,7 +71,7 @@ fn write_emem(&mut self, bar: Bar0<'_>, data: &[u8]) -> Result {
|
|||
let value = u32::from_le_bytes([chunk[0], chunk[1], chunk[2], chunk[3]]);
|
||||
|
||||
// Write the next 32-bit `value`; hardware advances the offset.
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<Fsp>(),
|
||||
regs::NV_PFALCON_FALCON_EMEMD::zeroed().with_data(value),
|
||||
);
|
||||
|
|
@ -81,20 +84,23 @@ fn write_emem(&mut self, bar: Bar0<'_>, data: &[u8]) -> Result {
|
|||
///
|
||||
/// `data` is stored as little-endian 32-bit words. Returns `EINVAL` if
|
||||
/// the `data` length is not 4-byte aligned.
|
||||
fn read_emem(&mut self, bar: Bar0<'_>, data: &mut [u8]) -> Result {
|
||||
fn read_emem(&mut self, data: &mut [u8]) -> Result {
|
||||
if data.len() % 4 != 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// Begin a read burst at offset `0`, auto-incrementing on each read.
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
WithBase::of::<Fsp>(),
|
||||
regs::NV_PFALCON_FALCON_EMEMC::zeroed().with_aincr(true),
|
||||
);
|
||||
|
||||
for chunk in data.chunks_exact_mut(4) {
|
||||
// Read the next 32-bit word; hardware advances the offset.
|
||||
let value = bar.read(regs::NV_PFALCON_FALCON_EMEMD::of::<Fsp>()).data();
|
||||
let value = self
|
||||
.bar
|
||||
.read(regs::NV_PFALCON_FALCON_EMEMD::of::<Fsp>())
|
||||
.data();
|
||||
chunk.copy_from_slice(&value.to_le_bytes());
|
||||
}
|
||||
|
||||
|
|
@ -105,37 +111,41 @@ fn read_emem(&mut self, bar: Bar0<'_>, data: &mut [u8]) -> Result {
|
|||
///
|
||||
/// Returns the size of available data in bytes, or 0 if no data is available.
|
||||
///
|
||||
/// Returns [`EIO`] if the queue pointers are bogus (`tail < head`).
|
||||
///
|
||||
/// The FSP message queue is not circular. Pointers are reset to 0 after each
|
||||
/// message exchange, so `tail >= head` is always true when data is present.
|
||||
fn poll_msgq(&self, bar: Bar0<'_>) -> u32 {
|
||||
let head = bar.read(regs::NV_PFSP_MSGQ_HEAD::at(0)).val();
|
||||
let tail = bar.read(regs::NV_PFSP_MSGQ_TAIL::at(0)).val();
|
||||
fn poll_msgq(&self) -> Result<u32> {
|
||||
let head = self.bar.read(regs::NV_PFSP_MSGQ_HEAD::at(0)).val();
|
||||
let tail = self.bar.read(regs::NV_PFSP_MSGQ_TAIL::at(0)).val();
|
||||
|
||||
if head == tail {
|
||||
return 0;
|
||||
Ok(0)
|
||||
} else {
|
||||
// TAIL points at the last DWORD written, so the size is `tail - head + 4`.
|
||||
tail.checked_sub(head)
|
||||
.and_then(|delta| delta.checked_add(4))
|
||||
.ok_or(EIO)
|
||||
}
|
||||
|
||||
// TAIL points at last DWORD written, so add 4 to get total size.
|
||||
tail.saturating_sub(head).saturating_add(4)
|
||||
}
|
||||
|
||||
/// Writes `packet` to FSP EMEM and updates the queue pointers to notify FSP.
|
||||
///
|
||||
/// Returns `EINVAL` if `packet` is empty or its length is not 4-byte aligned.
|
||||
pub(crate) fn send_msg(&mut self, bar: Bar0<'_>, packet: &[u8]) -> Result {
|
||||
pub(crate) fn send_msg(&mut self, packet: &[u8]) -> Result {
|
||||
if packet.is_empty() {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
self.write_emem(bar, packet)?;
|
||||
self.write_emem(packet)?;
|
||||
|
||||
// Update queue pointers. TAIL points at the last DWORD written.
|
||||
let tail_offset = u32::try_from(packet.len() - 4).map_err(|_| EINVAL)?;
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
Array::at(0),
|
||||
regs::NV_PFSP_QUEUE_TAIL::zeroed().with_address(tail_offset),
|
||||
);
|
||||
bar.write(
|
||||
self.bar.write(
|
||||
Array::at(0),
|
||||
regs::NV_PFSP_QUEUE_HEAD::zeroed().with_address(0),
|
||||
);
|
||||
|
|
@ -148,23 +158,30 @@ pub(crate) fn send_msg(&mut self, bar: Bar0<'_>, packet: &[u8]) -> Result {
|
|||
///
|
||||
/// Returns `ETIMEDOUT` if no message was available until timeout, or a regular error code if a
|
||||
/// memory allocation error occurred.
|
||||
pub(crate) fn recv_msg(&mut self, bar: Bar0<'_>) -> Result<KVec<u8>> {
|
||||
pub(crate) fn recv_msg(&mut self) -> Result<KVec<u8>> {
|
||||
let msg_size = read_poll_timeout(
|
||||
|| Ok(self.poll_msgq(bar)),
|
||||
|| self.poll_msgq(),
|
||||
|&size| size > 0,
|
||||
Delta::from_millis(10),
|
||||
Delta::from_millis(FSP_MSG_TIMEOUT_MS),
|
||||
)
|
||||
.map(num::u32_as_usize)?;
|
||||
|
||||
// Don't blindly allocate more than the maximum we expect from FSP.
|
||||
if msg_size > FSP_EMEM_CHANNEL_0_SIZE {
|
||||
return Err(EMSGSIZE);
|
||||
}
|
||||
|
||||
let mut buffer = KVec::<u8>::new();
|
||||
buffer.resize(msg_size, 0, GFP_KERNEL)?;
|
||||
|
||||
self.read_emem(bar, &mut buffer)?;
|
||||
self.read_emem(&mut buffer)?;
|
||||
|
||||
// Reset message queue pointers after reading.
|
||||
bar.write(Array::at(0), regs::NV_PFSP_MSGQ_TAIL::zeroed().with_val(0));
|
||||
bar.write(Array::at(0), regs::NV_PFSP_MSGQ_HEAD::zeroed().with_val(0));
|
||||
self.bar
|
||||
.write(Array::at(0), regs::NV_PFSP_MSGQ_TAIL::zeroed().with_val(0));
|
||||
self.bar
|
||||
.write(Array::at(0), regs::NV_PFSP_MSGQ_HEAD::zeroed().with_val(0));
|
||||
|
||||
Ok(buffer)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -14,7 +14,6 @@
|
|||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
Falcon,
|
||||
FalconEngine,
|
||||
|
|
@ -24,10 +23,6 @@
|
|||
regs,
|
||||
};
|
||||
|
||||
/// Pattern returned by GSP register reads while the PRIV target mask still blocks CPU access.
|
||||
const GSP_TARGET_MASK_LOCKED_PATTERN: u32 = 0xbadf_4100;
|
||||
const GSP_TARGET_MASK_LOCKED_MASK: u32 = 0xffff_ff00;
|
||||
|
||||
/// Type specifying the `Gsp` falcon engine. Cannot be instantiated.
|
||||
pub(crate) struct Gsp(());
|
||||
|
||||
|
|
@ -41,20 +36,20 @@ impl RegisterBase<PFalcon2Base> for Gsp {
|
|||
|
||||
impl FalconEngine for Gsp {}
|
||||
|
||||
impl Falcon<Gsp> {
|
||||
impl<'a> Falcon<'a, Gsp> {
|
||||
/// Clears the SWGEN0 bit in the Falcon's IRQ status clear register to
|
||||
/// allow GSP to signal CPU for processing new messages in message queue.
|
||||
pub(crate) fn clear_swgen0_intr(&self, bar: Bar0<'_>) {
|
||||
bar.write(
|
||||
pub(crate) fn clear_swgen0_intr(&self) {
|
||||
self.bar.write(
|
||||
WithBase::of::<Gsp>(),
|
||||
regs::NV_PFALCON_FALCON_IRQSCLR::zeroed().with_swgen0(true),
|
||||
);
|
||||
}
|
||||
|
||||
/// Checks if GSP reload/resume has completed during the boot process.
|
||||
pub(crate) fn check_reload_completed(&self, bar: Bar0<'_>, timeout: Delta) -> Result<bool> {
|
||||
pub(crate) fn check_reload_completed(&self, timeout: Delta) -> Result<bool> {
|
||||
read_poll_timeout(
|
||||
|| Ok(bar.read(regs::NV_PGC6_BSI_SECURE_SCRATCH_14)),
|
||||
|| Ok(self.bar.read(regs::NV_PGC6_BSI_SECURE_SCRATCH_14)),
|
||||
|val| val.boot_stage_3_handoff(),
|
||||
Delta::ZERO,
|
||||
timeout,
|
||||
|
|
@ -63,17 +58,24 @@ pub(crate) fn check_reload_completed(&self, bar: Bar0<'_>, timeout: Delta) -> Re
|
|||
}
|
||||
|
||||
/// Returns whether the RISC-V branch privilege lockdown bit is set.
|
||||
pub(crate) fn riscv_branch_privilege_lockdown(&self, bar: Bar0<'_>) -> bool {
|
||||
bar.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<Gsp>())
|
||||
pub(crate) fn riscv_branch_privilege_lockdown(&self) -> bool {
|
||||
self.bar
|
||||
.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<Gsp>())
|
||||
.riscv_br_priv_lockdown()
|
||||
}
|
||||
|
||||
/// Returns whether GSP registers can be read by the CPU.
|
||||
pub(crate) fn priv_target_mask_released(&self, bar: Bar0<'_>) -> bool {
|
||||
let hwcfg2 = bar
|
||||
pub(crate) fn priv_target_mask_released(&self) -> bool {
|
||||
/// Pattern returned by GSP register reads while the PRIV target mask still blocks CPU
|
||||
/// access. The low byte varies; the upper 24 bits are fixed.
|
||||
const LOCKED_PATTERN: u32 = 0xbadf_4100;
|
||||
const LOCKED_MASK: u32 = 0xffff_ff00;
|
||||
|
||||
let hwcfg2 = self
|
||||
.bar
|
||||
.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<Gsp>())
|
||||
.into_raw();
|
||||
|
||||
hwcfg2 != 0 && (hwcfg2 & GSP_TARGET_MASK_LOCKED_MASK) != GSP_TARGET_MASK_LOCKED_PATTERN
|
||||
hwcfg2 != 0 && (hwcfg2 & LOCKED_MASK) != LOCKED_PATTERN
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -3,7 +3,6 @@
|
|||
use kernel::prelude::*;
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
Falcon,
|
||||
FalconBromParams,
|
||||
|
|
@ -34,7 +33,7 @@ pub(crate) enum LoadMethod {
|
|||
/// registers.
|
||||
pub(crate) trait FalconHal<E: FalconEngine>: Send + Sync {
|
||||
/// Activates the Falcon core if the engine is a risvc/falcon dual engine.
|
||||
fn select_core(&self, _falcon: &Falcon<E>, _bar: Bar0<'_>) -> Result {
|
||||
fn select_core(&self, _falcon: &Falcon<'_, E>) -> Result {
|
||||
Ok(())
|
||||
}
|
||||
|
||||
|
|
@ -42,24 +41,28 @@ fn select_core(&self, _falcon: &Falcon<E>, _bar: Bar0<'_>) -> Result {
|
|||
/// falcon instance. `engine_id_mask` and `ucode_id` are obtained from the firmware header.
|
||||
fn signature_reg_fuse_version(
|
||||
&self,
|
||||
falcon: &Falcon<E>,
|
||||
bar: Bar0<'_>,
|
||||
falcon: &Falcon<'_, E>,
|
||||
engine_id_mask: u16,
|
||||
ucode_id: u8,
|
||||
) -> Result<u32>;
|
||||
|
||||
/// Program the boot ROM registers prior to starting a secure firmware.
|
||||
fn program_brom(&self, falcon: &Falcon<E>, bar: Bar0<'_>, params: &FalconBromParams);
|
||||
fn program_brom(&self, falcon: &Falcon<'_, E>, params: &FalconBromParams);
|
||||
|
||||
/// Check if the RISC-V core is active.
|
||||
/// Returns `true` if the RISC-V core is active, `false` otherwise.
|
||||
fn is_riscv_active(&self, bar: Bar0<'_>) -> bool;
|
||||
fn is_riscv_active(&self, falcon: &Falcon<'_, E>) -> bool;
|
||||
|
||||
/// Checks whether the RISC-V core is halted.
|
||||
///
|
||||
/// Returns [`ENOTSUPP`] if the chipset does not expose RISC-V halt status.
|
||||
fn is_riscv_halted(&self, falcon: &Falcon<'_, E>) -> Result<bool>;
|
||||
|
||||
/// Wait for memory scrubbing to complete.
|
||||
fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result;
|
||||
fn reset_wait_mem_scrubbing(&self, falcon: &Falcon<'_, E>) -> Result;
|
||||
|
||||
/// Reset the falcon engine.
|
||||
fn reset_eng(&self, bar: Bar0<'_>) -> Result;
|
||||
fn reset_eng(&self, falcon: &Falcon<'_, E>) -> Result;
|
||||
|
||||
/// Returns the method used to load data into the falcon's memory.
|
||||
///
|
||||
|
|
|
|||
|
|
@ -115,33 +115,41 @@ pub(super) fn new() -> Self {
|
|||
}
|
||||
|
||||
impl<E: FalconEngine> FalconHal<E> for Ga102<E> {
|
||||
fn select_core(&self, _falcon: &Falcon<E>, bar: Bar0<'_>) -> Result {
|
||||
select_core_ga102::<E>(bar)
|
||||
fn select_core(&self, falcon: &Falcon<'_, E>) -> Result {
|
||||
select_core_ga102::<E>(falcon.bar)
|
||||
}
|
||||
|
||||
fn signature_reg_fuse_version(
|
||||
&self,
|
||||
falcon: &Falcon<E>,
|
||||
bar: Bar0<'_>,
|
||||
falcon: &Falcon<'_, E>,
|
||||
engine_id_mask: u16,
|
||||
ucode_id: u8,
|
||||
) -> Result<u32> {
|
||||
signature_reg_fuse_version_ga102(&falcon.dev, bar, engine_id_mask, ucode_id)
|
||||
signature_reg_fuse_version_ga102(falcon.dev, falcon.bar, engine_id_mask, ucode_id)
|
||||
}
|
||||
|
||||
fn program_brom(&self, _falcon: &Falcon<E>, bar: Bar0<'_>, params: &FalconBromParams) {
|
||||
program_brom_ga102::<E>(bar, params);
|
||||
fn program_brom(&self, falcon: &Falcon<'_, E>, params: &FalconBromParams) {
|
||||
program_brom_ga102::<E>(falcon.bar, params);
|
||||
}
|
||||
|
||||
fn is_riscv_active(&self, bar: Bar0<'_>) -> bool {
|
||||
bar.read(regs::NV_PRISCV_RISCV_CPUCTL::of::<E>())
|
||||
fn is_riscv_active(&self, falcon: &Falcon<'_, E>) -> bool {
|
||||
falcon
|
||||
.bar
|
||||
.read(regs::NV_PRISCV_RISCV_CPUCTL::of::<E>())
|
||||
.active_stat()
|
||||
}
|
||||
|
||||
fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result {
|
||||
fn is_riscv_halted(&self, falcon: &Falcon<'_, E>) -> Result<bool> {
|
||||
Ok(falcon
|
||||
.bar
|
||||
.read(regs::NV_PRISCV_RISCV_CPUCTL::of::<E>())
|
||||
.halted())
|
||||
}
|
||||
|
||||
fn reset_wait_mem_scrubbing(&self, falcon: &Falcon<'_, E>) -> Result {
|
||||
// TIMEOUT: memory scrubbing should complete in less than 20ms.
|
||||
read_poll_timeout(
|
||||
|| Ok(bar.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<E>())),
|
||||
|| Ok(falcon.bar.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<E>())),
|
||||
|r| r.mem_scrubbing_done(),
|
||||
Delta::ZERO,
|
||||
Delta::from_millis(20),
|
||||
|
|
@ -149,7 +157,9 @@ fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result {
|
|||
.map(|_| ())
|
||||
}
|
||||
|
||||
fn reset_eng(&self, bar: Bar0<'_>) -> Result {
|
||||
fn reset_eng(&self, falcon: &Falcon<'_, E>) -> Result {
|
||||
let bar = falcon.bar;
|
||||
|
||||
let _ = bar.read(regs::NV_PFALCON_FALCON_HWCFG2::of::<E>());
|
||||
|
||||
// According to OpenRM's `kflcnPreResetWait_GA102` documentation, HW sometimes does not set
|
||||
|
|
@ -162,7 +172,7 @@ fn reset_eng(&self, bar: Bar0<'_>) -> Result {
|
|||
);
|
||||
|
||||
regs::NV_PFALCON_FALCON_ENGINE::reset_engine::<E>(bar);
|
||||
self.reset_wait_mem_scrubbing(bar)?;
|
||||
self.reset_wait_mem_scrubbing(falcon)?;
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
|
|
|||
|
|
@ -13,7 +13,6 @@
|
|||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
hal::LoadMethod,
|
||||
Falcon,
|
||||
|
|
@ -34,31 +33,36 @@ pub(super) fn new() -> Self {
|
|||
}
|
||||
|
||||
impl<E: FalconEngine> FalconHal<E> for Tu102<E> {
|
||||
fn select_core(&self, _falcon: &Falcon<E>, _bar: Bar0<'_>) -> Result {
|
||||
fn select_core(&self, _falcon: &Falcon<'_, E>) -> Result {
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn signature_reg_fuse_version(
|
||||
&self,
|
||||
_falcon: &Falcon<E>,
|
||||
_bar: Bar0<'_>,
|
||||
_falcon: &Falcon<'_, E>,
|
||||
_engine_id_mask: u16,
|
||||
_ucode_id: u8,
|
||||
) -> Result<u32> {
|
||||
Ok(0)
|
||||
}
|
||||
|
||||
fn program_brom(&self, _falcon: &Falcon<E>, _bar: Bar0<'_>, _params: &FalconBromParams) {}
|
||||
fn program_brom(&self, _falcon: &Falcon<'_, E>, _params: &FalconBromParams) {}
|
||||
|
||||
fn is_riscv_active(&self, bar: Bar0<'_>) -> bool {
|
||||
bar.read(regs::NV_PRISCV_RISCV_CORE_SWITCH_RISCV_STATUS::of::<E>())
|
||||
fn is_riscv_active(&self, falcon: &Falcon<'_, E>) -> bool {
|
||||
falcon
|
||||
.bar
|
||||
.read(regs::NV_PRISCV_RISCV_CORE_SWITCH_RISCV_STATUS::of::<E>())
|
||||
.active_stat()
|
||||
}
|
||||
|
||||
fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result {
|
||||
fn is_riscv_halted(&self, _falcon: &Falcon<'_, E>) -> Result<bool> {
|
||||
Err(ENOTSUPP)
|
||||
}
|
||||
|
||||
fn reset_wait_mem_scrubbing(&self, falcon: &Falcon<'_, E>) -> Result {
|
||||
// TIMEOUT: memory scrubbing should complete in less than 10ms.
|
||||
read_poll_timeout(
|
||||
|| Ok(bar.read(regs::NV_PFALCON_FALCON_DMACTL::of::<E>())),
|
||||
|| Ok(falcon.bar.read(regs::NV_PFALCON_FALCON_DMACTL::of::<E>())),
|
||||
|r| r.mem_scrubbing_done(),
|
||||
Delta::ZERO,
|
||||
Delta::from_millis(10),
|
||||
|
|
@ -66,9 +70,9 @@ fn reset_wait_mem_scrubbing(&self, bar: Bar0<'_>) -> Result {
|
|||
.map(|_| ())
|
||||
}
|
||||
|
||||
fn reset_eng(&self, bar: Bar0<'_>) -> Result {
|
||||
regs::NV_PFALCON_FALCON_ENGINE::reset_engine::<E>(bar);
|
||||
self.reset_wait_mem_scrubbing(bar)?;
|
||||
fn reset_eng(&self, falcon: &Falcon<'_, E>) -> Result {
|
||||
regs::NV_PFALCON_FALCON_ENGINE::reset_engine::<E>(falcon.bar);
|
||||
self.reset_wait_mem_scrubbing(falcon)?;
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
|
|
|||
|
|
@ -24,10 +24,11 @@
|
|||
gpu::Chipset,
|
||||
gsp,
|
||||
num::FromSafeCast,
|
||||
regs, //
|
||||
vgpu::VgpuState, //
|
||||
};
|
||||
|
||||
mod hal;
|
||||
mod regs;
|
||||
|
||||
/// Type holding the sysmem flush memory page, a page of memory to be written into the
|
||||
/// `NV_PFB_NISO_FLUSH_SYSMEM_ADDR*` registers and used to maintain memory coherency.
|
||||
|
|
@ -60,7 +61,7 @@ pub(crate) fn register(
|
|||
) -> Result<Self> {
|
||||
let page = CoherentHandle::alloc(dev, kernel::page::PAGE_SIZE, GFP_KERNEL)?;
|
||||
|
||||
hal::fb_hal(chipset).write_sysmem_flush_page(bar, page.dma_handle())?;
|
||||
hal::fb_hal(chipset).write_sysmem_flush_page(bar, page.dma_address())?;
|
||||
|
||||
Ok(Self {
|
||||
chipset,
|
||||
|
|
@ -75,7 +76,7 @@ impl Drop for SysmemFlush<'_> {
|
|||
fn drop(&mut self) {
|
||||
let hal = hal::fb_hal(self.chipset);
|
||||
|
||||
if hal.read_sysmem_flush_page(self.bar) == self.page.dma_handle() {
|
||||
if hal.read_sysmem_flush_page(self.bar) == self.page.dma_address() {
|
||||
let _ = hal.write_sysmem_flush_page(self.bar, 0).inspect_err(|e| {
|
||||
dev_warn!(
|
||||
&self.device,
|
||||
|
|
@ -148,7 +149,7 @@ fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
|||
///
|
||||
/// Contains ranges of GPU memory reserved for a given purpose during the GSP boot process.
|
||||
#[derive(Debug)]
|
||||
pub(crate) struct FbLayout {
|
||||
pub(crate) struct FbRanges {
|
||||
/// Range of the framebuffer. Starts at `0`.
|
||||
pub(crate) fb: FbRange,
|
||||
/// VGA workspace, small area of reserved memory at the end of the framebuffer.
|
||||
|
|
@ -163,15 +164,22 @@ pub(crate) struct FbLayout {
|
|||
pub(crate) wpr2_heap: FbRange,
|
||||
/// WPR2 region range, starting with an instance of `GspFwWprMeta`.
|
||||
pub(crate) wpr2: FbRange,
|
||||
pub(crate) heap: FbRange,
|
||||
/// Non-WPR heap, located just below WPR2.
|
||||
pub(crate) non_wpr_heap: FbRange,
|
||||
/// Number of VF partitions.
|
||||
pub(crate) vf_partition_count: u8,
|
||||
/// PMU reserved memory size, in bytes.
|
||||
pub(crate) pmu_reserved_size: u32,
|
||||
}
|
||||
|
||||
impl FbLayout {
|
||||
/// Computes the FB layout for `chipset` required to run the `gsp_fw` GSP firmware.
|
||||
pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, gsp_fw: &GspFirmware) -> Result<Self> {
|
||||
impl FbRanges {
|
||||
/// Computes concrete framebuffer ranges required on non-FSP booting architectures.
|
||||
pub(crate) fn new(
|
||||
chipset: Chipset,
|
||||
bar: Bar0<'_>,
|
||||
gsp_fw: &GspFirmware,
|
||||
vgpu_state: VgpuState,
|
||||
) -> Result<Self> {
|
||||
let hal = hal::fb_hal(chipset);
|
||||
|
||||
let fb = {
|
||||
|
|
@ -234,11 +242,15 @@ pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, gsp_fw: &GspFirmware) -> Resu
|
|||
FbRange(elf_addr..elf_addr + elf_size)
|
||||
};
|
||||
|
||||
let (vf_partition_count, wpr2_heap_size) = wpr2_heap_params(chipset, vgpu_state, fb.end)?;
|
||||
|
||||
let wpr2_heap = {
|
||||
const WPR2_HEAP_DOWN_ALIGN: Alignment = Alignment::new::<SZ_1M>();
|
||||
let wpr2_heap_size =
|
||||
gsp::LibosParams::from_chipset(chipset).wpr_heap_size(chipset, fb.end)?;
|
||||
let wpr2_heap_addr = (elf.start - wpr2_heap_size).align_down(WPR2_HEAP_DOWN_ALIGN);
|
||||
let wpr2_heap_addr = elf
|
||||
.start
|
||||
.checked_sub(wpr2_heap_size)
|
||||
.ok_or(EOVERFLOW)?
|
||||
.align_down(WPR2_HEAP_DOWN_ALIGN);
|
||||
|
||||
FbRange(wpr2_heap_addr..(elf.start).align_down(WPR2_HEAP_DOWN_ALIGN))
|
||||
};
|
||||
|
|
@ -251,9 +263,9 @@ pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, gsp_fw: &GspFirmware) -> Resu
|
|||
FbRange(wpr2_addr..frts.end)
|
||||
};
|
||||
|
||||
let heap = {
|
||||
let heap_size = u64::from(hal.non_wpr_heap_size());
|
||||
FbRange(wpr2.start - heap_size..wpr2.start)
|
||||
let non_wpr_heap = {
|
||||
let non_wpr_heap_size = hal.non_wpr_heap_size();
|
||||
FbRange(wpr2.start - non_wpr_heap_size..wpr2.start)
|
||||
};
|
||||
|
||||
Ok(Self {
|
||||
|
|
@ -264,9 +276,69 @@ pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, gsp_fw: &GspFirmware) -> Resu
|
|||
elf,
|
||||
wpr2_heap,
|
||||
wpr2,
|
||||
heap,
|
||||
vf_partition_count: 0,
|
||||
non_wpr_heap,
|
||||
vf_partition_count,
|
||||
pmu_reserved_size: hal.pmu_reserved_size(),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// Reads the WPR2 memory region registers and returns the range if set.
|
||||
/// Returns `None` if the WPR2 region is not set.
|
||||
pub(crate) fn wpr2_range(bar: Bar0<'_>) -> Option<Range<u64>> {
|
||||
let wpr2_hi = bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI);
|
||||
|
||||
if !wpr2_hi.is_wpr2_set() {
|
||||
return None;
|
||||
}
|
||||
|
||||
let wpr2_lo = bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_LO);
|
||||
|
||||
Some(wpr2_lo.lower_bound()..wpr2_hi.higher_bound())
|
||||
}
|
||||
|
||||
/// Computes the number of VF partitions and the WPR2 heap size from the vGPU state.
|
||||
fn wpr2_heap_params(chipset: Chipset, vgpu_state: VgpuState, fb_size: u64) -> Result<(u8, u64)> {
|
||||
Ok(match vgpu_state {
|
||||
VgpuState::Disabled => (
|
||||
0,
|
||||
gsp::LibosParams::from_chipset(chipset).wpr_heap_size(chipset, fb_size)?,
|
||||
),
|
||||
VgpuState::Enabled { total_vfs } => (
|
||||
u8::try_from(total_vfs.get()).map_err(|_| EINVAL)?,
|
||||
gsp::LibosParams::vgpu_wpr_heap_size(),
|
||||
),
|
||||
})
|
||||
}
|
||||
|
||||
/// Framebuffer region sizes needed for GSP-FMC boot.
|
||||
#[derive(Debug)]
|
||||
pub(crate) struct FbSizes {
|
||||
/// FRTS size, in bytes.
|
||||
pub(crate) frts_size: u64,
|
||||
/// WPR2 heap size, in bytes.
|
||||
pub(crate) wpr2_heap_size: u64,
|
||||
/// Non-WPR heap size, in bytes.
|
||||
pub(crate) non_wpr_heap_size: u64,
|
||||
/// PMU reserved memory size, in bytes.
|
||||
pub(crate) pmu_reserved_size: u32,
|
||||
/// Number of VF partitions.
|
||||
pub(crate) vf_partition_count: u8,
|
||||
}
|
||||
|
||||
impl FbSizes {
|
||||
/// Computes the framebuffer region sizes for GSP-FMC boot.
|
||||
pub(crate) fn new(chipset: Chipset, bar: Bar0<'_>, vgpu_state: VgpuState) -> Result<Self> {
|
||||
let hal = hal::fb_hal(chipset);
|
||||
let fb_size = hal.vidmem_size(bar);
|
||||
let (vf_partition_count, wpr2_heap_size) = wpr2_heap_params(chipset, vgpu_state, fb_size)?;
|
||||
|
||||
Ok(Self {
|
||||
frts_size: hal.frts_size(),
|
||||
wpr2_heap_size,
|
||||
non_wpr_heap_size: hal.non_wpr_heap_size(),
|
||||
pmu_reserved_size: hal.pmu_reserved_size(),
|
||||
vf_partition_count,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -37,7 +37,7 @@ pub(crate) trait FbHal {
|
|||
fn pmu_reserved_size(&self) -> u32;
|
||||
|
||||
/// Returns the non-WPR heap size for this chipset, in bytes.
|
||||
fn non_wpr_heap_size(&self) -> u32;
|
||||
fn non_wpr_heap_size(&self) -> u64;
|
||||
|
||||
/// Returns the FRTS size, in bytes.
|
||||
fn frts_size(&self) -> u64;
|
||||
|
|
|
|||
|
|
@ -9,8 +9,10 @@
|
|||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
fb::hal::FbHal,
|
||||
regs, //
|
||||
fb::{
|
||||
hal::FbHal,
|
||||
regs, //
|
||||
},
|
||||
};
|
||||
|
||||
use super::tu102::FLUSH_SYSMEM_ADDR_SHIFT;
|
||||
|
|
@ -41,7 +43,7 @@ pub(super) fn write_sysmem_flush_page_ga100(bar: Bar0<'_>, addr: u64) {
|
|||
}
|
||||
|
||||
pub(super) fn display_enabled_ga100(bar: Bar0<'_>) -> bool {
|
||||
!bar.read(regs::ga100::NV_FUSE_STATUS_OPT_DISPLAY)
|
||||
!bar.read(crate::regs::ga100::NV_FUSE_STATUS_OPT_DISPLAY)
|
||||
.display_disabled()
|
||||
}
|
||||
|
||||
|
|
@ -72,7 +74,7 @@ fn pmu_reserved_size(&self) -> u32 {
|
|||
super::tu102::pmu_reserved_size_tu102()
|
||||
}
|
||||
|
||||
fn non_wpr_heap_size(&self) -> u32 {
|
||||
fn non_wpr_heap_size(&self) -> u64 {
|
||||
super::tu102::non_wpr_heap_size_tu102()
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -41,7 +41,7 @@ fn pmu_reserved_size(&self) -> u32 {
|
|||
super::tu102::pmu_reserved_size_tu102()
|
||||
}
|
||||
|
||||
fn non_wpr_heap_size(&self) -> u32 {
|
||||
fn non_wpr_heap_size(&self) -> u64 {
|
||||
super::tu102::non_wpr_heap_size_tu102()
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -22,9 +22,11 @@
|
|||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
fb::hal::FbHal,
|
||||
fb::{
|
||||
hal::FbHal,
|
||||
regs, //
|
||||
},
|
||||
num::usize_into_u32,
|
||||
regs, //
|
||||
};
|
||||
|
||||
struct Gb100;
|
||||
|
|
@ -78,6 +80,7 @@ fn write_sysmem_flush_page_gb100(bar: Bar0<'_>, addr: Bounded<u64, 52>) {
|
|||
);
|
||||
}
|
||||
|
||||
// This PMU reservation size is r570-specific.
|
||||
pub(super) const fn pmu_reserved_size_gb100() -> u32 {
|
||||
usize_into_u32::<{ const_align_up(SZ_8M + SZ_16M + SZ_4K, Alignment::new::<SZ_128K>()).unwrap() }>(
|
||||
)
|
||||
|
|
@ -108,9 +111,9 @@ fn pmu_reserved_size(&self) -> u32 {
|
|||
pmu_reserved_size_gb100()
|
||||
}
|
||||
|
||||
fn non_wpr_heap_size(&self) -> u32 {
|
||||
fn non_wpr_heap_size(&self) -> u64 {
|
||||
// Non-WPR heap for GB10x (see Open RM: kgspGetNonWprHeapSize, GB100/GB102).
|
||||
u32::SZ_2M
|
||||
u64::SZ_2M
|
||||
}
|
||||
|
||||
fn frts_size(&self) -> u64 {
|
||||
|
|
|
|||
|
|
@ -4,13 +4,7 @@
|
|||
//! Blackwell GB20x framebuffer HAL.
|
||||
|
||||
use kernel::{
|
||||
io::{
|
||||
register::{
|
||||
RegisterBase,
|
||||
WithBase, //
|
||||
},
|
||||
Io, //
|
||||
},
|
||||
io::Io,
|
||||
num::Bounded,
|
||||
prelude::*,
|
||||
sizes::SizeConstants, //
|
||||
|
|
@ -18,41 +12,37 @@
|
|||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
fb::hal::FbHal,
|
||||
regs, //
|
||||
fb::{
|
||||
hal::FbHal,
|
||||
regs, //
|
||||
},
|
||||
};
|
||||
|
||||
struct Gb202;
|
||||
|
||||
impl RegisterBase<regs::Fbhub0Base> for Gb202 {
|
||||
const BASE: usize = 0x008a_0000;
|
||||
}
|
||||
|
||||
fn read_sysmem_flush_page_gb202(bar: Bar0<'_>) -> u64 {
|
||||
let lo = u64::from(
|
||||
bar.read(regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO::of::<Gb202>())
|
||||
bar.read(regs::NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_LO)
|
||||
.adr(),
|
||||
);
|
||||
let hi = u64::from(
|
||||
bar.read(regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI::of::<Gb202>())
|
||||
bar.read(regs::NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_HI)
|
||||
.adr(),
|
||||
);
|
||||
|
||||
lo | (hi << 32)
|
||||
(hi << 32) | lo
|
||||
}
|
||||
|
||||
/// Write the sysmem flush page address through the GB20x FBHUB0 registers.
|
||||
fn write_sysmem_flush_page_gb202(bar: Bar0<'_>, addr: Bounded<u64, 52>) {
|
||||
// Write HI first. The hardware will trigger the flush on the LO write.
|
||||
bar.write(
|
||||
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI::of::<Gb202>(),
|
||||
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI::zeroed()
|
||||
bar.write_reg(
|
||||
regs::NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_HI::zeroed()
|
||||
.with_adr(addr.shr::<32, 20>().cast::<u32>()),
|
||||
);
|
||||
bar.write(
|
||||
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO::of::<Gb202>(),
|
||||
bar.write_reg(
|
||||
// CAST: lower 32 bits. Hardware ignores bits 7:0.
|
||||
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO::zeroed().with_adr(*addr as u32),
|
||||
regs::NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_LO::zeroed().with_adr(*addr as u32),
|
||||
);
|
||||
}
|
||||
|
||||
|
|
@ -81,9 +71,10 @@ fn pmu_reserved_size(&self) -> u32 {
|
|||
super::gb100::pmu_reserved_size_gb100()
|
||||
}
|
||||
|
||||
fn non_wpr_heap_size(&self) -> u32 {
|
||||
fn non_wpr_heap_size(&self) -> u64 {
|
||||
// Non-WPR heap for GB20x (see Open RM: kgspGetNonWprHeapSize, GB202+).
|
||||
u32::SZ_2M + u32::SZ_128K
|
||||
// This size is r570-specific.
|
||||
u64::SZ_2M + u64::SZ_128K
|
||||
}
|
||||
|
||||
fn frts_size(&self) -> u64 {
|
||||
|
|
|
|||
|
|
@ -2,24 +2,51 @@
|
|||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use kernel::{
|
||||
io::Io,
|
||||
num::Bounded,
|
||||
prelude::*,
|
||||
sizes::SizeConstants, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
fb::hal::FbHal, //
|
||||
fb::{
|
||||
hal::FbHal,
|
||||
regs, //
|
||||
},
|
||||
};
|
||||
|
||||
struct Gh100;
|
||||
|
||||
fn read_sysmem_flush_page_gh100(bar: Bar0<'_>) -> u64 {
|
||||
let lo = u64::from(bar.read(regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO).adr());
|
||||
let hi = u64::from(bar.read(regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI).adr());
|
||||
|
||||
(hi << 32) | lo
|
||||
}
|
||||
|
||||
/// Write the sysmem flush page address through the Hopper FBHUB registers.
|
||||
fn write_sysmem_flush_page_gh100(bar: Bar0<'_>, addr: Bounded<u64, 52>) {
|
||||
// Write HI first. The hardware will trigger the flush on the LO write.
|
||||
bar.write_reg(
|
||||
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI::zeroed()
|
||||
.with_adr(addr.shr::<32, 20>().cast::<u32>()),
|
||||
);
|
||||
bar.write_reg(
|
||||
// CAST: lower 32 bits. Hardware ignores bits 7:0.
|
||||
regs::NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO::zeroed().with_adr(*addr as u32),
|
||||
);
|
||||
}
|
||||
|
||||
impl FbHal for Gh100 {
|
||||
fn read_sysmem_flush_page(&self, bar: Bar0<'_>) -> u64 {
|
||||
super::ga100::read_sysmem_flush_page_ga100(bar)
|
||||
read_sysmem_flush_page_gh100(bar)
|
||||
}
|
||||
|
||||
fn write_sysmem_flush_page(&self, bar: Bar0<'_>, addr: u64) -> Result {
|
||||
super::ga100::write_sysmem_flush_page_ga100(bar, addr);
|
||||
let addr = Bounded::<u64, 52>::try_new(addr).ok_or(EINVAL)?;
|
||||
|
||||
write_sysmem_flush_page_gh100(bar, addr);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
|
@ -36,9 +63,9 @@ fn pmu_reserved_size(&self) -> u32 {
|
|||
super::tu102::pmu_reserved_size_tu102()
|
||||
}
|
||||
|
||||
fn non_wpr_heap_size(&self) -> u32 {
|
||||
fn non_wpr_heap_size(&self) -> u64 {
|
||||
// Non-WPR heap for Hopper (see Open RM: kgspCalculateFbLayout_GH100).
|
||||
u32::SZ_2M
|
||||
u64::SZ_2M
|
||||
}
|
||||
|
||||
fn frts_size(&self) -> u64 {
|
||||
|
|
|
|||
|
|
@ -9,8 +9,10 @@
|
|||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
fb::hal::FbHal,
|
||||
regs, //
|
||||
fb::{
|
||||
hal::FbHal,
|
||||
regs, //
|
||||
},
|
||||
};
|
||||
|
||||
/// Shift applied to the sysmem address before it is written into `NV_PFB_NISO_FLUSH_SYSMEM_ADDR`,
|
||||
|
|
@ -31,7 +33,7 @@ pub(super) fn write_sysmem_flush_page_gm107(bar: Bar0<'_>, addr: u64) -> Result
|
|||
}
|
||||
|
||||
pub(super) fn display_enabled_gm107(bar: Bar0<'_>) -> bool {
|
||||
!bar.read(regs::gm107::NV_FUSE_STATUS_OPT_DISPLAY)
|
||||
!bar.read(crate::regs::gm107::NV_FUSE_STATUS_OPT_DISPLAY)
|
||||
.display_disabled()
|
||||
}
|
||||
|
||||
|
|
@ -44,8 +46,8 @@ pub(super) const fn pmu_reserved_size_tu102() -> u32 {
|
|||
0
|
||||
}
|
||||
|
||||
pub(super) const fn non_wpr_heap_size_tu102() -> u32 {
|
||||
u32::SZ_1M
|
||||
pub(super) const fn non_wpr_heap_size_tu102() -> u64 {
|
||||
u64::SZ_1M
|
||||
}
|
||||
|
||||
pub(super) const fn frts_size_tu102() -> u64 {
|
||||
|
|
@ -75,7 +77,7 @@ fn pmu_reserved_size(&self) -> u32 {
|
|||
pmu_reserved_size_tu102()
|
||||
}
|
||||
|
||||
fn non_wpr_heap_size(&self) -> u32 {
|
||||
fn non_wpr_heap_size(&self) -> u64 {
|
||||
non_wpr_heap_size_tu102()
|
||||
}
|
||||
|
||||
|
|
|
|||
155
drivers/gpu/nova-core/fb/regs.rs
Normal file
155
drivers/gpu/nova-core/fb/regs.rs
Normal file
|
|
@ -0,0 +1,155 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
use kernel::{
|
||||
io::register,
|
||||
sizes::SizeConstants, //
|
||||
};
|
||||
|
||||
// PDISP
|
||||
|
||||
register! {
|
||||
pub(super) NV_PDISP_VGA_WORKSPACE_BASE(u32) @ 0x00625f04 {
|
||||
/// VGA workspace base address divided by 0x10000.
|
||||
31:8 addr;
|
||||
/// Set if the `addr` field is valid.
|
||||
3:3 status_valid => bool;
|
||||
}
|
||||
}
|
||||
|
||||
impl NV_PDISP_VGA_WORKSPACE_BASE {
|
||||
/// Returns the base address of the VGA workspace, or `None` if none exists.
|
||||
pub(super) fn vga_workspace_addr(self) -> Option<u64> {
|
||||
if self.status_valid() {
|
||||
Some(u64::from(self.addr()) << 16)
|
||||
} else {
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// PFB
|
||||
|
||||
register! {
|
||||
/// Low bits of the physical system memory address used by the GPU to perform sysmembar
|
||||
/// operations (see [`crate::fb::SysmemFlush`]).
|
||||
pub(super) NV_PFB_NISO_FLUSH_SYSMEM_ADDR(u32) @ 0x00100c10 {
|
||||
31:0 adr_39_08;
|
||||
}
|
||||
|
||||
/// High bits of the physical system memory address used by the GPU to perform sysmembar
|
||||
/// operations.
|
||||
pub(super) NV_PFB_NISO_FLUSH_SYSMEM_ADDR_HI(u32) @ 0x00100c40 {
|
||||
23:0 adr_63_40;
|
||||
}
|
||||
|
||||
pub(super) NV_PFB_PRI_MMU_LOCAL_MEMORY_RANGE(u32) @ 0x00100ce0 {
|
||||
30:30 ecc_mode_enabled => bool;
|
||||
9:4 lower_mag;
|
||||
3:0 lower_scale;
|
||||
}
|
||||
|
||||
pub(super) NV_PFB_PRI_MMU_WPR2_ADDR_LO(u32) @ 0x001fa824 {
|
||||
/// Bits 12..40 of the lower (inclusive) bound of the WPR2 region.
|
||||
31:4 lo_val;
|
||||
}
|
||||
|
||||
pub(super) NV_PFB_PRI_MMU_WPR2_ADDR_HI(u32) @ 0x001fa828 {
|
||||
/// Bits 12..40 of the higher (exclusive) bound of the WPR2 region.
|
||||
31:4 hi_val;
|
||||
}
|
||||
}
|
||||
|
||||
/// Base of the GB10x HSHUB0 register window (`NV_HSHUB0_PRIV_BASE` in Open RM).
|
||||
///
|
||||
/// The base is provided by the GB10x framebuffer HAL.
|
||||
pub(super) struct Hshub0Base(());
|
||||
|
||||
register! {
|
||||
// GB10x sysmem flush registers, relative to the HSHUB0 base. GB10x routes sysmembar
|
||||
// through a primary and an EG (egress) pair that must both be programmed to the same
|
||||
// address. Hardware ignores bits 7:0 of each LO register. The boot path uses a fixed
|
||||
// HSHUB0 base, so the multiple runtime-discovered HSHUB bases are not needed here.
|
||||
pub(super) NV_PFB_HSHUB_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Hshub0Base + 0x00000e50 {
|
||||
31:0 adr => u32;
|
||||
}
|
||||
|
||||
pub(super) NV_PFB_HSHUB_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Hshub0Base + 0x00000e54 {
|
||||
19:0 adr;
|
||||
}
|
||||
|
||||
pub(super) NV_PFB_HSHUB_EG_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Hshub0Base + 0x000006c0 {
|
||||
31:0 adr => u32;
|
||||
}
|
||||
|
||||
pub(super) NV_PFB_HSHUB_EG_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Hshub0Base + 0x000006c4 {
|
||||
19:0 adr;
|
||||
}
|
||||
}
|
||||
|
||||
register! {
|
||||
// GB20x FBHUB0 sysmem flush registers. Unlike the older
|
||||
// NV_PFB_NISO_FLUSH_SYSMEM_ADDR registers, which encode the address with an
|
||||
// 8-bit right-shift, these take the raw address split into lower and upper
|
||||
// halves. Hardware ignores bits 7:0 of the LO register.
|
||||
pub(super) NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ 0x008a1d58 {
|
||||
31:0 adr => u32;
|
||||
}
|
||||
|
||||
pub(super) NV_PFB_FBHUB0_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ 0x008a1d5c {
|
||||
19:0 adr;
|
||||
}
|
||||
}
|
||||
|
||||
register! {
|
||||
/// Low bits of the physical system memory address used by the GPU to perform
|
||||
/// sysmembar operations on Hopper.
|
||||
///
|
||||
/// Like the GB20x FBHUB0 registers, and unlike the Ampere
|
||||
/// `NV_PFB_NISO_FLUSH_SYSMEM_ADDR` registers (which encode the address with an
|
||||
/// 8-bit right-shift), these take the raw address split into lower and upper
|
||||
/// halves. Hardware ignores bits 7:0 of the LO register.
|
||||
pub(super) NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ 0x00100a34 {
|
||||
31:0 adr => u32;
|
||||
}
|
||||
|
||||
/// High bits of the physical system memory address used by the GPU to perform
|
||||
/// sysmembar operations on Hopper.
|
||||
pub(super) NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ 0x00100a38 {
|
||||
19:0 adr;
|
||||
}
|
||||
}
|
||||
|
||||
impl NV_PFB_PRI_MMU_LOCAL_MEMORY_RANGE {
|
||||
/// Returns the usable framebuffer size, in bytes.
|
||||
pub(super) fn usable_fb_size(self) -> u64 {
|
||||
let size = (u64::from(self.lower_mag()) << u64::from(self.lower_scale())) * u64::SZ_1M;
|
||||
|
||||
if self.ecc_mode_enabled() {
|
||||
// Remove the amount of memory reserved for ECC (one per 16 units).
|
||||
size / 16 * 15
|
||||
} else {
|
||||
size
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl NV_PFB_PRI_MMU_WPR2_ADDR_LO {
|
||||
/// Returns the lower (inclusive) bound of the WPR2 region.
|
||||
pub(super) fn lower_bound(self) -> u64 {
|
||||
u64::from(self.lo_val()) << 12
|
||||
}
|
||||
}
|
||||
|
||||
impl NV_PFB_PRI_MMU_WPR2_ADDR_HI {
|
||||
/// Returns the higher (exclusive) bound of the WPR2 region.
|
||||
///
|
||||
/// A value of zero means the WPR2 region is not set.
|
||||
pub(super) fn higher_bound(self) -> u64 {
|
||||
u64::from(self.hi_val()) << 12
|
||||
}
|
||||
|
||||
/// Returns whether the WPR2 region is currently set.
|
||||
pub(super) fn is_wpr2_set(self) -> bool {
|
||||
self.hi_val() != 0
|
||||
}
|
||||
}
|
||||
|
|
@ -8,11 +8,8 @@
|
|||
use core::ops::Deref;
|
||||
|
||||
use kernel::{
|
||||
device,
|
||||
firmware,
|
||||
prelude::*,
|
||||
str::CString,
|
||||
transmute::FromBytes, //
|
||||
prelude::*, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
|
|
@ -21,10 +18,8 @@
|
|||
FalconFirmware, //
|
||||
},
|
||||
gpu,
|
||||
num::{
|
||||
FromSafeCast,
|
||||
IntoSafeCast, //
|
||||
},
|
||||
gsp::boot_firmware_files,
|
||||
num::IntoSafeCast, //
|
||||
};
|
||||
|
||||
pub(crate) mod booter;
|
||||
|
|
@ -32,21 +27,7 @@
|
|||
pub(crate) mod fwsec;
|
||||
pub(crate) mod gsp;
|
||||
pub(crate) mod riscv;
|
||||
|
||||
pub(crate) const FIRMWARE_VERSION: &str = "570.144";
|
||||
|
||||
/// Requests the GPU firmware `name` suitable for `chipset`, with version `ver`.
|
||||
fn request_firmware(
|
||||
dev: &device::Device,
|
||||
chipset: gpu::Chipset,
|
||||
name: &str,
|
||||
ver: &str,
|
||||
) -> Result<firmware::Firmware> {
|
||||
let chip_name = chipset.name();
|
||||
|
||||
CString::try_from_fmt(fmt!("nvidia/{chip_name}/gsp/{name}-{ver}.bin"))
|
||||
.and_then(|path| firmware::Firmware::request(&path, dev))
|
||||
}
|
||||
pub(crate) mod tlv;
|
||||
|
||||
/// Structure used to describe some firmwares, notably FWSEC-FRTS.
|
||||
#[repr(C)]
|
||||
|
|
@ -88,7 +69,7 @@ pub(crate) struct FalconUCodeDescV2 {
|
|||
|
||||
/// Structure used to describe some firmwares, notably FWSEC-FRTS.
|
||||
#[repr(C)]
|
||||
#[derive(Debug, Clone)]
|
||||
#[derive(Debug, Clone, FromBytes)]
|
||||
pub(crate) struct FalconUCodeDescV3 {
|
||||
/// Header defined by `NV_BIT_FALCON_UCODE_DESC_HEADER_VDESC*` in OpenRM.
|
||||
hdr: u32,
|
||||
|
|
@ -119,10 +100,6 @@ pub(crate) struct FalconUCodeDescV3 {
|
|||
_reserved: u16,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use
|
||||
// interior mutability.
|
||||
unsafe impl FromBytes for FalconUCodeDescV3 {}
|
||||
|
||||
/// Enum wrapping the different versions of Falcon microcode descriptors.
|
||||
///
|
||||
/// This allows handling both V2 and V3 descriptor formats through a
|
||||
|
|
@ -349,60 +326,6 @@ fn no_patch_signature(self) -> FirmwareObject<F, Signed> {
|
|||
}
|
||||
}
|
||||
|
||||
/// Header common to most firmware files.
|
||||
#[repr(C)]
|
||||
#[derive(Debug, Clone)]
|
||||
struct BinHdr {
|
||||
/// Magic number, must be `0x10de`.
|
||||
bin_magic: u32,
|
||||
/// Version of the header.
|
||||
bin_ver: u32,
|
||||
/// Size in bytes of the binary (to be ignored).
|
||||
bin_size: u32,
|
||||
/// Offset of the start of the application-specific header.
|
||||
header_offset: u32,
|
||||
/// Offset of the start of the data payload.
|
||||
data_offset: u32,
|
||||
/// Size in bytes of the data payload.
|
||||
data_size: u32,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for BinHdr {}
|
||||
|
||||
// A firmware blob starting with a `BinHdr`.
|
||||
struct BinFirmware<'a> {
|
||||
hdr: BinHdr,
|
||||
fw: &'a [u8],
|
||||
}
|
||||
|
||||
impl<'a> BinFirmware<'a> {
|
||||
/// Interpret `fw` as a firmware image starting with a [`BinHdr`], and returns the
|
||||
/// corresponding [`BinFirmware`] that can be used to extract its payload.
|
||||
fn new(fw: &'a firmware::Firmware) -> Result<Self> {
|
||||
const BIN_MAGIC: u32 = 0x10de;
|
||||
let fw = fw.data();
|
||||
|
||||
fw.get(0..size_of::<BinHdr>())
|
||||
// Extract header.
|
||||
.and_then(BinHdr::from_bytes_copy)
|
||||
// Validate header.
|
||||
.filter(|hdr| hdr.bin_magic == BIN_MAGIC)
|
||||
.map(|hdr| Self { hdr, fw })
|
||||
.ok_or(EINVAL)
|
||||
}
|
||||
|
||||
/// Returns the data payload of the firmware, or `None` if the data range is out of bounds of
|
||||
/// the firmware image.
|
||||
fn data(&self) -> Option<&[u8]> {
|
||||
let fw_start = usize::from_safe_cast(self.hdr.data_offset);
|
||||
let fw_size = usize::from_safe_cast(self.hdr.data_size);
|
||||
let fw_end = fw_start.checked_add(fw_size)?;
|
||||
|
||||
self.fw.get(fw_start..fw_end)
|
||||
}
|
||||
}
|
||||
|
||||
pub(crate) struct ModInfoBuilder<const N: usize>(firmware::ModInfoBuilder<N>);
|
||||
|
||||
impl<const N: usize> ModInfoBuilder<N> {
|
||||
|
|
@ -413,33 +336,28 @@ const fn make_entry_file(self, chipset: &str, fw: &str) -> Self {
|
|||
.push("nvidia/")
|
||||
.push(chipset)
|
||||
.push("/gsp/")
|
||||
.push(fw)
|
||||
.push("-")
|
||||
.push(FIRMWARE_VERSION)
|
||||
.push(".bin"),
|
||||
.push(fw),
|
||||
)
|
||||
}
|
||||
|
||||
const fn make_entry_chipset(self, chipset: gpu::Chipset) -> Self {
|
||||
let name = chipset.name();
|
||||
|
||||
let this = self
|
||||
.make_entry_file(name, "booter_load")
|
||||
.make_entry_file(name, "booter_unload")
|
||||
.make_entry_file(name, "bootloader")
|
||||
.make_entry_file(name, "gsp");
|
||||
// GSP firmware files are always present.
|
||||
let mut this = self
|
||||
.make_entry_file(name, "gsp_bootloader.tlv")
|
||||
.make_entry_file(name, "gsp.tlv")
|
||||
.make_entry_file(name, "gsp.bin");
|
||||
|
||||
let this = if chipset.needs_fwsec_bootloader() {
|
||||
this.make_entry_file(name, "gen_bootloader")
|
||||
} else {
|
||||
this
|
||||
};
|
||||
|
||||
if chipset.uses_fsp() {
|
||||
this.make_entry_file(name, "fmc")
|
||||
} else {
|
||||
this
|
||||
// Add the firmware files specific to the GSP boot method of `chipset`.
|
||||
let boot_files = boot_firmware_files(chipset);
|
||||
let mut i = 0;
|
||||
while i < boot_files.len() {
|
||||
this = this.make_entry_file(name, boot_files[i]);
|
||||
i += 1;
|
||||
}
|
||||
|
||||
this
|
||||
}
|
||||
|
||||
pub(crate) const fn create(
|
||||
|
|
@ -456,210 +374,3 @@ pub(crate) const fn create(
|
|||
this.0
|
||||
}
|
||||
}
|
||||
|
||||
/// Ad-hoc and temporary module to extract sections from ELF images.
|
||||
///
|
||||
/// Some firmware images are currently packaged as ELF files, where sections names are used as keys
|
||||
/// to specific and related bits of data. Future firmware versions are scheduled to move away from
|
||||
/// that scheme before nova-core becomes stable, which means this module will eventually be
|
||||
/// removed.
|
||||
mod elf {
|
||||
use core::mem::size_of;
|
||||
|
||||
use kernel::{
|
||||
bindings,
|
||||
str::CStr,
|
||||
transmute::FromBytes, //
|
||||
};
|
||||
|
||||
/// Trait to abstract over ELF header differences.
|
||||
trait ElfHeader: FromBytes {
|
||||
fn shnum(&self) -> u16;
|
||||
fn shoff(&self) -> u64;
|
||||
fn shstrndx(&self) -> u16;
|
||||
}
|
||||
|
||||
/// Trait to abstract over ELF section-header differences.
|
||||
trait ElfSectionHeader: FromBytes {
|
||||
fn name(&self) -> u32;
|
||||
fn offset(&self) -> u64;
|
||||
fn size(&self) -> u64;
|
||||
}
|
||||
|
||||
/// Trait describing a matching ELF header and section-header format.
|
||||
trait ElfFormat {
|
||||
type Header: ElfHeader;
|
||||
type SectionHeader: ElfSectionHeader;
|
||||
}
|
||||
|
||||
/// Newtype to provide a [`FromBytes`] implementation.
|
||||
#[repr(transparent)]
|
||||
struct Elf64Hdr(bindings::elf64_hdr);
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for Elf64Hdr {}
|
||||
|
||||
impl ElfHeader for Elf64Hdr {
|
||||
fn shnum(&self) -> u16 {
|
||||
self.0.e_shnum
|
||||
}
|
||||
|
||||
fn shoff(&self) -> u64 {
|
||||
self.0.e_shoff
|
||||
}
|
||||
|
||||
fn shstrndx(&self) -> u16 {
|
||||
self.0.e_shstrndx
|
||||
}
|
||||
}
|
||||
|
||||
#[repr(transparent)]
|
||||
struct Elf64SHdr(bindings::elf64_shdr);
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for Elf64SHdr {}
|
||||
|
||||
impl ElfSectionHeader for Elf64SHdr {
|
||||
fn name(&self) -> u32 {
|
||||
self.0.sh_name
|
||||
}
|
||||
|
||||
fn offset(&self) -> u64 {
|
||||
self.0.sh_offset
|
||||
}
|
||||
|
||||
fn size(&self) -> u64 {
|
||||
self.0.sh_size
|
||||
}
|
||||
}
|
||||
|
||||
struct Elf64Format;
|
||||
|
||||
impl ElfFormat for Elf64Format {
|
||||
type Header = Elf64Hdr;
|
||||
type SectionHeader = Elf64SHdr;
|
||||
}
|
||||
|
||||
/// Newtype to provide [`FromBytes`] and [`ElfHeader`] implementations for ELF32.
|
||||
#[repr(transparent)]
|
||||
struct Elf32Hdr(bindings::elf32_hdr);
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for Elf32Hdr {}
|
||||
|
||||
impl ElfHeader for Elf32Hdr {
|
||||
fn shnum(&self) -> u16 {
|
||||
self.0.e_shnum
|
||||
}
|
||||
|
||||
fn shoff(&self) -> u64 {
|
||||
u64::from(self.0.e_shoff)
|
||||
}
|
||||
|
||||
fn shstrndx(&self) -> u16 {
|
||||
self.0.e_shstrndx
|
||||
}
|
||||
}
|
||||
|
||||
/// Newtype to provide [`FromBytes`] and [`ElfSectionHeader`] implementations for ELF32.
|
||||
#[repr(transparent)]
|
||||
struct Elf32SHdr(bindings::elf32_shdr);
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for Elf32SHdr {}
|
||||
|
||||
impl ElfSectionHeader for Elf32SHdr {
|
||||
fn name(&self) -> u32 {
|
||||
self.0.sh_name
|
||||
}
|
||||
|
||||
fn offset(&self) -> u64 {
|
||||
u64::from(self.0.sh_offset)
|
||||
}
|
||||
|
||||
fn size(&self) -> u64 {
|
||||
u64::from(self.0.sh_size)
|
||||
}
|
||||
}
|
||||
|
||||
struct Elf32Format;
|
||||
|
||||
impl ElfFormat for Elf32Format {
|
||||
type Header = Elf32Hdr;
|
||||
type SectionHeader = Elf32SHdr;
|
||||
}
|
||||
|
||||
/// Returns a NULL-terminated string from the ELF image at `offset`.
|
||||
fn elf_str(elf: &[u8], offset: u64) -> Option<&str> {
|
||||
let idx = usize::try_from(offset).ok()?;
|
||||
let bytes = elf.get(idx..)?;
|
||||
CStr::from_bytes_until_nul(bytes).ok()?.to_str().ok()
|
||||
}
|
||||
|
||||
fn elf_section_generic<'a, F>(elf: &'a [u8], name: &str) -> Option<&'a [u8]>
|
||||
where
|
||||
F: ElfFormat,
|
||||
{
|
||||
let hdr = F::Header::from_bytes(elf.get(0..size_of::<F::Header>())?)?;
|
||||
|
||||
let shdr_num = usize::from(hdr.shnum());
|
||||
let shdr_start = usize::try_from(hdr.shoff()).ok()?;
|
||||
let shdr_end = shdr_num
|
||||
.checked_mul(size_of::<F::SectionHeader>())
|
||||
.and_then(|v| v.checked_add(shdr_start))?;
|
||||
|
||||
// Get all the section headers as an iterator over byte chunks.
|
||||
let shdr_bytes = elf.get(shdr_start..shdr_end)?;
|
||||
let mut shdr_iter = shdr_bytes.chunks_exact(size_of::<F::SectionHeader>());
|
||||
|
||||
// Get the strings table.
|
||||
let strhdr = shdr_iter
|
||||
.clone()
|
||||
.nth(usize::from(hdr.shstrndx()))
|
||||
.and_then(F::SectionHeader::from_bytes)?;
|
||||
|
||||
// Find the section which name matches `name` and return it.
|
||||
shdr_iter.find_map(|sh_bytes| {
|
||||
let sh = F::SectionHeader::from_bytes(sh_bytes)?;
|
||||
let name_offset = strhdr.offset().checked_add(u64::from(sh.name()))?;
|
||||
let section_name = elf_str(elf, name_offset)?;
|
||||
|
||||
if section_name != name {
|
||||
return None;
|
||||
}
|
||||
|
||||
let start = usize::try_from(sh.offset()).ok()?;
|
||||
let end = usize::try_from(sh.size())
|
||||
.ok()
|
||||
.and_then(|sz| start.checked_add(sz))?;
|
||||
|
||||
elf.get(start..end)
|
||||
})
|
||||
}
|
||||
|
||||
/// Extract the section with name `name` from the ELF64 image `elf`.
|
||||
fn elf64_section<'a>(elf: &'a [u8], name: &str) -> Option<&'a [u8]> {
|
||||
elf_section_generic::<Elf64Format>(elf, name)
|
||||
}
|
||||
|
||||
/// Extract the section with name `name` from the ELF32 image `elf`.
|
||||
fn elf32_section<'a>(elf: &'a [u8], name: &str) -> Option<&'a [u8]> {
|
||||
elf_section_generic::<Elf32Format>(elf, name)
|
||||
}
|
||||
|
||||
/// Automatically detects ELF32 vs ELF64 based on the ELF header.
|
||||
pub(super) fn elf_section<'a>(elf: &'a [u8], name: &str) -> Option<&'a [u8]> {
|
||||
// ELF identification: a 4-byte magic followed by a class byte (32- vs 64-bit).
|
||||
const ELFMAG: &[u8] = b"\x7fELF";
|
||||
const SELFMAG: usize = ELFMAG.len();
|
||||
const EI_CLASS: usize = 4;
|
||||
const ELFCLASS32: u8 = 1;
|
||||
const ELFCLASS64: u8 = 2;
|
||||
|
||||
if elf.get(0..SELFMAG) != Some(ELFMAG) {
|
||||
return None;
|
||||
}
|
||||
|
||||
match *elf.get(EI_CLASS)? {
|
||||
ELFCLASS32 => elf32_section(elf, name),
|
||||
ELFCLASS64 => elf64_section(elf, name),
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -10,12 +10,10 @@
|
|||
use kernel::{
|
||||
device,
|
||||
dma::Coherent,
|
||||
prelude::*,
|
||||
transmute::FromBytes, //
|
||||
prelude::*, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
sec2::Sec2,
|
||||
Falcon,
|
||||
|
|
@ -25,224 +23,19 @@
|
|||
FalconFirmware, //
|
||||
},
|
||||
firmware::{
|
||||
BinFirmware,
|
||||
tlv::{
|
||||
request_tlv, //
|
||||
Tlv,
|
||||
},
|
||||
FirmwareObject,
|
||||
FirmwareSignature,
|
||||
Signed,
|
||||
Unsigned, //
|
||||
},
|
||||
gpu::Chipset,
|
||||
num::{
|
||||
FromSafeCast,
|
||||
IntoSafeCast, //
|
||||
},
|
||||
num::IntoSafeCast,
|
||||
};
|
||||
|
||||
/// Local convenience function to return a copy of `S` by reinterpreting the bytes starting at
|
||||
/// `offset` in `slice`.
|
||||
fn frombytes_at<S: FromBytes + Sized>(slice: &[u8], offset: usize) -> Result<S> {
|
||||
let end = offset.checked_add(size_of::<S>()).ok_or(EINVAL)?;
|
||||
slice
|
||||
.get(offset..end)
|
||||
.and_then(S::from_bytes_copy)
|
||||
.ok_or(EINVAL)
|
||||
}
|
||||
|
||||
/// Heavy-Secured firmware header.
|
||||
///
|
||||
/// Such firmwares have an application-specific payload that needs to be patched with a given
|
||||
/// signature.
|
||||
#[repr(C)]
|
||||
#[derive(Debug, Clone)]
|
||||
struct HsHeaderV2 {
|
||||
/// Offset to the start of the signatures.
|
||||
sig_prod_offset: u32,
|
||||
/// Size in bytes of the signatures.
|
||||
sig_prod_size: u32,
|
||||
/// Offset to a `u32` containing the location at which to patch the signature in the microcode
|
||||
/// image.
|
||||
patch_loc_offset: u32,
|
||||
/// Offset to a `u32` containing the index of the signature to patch.
|
||||
patch_sig_offset: u32,
|
||||
/// Start offset to the signature metadata.
|
||||
meta_data_offset: u32,
|
||||
/// Size in bytes of the signature metadata.
|
||||
meta_data_size: u32,
|
||||
/// Offset to a `u32` containing the number of signatures in the signatures section.
|
||||
num_sig_offset: u32,
|
||||
/// Offset of the application-specific header.
|
||||
header_offset: u32,
|
||||
/// Size in bytes of the application-specific header.
|
||||
header_size: u32,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for HsHeaderV2 {}
|
||||
|
||||
/// Heavy-Secured Firmware image container.
|
||||
///
|
||||
/// This provides convenient access to the fields of [`HsHeaderV2`] that are actually indices to
|
||||
/// read from in the firmware data.
|
||||
struct HsFirmwareV2<'a> {
|
||||
hdr: HsHeaderV2,
|
||||
fw: &'a [u8],
|
||||
}
|
||||
|
||||
impl<'a> HsFirmwareV2<'a> {
|
||||
/// Interprets the header of `bin_fw` as a [`HsHeaderV2`] and returns an instance of
|
||||
/// `HsFirmwareV2` for further parsing.
|
||||
///
|
||||
/// Fails if the header pointed at by `bin_fw` is not within the bounds of the firmware image.
|
||||
fn new(bin_fw: &BinFirmware<'a>) -> Result<Self> {
|
||||
frombytes_at::<HsHeaderV2>(bin_fw.fw, bin_fw.hdr.header_offset.into_safe_cast())
|
||||
.map(|hdr| Self { hdr, fw: bin_fw.fw })
|
||||
}
|
||||
|
||||
/// Returns the location at which the signatures should be patched in the microcode image.
|
||||
///
|
||||
/// Fails if the offset of the patch location is outside the bounds of the firmware
|
||||
/// image.
|
||||
fn patch_location(&self) -> Result<u32> {
|
||||
frombytes_at::<u32>(self.fw, self.hdr.patch_loc_offset.into_safe_cast())
|
||||
}
|
||||
|
||||
/// Returns an iterator to the signatures of the firmware. The iterator can be empty if the
|
||||
/// firmware is unsigned.
|
||||
///
|
||||
/// Fails if the pointed signatures are outside the bounds of the firmware image.
|
||||
fn signatures_iter(&'a self) -> Result<impl Iterator<Item = BooterSignature<'a>>> {
|
||||
let num_sig = frombytes_at::<u32>(self.fw, self.hdr.num_sig_offset.into_safe_cast())?;
|
||||
let iter = match self.hdr.sig_prod_size.checked_div(num_sig) {
|
||||
// If there are no signatures, return an iterator that will yield zero elements.
|
||||
None => (&[] as &[u8]).chunks_exact(1),
|
||||
Some(sig_size) => {
|
||||
let patch_sig =
|
||||
frombytes_at::<u32>(self.fw, self.hdr.patch_sig_offset.into_safe_cast())?;
|
||||
|
||||
let signatures_start = self
|
||||
.hdr
|
||||
.sig_prod_offset
|
||||
.checked_add(patch_sig)
|
||||
.map(usize::from_safe_cast)
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
let signatures_end = signatures_start
|
||||
.checked_add(usize::from_safe_cast(self.hdr.sig_prod_size))
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
self.fw
|
||||
// Get signatures range.
|
||||
.get(signatures_start..signatures_end)
|
||||
.ok_or(EINVAL)?
|
||||
.chunks_exact(sig_size.into_safe_cast())
|
||||
}
|
||||
};
|
||||
|
||||
// Map the byte slices into signatures.
|
||||
Ok(iter.map(BooterSignature))
|
||||
}
|
||||
}
|
||||
|
||||
/// Signature parameters, as defined in the firmware.
|
||||
#[repr(C)]
|
||||
struct HsSignatureParams {
|
||||
/// Fuse version to use.
|
||||
fuse_ver: u32,
|
||||
/// Mask of engine IDs this firmware applies to.
|
||||
engine_id_mask: u32,
|
||||
/// ID of the microcode.
|
||||
ucode_id: u32,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for HsSignatureParams {}
|
||||
|
||||
impl HsSignatureParams {
|
||||
/// Returns the signature parameters contained in `hs_fw`.
|
||||
///
|
||||
/// Fails if the meta data parameter of `hs_fw` is outside the bounds of the firmware image, or
|
||||
/// if its size doesn't match that of [`HsSignatureParams`].
|
||||
fn new(hs_fw: &HsFirmwareV2<'_>) -> Result<Self> {
|
||||
let start = usize::from_safe_cast(hs_fw.hdr.meta_data_offset);
|
||||
let end = start
|
||||
.checked_add(hs_fw.hdr.meta_data_size.into_safe_cast())
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
hs_fw
|
||||
.fw
|
||||
.get(start..end)
|
||||
.and_then(Self::from_bytes_copy)
|
||||
.ok_or(EINVAL)
|
||||
}
|
||||
}
|
||||
|
||||
/// Header for code and data load offsets.
|
||||
#[repr(C)]
|
||||
#[derive(Debug, Clone)]
|
||||
struct HsLoadHeaderV2 {
|
||||
// Offset at which the code starts.
|
||||
os_code_offset: u32,
|
||||
// Total size of the code, for all apps.
|
||||
os_code_size: u32,
|
||||
// Offset at which the data starts.
|
||||
os_data_offset: u32,
|
||||
// Size of the data.
|
||||
os_data_size: u32,
|
||||
// Number of apps following this header. Each app is described by a [`HsLoadHeaderV2App`].
|
||||
num_apps: u32,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for HsLoadHeaderV2 {}
|
||||
|
||||
impl HsLoadHeaderV2 {
|
||||
/// Returns the load header contained in `hs_fw`.
|
||||
///
|
||||
/// Fails if the header pointed at by `hs_fw` is not within the bounds of the firmware image.
|
||||
fn new(hs_fw: &HsFirmwareV2<'_>) -> Result<Self> {
|
||||
frombytes_at::<Self>(hs_fw.fw, hs_fw.hdr.header_offset.into_safe_cast())
|
||||
}
|
||||
}
|
||||
|
||||
/// Header for app code loader.
|
||||
#[repr(C)]
|
||||
#[derive(Debug, Clone)]
|
||||
struct HsLoadHeaderV2App {
|
||||
/// Offset at which to load the app code.
|
||||
offset: u32,
|
||||
/// Length in bytes of the app code.
|
||||
len: u32,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for HsLoadHeaderV2App {}
|
||||
|
||||
impl HsLoadHeaderV2App {
|
||||
/// Returns the [`HsLoadHeaderV2App`] for app `idx` of `hs_fw`.
|
||||
///
|
||||
/// Fails if `idx` is larger than the number of apps declared in `hs_fw`, or if the header is
|
||||
/// not within the bounds of the firmware image.
|
||||
fn new(hs_fw: &HsFirmwareV2<'_>, idx: u32) -> Result<Self> {
|
||||
let load_hdr = HsLoadHeaderV2::new(hs_fw)?;
|
||||
if idx >= load_hdr.num_apps {
|
||||
Err(EINVAL)
|
||||
} else {
|
||||
frombytes_at::<Self>(
|
||||
hs_fw.fw,
|
||||
usize::from_safe_cast(hs_fw.hdr.header_offset)
|
||||
// Skip the load header...
|
||||
.checked_add(size_of::<HsLoadHeaderV2>())
|
||||
// ... and jump to app header `idx`.
|
||||
.and_then(|offset| {
|
||||
offset
|
||||
.checked_add(usize::from_safe_cast(idx).checked_mul(size_of::<Self>())?)
|
||||
})
|
||||
.ok_or(EINVAL)?,
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Signature for Booter firmware. Their size is encoded into the header and not known a compile
|
||||
/// time, so we just wrap a byte slices on which we can implement [`FirmwareSignature`].
|
||||
struct BooterSignature<'a>(&'a [u8]);
|
||||
|
|
@ -292,89 +85,76 @@ pub(crate) fn new(
|
|||
dev: &device::Device<device::Bound>,
|
||||
kind: BooterKind,
|
||||
chipset: Chipset,
|
||||
ver: &str,
|
||||
falcon: &Falcon<<Self as FalconFirmware>::Target>,
|
||||
bar: Bar0<'_>,
|
||||
falcon: &Falcon<'_, <Self as FalconFirmware>::Target>,
|
||||
) -> Result<Self> {
|
||||
let fw_name = match kind {
|
||||
BooterKind::Loader => "booter_load",
|
||||
BooterKind::Unloader => "booter_unload",
|
||||
};
|
||||
let fw = super::request_firmware(dev, chipset, fw_name, ver)?;
|
||||
let bin_fw = BinFirmware::new(&fw)?;
|
||||
let fw = request_tlv(dev, chipset, fw_name)?;
|
||||
let tlv = Tlv::new(fw.data())?;
|
||||
dev_dbg!(
|
||||
dev,
|
||||
"loaded {} firmware v{}\n",
|
||||
fw_name,
|
||||
tlv.get_string(b"VERS")?
|
||||
);
|
||||
|
||||
// The binary firmware embeds a Heavy-Secured firmware.
|
||||
let hs_fw = HsFirmwareV2::new(&bin_fw)?;
|
||||
let os_data_offset = tlv.get_u32(b"DAOF")?;
|
||||
let os_data_size = tlv.get_u32(b"DASZ")?;
|
||||
let os_code_offset = tlv.get_u32(b"CDOF")?;
|
||||
let os_code_size = tlv.get_u32(b"CDSZ")?;
|
||||
let patch_loc = tlv.get_u32(b"PLOC")?;
|
||||
let fuse_version: usize = tlv.get_u32(b"FUSE")?.into_safe_cast();
|
||||
let engine_id = tlv.get_u32(b"ENID")?;
|
||||
let ucode_id = tlv.get_u32(b"UCID")?;
|
||||
let app0_code_offset = tlv.get_u32(b"A0CO")?;
|
||||
let app0_code_size = tlv.get_u32(b"A0CS")?;
|
||||
|
||||
// The Heavy-Secured firmware embeds a firmware load descriptor.
|
||||
let load_hdr = HsLoadHeaderV2::new(&hs_fw)?;
|
||||
|
||||
// Offset in `ucode` where to patch the signature.
|
||||
let patch_loc = hs_fw.patch_location()?;
|
||||
|
||||
let sig_params = HsSignatureParams::new(&hs_fw)?;
|
||||
let brom_params = FalconBromParams {
|
||||
// `load_hdr.os_data_offset` is an absolute index, but `pkc_data_offset` is from the
|
||||
// `os_data_offset` is an absolute index, but `pkc_data_offset` is from the
|
||||
// signature patch location.
|
||||
pkc_data_offset: patch_loc
|
||||
.checked_sub(load_hdr.os_data_offset)
|
||||
.ok_or(EINVAL)?,
|
||||
engine_id_mask: u16::try_from(sig_params.engine_id_mask).map_err(|_| EINVAL)?,
|
||||
ucode_id: u8::try_from(sig_params.ucode_id).map_err(|_| EINVAL)?,
|
||||
pkc_data_offset: patch_loc.checked_sub(os_data_offset).ok_or(EINVAL)?,
|
||||
engine_id_mask: u16::try_from(engine_id).map_err(|_| EINVAL)?,
|
||||
ucode_id: u8::try_from(ucode_id).map_err(|_| EINVAL)?,
|
||||
};
|
||||
let app0 = HsLoadHeaderV2App::new(&hs_fw, 0)?;
|
||||
|
||||
// Object containing the firmware microcode to be signature-patched.
|
||||
let ucode = bin_fw
|
||||
.data()
|
||||
.ok_or(EINVAL)
|
||||
let ucode = tlv
|
||||
.get_bytes(b"BLOB")
|
||||
.and_then(FirmwareObject::<Self, _>::new_booter)?;
|
||||
|
||||
let ucode_signed = {
|
||||
let mut signatures = hs_fw.signatures_iter()?.peekable();
|
||||
// Obtain the version from the fuse register, and extract the corresponding
|
||||
// signature.
|
||||
let reg_fuse_version: usize = falcon
|
||||
.signature_reg_fuse_version(brom_params.engine_id_mask, brom_params.ucode_id)?
|
||||
.into_safe_cast();
|
||||
|
||||
if signatures.peek().is_none() {
|
||||
// If there are no signatures, then the firmware is unsigned.
|
||||
ucode.no_patch_signature()
|
||||
} else {
|
||||
// Obtain the version from the fuse register, and extract the corresponding
|
||||
// signature.
|
||||
let reg_fuse_version = falcon.signature_reg_fuse_version(
|
||||
bar,
|
||||
brom_params.engine_id_mask,
|
||||
brom_params.ucode_id,
|
||||
)?;
|
||||
const FUSE_VERSION_USE_LAST_SIG: usize = 0;
|
||||
|
||||
// `0` means the last signature should be used.
|
||||
const FUSE_VERSION_USE_LAST_SIG: u32 = 0;
|
||||
let signature = match reg_fuse_version {
|
||||
FUSE_VERSION_USE_LAST_SIG => signatures.last(),
|
||||
// Otherwise hardware fuse version needs to be subtracted to obtain the index.
|
||||
reg_fuse_version => {
|
||||
let Some(idx) = sig_params.fuse_ver.checked_sub(reg_fuse_version) else {
|
||||
dev_err!(dev, "invalid fuse version for Booter firmware\n");
|
||||
return Err(EINVAL);
|
||||
};
|
||||
signatures.nth(idx.into_safe_cast())
|
||||
}
|
||||
}
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
ucode.patch_signature(&signature, patch_loc.into_safe_cast())?
|
||||
}
|
||||
let index = match reg_fuse_version {
|
||||
// `0` means the last signature should be used.
|
||||
FUSE_VERSION_USE_LAST_SIG => None,
|
||||
// Otherwise, hardware fuse version needs to be subtracted to obtain the index.
|
||||
_ => Some(fuse_version.checked_sub(reg_fuse_version).ok_or(EINVAL)?),
|
||||
};
|
||||
|
||||
// Extract the nth signature. Booter is always signed.
|
||||
let sig_chunk = tlv.get_signature(index)?;
|
||||
|
||||
let signature = BooterSignature(sig_chunk);
|
||||
let ucode_signed = ucode.patch_signature(&signature, patch_loc.into_safe_cast())?;
|
||||
|
||||
// There are two versions of Booter, one for Turing/GA100, and another for
|
||||
// GA102+. The extraction of the IMEM sections differs between the two
|
||||
// versions. Unfortunately, the file names are the same, and the headers
|
||||
// don't indicate the versions. The only way to differentiate is by the Chipset.
|
||||
let (imem_sec_dst_start, imem_ns_load_target) = if chipset <= Chipset::GA100 {
|
||||
(
|
||||
app0.offset,
|
||||
app0_code_offset,
|
||||
Some(FalconDmaLoadTarget {
|
||||
src_start: 0,
|
||||
dst_start: load_hdr.os_code_offset,
|
||||
len: load_hdr.os_code_size,
|
||||
dst_start: os_code_offset,
|
||||
len: os_code_size,
|
||||
}),
|
||||
)
|
||||
} else {
|
||||
|
|
@ -383,15 +163,15 @@ pub(crate) fn new(
|
|||
|
||||
Ok(Self {
|
||||
imem_sec_load_target: FalconDmaLoadTarget {
|
||||
src_start: app0.offset,
|
||||
src_start: app0_code_offset,
|
||||
dst_start: imem_sec_dst_start,
|
||||
len: app0.len,
|
||||
len: app0_code_size,
|
||||
},
|
||||
imem_ns_load_target,
|
||||
dmem_load_target: FalconDmaLoadTarget {
|
||||
src_start: load_hdr.os_data_offset,
|
||||
src_start: os_data_offset,
|
||||
dst_start: 0,
|
||||
len: load_hdr.os_data_size,
|
||||
len: os_data_size,
|
||||
},
|
||||
brom_params,
|
||||
ucode: ucode_signed,
|
||||
|
|
@ -405,17 +185,15 @@ pub(crate) fn new(
|
|||
pub(crate) fn run<T>(
|
||||
&self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
sec2_falcon: &Falcon<Sec2>,
|
||||
sec2_falcon: &Falcon<'_, Sec2>,
|
||||
wpr_meta: &Coherent<T>,
|
||||
) -> Result {
|
||||
sec2_falcon.reset(bar)?;
|
||||
sec2_falcon.load(dev, bar, self)?;
|
||||
let wpr_handle = wpr_meta.dma_handle();
|
||||
sec2_falcon.reset()?;
|
||||
sec2_falcon.load(self)?;
|
||||
let wpr_dma_address = wpr_meta.dma_address();
|
||||
let (mbox0, mbox1) = sec2_falcon.boot(
|
||||
bar,
|
||||
Some(wpr_handle as u32),
|
||||
Some((wpr_handle >> 32) as u32),
|
||||
Some(wpr_dma_address as u32),
|
||||
Some((wpr_dma_address >> 32) as u32),
|
||||
)?;
|
||||
dev_dbg!(dev, "SEC2 MBOX0: {:#x}, MBOX1: {:#x}\n", mbox0, mbox1);
|
||||
|
||||
|
|
|
|||
|
|
@ -6,12 +6,14 @@
|
|||
use kernel::{
|
||||
device,
|
||||
dma::Coherent,
|
||||
firmware::Firmware,
|
||||
prelude::*, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
firmware::elf,
|
||||
firmware::tlv::{
|
||||
request_tlv, //
|
||||
Tlv,
|
||||
},
|
||||
gpu::Chipset, //
|
||||
};
|
||||
|
||||
|
|
@ -19,11 +21,11 @@
|
|||
const FSP_HASH_SIZE: usize = 48;
|
||||
/// Maximum size of the FSP public key (RSA-3072), in bytes.
|
||||
///
|
||||
/// The FMC ELF `publickey` section may be shorter, so the remaining bytes are zero-padded.
|
||||
/// The FMC `PKEY` tag may be shorter, so the remaining bytes are zero-padded.
|
||||
const FSP_PKEY_SIZE: usize = 384;
|
||||
/// Maximum size of the FSP signature (RSA-3072), in bytes.
|
||||
///
|
||||
/// The FMC ELF `signature` section may be shorter, so the remaining bytes are zero-padded.
|
||||
/// The FMC `SIGN` tag may be shorter, so the remaining bytes are zero-padded.
|
||||
const FSP_SIG_SIZE: usize = 384;
|
||||
|
||||
/// Structure to hold FMC signatures.
|
||||
|
|
@ -38,64 +40,35 @@ pub(crate) struct FmcSignatures {
|
|||
}
|
||||
|
||||
pub(crate) struct FspFirmware {
|
||||
/// FMC firmware image data (only the "image" ELF section).
|
||||
/// FMC firmware image data
|
||||
pub(crate) fmc_image: Coherent<[u8]>,
|
||||
/// FMC firmware signatures.
|
||||
pub(crate) fmc_sigs: KBox<FmcSignatures>,
|
||||
}
|
||||
|
||||
impl FspFirmware {
|
||||
pub(crate) fn new(
|
||||
dev: &device::Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
ver: &str,
|
||||
) -> Result<Self> {
|
||||
let fw = super::request_firmware(dev, chipset, "fmc", ver)?;
|
||||
pub(crate) fn new(dev: &device::Device<device::Bound>, chipset: Chipset) -> Result<Self> {
|
||||
let fw = request_tlv(dev, chipset, "fmc")?;
|
||||
let tlv = Tlv::new(fw.data())?;
|
||||
dev_dbg!(dev, "loaded fsp firmware v{}\n", tlv.get_string(b"VERS")?);
|
||||
|
||||
// FSP expects only the "image" section, not the entire ELF file.
|
||||
let fmc_image_data = elf::elf_section(fw.data(), "image").ok_or_else(|| {
|
||||
dev_err!(dev, "FMC ELF file missing 'image' section\n");
|
||||
EINVAL
|
||||
})?;
|
||||
let fmc_image_data = tlv.get_bytes(b"BLOB")?;
|
||||
let fmc_image = Coherent::from_slice(dev, fmc_image_data, GFP_KERNEL)?;
|
||||
|
||||
Ok(Self {
|
||||
fmc_image,
|
||||
fmc_sigs: Self::extract_fmc_signatures(&fw, dev)?,
|
||||
fmc_sigs: Self::extract_fmc_signatures(&tlv, dev)?,
|
||||
})
|
||||
}
|
||||
|
||||
/// Extract FMC firmware signatures for Chain of Trust verification.
|
||||
///
|
||||
/// Extracts real cryptographic signatures from FMC ELF32 firmware sections.
|
||||
/// Extracts real cryptographic signatures from FMC TLV firmware tags.
|
||||
/// Returns signatures in a heap-allocated structure to prevent stack overflow.
|
||||
fn extract_fmc_signatures(
|
||||
fmc_fw: &Firmware,
|
||||
dev: &device::Device,
|
||||
) -> Result<KBox<FmcSignatures>> {
|
||||
let get_section = |name: &str, max_len: usize| {
|
||||
elf::elf_section(fmc_fw.data(), name)
|
||||
.ok_or(EINVAL)
|
||||
.inspect_err(|_| dev_err!(dev, "FMC firmware missing '{}' section\n", name))
|
||||
.and_then(|section| {
|
||||
if section.len() > max_len {
|
||||
dev_err!(
|
||||
dev,
|
||||
"FMC {} section size {} > maximum {}\n",
|
||||
name,
|
||||
section.len(),
|
||||
max_len
|
||||
);
|
||||
Err(EINVAL)
|
||||
} else {
|
||||
Ok(section)
|
||||
}
|
||||
})
|
||||
};
|
||||
|
||||
let hash_section = get_section("hash", FSP_HASH_SIZE)?;
|
||||
let pkey_section = get_section("publickey", FSP_PKEY_SIZE)?;
|
||||
let sig_section = get_section("signature", FSP_SIG_SIZE)?;
|
||||
fn extract_fmc_signatures(tlv: &Tlv<'_>, dev: &device::Device) -> Result<KBox<FmcSignatures>> {
|
||||
let hash_section = tlv.get_bytes(b"HASH")?;
|
||||
let pkey_section = tlv.get_bytes(b"PKEY")?;
|
||||
let sig_section = tlv.get_bytes(b"SIGN")?;
|
||||
|
||||
// The hash section is a SHA-384 output: it must be exactly FSP_HASH_SIZE bytes.
|
||||
if hash_section.len() != FSP_HASH_SIZE {
|
||||
|
|
@ -108,15 +81,36 @@ fn extract_fmc_signatures(
|
|||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// The key and signature sections are zero-padded to a fixed maximum, so they may be
|
||||
// shorter, but must not exceed the destination buffers.
|
||||
if pkey_section.len() > FSP_PKEY_SIZE {
|
||||
dev_err!(
|
||||
dev,
|
||||
"FMC public key section size {} > maximum {}\n",
|
||||
pkey_section.len(),
|
||||
FSP_PKEY_SIZE
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
if sig_section.len() > FSP_SIG_SIZE {
|
||||
dev_err!(
|
||||
dev,
|
||||
"FMC signature section size {} > maximum {}\n",
|
||||
sig_section.len(),
|
||||
FSP_SIG_SIZE
|
||||
);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// Initialize the signatures in place to avoid building the large `FmcSignatures` on the
|
||||
// stack, then fill each section from the firmware.
|
||||
let signatures = KBox::init(
|
||||
pin_init::init_zeroed::<FmcSignatures>().chain(|sigs| {
|
||||
// PANIC: src and dst lengths are both FSP_HASH_SIZE (verified above).
|
||||
sigs.hash384.copy_from_slice(hash_section);
|
||||
// PANIC: dst is sliced to src.len(); src.len() <= FSP_PKEY_SIZE per `get_section`.
|
||||
// PANIC: dst is sliced to src.len(); src.len() <= FSP_PKEY_SIZE (verified above).
|
||||
sigs.public_key[..pkey_section.len()].copy_from_slice(pkey_section);
|
||||
// PANIC: dst is sliced to src.len(); src.len() <= FSP_SIG_SIZE per `get_section`.
|
||||
// PANIC: dst is sliced to src.len(); src.len() <= FSP_SIG_SIZE (verified above).
|
||||
sigs.signature[..sig_section.len()].copy_from_slice(sig_section);
|
||||
Ok(())
|
||||
}),
|
||||
|
|
|
|||
|
|
@ -27,7 +27,6 @@
|
|||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
gsp::Gsp,
|
||||
Falcon,
|
||||
|
|
@ -320,8 +319,7 @@ impl FwsecFirmware {
|
|||
/// command.
|
||||
pub(crate) fn new(
|
||||
dev: &Device<device::Bound>,
|
||||
falcon: &Falcon<Gsp>,
|
||||
bar: Bar0<'_>,
|
||||
falcon: &Falcon<'_, Gsp>,
|
||||
bios: &Vbios,
|
||||
cmd: FwsecCommand,
|
||||
) -> Result<Self> {
|
||||
|
|
@ -337,7 +335,7 @@ pub(crate) fn new(
|
|||
.ok_or(EINVAL)?;
|
||||
let desc_sig_versions = u32::from(desc.signature_versions());
|
||||
let reg_fuse_version =
|
||||
falcon.signature_reg_fuse_version(bar, desc.engine_id_mask(), desc.ucode_id())?;
|
||||
falcon.signature_reg_fuse_version(desc.engine_id_mask(), desc.ucode_id())?;
|
||||
dev_dbg!(
|
||||
dev,
|
||||
"desc_sig_versions: {:#x}, reg_fuse_version: {}\n",
|
||||
|
|
@ -387,24 +385,18 @@ pub(crate) fn new(
|
|||
|
||||
/// Loads the FWSEC firmware into `falcon` and execute it.
|
||||
///
|
||||
/// This must only be called on chipsets that do not need the FWSEC bootloader (i.e., where
|
||||
/// [`Chipset::needs_fwsec_bootloader()`](crate::gpu::Chipset::needs_fwsec_bootloader) returns
|
||||
/// `false`). On chipsets that do, use [`bootloader::FwsecFirmwareWithBl`] instead.
|
||||
pub(crate) fn run(
|
||||
&self,
|
||||
dev: &Device<device::Bound>,
|
||||
falcon: &Falcon<Gsp>,
|
||||
bar: Bar0<'_>,
|
||||
) -> Result<()> {
|
||||
/// This must only be called on chipsets that do not need the FWSEC bootloader. On chipsets
|
||||
/// where the bootloader is required, use [`bootloader::FwsecFirmwareWithBl`] instead.
|
||||
pub(crate) fn run(&self, dev: &Device<device::Bound>, falcon: &Falcon<'_, Gsp>) -> Result<()> {
|
||||
// Reset falcon, load the firmware, and run it.
|
||||
falcon
|
||||
.reset(bar)
|
||||
.reset()
|
||||
.inspect_err(|e| dev_err!(dev, "Failed to reset GSP falcon: {:?}\n", e))?;
|
||||
falcon
|
||||
.load(dev, bar, self)
|
||||
.load(self)
|
||||
.inspect_err(|e| dev_err!(dev, "Failed to load FWSEC firmware: {:?}\n", e))?;
|
||||
let (mbox0, _) = falcon
|
||||
.boot(bar, Some(0), None)
|
||||
.boot(Some(0), None)
|
||||
.inspect_err(|e| dev_err!(dev, "Failed to boot FWSEC firmware: {:?}\n", e))?;
|
||||
if mbox0 != 0 {
|
||||
dev_err!(dev, "FWSEC firmware returned error {}\n", mbox0);
|
||||
|
|
|
|||
|
|
@ -7,26 +7,19 @@
|
|||
//! be loaded using PIO.
|
||||
|
||||
use kernel::{
|
||||
alloc::KVec,
|
||||
device::{
|
||||
self,
|
||||
Device, //
|
||||
},
|
||||
dma::Coherent,
|
||||
io::{
|
||||
register::WithBase, //
|
||||
Io,
|
||||
},
|
||||
io::{register::WithBase, Io},
|
||||
prelude::*,
|
||||
ptr::{
|
||||
Alignable,
|
||||
Alignment, //
|
||||
},
|
||||
sizes,
|
||||
transmute::{
|
||||
AsBytes,
|
||||
FromBytes, //
|
||||
},
|
||||
transmute::AsBytes,
|
||||
};
|
||||
|
||||
use crate::{
|
||||
|
|
@ -46,38 +39,16 @@
|
|||
},
|
||||
firmware::{
|
||||
fwsec::FwsecFirmware,
|
||||
request_firmware,
|
||||
BinHdr,
|
||||
FIRMWARE_VERSION, //
|
||||
tlv::{
|
||||
request_tlv, //
|
||||
Tlv,
|
||||
},
|
||||
},
|
||||
gpu::Chipset,
|
||||
num::FromSafeCast,
|
||||
num::FromSafeCast, //
|
||||
regs,
|
||||
};
|
||||
|
||||
/// Descriptor used by RM to figure out the requirements of the boot loader.
|
||||
///
|
||||
/// Most of its fields appear to be legacy and carry incorrect values, so they are left unused.
|
||||
#[repr(C)]
|
||||
#[derive(Debug, Clone)]
|
||||
struct BootloaderDesc {
|
||||
/// Starting tag of bootloader.
|
||||
start_tag: u32,
|
||||
/// DMEM load offset - unused here as we always load at offset `0`.
|
||||
_dmem_load_off: u32,
|
||||
/// Offset of code section in the image. Unused as there is only one section in the bootloader
|
||||
/// binary.
|
||||
_code_off: u32,
|
||||
/// Size of code section in the image.
|
||||
code_size: u32,
|
||||
/// Offset of data section in the image. Unused as we build the data section ourselves.
|
||||
_data_off: u32,
|
||||
/// Size of data section in the image. Unused as we build the data section ourselves.
|
||||
_data_size: u32,
|
||||
}
|
||||
// SAFETY: any byte sequence is valid for this struct.
|
||||
unsafe impl FromBytes for BootloaderDesc {}
|
||||
|
||||
/// Structure used by the boot-loader to load the rest of the code.
|
||||
///
|
||||
/// This has to be filled by the GPU driver and copied into DMEM at offset
|
||||
|
|
@ -150,38 +121,24 @@ pub(crate) fn new(
|
|||
dev: &Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
) -> Result<Self> {
|
||||
let fw = request_firmware(dev, chipset, "gen_bootloader", FIRMWARE_VERSION)?;
|
||||
let hdr = fw
|
||||
.data()
|
||||
.get(0..size_of::<BinHdr>())
|
||||
.and_then(BinHdr::from_bytes_copy)
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
let desc = {
|
||||
let desc_offset = usize::from_safe_cast(hdr.header_offset);
|
||||
|
||||
fw.data()
|
||||
.get(desc_offset..)
|
||||
.and_then(BootloaderDesc::from_bytes_copy_prefix)
|
||||
.ok_or(EINVAL)?
|
||||
.0
|
||||
};
|
||||
let fw = request_tlv(dev, chipset, "gen_bootloader")?;
|
||||
let tlv = Tlv::new(fw.data())?;
|
||||
dev_dbg!(
|
||||
dev,
|
||||
"loaded generic bootloader firmware v{}\n",
|
||||
tlv.get_string(b"VERS")?
|
||||
);
|
||||
|
||||
let ucode = {
|
||||
let ucode_start = usize::from_safe_cast(hdr.data_offset);
|
||||
let code_size = usize::from_safe_cast(desc.code_size);
|
||||
// Align to falcon block size (256 bytes).
|
||||
let blob = tlv.get_bytes(b"BLOB")?;
|
||||
let code_size = usize::from_safe_cast(tlv.get_u32(b"CDSZ")?);
|
||||
let code = blob.get(..code_size).ok_or(EINVAL)?;
|
||||
let aligned_code_size = code_size
|
||||
.align_up(Alignment::new::<{ falcon::MEM_BLOCK_ALIGNMENT }>())
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
let mut ucode = KVec::with_capacity(aligned_code_size, GFP_KERNEL)?;
|
||||
ucode.extend_from_slice(
|
||||
fw.data()
|
||||
.get(ucode_start..ucode_start + code_size)
|
||||
.ok_or(EINVAL)?,
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
ucode.extend_from_slice(code, GFP_KERNEL)?;
|
||||
ucode.resize(aligned_code_size, 0, GFP_KERNEL)?;
|
||||
|
||||
ucode
|
||||
|
|
@ -234,7 +191,7 @@ pub(crate) fn new(
|
|||
reserved: [0; 4],
|
||||
signature: [0; 4],
|
||||
ctx_dma: FALCON_DMAIDX_PHYS_SYS_NCOH,
|
||||
code_dma_base: firmware_dma.dma_handle(),
|
||||
code_dma_base: firmware_dma.dma_address(),
|
||||
// `dst_start` is also valid as the source offset since the firmware DMA object is
|
||||
// a mirror image of the target IMEM layout.
|
||||
non_sec_code_off: imem_ns.dst_start,
|
||||
|
|
@ -246,7 +203,7 @@ pub(crate) fn new(
|
|||
code_entry_point: 0,
|
||||
// Start of data section is the added padding + the DMEM `src_start` field.
|
||||
data_dma_base: firmware_dma
|
||||
.dma_handle()
|
||||
.dma_address()
|
||||
.checked_add(u64::from_safe_cast(align_padding))
|
||||
.and_then(|offset| offset.checked_add(dmem.src_start.into()))
|
||||
.ok_or(EOVERFLOW)?,
|
||||
|
|
@ -262,13 +219,15 @@ pub(crate) fn new(
|
|||
.checked_sub(ucode.len())
|
||||
.ok_or(EOVERFLOW)?;
|
||||
|
||||
let start_tag = u16::try_from(tlv.get_u32(b"STRT")?)?;
|
||||
|
||||
Ok(Self {
|
||||
_firmware_dma: firmware_dma,
|
||||
ucode,
|
||||
dmem_desc,
|
||||
brom_params: firmware.brom_params(),
|
||||
imem_dst_start: u16::try_from(imem_dst_start)?,
|
||||
start_tag: u16::try_from(desc.start_tag)?,
|
||||
start_tag,
|
||||
})
|
||||
}
|
||||
|
||||
|
|
@ -279,15 +238,15 @@ pub(crate) fn new(
|
|||
pub(crate) fn run(
|
||||
&self,
|
||||
dev: &Device<device::Bound>,
|
||||
falcon: &Falcon<Gsp>,
|
||||
falcon: &Falcon<'_, Gsp>,
|
||||
bar: Bar0<'_>,
|
||||
) -> Result<()> {
|
||||
// Reset falcon, load the firmware, and run it.
|
||||
falcon
|
||||
.reset(bar)
|
||||
.reset()
|
||||
.inspect_err(|e| dev_err!(dev, "Failed to reset GSP falcon: {:?}\n", e))?;
|
||||
falcon
|
||||
.pio_load(bar, self)
|
||||
.pio_load(self)
|
||||
.inspect_err(|e| dev_err!(dev, "Failed to load FWSEC firmware: {:?}\n", e))?;
|
||||
|
||||
// Configure DMA index for the bootloader to fetch the FWSEC firmware from system memory.
|
||||
|
|
@ -302,7 +261,7 @@ pub(crate) fn run(
|
|||
);
|
||||
|
||||
let (mbox0, _) = falcon
|
||||
.boot(bar, Some(0), None)
|
||||
.boot(Some(0), None)
|
||||
.inspect_err(|e| dev_err!(dev, "Failed to boot FWSEC firmware: {:?}\n", e))?;
|
||||
if mbox0 != 0 {
|
||||
dev_err!(dev, "FWSEC firmware returned error {}\n", mbox0);
|
||||
|
|
|
|||
|
|
@ -8,22 +8,24 @@
|
|||
DataDirection,
|
||||
DmaAddress, //
|
||||
},
|
||||
firmware,
|
||||
prelude::*,
|
||||
scatterlist::{
|
||||
Owned,
|
||||
SGTable, //
|
||||
},
|
||||
str::CString,
|
||||
};
|
||||
|
||||
use crate::{
|
||||
firmware::{
|
||||
elf,
|
||||
riscv::RiscvFirmware, //
|
||||
tlv::{
|
||||
request_tlv, //
|
||||
Tlv,
|
||||
},
|
||||
},
|
||||
gpu::{
|
||||
Architecture,
|
||||
Chipset, //
|
||||
},
|
||||
gpu::Chipset,
|
||||
gsp::GSP_PAGE_SIZE,
|
||||
num::FromSafeCast,
|
||||
};
|
||||
|
|
@ -63,43 +65,26 @@ pub(crate) struct GspFirmware {
|
|||
}
|
||||
|
||||
impl GspFirmware {
|
||||
fn find_gsp_sigs_section(chipset: Chipset) -> &'static str {
|
||||
match chipset.arch() {
|
||||
Architecture::Turing if matches!(chipset, Chipset::TU116 | Chipset::TU117) => {
|
||||
".fwsignature_tu11x"
|
||||
}
|
||||
Architecture::Turing => ".fwsignature_tu10x",
|
||||
Architecture::Ampere if chipset == Chipset::GA100 => ".fwsignature_ga100",
|
||||
Architecture::Ampere => ".fwsignature_ga10x",
|
||||
Architecture::Ada => ".fwsignature_ad10x",
|
||||
Architecture::Hopper => ".fwsignature_gh10x",
|
||||
Architecture::BlackwellGB10x => ".fwsignature_gb10x",
|
||||
Architecture::BlackwellGB20x => ".fwsignature_gb20x",
|
||||
}
|
||||
}
|
||||
|
||||
/// Loads the GSP firmware binaries, map them into `dev`'s address-space, and creates the page
|
||||
/// tables expected by the GSP bootloader to load it.
|
||||
pub(crate) fn new<'a>(
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
ver: &'a str,
|
||||
) -> impl PinInit<Self, Error> + 'a {
|
||||
pin_init::pin_init_scope(move || {
|
||||
let firmware = super::request_firmware(dev, chipset, "gsp", ver)?;
|
||||
let firmware = request_tlv(dev, chipset, "gsp")?;
|
||||
let tlv = Tlv::new(firmware.data())?;
|
||||
dev_dbg!(dev, "loaded gsp firmware v{}\n", tlv.get_string(b"VERS")?);
|
||||
|
||||
let fw_section = elf::elf_section(firmware.data(), ".fwimage").ok_or(EINVAL)?;
|
||||
let size = usize::from_safe_cast(tlv.get_u32(b"SIZE")?);
|
||||
let mut fw_vvec = VVec::zeroed(size, GFP_KERNEL).map_err(|_| ENOMEM)?;
|
||||
|
||||
let size = fw_section.len();
|
||||
let chip_name = chipset.name();
|
||||
let file = tlv.get_string(b"FILE")?;
|
||||
let filename = CString::try_from_fmt(fmt!("nvidia/{chip_name}/gsp/{file}"))?;
|
||||
firmware::request_into_buf(&filename, dev, fw_vvec.as_mut_slice())?;
|
||||
|
||||
// Move the firmware into a vmalloc'd vector and map it into the device address
|
||||
// space.
|
||||
let fw_vvec = VVec::with_capacity(fw_section.len(), GFP_KERNEL)
|
||||
.and_then(|mut v| {
|
||||
v.extend_from_slice(fw_section, GFP_KERNEL)?;
|
||||
Ok(v)
|
||||
})
|
||||
.map_err(|_| ENOMEM)?;
|
||||
let signatures = Coherent::from_slice(dev, tlv.get_bytes(b"SIGN")?, GFP_KERNEL)?;
|
||||
|
||||
Ok(try_pin_init!(Self {
|
||||
fw <- SGTable::new(dev, fw_vvec, DataDirection::ToDevice, GFP_KERNEL),
|
||||
|
|
@ -145,15 +130,9 @@ pub(crate) fn new<'a>(
|
|||
level0.into()
|
||||
},
|
||||
size,
|
||||
signatures: {
|
||||
let sigs_section = Self::find_gsp_sigs_section(chipset);
|
||||
|
||||
elf::elf_section(firmware.data(), sigs_section)
|
||||
.ok_or(EINVAL)
|
||||
.and_then(|data| Coherent::from_slice(dev, data, GFP_KERNEL))?
|
||||
},
|
||||
signatures,
|
||||
bootloader: {
|
||||
let bl = super::request_firmware(dev, chipset, "bootloader", ver)?;
|
||||
let bl = request_tlv(dev, chipset, "gsp_bootloader")?;
|
||||
|
||||
RiscvFirmware::new(dev, &bl)?
|
||||
},
|
||||
|
|
@ -161,9 +140,9 @@ pub(crate) fn new<'a>(
|
|||
})
|
||||
}
|
||||
|
||||
/// Returns the DMA handle of the radix3 level 0 page table.
|
||||
pub(crate) fn radix3_dma_handle(&self) -> DmaAddress {
|
||||
self.level0.dma_handle()
|
||||
/// Returns the DMA address of the radix3 level 0 page table.
|
||||
pub(crate) fn radix3_dma_address(&self) -> DmaAddress {
|
||||
self.level0.dma_address()
|
||||
}
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -7,53 +7,10 @@
|
|||
device,
|
||||
dma::Coherent,
|
||||
firmware::Firmware,
|
||||
prelude::*,
|
||||
transmute::FromBytes, //
|
||||
prelude::*, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
firmware::BinFirmware,
|
||||
num::FromSafeCast, //
|
||||
};
|
||||
|
||||
/// Descriptor for microcode running on a RISC-V core.
|
||||
#[repr(C)]
|
||||
#[derive(Debug)]
|
||||
struct RmRiscvUCodeDesc {
|
||||
version: u32,
|
||||
bootloader_offset: u32,
|
||||
bootloader_size: u32,
|
||||
bootloader_param_offset: u32,
|
||||
bootloader_param_size: u32,
|
||||
riscv_elf_offset: u32,
|
||||
riscv_elf_size: u32,
|
||||
app_version: u32,
|
||||
manifest_offset: u32,
|
||||
manifest_size: u32,
|
||||
monitor_data_offset: u32,
|
||||
monitor_data_size: u32,
|
||||
monitor_code_offset: u32,
|
||||
monitor_code_size: u32,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for this type, and it doesn't use interior mutability.
|
||||
unsafe impl FromBytes for RmRiscvUCodeDesc {}
|
||||
|
||||
impl RmRiscvUCodeDesc {
|
||||
/// Interprets the header of `bin_fw` as a [`RmRiscvUCodeDesc`] and returns it.
|
||||
///
|
||||
/// Fails if the header pointed at by `bin_fw` is not within the bounds of the firmware image.
|
||||
fn new(bin_fw: &BinFirmware<'_>) -> Result<Self> {
|
||||
let offset = usize::from_safe_cast(bin_fw.hdr.header_offset);
|
||||
let end = offset.checked_add(size_of::<Self>()).ok_or(EINVAL)?;
|
||||
|
||||
bin_fw
|
||||
.fw
|
||||
.get(offset..end)
|
||||
.and_then(Self::from_bytes_copy)
|
||||
.ok_or(EINVAL)
|
||||
}
|
||||
}
|
||||
use crate::firmware::tlv::Tlv;
|
||||
|
||||
/// A parsed firmware for a RISC-V core, ready to be loaded and run.
|
||||
pub(crate) struct RiscvFirmware {
|
||||
|
|
@ -72,24 +29,26 @@ pub(crate) struct RiscvFirmware {
|
|||
impl RiscvFirmware {
|
||||
/// Parses the RISC-V firmware image contained in `fw`.
|
||||
pub(crate) fn new(dev: &device::Device<device::Bound>, fw: &Firmware) -> Result<Self> {
|
||||
let bin_fw = BinFirmware::new(fw)?;
|
||||
let tlv = Tlv::new(fw.data())?;
|
||||
dev_dbg!(
|
||||
dev,
|
||||
"loaded gsp bootloader firmware v{}\n",
|
||||
tlv.get_string(b"VERS")?
|
||||
);
|
||||
|
||||
let riscv_desc = RmRiscvUCodeDesc::new(&bin_fw)?;
|
||||
let code_offset = tlv.get_u32(b"CDOF")?;
|
||||
let data_offset = tlv.get_u32(b"DAOF")?;
|
||||
let manifest_offset = tlv.get_u32(b"MFOF")?;
|
||||
let app_version = tlv.get_u32(b"APPV")?;
|
||||
|
||||
let ucode = {
|
||||
let start = usize::from_safe_cast(bin_fw.hdr.data_offset);
|
||||
let len = usize::from_safe_cast(bin_fw.hdr.data_size);
|
||||
let end = start.checked_add(len).ok_or(EINVAL)?;
|
||||
|
||||
Coherent::from_slice(dev, fw.data().get(start..end).ok_or(EINVAL)?, GFP_KERNEL)?
|
||||
};
|
||||
let ucode = Coherent::from_slice(dev, tlv.get_bytes(b"BLOB")?, GFP_KERNEL)?;
|
||||
|
||||
Ok(Self {
|
||||
ucode,
|
||||
code_offset: riscv_desc.monitor_code_offset,
|
||||
data_offset: riscv_desc.monitor_data_offset,
|
||||
manifest_offset: riscv_desc.manifest_offset,
|
||||
app_version: riscv_desc.app_version,
|
||||
code_offset,
|
||||
data_offset,
|
||||
manifest_offset,
|
||||
app_version,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
|
|
|||
262
drivers/gpu/nova-core/firmware/tlv.rs
Normal file
262
drivers/gpu/nova-core/firmware/tlv.rs
Normal file
|
|
@ -0,0 +1,262 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use kernel::{
|
||||
device,
|
||||
firmware,
|
||||
prelude::*,
|
||||
str::CString, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
gpu,
|
||||
num::*, //
|
||||
};
|
||||
|
||||
/// Requests the GPU firmware TLV `name` suitable for `chipset`.
|
||||
pub(crate) fn request_tlv(
|
||||
dev: &device::Device,
|
||||
chipset: gpu::Chipset,
|
||||
name: &str,
|
||||
) -> Result<firmware::Firmware> {
|
||||
let chip_name = chipset.name();
|
||||
|
||||
let filename = CString::try_from_fmt(fmt!("nvidia/{chip_name}/gsp/{name}.tlv"))?;
|
||||
|
||||
dev_dbg!(dev, "loading firmware image {:?}\n", &filename);
|
||||
|
||||
firmware::Firmware::request(&filename, dev)
|
||||
}
|
||||
|
||||
struct TlvBlock<'a> {
|
||||
tag: [u8; 4],
|
||||
value: &'a [u8],
|
||||
}
|
||||
|
||||
/// On-wire TLV block header: 4-byte ASCII tag + little-endian payload length (bytes, excluding
|
||||
/// padding to a 4-byte boundary).
|
||||
struct TlvBlockHeader {
|
||||
tag: [u8; 4],
|
||||
length: usize,
|
||||
}
|
||||
|
||||
impl TlvBlockHeader {
|
||||
const SIZE: usize = size_of::<[u8; 4]>() + size_of::<u32>();
|
||||
|
||||
/// Parses the first [`Self::SIZE`] bytes of `hdr` (caller may pass a longer slice).
|
||||
fn parse(hdr: &[u8]) -> Option<Self> {
|
||||
let hdr = hdr.get(..Self::SIZE)?;
|
||||
let tag = <[u8; 4]>::try_from(hdr.get(..4)?).ok()?;
|
||||
if !tag.is_ascii() {
|
||||
return None;
|
||||
}
|
||||
let len_arr = <[u8; 4]>::try_from(hdr.get(4..Self::SIZE)?).ok()?;
|
||||
let length = u32_as_usize(u32::from_le_bytes(len_arr));
|
||||
Some(Self { tag, length })
|
||||
}
|
||||
}
|
||||
|
||||
/// Iterator over the [`TlvBlock`]s of a [`Tlv`].
|
||||
///
|
||||
/// # Invariants
|
||||
///
|
||||
/// `pos` is a byte offset into `tlv.data` that always lies on a block boundary (in the sense
|
||||
/// of the [`Tlv`] invariant): it is either the start of a well-formed block, or equal to
|
||||
/// `tlv.data.len()` (end of iteration).
|
||||
struct TlvIter<'tlv, 'a> {
|
||||
tlv: &'tlv Tlv<'a>,
|
||||
pos: usize,
|
||||
}
|
||||
|
||||
impl<'tlv, 'a> Iterator for TlvIter<'tlv, 'a> {
|
||||
type Item = TlvBlock<'a>;
|
||||
|
||||
/// Returns the block starting at `self.pos` and advances the cursor past it, or [`None`]
|
||||
/// once the cursor reaches the end of the data or encounters an error.
|
||||
///
|
||||
/// Note that errors cannot actually occur because the data is validated in the constructor.
|
||||
fn next(&mut self) -> Option<Self::Item> {
|
||||
if self.pos >= self.tlv.data.len() {
|
||||
return None;
|
||||
}
|
||||
|
||||
let tail = self.tlv.data.get(self.pos..)?;
|
||||
|
||||
let hdr = tail.get(..TlvBlockHeader::SIZE)?;
|
||||
let header = TlvBlockHeader::parse(hdr)?;
|
||||
|
||||
let stored_size = header.length.checked_next_multiple_of(4)?;
|
||||
let advance = TlvBlockHeader::SIZE.checked_add(stored_size)?;
|
||||
let payload_end = TlvBlockHeader::SIZE.checked_add(header.length)?;
|
||||
|
||||
let value = tail
|
||||
.get(..advance)?
|
||||
.get(TlvBlockHeader::SIZE..payload_end)?;
|
||||
|
||||
// INVARIANT: by the `Tlv` invariant the block at `self.pos` occupies exactly `advance`
|
||||
// bytes, so `self.pos + advance` is the next block boundary (or `data.len()`).
|
||||
self.pos = self.pos.checked_add(advance)?;
|
||||
|
||||
Some(TlvBlock {
|
||||
tag: header.tag,
|
||||
value,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// The post-header part of a validated TLV (type, length, value) firmware image.
|
||||
///
|
||||
/// TLV firmware images start with a 4-byte "NVFW" magic header, followed by a sequence of
|
||||
/// blocks. Each block has a 4-byte type tag, a 4-byte length field, and a data payload
|
||||
/// (value) whose stored size is the length rounded up to the nearest multiple of 4.
|
||||
///
|
||||
/// [`Self::new`] checks the magic header and walks every block: tags must be ASCII,
|
||||
/// lengths and padding must fit without overflow, and the byte stream after `NVFW` must
|
||||
/// be exactly partitionable into blocks (no trailing partial header or slack). After
|
||||
/// that, [`TlvIter`] only signals end-of-stream via [`None`], not parse failure.
|
||||
///
|
||||
/// Although the spec forbids duplicate tags, neither the constructor nor the iterator
|
||||
/// enforces this restriction. Instead, duplicate tags are simply ignored.
|
||||
///
|
||||
/// # Invariants
|
||||
///
|
||||
/// `data` is a validated TLV payload (the bytes *after* the `NVFW` magic): it is the exact
|
||||
/// concatenation of zero or more well-formed blocks, with no trailing partial header or slack.
|
||||
/// Consequently, any offset `o` into `data` that is a block boundary and satisfies
|
||||
/// `o < data.len()` is the start of a complete block whose header parses and whose stored
|
||||
/// extent (`TlvBlockHeader::SIZE + header.length.next_multiple_of(4)` bytes) lies within
|
||||
/// `data`. `data.len()` is itself a boundary.
|
||||
pub(crate) struct Tlv<'a> {
|
||||
data: &'a [u8],
|
||||
}
|
||||
|
||||
impl<'a> Tlv<'a> {
|
||||
const MAGIC: &'static [u8; 4] = b"NVFW";
|
||||
|
||||
/// Parses `data` as a TLV firmware image, returning [`EINVAL`] if the image is malformed.
|
||||
pub(crate) fn new(data: &'a [u8]) -> Result<Self> {
|
||||
// Verify that the magic bytes exist and are the correct value
|
||||
let magic_len = Self::MAGIC.len();
|
||||
if data
|
||||
.get(..magic_len)
|
||||
.is_none_or(|magic| magic != Self::MAGIC)
|
||||
{
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// The payload is the contiguous sequence of TLV blocks after the magic.
|
||||
let payload = data.get(magic_len..).ok_or(EINVAL)?;
|
||||
|
||||
// The spec says every TLV must have a VERS tag.
|
||||
let mut has_vers = false;
|
||||
|
||||
let mut rest = payload;
|
||||
while !rest.is_empty() {
|
||||
// Validate and extract the header (type, length).
|
||||
let Some(header): Option<TlvBlockHeader> = rest
|
||||
.get(..TlvBlockHeader::SIZE)
|
||||
.and_then(TlvBlockHeader::parse)
|
||||
else {
|
||||
return Err(EINVAL);
|
||||
};
|
||||
|
||||
has_vers |= header.tag == *b"VERS";
|
||||
|
||||
// The `length` field of a TLV block contains the actual byte length of the
|
||||
// value, but each TLV block is aligned to a 4-byte boundary.
|
||||
let Some(stored_size) = header.length.checked_next_multiple_of(4) else {
|
||||
return Err(EINVAL);
|
||||
};
|
||||
|
||||
let length = TlvBlockHeader::SIZE
|
||||
.checked_add(stored_size)
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
rest = rest.split_at_checked(length).ok_or(EINVAL)?.1;
|
||||
}
|
||||
|
||||
if !has_vers {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// INVARIANT: the loop above walked `payload` block-by-block. For each block, the
|
||||
// header is parsed (`TlvBlockHeader::parse` rejects non-ASCII tags), and the
|
||||
// stored extent (`SIZE + length.next_multiple_of(4)`) is computed without
|
||||
// overflow and split off `rest` only when it fits. The loop ends only when `rest`
|
||||
// is empty, so the byte stream is an exact concatenation of blocks with no
|
||||
// trailing partial header or slack.
|
||||
Ok(Self { data: payload })
|
||||
}
|
||||
|
||||
fn iter(&self) -> TlvIter<'_, 'a> {
|
||||
// INVARIANT: 0 is a block boundary, either the start of the first block,
|
||||
// or `data.len()` when `data` is empty.
|
||||
TlvIter { tlv: self, pos: 0 }
|
||||
}
|
||||
|
||||
fn find(&self, tag: &[u8; 4]) -> Result<TlvBlock<'a>> {
|
||||
self.iter().find(|b| b.tag == *tag).ok_or(EINVAL)
|
||||
}
|
||||
|
||||
/// Return a slice of bytes.
|
||||
///
|
||||
/// Returns `EINVAL` if the value is empty.
|
||||
pub(crate) fn get_bytes(&self, tag: &[u8; 4]) -> Result<&'a [u8]> {
|
||||
let tlv = self.find(tag)?;
|
||||
|
||||
// Treat empty value as an error, to avoid trying to parse nothing.
|
||||
if tlv.value.is_empty() {
|
||||
return Err(EINVAL); // TODO: Use ENODATA once available.
|
||||
}
|
||||
|
||||
Ok(tlv.value)
|
||||
}
|
||||
|
||||
/// Return a little-endian u32.
|
||||
pub(crate) fn get_u32(&self, tag: &[u8; 4]) -> Result<u32> {
|
||||
let tlv = self.find(tag)?;
|
||||
|
||||
tlv.value
|
||||
.try_into()
|
||||
.ok()
|
||||
.map(u32::from_le_bytes)
|
||||
.ok_or(EINVAL)
|
||||
}
|
||||
|
||||
/// Return a string value.
|
||||
pub(crate) fn get_string(&self, tag: &[u8; 4]) -> Result<&'a str> {
|
||||
let tlv = self.find(tag)?;
|
||||
|
||||
let bytes = tlv.value;
|
||||
|
||||
// Strings can only contain printable ASCII characters.
|
||||
if bytes.iter().any(|&b| !(32..127).contains(&b)) {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
core::str::from_utf8(bytes).map_err(|_| EINVAL)
|
||||
}
|
||||
|
||||
/// Obtain the nth signature from a SIGN tag. If `index` is None,
|
||||
/// then return the last signature.
|
||||
pub(crate) fn get_signature(&self, index: Option<usize>) -> Result<&'a [u8]> {
|
||||
let num_sigs: usize = match self.get_u32(b"NSIG")? {
|
||||
0 => return Err(EINVAL),
|
||||
n => n.into_safe_cast(),
|
||||
};
|
||||
|
||||
let sig_bytes = self.get_bytes(b"SIGN")?;
|
||||
|
||||
// Ensure that sig_bytes can be divided evenly into chunks.
|
||||
if sig_bytes.len() % num_sigs != 0 {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// num_sigs cannot be 0, and sig_bytes cannot be empty, so this cannot panic.
|
||||
let sig_size = sig_bytes.len() / num_sigs;
|
||||
|
||||
let index = index.unwrap_or(num_sigs - 1);
|
||||
|
||||
sig_bytes.chunks_exact(sig_size).nth(index).ok_or(EINVAL)
|
||||
}
|
||||
}
|
||||
|
|
@ -11,6 +11,7 @@
|
|||
device,
|
||||
dma::Coherent,
|
||||
io::poll::read_poll_timeout,
|
||||
num::TryIntoBounded,
|
||||
prelude::*,
|
||||
ptr::{
|
||||
Alignable,
|
||||
|
|
@ -30,13 +31,17 @@
|
|||
fsp::Fsp as FspEngine,
|
||||
Falcon, //
|
||||
},
|
||||
fb::FbLayout,
|
||||
fb::FbSizes,
|
||||
firmware::fsp::{
|
||||
FmcSignatures,
|
||||
FspFirmware, //
|
||||
},
|
||||
gpu::Chipset,
|
||||
gsp::GspFmcBootParams,
|
||||
gsp::{
|
||||
GspFmcBootParams,
|
||||
GspFwWprMeta,
|
||||
LibosMemoryRegionInitArgument, //
|
||||
},
|
||||
mctp::{
|
||||
MctpHeader,
|
||||
NvdmHeader,
|
||||
|
|
@ -48,6 +53,56 @@
|
|||
|
||||
mod hal;
|
||||
|
||||
/// PRC message sub-command.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
#[repr(u8)]
|
||||
enum PrcMessageSubcmd {
|
||||
/// Read a PRC knob value.
|
||||
Read = 0x0c,
|
||||
}
|
||||
|
||||
impl From<PrcMessageSubcmd> for u8 {
|
||||
fn from(value: PrcMessageSubcmd) -> Self {
|
||||
value as u8
|
||||
}
|
||||
}
|
||||
|
||||
/// PRC object identifier.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
#[repr(u8)]
|
||||
enum PrcObjectId {
|
||||
/// vGPU mode configuration knob.
|
||||
VgpuMode = 0x29,
|
||||
}
|
||||
|
||||
impl From<PrcObjectId> for u8 {
|
||||
fn from(value: PrcObjectId) -> Self {
|
||||
value as u8
|
||||
}
|
||||
}
|
||||
|
||||
kernel::impl_flags!(
|
||||
/// PRC request flags.
|
||||
#[derive(Clone, Copy, Default, PartialEq, Eq)]
|
||||
struct PrcFlags(u8);
|
||||
|
||||
/// Individual PRC request flag.
|
||||
#[derive(Clone, Copy, PartialEq, Eq)]
|
||||
enum PrcFlag {
|
||||
/// Request the active knob value for the current boot.
|
||||
Active = 1 << 1,
|
||||
}
|
||||
);
|
||||
|
||||
/// vGPU operating mode as reported by FSP via the PRC protocol.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub(crate) enum VgpuMode {
|
||||
/// vGPU support is disabled on this GPU.
|
||||
Disabled,
|
||||
/// vGPU support is enabled on this GPU.
|
||||
Enabled,
|
||||
}
|
||||
|
||||
/// FSP command response payload (`NVDM_PAYLOAD_COMMAND_RESPONSE`).
|
||||
#[repr(C, packed)]
|
||||
#[derive(Clone, Copy)]
|
||||
|
|
@ -57,17 +112,107 @@ struct NvdmPayloadCommandResponse {
|
|||
error_code: u32,
|
||||
}
|
||||
|
||||
/// Complete FSP response structure with MCTP and NVDM headers.
|
||||
/// PRC message payload.
|
||||
///
|
||||
/// Sent to FSP to query or modify a device configuration knob.
|
||||
#[repr(C, packed)]
|
||||
#[derive(Clone, Copy)]
|
||||
struct FspResponse {
|
||||
struct NvdmPayloadPrc {
|
||||
sub_message_id: u8,
|
||||
flags: u8,
|
||||
object_id: u8,
|
||||
reserved: u8,
|
||||
}
|
||||
|
||||
impl NvdmPayloadPrc {
|
||||
/// Constructs a PRC payload from typed protocol fields.
|
||||
fn new(subcmd: PrcMessageSubcmd, object_id: PrcObjectId, flags: PrcFlags) -> Self {
|
||||
Self {
|
||||
sub_message_id: subcmd.into(),
|
||||
flags: flags.into(),
|
||||
object_id: object_id.into(),
|
||||
reserved: 0,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: NvdmPayloadPrc is a packed C struct with only integral fields.
|
||||
unsafe impl AsBytes for NvdmPayloadPrc {}
|
||||
|
||||
/// PRC response payload containing the knob state value.
|
||||
#[repr(C, packed)]
|
||||
#[derive(Clone, Copy)]
|
||||
struct NvdmPayloadPrcResponse {
|
||||
value_low: u8,
|
||||
value_high: u8,
|
||||
reserved1: u8,
|
||||
reserved2: u8,
|
||||
}
|
||||
|
||||
impl NvdmPayloadPrcResponse {
|
||||
/// Returns the PRC knob value as a little-endian 16-bit integer.
|
||||
fn value(self) -> u16 {
|
||||
u16::from(self.value_low) | (u16::from(self.value_high) << 8)
|
||||
}
|
||||
}
|
||||
|
||||
impl TryFrom<NvdmPayloadPrcResponse> for VgpuMode {
|
||||
type Error = kernel::error::Error;
|
||||
|
||||
fn try_from(value: NvdmPayloadPrcResponse) -> Result<Self> {
|
||||
match value.value() {
|
||||
0 => Ok(VgpuMode::Disabled),
|
||||
1 => Ok(VgpuMode::Enabled),
|
||||
_ => Err(EINVAL),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Common MCTP and NVDM headers shared by all FSP messages.
|
||||
#[repr(C, packed)]
|
||||
#[derive(Clone, Copy)]
|
||||
struct FspMessageHeader {
|
||||
mctp_header: MctpHeader,
|
||||
nvdm_header: NvdmHeader,
|
||||
}
|
||||
|
||||
// SAFETY: FspMessageHeader is a packed C struct with only integral fields.
|
||||
unsafe impl AsBytes for FspMessageHeader {}
|
||||
|
||||
// SAFETY: FspMessageHeader is a packed C struct with only integral fields.
|
||||
unsafe impl FromBytes for FspMessageHeader {}
|
||||
|
||||
impl FspMessageHeader {
|
||||
/// Construct a standard FSP message header for the given NVDM type.
|
||||
fn new(nvdm_type: NvdmType) -> Self {
|
||||
Self {
|
||||
mctp_header: MctpHeader::single_packet(),
|
||||
nvdm_header: NvdmHeader::new(nvdm_type),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Common FSP response header with MCTP, NVDM and command response payloads.
|
||||
#[repr(C, packed)]
|
||||
#[derive(Clone, Copy)]
|
||||
struct FspResponseHeader {
|
||||
header: FspMessageHeader,
|
||||
response: NvdmPayloadCommandResponse,
|
||||
}
|
||||
|
||||
// SAFETY: FspResponse is a packed C struct with only integral fields.
|
||||
unsafe impl FromBytes for FspResponse {}
|
||||
// SAFETY: FspResponseHeader is a packed C struct with only integral fields.
|
||||
unsafe impl FromBytes for FspResponseHeader {}
|
||||
|
||||
/// Complete FSP PRC response including the knob state payload.
|
||||
#[repr(C, packed)]
|
||||
#[derive(Clone, Copy)]
|
||||
struct FspPrcResponse {
|
||||
header: FspResponseHeader,
|
||||
prc_data: NvdmPayloadPrcResponse,
|
||||
}
|
||||
|
||||
// SAFETY: FspPrcResponse is a packed C struct with only integral fields.
|
||||
unsafe impl FromBytes for FspPrcResponse {}
|
||||
|
||||
/// Trait implemented by types representing a message to send to FSP.
|
||||
///
|
||||
|
|
@ -94,45 +239,56 @@ struct NvdmPayloadCot {
|
|||
gsp_boot_args_sysmem_offset: u64,
|
||||
}
|
||||
|
||||
/// Complete FSP message structure with MCTP and NVDM headers.
|
||||
/// Complete FSP COT (Chain of Trust) message structure.
|
||||
#[repr(C)]
|
||||
#[derive(Clone, Copy)]
|
||||
struct FspMessage {
|
||||
mctp_header: MctpHeader,
|
||||
nvdm_header: NvdmHeader,
|
||||
struct FspCotMessage {
|
||||
header: FspMessageHeader,
|
||||
cot: NvdmPayloadCot,
|
||||
}
|
||||
|
||||
impl FspMessage {
|
||||
/// Returns an in-place initializer for [`FspMessage`].
|
||||
fn new<'a>(
|
||||
fb_layout: &FbLayout,
|
||||
fsp_fw: &'a FspFirmware,
|
||||
args: &'a FmcBootArgs,
|
||||
) -> Result<impl Init<Self> + 'a> {
|
||||
// frts_offset is relative to FB end: FRTS_location = FB_END - frts_offset
|
||||
let frts_vidmem_offset = if !args.resume {
|
||||
let frts_reserved_size = fb_layout.heap.len() + u64::from(fb_layout.pmu_reserved_size);
|
||||
impl FspCotMessage {
|
||||
/// Computes the FRTS vidmem offset for the Chain-of-Trust message. It is measured backwards
|
||||
/// from the end of the framebuffer.
|
||||
fn frts_vidmem_offset(hal: &dyn hal::FspHal, fb_info: &FbSizes) -> Result<u64> {
|
||||
let mut offset = hal.fb_end_reserved_size();
|
||||
|
||||
frts_reserved_size
|
||||
// As per OpenRM's `kfspPrepareBootCommands_GH100`.
|
||||
if fb_info.pmu_reserved_size != 0 {
|
||||
offset = (offset + u64::from(fb_info.pmu_reserved_size))
|
||||
// The 2 MiB alignment is r570-specific.
|
||||
.align_up(Alignment::new::<SZ_2M>())
|
||||
.ok_or(EINVAL)?
|
||||
.ok_or(EINVAL)?;
|
||||
}
|
||||
|
||||
Ok(offset)
|
||||
}
|
||||
|
||||
/// Returns an in-place initializer for [`FspCotMessage`].
|
||||
fn new<'a>(
|
||||
fb_info: &FbSizes,
|
||||
fsp_fw: &'a FspFirmware,
|
||||
args: &'a FmcBootArgs<'_>,
|
||||
) -> Result<impl Init<Self> + 'a> {
|
||||
let hal = hal::fsp_hal(args.chipset).ok_or(ENOTSUPP)?;
|
||||
|
||||
let frts_vidmem_offset = if !args.resume {
|
||||
Self::frts_vidmem_offset(hal, fb_info)?
|
||||
} else {
|
||||
0
|
||||
};
|
||||
|
||||
let frts_size: u32 = if !args.resume {
|
||||
fb_layout.frts.len().try_into()?
|
||||
fb_info.frts_size.try_into()?
|
||||
} else {
|
||||
0
|
||||
};
|
||||
|
||||
let version = hal::fsp_hal(args.chipset).ok_or(ENOTSUPP)?.cot_version();
|
||||
let version = hal.cot_version();
|
||||
let size = num::usize_into_u16::<{ core::mem::size_of::<NvdmPayloadCot>() }>();
|
||||
|
||||
Ok(init!(Self {
|
||||
mctp_header: MctpHeader::single_packet(),
|
||||
nvdm_header: NvdmHeader::new(NvdmType::Cot),
|
||||
header: FspMessageHeader::new(NvdmType::Cot),
|
||||
// The payload is packed, so we cannot use `init!`. Initialize it member-by-member using
|
||||
// `chain`.
|
||||
cot <- pin_init::init_zeroed(),
|
||||
|
|
@ -140,12 +296,12 @@ fn new<'a>(
|
|||
.chain(move |msg| {
|
||||
msg.cot.version = version;
|
||||
msg.cot.size = size;
|
||||
msg.cot.gsp_fmc_sysmem_offset = fsp_fw.fmc_image.dma_handle();
|
||||
msg.cot.gsp_fmc_sysmem_offset = fsp_fw.fmc_image.dma_address();
|
||||
msg.cot.frts_vidmem_offset = frts_vidmem_offset;
|
||||
msg.cot.frts_vidmem_size = frts_size;
|
||||
// frts_sysmem_* intentionally left at zero for now, but will be needed for e.g.
|
||||
// systems without VRAM.
|
||||
msg.cot.gsp_boot_args_sysmem_offset = args.fmc_boot_params.dma_handle();
|
||||
// frts_sysmem_* are left at zero because this path places FRTS in vidmem. The sysmem
|
||||
// fields point to an FRTS buffer in sysmem instead, for systems without VRAM.
|
||||
msg.cot.gsp_boot_args_sysmem_offset = args.fmc_boot_params.dma_address();
|
||||
msg.cot.sigs = *fsp_fw.fmc_sigs;
|
||||
|
||||
Ok(())
|
||||
|
|
@ -153,44 +309,73 @@ fn new<'a>(
|
|||
}
|
||||
}
|
||||
|
||||
// SAFETY: `FspMessage` is `#[repr(C)]` with no padding, so all of its
|
||||
// SAFETY: `FspCotMessage` is `#[repr(C)]` with no padding, so all of its
|
||||
// bytes are initialized.
|
||||
unsafe impl AsBytes for FspMessage {}
|
||||
unsafe impl AsBytes for FspCotMessage {}
|
||||
|
||||
impl MessageToFsp for FspMessage {
|
||||
/// Complete FSP PRC message.
|
||||
#[repr(C, packed)]
|
||||
#[derive(Clone, Copy)]
|
||||
struct FspPrcMessage {
|
||||
header: FspMessageHeader,
|
||||
prc: NvdmPayloadPrc,
|
||||
}
|
||||
|
||||
impl FspPrcMessage {
|
||||
/// Constructs a PRC message.
|
||||
fn new(subcmd: PrcMessageSubcmd, object_id: PrcObjectId, flags: PrcFlags) -> Self {
|
||||
Self {
|
||||
header: FspMessageHeader::new(NvdmType::Prc),
|
||||
prc: NvdmPayloadPrc::new(subcmd, object_id, flags),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: FspPrcMessage is a packed C struct with only integral fields.
|
||||
unsafe impl AsBytes for FspPrcMessage {}
|
||||
|
||||
impl MessageToFsp for FspCotMessage {
|
||||
const NVDM_TYPE: NvdmType = NvdmType::Cot;
|
||||
}
|
||||
|
||||
impl MessageToFsp for FspPrcMessage {
|
||||
const NVDM_TYPE: NvdmType = NvdmType::Prc;
|
||||
}
|
||||
|
||||
/// Bundled arguments for FMC boot via FSP Chain of Trust.
|
||||
pub(crate) struct FmcBootArgs {
|
||||
pub(crate) struct FmcBootArgs<'a> {
|
||||
chipset: Chipset,
|
||||
fmc_boot_params: Coherent<GspFmcBootParams>,
|
||||
resume: bool,
|
||||
// Additional dependencies required to be kept alive for FMC boot.
|
||||
_wpr_meta: Coherent<GspFwWprMeta>,
|
||||
_libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
|
||||
}
|
||||
|
||||
impl FmcBootArgs {
|
||||
impl<'a> FmcBootArgs<'a> {
|
||||
/// Builds FMC boot arguments, allocating the DMA-coherent boot parameter
|
||||
/// structure that FSP will read.
|
||||
pub(crate) fn new(
|
||||
dev: &device::Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
wpr_meta_addr: u64,
|
||||
libos_addr: u64,
|
||||
wpr_meta: Coherent<GspFwWprMeta>,
|
||||
libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
|
||||
resume: bool,
|
||||
) -> Result<Self> {
|
||||
let init = GspFmcBootParams::new(wpr_meta_addr, libos_addr);
|
||||
let init = GspFmcBootParams::new(wpr_meta.dma_address(), libos.dma_address());
|
||||
|
||||
Ok(Self {
|
||||
chipset,
|
||||
fmc_boot_params: Coherent::<GspFmcBootParams>::init(dev, GFP_KERNEL, init)?,
|
||||
resume,
|
||||
_wpr_meta: wpr_meta,
|
||||
_libos: libos,
|
||||
})
|
||||
}
|
||||
|
||||
/// DMA address of the FMC boot parameters, needed after boot for lockdown
|
||||
/// release polling.
|
||||
pub(crate) fn boot_params_dma_handle(&self) -> u64 {
|
||||
self.fmc_boot_params.dma_handle()
|
||||
/// Returns the FMC boot parameters allocation.
|
||||
pub(crate) fn boot_params(&self) -> &Coherent<GspFmcBootParams> {
|
||||
&self.fmc_boot_params
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -199,28 +384,45 @@ pub(crate) fn boot_params_dma_handle(&self) -> u64 {
|
|||
/// An `Fsp` is produced by [`Fsp::wait_secure_boot`], which only returns once FSP secure boot
|
||||
/// has completed. It owns the FSP falcon and the FMC firmware, which are used for the subsequent
|
||||
/// Chain of Trust boot.
|
||||
pub(crate) struct Fsp {
|
||||
falcon: Falcon<FspEngine>,
|
||||
pub(crate) struct Fsp<'a> {
|
||||
falcon: Falcon<'a, FspEngine>,
|
||||
fsp_fw: FspFirmware,
|
||||
}
|
||||
|
||||
impl Fsp {
|
||||
impl<'a> Fsp<'a> {
|
||||
/// Attempts to create a `Fsp` instance.
|
||||
///
|
||||
/// This can involve waiting for FSP secure boot completion, but should be instantaneous in
|
||||
/// practice.
|
||||
///
|
||||
/// If `chipset` doesn't support FSP, `Ok(None)` is returned.
|
||||
pub(crate) fn try_new(
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
bar: Bar0<'a>,
|
||||
chipset: Chipset,
|
||||
) -> Result<Option<Self>> {
|
||||
match hal::fsp_hal(chipset) {
|
||||
None => Ok(None),
|
||||
Some(hal) => Self::wait_secure_boot(dev, bar, chipset, hal).map(Option::Some),
|
||||
}
|
||||
}
|
||||
|
||||
/// Waits for FSP secure boot completion, then returns the [`Fsp`] interface.
|
||||
///
|
||||
/// Polls the thermal scratch register until FSP signals boot completion or the timeout
|
||||
/// elapses. Returning an [`Fsp`] only on success guarantees, at the API level, that the
|
||||
/// interface is not used before secure boot has completed.
|
||||
pub(crate) fn wait_secure_boot(
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
fn wait_secure_boot(
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
bar: Bar0<'a>,
|
||||
chipset: Chipset,
|
||||
fsp_fw: FspFirmware,
|
||||
) -> Result<Fsp> {
|
||||
hal: &'static dyn hal::FspHal,
|
||||
) -> Result<Fsp<'a>> {
|
||||
/// FSP secure boot completion timeout in milliseconds.
|
||||
const FSP_SECURE_BOOT_TIMEOUT_MS: i64 = 5000;
|
||||
|
||||
let hal = hal::fsp_hal(chipset).ok_or(ENOTSUPP)?;
|
||||
let falcon = Falcon::<FspEngine>::new(dev, chipset)?;
|
||||
let falcon = Falcon::<FspEngine>::new(dev, chipset, bar)?;
|
||||
let fsp_fw = FspFirmware::new(dev, chipset)?;
|
||||
|
||||
read_poll_timeout(
|
||||
|| Ok(hal.fsp_boot_status(bar)),
|
||||
|
|
@ -236,23 +438,25 @@ pub(crate) fn wait_secure_boot(
|
|||
}
|
||||
|
||||
/// Sends a message to FSP and waits for the response.
|
||||
fn send_sync_fsp<M>(&mut self, dev: &device::Device, bar: Bar0<'_>, msg: &M) -> Result
|
||||
/// Returns the full response buffer on success.
|
||||
fn send_sync_fsp<M>(&mut self, dev: &device::Device, msg: &M) -> Result<KVec<u8>>
|
||||
where
|
||||
M: MessageToFsp,
|
||||
{
|
||||
self.falcon.send_msg(bar, msg.as_bytes())?;
|
||||
self.falcon.send_msg(msg.as_bytes())?;
|
||||
|
||||
let response_buf = self.falcon.recv_msg(bar).inspect_err(|e| {
|
||||
let response_buf = self.falcon.recv_msg().inspect_err(|e| {
|
||||
dev_err!(dev, "FSP response error: {:?}\n", e);
|
||||
})?;
|
||||
|
||||
let (response, _) = FspResponse::from_bytes_prefix(&response_buf[..]).ok_or_else(|| {
|
||||
dev_err!(dev, "FSP response too small: {}\n", response_buf.len());
|
||||
EIO
|
||||
})?;
|
||||
let (response, _) =
|
||||
FspResponseHeader::from_bytes_prefix(&response_buf[..]).ok_or_else(|| {
|
||||
dev_err!(dev, "FSP response too small: {}\n", response_buf.len());
|
||||
EIO
|
||||
})?;
|
||||
|
||||
let mctp_header = response.mctp_header;
|
||||
let nvdm_header = response.nvdm_header;
|
||||
let mctp_header = response.header.mctp_header;
|
||||
let nvdm_header = response.header.nvdm_header;
|
||||
let command_nvdm_type = response.response.command_nvdm_type;
|
||||
let error_code = response.response.error_code;
|
||||
|
||||
|
|
@ -274,7 +478,7 @@ fn send_sync_fsp<M>(&mut self, dev: &device::Device, bar: Bar0<'_>, msg: &M) ->
|
|||
return Err(EIO);
|
||||
}
|
||||
|
||||
if command_nvdm_type != u8::from(M::NVDM_TYPE).into() {
|
||||
if command_nvdm_type.try_into_bounded() != Some(M::NVDM_TYPE.into()) {
|
||||
dev_err!(
|
||||
dev,
|
||||
"Expected NVDM type {:?} in reply, got {:#x}\n",
|
||||
|
|
@ -294,7 +498,34 @@ fn send_sync_fsp<M>(&mut self, dev: &device::Device, bar: Bar0<'_>, msg: &M) ->
|
|||
return Err(EIO);
|
||||
}
|
||||
|
||||
Ok(())
|
||||
Ok(response_buf)
|
||||
}
|
||||
|
||||
/// Reads the active vGPU mode from FSP using the PRC protocol.
|
||||
///
|
||||
/// Queries FSP's Management Partition for the active vGPU mode knob value.
|
||||
pub(crate) fn read_vgpu_mode(
|
||||
&mut self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
) -> Result<VgpuMode> {
|
||||
let msg = FspPrcMessage::new(
|
||||
PrcMessageSubcmd::Read,
|
||||
PrcObjectId::VgpuMode,
|
||||
PrcFlags::from(PrcFlag::Active),
|
||||
);
|
||||
|
||||
let response_buf = self.send_sync_fsp(dev, &msg)?;
|
||||
let (prc_response, _) =
|
||||
FspPrcResponse::from_bytes_prefix(&response_buf[..]).ok_or_else(|| {
|
||||
dev_err!(dev, "PRC response too small: {}\n", response_buf.len());
|
||||
EIO
|
||||
})?;
|
||||
|
||||
let prc_data = prc_response.prc_data;
|
||||
|
||||
VgpuMode::try_from(prc_data).inspect_err(|_| {
|
||||
dev_err!(dev, "Unexpected vGPU mode value: {:#x}\n", prc_data.value());
|
||||
})
|
||||
}
|
||||
|
||||
/// Boots GSP FMC via FSP Chain of Trust.
|
||||
|
|
@ -304,15 +535,14 @@ fn send_sync_fsp<M>(&mut self, dev: &device::Device, bar: Bar0<'_>, msg: &M) ->
|
|||
pub(crate) fn boot_fmc(
|
||||
&mut self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
fb_layout: &FbLayout,
|
||||
args: &FmcBootArgs,
|
||||
fb_info: &FbSizes,
|
||||
args: &FmcBootArgs<'_>,
|
||||
) -> Result {
|
||||
dev_dbg!(dev, "Starting FSP boot sequence for {}\n", args.chipset);
|
||||
|
||||
let msg = KBox::init(FspMessage::new(fb_layout, &self.fsp_fw, args)?, GFP_KERNEL)?;
|
||||
let msg = KBox::init(FspCotMessage::new(fb_info, &self.fsp_fw, args)?, GFP_KERNEL)?;
|
||||
|
||||
self.send_sync_fsp(dev, bar, &*msg)?;
|
||||
let _response_buf = self.send_sync_fsp(dev, &*msg)?;
|
||||
|
||||
dev_dbg!(dev, "FSP Chain of Trust completed successfully\n");
|
||||
Ok(())
|
||||
|
|
|
|||
|
|
@ -19,6 +19,10 @@ pub(super) trait FspHal {
|
|||
|
||||
/// Returns the FSP Chain of Trust protocol version this chipset advertises.
|
||||
fn cot_version(&self) -> u16;
|
||||
|
||||
// TODO: consider moving this into the TLV firmware metadata when ready
|
||||
/// Returns the size reserved at the end of the framebuffer, in bytes.
|
||||
fn fb_end_reserved_size(&self) -> u64;
|
||||
}
|
||||
|
||||
/// Returns the FSP HAL, or `None` if the architecture doesn't support FSP.
|
||||
|
|
|
|||
|
|
@ -1,6 +1,8 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use kernel::sizes::SizeConstants;
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
fsp::hal::FspHal, //
|
||||
|
|
@ -17,6 +19,10 @@ fn fsp_boot_status(&self, bar: Bar0<'_>) -> u32 {
|
|||
fn cot_version(&self) -> u16 {
|
||||
2
|
||||
}
|
||||
|
||||
fn fb_end_reserved_size(&self) -> u64 {
|
||||
u64::SZ_2M + u64::SZ_128K
|
||||
}
|
||||
}
|
||||
|
||||
const GB100: Gb100 = Gb100;
|
||||
|
|
|
|||
|
|
@ -1,7 +1,10 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use kernel::io::Io;
|
||||
use kernel::{
|
||||
io::Io,
|
||||
sizes::SizeConstants, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
|
|
@ -21,6 +24,10 @@ fn fsp_boot_status(&self, bar: Bar0<'_>) -> u32 {
|
|||
fn cot_version(&self) -> u16 {
|
||||
2
|
||||
}
|
||||
|
||||
fn fb_end_reserved_size(&self) -> u64 {
|
||||
u64::SZ_2M + u64::SZ_128K
|
||||
}
|
||||
}
|
||||
|
||||
const GB202: Gb202 = Gb202;
|
||||
|
|
|
|||
|
|
@ -1,7 +1,10 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use kernel::io::Io;
|
||||
use kernel::{
|
||||
io::Io,
|
||||
sizes::SizeConstants, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
|
|
@ -26,6 +29,10 @@ fn fsp_boot_status(&self, bar: Bar0<'_>) -> u32 {
|
|||
fn cot_version(&self) -> u16 {
|
||||
1
|
||||
}
|
||||
|
||||
fn fb_end_reserved_size(&self) -> u64 {
|
||||
u64::SZ_2M
|
||||
}
|
||||
}
|
||||
|
||||
const GH100: Gh100 = Gh100;
|
||||
|
|
|
|||
|
|
@ -9,7 +9,8 @@
|
|||
io::Io,
|
||||
num::Bounded,
|
||||
pci,
|
||||
prelude::*, //
|
||||
prelude::*,
|
||||
sizes::SizeConstants, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
|
|
@ -21,11 +22,15 @@
|
|||
Falcon, //
|
||||
},
|
||||
fb::SysmemFlush,
|
||||
fsp::Fsp,
|
||||
gsp::{
|
||||
self,
|
||||
Gsp, //
|
||||
commands::GetGspStaticInfoReply,
|
||||
Gsp,
|
||||
GspBootContext, //
|
||||
},
|
||||
regs,
|
||||
vgpu::VgpuManager, //
|
||||
};
|
||||
|
||||
mod hal;
|
||||
|
|
@ -130,22 +135,6 @@ pub(crate) const fn arch(self) -> Architecture {
|
|||
}
|
||||
}
|
||||
|
||||
/// Returns `true` if this chipset requires the PIO-loaded bootloader in order to boot FWSEC.
|
||||
///
|
||||
/// This includes all chipsets < GA102.
|
||||
pub(crate) const fn needs_fwsec_bootloader(self) -> bool {
|
||||
matches!(self.arch(), Architecture::Turing) || matches!(self, Self::GA100)
|
||||
}
|
||||
|
||||
/// Returns `true` if this chipset boots via FSP (Hopper and later), which requires the FMC
|
||||
/// firmware image.
|
||||
pub(crate) const fn uses_fsp(self) -> bool {
|
||||
matches!(
|
||||
self.arch(),
|
||||
Architecture::Hopper | Architecture::BlackwellGB10x | Architecture::BlackwellGB20x
|
||||
)
|
||||
}
|
||||
|
||||
/// Returns the address range of the PCI config mirror space.
|
||||
pub(crate) fn pci_config_mirror_range(self) -> Range<u32> {
|
||||
hal::gpu_hal(self).pci_config_mirror_range()
|
||||
|
|
@ -262,37 +251,87 @@ fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
|||
}
|
||||
}
|
||||
|
||||
/// Structure holding the resources required to operate the GPU.
|
||||
/// Self-contained resources to operate and drop the GSP.
|
||||
#[pin_data(PinnedDrop)]
|
||||
pub(crate) struct Gpu<'gpu> {
|
||||
struct GspResources<'gpu> {
|
||||
/// Device owning the GPU.
|
||||
device: &'gpu device::Device<device::Bound>,
|
||||
device: &'gpu pci::Device<device::Bound>,
|
||||
/// Details about the chipset.
|
||||
spec: Spec,
|
||||
/// MMIO mapping of PCI BAR 0.
|
||||
bar: Bar0<'gpu>,
|
||||
/// System memory page required for flushing all pending GPU-side memory writes done through
|
||||
/// PCIE into system memory, via sysmembar (A GPU-initiated HW memory-barrier operation).
|
||||
sysmem_flush: SysmemFlush<'gpu>,
|
||||
/// GSP falcon instance, used for GSP boot up and cleanup.
|
||||
gsp_falcon: Falcon<GspFalcon>,
|
||||
gsp_falcon: Falcon<'gpu, GspFalcon>,
|
||||
/// SEC2 falcon instance, used for GSP boot up and cleanup.
|
||||
sec2_falcon: Falcon<Sec2Falcon>,
|
||||
/// GSP runtime data. Temporarily an empty placeholder.
|
||||
sec2_falcon: Falcon<'gpu, Sec2Falcon>,
|
||||
/// FSP instance, if on an arch that supports it.
|
||||
// TODO: use different resource types for each boot method, and make the relevant Gsp methods
|
||||
// generic against them.
|
||||
fsp: Option<Fsp<'gpu>>,
|
||||
/// vGPU state detected before GSP boot.
|
||||
vgpu: VgpuManager,
|
||||
/// GSP runtime data.
|
||||
#[pin]
|
||||
gsp: Gsp,
|
||||
/// GSP unload firmware bundle, if any.
|
||||
unload_bundle: Option<gsp::UnloadBundle>,
|
||||
}
|
||||
|
||||
/// Structure holding the resources required to operate the GPU.
|
||||
#[pin_data]
|
||||
pub(crate) struct Gpu<'gpu> {
|
||||
spec: Spec,
|
||||
/// Static GPU information as provided by the GSP.
|
||||
gsp_static_info: GetGspStaticInfoReply,
|
||||
/// GSP and its resources.
|
||||
#[pin]
|
||||
gsp_resources: GspResources<'gpu>,
|
||||
/// System memory page required for flushing all pending GPU-side memory writes done through
|
||||
/// PCIE into system memory, via sysmembar (A GPU-initiated HW memory-barrier operation).
|
||||
///
|
||||
/// Must be kept declared *after* `gsp_resources`, as the latter's `PinnedDrop` implementation
|
||||
/// requires the sysmem flush page to be in place.
|
||||
sysmem_flush: SysmemFlush<'gpu>,
|
||||
}
|
||||
|
||||
#[pinned_drop]
|
||||
impl PinnedDrop for GspResources<'_> {
|
||||
fn drop(self: Pin<&mut Self>) {
|
||||
let this = self.project();
|
||||
let device = *this.device;
|
||||
let bar = *this.bar;
|
||||
let bundle = this.unload_bundle.take();
|
||||
|
||||
let _ = this
|
||||
.gsp
|
||||
.as_ref()
|
||||
.get_ref()
|
||||
.unload(
|
||||
GspBootContext {
|
||||
pdev: device,
|
||||
bar,
|
||||
chipset: this.spec.chipset,
|
||||
gsp_falcon: &*this.gsp_falcon,
|
||||
sec2_falcon: &*this.sec2_falcon,
|
||||
fsp: this.fsp.as_mut(),
|
||||
vgpu: &*this.vgpu,
|
||||
},
|
||||
bundle,
|
||||
)
|
||||
.inspect_err(|e| dev_err!(device, "failed to unload GSP: {:?}\n", e));
|
||||
}
|
||||
}
|
||||
|
||||
impl<'gpu> Gpu<'gpu> {
|
||||
pub(crate) fn new(
|
||||
pdev: &'gpu pci::Device<device::Core<'_>>,
|
||||
bar: Bar0<'gpu>,
|
||||
) -> impl PinInit<Self, Error> + 'gpu {
|
||||
let dev = pdev.as_ref();
|
||||
|
||||
try_pin_init!(Self {
|
||||
device: pdev.as_ref(),
|
||||
spec: Spec::new(pdev.as_ref(), bar).inspect(|spec| {
|
||||
dev_info!(pdev,"NVIDIA ({})\n", spec);
|
||||
spec: Spec::new(dev, bar).inspect(|spec| {
|
||||
dev_info!(dev,"NVIDIA ({})\n", spec);
|
||||
})?,
|
||||
|
||||
// We must wait for GFW_BOOT completion before doing any significant setup on the GPU.
|
||||
|
|
@ -305,43 +344,73 @@ pub(crate) fn new(
|
|||
unsafe { pdev.dma_set_mask_and_coherent(dma_mask)? };
|
||||
|
||||
hal.wait_gfw_boot_completion(bar)
|
||||
.inspect_err(|_| dev_err!(pdev, "GFW boot did not complete\n"))?;
|
||||
.inspect_err(|_| dev_err!(dev, "GFW boot did not complete\n"))?;
|
||||
},
|
||||
|
||||
sysmem_flush: SysmemFlush::register(pdev.as_ref(), bar, spec.chipset)?,
|
||||
// Initialize this early because `gsp_resources` depends on it.
|
||||
sysmem_flush: SysmemFlush::register(dev, bar, spec.chipset)?,
|
||||
|
||||
gsp_falcon: Falcon::new(
|
||||
pdev.as_ref(),
|
||||
spec.chipset,
|
||||
)
|
||||
.inspect(|falcon| falcon.clear_swgen0_intr(bar))?,
|
||||
gsp_resources <- try_pin_init!(GspResources {
|
||||
device: pdev,
|
||||
|
||||
sec2_falcon: Falcon::new(pdev.as_ref(), spec.chipset)?,
|
||||
spec: *spec,
|
||||
|
||||
gsp <- Gsp::new(pdev),
|
||||
bar,
|
||||
|
||||
// This member must be initialized last, so the `UnloadBundle` can never be dropped from
|
||||
// outside of the constructed `Gpu`, ensuring that the unload sequence is properly run
|
||||
// in case of failure.
|
||||
unload_bundle: gsp.boot(pdev, bar, spec.chipset, gsp_falcon, sec2_falcon)?,
|
||||
bar,
|
||||
gsp_falcon: Falcon::new(
|
||||
dev,
|
||||
spec.chipset,
|
||||
bar
|
||||
)
|
||||
.inspect(|falcon| falcon.clear_swgen0_intr())?,
|
||||
|
||||
sec2_falcon: Falcon::new(dev, spec.chipset, bar)?,
|
||||
|
||||
fsp: Fsp::try_new(dev, bar, spec.chipset)?,
|
||||
|
||||
vgpu: VgpuManager::new(pdev, spec.chipset, fsp.as_mut()),
|
||||
|
||||
gsp <- Gsp::new(pdev),
|
||||
|
||||
// This member must be initialized last, so the `UnloadBundle` can never be dropped
|
||||
// from outside of the constructed `GspResources`, ensuring that the unload sequence
|
||||
// is properly run in case of failure.
|
||||
unload_bundle: gsp.boot(GspBootContext {
|
||||
pdev,
|
||||
bar,
|
||||
chipset: spec.chipset,
|
||||
gsp_falcon,
|
||||
sec2_falcon,
|
||||
fsp: fsp.as_mut(),
|
||||
vgpu,
|
||||
})?,
|
||||
}),
|
||||
|
||||
gsp_static_info: {
|
||||
// Obtain and display basic GPU information.
|
||||
let info = gsp_resources.gsp.get_static_info(bar)?;
|
||||
match info.gpu_name() {
|
||||
Ok(name) => dev_info!(dev, "GPU name: {}\n", name),
|
||||
Err(e) => dev_warn!(dev, "GPU name unavailable: {:?}\n", e),
|
||||
}
|
||||
|
||||
if !info.usable_fb_regions.is_empty() {
|
||||
dev_dbg!(dev, "Usable FB regions:\n");
|
||||
for region in &info.usable_fb_regions {
|
||||
dev_dbg!(dev, " - {:#x?}\n", region);
|
||||
}
|
||||
|
||||
dev_dbg!(
|
||||
dev,
|
||||
"Total usable VRAM: {} MiB\n",
|
||||
info.usable_fb_regions.iter().fold(0u64, |res, region| res
|
||||
.saturating_add(region.end - region.start))
|
||||
/ u64::SZ_1M
|
||||
);
|
||||
}
|
||||
|
||||
info
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
#[pinned_drop]
|
||||
impl PinnedDrop for Gpu<'_> {
|
||||
fn drop(self: Pin<&mut Self>) {
|
||||
let this = self.project();
|
||||
let device = *this.device;
|
||||
let bar = *this.bar;
|
||||
let bundle = this.unload_bundle.take();
|
||||
|
||||
let _ = this
|
||||
.gsp
|
||||
.as_ref()
|
||||
.get_ref()
|
||||
.unload(device, bar, &*this.gsp_falcon, &*this.sec2_falcon, bundle)
|
||||
.inspect_err(|e| dev_err!(device, "failed to unload GSP: {:?}\n", e));
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -9,60 +9,95 @@
|
|||
dma::{
|
||||
Coherent,
|
||||
CoherentBox,
|
||||
CoherentView,
|
||||
DmaAddress, //
|
||||
},
|
||||
io::{
|
||||
io_project,
|
||||
io_write,
|
||||
Io, //
|
||||
},
|
||||
pci,
|
||||
prelude::*,
|
||||
transmute::{
|
||||
AsBytes,
|
||||
FromBytes, //
|
||||
}, //
|
||||
prelude::*, //
|
||||
};
|
||||
|
||||
pub(crate) mod cmdq;
|
||||
pub(crate) mod commands;
|
||||
mod fw;
|
||||
mod regs;
|
||||
mod sequencer;
|
||||
|
||||
pub(crate) use fw::{
|
||||
GspFmcBootParams,
|
||||
GspFwWprMeta,
|
||||
LibosMemoryRegionInitArgument,
|
||||
LibosParams, //
|
||||
};
|
||||
pub(crate) use hal::boot_firmware_files;
|
||||
|
||||
use crate::{
|
||||
gsp::cmdq::Cmdq,
|
||||
gsp::fw::{
|
||||
GspArgumentsPadded,
|
||||
LibosMemoryRegionInitArgument, //
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
gsp::Gsp as GspFalcon,
|
||||
sec2::Sec2 as Sec2Falcon,
|
||||
Falcon, //
|
||||
},
|
||||
fsp::Fsp,
|
||||
gpu::Chipset,
|
||||
gsp::{
|
||||
cmdq::Cmdq,
|
||||
fw::GspArgumentsPadded, //
|
||||
},
|
||||
num,
|
||||
vgpu::VgpuManager, //
|
||||
};
|
||||
|
||||
pub(crate) const GSP_PAGE_SHIFT: usize = 12;
|
||||
pub(crate) const GSP_PAGE_SIZE: usize = 1 << GSP_PAGE_SHIFT;
|
||||
|
||||
/// Common context for the GSP boot process.
|
||||
///
|
||||
/// It carries two distinct lifetimes:
|
||||
///
|
||||
/// - `'gpu` is the lifetime of the bound GPU device, as captured by the GPU subdevices.
|
||||
/// - `'ctx` is a shorter lifetime during which this context borrows those subdevices.
|
||||
pub(crate) struct GspBootContext<'ctx, 'gpu> {
|
||||
pub(crate) pdev: &'gpu pci::Device<device::Bound>,
|
||||
pub(crate) bar: Bar0<'gpu>,
|
||||
pub(crate) chipset: Chipset,
|
||||
pub(crate) gsp_falcon: &'ctx Falcon<'gpu, GspFalcon>,
|
||||
pub(crate) sec2_falcon: &'ctx Falcon<'gpu, Sec2Falcon>,
|
||||
pub(crate) fsp: Option<&'ctx mut Fsp<'gpu>>,
|
||||
pub(crate) vgpu: &'ctx VgpuManager,
|
||||
}
|
||||
|
||||
impl<'ctx, 'gpu> GspBootContext<'ctx, 'gpu> {
|
||||
pub(crate) fn dev(&self) -> &'gpu device::Device<device::Bound> {
|
||||
self.pdev.as_ref()
|
||||
}
|
||||
}
|
||||
|
||||
/// Number of GSP pages to use in a RM log buffer.
|
||||
const RM_LOG_BUFFER_NUM_PAGES: usize = 0x10;
|
||||
const LOG_BUFFER_SIZE: usize = RM_LOG_BUFFER_NUM_PAGES * GSP_PAGE_SIZE;
|
||||
|
||||
/// Array of page table entries, as understood by the GSP bootloader.
|
||||
#[repr(C)]
|
||||
#[derive(FromBytes, IntoBytes)]
|
||||
struct PteArray<const NUM_ENTRIES: usize>([u64; NUM_ENTRIES]);
|
||||
|
||||
/// SAFETY: arrays of `u64` implement `FromBytes` and we are but a wrapper around one.
|
||||
unsafe impl<const NUM_ENTRIES: usize> FromBytes for PteArray<NUM_ENTRIES> {}
|
||||
|
||||
/// SAFETY: arrays of `u64` implement `AsBytes` and we are but a wrapper around one.
|
||||
unsafe impl<const NUM_ENTRIES: usize> AsBytes for PteArray<NUM_ENTRIES> {}
|
||||
|
||||
impl<const NUM_PAGES: usize> PteArray<NUM_PAGES> {
|
||||
/// Returns the page table entry for `index`, for a mapping starting at `start`.
|
||||
// TODO: Replace with `IoView` projection once available.
|
||||
fn entry(start: DmaAddress, index: usize) -> Result<u64> {
|
||||
start
|
||||
.checked_add(num::usize_as_u64(index) << GSP_PAGE_SHIFT)
|
||||
.ok_or(EOVERFLOW)
|
||||
/// Initialize a new page table array mapping `NUM_PAGES` GSP pages starting at address `start`.
|
||||
fn init(view: CoherentView<'_, Self>, start: DmaAddress) -> Result<()> {
|
||||
for i in 0..NUM_PAGES {
|
||||
io_write!(view, .0[build: i],
|
||||
start
|
||||
.checked_add(num::usize_as_u64(i) << GSP_PAGE_SHIFT)
|
||||
.ok_or(EOVERFLOW)?
|
||||
);
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -87,19 +122,14 @@ impl LogBuffer {
|
|||
fn new(dev: &device::Device<device::Bound>) -> Result<Self> {
|
||||
let obj = Self(Coherent::zeroed(dev, GFP_KERNEL)?);
|
||||
|
||||
let start_addr = obj.0.dma_handle();
|
||||
let start_addr = obj.0.dma_address();
|
||||
|
||||
// SAFETY: `obj` has just been created and we are its sole user.
|
||||
let pte_region = unsafe {
|
||||
&mut obj.0.as_mut()[size_of::<u64>()..][..RM_LOG_BUFFER_NUM_PAGES * size_of::<u64>()]
|
||||
};
|
||||
|
||||
// Write values one by one to avoid an on-stack instance of `PteArray`.
|
||||
for (i, chunk) in pte_region.chunks_exact_mut(size_of::<u64>()).enumerate() {
|
||||
let pte_value = PteArray::<0>::entry(start_addr, i)?;
|
||||
|
||||
chunk.copy_from_slice(&pte_value.to_ne_bytes());
|
||||
}
|
||||
let pte_view = io_project!(
|
||||
obj.0,
|
||||
[build: size_of::<u64>()..][build: ..RM_LOG_BUFFER_NUM_PAGES * size_of::<u64>()]
|
||||
)
|
||||
.try_cast::<PteArray<RM_LOG_BUFFER_NUM_PAGES>>()?;
|
||||
PteArray::init(pte_view, start_addr)?;
|
||||
|
||||
Ok(obj)
|
||||
}
|
||||
|
|
@ -185,6 +215,11 @@ pub(crate) fn new(pdev: &pci::Device<device::Bound>) -> impl PinInit<Self, Error
|
|||
}))
|
||||
})
|
||||
}
|
||||
|
||||
/// Query the GSP for the static GPU information.
|
||||
pub(crate) fn get_static_info(&self, bar: Bar0<'_>) -> Result<commands::GetGspStaticInfoReply> {
|
||||
self.cmdq.send_command(bar, commands::GetGspStaticInfo)
|
||||
}
|
||||
}
|
||||
|
||||
/// Opaque bundle required to unload the GSP. Created by [`Gsp::boot`], consumed by [`Gsp::unload`].
|
||||
|
|
|
|||
|
|
@ -3,10 +3,7 @@
|
|||
|
||||
use kernel::{
|
||||
bits,
|
||||
device,
|
||||
dma::Coherent,
|
||||
io::poll::read_poll_timeout,
|
||||
pci,
|
||||
prelude::*,
|
||||
time::Delta,
|
||||
types::ScopeGuard, //
|
||||
|
|
@ -16,82 +13,15 @@
|
|||
driver::Bar0,
|
||||
falcon::{
|
||||
gsp::Gsp,
|
||||
sec2::Sec2,
|
||||
Falcon, //
|
||||
},
|
||||
fb::FbLayout,
|
||||
firmware::{
|
||||
gsp::GspFirmware,
|
||||
FIRMWARE_VERSION, //
|
||||
},
|
||||
gpu::Chipset,
|
||||
firmware::gsp::GspFirmware,
|
||||
gsp::{
|
||||
cmdq::Cmdq,
|
||||
commands,
|
||||
GspFwWprMeta, //
|
||||
commands, //
|
||||
},
|
||||
};
|
||||
|
||||
/// Arguments required to call [`Gsp::unload`](super::Gsp::unload).
|
||||
///
|
||||
/// Stored as their own type to avoid repeating a long and tedious list in [`BootUnloadGuard`].
|
||||
pub(super) struct BootUnloadArgs<'a> {
|
||||
gsp: &'a super::Gsp,
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
bar: Bar0<'a>,
|
||||
gsp_falcon: &'a Falcon<Gsp>,
|
||||
sec2_falcon: &'a Falcon<Sec2>,
|
||||
unload_bundle: Option<super::UnloadBundle>,
|
||||
}
|
||||
|
||||
/// Guard that calls [`Gsp::unload`](super::Gsp::unload) with a
|
||||
/// [`UnloadBundle`](super::UnloadBundle) when dropped.
|
||||
///
|
||||
/// Used to ensure the `UnloadBundle` is run during failure paths.
|
||||
pub(super) struct BootUnloadGuard<'a> {
|
||||
guard: ScopeGuard<BootUnloadArgs<'a>, fn(BootUnloadArgs<'a>)>,
|
||||
}
|
||||
|
||||
impl<'a> BootUnloadGuard<'a> {
|
||||
/// Wraps `unload_bundle` into a guard that executes it when dropped.
|
||||
pub(super) fn new(
|
||||
gsp: &'a super::Gsp,
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
bar: Bar0<'a>,
|
||||
gsp_falcon: &'a Falcon<Gsp>,
|
||||
sec2_falcon: &'a Falcon<Sec2>,
|
||||
unload_bundle: Option<super::UnloadBundle>,
|
||||
) -> Self {
|
||||
Self {
|
||||
guard: ScopeGuard::new_with_data(
|
||||
BootUnloadArgs {
|
||||
gsp,
|
||||
dev,
|
||||
bar,
|
||||
gsp_falcon,
|
||||
sec2_falcon,
|
||||
unload_bundle,
|
||||
},
|
||||
|args| {
|
||||
let _ = super::Gsp::unload(
|
||||
args.gsp,
|
||||
args.dev,
|
||||
args.bar,
|
||||
args.gsp_falcon,
|
||||
args.sec2_falcon,
|
||||
args.unload_bundle,
|
||||
);
|
||||
},
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
/// Disarms the guard and returns the [`UnloadBundle`](super::UnloadBundle) it contains.
|
||||
pub(super) fn dismiss(self) -> Option<super::UnloadBundle> {
|
||||
self.guard.dismiss().unload_bundle
|
||||
}
|
||||
}
|
||||
|
||||
impl super::Gsp {
|
||||
/// Attempt to boot the GSP.
|
||||
///
|
||||
|
|
@ -103,71 +33,64 @@ impl super::Gsp {
|
|||
/// [`Self::unload`]) returned.
|
||||
pub(crate) fn boot(
|
||||
self: Pin<&mut Self>,
|
||||
pdev: &pci::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
chipset: Chipset,
|
||||
gsp_falcon: &Falcon<Gsp>,
|
||||
sec2_falcon: &Falcon<Sec2>,
|
||||
mut ctx: super::GspBootContext<'_, '_>,
|
||||
) -> Result<Option<super::UnloadBundle>> {
|
||||
let pdev = ctx.pdev;
|
||||
let bar = ctx.bar;
|
||||
let chipset = ctx.chipset;
|
||||
let gsp_falcon = ctx.gsp_falcon;
|
||||
let dev = pdev.as_ref();
|
||||
let hal = super::hal::gsp_hal(chipset);
|
||||
|
||||
let gsp_fw = KBox::pin_init(GspFirmware::new(dev, chipset, FIRMWARE_VERSION), GFP_KERNEL)?;
|
||||
|
||||
let fb_layout = FbLayout::new(chipset, bar, &gsp_fw)?;
|
||||
dev_dbg!(dev, "{:#x?}\n", fb_layout);
|
||||
|
||||
let wpr_meta = Coherent::init(dev, GFP_KERNEL, GspFwWprMeta::new(&gsp_fw, &fb_layout))?;
|
||||
let gsp_fw = KBox::pin_init(GspFirmware::new(dev, chipset), GFP_KERNEL)?;
|
||||
|
||||
// Perform the chipset-specific boot sequence, and retrieve the unload bundle.
|
||||
let unload_guard = hal.boot(
|
||||
&self,
|
||||
dev,
|
||||
bar,
|
||||
chipset,
|
||||
&fb_layout,
|
||||
&wpr_meta,
|
||||
gsp_falcon,
|
||||
sec2_falcon,
|
||||
)?;
|
||||
let unload_bundle = hal.boot(&self, &mut ctx, &gsp_fw)?.or_else(|| {
|
||||
dev_warn!(dev, "The GSP won't be able to unload properly on unbind.\n");
|
||||
dev_warn!(
|
||||
dev,
|
||||
"The GPU will need to be reset before the driver can bind again.\n"
|
||||
);
|
||||
|
||||
gsp_falcon.write_os_version(bar, gsp_fw.bootloader.app_version);
|
||||
None
|
||||
});
|
||||
|
||||
let mut unload_guard =
|
||||
ScopeGuard::new_with_data((ctx, unload_bundle), |(ctx, unload_bundle)| {
|
||||
let _ = self.unload(ctx, unload_bundle);
|
||||
});
|
||||
let ctx = &mut unload_guard.0;
|
||||
|
||||
gsp_falcon.write_os_version(gsp_fw.bootloader.app_version);
|
||||
|
||||
// Poll for RISC-V to become active before continuing.
|
||||
read_poll_timeout(
|
||||
|| Ok(gsp_falcon.is_riscv_active(bar)),
|
||||
|| Ok(gsp_falcon.is_riscv_active()),
|
||||
|val: &bool| *val,
|
||||
Delta::from_millis(10),
|
||||
Delta::from_secs(5),
|
||||
)?;
|
||||
|
||||
dev_dbg!(pdev, "RISC-V active? {}\n", gsp_falcon.is_riscv_active(bar),);
|
||||
dev_dbg!(pdev, "RISC-V active? {}\n", gsp_falcon.is_riscv_active(),);
|
||||
|
||||
self.cmdq
|
||||
.send_command_no_wait(bar, commands::SetSystemInfo::new(pdev, chipset))?;
|
||||
self.cmdq
|
||||
.send_command_no_wait(bar, commands::SetRegistry::new())?;
|
||||
.send_command_no_wait(bar, commands::SetRegistry::new(ctx.vgpu.state())?)?;
|
||||
|
||||
hal.post_boot(&self, dev, bar, &gsp_fw, gsp_falcon, sec2_falcon)?;
|
||||
hal.post_boot(&self, ctx, &gsp_fw)?;
|
||||
|
||||
// Wait until GSP is fully initialized.
|
||||
commands::wait_gsp_init_done(&self.cmdq)?;
|
||||
|
||||
// Obtain and display basic GPU information.
|
||||
let info = self.cmdq.send_command(bar, commands::GetGspStaticInfo)?;
|
||||
match info.gpu_name() {
|
||||
Ok(name) => dev_info!(pdev, "GPU name: {}\n", name),
|
||||
Err(e) => dev_warn!(pdev, "GPU name unavailable: {:?}\n", e),
|
||||
}
|
||||
|
||||
Ok(unload_guard.dismiss())
|
||||
Ok(unload_guard.dismiss().1)
|
||||
}
|
||||
|
||||
/// Shut down the GSP and wait until it is offline.
|
||||
fn shutdown_gsp(
|
||||
cmdq: &Cmdq,
|
||||
bar: Bar0<'_>,
|
||||
gsp_falcon: &Falcon<Gsp>,
|
||||
gsp_falcon: &Falcon<'_, Gsp>,
|
||||
mode: commands::PowerStateLevel,
|
||||
) -> Result {
|
||||
// Command to shut the GSP down.
|
||||
|
|
@ -176,7 +99,7 @@ fn shutdown_gsp(
|
|||
// Wait until GSP signals it is suspended.
|
||||
const LIBOS_INTERRUPT_PROCESSOR_SUSPENDED: u32 = bits::bit_u32(31);
|
||||
read_poll_timeout(
|
||||
|| Ok(gsp_falcon.read_mailbox0(bar)),
|
||||
|| Ok(gsp_falcon.read_mailbox0()),
|
||||
|&mb0| mb0 & LIBOS_INTERRUPT_PROCESSOR_SUSPENDED != 0,
|
||||
Delta::from_millis(10),
|
||||
Delta::from_secs(5),
|
||||
|
|
@ -189,17 +112,16 @@ fn shutdown_gsp(
|
|||
/// This stops all activity on the GSP.
|
||||
pub(crate) fn unload(
|
||||
&self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
gsp_falcon: &Falcon<Gsp>,
|
||||
sec2_falcon: &Falcon<Sec2>,
|
||||
mut ctx: super::GspBootContext<'_, '_>,
|
||||
unload_bundle: Option<super::UnloadBundle>,
|
||||
) -> Result {
|
||||
let dev = ctx.dev();
|
||||
|
||||
// Shut down the GSP. Keep going even in case of error.
|
||||
let mut res = Self::shutdown_gsp(
|
||||
&self.cmdq,
|
||||
bar,
|
||||
gsp_falcon,
|
||||
ctx.bar,
|
||||
ctx.gsp_falcon,
|
||||
commands::PowerStateLevel::Level0,
|
||||
)
|
||||
.inspect_err(|e| dev_err!(dev, "GSP shutdown failed: {:?}\n", e));
|
||||
|
|
@ -209,7 +131,7 @@ pub(crate) fn unload(
|
|||
res = res.and(
|
||||
unload_bundle
|
||||
.0
|
||||
.run(dev, bar, gsp_falcon, sec2_falcon)
|
||||
.run(&mut ctx)
|
||||
.inspect_err(|e| dev_err!(dev, "Unload bundle failed: {:?}\n", e)),
|
||||
);
|
||||
} else {
|
||||
|
|
|
|||
|
|
@ -2,16 +2,23 @@
|
|||
|
||||
mod continuation;
|
||||
|
||||
use core::mem;
|
||||
use core::{
|
||||
mem,
|
||||
sync::atomic::{
|
||||
fence,
|
||||
Ordering, //
|
||||
},
|
||||
};
|
||||
|
||||
use kernel::{
|
||||
device,
|
||||
dma::{
|
||||
Coherent,
|
||||
CoherentBox,
|
||||
DmaAddress, //
|
||||
},
|
||||
dma_write,
|
||||
io::{
|
||||
io_project,
|
||||
poll::read_poll_timeout,
|
||||
Io, //
|
||||
},
|
||||
|
|
@ -51,10 +58,11 @@
|
|||
GSP_PAGE_SIZE, //
|
||||
},
|
||||
num,
|
||||
regs,
|
||||
sbuffer::SBufferIter, //
|
||||
};
|
||||
|
||||
use super::regs;
|
||||
|
||||
/// Marker type representing the absence of a reply for a command. Commands using this as their
|
||||
/// reply type are sent using [`Cmdq::send_command_no_wait`].
|
||||
pub(crate) struct NoReply;
|
||||
|
|
@ -171,20 +179,18 @@ struct MsgqData {
|
|||
#[repr(C)]
|
||||
// There is no struct defined for this in the open-gpu-kernel-source headers.
|
||||
// Instead it is defined by code in `GspMsgQueuesInit()`.
|
||||
// TODO: Revert to private once `IoView` projections replace the `gsp_mem` module.
|
||||
pub(super) struct Msgq {
|
||||
struct Msgq {
|
||||
/// Header for sending messages, including the write pointer.
|
||||
pub(super) tx: MsgqTxHeader,
|
||||
tx: MsgqTxHeader,
|
||||
/// Header for receiving messages, including the read pointer.
|
||||
pub(super) rx: MsgqRxHeader,
|
||||
rx: MsgqRxHeader,
|
||||
/// The message queue proper.
|
||||
msgq: MsgqData,
|
||||
}
|
||||
|
||||
/// Structure shared between the driver and the GSP and containing the command and message queues.
|
||||
#[repr(C)]
|
||||
// TODO: Revert to private once `IoView` projections replace the `gsp_mem` module.
|
||||
pub(super) struct GspMem {
|
||||
struct GspMem {
|
||||
/// Self-mapping page table entries.
|
||||
ptes: PteArray<{ Self::PTE_ARRAY_SIZE }>,
|
||||
/// CPU queue: the driver writes commands here, and the GSP reads them. It also contains the
|
||||
|
|
@ -192,13 +198,13 @@ pub(super) struct GspMem {
|
|||
/// index into the GSP queue.
|
||||
///
|
||||
/// This member is read-only for the GSP.
|
||||
pub(super) cpuq: Msgq,
|
||||
cpuq: Msgq,
|
||||
/// GSP queue: the GSP writes messages here, and the driver reads them. It also contains the
|
||||
/// write and read pointers that the GSP updates. This means that the read pointer here is an
|
||||
/// index into the CPU queue.
|
||||
///
|
||||
/// This member is read-only for the driver.
|
||||
pub(super) gspq: Msgq,
|
||||
gspq: Msgq,
|
||||
}
|
||||
|
||||
impl GspMem {
|
||||
|
|
@ -232,20 +238,12 @@ fn new(dev: &device::Device<device::Bound>) -> Result<Self> {
|
|||
const MSGQ_SIZE: u32 = num::usize_into_u32::<{ size_of::<Msgq>() }>();
|
||||
const RX_HDR_OFF: u32 = num::usize_into_u32::<{ mem::offset_of!(Msgq, rx) }>();
|
||||
|
||||
let gsp_mem = Coherent::<GspMem>::zeroed(dev, GFP_KERNEL)?;
|
||||
let mut gsp_mem = CoherentBox::<GspMem>::zeroed(dev, GFP_KERNEL)?;
|
||||
gsp_mem.cpuq.tx = MsgqTxHeader::new(MSGQ_SIZE, RX_HDR_OFF, MSGQ_NUM_PAGES);
|
||||
gsp_mem.cpuq.rx = MsgqRxHeader::new();
|
||||
|
||||
let start = gsp_mem.dma_handle();
|
||||
// Write values one by one to avoid an on-stack instance of `PteArray`.
|
||||
for i in 0..GspMem::PTE_ARRAY_SIZE {
|
||||
dma_write!(gsp_mem, .ptes.0[build: i], PteArray::<0>::entry(start, i)?);
|
||||
}
|
||||
|
||||
dma_write!(
|
||||
gsp_mem,
|
||||
.cpuq.tx,
|
||||
MsgqTxHeader::new(MSGQ_SIZE, RX_HDR_OFF, MSGQ_NUM_PAGES)
|
||||
);
|
||||
dma_write!(gsp_mem, .cpuq.rx, MsgqRxHeader::new());
|
||||
let gsp_mem: Coherent<_> = gsp_mem.into();
|
||||
PteArray::init(io_project!(gsp_mem, .ptes), gsp_mem.dma_address())?;
|
||||
|
||||
Ok(Self(gsp_mem))
|
||||
}
|
||||
|
|
@ -406,7 +404,7 @@ fn allocate_command(&mut self, size: usize, timeout: Delta) -> Result<GspCommand
|
|||
//
|
||||
// - The returned value is within `0..MSGQ_NUM_PAGES`.
|
||||
fn gsp_write_ptr(&self) -> u32 {
|
||||
super::fw::gsp_mem::gsp_write_ptr(&self.0)
|
||||
MsgqTxHeader::write_ptr(io_project!(self.0, .gspq.tx)) % MSGQ_NUM_PAGES
|
||||
}
|
||||
|
||||
// Returns the index of the memory page the GSP will read the next command from.
|
||||
|
|
@ -415,7 +413,7 @@ fn gsp_write_ptr(&self) -> u32 {
|
|||
//
|
||||
// - The returned value is within `0..MSGQ_NUM_PAGES`.
|
||||
fn gsp_read_ptr(&self) -> u32 {
|
||||
super::fw::gsp_mem::gsp_read_ptr(&self.0)
|
||||
MsgqRxHeader::read_ptr(io_project!(self.0, .gspq.rx)) % MSGQ_NUM_PAGES
|
||||
}
|
||||
|
||||
// Returns the index of the memory page the CPU can read the next message from.
|
||||
|
|
@ -424,12 +422,18 @@ fn gsp_read_ptr(&self) -> u32 {
|
|||
//
|
||||
// - The returned value is within `0..MSGQ_NUM_PAGES`.
|
||||
fn cpu_read_ptr(&self) -> u32 {
|
||||
super::fw::gsp_mem::cpu_read_ptr(&self.0)
|
||||
MsgqRxHeader::read_ptr(io_project!(self.0, .cpuq.rx)) % MSGQ_NUM_PAGES
|
||||
}
|
||||
|
||||
// Informs the GSP that it can send `elem_count` new pages into the message queue.
|
||||
fn advance_cpu_read_ptr(&mut self, elem_count: u32) {
|
||||
super::fw::gsp_mem::advance_cpu_read_ptr(&self.0, elem_count)
|
||||
let rx = io_project!(self.0, .cpuq.rx);
|
||||
let rptr = MsgqRxHeader::read_ptr(rx).wrapping_add(elem_count) % MSGQ_NUM_PAGES;
|
||||
|
||||
// Ensure read pointer is properly ordered.
|
||||
fence(Ordering::SeqCst);
|
||||
|
||||
MsgqRxHeader::set_read_ptr(rx, rptr)
|
||||
}
|
||||
|
||||
// Returns the index of the memory page the CPU can write the next command to.
|
||||
|
|
@ -438,12 +442,17 @@ fn advance_cpu_read_ptr(&mut self, elem_count: u32) {
|
|||
//
|
||||
// - The returned value is within `0..MSGQ_NUM_PAGES`.
|
||||
fn cpu_write_ptr(&self) -> u32 {
|
||||
super::fw::gsp_mem::cpu_write_ptr(&self.0)
|
||||
MsgqTxHeader::write_ptr(io_project!(self.0, .cpuq.tx)) % MSGQ_NUM_PAGES
|
||||
}
|
||||
|
||||
// Informs the GSP that it can process `elem_count` new pages from the command queue.
|
||||
fn advance_cpu_write_ptr(&mut self, elem_count: u32) {
|
||||
super::fw::gsp_mem::advance_cpu_write_ptr(&self.0, elem_count)
|
||||
let tx = io_project!(self.0, .cpuq.tx);
|
||||
let wptr = MsgqTxHeader::write_ptr(tx).wrapping_add(elem_count) % MSGQ_NUM_PAGES;
|
||||
MsgqTxHeader::set_write_ptr(tx, wptr);
|
||||
|
||||
// Ensure all command data is visible before triggering the GSP read.
|
||||
fence(Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -478,8 +487,8 @@ pub(crate) struct Cmdq {
|
|||
/// Inner mutex-protected state.
|
||||
#[pin]
|
||||
inner: Mutex<CmdqInner>,
|
||||
/// DMA handle of the command queue's shared memory region.
|
||||
pub(super) dma_handle: DmaAddress,
|
||||
/// DMA address of the command queue's shared memory region.
|
||||
pub(super) dma_addr: DmaAddress,
|
||||
}
|
||||
|
||||
impl Cmdq {
|
||||
|
|
@ -508,7 +517,7 @@ pub(crate) fn new(dev: &device::Device<device::Bound>) -> impl PinInit<Self, Err
|
|||
let gsp_mem = DmaGspMem::new(dev)?;
|
||||
|
||||
Ok(try_pin_init!(Self {
|
||||
dma_handle: gsp_mem.0.dma_handle(),
|
||||
dma_addr: gsp_mem.0.dma_address(),
|
||||
inner <- new_mutex!(CmdqInner {
|
||||
dev: dev.into(),
|
||||
gsp_mem,
|
||||
|
|
|
|||
|
|
@ -5,6 +5,7 @@
|
|||
array,
|
||||
convert::Infallible,
|
||||
ffi::FromBytesUntilNulError,
|
||||
ops::Range,
|
||||
str::Utf8Error, //
|
||||
};
|
||||
|
||||
|
|
@ -33,6 +34,7 @@
|
|||
},
|
||||
},
|
||||
sbuffer::SBufferIter,
|
||||
vgpu::VgpuState, //
|
||||
};
|
||||
|
||||
/// The `GspSetSystemInfo` command.
|
||||
|
|
@ -66,37 +68,55 @@ struct RegistryEntry {
|
|||
|
||||
/// The `SetRegistry` command.
|
||||
pub(crate) struct SetRegistry {
|
||||
entries: [RegistryEntry; Self::NUM_ENTRIES],
|
||||
entries: KVec<RegistryEntry>,
|
||||
}
|
||||
|
||||
impl SetRegistry {
|
||||
// For now we hard-code the registry entries. Future work will allow others to
|
||||
// be added as module parameters.
|
||||
const NUM_ENTRIES: usize = 3;
|
||||
|
||||
/// Creates a new `SetRegistry` command, using a set of hardcoded entries.
|
||||
pub(crate) fn new() -> Self {
|
||||
Self {
|
||||
entries: [
|
||||
// RMSecBusResetEnable - enables PCI secondary bus reset
|
||||
pub(crate) fn new(vgpu_state: VgpuState) -> Result<Self> {
|
||||
let mut entries = KVec::new();
|
||||
|
||||
// RMSecBusResetEnable - enables PCI secondary bus reset
|
||||
entries.push(
|
||||
RegistryEntry {
|
||||
key: "RMSecBusResetEnable",
|
||||
value: 1,
|
||||
},
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
|
||||
// RMForcePcieConfigSave - forces GSP-RM to preserve PCI configuration registers on
|
||||
// any PCI reset.
|
||||
entries.push(
|
||||
RegistryEntry {
|
||||
key: "RMForcePcieConfigSave",
|
||||
value: 1,
|
||||
},
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
|
||||
// RMDevidCheckIgnore - allows GSP-RM to boot even if the PCI dev ID is not found
|
||||
// in the internal product name database.
|
||||
entries.push(
|
||||
RegistryEntry {
|
||||
key: "RMDevidCheckIgnore",
|
||||
value: 1,
|
||||
},
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
|
||||
if matches!(vgpu_state, VgpuState::Enabled { .. }) {
|
||||
// RMSetSriovMode - required when vGPU is enabled.
|
||||
entries.push(
|
||||
RegistryEntry {
|
||||
key: "RMSecBusResetEnable",
|
||||
key: "RMSetSriovMode",
|
||||
value: 1,
|
||||
},
|
||||
// RMForcePcieConfigSave - forces GSP-RM to preserve PCI configuration registers on
|
||||
// any PCI reset.
|
||||
RegistryEntry {
|
||||
key: "RMForcePcieConfigSave",
|
||||
value: 1,
|
||||
},
|
||||
// RMDevidCheckIgnore - allows GSP-RM to boot even if the PCI dev ID is not found
|
||||
// in the internal product name database.
|
||||
RegistryEntry {
|
||||
key: "RMDevidCheckIgnore",
|
||||
value: 1,
|
||||
},
|
||||
],
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
}
|
||||
|
||||
Ok(Self { entries })
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -107,15 +127,15 @@ impl CommandToGsp for SetRegistry {
|
|||
type InitError = Infallible;
|
||||
|
||||
fn init(&self) -> impl Init<Self::Command, Self::InitError> {
|
||||
Self::Command::init(Self::NUM_ENTRIES as u32, self.variable_payload_len() as u32)
|
||||
Self::Command::init(self.entries.len() as u32, self.size() as u32)
|
||||
}
|
||||
|
||||
fn variable_payload_len(&self) -> usize {
|
||||
let mut key_size = 0;
|
||||
for i in 0..Self::NUM_ENTRIES {
|
||||
key_size += self.entries[i].key.len() + 1; // +1 for NULL terminator
|
||||
for entry in self.entries.iter() {
|
||||
key_size += entry.key.len() + 1; // +1 for NULL terminator
|
||||
}
|
||||
Self::NUM_ENTRIES * size_of::<fw::commands::PackedRegistryEntry>() + key_size
|
||||
self.entries.len() * size_of::<fw::commands::PackedRegistryEntry>() + key_size
|
||||
}
|
||||
|
||||
fn init_variable_payload(
|
||||
|
|
@ -123,12 +143,12 @@ fn init_variable_payload(
|
|||
dst: &mut SBufferIter<core::array::IntoIter<&mut [u8], 2>>,
|
||||
) -> Result {
|
||||
let string_data_start_offset = size_of::<Self::Command>()
|
||||
+ Self::NUM_ENTRIES * size_of::<fw::commands::PackedRegistryEntry>();
|
||||
+ self.entries.len() * size_of::<fw::commands::PackedRegistryEntry>();
|
||||
|
||||
// Array for string data.
|
||||
let mut string_data = KVec::new();
|
||||
|
||||
for entry in self.entries.iter().take(Self::NUM_ENTRIES) {
|
||||
for entry in self.entries.iter() {
|
||||
dst.write_all(
|
||||
fw::commands::PackedRegistryEntry::new(
|
||||
(string_data_start_offset + string_data.len()) as u32,
|
||||
|
|
@ -191,22 +211,30 @@ fn init(&self) -> impl Init<Self::Command, Self::InitError> {
|
|||
}
|
||||
}
|
||||
|
||||
/// The reply from the GSP to the [`GetGspInfo`] command.
|
||||
/// The reply from the GSP to the [`GetGspStaticInfo`] command.
|
||||
pub(crate) struct GetGspStaticInfoReply {
|
||||
gpu_name: [u8; 64],
|
||||
/// Usable FB (VRAM) regions for driver memory allocation.
|
||||
pub(crate) usable_fb_regions: KVec<Range<u64>>,
|
||||
}
|
||||
|
||||
impl MessageFromGsp for GetGspStaticInfoReply {
|
||||
const FUNCTION: MsgFunction = MsgFunction::GetGspStaticInfo;
|
||||
type Message = fw::commands::GspStaticConfigInfo;
|
||||
type InitError = Infallible;
|
||||
type InitError = Error;
|
||||
|
||||
fn read(
|
||||
msg: &Self::Message,
|
||||
_sbuffer: &mut SBufferIter<array::IntoIter<&[u8], 2>>,
|
||||
) -> Result<Self, Self::InitError> {
|
||||
let mut usable_fb_regions = KVec::new();
|
||||
for region in msg.usable_fb_regions() {
|
||||
usable_fb_regions.push(region, GFP_KERNEL)?;
|
||||
}
|
||||
|
||||
Ok(GetGspStaticInfoReply {
|
||||
gpu_name: msg.gpu_name_str(),
|
||||
usable_fb_regions,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -10,7 +10,15 @@
|
|||
use core::ops::Range;
|
||||
|
||||
use kernel::{
|
||||
dma::Coherent,
|
||||
bitfield,
|
||||
dma::{
|
||||
Coherent,
|
||||
CoherentView, //
|
||||
},
|
||||
io::{
|
||||
io_read,
|
||||
io_write, //
|
||||
},
|
||||
prelude::*,
|
||||
ptr::{
|
||||
Alignable,
|
||||
|
|
@ -28,7 +36,10 @@
|
|||
};
|
||||
|
||||
use crate::{
|
||||
fb::FbLayout,
|
||||
fb::{
|
||||
FbRanges,
|
||||
FbSizes, //
|
||||
},
|
||||
firmware::gsp::GspFirmware,
|
||||
gpu::{
|
||||
Architecture,
|
||||
|
|
@ -44,59 +55,6 @@
|
|||
},
|
||||
};
|
||||
|
||||
// TODO: Replace with `IoView` projections once available.
|
||||
pub(super) mod gsp_mem {
|
||||
use core::sync::atomic::{
|
||||
fence,
|
||||
Ordering, //
|
||||
};
|
||||
|
||||
use kernel::{
|
||||
dma::Coherent,
|
||||
dma_read,
|
||||
dma_write, //
|
||||
};
|
||||
|
||||
use crate::gsp::cmdq::{
|
||||
GspMem,
|
||||
MSGQ_NUM_PAGES, //
|
||||
};
|
||||
|
||||
pub(in crate::gsp) fn gsp_write_ptr(qs: &Coherent<GspMem>) -> u32 {
|
||||
dma_read!(qs, .gspq.tx.0.writePtr) % MSGQ_NUM_PAGES
|
||||
}
|
||||
|
||||
pub(in crate::gsp) fn gsp_read_ptr(qs: &Coherent<GspMem>) -> u32 {
|
||||
dma_read!(qs, .gspq.rx.0.readPtr) % MSGQ_NUM_PAGES
|
||||
}
|
||||
|
||||
pub(in crate::gsp) fn cpu_read_ptr(qs: &Coherent<GspMem>) -> u32 {
|
||||
dma_read!(qs, .cpuq.rx.0.readPtr) % MSGQ_NUM_PAGES
|
||||
}
|
||||
|
||||
pub(in crate::gsp) fn advance_cpu_read_ptr(qs: &Coherent<GspMem>, count: u32) {
|
||||
let rptr = cpu_read_ptr(qs).wrapping_add(count) % MSGQ_NUM_PAGES;
|
||||
|
||||
// Ensure read pointer is properly ordered.
|
||||
fence(Ordering::SeqCst);
|
||||
|
||||
dma_write!(qs, .cpuq.rx.0.readPtr, rptr);
|
||||
}
|
||||
|
||||
pub(in crate::gsp) fn cpu_write_ptr(qs: &Coherent<GspMem>) -> u32 {
|
||||
dma_read!(qs, .cpuq.tx.0.writePtr) % MSGQ_NUM_PAGES
|
||||
}
|
||||
|
||||
pub(in crate::gsp) fn advance_cpu_write_ptr(qs: &Coherent<GspMem>, count: u32) {
|
||||
let wptr = cpu_write_ptr(qs).wrapping_add(count) % MSGQ_NUM_PAGES;
|
||||
|
||||
dma_write!(qs, .cpuq.tx.0.writePtr, wptr);
|
||||
|
||||
// Ensure all command data is visible before triggering the GSP read.
|
||||
fence(Ordering::SeqCst);
|
||||
}
|
||||
}
|
||||
|
||||
/// Maximum size of a single GSP message queue element in bytes.
|
||||
pub(crate) const GSP_MSG_QUEUE_ELEMENT_SIZE_MAX: usize =
|
||||
num::u32_as_usize(bindings::GSP_MSG_QUEUE_ELEMENT_SIZE_MAX);
|
||||
|
|
@ -177,6 +135,11 @@ pub(crate) fn from_chipset(chipset: Chipset) -> &'static LibosParams {
|
|||
}
|
||||
}
|
||||
|
||||
/// Returns the WPR heap size to reserve when vGPU is enabled.
|
||||
pub(crate) fn vgpu_wpr_heap_size() -> u64 {
|
||||
u64::from(bindings::GSP_FW_HEAP_SIZE_VGPU_DEFAULT)
|
||||
}
|
||||
|
||||
/// Returns the amount of memory (in bytes) to allocate for the WPR heap for a framebuffer size
|
||||
/// of `fb_size` (in bytes) for `chipset`.
|
||||
pub(crate) fn wpr_heap_size(&self, chipset: Chipset, fb_size: u64) -> Result<u64> {
|
||||
|
|
@ -214,48 +177,89 @@ unsafe impl FromBytes for GspFwWprMeta {}
|
|||
|
||||
impl GspFwWprMeta {
|
||||
/// Returns an initializer for a `GspFwWprMeta` suitable for booting `gsp_firmware` using the
|
||||
/// `fb_layout` layout.
|
||||
pub(crate) fn new<'a>(
|
||||
/// framebuffer ranges `ranges`.
|
||||
pub(crate) fn from_ranges<'a>(
|
||||
gsp_firmware: &'a GspFirmware,
|
||||
fb_layout: &'a FbLayout,
|
||||
ranges: &'a FbRanges,
|
||||
) -> impl Init<Self> + 'a {
|
||||
#[allow(non_snake_case)]
|
||||
let init_inner = init!(bindings::GspFwWprMeta {
|
||||
// CAST: we want to store the bits of `GSP_FW_WPR_META_MAGIC` unmodified.
|
||||
magic: bindings::GSP_FW_WPR_META_MAGIC as u64,
|
||||
revision: u64::from(bindings::GSP_FW_WPR_META_REVISION),
|
||||
sysmemAddrOfRadix3Elf: gsp_firmware.radix3_dma_handle(),
|
||||
sysmemAddrOfRadix3Elf: gsp_firmware.radix3_dma_address(),
|
||||
sizeOfRadix3Elf: u64::from_safe_cast(gsp_firmware.size),
|
||||
sysmemAddrOfBootloader: gsp_firmware.bootloader.ucode.dma_handle(),
|
||||
sysmemAddrOfBootloader: gsp_firmware.bootloader.ucode.dma_address(),
|
||||
sizeOfBootloader: u64::from_safe_cast(gsp_firmware.bootloader.ucode.size()),
|
||||
bootloaderCodeOffset: u64::from(gsp_firmware.bootloader.code_offset),
|
||||
bootloaderDataOffset: u64::from(gsp_firmware.bootloader.data_offset),
|
||||
bootloaderManifestOffset: u64::from(gsp_firmware.bootloader.manifest_offset),
|
||||
__bindgen_anon_1: GspFwWprMetaBootResumeInfo {
|
||||
__bindgen_anon_1: GspFwWprMetaBootInfo {
|
||||
sysmemAddrOfSignature: gsp_firmware.signatures.dma_handle(),
|
||||
sysmemAddrOfSignature: gsp_firmware.signatures.dma_address(),
|
||||
sizeOfSignature: u64::from_safe_cast(gsp_firmware.signatures.size()),
|
||||
},
|
||||
},
|
||||
gspFwRsvdStart: fb_layout.heap.start,
|
||||
nonWprHeapOffset: fb_layout.heap.start,
|
||||
nonWprHeapSize: fb_layout.heap.end - fb_layout.heap.start,
|
||||
gspFwWprStart: fb_layout.wpr2.start,
|
||||
gspFwHeapOffset: fb_layout.wpr2_heap.start,
|
||||
gspFwHeapSize: fb_layout.wpr2_heap.end - fb_layout.wpr2_heap.start,
|
||||
gspFwOffset: fb_layout.elf.start,
|
||||
bootBinOffset: fb_layout.boot.start,
|
||||
frtsOffset: fb_layout.frts.start,
|
||||
frtsSize: fb_layout.frts.end - fb_layout.frts.start,
|
||||
gspFwWprEnd: fb_layout
|
||||
gspFwRsvdStart: ranges.non_wpr_heap.start,
|
||||
nonWprHeapOffset: ranges.non_wpr_heap.start,
|
||||
nonWprHeapSize: ranges.non_wpr_heap.len(),
|
||||
gspFwWprStart: ranges.wpr2.start,
|
||||
gspFwHeapOffset: ranges.wpr2_heap.start,
|
||||
gspFwHeapSize: ranges.wpr2_heap.len(),
|
||||
gspFwOffset: ranges.elf.start,
|
||||
bootBinOffset: ranges.boot.start,
|
||||
frtsOffset: ranges.frts.start,
|
||||
frtsSize: ranges.frts.len(),
|
||||
gspFwWprEnd: ranges
|
||||
.vga_workspace
|
||||
.start
|
||||
.align_down(Alignment::new::<SZ_128K>()),
|
||||
gspFwHeapVfPartitionCount: fb_layout.vf_partition_count,
|
||||
fbSize: fb_layout.fb.end - fb_layout.fb.start,
|
||||
vgaWorkspaceOffset: fb_layout.vga_workspace.start,
|
||||
vgaWorkspaceSize: fb_layout.vga_workspace.end - fb_layout.vga_workspace.start,
|
||||
pmuReservedSize: fb_layout.pmu_reserved_size,
|
||||
gspFwHeapVfPartitionCount: ranges.vf_partition_count,
|
||||
fbSize: ranges.fb.len(),
|
||||
vgaWorkspaceOffset: ranges.vga_workspace.start,
|
||||
vgaWorkspaceSize: ranges.vga_workspace.len(),
|
||||
pmuReservedSize: ranges.pmu_reserved_size,
|
||||
..Zeroable::init_zeroed()
|
||||
});
|
||||
|
||||
init!(GspFwWprMeta {
|
||||
inner <- init_inner,
|
||||
})
|
||||
}
|
||||
|
||||
/// Returns an initializer for a `GspFwWprMeta` suitable for booting `gsp_firmware` using the
|
||||
/// framebuffer region sizes `sizes`.
|
||||
///
|
||||
/// The region offsets are left at zero: the ACR ucode computes them when it sets up WPR2.
|
||||
pub(crate) fn from_sizes<'a>(
|
||||
gsp_firmware: &'a GspFirmware,
|
||||
sizes: &'a FbSizes,
|
||||
) -> impl Init<Self> + 'a {
|
||||
/// VGA workspace size to reserve at the end of the framebuffer, in bytes.
|
||||
const VGA_WORKSPACE_SIZE: u64 = u64::SZ_128K;
|
||||
|
||||
let init_inner = init!(bindings::GspFwWprMeta {
|
||||
// CAST: we want to store the bits of `GSP_FW_WPR_META_MAGIC` unmodified.
|
||||
magic: bindings::GSP_FW_WPR_META_MAGIC as u64,
|
||||
revision: u64::from(bindings::GSP_FW_WPR_META_REVISION),
|
||||
sysmemAddrOfRadix3Elf: gsp_firmware.radix3_dma_address(),
|
||||
sizeOfRadix3Elf: u64::from_safe_cast(gsp_firmware.size),
|
||||
sysmemAddrOfBootloader: gsp_firmware.bootloader.ucode.dma_address(),
|
||||
sizeOfBootloader: u64::from_safe_cast(gsp_firmware.bootloader.ucode.size()),
|
||||
bootloaderCodeOffset: u64::from(gsp_firmware.bootloader.code_offset),
|
||||
bootloaderDataOffset: u64::from(gsp_firmware.bootloader.data_offset),
|
||||
bootloaderManifestOffset: u64::from(gsp_firmware.bootloader.manifest_offset),
|
||||
__bindgen_anon_1: GspFwWprMetaBootResumeInfo {
|
||||
__bindgen_anon_1: GspFwWprMetaBootInfo {
|
||||
sysmemAddrOfSignature: gsp_firmware.signatures.dma_address(),
|
||||
sizeOfSignature: u64::from_safe_cast(gsp_firmware.signatures.size()),
|
||||
},
|
||||
},
|
||||
nonWprHeapSize: sizes.non_wpr_heap_size,
|
||||
gspFwHeapSize: sizes.wpr2_heap_size,
|
||||
frtsSize: sizes.frts_size,
|
||||
gspFwHeapVfPartitionCount: sizes.vf_partition_count,
|
||||
vgaWorkspaceSize: VGA_WORKSPACE_SIZE,
|
||||
pmuReservedSize: sizes.pmu_reserved_size,
|
||||
..Zeroable::init_zeroed()
|
||||
});
|
||||
|
||||
|
|
@ -674,10 +678,9 @@ fn id8(name: &str) -> u64 {
|
|||
u64::from_ne_bytes(bytes)
|
||||
}
|
||||
|
||||
#[allow(non_snake_case)]
|
||||
let init_inner = init!(bindings::LibosMemoryRegionInitArgument {
|
||||
id8: id8(name),
|
||||
pa: obj.dma_handle(),
|
||||
pa: obj.dma_address(),
|
||||
size: num::usize_as_u64(obj.size()),
|
||||
kind: num::u32_into_u8::<
|
||||
{ bindings::LibosMemoryRegionKind_LIBOS_MEMORY_REGION_CONTIGUOUS },
|
||||
|
|
@ -720,6 +723,16 @@ pub(crate) fn new(msgq_size: u32, rx_hdr_offset: u32, msg_count: u32) -> Self {
|
|||
entryOff: num::usize_into_u32::<GSP_PAGE_SIZE>(),
|
||||
})
|
||||
}
|
||||
|
||||
/// Returns the value of the write pointer for this queue.
|
||||
pub(crate) fn write_ptr(this: CoherentView<'_, Self>) -> u32 {
|
||||
io_read!(this, .0.writePtr)
|
||||
}
|
||||
|
||||
/// Sets the value of the write pointer for this queue.
|
||||
pub(crate) fn set_write_ptr(this: CoherentView<'_, Self>, val: u32) {
|
||||
io_write!(this, .0.writePtr, val)
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: Padding is explicit and does not contain uninitialized data.
|
||||
|
|
@ -735,6 +748,16 @@ impl MsgqRxHeader {
|
|||
pub(crate) fn new() -> Self {
|
||||
Self(Default::default())
|
||||
}
|
||||
|
||||
/// Returns the value of the read pointer for this queue.
|
||||
pub(crate) fn read_ptr(this: CoherentView<'_, Self>) -> u32 {
|
||||
io_read!(this, .0.readPtr)
|
||||
}
|
||||
|
||||
/// Sets the value of the read pointer for this queue.
|
||||
pub(crate) fn set_read_ptr(this: CoherentView<'_, Self>, val: u32) {
|
||||
io_write!(this, .0.readPtr, val)
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: Padding is explicit and does not contain uninitialized data.
|
||||
|
|
@ -742,8 +765,8 @@ unsafe impl AsBytes for MsgqRxHeader {}
|
|||
|
||||
bitfield! {
|
||||
struct MsgHeaderVersion(u32) {
|
||||
31:24 major as u8;
|
||||
23:16 minor as u8;
|
||||
31:24 major;
|
||||
23:16 minor;
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -752,9 +775,9 @@ impl MsgHeaderVersion {
|
|||
const MINOR_TOT: u8 = 0;
|
||||
|
||||
fn new() -> Self {
|
||||
Self::default()
|
||||
.set_major(Self::MAJOR_TOT)
|
||||
.set_minor(Self::MINOR_TOT)
|
||||
Self::zeroed()
|
||||
.with_major(Self::MAJOR_TOT)
|
||||
.with_minor(Self::MINOR_TOT)
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -793,7 +816,6 @@ impl GspMsgElement {
|
|||
/// * `sequence` - Sequence number of the message.
|
||||
/// * `cmd_size` - Size of the command (not including the message element), in bytes.
|
||||
/// * `function` - Function of the message.
|
||||
#[allow(non_snake_case)]
|
||||
pub(crate) fn init(
|
||||
sequence: u32,
|
||||
cmd_size: usize,
|
||||
|
|
@ -876,7 +898,6 @@ pub(crate) struct GspArgumentsCached {
|
|||
impl GspArgumentsCached {
|
||||
/// Creates the arguments for starting the GSP up using `cmdq` as its command queue.
|
||||
pub(crate) fn new(cmdq: &Cmdq) -> impl Init<Self> + '_ {
|
||||
#[allow(non_snake_case)]
|
||||
let init_inner = init!(bindings::GSP_ARGUMENTS_CACHED {
|
||||
messageQueueInitArguments <- MessageQueueInitArguments::new(cmdq),
|
||||
bDmemStack: 1,
|
||||
|
|
@ -923,10 +944,9 @@ unsafe impl FromBytes for GspArgumentsPadded {}
|
|||
|
||||
impl MessageQueueInitArguments {
|
||||
/// Creates a new init arguments structure for `cmdq`.
|
||||
#[allow(non_snake_case)]
|
||||
fn new(cmdq: &Cmdq) -> impl Init<Self> + '_ {
|
||||
init!(MessageQueueInitArguments {
|
||||
sharedMemPhysAddr: cmdq.dma_handle,
|
||||
sharedMemPhysAddr: cmdq.dma_addr,
|
||||
pageTableEntryCount: num::usize_into_u32::<{ Cmdq::NUM_PTES }>(),
|
||||
cmdQueueOffset: num::usize_as_u64(Cmdq::CMDQ_OFFSET),
|
||||
statQueueOffset: num::usize_as_u64(Cmdq::STATQ_OFFSET),
|
||||
|
|
@ -947,7 +967,6 @@ pub(crate) enum GspDmaTarget {
|
|||
|
||||
impl GspAcrBootGspRmParams {
|
||||
fn new(target: GspDmaTarget, wpr_meta_addr: u64) -> impl Init<Self> {
|
||||
#[allow(non_snake_case)]
|
||||
let params = init!(Self {
|
||||
target: target as u32,
|
||||
gspRmDescSize: num::usize_into_u32::<{ size_of::<GspFwWprMeta>() }>(),
|
||||
|
|
@ -966,7 +985,6 @@ fn new(target: GspDmaTarget, wpr_meta_addr: u64) -> impl Init<Self> {
|
|||
|
||||
impl GspRmParams {
|
||||
fn new(target: GspDmaTarget, libos_addr: u64) -> impl Init<Self> {
|
||||
#[allow(non_snake_case)]
|
||||
let params = init!(Self {
|
||||
target: target as u32,
|
||||
bootArgsOffset: libos_addr,
|
||||
|
|
@ -986,7 +1004,6 @@ unsafe impl FromBytes for GspFmcBootParams {}
|
|||
|
||||
impl GspFmcBootParams {
|
||||
pub(crate) fn new(wpr_meta_addr: u64, libos_addr: u64) -> impl Init<Self> {
|
||||
#[allow(non_snake_case)]
|
||||
let init = init!(Self {
|
||||
// Blackwell FSP obtains WPR info from other sources, so
|
||||
// wprCarveoutOffset and wprCarveoutSize are left zero.
|
||||
|
|
|
|||
|
|
@ -1,6 +1,8 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use core::ops::Range;
|
||||
|
||||
use kernel::{
|
||||
device,
|
||||
pci,
|
||||
|
|
@ -13,7 +15,8 @@
|
|||
|
||||
use crate::{
|
||||
gpu::Chipset,
|
||||
gsp::GSP_PAGE_SIZE, //
|
||||
gsp::GSP_PAGE_SIZE,
|
||||
num::IntoSafeCast, //
|
||||
};
|
||||
|
||||
use super::bindings;
|
||||
|
|
@ -27,7 +30,6 @@ pub(crate) struct GspSetSystemInfo {
|
|||
|
||||
impl GspSetSystemInfo {
|
||||
/// Returns an in-place initializer for the `GspSetSystemInfo` command.
|
||||
#[allow(non_snake_case)]
|
||||
pub(crate) fn init<'a>(
|
||||
dev: &'a pci::Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
|
|
@ -99,7 +101,6 @@ pub(crate) struct PackedRegistryTable {
|
|||
}
|
||||
|
||||
impl PackedRegistryTable {
|
||||
#[allow(non_snake_case)]
|
||||
pub(crate) fn init(num_entries: u32, size: u32) -> impl Init<Self> {
|
||||
type InnerPackedRegistryTable = bindings::PACKED_REGISTRY_TABLE;
|
||||
let init_inner = init!(InnerPackedRegistryTable {
|
||||
|
|
@ -129,6 +130,41 @@ impl GspStaticConfigInfo {
|
|||
pub(crate) fn gpu_name_str(&self) -> [u8; 64] {
|
||||
self.0.gpuNameString
|
||||
}
|
||||
|
||||
/// Returns an iterator over valid FB regions from GSP firmware data.
|
||||
fn fb_regions(
|
||||
&self,
|
||||
) -> impl Iterator<Item = &bindings::NV2080_CTRL_CMD_FB_GET_FB_REGION_FB_REGION_INFO> {
|
||||
let fb_info = &self.0.fbRegionInfoParams;
|
||||
fb_info
|
||||
.fbRegion
|
||||
.iter()
|
||||
.take(fb_info.numFBRegions.into_safe_cast())
|
||||
.filter(|reg| reg.limit >= reg.base)
|
||||
}
|
||||
|
||||
/// Iterates over usable FB regions from GSP firmware data.
|
||||
///
|
||||
/// Each yielded region is a [`Range<u64>`] suitable for driver memory allocation.
|
||||
/// Usable regions are those that satisfy all the following properties:
|
||||
/// - Are not reserved for firmware internal use.
|
||||
/// - Are not protected (hardware-enforced access restrictions).
|
||||
/// - Support compression (can use GPU memory compression for bandwidth).
|
||||
/// - Support ISO (isochronous memory for display requiring guaranteed bandwidth).
|
||||
pub(crate) fn usable_fb_regions(&self) -> impl Iterator<Item = Range<u64>> + '_ {
|
||||
self.fb_regions().filter_map(|reg| {
|
||||
// Filter: not reserved, not protected, supports compression and ISO.
|
||||
if reg.reserved == 0
|
||||
&& reg.bProtected == 0
|
||||
&& reg.supportCompressed != 0
|
||||
&& reg.supportISO != 0
|
||||
{
|
||||
reg.limit.checked_add(1).map(|end| reg.base..end)
|
||||
} else {
|
||||
None
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: Padding is explicit and will not contain uninitialized data.
|
||||
|
|
|
|||
|
|
@ -40,6 +40,7 @@ fn fmt(&self, fmt: &mut ::core::fmt::Formatter<'_>) -> ::core::fmt::Result {
|
|||
pub const GSP_FW_HEAP_PARAM_BASE_RM_SIZE_GH100: u32 = 14680064;
|
||||
pub const GSP_FW_HEAP_PARAM_SIZE_PER_GB_FB: u32 = 98304;
|
||||
pub const GSP_FW_HEAP_PARAM_CLIENT_ALLOC_SIZE: u32 = 100663296;
|
||||
pub const GSP_FW_HEAP_SIZE_VGPU_DEFAULT: u32 = 609222656;
|
||||
pub const GSP_FW_HEAP_SIZE_OVERRIDE_LIBOS2_MIN_MB: u32 = 64;
|
||||
pub const GSP_FW_HEAP_SIZE_OVERRIDE_LIBOS2_MAX_MB: u32 = 256;
|
||||
pub const GSP_FW_HEAP_SIZE_OVERRIDE_LIBOS3_BAREMETAL_MIN_MB: u32 = 88;
|
||||
|
|
|
|||
|
|
@ -1,33 +1,21 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
mod ga102;
|
||||
mod gh100;
|
||||
mod tu102;
|
||||
|
||||
use kernel::prelude::*;
|
||||
|
||||
use kernel::{
|
||||
device,
|
||||
dma::Coherent, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
gsp::Gsp as GspEngine,
|
||||
sec2::Sec2,
|
||||
Falcon, //
|
||||
},
|
||||
fb::FbLayout,
|
||||
firmware::gsp::GspFirmware,
|
||||
gpu::{
|
||||
Architecture,
|
||||
Chipset, //
|
||||
},
|
||||
gsp::{
|
||||
boot::BootUnloadGuard,
|
||||
Gsp,
|
||||
GspFwWprMeta, //
|
||||
GspBootContext, //
|
||||
},
|
||||
};
|
||||
|
||||
|
|
@ -38,33 +26,21 @@
|
|||
/// required for unloading is prepared at load time, and stored here until it needs to be run.
|
||||
pub(super) trait UnloadBundle: Send {
|
||||
/// Performs the steps required to properly reset the GSP after it has been stopped.
|
||||
fn run(
|
||||
&self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
sec2_falcon: &Falcon<Sec2>,
|
||||
) -> Result;
|
||||
fn run(&self, ctx: &mut GspBootContext<'_, '_>) -> Result;
|
||||
}
|
||||
|
||||
/// Trait implemented by GSP HALs.
|
||||
pub(super) trait GspHal: Send {
|
||||
/// Performs the GSP boot process, loading and running the required firmwares as needed.
|
||||
///
|
||||
/// Upon success, returns a guard that runs the GSP unload sequence if GSP boot does not
|
||||
/// complete.
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
fn boot<'a>(
|
||||
/// Upon success, returns the [`crate::gsp::UnloadBundle`] to use with [`Gsp::unload`], if one
|
||||
/// could be created.
|
||||
fn boot(
|
||||
&self,
|
||||
gsp: &'a Gsp,
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
bar: Bar0<'a>,
|
||||
chipset: Chipset,
|
||||
fb_layout: &FbLayout,
|
||||
wpr_meta: &Coherent<GspFwWprMeta>,
|
||||
gsp_falcon: &'a Falcon<GspEngine>,
|
||||
sec2_falcon: &'a Falcon<Sec2>,
|
||||
) -> Result<BootUnloadGuard<'a>>;
|
||||
gsp: &Gsp,
|
||||
ctx: &mut GspBootContext<'_, '_>,
|
||||
gsp_fw: &GspFirmware,
|
||||
) -> Result<Option<crate::gsp::UnloadBundle>>;
|
||||
|
||||
/// Performs HAL-specific post-GSP boot tasks.
|
||||
///
|
||||
|
|
@ -73,20 +49,38 @@ fn boot<'a>(
|
|||
fn post_boot(
|
||||
&self,
|
||||
_gsp: &Gsp,
|
||||
_dev: &device::Device<device::Bound>,
|
||||
_bar: Bar0<'_>,
|
||||
_ctx: &mut GspBootContext<'_, '_>,
|
||||
_gsp_fw: &GspFirmware,
|
||||
_gsp_falcon: &Falcon<GspEngine>,
|
||||
_sec2_falcon: &Falcon<Sec2>,
|
||||
) -> Result {
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
/// Returns the names of the firmware files required to boot the GSP of `chipset`, in addition to
|
||||
/// the "bootloader" and "gsp" images required by all chipsets.
|
||||
pub(crate) const fn boot_firmware_files(chipset: Chipset) -> &'static [&'static str] {
|
||||
match chipset.arch() {
|
||||
// Turing chipsets boot the GSP via the SEC2 Booter, and require the FWSEC bootloader.
|
||||
Architecture::Turing => &["booter_load.tlv", "booter_unload.tlv", "gen_bootloader.tlv"],
|
||||
// GA100 also requires the FWSEC bootloader.
|
||||
Architecture::Ampere if matches!(chipset, Chipset::GA100) => {
|
||||
&["booter_load.tlv", "booter_unload.tlv", "gen_bootloader.tlv"]
|
||||
}
|
||||
// Other Ampere chipsets, as well as Ada chipsets, run FWSEC directly.
|
||||
Architecture::Ampere | Architecture::Ada => &["booter_load.tlv", "booter_unload.tlv"],
|
||||
// Hopper and later chipsets boot the GSP via the FMC image loaded by FSP.
|
||||
Architecture::Hopper | Architecture::BlackwellGB10x | Architecture::BlackwellGB20x => {
|
||||
&["fmc.tlv"]
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Returns the GSP HAL to be used for `chipset`.
|
||||
pub(super) fn gsp_hal(chipset: Chipset) -> &'static dyn GspHal {
|
||||
match chipset.arch() {
|
||||
Architecture::Turing | Architecture::Ampere | Architecture::Ada => tu102::TU102_HAL,
|
||||
Architecture::Turing => tu102::TU102_HAL,
|
||||
Architecture::Ampere if matches!(chipset, Chipset::GA100) => tu102::TU102_HAL,
|
||||
Architecture::Ampere | Architecture::Ada => ga102::GA102_HAL,
|
||||
Architecture::Hopper | Architecture::BlackwellGB10x | Architecture::BlackwellGB20x => {
|
||||
gh100::GH100_HAL
|
||||
}
|
||||
|
|
|
|||
14
drivers/gpu/nova-core/gsp/hal/ga102.rs
Normal file
14
drivers/gpu/nova-core/gsp/hal/ga102.rs
Normal file
|
|
@ -0,0 +1,14 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use crate::gsp::hal::{
|
||||
tu102::Tu102,
|
||||
GspHal, //
|
||||
};
|
||||
|
||||
/// The GA102 HAL is like the TU102 one, except it doesn't use the bootloader.
|
||||
const GA102: Tu102 = Tu102 {
|
||||
needs_fwsec_bootloader: false,
|
||||
};
|
||||
|
||||
pub(super) const GA102_HAL: &dyn GspHal = &GA102;
|
||||
|
|
@ -7,33 +7,26 @@
|
|||
device,
|
||||
dma::Coherent,
|
||||
io::poll::read_poll_timeout,
|
||||
time::Delta, //
|
||||
time::Delta,
|
||||
types::ScopeGuard, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
falcon::{
|
||||
gsp::Gsp as GspEngine,
|
||||
sec2::Sec2,
|
||||
Falcon, //
|
||||
},
|
||||
fb::FbLayout,
|
||||
firmware::{
|
||||
fsp::FspFirmware,
|
||||
FIRMWARE_VERSION, //
|
||||
},
|
||||
fsp::{
|
||||
FmcBootArgs,
|
||||
Fsp, //
|
||||
},
|
||||
gpu::Chipset,
|
||||
fb::FbSizes,
|
||||
firmware::gsp::GspFirmware,
|
||||
fsp::FmcBootArgs,
|
||||
gsp::{
|
||||
boot::BootUnloadGuard,
|
||||
hal::{
|
||||
GspHal,
|
||||
UnloadBundle, //
|
||||
},
|
||||
Gsp,
|
||||
GspBootContext,
|
||||
GspFmcBootParams,
|
||||
GspFwWprMeta, //
|
||||
},
|
||||
};
|
||||
|
|
@ -46,10 +39,10 @@ struct GspMbox {
|
|||
|
||||
impl GspMbox {
|
||||
/// Reads both mailboxes from the GSP falcon.
|
||||
fn read(gsp_falcon: &Falcon<GspEngine>, bar: Bar0<'_>) -> Self {
|
||||
fn read(gsp_falcon: &Falcon<'_, GspEngine>) -> Self {
|
||||
Self {
|
||||
mbox0: gsp_falcon.read_mailbox0(bar),
|
||||
mbox1: gsp_falcon.read_mailbox1(bar),
|
||||
mbox0: gsp_falcon.read_mailbox0(),
|
||||
mbox1: gsp_falcon.read_mailbox1(),
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -64,27 +57,25 @@ fn combined_addr(&self) -> u64 {
|
|||
/// either condition should stop the poll loop.
|
||||
fn lockdown_released_or_error(
|
||||
&self,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
bar: Bar0<'_>,
|
||||
fmc_boot_params_addr: u64,
|
||||
gsp_falcon: &Falcon<'_, GspEngine>,
|
||||
fmc_boot_params: &Coherent<GspFmcBootParams>,
|
||||
) -> bool {
|
||||
// GSP-FMC normally clears the boot parameters address from the mailboxes early during
|
||||
// boot. If the address is still there, keep polling rather than treating it as an error.
|
||||
// Any other non-zero mailbox0 value is a GSP-FMC error code.
|
||||
if self.mbox0 != 0 {
|
||||
return self.combined_addr() != fmc_boot_params_addr;
|
||||
return self.combined_addr() != fmc_boot_params.dma_address();
|
||||
}
|
||||
|
||||
!gsp_falcon.riscv_branch_privilege_lockdown(bar)
|
||||
!gsp_falcon.riscv_branch_privilege_lockdown()
|
||||
}
|
||||
}
|
||||
|
||||
/// Waits for GSP lockdown to be released after FSP Chain of Trust.
|
||||
fn wait_for_gsp_lockdown_release(
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
fmc_boot_params_addr: u64,
|
||||
gsp_falcon: &Falcon<'_, GspEngine>,
|
||||
fmc_boot_params: &Coherent<GspFmcBootParams>,
|
||||
) -> Result {
|
||||
dev_dbg!(dev, "Waiting for GSP lockdown release\n");
|
||||
|
||||
|
|
@ -92,14 +83,14 @@ fn wait_for_gsp_lockdown_release(
|
|||
|| {
|
||||
// While the PRIV target mask is still locked to FSP, GSP register and mailbox reads
|
||||
// are not meaningful. Wait until HWCFG2 says the CPU can read them.
|
||||
Ok(match gsp_falcon.priv_target_mask_released(bar) {
|
||||
Ok(match gsp_falcon.priv_target_mask_released() {
|
||||
false => None,
|
||||
true => Some(GspMbox::read(gsp_falcon, bar)),
|
||||
true => Some(GspMbox::read(gsp_falcon)),
|
||||
})
|
||||
},
|
||||
|mbox| match mbox {
|
||||
None => false,
|
||||
Some(mbox) => mbox.lockdown_released_or_error(gsp_falcon, bar, fmc_boot_params_addr),
|
||||
Some(mbox) => mbox.lockdown_released_or_error(gsp_falcon, fmc_boot_params),
|
||||
},
|
||||
Delta::from_millis(10),
|
||||
Delta::from_secs(30),
|
||||
|
|
@ -123,22 +114,23 @@ fn wait_for_gsp_lockdown_release(
|
|||
struct FspUnloadBundle;
|
||||
|
||||
impl UnloadBundle for FspUnloadBundle {
|
||||
fn run(
|
||||
&self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
_sec2_falcon: &Falcon<Sec2>,
|
||||
) -> Result {
|
||||
fn run(&self, ctx: &mut GspBootContext<'_, '_>) -> Result {
|
||||
// GSP falcon does most of the work of resetting, so just wait for it to finish.
|
||||
read_poll_timeout(
|
||||
|| Ok(gsp_falcon.is_riscv_active(bar)),
|
||||
|&active| !active,
|
||||
|| {
|
||||
// GSP register reads are not meaningful until the PRIV target mask is released.
|
||||
if !ctx.gsp_falcon.priv_target_mask_released() {
|
||||
return Ok(false);
|
||||
}
|
||||
|
||||
ctx.gsp_falcon.is_riscv_halted()
|
||||
},
|
||||
|&halted| halted,
|
||||
Delta::from_millis(10),
|
||||
Delta::from_secs(5),
|
||||
)
|
||||
.map(|_| ())
|
||||
.inspect_err(|_| dev_err!(dev, "GSP falcon failed to halt\n"))
|
||||
.inspect_err(|_| dev_err!(ctx.dev(), "GSP falcon failed to halt\n"))
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -149,42 +141,44 @@ impl GspHal for Gh100 {
|
|||
///
|
||||
/// This path uses FSP to establish a chain of trust and boot GSP-FMC. FSP handles
|
||||
/// the GSP boot internally - no manual GSP reset/boot is needed.
|
||||
fn boot<'a>(
|
||||
fn boot(
|
||||
&self,
|
||||
gsp: &'a Gsp,
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
bar: Bar0<'a>,
|
||||
chipset: Chipset,
|
||||
fb_layout: &FbLayout,
|
||||
wpr_meta: &Coherent<GspFwWprMeta>,
|
||||
gsp_falcon: &'a Falcon<GspEngine>,
|
||||
sec2_falcon: &'a Falcon<Sec2>,
|
||||
) -> Result<BootUnloadGuard<'a>> {
|
||||
let fsp_fw = FspFirmware::new(dev, chipset, FIRMWARE_VERSION)?;
|
||||
gsp: &Gsp,
|
||||
ctx: &mut GspBootContext<'_, '_>,
|
||||
gsp_fw: &GspFirmware,
|
||||
) -> Result<Option<crate::gsp::UnloadBundle>> {
|
||||
let dev = ctx.dev();
|
||||
let chipset = ctx.chipset;
|
||||
let gsp_falcon = ctx.gsp_falcon;
|
||||
|
||||
let fb_sizes = FbSizes::new(chipset, ctx.bar, ctx.vgpu.state())?;
|
||||
dev_dbg!(dev, "{:#x?}\n", fb_sizes);
|
||||
|
||||
let wpr_meta =
|
||||
Coherent::init(dev, GFP_KERNEL, GspFwWprMeta::from_sizes(gsp_fw, &fb_sizes))?;
|
||||
let args = FmcBootArgs::new(dev, chipset, wpr_meta, &gsp.libos, false)?;
|
||||
|
||||
let unload_bundle = crate::gsp::UnloadBundle(
|
||||
KBox::new(FspUnloadBundle, GFP_KERNEL)? as KBox<dyn UnloadBundle>
|
||||
);
|
||||
|
||||
// Wrap the unload bundle into a drop guard so it is automatically run upon failure.
|
||||
let unload_guard =
|
||||
BootUnloadGuard::new(gsp, dev, bar, gsp_falcon, sec2_falcon, Some(unload_bundle));
|
||||
// Wait for the GSP RISC-V core to halt in case of error. We create this guard after `args`
|
||||
// to make sure that the boot args and the WPR metadata they own are kept alive until halt,
|
||||
// in case they are still being accessed.
|
||||
let mut unload_guard =
|
||||
ScopeGuard::new_with_data((unload_bundle, ctx), |(unload_bundle, ctx)| {
|
||||
let _ = unload_bundle.0.run(ctx);
|
||||
});
|
||||
|
||||
let mut fsp = Fsp::wait_secure_boot(dev, bar, chipset, fsp_fw)?;
|
||||
let fsp = unload_guard.1.fsp.as_mut().ok_or(ENODEV)?;
|
||||
|
||||
let args = FmcBootArgs::new(
|
||||
dev,
|
||||
chipset,
|
||||
wpr_meta.dma_handle(),
|
||||
gsp.libos.dma_handle(),
|
||||
false,
|
||||
)?;
|
||||
fsp.boot_fmc(dev, &fb_sizes, &args)?;
|
||||
|
||||
fsp.boot_fmc(dev, bar, fb_layout, &args)?;
|
||||
// Wait for GSP-FMC to release the GSP lockdown, indicating that `args` is not accessed
|
||||
// anymore.
|
||||
wait_for_gsp_lockdown_release(dev, gsp_falcon, args.boot_params())?;
|
||||
|
||||
wait_for_gsp_lockdown_release(dev, bar, gsp_falcon, args.boot_params_dma_handle())?;
|
||||
|
||||
Ok(unload_guard)
|
||||
Ok(Some(unload_guard.dismiss().0))
|
||||
}
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -6,7 +6,8 @@
|
|||
use kernel::{
|
||||
device,
|
||||
dma::Coherent,
|
||||
io::Io, //
|
||||
io::Io,
|
||||
types::ScopeGuard, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
|
|
@ -16,7 +17,10 @@
|
|||
sec2::Sec2,
|
||||
Falcon, //
|
||||
},
|
||||
fb::FbLayout,
|
||||
fb::{
|
||||
wpr2_range,
|
||||
FbRanges, //
|
||||
},
|
||||
firmware::{
|
||||
booter::{
|
||||
BooterFirmware,
|
||||
|
|
@ -27,24 +31,20 @@
|
|||
FwsecCommand,
|
||||
FwsecFirmware, //
|
||||
},
|
||||
gsp::GspFirmware,
|
||||
FIRMWARE_VERSION, //
|
||||
gsp::GspFirmware, //
|
||||
},
|
||||
gpu::Chipset,
|
||||
gsp::{
|
||||
boot::BootUnloadGuard,
|
||||
hal::{
|
||||
GspHal,
|
||||
UnloadBundle, //
|
||||
},
|
||||
sequencer::{
|
||||
GspSequencer,
|
||||
GspSequencerParams, //
|
||||
},
|
||||
regs,
|
||||
sequencer::GspSequencer,
|
||||
Gsp,
|
||||
GspBootContext,
|
||||
GspFwWprMeta, //
|
||||
},
|
||||
regs,
|
||||
vbios::Vbios, //
|
||||
};
|
||||
|
||||
|
|
@ -58,32 +58,15 @@ enum FwsecUnloadFirmware {
|
|||
}
|
||||
|
||||
impl FwsecUnloadFirmware {
|
||||
/// Loads the FWSEC SB firmware, as well as its bootloader if `chipset` requires it.
|
||||
fn new(
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
chipset: Chipset,
|
||||
bios: &Vbios,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
) -> Result<Self> {
|
||||
let fwsec_sb = FwsecFirmware::new(dev, gsp_falcon, bar, bios, FwsecCommand::Sb)?;
|
||||
|
||||
Ok(if chipset.needs_fwsec_bootloader() {
|
||||
Self::WithBl(FwsecFirmwareWithBl::new(fwsec_sb, dev, chipset)?)
|
||||
} else {
|
||||
Self::WithoutBl(fwsec_sb)
|
||||
})
|
||||
}
|
||||
|
||||
/// Runs the FWSEC SB firmware.
|
||||
fn run(
|
||||
&self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
gsp_falcon: &Falcon<'_, GspEngine>,
|
||||
) -> Result {
|
||||
match self {
|
||||
Self::WithoutBl(fw) => fw.run(dev, gsp_falcon, bar),
|
||||
Self::WithoutBl(fw) => fw.run(dev, gsp_falcon),
|
||||
Self::WithBl(fw) => fw.run(dev, gsp_falcon, bar),
|
||||
}
|
||||
}
|
||||
|
|
@ -96,210 +79,225 @@ struct Sec2UnloadBundle {
|
|||
booter_unloader: BooterFirmware,
|
||||
}
|
||||
|
||||
impl Sec2UnloadBundle {
|
||||
/// Load and prepare the resources required to properly reset the GSP after it has been stopped.
|
||||
fn build(
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
chipset: Chipset,
|
||||
bios: &Vbios,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
sec2_falcon: &Falcon<Sec2>,
|
||||
) -> Result<KBox<dyn UnloadBundle>> {
|
||||
KBox::new(
|
||||
Self {
|
||||
fwsec_sb: FwsecUnloadFirmware::new(dev, bar, chipset, bios, gsp_falcon)?,
|
||||
booter_unloader: BooterFirmware::new(
|
||||
dev,
|
||||
BooterKind::Unloader,
|
||||
chipset,
|
||||
FIRMWARE_VERSION,
|
||||
sec2_falcon,
|
||||
bar,
|
||||
)?,
|
||||
},
|
||||
GFP_KERNEL,
|
||||
)
|
||||
.map(|b| b as KBox<dyn UnloadBundle>)
|
||||
.map_err(Into::into)
|
||||
}
|
||||
}
|
||||
|
||||
impl UnloadBundle for Sec2UnloadBundle {
|
||||
fn run(
|
||||
&self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
sec2_falcon: &Falcon<Sec2>,
|
||||
) -> Result {
|
||||
fn run(&self, ctx: &mut GspBootContext<'_, '_>) -> Result {
|
||||
let dev = ctx.dev();
|
||||
let bar = ctx.bar;
|
||||
|
||||
// Run FWSEC-SB to reset the GSP falcon to its pre-libos state.
|
||||
self.fwsec_sb.run(dev, bar, gsp_falcon)?;
|
||||
// Log errors but keep going if it fails.
|
||||
let fwsec_sb_res = self
|
||||
.fwsec_sb
|
||||
.run(dev, bar, ctx.gsp_falcon)
|
||||
.inspect_err(|e| dev_err!(dev, "FWSEC-SB failed to run: {:?}\n", e));
|
||||
|
||||
// Remove WPR2 region if set.
|
||||
let wpr2_hi = bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI);
|
||||
if wpr2_hi.is_wpr2_set() {
|
||||
sec2_falcon.reset(bar)?;
|
||||
sec2_falcon.load(dev, bar, &self.booter_unloader)?;
|
||||
let booter_unloader_res = (|| {
|
||||
if wpr2_range(bar).is_none() {
|
||||
return Ok(());
|
||||
}
|
||||
|
||||
ctx.sec2_falcon.reset()?;
|
||||
ctx.sec2_falcon.load(&self.booter_unloader)?;
|
||||
|
||||
// Sentinel value to confirm that Booter Unloader has run.
|
||||
const MAILBOX_SENTINEL: u32 = 0xff;
|
||||
let (mbox0, _) =
|
||||
sec2_falcon.boot(bar, Some(MAILBOX_SENTINEL), Some(MAILBOX_SENTINEL))?;
|
||||
let (mbox0, _) = ctx
|
||||
.sec2_falcon
|
||||
.boot(Some(MAILBOX_SENTINEL), Some(MAILBOX_SENTINEL))?;
|
||||
if mbox0 != 0 {
|
||||
dev_err!(dev, "Booter Unloader returned error 0x{:x}\n", mbox0);
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
// Confirm that the WPR2 region has been removed.
|
||||
let wpr2_hi = bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI);
|
||||
if wpr2_hi.is_wpr2_set() {
|
||||
if wpr2_range(bar).is_some() {
|
||||
dev_err!(
|
||||
dev,
|
||||
"WPR2 region still set after Booter Unloader returned\n"
|
||||
);
|
||||
return Err(EBUSY);
|
||||
}
|
||||
}
|
||||
|
||||
Ok(())
|
||||
Ok(())
|
||||
})()
|
||||
.inspect_err(|e| dev_err!(dev, "Booter Unloader failed to run: {:?}\n", e));
|
||||
|
||||
fwsec_sb_res.and(booter_unloader_res)
|
||||
}
|
||||
}
|
||||
|
||||
/// Helper function to load and run the FWSEC-FRTS firmware and confirm that it has properly
|
||||
/// created the WPR2 region.
|
||||
fn run_fwsec_frts(
|
||||
dev: &device::Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
falcon: &Falcon<GspEngine>,
|
||||
bar: Bar0<'_>,
|
||||
bios: &Vbios,
|
||||
fb_layout: &FbLayout,
|
||||
) -> Result {
|
||||
// Check that the WPR2 region does not already exist - if it does, we cannot run
|
||||
// FWSEC-FRTS until the GPU is reset.
|
||||
if bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI).higher_bound() != 0 {
|
||||
dev_err!(
|
||||
pub(super) struct Tu102 {
|
||||
/// If `true`, then the FWSEC-FRTS bootloader will be used to load the actual firmware.
|
||||
pub(super) needs_fwsec_bootloader: bool,
|
||||
}
|
||||
|
||||
impl Tu102 {
|
||||
/// Helper method to load and run the FWSEC-FRTS firmware and confirm that it has properly
|
||||
/// created the WPR2 region.
|
||||
fn run_fwsec_frts(
|
||||
&self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
falcon: &Falcon<'_, GspEngine>,
|
||||
bar: Bar0<'_>,
|
||||
bios: &Vbios,
|
||||
fb_ranges: &FbRanges,
|
||||
) -> Result {
|
||||
// Check that the WPR2 region does not already exist - if it does, we cannot run
|
||||
// FWSEC-FRTS until the GPU is reset.
|
||||
if wpr2_range(bar).is_some() {
|
||||
dev_err!(
|
||||
dev,
|
||||
"WPR2 region already exists - GPU needs to be reset to proceed\n"
|
||||
);
|
||||
return Err(EBUSY);
|
||||
}
|
||||
|
||||
// FWSEC-FRTS will create the WPR2 region.
|
||||
let fwsec_frts = FwsecFirmware::new(
|
||||
dev,
|
||||
"WPR2 region already exists - GPU needs to be reset to proceed\n"
|
||||
);
|
||||
return Err(EBUSY);
|
||||
}
|
||||
falcon,
|
||||
bios,
|
||||
FwsecCommand::Frts {
|
||||
frts_addr: fb_ranges.frts.start,
|
||||
frts_size: fb_ranges.frts.len(),
|
||||
},
|
||||
)?;
|
||||
|
||||
// FWSEC-FRTS will create the WPR2 region.
|
||||
let fwsec_frts = FwsecFirmware::new(
|
||||
dev,
|
||||
falcon,
|
||||
bar,
|
||||
bios,
|
||||
FwsecCommand::Frts {
|
||||
frts_addr: fb_layout.frts.start,
|
||||
frts_size: fb_layout.frts.len(),
|
||||
},
|
||||
)?;
|
||||
if self.needs_fwsec_bootloader {
|
||||
let fwsec_frts_bl = FwsecFirmwareWithBl::new(fwsec_frts, dev, chipset)?;
|
||||
// Load and run the bootloader, which will load FWSEC-FRTS and run it.
|
||||
fwsec_frts_bl.run(dev, falcon, bar)?;
|
||||
} else {
|
||||
// Load and run FWSEC-FRTS directly.
|
||||
fwsec_frts.run(dev, falcon)?;
|
||||
}
|
||||
|
||||
if chipset.needs_fwsec_bootloader() {
|
||||
let fwsec_frts_bl = FwsecFirmwareWithBl::new(fwsec_frts, dev, chipset)?;
|
||||
// Load and run the bootloader, which will load FWSEC-FRTS and run it.
|
||||
fwsec_frts_bl.run(dev, falcon, bar)?;
|
||||
} else {
|
||||
// Load and run FWSEC-FRTS directly.
|
||||
fwsec_frts.run(dev, falcon, bar)?;
|
||||
}
|
||||
// SCRATCH_E contains the error code for FWSEC-FRTS.
|
||||
let frts_status = bar
|
||||
.read(regs::NV_PBUS_SW_SCRATCH_0E_FRTS_ERR)
|
||||
.frts_err_code();
|
||||
if frts_status != 0 {
|
||||
dev_err!(
|
||||
dev,
|
||||
"FWSEC-FRTS returned with error code {:#x}\n",
|
||||
frts_status
|
||||
);
|
||||
|
||||
// SCRATCH_E contains the error code for FWSEC-FRTS.
|
||||
let frts_status = bar
|
||||
.read(regs::NV_PBUS_SW_SCRATCH_0E_FRTS_ERR)
|
||||
.frts_err_code();
|
||||
if frts_status != 0 {
|
||||
dev_err!(
|
||||
dev,
|
||||
"FWSEC-FRTS returned with error code {:#x}\n",
|
||||
frts_status
|
||||
);
|
||||
return Err(EIO);
|
||||
}
|
||||
|
||||
return Err(EIO);
|
||||
}
|
||||
|
||||
// Check that the WPR2 region has been created as we requested.
|
||||
let (wpr2_lo, wpr2_hi) = (
|
||||
bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_LO).lower_bound(),
|
||||
bar.read(regs::NV_PFB_PRI_MMU_WPR2_ADDR_HI).higher_bound(),
|
||||
);
|
||||
|
||||
match (wpr2_lo, wpr2_hi) {
|
||||
(_, 0) => {
|
||||
// Check that the WPR2 region has been created as we requested.
|
||||
let Some(wpr2_range) = wpr2_range(bar) else {
|
||||
dev_err!(dev, "WPR2 region not created after running FWSEC-FRTS\n");
|
||||
|
||||
Err(EIO)
|
||||
}
|
||||
(wpr2_lo, _) if wpr2_lo != fb_layout.frts.start => {
|
||||
return Err(EIO);
|
||||
};
|
||||
|
||||
if wpr2_range.start != fb_ranges.frts.start {
|
||||
dev_err!(
|
||||
dev,
|
||||
"WPR2 region created at unexpected address {:#x}; expected {:#x}\n",
|
||||
wpr2_lo,
|
||||
fb_layout.frts.start,
|
||||
wpr2_range.start,
|
||||
fb_ranges.frts.start,
|
||||
);
|
||||
|
||||
Err(EIO)
|
||||
return Err(EIO);
|
||||
}
|
||||
(wpr2_lo, wpr2_hi) => {
|
||||
dev_dbg!(dev, "WPR2: {:#x}-{:#x}\n", wpr2_lo, wpr2_hi);
|
||||
dev_dbg!(dev, "GPU instance built\n");
|
||||
|
||||
Ok(())
|
||||
}
|
||||
dev_dbg!(dev, "WPR2: {:#x}-{:#x}\n", wpr2_range.start, wpr2_range.end);
|
||||
dev_dbg!(dev, "GPU instance built\n");
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Load and prepare the resources required to properly reset the GSP after it has been stopped.
|
||||
fn build_unload_bundle(
|
||||
&self,
|
||||
dev: &device::Device<device::Bound>,
|
||||
chipset: Chipset,
|
||||
bios: &Vbios,
|
||||
gsp_falcon: &Falcon<'_, GspEngine>,
|
||||
sec2_falcon: &Falcon<'_, Sec2>,
|
||||
) -> Result<crate::gsp::UnloadBundle> {
|
||||
// Load the FWSEC SB firmware, as well as its bootloader if required.
|
||||
let fwsec_sb = FwsecFirmware::new(dev, gsp_falcon, bios, FwsecCommand::Sb)?;
|
||||
let fwsec_sb = if self.needs_fwsec_bootloader {
|
||||
FwsecUnloadFirmware::WithBl(FwsecFirmwareWithBl::new(fwsec_sb, dev, chipset)?)
|
||||
} else {
|
||||
FwsecUnloadFirmware::WithoutBl(fwsec_sb)
|
||||
};
|
||||
|
||||
KBox::new(
|
||||
Sec2UnloadBundle {
|
||||
fwsec_sb,
|
||||
booter_unloader: BooterFirmware::new(
|
||||
dev,
|
||||
BooterKind::Unloader,
|
||||
chipset,
|
||||
sec2_falcon,
|
||||
)?,
|
||||
},
|
||||
GFP_KERNEL,
|
||||
)
|
||||
.map(|b| crate::gsp::UnloadBundle(b))
|
||||
.map_err(Into::into)
|
||||
}
|
||||
}
|
||||
|
||||
struct Tu102;
|
||||
|
||||
impl GspHal for Tu102 {
|
||||
fn boot<'a>(
|
||||
fn boot(
|
||||
&self,
|
||||
gsp: &'a Gsp,
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
bar: Bar0<'a>,
|
||||
chipset: Chipset,
|
||||
fb_layout: &FbLayout,
|
||||
wpr_meta: &Coherent<GspFwWprMeta>,
|
||||
gsp_falcon: &'a Falcon<GspEngine>,
|
||||
sec2_falcon: &'a Falcon<Sec2>,
|
||||
) -> Result<BootUnloadGuard<'a>> {
|
||||
gsp: &Gsp,
|
||||
ctx: &mut GspBootContext<'_, '_>,
|
||||
gsp_fw: &GspFirmware,
|
||||
) -> Result<Option<crate::gsp::UnloadBundle>> {
|
||||
let dev = ctx.dev();
|
||||
let bar = ctx.bar;
|
||||
let chipset = ctx.chipset;
|
||||
let gsp_falcon = ctx.gsp_falcon;
|
||||
let sec2_falcon = ctx.sec2_falcon;
|
||||
|
||||
let fb_ranges = FbRanges::new(chipset, bar, gsp_fw, ctx.vgpu.state())?;
|
||||
dev_dbg!(dev, "{:#x?}\n", fb_ranges);
|
||||
|
||||
// Declared before the unload guard so that if Booter fails while running, SEC2 is reset
|
||||
// by the guard before this allocation is freed.
|
||||
let wpr_meta = Coherent::init(
|
||||
dev,
|
||||
GFP_KERNEL,
|
||||
GspFwWprMeta::from_ranges(gsp_fw, &fb_ranges),
|
||||
)?;
|
||||
|
||||
let bios = Vbios::new(dev, bar)?;
|
||||
|
||||
// Try and prepare the unload bundle.
|
||||
//
|
||||
// If the unload bundle creation fails, the GPU will need to be reset before the driver can
|
||||
// be probed again.
|
||||
let unload_bundle =
|
||||
Sec2UnloadBundle::build(dev, bar, chipset, &bios, gsp_falcon, sec2_falcon)
|
||||
.inspect_err(|e| {
|
||||
dev_warn!(dev, "Failed to prepare unload firmware: {:?}\n", e);
|
||||
dev_warn!(dev, "The GSP won't be able to unload properly on unbind.\n");
|
||||
dev_warn!(
|
||||
dev,
|
||||
"The GPU will need to be reset before the driver can bind again.\n"
|
||||
);
|
||||
})
|
||||
.ok()
|
||||
.map(crate::gsp::UnloadBundle);
|
||||
let unload_bundle = self
|
||||
.build_unload_bundle(dev, chipset, &bios, gsp_falcon, sec2_falcon)
|
||||
.inspect_err(|e| dev_warn!(dev, "Failed to prepare unload firmware: {:?}\n", e))
|
||||
.ok();
|
||||
|
||||
// Wrap the unload bundle into a drop guard so it is automatically run upon failure.
|
||||
let unload_guard =
|
||||
BootUnloadGuard::new(gsp, dev, bar, gsp_falcon, sec2_falcon, unload_bundle);
|
||||
// Run the unload bundle to try and recover the GSP if an error occurs.
|
||||
let unload_guard = ScopeGuard::new_with_data(unload_bundle, |unload_bundle| {
|
||||
if let Some(unload_bundle) = unload_bundle {
|
||||
let _ = unload_bundle.0.run(ctx);
|
||||
}
|
||||
});
|
||||
|
||||
// FWSEC-FRTS is not executed on chips where the FRTS region size is 0 (e.g. GA100).
|
||||
if !fb_layout.frts.is_empty() {
|
||||
run_fwsec_frts(dev, chipset, gsp_falcon, bar, &bios, fb_layout)?;
|
||||
if !fb_ranges.frts.is_empty() {
|
||||
self.run_fwsec_frts(dev, chipset, gsp_falcon, bar, &bios, &fb_ranges)?;
|
||||
}
|
||||
|
||||
gsp_falcon.reset(bar)?;
|
||||
let libos_handle = gsp.libos.dma_handle();
|
||||
gsp_falcon.reset()?;
|
||||
let libos_dma_address = gsp.libos.dma_address();
|
||||
let (mbox0, mbox1) = gsp_falcon.boot(
|
||||
bar,
|
||||
Some(libos_handle as u32),
|
||||
Some((libos_handle >> 32) as u32),
|
||||
Some(libos_dma_address as u32),
|
||||
Some((libos_dma_address >> 32) as u32),
|
||||
)?;
|
||||
dev_dbg!(dev, "GSP MBOX0: {:#x}, MBOX1: {:#x}\n", mbox0, mbox1);
|
||||
|
||||
|
|
@ -308,42 +306,30 @@ fn boot<'a>(
|
|||
"Using SEC2 to load and run the booter_load firmware...\n"
|
||||
);
|
||||
|
||||
BooterFirmware::new(
|
||||
BooterFirmware::new(dev, BooterKind::Loader, chipset, sec2_falcon)?.run(
|
||||
dev,
|
||||
BooterKind::Loader,
|
||||
chipset,
|
||||
FIRMWARE_VERSION,
|
||||
sec2_falcon,
|
||||
bar,
|
||||
)?
|
||||
.run(dev, bar, sec2_falcon, wpr_meta)?;
|
||||
&wpr_meta,
|
||||
)?;
|
||||
|
||||
Ok(unload_guard)
|
||||
Ok(unload_guard.dismiss())
|
||||
}
|
||||
|
||||
fn post_boot(
|
||||
&self,
|
||||
gsp: &Gsp,
|
||||
dev: &device::Device<device::Bound>,
|
||||
bar: Bar0<'_>,
|
||||
ctx: &mut GspBootContext<'_, '_>,
|
||||
gsp_fw: &GspFirmware,
|
||||
gsp_falcon: &Falcon<GspEngine>,
|
||||
sec2_falcon: &Falcon<Sec2>,
|
||||
) -> Result {
|
||||
// Create and run the GSP sequencer.
|
||||
let seq_params = GspSequencerParams {
|
||||
bootloader_app_version: gsp_fw.bootloader.app_version,
|
||||
libos_dma_handle: gsp.libos.dma_handle(),
|
||||
gsp_falcon,
|
||||
sec2_falcon,
|
||||
dev,
|
||||
bar,
|
||||
};
|
||||
GspSequencer::run(&gsp.cmdq, seq_params)?;
|
||||
GspSequencer::run(&gsp.cmdq, ctx, &gsp.libos, gsp_fw.bootloader.app_version)?;
|
||||
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
const TU102: Tu102 = Tu102;
|
||||
/// The TU102 HAL requires the use of the FWSEC bootloader.
|
||||
const TU102: Tu102 = Tu102 {
|
||||
needs_fwsec_bootloader: true,
|
||||
};
|
||||
|
||||
pub(super) const TU102_HAL: &dyn GspHal = &TU102;
|
||||
|
|
|
|||
22
drivers/gpu/nova-core/gsp/regs.rs
Normal file
22
drivers/gpu/nova-core/gsp/regs.rs
Normal file
|
|
@ -0,0 +1,22 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
use kernel::io::register;
|
||||
|
||||
use crate::regs::NV_PBUS_SW_SCRATCH;
|
||||
|
||||
// PGSP
|
||||
|
||||
register! {
|
||||
pub(super) NV_PGSP_QUEUE_HEAD(u32) @ 0x00110c00 {
|
||||
31:0 address;
|
||||
}
|
||||
}
|
||||
|
||||
// PBUS
|
||||
|
||||
register! {
|
||||
/// Scratch register 0xe used as FRTS firmware error code.
|
||||
pub(super) NV_PBUS_SW_SCRATCH_0E_FRTS_ERR(u32) => NV_PBUS_SW_SCRATCH[0xe] {
|
||||
31:16 frts_err_code;
|
||||
}
|
||||
}
|
||||
|
|
@ -6,6 +6,7 @@
|
|||
|
||||
use kernel::{
|
||||
device,
|
||||
dma::Coherent,
|
||||
io::{
|
||||
poll::read_poll_timeout,
|
||||
Io, //
|
||||
|
|
@ -31,6 +32,8 @@
|
|||
MessageFromGsp, //
|
||||
},
|
||||
fw,
|
||||
GspBootContext,
|
||||
LibosMemoryRegionInitArgument, //
|
||||
},
|
||||
num::FromSafeCast,
|
||||
sbuffer::SBufferIter,
|
||||
|
|
@ -128,16 +131,14 @@ pub(crate) fn new(data: &[u8], dev: &device::Device) -> Result<(Self, usize)> {
|
|||
|
||||
/// GSP Sequencer for executing firmware commands during boot.
|
||||
pub(crate) struct GspSequencer<'a> {
|
||||
/// Sequencer information with command data.
|
||||
seq_info: GspSequence,
|
||||
/// `Bar0` for register access.
|
||||
bar: Bar0<'a>,
|
||||
/// SEC2 falcon for core operations.
|
||||
sec2_falcon: &'a Falcon<Sec2>,
|
||||
sec2_falcon: &'a Falcon<'a, Sec2>,
|
||||
/// GSP falcon for core operations.
|
||||
gsp_falcon: &'a Falcon<Gsp>,
|
||||
/// LibOS DMA handle address.
|
||||
libos_dma_handle: u64,
|
||||
gsp_falcon: &'a Falcon<'a, Gsp>,
|
||||
/// LibOS memory region init arguments.
|
||||
libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
|
||||
/// Bootloader application version.
|
||||
bootloader_app_version: u32,
|
||||
/// Device for logging.
|
||||
|
|
@ -213,16 +214,16 @@ fn run(&self, seq: &GspSequencer<'_>) -> Result {
|
|||
GspSeqCmd::DelayUs(cmd) => cmd.run(seq),
|
||||
GspSeqCmd::RegStore(cmd) => cmd.run(seq),
|
||||
GspSeqCmd::CoreReset => {
|
||||
seq.gsp_falcon.reset(seq.bar)?;
|
||||
seq.gsp_falcon.dma_reset(seq.bar);
|
||||
seq.gsp_falcon.reset()?;
|
||||
seq.gsp_falcon.dma_reset();
|
||||
Ok(())
|
||||
}
|
||||
GspSeqCmd::CoreStart => {
|
||||
seq.gsp_falcon.start(seq.bar)?;
|
||||
seq.gsp_falcon.start()?;
|
||||
Ok(())
|
||||
}
|
||||
GspSeqCmd::CoreWaitForHalt => {
|
||||
seq.gsp_falcon.wait_till_halted(seq.bar)?;
|
||||
seq.gsp_falcon.wait_till_halted()?;
|
||||
Ok(())
|
||||
}
|
||||
GspSeqCmd::CoreResume => {
|
||||
|
|
@ -231,35 +232,34 @@ fn run(&self, seq: &GspSequencer<'_>) -> Result {
|
|||
// sequencer will start both.
|
||||
|
||||
// Reset the GSP to prepare it for resuming.
|
||||
seq.gsp_falcon.reset(seq.bar)?;
|
||||
seq.gsp_falcon.reset()?;
|
||||
|
||||
// Write the libOS DMA handle to GSP mailboxes.
|
||||
let libos_dma_address = seq.libos.dma_address();
|
||||
|
||||
// Write the libOS DMA address to GSP mailboxes.
|
||||
seq.gsp_falcon.write_mailboxes(
|
||||
seq.bar,
|
||||
Some(seq.libos_dma_handle as u32),
|
||||
Some((seq.libos_dma_handle >> 32) as u32),
|
||||
Some(libos_dma_address as u32),
|
||||
Some((libos_dma_address >> 32) as u32),
|
||||
);
|
||||
|
||||
// Start the SEC2 falcon which will trigger GSP-RM to resume on the GSP.
|
||||
seq.sec2_falcon.start(seq.bar)?;
|
||||
seq.sec2_falcon.start()?;
|
||||
|
||||
// Poll until GSP-RM reload/resume has completed (up to 2 seconds).
|
||||
seq.gsp_falcon
|
||||
.check_reload_completed(seq.bar, Delta::from_secs(2))?;
|
||||
seq.gsp_falcon.check_reload_completed(Delta::from_secs(2))?;
|
||||
|
||||
// Verify SEC2 completed successfully by checking its mailbox for errors.
|
||||
let mbox0 = seq.sec2_falcon.read_mailbox0(seq.bar);
|
||||
let mbox0 = seq.sec2_falcon.read_mailbox0();
|
||||
if mbox0 != 0 {
|
||||
dev_err!(seq.dev, "Sequencer: sec2 errors: {:?}\n", mbox0);
|
||||
return Err(EIO);
|
||||
}
|
||||
|
||||
// Configure GSP with the bootloader version.
|
||||
seq.gsp_falcon
|
||||
.write_os_version(seq.bar, seq.bootloader_app_version);
|
||||
seq.gsp_falcon.write_os_version(seq.bootloader_app_version);
|
||||
|
||||
// Verify the GSP's RISC-V core is active indicating successful GSP boot.
|
||||
if !seq.gsp_falcon.is_riscv_active(seq.bar) {
|
||||
if !seq.gsp_falcon.is_riscv_active() {
|
||||
dev_err!(seq.dev, "Sequencer: RISC-V core is not active\n");
|
||||
return Err(EIO);
|
||||
}
|
||||
|
|
@ -270,7 +270,7 @@ fn run(&self, seq: &GspSequencer<'_>) -> Result {
|
|||
}
|
||||
|
||||
/// Iterator over GSP sequencer commands.
|
||||
pub(crate) struct GspSeqIter<'a> {
|
||||
struct GspSeqIter<'a> {
|
||||
/// Command data buffer.
|
||||
cmd_data: &'a [u8],
|
||||
/// Current position in the buffer.
|
||||
|
|
@ -283,6 +283,18 @@ pub(crate) struct GspSeqIter<'a> {
|
|||
dev: &'a device::Device,
|
||||
}
|
||||
|
||||
impl<'a> GspSeqIter<'a> {
|
||||
fn new(seq: &'a GspSequence, dev: &'a device::Device) -> Self {
|
||||
Self {
|
||||
cmd_data: &seq.cmd_data,
|
||||
current_offset: 0,
|
||||
total_cmds: seq.cmd_index,
|
||||
cmds_processed: 0,
|
||||
dev,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl<'a> Iterator for GspSeqIter<'a> {
|
||||
type Item = Result<GspSeqCmd>;
|
||||
|
||||
|
|
@ -325,37 +337,12 @@ fn next(&mut self) -> Option<Self::Item> {
|
|||
}
|
||||
|
||||
impl<'a> GspSequencer<'a> {
|
||||
fn iter(&self) -> GspSeqIter<'_> {
|
||||
let cmd_data = &self.seq_info.cmd_data[..];
|
||||
|
||||
GspSeqIter {
|
||||
cmd_data,
|
||||
current_offset: 0,
|
||||
total_cmds: self.seq_info.cmd_index,
|
||||
cmds_processed: 0,
|
||||
dev: self.dev,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Parameters for running the GSP sequencer.
|
||||
pub(crate) struct GspSequencerParams<'a> {
|
||||
/// Bootloader application version.
|
||||
pub(crate) bootloader_app_version: u32,
|
||||
/// LibOS DMA handle address.
|
||||
pub(crate) libos_dma_handle: u64,
|
||||
/// GSP falcon for core operations.
|
||||
pub(crate) gsp_falcon: &'a Falcon<Gsp>,
|
||||
/// SEC2 falcon for core operations.
|
||||
pub(crate) sec2_falcon: &'a Falcon<Sec2>,
|
||||
/// Device for logging.
|
||||
pub(crate) dev: &'a device::Device,
|
||||
/// BAR0 for register access.
|
||||
pub(crate) bar: Bar0<'a>,
|
||||
}
|
||||
|
||||
impl<'a> GspSequencer<'a> {
|
||||
pub(crate) fn run(cmdq: &Cmdq, params: GspSequencerParams<'a>) -> Result {
|
||||
pub(crate) fn run(
|
||||
cmdq: &Cmdq,
|
||||
ctx: &'a GspBootContext<'_, '_>,
|
||||
libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
|
||||
bootloader_app_version: u32,
|
||||
) -> Result {
|
||||
let seq_info = loop {
|
||||
match cmdq.receive_msg::<GspSequence>(Cmdq::RECEIVE_TIMEOUT) {
|
||||
Ok(seq_info) => break seq_info,
|
||||
|
|
@ -365,25 +352,24 @@ pub(crate) fn run(cmdq: &Cmdq, params: GspSequencerParams<'a>) -> Result {
|
|||
};
|
||||
|
||||
let sequencer = GspSequencer {
|
||||
seq_info,
|
||||
bar: params.bar,
|
||||
sec2_falcon: params.sec2_falcon,
|
||||
gsp_falcon: params.gsp_falcon,
|
||||
libos_dma_handle: params.libos_dma_handle,
|
||||
bootloader_app_version: params.bootloader_app_version,
|
||||
dev: params.dev,
|
||||
bar: ctx.bar,
|
||||
sec2_falcon: ctx.sec2_falcon,
|
||||
gsp_falcon: ctx.gsp_falcon,
|
||||
libos,
|
||||
bootloader_app_version,
|
||||
dev: ctx.dev(),
|
||||
};
|
||||
|
||||
dev_dbg!(sequencer.dev, "Running CPU Sequencer commands\n");
|
||||
|
||||
for cmd_result in sequencer.iter() {
|
||||
for cmd_result in GspSeqIter::new(&seq_info, sequencer.dev) {
|
||||
match cmd_result {
|
||||
Ok(cmd) => cmd.run(&sequencer)?,
|
||||
Err(e) => {
|
||||
dev_err!(
|
||||
sequencer.dev,
|
||||
"Error running command at index {}\n",
|
||||
sequencer.seq_info.cmd_index
|
||||
seq_info.cmd_index
|
||||
);
|
||||
return Err(e);
|
||||
}
|
||||
|
|
|
|||
|
|
@ -7,55 +7,53 @@
|
|||
//! Data Model) messages between the kernel driver and GPU firmware processors
|
||||
//! such as FSP and GSP.
|
||||
|
||||
use kernel::pci::Vendor;
|
||||
use kernel::{
|
||||
bitfield,
|
||||
pci::Vendor,
|
||||
prelude::*, //
|
||||
};
|
||||
|
||||
/// NVDM message type identifiers carried over MCTP.
|
||||
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
|
||||
#[repr(u8)]
|
||||
pub(crate) enum NvdmType {
|
||||
#[default]
|
||||
/// Chain of Trust boot message.
|
||||
Cot = 0x14,
|
||||
/// FSP command response.
|
||||
FspResponse = 0x15,
|
||||
}
|
||||
use crate::{
|
||||
bounded_enum,
|
||||
num, //
|
||||
};
|
||||
|
||||
impl TryFrom<u8> for NvdmType {
|
||||
type Error = u8;
|
||||
|
||||
fn try_from(value: u8) -> Result<Self, Self::Error> {
|
||||
match value {
|
||||
x if x == u8::from(Self::Cot) => Ok(Self::Cot),
|
||||
x if x == u8::from(Self::FspResponse) => Ok(Self::FspResponse),
|
||||
_ => Err(value),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl From<NvdmType> for u8 {
|
||||
fn from(value: NvdmType) -> Self {
|
||||
value as u8
|
||||
bounded_enum! {
|
||||
/// NVDM message type identifiers carried over MCTP.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub(crate) enum NvdmType with TryFrom<Bounded<u32, 8>> {
|
||||
/// PRC (Product Reconfiguration Control) message.
|
||||
Prc = 0x13,
|
||||
/// Chain of Trust boot message.
|
||||
Cot = 0x14,
|
||||
/// FSP command response.
|
||||
FspResponse = 0x15,
|
||||
}
|
||||
}
|
||||
|
||||
bitfield! {
|
||||
pub(crate) struct MctpHeader(u32), "MCTP transport header for NVIDIA firmware messages." {
|
||||
31:31 som as bool, "Start-of-message bit.";
|
||||
30:30 eom as bool, "End-of-message bit.";
|
||||
29:28 seq as u8, "Packet sequence number.";
|
||||
23:16 seid as u8, "Source endpoint ID.";
|
||||
/// MCTP transport header for NVIDIA firmware messages.
|
||||
pub(crate) struct MctpHeader(u32) {
|
||||
/// Start-of-message bit.
|
||||
31:31 som;
|
||||
/// End-of-message bit.
|
||||
30:30 eom;
|
||||
/// Packet sequence number.
|
||||
29:28 seq;
|
||||
/// Source endpoint ID.
|
||||
23:16 seid;
|
||||
}
|
||||
}
|
||||
|
||||
impl MctpHeader {
|
||||
/// Builds a single-packet MCTP header (`SOM=1`, `EOM=1`, `SEQ=0`, `SEID=0`).
|
||||
pub(crate) fn single_packet() -> Self {
|
||||
Self::default().set_som(true).set_eom(true)
|
||||
Self::zeroed().with_som(true).with_eom(true)
|
||||
}
|
||||
|
||||
/// Returns whether this is a complete single-packet message (`SOM=1` and `EOM=1`).
|
||||
pub(crate) fn is_single_packet(self) -> bool {
|
||||
self.som() && self.eom()
|
||||
self.som().into_bool() && self.eom().into_bool()
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -63,26 +61,30 @@ pub(crate) fn is_single_packet(self) -> bool {
|
|||
const MSG_TYPE_VENDOR_PCI: u8 = 0x7e;
|
||||
|
||||
bitfield! {
|
||||
pub(crate) struct NvdmHeader(u32), "NVIDIA Vendor-Defined Message header over MCTP." {
|
||||
31:24 nvdm_type as u8 ?=> NvdmType, "NVDM message type.";
|
||||
23:8 vendor_id as u16, "PCI vendor ID.";
|
||||
6:0 msg_type as u8, "MCTP vendor-defined message type.";
|
||||
/// NVIDIA Vendor-Defined Message header over MCTP.
|
||||
pub(crate) struct NvdmHeader(u32) {
|
||||
/// NVDM message type.
|
||||
31:24 nvdm_type ?=> NvdmType;
|
||||
/// PCI vendor ID.
|
||||
23:8 vendor_id;
|
||||
/// MCTP vendor-defined message type.
|
||||
6:0 msg_type;
|
||||
}
|
||||
}
|
||||
|
||||
impl NvdmHeader {
|
||||
/// Builds an NVDM header for the given message type.
|
||||
pub(crate) fn new(nvdm_type: NvdmType) -> Self {
|
||||
Self::default()
|
||||
.set_msg_type(MSG_TYPE_VENDOR_PCI)
|
||||
.set_vendor_id(Vendor::NVIDIA.as_raw())
|
||||
.set_nvdm_type(nvdm_type)
|
||||
Self::zeroed()
|
||||
.with_const_msg_type::<{ num::u8_as_u32(MSG_TYPE_VENDOR_PCI) }>()
|
||||
.with_vendor_id(Vendor::NVIDIA.as_raw())
|
||||
.with_nvdm_type(nvdm_type)
|
||||
}
|
||||
|
||||
/// Validates this header against the expected NVIDIA NVDM format and type.
|
||||
pub(crate) fn validate(self, expected_type: NvdmType) -> bool {
|
||||
self.msg_type() == MSG_TYPE_VENDOR_PCI
|
||||
&& self.vendor_id() == Vendor::NVIDIA.as_raw()
|
||||
u8::from(self.msg_type()) == MSG_TYPE_VENDOR_PCI
|
||||
&& u16::from(self.vendor_id()) == Vendor::NVIDIA.as_raw()
|
||||
&& matches!(self.nvdm_type(), Ok(nvdm_type) if nvdm_type == expected_type)
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -10,9 +10,6 @@
|
|||
InPlaceModule, //
|
||||
};
|
||||
|
||||
#[macro_use]
|
||||
mod bitfield;
|
||||
|
||||
mod driver;
|
||||
mod falcon;
|
||||
mod fb;
|
||||
|
|
@ -26,6 +23,7 @@
|
|||
mod regs;
|
||||
mod sbuffer;
|
||||
mod vbios;
|
||||
mod vgpu;
|
||||
|
||||
pub(crate) const MODULE_NAME: &core::ffi::CStr = <LocalModule as kernel::ModuleMetadata>::NAME;
|
||||
|
||||
|
|
@ -54,7 +52,7 @@ struct NovaCoreModule {
|
|||
|
||||
impl InPlaceModule for NovaCoreModule {
|
||||
fn init(module: &'static kernel::ThisModule) -> impl PinInit<Self, Error> {
|
||||
let dir = debugfs::Dir::new(kernel::c_str!("nova-core"));
|
||||
let dir = debugfs::Dir::new(c"nova-core");
|
||||
|
||||
// SAFETY: We are the only driver code running during init, so there
|
||||
// cannot be any concurrent access to `DEBUGFS_ROOT`.
|
||||
|
|
|
|||
15
drivers/gpu/nova-core/nova_core_exports.c
Normal file
15
drivers/gpu/nova-core/nova_core_exports.c
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
/*
|
||||
* Exports Rust symbols from the `nova_core` crate for use by dependent modules.
|
||||
*
|
||||
* This is a workaround until the build system supports Rust cross-module
|
||||
* dependencies natively.
|
||||
*/
|
||||
|
||||
#include <linux/export.h>
|
||||
|
||||
#define EXPORT_SYMBOL_RUST_GPL(sym) extern int sym; EXPORT_SYMBOL_GPL(sym)
|
||||
|
||||
#include "exports_nova_core_generated.h"
|
||||
|
|
@ -109,130 +109,6 @@ fn fmt(&self, f: &mut kernel::fmt::Formatter<'_>) -> kernel::fmt::Result {
|
|||
|
||||
register! {
|
||||
pub(crate) NV_PBUS_SW_SCRATCH(u32)[64] @ 0x00001400 {}
|
||||
|
||||
/// Scratch register 0xe used as FRTS firmware error code.
|
||||
pub(crate) NV_PBUS_SW_SCRATCH_0E_FRTS_ERR(u32) => NV_PBUS_SW_SCRATCH[0xe] {
|
||||
31:16 frts_err_code;
|
||||
}
|
||||
}
|
||||
|
||||
// PFB
|
||||
|
||||
register! {
|
||||
/// Low bits of the physical system memory address used by the GPU to perform sysmembar
|
||||
/// operations (see [`crate::fb::SysmemFlush`]).
|
||||
pub(crate) NV_PFB_NISO_FLUSH_SYSMEM_ADDR(u32) @ 0x00100c10 {
|
||||
31:0 adr_39_08;
|
||||
}
|
||||
|
||||
/// High bits of the physical system memory address used by the GPU to perform sysmembar
|
||||
/// operations (see [`crate::fb::SysmemFlush`]).
|
||||
pub(crate) NV_PFB_NISO_FLUSH_SYSMEM_ADDR_HI(u32) @ 0x00100c40 {
|
||||
23:0 adr_63_40;
|
||||
}
|
||||
|
||||
pub(crate) NV_PFB_PRI_MMU_LOCAL_MEMORY_RANGE(u32) @ 0x00100ce0 {
|
||||
30:30 ecc_mode_enabled => bool;
|
||||
9:4 lower_mag;
|
||||
3:0 lower_scale;
|
||||
}
|
||||
|
||||
pub(crate) NV_PFB_PRI_MMU_WPR2_ADDR_LO(u32) @ 0x001fa824 {
|
||||
/// Bits 12..40 of the lower (inclusive) bound of the WPR2 region.
|
||||
31:4 lo_val;
|
||||
}
|
||||
|
||||
pub(crate) NV_PFB_PRI_MMU_WPR2_ADDR_HI(u32) @ 0x001fa828 {
|
||||
/// Bits 12..40 of the higher (exclusive) bound of the WPR2 region.
|
||||
31:4 hi_val;
|
||||
}
|
||||
}
|
||||
|
||||
/// Base of the GB10x HSHUB0 register window (`NV_HSHUB0_PRIV_BASE` in Open RM).
|
||||
///
|
||||
/// The base is provided by the GB10x framebuffer HAL.
|
||||
pub(crate) struct Hshub0Base(());
|
||||
|
||||
/// Base of the GB20x FBHUB0 register window (`NV_FBHUB0_PRI_BASE` in Open RM).
|
||||
///
|
||||
/// The base is provided by the GB20x framebuffer HAL.
|
||||
pub(crate) struct Fbhub0Base(());
|
||||
|
||||
register! {
|
||||
// GB10x sysmem flush registers, relative to the HSHUB0 base. GB10x routes sysmembar
|
||||
// through a primary and an EG (egress) pair that must both be programmed to the same
|
||||
// address. Hardware ignores bits 7:0 of each LO register. The boot path uses a fixed
|
||||
// HSHUB0 base, so the multiple runtime-discovered HSHUB bases are not needed here.
|
||||
pub(crate) NV_PFB_HSHUB_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Hshub0Base + 0x00000e50 {
|
||||
31:0 adr => u32;
|
||||
}
|
||||
|
||||
pub(crate) NV_PFB_HSHUB_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Hshub0Base + 0x00000e54 {
|
||||
19:0 adr;
|
||||
}
|
||||
|
||||
pub(crate) NV_PFB_HSHUB_EG_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Hshub0Base + 0x000006c0 {
|
||||
31:0 adr => u32;
|
||||
}
|
||||
|
||||
pub(crate) NV_PFB_HSHUB_EG_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Hshub0Base + 0x000006c4 {
|
||||
19:0 adr;
|
||||
}
|
||||
|
||||
// GB20x sysmem flush registers, relative to the FBHUB0 base. Unlike the older
|
||||
// NV_PFB_NISO_FLUSH_SYSMEM_ADDR registers which encode the address with an 8-bit
|
||||
// right-shift, these take the raw address split into lower and upper halves. Hardware
|
||||
// ignores bits 7:0 of the LO register.
|
||||
pub(crate) NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_LO(u32) @ Fbhub0Base + 0x00001d58 {
|
||||
31:0 adr => u32;
|
||||
}
|
||||
|
||||
pub(crate) NV_PFB_FBHUB_PCIE_FLUSH_SYSMEM_ADDR_HI(u32) @ Fbhub0Base + 0x00001d5c {
|
||||
19:0 adr;
|
||||
}
|
||||
}
|
||||
|
||||
impl NV_PFB_PRI_MMU_LOCAL_MEMORY_RANGE {
|
||||
/// Returns the usable framebuffer size, in bytes.
|
||||
pub(crate) fn usable_fb_size(self) -> u64 {
|
||||
let size = (u64::from(self.lower_mag()) << u64::from(self.lower_scale())) * u64::SZ_1M;
|
||||
|
||||
if self.ecc_mode_enabled() {
|
||||
// Remove the amount of memory reserved for ECC (one per 16 units).
|
||||
size / 16 * 15
|
||||
} else {
|
||||
size
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl NV_PFB_PRI_MMU_WPR2_ADDR_LO {
|
||||
/// Returns the lower (inclusive) bound of the WPR2 region.
|
||||
pub(crate) fn lower_bound(self) -> u64 {
|
||||
u64::from(self.lo_val()) << 12
|
||||
}
|
||||
}
|
||||
|
||||
impl NV_PFB_PRI_MMU_WPR2_ADDR_HI {
|
||||
/// Returns the higher (exclusive) bound of the WPR2 region.
|
||||
///
|
||||
/// A value of zero means the WPR2 region is not set.
|
||||
pub(crate) fn higher_bound(self) -> u64 {
|
||||
u64::from(self.hi_val()) << 12
|
||||
}
|
||||
|
||||
/// Returns whether the WPR2 region is currently set.
|
||||
pub(crate) fn is_wpr2_set(self) -> bool {
|
||||
self.hi_val() != 0
|
||||
}
|
||||
}
|
||||
|
||||
// PGSP
|
||||
|
||||
register! {
|
||||
pub(crate) NV_PGSP_QUEUE_HEAD(u32) @ 0x00110c00 {
|
||||
31:0 address;
|
||||
}
|
||||
}
|
||||
|
||||
// PGC6 register space.
|
||||
|
|
@ -294,28 +170,6 @@ pub(crate) fn usable_fb_size(self) -> u64 {
|
|||
}
|
||||
}
|
||||
|
||||
// PDISP
|
||||
|
||||
register! {
|
||||
pub(crate) NV_PDISP_VGA_WORKSPACE_BASE(u32) @ 0x00625f04 {
|
||||
/// VGA workspace base address divided by 0x10000.
|
||||
31:8 addr;
|
||||
/// Set if the `addr` field is valid.
|
||||
3:3 status_valid => bool;
|
||||
}
|
||||
}
|
||||
|
||||
impl NV_PDISP_VGA_WORKSPACE_BASE {
|
||||
/// Returns the base address of the VGA workspace, or `None` if none exists.
|
||||
pub(crate) fn vga_workspace_addr(self) -> Option<u64> {
|
||||
if self.status_valid() {
|
||||
Some(u64::from(self.addr()) << 16)
|
||||
} else {
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// FUSE
|
||||
|
||||
pub(crate) const NV_FUSE_OPT_FPF_SIZE: usize = 16;
|
||||
|
|
@ -570,7 +424,7 @@ pub(crate) fn mem_scrubbing_done(self) -> bool {
|
|||
/// GA102 and later.
|
||||
pub(crate) NV_PRISCV_RISCV_CPUCTL(u32) @ PFalcon2Base + 0x00000388 {
|
||||
7:7 active_stat => bool;
|
||||
0:0 halted => bool;
|
||||
4:4 halted => bool;
|
||||
}
|
||||
|
||||
/// GA102 and later.
|
||||
|
|
|
|||
|
|
@ -13,11 +13,8 @@
|
|||
register,
|
||||
sizes::SZ_4K,
|
||||
sync::aref::ARef,
|
||||
transmute::FromBytes,
|
||||
};
|
||||
|
||||
use zerocopy::FromBytes as _;
|
||||
|
||||
use crate::{
|
||||
driver::Bar0,
|
||||
firmware::{
|
||||
|
|
@ -359,7 +356,7 @@ pub(crate) fn fwsec_image(&self) -> &FwSecBiosImage {
|
|||
}
|
||||
|
||||
/// PCI Data Structure as defined in PCI Firmware Specification
|
||||
#[derive(Debug, Clone)]
|
||||
#[derive(Debug, Clone, FromBytes)]
|
||||
#[repr(C)]
|
||||
struct PcirStruct {
|
||||
/// PCI Data Structure signature ("PCIR" or "NPDS")
|
||||
|
|
@ -388,15 +385,12 @@ struct PcirStruct {
|
|||
max_runtime_image_len: u16,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for `PcirStruct`.
|
||||
unsafe impl FromBytes for PcirStruct {}
|
||||
|
||||
impl PcirStruct {
|
||||
/// The bit in `last_image` that indicates the last image.
|
||||
const LAST_IMAGE_BIT_MASK: u8 = 0x80;
|
||||
|
||||
fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
|
||||
let (pcir, _) = PcirStruct::from_bytes_copy_prefix(data).ok_or(EINVAL)?;
|
||||
let (pcir, _) = PcirStruct::read_from_prefix(data).map_err(|_| EINVAL)?;
|
||||
|
||||
// Signature should be "PCIR" (0x52494350) or "NPDS" (0x5344504e).
|
||||
if &pcir.signature != b"PCIR" && &pcir.signature != b"NPDS" {
|
||||
|
|
@ -432,7 +426,7 @@ fn image_size_bytes(&self) -> usize {
|
|||
/// This is the head of the BIT table, that is used to locate the Falcon data. The BIT table (with
|
||||
/// its header) is in the [`PciAtBiosImage`] and the falcon data it is pointing to is in the
|
||||
/// [`FwSecBiosImage`].
|
||||
#[derive(Debug, Clone, Copy)]
|
||||
#[derive(Debug, Clone, Copy, FromBytes)]
|
||||
#[repr(C)]
|
||||
struct BitHeader {
|
||||
/// 0h: BIT Header Identifier (BMP=0x7FFF/BIT=0xB8FF)
|
||||
|
|
@ -451,12 +445,9 @@ struct BitHeader {
|
|||
checksum: u8,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for `BitHeader`.
|
||||
unsafe impl FromBytes for BitHeader {}
|
||||
|
||||
impl BitHeader {
|
||||
fn new(data: &[u8]) -> Result<Self> {
|
||||
let (header, _) = BitHeader::from_bytes_copy_prefix(data).ok_or(EINVAL)?;
|
||||
let (header, _) = BitHeader::read_from_prefix(data).map_err(|_| EINVAL)?;
|
||||
|
||||
// Check header ID and signature
|
||||
if header.id != 0xB8FF || &header.signature != b"BIT\0" {
|
||||
|
|
@ -468,7 +459,7 @@ fn new(data: &[u8]) -> Result<Self> {
|
|||
}
|
||||
|
||||
/// BIT Token Entry: Records in the BIT table followed by the BIT header.
|
||||
#[derive(Debug, Clone, Copy)]
|
||||
#[derive(Debug, Clone, Copy, FromBytes)]
|
||||
#[repr(C)]
|
||||
struct BitToken {
|
||||
/// 00h: Token identifier
|
||||
|
|
@ -481,9 +472,6 @@ struct BitToken {
|
|||
data_offset: u16,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for `BitToken`.
|
||||
unsafe impl FromBytes for BitToken {}
|
||||
|
||||
impl BitToken {
|
||||
/// BIT token ID for Falcon data.
|
||||
const ID_FALCON_DATA: u8 = 0x70;
|
||||
|
|
@ -508,7 +496,7 @@ fn from_id(image: &PciAtBiosImage, token_id: u8) -> Result<Self> {
|
|||
.and_then(|data| data.get(..entry_size))
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
let (token, _) = BitToken::from_bytes_copy_prefix(entry).ok_or(EINVAL)?;
|
||||
let (token, _) = BitToken::read_from_prefix(entry).map_err(|_| EINVAL)?;
|
||||
|
||||
// Check if this token has the requested ID
|
||||
if token.id == token_id {
|
||||
|
|
@ -525,7 +513,7 @@ fn from_id(image: &PciAtBiosImage, token_id: u8) -> Result<Self> {
|
|||
///
|
||||
/// This header is at the beginning of every image in the set of images in the ROM. It contains a
|
||||
/// pointer to the PCI Data Structure which describes the image.
|
||||
#[derive(Debug, Clone, Copy)]
|
||||
#[derive(Debug, Clone, Copy, FromBytes)]
|
||||
#[repr(C)]
|
||||
struct PciRomHeader {
|
||||
/// 00h: Signature (0xAA55)
|
||||
|
|
@ -536,13 +524,10 @@ struct PciRomHeader {
|
|||
pci_data_struct_offset: u16,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for `PciRomHeader`.
|
||||
unsafe impl FromBytes for PciRomHeader {}
|
||||
|
||||
impl PciRomHeader {
|
||||
fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
|
||||
let (rom_header, _) = PciRomHeader::from_bytes_copy_prefix(data)
|
||||
.ok_or(EINVAL)
|
||||
let (rom_header, _) = PciRomHeader::read_from_prefix(data)
|
||||
.map_err(|_| EINVAL)
|
||||
.inspect_err(|_| dev_err!(dev, "Not enough data for ROM header\n"))?;
|
||||
|
||||
// Check for valid ROM signatures.
|
||||
|
|
@ -564,7 +549,7 @@ fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
|
|||
/// PCI Data Structure. It contains some fields that are redundant with the PCI Data Structure, but
|
||||
/// are needed for traversing the BIOS images. It is expected to be present in all BIOS images
|
||||
/// except for NBSI images.
|
||||
#[derive(Debug, Clone)]
|
||||
#[derive(Debug, Clone, FromBytes)]
|
||||
#[repr(C)]
|
||||
struct NpdeStruct {
|
||||
/// 00h: Signature ("NPDE")
|
||||
|
|
@ -579,15 +564,12 @@ struct NpdeStruct {
|
|||
last_image: u8,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for `NpdeStruct`.
|
||||
unsafe impl FromBytes for NpdeStruct {}
|
||||
|
||||
impl NpdeStruct {
|
||||
/// The bit in `last_image` that indicates the last image.
|
||||
const LAST_IMAGE_BIT_MASK: u8 = 0x80;
|
||||
|
||||
fn new(dev: &device::Device, data: &[u8]) -> Option<Self> {
|
||||
let (npde, _) = NpdeStruct::from_bytes_copy_prefix(data)?;
|
||||
let (npde, _) = NpdeStruct::read_from_prefix(data).ok()?;
|
||||
|
||||
// Signature should be "NPDE" (0x4544504E).
|
||||
if &npde.signature != b"NPDE" {
|
||||
|
|
@ -784,7 +766,7 @@ fn falcon_data_offset(&self, dev: &device::Device) -> Result<usize> {
|
|||
let data = &self.base.data;
|
||||
let (ptr, _) = data
|
||||
.get(offset..)
|
||||
.and_then(u32::from_bytes_copy_prefix)
|
||||
.and_then(|p| u32::read_from_prefix(p).ok())
|
||||
.ok_or(EINVAL)?;
|
||||
|
||||
usize::from_safe_cast(ptr)
|
||||
|
|
@ -814,6 +796,7 @@ fn try_from(base: BiosImage) -> Result<Self> {
|
|||
/// The [`PmuLookupTableEntry`] structure is a single entry in the [`PmuLookupTable`].
|
||||
///
|
||||
/// See the [`PmuLookupTable`] description for more information.
|
||||
#[derive(FromBytes)]
|
||||
#[repr(C, packed)]
|
||||
struct PmuLookupTableEntry {
|
||||
application_id: u8,
|
||||
|
|
@ -821,9 +804,6 @@ struct PmuLookupTableEntry {
|
|||
data: u32,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for `PmuLookupTableEntry`.
|
||||
unsafe impl FromBytes for PmuLookupTableEntry {}
|
||||
|
||||
impl PmuLookupTableEntry {
|
||||
/// PMU lookup table application ID for firmware security license ucode.
|
||||
#[expect(dead_code)]
|
||||
|
|
@ -836,6 +816,7 @@ impl PmuLookupTableEntry {
|
|||
}
|
||||
|
||||
#[repr(C)]
|
||||
#[derive(FromBytes)]
|
||||
struct PmuLookupTableHeader {
|
||||
version: u8,
|
||||
header_len: u8,
|
||||
|
|
@ -843,9 +824,6 @@ struct PmuLookupTableHeader {
|
|||
entry_count: u8,
|
||||
}
|
||||
|
||||
// SAFETY: all bit patterns are valid for `PmuLookupTableHeader`.
|
||||
unsafe impl FromBytes for PmuLookupTableHeader {}
|
||||
|
||||
/// The [`PmuLookupTableEntry`] structure is used to find the [`PmuLookupTableEntry`] for a given
|
||||
/// application ID.
|
||||
///
|
||||
|
|
@ -857,7 +835,7 @@ struct PmuLookupTable {
|
|||
|
||||
impl PmuLookupTable {
|
||||
fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
|
||||
let (header, _) = PmuLookupTableHeader::from_bytes_copy_prefix(data).ok_or(EINVAL)?;
|
||||
let (header, _) = PmuLookupTableHeader::read_from_prefix(data).map_err(|_| EINVAL)?;
|
||||
|
||||
let header_len = usize::from(header.header_len);
|
||||
let entry_len = usize::from(header.entry_len);
|
||||
|
|
@ -872,8 +850,8 @@ fn new(dev: &device::Device, data: &[u8]) -> Result<Self> {
|
|||
|
||||
let mut entries = KVVec::with_capacity(entry_count, GFP_KERNEL)?;
|
||||
for i in 0..entry_count {
|
||||
let (entry, _) = PmuLookupTableEntry::from_bytes_copy_prefix(&data[i * entry_len..])
|
||||
.ok_or(EINVAL)?;
|
||||
let (entry, _) = PmuLookupTableEntry::read_from_prefix(&data[i * entry_len..])
|
||||
.map_err(|_| EINVAL)?;
|
||||
entries.push(entry, GFP_KERNEL)?;
|
||||
}
|
||||
|
||||
|
|
@ -929,15 +907,11 @@ pub(crate) fn header(&self) -> Result<FalconUCodeDesc> {
|
|||
let ver = data.get(1).copied().ok_or(EINVAL)?;
|
||||
match ver {
|
||||
2 => {
|
||||
let v2 = FalconUCodeDescV2::read_from_prefix(data)
|
||||
.map_err(|_| EINVAL)?
|
||||
.0;
|
||||
let (v2, _) = FalconUCodeDescV2::read_from_prefix(data).map_err(|_| EINVAL)?;
|
||||
Ok(FalconUCodeDesc::V2(v2))
|
||||
}
|
||||
3 => {
|
||||
let v3 = FalconUCodeDescV3::from_bytes_copy_prefix(data)
|
||||
.ok_or(EINVAL)?
|
||||
.0;
|
||||
let (v3, _) = FalconUCodeDescV3::read_from_prefix(data).map_err(|_| EINVAL)?;
|
||||
Ok(FalconUCodeDesc::V3(v3))
|
||||
}
|
||||
_ => {
|
||||
|
|
|
|||
91
drivers/gpu/nova-core/vgpu.rs
Normal file
91
drivers/gpu/nova-core/vgpu.rs
Normal file
|
|
@ -0,0 +1,91 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
use core::num::NonZero;
|
||||
|
||||
use kernel::{
|
||||
device,
|
||||
pci,
|
||||
prelude::*, //
|
||||
};
|
||||
|
||||
use crate::{
|
||||
fsp::{
|
||||
Fsp,
|
||||
VgpuMode, //
|
||||
},
|
||||
gpu::Chipset, //
|
||||
};
|
||||
|
||||
mod hal;
|
||||
|
||||
/// vGPU state detected during GPU construction.
|
||||
#[derive(Debug, Clone, Copy)]
|
||||
pub(crate) enum VgpuState {
|
||||
/// vGPU mode is not enabled for this boot.
|
||||
Disabled,
|
||||
/// vGPU mode is enabled for this boot.
|
||||
Enabled {
|
||||
/// Total number of SR-IOV VFs supported by this device.
|
||||
total_vfs: NonZero<u16>,
|
||||
},
|
||||
}
|
||||
|
||||
/// vGPU state manager.
|
||||
pub(crate) struct VgpuManager {
|
||||
state: VgpuState,
|
||||
}
|
||||
|
||||
impl VgpuManager {
|
||||
/// Creates a vGPU manager by querying SR-IOV and the FSP PRC vGPU knob.
|
||||
pub(crate) fn new(
|
||||
pdev: &pci::Device<device::Core<'_>>,
|
||||
chipset: Chipset,
|
||||
fsp: Option<&mut Fsp<'_>>,
|
||||
) -> Self {
|
||||
let state = Self::detect_state(pdev, chipset, fsp).unwrap_or_else(|e| {
|
||||
dev_warn!(
|
||||
pdev,
|
||||
"vGPU state detection failed: {:?}; disabling vGPU\n",
|
||||
e
|
||||
);
|
||||
VgpuState::Disabled
|
||||
});
|
||||
dev_dbg!(pdev, "vGPU state: {:?}\n", state);
|
||||
|
||||
Self { state }
|
||||
}
|
||||
|
||||
/// Detects the vGPU state from the chipset, SR-IOV capability and FSP PRC knob.
|
||||
fn detect_state(
|
||||
pdev: &pci::Device<device::Core<'_>>,
|
||||
chipset: Chipset,
|
||||
fsp: Option<&mut Fsp<'_>>,
|
||||
) -> Result<VgpuState> {
|
||||
if !hal::vgpu_hal(chipset).supports_vgpu() {
|
||||
return Ok(VgpuState::Disabled);
|
||||
}
|
||||
|
||||
let Some(total_vfs) = pdev.sriov_get_totalvfs() else {
|
||||
return Ok(VgpuState::Disabled);
|
||||
};
|
||||
|
||||
if total_vfs.get() < 2 {
|
||||
// The current vGPU path does not support single-VF SR-IOV devices yet.
|
||||
// Treat one total VF as vGPU-disabled for now; single-VF support can relax
|
||||
// this gate once the manager handles that topology.
|
||||
return Ok(VgpuState::Disabled);
|
||||
}
|
||||
|
||||
let fsp = fsp.ok_or(ENODEV)?;
|
||||
|
||||
match fsp.read_vgpu_mode(pdev.as_ref())? {
|
||||
VgpuMode::Enabled => Ok(VgpuState::Enabled { total_vfs }),
|
||||
VgpuMode::Disabled => Ok(VgpuState::Disabled),
|
||||
}
|
||||
}
|
||||
|
||||
/// Returns the detected vGPU state for this boot.
|
||||
pub(crate) fn state(&self) -> VgpuState {
|
||||
self.state
|
||||
}
|
||||
}
|
||||
25
drivers/gpu/nova-core/vgpu/hal.rs
Normal file
25
drivers/gpu/nova-core/vgpu/hal.rs
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
use crate::gpu::{
|
||||
Architecture,
|
||||
Chipset, //
|
||||
};
|
||||
|
||||
mod gb202;
|
||||
mod tu102;
|
||||
|
||||
pub(super) trait VgpuHal {
|
||||
/// Returns whether this chipset can support vGPU.
|
||||
fn supports_vgpu(&self) -> bool;
|
||||
}
|
||||
|
||||
pub(super) fn vgpu_hal(chipset: Chipset) -> &'static dyn VgpuHal {
|
||||
match chipset.arch() {
|
||||
Architecture::BlackwellGB20x => gb202::GB202_HAL,
|
||||
Architecture::Turing
|
||||
| Architecture::Ampere
|
||||
| Architecture::Hopper
|
||||
| Architecture::Ada
|
||||
| Architecture::BlackwellGB10x => tu102::TU102_HAL,
|
||||
}
|
||||
}
|
||||
15
drivers/gpu/nova-core/vgpu/hal/gb202.rs
Normal file
15
drivers/gpu/nova-core/vgpu/hal/gb202.rs
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use crate::vgpu::hal::VgpuHal;
|
||||
|
||||
struct Gb202;
|
||||
|
||||
impl VgpuHal for Gb202 {
|
||||
fn supports_vgpu(&self) -> bool {
|
||||
true
|
||||
}
|
||||
}
|
||||
|
||||
const GB202: Gb202 = Gb202;
|
||||
pub(super) const GB202_HAL: &dyn VgpuHal = &GB202;
|
||||
15
drivers/gpu/nova-core/vgpu/hal/tu102.rs
Normal file
15
drivers/gpu/nova-core/vgpu/hal/tu102.rs
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
|
||||
use crate::vgpu::hal::VgpuHal;
|
||||
|
||||
struct Tu102;
|
||||
|
||||
impl VgpuHal for Tu102 {
|
||||
fn supports_vgpu(&self) -> bool {
|
||||
false
|
||||
}
|
||||
}
|
||||
|
||||
const TU102: Tu102 = Tu102;
|
||||
pub(super) const TU102_HAL: &dyn VgpuHal = &TU102;
|
||||
|
|
@ -1283,7 +1283,7 @@ EXPORT_SYMBOL_GPL(pci_sriov_set_totalvfs);
|
|||
* SRIOV capability value of TotalVFs or the value of driver_max_VFs
|
||||
* if the driver reduced it. Otherwise 0.
|
||||
*/
|
||||
int pci_sriov_get_totalvfs(struct pci_dev *dev)
|
||||
unsigned int pci_sriov_get_totalvfs(struct pci_dev *dev)
|
||||
{
|
||||
if (!dev->is_physfn)
|
||||
return 0;
|
||||
|
|
|
|||
|
|
@ -20,7 +20,6 @@
|
|||
//! this method is not used in this driver.
|
||||
//!
|
||||
|
||||
use core::ops::Deref;
|
||||
use kernel::{
|
||||
clk::Clk,
|
||||
device::{Bound, Core, Device},
|
||||
|
|
@ -213,8 +212,7 @@ fn read_waveform(
|
|||
) -> Result<Self::WfHw> {
|
||||
let data = chip.drvdata();
|
||||
let hwpwm = pwm.hwpwm();
|
||||
let iomem_accessor = data.iomem.access(parent_dev)?;
|
||||
let iomap = iomem_accessor.deref();
|
||||
let iomap = data.iomem.access(parent_dev)?;
|
||||
|
||||
let ctrl = iomap.try_read32(th1520_pwm_ctrl(hwpwm))?;
|
||||
let period_cycles = iomap.try_read32(th1520_pwm_per(hwpwm))?;
|
||||
|
|
@ -248,8 +246,7 @@ fn write_waveform(
|
|||
) -> Result {
|
||||
let data = chip.drvdata();
|
||||
let hwpwm = pwm.hwpwm();
|
||||
let iomem_accessor = data.iomem.access(parent_dev)?;
|
||||
let iomap = iomem_accessor.deref();
|
||||
let iomap = data.iomem.access(parent_dev)?;
|
||||
let duty_cycles = iomap.try_read32(th1520_pwm_fp(hwpwm))?;
|
||||
let was_enabled = duty_cycles != 0;
|
||||
|
||||
|
|
|
|||
|
|
@ -2569,7 +2569,7 @@ void pci_iov_remove_virtfn(struct pci_dev *dev, int id);
|
|||
int pci_num_vf(struct pci_dev *dev);
|
||||
int pci_vfs_assigned(struct pci_dev *dev);
|
||||
int pci_sriov_set_totalvfs(struct pci_dev *dev, u16 numvfs);
|
||||
int pci_sriov_get_totalvfs(struct pci_dev *dev);
|
||||
unsigned int pci_sriov_get_totalvfs(struct pci_dev *dev);
|
||||
int pci_sriov_configure_simple(struct pci_dev *dev, int nr_virtfn);
|
||||
resource_size_t pci_iov_resource_size(const struct pci_dev *dev, int resno);
|
||||
int pci_iov_vf_bar_set_size(struct pci_dev *dev, int resno, int size);
|
||||
|
|
@ -2622,7 +2622,7 @@ static inline int pci_vfs_assigned(struct pci_dev *dev)
|
|||
{ return 0; }
|
||||
static inline int pci_sriov_set_totalvfs(struct pci_dev *dev, u16 numvfs)
|
||||
{ return 0; }
|
||||
static inline int pci_sriov_get_totalvfs(struct pci_dev *dev)
|
||||
static inline unsigned int pci_sriov_get_totalvfs(struct pci_dev *dev)
|
||||
{ return 0; }
|
||||
#define pci_sriov_configure_simple NULL
|
||||
static inline resource_size_t pci_iov_resource_size(const struct pci_dev *dev,
|
||||
|
|
|
|||
|
|
@ -19,6 +19,19 @@ __rust_helper void rust_helper_iounmap(void __iomem *addr)
|
|||
iounmap(addr);
|
||||
}
|
||||
|
||||
__rust_helper void rust_helper_memcpy_fromio(void *dst,
|
||||
const volatile void __iomem *src,
|
||||
size_t count)
|
||||
{
|
||||
memcpy_fromio(dst, src, count);
|
||||
}
|
||||
|
||||
__rust_helper void rust_helper_memcpy_toio(volatile void __iomem *dst,
|
||||
const void *src, size_t count)
|
||||
{
|
||||
memcpy_toio(dst, src, count);
|
||||
}
|
||||
|
||||
__rust_helper u8 rust_helper_readb(const void __iomem *addr)
|
||||
{
|
||||
return readb(addr);
|
||||
|
|
|
|||
|
|
@ -24,6 +24,14 @@ __rust_helper bool rust_helper_dev_is_pci(const struct device *dev)
|
|||
return dev_is_pci(dev);
|
||||
}
|
||||
|
||||
#ifndef CONFIG_PCI_IOV
|
||||
__rust_helper unsigned int
|
||||
rust_helper_pci_sriov_get_totalvfs(struct pci_dev *pdev)
|
||||
{
|
||||
return pci_sriov_get_totalvfs(pdev);
|
||||
}
|
||||
#endif
|
||||
|
||||
#ifndef CONFIG_PCI_MSI
|
||||
__rust_helper int rust_helper_pci_alloc_irq_vectors(struct pci_dev *dev,
|
||||
unsigned int min_vecs,
|
||||
|
|
|
|||
|
|
@ -9,6 +9,7 @@
|
|||
Vmalloc,
|
||||
VmallocPageIter, //
|
||||
},
|
||||
flags::__GFP_ZERO,
|
||||
layout::ArrayLayout,
|
||||
AllocError,
|
||||
Allocator,
|
||||
|
|
@ -51,6 +52,8 @@
|
|||
}, //
|
||||
};
|
||||
|
||||
use pin_init::Zeroable;
|
||||
|
||||
mod errors;
|
||||
pub use self::errors::{InsertError, PushError, RemoveError};
|
||||
|
||||
|
|
@ -532,6 +535,30 @@ pub fn with_capacity(capacity: usize, flags: Flags) -> Result<Self, AllocError>
|
|||
Ok(v)
|
||||
}
|
||||
|
||||
/// Creates a new [`Vec`] with `n` zero-initialized elements.
|
||||
///
|
||||
/// # Examples
|
||||
///
|
||||
/// ```
|
||||
/// let v = KVec::<u32>::zeroed(20, GFP_KERNEL)?;
|
||||
///
|
||||
/// assert!(v.iter().all(|&x| x == 0));
|
||||
/// # Ok::<(), Error>(())
|
||||
/// ```
|
||||
pub fn zeroed(n: usize, flags: Flags) -> Result<Self, AllocError>
|
||||
where
|
||||
T: Zeroable,
|
||||
{
|
||||
let mut v = Self::with_capacity(n, flags | __GFP_ZERO)?;
|
||||
|
||||
// SAFETY:
|
||||
// - `n <= capacity - len`: `with_capacity(n)` guarantees capacity >= n, len is 0.
|
||||
// - All elements in `[0, n)` are initialized: `__GFP_ZERO` zeroes the allocation,
|
||||
// and `T: Zeroable` guarantees all-zeroes is a valid bit pattern.
|
||||
unsafe { v.inc_len(n) };
|
||||
Ok(v)
|
||||
}
|
||||
|
||||
/// Creates a `Vec<T, A>` from a pointer, a length and a capacity using the allocator `A`.
|
||||
///
|
||||
/// # Examples
|
||||
|
|
|
|||
|
|
@ -68,17 +68,19 @@ struct Inner<T> {
|
|||
/// devres::Devres,
|
||||
/// io::{
|
||||
/// Io,
|
||||
/// IoKnownSize,
|
||||
/// IoBase,
|
||||
/// Mmio,
|
||||
/// MmioRaw,
|
||||
/// PhysAddr, //
|
||||
/// MmioBackend,
|
||||
/// PhysAddr,
|
||||
/// Region, //
|
||||
/// },
|
||||
/// prelude::*,
|
||||
/// };
|
||||
/// use core::ops::Deref;
|
||||
///
|
||||
/// // See also [`pci::Bar`] for a real example.
|
||||
/// struct IoMem<const SIZE: usize>(MmioRaw<SIZE>);
|
||||
/// struct IoMem<const SIZE: usize>(MmioRaw<Region<SIZE>>);
|
||||
///
|
||||
/// impl<const SIZE: usize> IoMem<SIZE> {
|
||||
/// /// # Safety
|
||||
|
|
@ -93,7 +95,7 @@ struct Inner<T> {
|
|||
/// return Err(ENOMEM);
|
||||
/// }
|
||||
///
|
||||
/// Ok(IoMem(MmioRaw::new(addr as usize, SIZE)?))
|
||||
/// Ok(IoMem(MmioRaw::new_region(addr as usize, SIZE)?))
|
||||
/// }
|
||||
/// }
|
||||
///
|
||||
|
|
@ -104,12 +106,13 @@ struct Inner<T> {
|
|||
/// }
|
||||
/// }
|
||||
///
|
||||
/// impl<const SIZE: usize> Deref for IoMem<SIZE> {
|
||||
/// type Target = Mmio<SIZE>;
|
||||
/// impl<'a, const SIZE: usize> IoBase<'a> for &'a IoMem<SIZE> {
|
||||
/// type Backend = MmioBackend;
|
||||
/// type Target = Region<SIZE>;
|
||||
///
|
||||
/// fn deref(&self) -> &Self::Target {
|
||||
/// fn as_view(self) -> Mmio<'a, Region<SIZE>> {
|
||||
/// // SAFETY: The memory range stored in `self` has been properly mapped in `Self::new`.
|
||||
/// unsafe { Mmio::from_raw(&self.0) }
|
||||
/// unsafe { Mmio::from_raw(self.0) }
|
||||
/// }
|
||||
/// }
|
||||
/// # fn no_run(dev: &Device<Bound>) -> Result<(), Error> {
|
||||
|
|
@ -297,10 +300,7 @@ pub fn device(&self) -> &Device {
|
|||
/// use kernel::{
|
||||
/// device::Core,
|
||||
/// devres::Devres,
|
||||
/// io::{
|
||||
/// Io,
|
||||
/// IoKnownSize, //
|
||||
/// },
|
||||
/// io::Io,
|
||||
/// pci, //
|
||||
/// };
|
||||
///
|
||||
|
|
|
|||
|
|
@ -14,14 +14,22 @@
|
|||
},
|
||||
error::to_result,
|
||||
fs::file,
|
||||
io::{
|
||||
IoBackend,
|
||||
IoBase,
|
||||
IoCapable,
|
||||
IoCopyable,
|
||||
SysMem,
|
||||
SysMemBackend, //
|
||||
},
|
||||
prelude::*,
|
||||
ptr::KnownSize,
|
||||
sync::aref::ARef,
|
||||
transmute::{
|
||||
AsBytes,
|
||||
FromBytes, //
|
||||
}, //
|
||||
uaccess::UserSliceWriter,
|
||||
},
|
||||
uaccess::UserSliceWriter, //
|
||||
};
|
||||
use core::{
|
||||
ops::{
|
||||
|
|
@ -577,7 +585,7 @@ fn from(value: CoherentBox<T>) -> Self {
|
|||
/// # Invariants
|
||||
///
|
||||
/// - For the lifetime of an instance of [`Coherent`], the `cpu_addr` is a valid pointer
|
||||
/// to an allocated region of coherent memory and `dma_handle` is the DMA address base of the
|
||||
/// to an allocated region of coherent memory and `dma_addr` is the DMA address base of the
|
||||
/// region.
|
||||
/// - The size in bytes of the allocation is equal to size information via pointer.
|
||||
// TODO
|
||||
|
|
@ -594,7 +602,7 @@ fn from(value: CoherentBox<T>) -> Self {
|
|||
// entire `Coherent` including the allocated memory itself.
|
||||
pub struct Coherent<T: KnownSize + ?Sized> {
|
||||
dev: ARef<device::Device>,
|
||||
dma_handle: DmaAddress,
|
||||
dma_addr: DmaAddress,
|
||||
cpu_addr: NonNull<T>,
|
||||
dma_attrs: Attrs,
|
||||
}
|
||||
|
|
@ -619,11 +627,10 @@ pub fn as_mut_ptr(&self) -> *mut T {
|
|||
self.cpu_addr.as_ptr()
|
||||
}
|
||||
|
||||
/// Returns a DMA handle which may be given to the device as the DMA address base of
|
||||
/// the region.
|
||||
/// Returns a DMA address which may be given to the device as the base of the region.
|
||||
#[inline]
|
||||
pub fn dma_handle(&self) -> DmaAddress {
|
||||
self.dma_handle
|
||||
pub fn dma_address(&self) -> DmaAddress {
|
||||
self.dma_addr
|
||||
}
|
||||
|
||||
/// Returns a reference to the data in the region.
|
||||
|
|
@ -654,52 +661,6 @@ pub unsafe fn as_mut(&self) -> &mut T {
|
|||
// SAFETY: per safety requirement.
|
||||
unsafe { &mut *self.as_mut_ptr() }
|
||||
}
|
||||
|
||||
/// Reads the value of `field` and ensures that its type is [`FromBytes`].
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// This must be called from the [`dma_read`] macro which ensures that the `field` pointer is
|
||||
/// validated beforehand.
|
||||
///
|
||||
/// Public but hidden since it should only be used from [`dma_read`] macro.
|
||||
#[doc(hidden)]
|
||||
pub unsafe fn field_read<F: FromBytes>(&self, field: *const F) -> F {
|
||||
// SAFETY:
|
||||
// - By the safety requirements field is valid.
|
||||
// - Using read_volatile() here is not sound as per the usual rules, the usage here is
|
||||
// a special exception with the following notes in place. When dealing with a potential
|
||||
// race from a hardware or code outside kernel (e.g. user-space program), we need that
|
||||
// read on a valid memory is not UB. Currently read_volatile() is used for this, and the
|
||||
// rationale behind is that it should generate the same code as READ_ONCE() which the
|
||||
// kernel already relies on to avoid UB on data races. Note that the usage of
|
||||
// read_volatile() is limited to this particular case, it cannot be used to prevent
|
||||
// the UB caused by racing between two kernel functions nor do they provide atomicity.
|
||||
unsafe { field.read_volatile() }
|
||||
}
|
||||
|
||||
/// Writes a value to `field` and ensures that its type is [`AsBytes`].
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// This must be called from the [`dma_write`] macro which ensures that the `field` pointer is
|
||||
/// validated beforehand.
|
||||
///
|
||||
/// Public but hidden since it should only be used from [`dma_write`] macro.
|
||||
#[doc(hidden)]
|
||||
pub unsafe fn field_write<F: AsBytes>(&self, field: *mut F, val: F) {
|
||||
// SAFETY:
|
||||
// - By the safety requirements field is valid.
|
||||
// - Using write_volatile() here is not sound as per the usual rules, the usage here is
|
||||
// a special exception with the following notes in place. When dealing with a potential
|
||||
// race from a hardware or code outside kernel (e.g. user-space program), we need that
|
||||
// write on a valid memory is not UB. Currently write_volatile() is used for this, and the
|
||||
// rationale behind is that it should generate the same code as WRITE_ONCE() which the
|
||||
// kernel already relies on to avoid UB on data races. Note that the usage of
|
||||
// write_volatile() is limited to this particular case, it cannot be used to prevent
|
||||
// the UB caused by racing between two kernel functions nor do they provide atomicity.
|
||||
unsafe { field.write_volatile(val) }
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: AsBytes + FromBytes> Coherent<T> {
|
||||
|
|
@ -716,13 +677,13 @@ fn alloc_with_attrs(
|
|||
);
|
||||
}
|
||||
|
||||
let mut dma_handle = 0;
|
||||
let mut dma_addr = 0;
|
||||
// SAFETY: Device pointer is guaranteed as valid by the type invariant on `Device`.
|
||||
let addr = unsafe {
|
||||
bindings::dma_alloc_attrs(
|
||||
dev.as_raw(),
|
||||
core::mem::size_of::<T>(),
|
||||
&mut dma_handle,
|
||||
&mut dma_addr,
|
||||
gfp_flags.as_raw(),
|
||||
dma_attrs.as_raw(),
|
||||
)
|
||||
|
|
@ -734,7 +695,7 @@ fn alloc_with_attrs(
|
|||
// - We also hold a refcounted reference to the device.
|
||||
Ok(Self {
|
||||
dev: dev.into(),
|
||||
dma_handle,
|
||||
dma_addr,
|
||||
cpu_addr,
|
||||
dma_attrs,
|
||||
})
|
||||
|
|
@ -833,13 +794,13 @@ fn alloc_slice_with_attrs(
|
|||
}
|
||||
|
||||
let size = core::mem::size_of::<T>().checked_mul(len).ok_or(ENOMEM)?;
|
||||
let mut dma_handle = 0;
|
||||
let mut dma_addr = 0;
|
||||
// SAFETY: Device pointer is guaranteed as valid by the type invariant on `Device`.
|
||||
let addr = unsafe {
|
||||
bindings::dma_alloc_attrs(
|
||||
dev.as_raw(),
|
||||
size,
|
||||
&mut dma_handle,
|
||||
&mut dma_addr,
|
||||
gfp_flags.as_raw(),
|
||||
dma_attrs.as_raw(),
|
||||
)
|
||||
|
|
@ -851,7 +812,7 @@ fn alloc_slice_with_attrs(
|
|||
// - We also hold a refcounted reference to the device.
|
||||
Ok(Coherent {
|
||||
dev: dev.into(),
|
||||
dma_handle,
|
||||
dma_addr,
|
||||
cpu_addr,
|
||||
dma_attrs,
|
||||
})
|
||||
|
|
@ -965,14 +926,14 @@ impl<T: KnownSize + ?Sized> Drop for Coherent<T> {
|
|||
fn drop(&mut self) {
|
||||
let size = T::size(self.cpu_addr.as_ptr());
|
||||
// SAFETY: Device pointer is guaranteed as valid by the type invariant on `Device`.
|
||||
// The cpu address, and the dma handle are valid due to the type invariants on
|
||||
// The cpu address, and the dma address are valid due to the type invariants on
|
||||
// `Coherent`.
|
||||
unsafe {
|
||||
bindings::dma_free_attrs(
|
||||
self.dev.as_raw(),
|
||||
size,
|
||||
self.cpu_addr.as_ptr().cast(),
|
||||
self.dma_handle,
|
||||
self.dma_addr,
|
||||
self.dma_attrs.as_raw(),
|
||||
)
|
||||
}
|
||||
|
|
@ -1027,13 +988,13 @@ fn write_to_slice(
|
|||
///
|
||||
/// - `cpu_handle` holds the opaque handle returned by `dma_alloc_attrs` with
|
||||
/// `DMA_ATTR_NO_KERNEL_MAPPING` set, and is only valid for passing back to `dma_free_attrs`.
|
||||
/// - `dma_handle` is the corresponding bus address for device DMA.
|
||||
/// - `dma_addr` is the corresponding bus address for device DMA.
|
||||
/// - `size` is the allocation size in bytes as passed to `dma_alloc_attrs`.
|
||||
/// - `dma_attrs` contains the attributes used for the allocation, always including
|
||||
/// `DMA_ATTR_NO_KERNEL_MAPPING`.
|
||||
pub struct CoherentHandle {
|
||||
dev: ARef<device::Device>,
|
||||
dma_handle: DmaAddress,
|
||||
dma_addr: DmaAddress,
|
||||
cpu_handle: NonNull<c_void>,
|
||||
size: usize,
|
||||
dma_attrs: Attrs,
|
||||
|
|
@ -1057,13 +1018,13 @@ pub fn alloc_with_attrs(
|
|||
}
|
||||
|
||||
let dma_attrs = dma_attrs | Attrs(bindings::DMA_ATTR_NO_KERNEL_MAPPING);
|
||||
let mut dma_handle = 0;
|
||||
let mut dma_addr = 0;
|
||||
// SAFETY: `dev.as_raw()` is valid by the type invariant on `device::Device`.
|
||||
let cpu_handle = unsafe {
|
||||
bindings::dma_alloc_attrs(
|
||||
dev.as_raw(),
|
||||
size,
|
||||
&mut dma_handle,
|
||||
&mut dma_addr,
|
||||
gfp_flags.as_raw(),
|
||||
dma_attrs.as_raw(),
|
||||
)
|
||||
|
|
@ -1072,11 +1033,11 @@ pub fn alloc_with_attrs(
|
|||
let cpu_handle = NonNull::new(cpu_handle).ok_or(ENOMEM)?;
|
||||
|
||||
// INVARIANT: `cpu_handle` is the opaque handle from a successful `dma_alloc_attrs` call
|
||||
// with `DMA_ATTR_NO_KERNEL_MAPPING`, `dma_handle` is the corresponding DMA address,
|
||||
// with `DMA_ATTR_NO_KERNEL_MAPPING`, `dma_addr` is the corresponding DMA address,
|
||||
// and we hold a refcounted reference to the device.
|
||||
Ok(Self {
|
||||
dev: dev.into(),
|
||||
dma_handle,
|
||||
dma_addr,
|
||||
cpu_handle,
|
||||
size,
|
||||
dma_attrs,
|
||||
|
|
@ -1093,12 +1054,12 @@ pub fn alloc(
|
|||
Self::alloc_with_attrs(dev, size, gfp_flags, Attrs(0))
|
||||
}
|
||||
|
||||
/// Returns the DMA handle for this allocation.
|
||||
/// Returns the DMA address for this allocation.
|
||||
///
|
||||
/// This address can be programmed into device hardware for DMA access.
|
||||
#[inline]
|
||||
pub fn dma_handle(&self) -> DmaAddress {
|
||||
self.dma_handle
|
||||
pub fn dma_address(&self) -> DmaAddress {
|
||||
self.dma_addr
|
||||
}
|
||||
|
||||
/// Returns the size in bytes of this allocation.
|
||||
|
|
@ -1117,100 +1078,170 @@ fn drop(&mut self) {
|
|||
self.dev.as_raw(),
|
||||
self.size,
|
||||
self.cpu_handle.as_ptr(),
|
||||
self.dma_handle,
|
||||
self.dma_addr,
|
||||
self.dma_attrs.as_raw(),
|
||||
)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: `CoherentHandle` only holds a device reference, a DMA handle, an opaque CPU handle,
|
||||
// SAFETY: `CoherentHandle` only holds a device reference, a DMA address, an opaque CPU handle,
|
||||
// and a size. None of these are tied to a specific thread.
|
||||
unsafe impl Send for CoherentHandle {}
|
||||
|
||||
// SAFETY: `CoherentHandle` provides no CPU access to the underlying allocation. The only
|
||||
// operations on `&CoherentHandle` are reading the DMA handle and size, both of which are
|
||||
// operations on `&CoherentHandle` are reading the DMA address and size, both of which are
|
||||
// plain `Copy` values.
|
||||
unsafe impl Sync for CoherentHandle {}
|
||||
|
||||
/// Reads a field of an item from an allocated region of structs.
|
||||
/// View type for `Coherent`.
|
||||
///
|
||||
/// The syntax is of the form `kernel::dma_read!(dma, proj)` where `dma` is an expression evaluating
|
||||
/// to a [`Coherent`] and `proj` is a [projection specification](kernel::ptr::project!).
|
||||
///
|
||||
/// # Examples
|
||||
///
|
||||
/// ```
|
||||
/// use kernel::device::Device;
|
||||
/// use kernel::dma::{attrs::*, Coherent};
|
||||
///
|
||||
/// struct MyStruct { field: u32, }
|
||||
///
|
||||
/// // SAFETY: All bit patterns are acceptable values for `MyStruct`.
|
||||
/// unsafe impl kernel::transmute::FromBytes for MyStruct{};
|
||||
/// // SAFETY: Instances of `MyStruct` have no uninitialized portions.
|
||||
/// unsafe impl kernel::transmute::AsBytes for MyStruct{};
|
||||
///
|
||||
/// # fn test(alloc: &kernel::dma::Coherent<[MyStruct]>) -> Result {
|
||||
/// let whole = kernel::dma_read!(alloc, [try: 2]);
|
||||
/// let field = kernel::dma_read!(alloc, [panic: 1].field);
|
||||
/// # Ok::<(), Error>(()) }
|
||||
/// ```
|
||||
#[macro_export]
|
||||
macro_rules! dma_read {
|
||||
($dma:expr, $($proj:tt)*) => {{
|
||||
let dma = &$dma;
|
||||
let ptr = $crate::ptr::project!(
|
||||
$crate::dma::Coherent::as_ptr(dma), $($proj)*
|
||||
);
|
||||
// SAFETY: The pointer created by the projection is within the DMA region.
|
||||
unsafe { $crate::dma::Coherent::field_read(dma, ptr) }
|
||||
}};
|
||||
/// This is same as [`SysMem`] but with additional information that allows handing out a DMA
|
||||
/// address.
|
||||
pub struct CoherentView<'a, T: ?Sized> {
|
||||
cpu_addr: SysMem<'a, T>,
|
||||
dma_addr: DmaAddress,
|
||||
}
|
||||
|
||||
/// Writes to a field of an item from an allocated region of structs.
|
||||
///
|
||||
/// The syntax is of the form `kernel::dma_write!(dma, proj, val)` where `dma` is an expression
|
||||
/// evaluating to a [`Coherent`], `proj` is a
|
||||
/// [projection specification](kernel::ptr::project!), and `val` is the value to be written to the
|
||||
/// projected location.
|
||||
///
|
||||
/// # Examples
|
||||
///
|
||||
/// ```
|
||||
/// use kernel::device::Device;
|
||||
/// use kernel::dma::{attrs::*, Coherent};
|
||||
///
|
||||
/// struct MyStruct { member: u32, }
|
||||
///
|
||||
/// // SAFETY: All bit patterns are acceptable values for `MyStruct`.
|
||||
/// unsafe impl kernel::transmute::FromBytes for MyStruct{};
|
||||
/// // SAFETY: Instances of `MyStruct` have no uninitialized portions.
|
||||
/// unsafe impl kernel::transmute::AsBytes for MyStruct{};
|
||||
///
|
||||
/// # fn test(alloc: &kernel::dma::Coherent<[MyStruct]>) -> Result {
|
||||
/// kernel::dma_write!(alloc, [try: 2].member, 0xf);
|
||||
/// kernel::dma_write!(alloc, [panic: 1], MyStruct { member: 0xf });
|
||||
/// # Ok::<(), Error>(()) }
|
||||
/// ```
|
||||
#[macro_export]
|
||||
macro_rules! dma_write {
|
||||
(@parse [$dma:expr] [$($proj:tt)*] [, $val:expr]) => {{
|
||||
let dma = &$dma;
|
||||
let ptr = $crate::ptr::project!(
|
||||
mut $crate::dma::Coherent::as_mut_ptr(dma), $($proj)*
|
||||
);
|
||||
let val = $val;
|
||||
// SAFETY: The pointer created by the projection is within the DMA region.
|
||||
unsafe { $crate::dma::Coherent::field_write(dma, ptr, val) }
|
||||
}};
|
||||
(@parse [$dma:expr] [$($proj:tt)*] [.$field:tt $($rest:tt)*]) => {
|
||||
$crate::dma_write!(@parse [$dma] [$($proj)* .$field] [$($rest)*])
|
||||
};
|
||||
(@parse [$dma:expr] [$($proj:tt)*] [[$flavor:ident: $index:expr] $($rest:tt)*]) => {
|
||||
$crate::dma_write!(@parse [$dma] [$($proj)* [$flavor: $index]] [$($rest)*])
|
||||
};
|
||||
($dma:expr, $($rest:tt)*) => {
|
||||
$crate::dma_write!(@parse [$dma] [] [$($rest)*])
|
||||
};
|
||||
impl<T: ?Sized> Copy for CoherentView<'_, T> {}
|
||||
impl<T: ?Sized> Clone for CoherentView<'_, T> {
|
||||
#[inline]
|
||||
fn clone(&self) -> Self {
|
||||
*self
|
||||
}
|
||||
}
|
||||
|
||||
impl<'a, T: ?Sized> CoherentView<'a, T> {
|
||||
/// Erase the DMA address information and obtain a [`SysMem`] view of the same memory region.
|
||||
#[inline]
|
||||
pub fn as_sys_mem(self) -> SysMem<'a, T> {
|
||||
self.cpu_addr
|
||||
}
|
||||
|
||||
/// Returns the DMA address which may be given to the device as base of the region.
|
||||
#[inline]
|
||||
pub fn dma_address(self) -> DmaAddress {
|
||||
self.dma_addr
|
||||
}
|
||||
|
||||
/// Returns a reference to the data in the region.
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// * Callers must ensure that the device does not read/write to/from memory while the returned
|
||||
/// reference is live.
|
||||
/// * Callers must ensure that this call does not race with a write (including call to `as_mut`)
|
||||
/// to the same region while the returned reference is live.
|
||||
#[inline]
|
||||
pub unsafe fn as_ref(self) -> &'a T {
|
||||
// SAFETY: pointer is aligned and valid per type invariant. Aliasing rule is satisfied per
|
||||
// safety requirement.
|
||||
unsafe { &*self.cpu_addr.as_ptr() }
|
||||
}
|
||||
|
||||
/// Returns a mutable reference to the data in the region.
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// * Callers must ensure that the device does not read/write to/from memory while the returned
|
||||
/// reference is live.
|
||||
/// * Callers must ensure that this call does not race with a read (including call to `as_ref`)
|
||||
/// or write (including call to `as_mut`) to the same region while the returned reference is
|
||||
/// live.
|
||||
#[inline]
|
||||
pub unsafe fn as_mut(self) -> &'a mut T {
|
||||
// SAFETY: pointer is aligned and valid per type invariant. Aliasing rule is satisfied per
|
||||
// safety requirement.
|
||||
unsafe { &mut *self.cpu_addr.as_ptr() }
|
||||
}
|
||||
}
|
||||
|
||||
/// `IoBackend` implementation for `Coherent`.
|
||||
pub struct CoherentIoBackend;
|
||||
|
||||
impl IoBackend for CoherentIoBackend {
|
||||
type View<'a, T: ?Sized + KnownSize> = CoherentView<'a, T>;
|
||||
|
||||
#[inline]
|
||||
fn as_ptr<'a, T: ?Sized + KnownSize>(view: Self::View<'a, T>) -> *mut T {
|
||||
SysMemBackend::as_ptr(view.cpu_addr)
|
||||
}
|
||||
|
||||
#[inline]
|
||||
unsafe fn project_view<'a, T: ?Sized + KnownSize, U: ?Sized + KnownSize>(
|
||||
view: Self::View<'a, T>,
|
||||
ptr: *mut U,
|
||||
) -> Self::View<'a, U> {
|
||||
let offset = ptr.addr() - view.cpu_addr.as_ptr().addr();
|
||||
// CAST: The offset DMA address can never overflow.
|
||||
let dma_addr = view.dma_addr + offset as DmaAddress;
|
||||
CoherentView {
|
||||
dma_addr,
|
||||
// SAFETY: Per safety requirement.
|
||||
cpu_addr: unsafe { SysMemBackend::project_view(view.cpu_addr, ptr) },
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl<T> IoCapable<T> for CoherentIoBackend
|
||||
where
|
||||
SysMemBackend: IoCapable<T>,
|
||||
{
|
||||
#[inline]
|
||||
fn io_read<'a>(view: Self::View<'a, T>) -> T {
|
||||
SysMemBackend::io_read(view.cpu_addr)
|
||||
}
|
||||
|
||||
#[inline]
|
||||
fn io_write<'a>(view: Self::View<'a, T>, value: T) {
|
||||
SysMemBackend::io_write(view.cpu_addr, value)
|
||||
}
|
||||
}
|
||||
|
||||
impl IoCopyable for CoherentIoBackend {
|
||||
#[inline]
|
||||
unsafe fn copy_from_io(view: Self::View<'_, [u8]>, buffer: *mut u8) {
|
||||
// SAFETY: Per safety requirement.
|
||||
unsafe { SysMemBackend::copy_from_io(view.cpu_addr, buffer) }
|
||||
}
|
||||
|
||||
#[inline]
|
||||
unsafe fn copy_to_io(view: Self::View<'_, [u8]>, buffer: *const u8) {
|
||||
// SAFETY: Per safety requirement.
|
||||
unsafe { SysMemBackend::copy_to_io(view.cpu_addr, buffer) }
|
||||
}
|
||||
|
||||
#[inline]
|
||||
fn copy_read<T: zerocopy::FromBytes>(view: Self::View<'_, T>) -> T {
|
||||
SysMemBackend::copy_read(view.cpu_addr)
|
||||
}
|
||||
|
||||
#[inline]
|
||||
fn copy_write<T: zerocopy::IntoBytes>(view: Self::View<'_, T>, value: T) {
|
||||
SysMemBackend::copy_write(view.cpu_addr, value)
|
||||
}
|
||||
}
|
||||
|
||||
impl<'a, T: ?Sized + KnownSize> IoBase<'a> for CoherentView<'a, T> {
|
||||
type Backend = CoherentIoBackend;
|
||||
type Target = T;
|
||||
|
||||
#[inline]
|
||||
fn as_view(self) -> CoherentView<'a, Self::Target> {
|
||||
self
|
||||
}
|
||||
}
|
||||
|
||||
impl<'a, T: ?Sized + KnownSize> IoBase<'a> for &'a Coherent<T> {
|
||||
type Backend = CoherentIoBackend;
|
||||
type Target = T;
|
||||
|
||||
#[inline]
|
||||
fn as_view(self) -> CoherentView<'a, Self::Target> {
|
||||
CoherentView {
|
||||
// SAFETY: `cpu_addr` is valid and aligned kernel accessible memory.
|
||||
cpu_addr: unsafe { SysMem::new(self.cpu_addr.as_ptr()) },
|
||||
dma_addr: self.dma_addr,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -32,6 +32,7 @@
|
|||
};
|
||||
use core::{
|
||||
alloc::Layout,
|
||||
cell::UnsafeCell,
|
||||
marker::PhantomData,
|
||||
mem,
|
||||
ops::Deref,
|
||||
|
|
@ -74,66 +75,59 @@ macro_rules! drm_legacy_fields {
|
|||
|
||||
/// A trait implemented by all possible contexts a [`Device`] can be used in.
|
||||
///
|
||||
/// Setting up a new [`Device`] is a multi-stage process. Each step of the process that a user
|
||||
/// interacts with in Rust has a respective [`DeviceContext`] typestate. For example,
|
||||
/// `Device<T, Registered>` would be a [`Device`] that reached the [`Registered`] [`DeviceContext`].
|
||||
/// A [`Device`] can be in one of the following contexts:
|
||||
///
|
||||
/// Each stage of this process is described below:
|
||||
/// - [`Normal`]: The general-purpose, reference-counted context. A [`Device`] in this context may
|
||||
/// or may not be registered with userspace.
|
||||
/// - [`Ioctl`]: The device has been registered with userspace at some point; used in ioctl
|
||||
/// dispatch context.
|
||||
/// - [`Registered`]: The device is currently registered with userspace and the parent bus device
|
||||
/// is bound.
|
||||
///
|
||||
/// ```text
|
||||
/// 1 2 3
|
||||
/// +--------------+ +------------------+ +-----------------------+
|
||||
/// |Device created| → |Device initialized| → |Registered w/ userspace|
|
||||
/// +--------------+ +------------------+ +-----------------------+
|
||||
/// (Uninit) (Registered)
|
||||
/// ```
|
||||
///
|
||||
/// 1. The [`Device`] is in the [`Uninit`] context and is not guaranteed to be initialized or
|
||||
/// registered with userspace. Only a limited subset of DRM core functionality is available.
|
||||
/// 2. The [`Device`] is guaranteed to be fully initialized, but is not guaranteed to be registered
|
||||
/// with userspace. All DRM core functionality which doesn't interact with userspace is
|
||||
/// available. We currently don't have a context for representing this.
|
||||
/// 3. The [`Device`] is guaranteed to be fully initialized, and is guaranteed to have been
|
||||
/// registered with userspace at some point - thus putting it in the [`Registered`] context.
|
||||
///
|
||||
/// An important caveat of [`DeviceContext`] which must be kept in mind: when used as a typestate
|
||||
/// for a reference type, it can only guarantee that a [`Device`] reached a particular stage in the
|
||||
/// initialization process _at the time the reference was taken_. No guarantee is made in regards to
|
||||
/// what stage of the process the [`Device`] is currently in. This means for instance that a
|
||||
/// `&Device<T, Uninit>` may actually be registered with userspace, it just wasn't known to be
|
||||
/// registered at the time the reference was taken.
|
||||
pub trait DeviceContext: Sealed + Send + Sync {}
|
||||
/// Both `Device<T, Ioctl>` and `Device<T, Registered>` dereference to `Device<T>` ([`Normal`]),
|
||||
/// so any method available on a [`Normal`] device is also available in the other contexts.
|
||||
pub trait DeviceContext: Sealed + Send + Sync + 'static {}
|
||||
|
||||
/// The [`DeviceContext`] of a [`Device`] that was registered with userspace at some point.
|
||||
/// The general-purpose, reference-counted [`DeviceContext`].
|
||||
///
|
||||
/// This represents a [`Device`] which is guaranteed to have been registered with userspace at
|
||||
/// some point in time. Such a DRM device is guaranteed to have been fully-initialized.
|
||||
/// A [`Device`] in this context may or may not be registered with userspace. This context is used
|
||||
/// for reference-counted device handles and during device setup via [`UnregisteredDevice`].
|
||||
///
|
||||
/// Note: A device in this context is not guaranteed to remain registered with userspace for its
|
||||
/// entire lifetime, as this is impossible to guarantee at compile-time.
|
||||
/// [`AlwaysRefCounted`] is only implemented for `Device<T, Normal>`, making this the required
|
||||
/// context for [`ARef`]-based device handles.
|
||||
pub struct Normal;
|
||||
|
||||
impl Sealed for Normal {}
|
||||
impl DeviceContext for Normal {}
|
||||
|
||||
/// The [`DeviceContext`] of a [`Device`] that is currently registered with userspace.
|
||||
///
|
||||
/// A [`Device`] in this context is guaranteed to be registered and its parent bus device is
|
||||
/// guaranteed to be bound. This is enforced at runtime by [`RegistrationGuard`], which holds a
|
||||
/// `drm_dev_enter()` / `drm_dev_exit()` SRCU critical section.
|
||||
///
|
||||
/// # Invariants
|
||||
///
|
||||
/// A [`Device`] in this [`DeviceContext`] is guaranteed to have been registered with userspace
|
||||
/// at some point in time.
|
||||
/// The parent bus device is bound for the duration of any reference to a `Device<T, Registered>`.
|
||||
pub struct Registered;
|
||||
|
||||
impl Sealed for Registered {}
|
||||
impl DeviceContext for Registered {}
|
||||
|
||||
/// The [`DeviceContext`] of a [`Device`] that may be unregistered and partly uninitialized.
|
||||
/// The [`DeviceContext`] of a [`Device`] that has been registered with userspace previously.
|
||||
///
|
||||
/// A [`Device`] in this context is only guaranteed to be partly initialized, and may or may not
|
||||
/// be registered with userspace. Thus operations which depend on the [`Device`] being fully
|
||||
/// initialized, or which depend on the [`Device`] being registered with userspace are not
|
||||
/// available through this [`DeviceContext`].
|
||||
/// A [`Device`] in this context has been registered at some point, but may be concurrently
|
||||
/// unregistering or already unregistered. `drm_dev_enter()` can guard against this, ensuring the
|
||||
/// device remains registered for the duration of the critical section.
|
||||
///
|
||||
/// A [`Device`] in this context can be used to create a
|
||||
/// [`Registration`](drm::driver::Registration).
|
||||
pub struct Uninit;
|
||||
/// # Invariants
|
||||
///
|
||||
/// A [`Device`] in this context has been registered with userspace via `drm_dev_register()` at
|
||||
/// some point.
|
||||
pub struct Ioctl;
|
||||
|
||||
impl Sealed for Uninit {}
|
||||
impl DeviceContext for Uninit {}
|
||||
impl Sealed for Ioctl {}
|
||||
impl DeviceContext for Ioctl {}
|
||||
|
||||
/// A [`Device`] which is known at compile-time to be unregistered with userspace.
|
||||
///
|
||||
|
|
@ -147,10 +141,10 @@ impl DeviceContext for Uninit {}
|
|||
///
|
||||
/// The device in `self.0` is guaranteed to be a newly created [`Device`] that has not yet been
|
||||
/// registered with userspace until this type is dropped.
|
||||
pub struct UnregisteredDevice<T: drm::Driver>(ARef<Device<T, Uninit>>, NotThreadSafe);
|
||||
pub struct UnregisteredDevice<T: drm::Driver>(ARef<Device<T, Normal>>, NotThreadSafe);
|
||||
|
||||
impl<T: drm::Driver> Deref for UnregisteredDevice<T> {
|
||||
type Target = Device<T, Uninit>;
|
||||
type Target = Device<T, Normal>;
|
||||
|
||||
fn deref(&self) -> &Self::Target {
|
||||
&self.0
|
||||
|
|
@ -178,15 +172,13 @@ const fn compute_features() -> u32 {
|
|||
master_drop: None,
|
||||
debugfs_init: None,
|
||||
|
||||
// Ignore the Uninit DeviceContext below. It is only provided because it is required by the
|
||||
// compiler, and it is not actually used by these functions.
|
||||
gem_create_object: T::Object::<Uninit>::ALLOC_OPS.gem_create_object,
|
||||
prime_handle_to_fd: T::Object::<Uninit>::ALLOC_OPS.prime_handle_to_fd,
|
||||
prime_fd_to_handle: T::Object::<Uninit>::ALLOC_OPS.prime_fd_to_handle,
|
||||
gem_prime_import: T::Object::<Uninit>::ALLOC_OPS.gem_prime_import,
|
||||
gem_prime_import_sg_table: T::Object::<Uninit>::ALLOC_OPS.gem_prime_import_sg_table,
|
||||
dumb_create: T::Object::<Uninit>::ALLOC_OPS.dumb_create,
|
||||
dumb_map_offset: T::Object::<Uninit>::ALLOC_OPS.dumb_map_offset,
|
||||
gem_create_object: T::Object::ALLOC_OPS.gem_create_object,
|
||||
prime_handle_to_fd: T::Object::ALLOC_OPS.prime_handle_to_fd,
|
||||
prime_fd_to_handle: T::Object::ALLOC_OPS.prime_fd_to_handle,
|
||||
gem_prime_import: T::Object::ALLOC_OPS.gem_prime_import,
|
||||
gem_prime_import_sg_table: T::Object::ALLOC_OPS.gem_prime_import_sg_table,
|
||||
dumb_create: T::Object::ALLOC_OPS.dumb_create,
|
||||
dumb_map_offset: T::Object::ALLOC_OPS.dumb_map_offset,
|
||||
|
||||
show_fdinfo: None,
|
||||
fbdev_probe: None,
|
||||
|
|
@ -208,10 +200,13 @@ const fn compute_features() -> u32 {
|
|||
/// Create a new `UnregisteredDevice` for a `drm::Driver`.
|
||||
///
|
||||
/// This can be used to create a [`Registration`](kernel::drm::Registration).
|
||||
pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<Self> {
|
||||
pub fn new(
|
||||
dev: &T::ParentDevice<device::Bound>,
|
||||
data: impl PinInit<T::Data, Error>,
|
||||
) -> Result<Self> {
|
||||
// `__drm_dev_alloc` uses `kmalloc()` to allocate memory, hence ensure a `kmalloc()`
|
||||
// compatible `Layout`.
|
||||
let layout = Kmalloc::aligned_layout(Layout::new::<Device<T, Uninit>>());
|
||||
let layout = Kmalloc::aligned_layout(Layout::new::<Device<T, Normal>>());
|
||||
|
||||
// Use a temporary vtable without a `release` callback until `data` is initialized, so
|
||||
// init failure can release the DRM device without dropping uninitialized fields.
|
||||
|
|
@ -223,12 +218,12 @@ pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<S
|
|||
// SAFETY:
|
||||
// - `alloc_vtable` reference remains valid until no longer used,
|
||||
// - `dev` is valid by its type invarants,
|
||||
let raw_drm: *mut Device<T, Uninit> = unsafe {
|
||||
let raw_drm: *mut Device<T, Normal> = unsafe {
|
||||
bindings::__drm_dev_alloc(
|
||||
dev.as_raw(),
|
||||
dev.as_ref().as_raw(),
|
||||
&alloc_vtable,
|
||||
layout.size(),
|
||||
mem::offset_of!(Device<T, Uninit>, dev),
|
||||
mem::offset_of!(Device<T, Normal>, dev),
|
||||
)
|
||||
}
|
||||
.cast();
|
||||
|
|
@ -253,6 +248,9 @@ pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<S
|
|||
// SAFETY: `drm_dev` is still private to this function.
|
||||
unsafe { (*drm_dev).driver = const { &Self::VTABLE } };
|
||||
|
||||
// SAFETY: `raw_drm` is valid; no concurrent access before registration.
|
||||
unsafe { (*raw_drm.as_ptr()).registration_data = UnsafeCell::new(NonNull::dangling()) };
|
||||
|
||||
// SAFETY: The reference count is one, and now we take ownership of that reference as a
|
||||
// `drm::Device`.
|
||||
// INVARIANT: We just created the device above, but have yet to call `drm_dev_register`.
|
||||
|
|
@ -264,16 +262,8 @@ pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<S
|
|||
|
||||
/// A typed DRM device with a specific [`drm::Driver`] implementation and [`DeviceContext`].
|
||||
///
|
||||
/// Since DRM devices can be used before being fully initialized and registered with userspace, `C`
|
||||
/// represents the furthest [`DeviceContext`] we can guarantee that this [`Device`] has reached.
|
||||
///
|
||||
/// Keep in mind: this means that an unregistered device can still have the registration state
|
||||
/// [`Registered`] as long as it was registered with userspace once in the past, and that the
|
||||
/// behavior of such a device is still well-defined. Additionally, a device with the registration
|
||||
/// state [`Uninit`] simply does not have a guaranteed registration state at compile time, and could
|
||||
/// be either registered or unregistered. Since there is no way to guarantee a long-lived reference
|
||||
/// to an unregistered device would remain unregistered, we do not provide a [`DeviceContext`] for
|
||||
/// this.
|
||||
/// A device in the [`Registered`] context is currently registered with userspace and its parent
|
||||
/// bus device is bound. The [`Normal`] context is the general-purpose, reference-counted context.
|
||||
///
|
||||
/// # Invariants
|
||||
///
|
||||
|
|
@ -281,9 +271,10 @@ pub fn new(dev: &device::Device, data: impl PinInit<T::Data, Error>) -> Result<S
|
|||
/// * The data layout of `Self` remains the same across all implementations of `C`.
|
||||
/// * Any invariants for `C` also apply.
|
||||
#[repr(C)]
|
||||
pub struct Device<T: drm::Driver, C: DeviceContext = Registered> {
|
||||
pub struct Device<T: drm::Driver, C: DeviceContext = Normal> {
|
||||
dev: Opaque<bindings::drm_device>,
|
||||
data: T::Data,
|
||||
pub(super) registration_data: UnsafeCell<NonNull<T::RegistrationData<'static>>>,
|
||||
_ctx: PhantomData<C>,
|
||||
}
|
||||
|
||||
|
|
@ -352,7 +343,111 @@ pub(crate) unsafe fn assume_ctx<NewCtx: DeviceContext>(&self) -> &Device<T, NewC
|
|||
}
|
||||
}
|
||||
|
||||
impl<T: drm::Driver, C: DeviceContext> Deref for Device<T, C> {
|
||||
impl<T: drm::Driver> Device<T, Ioctl> {
|
||||
/// Guard against the parent bus device being unbound.
|
||||
///
|
||||
/// Returns a [`RegistrationGuard`] if the device has not been unplugged, [`None`] otherwise.
|
||||
///
|
||||
/// While [`RegistrationGuard`] is held the parent device is guaranteed to be bound.
|
||||
#[must_use]
|
||||
pub fn registration_guard(&self) -> Option<RegistrationGuard<'_, T>> {
|
||||
let mut idx: i32 = 0;
|
||||
// SAFETY: `self.as_raw()` is a valid pointer to a `struct drm_device`.
|
||||
if unsafe { bindings::drm_dev_enter(self.as_raw(), &mut idx) } {
|
||||
// INVARIANT:
|
||||
// - `idx` is the SRCU index from the successful `drm_dev_enter()` above.
|
||||
// - The parent bus device is bound: `drm_dev_enter()` succeeded, meaning
|
||||
// `drm_dev_unplug()` has not completed; since it is only called from
|
||||
// `Registration::drop()` during parent unbind, the parent is still bound.
|
||||
Some(RegistrationGuard {
|
||||
// SAFETY: See INVARIANT above; the `Registered` context invariant holds.
|
||||
dev: unsafe { self.assume_ctx() },
|
||||
idx,
|
||||
_not_send: NotThreadSafe,
|
||||
})
|
||||
} else {
|
||||
None
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A guard proving the DRM device is registered and the parent bus device is bound.
|
||||
///
|
||||
/// The guard dereferences to [`Device<T, Registered>`], providing access to the DRM device with
|
||||
/// the guarantee that the parent bus device is bound for the entire duration of the critical
|
||||
/// section.
|
||||
///
|
||||
/// Internally this is backed by a `drm_dev_enter()` / `drm_dev_exit()` SRCU critical section.
|
||||
///
|
||||
/// # Invariants
|
||||
///
|
||||
/// - `idx` is the SRCU read lock index returned by a successful `drm_dev_enter()` call.
|
||||
/// - The parent bus device of `dev` is bound for the lifetime of this guard.
|
||||
#[must_use]
|
||||
pub struct RegistrationGuard<'a, T: drm::Driver> {
|
||||
dev: &'a Device<T, Registered>,
|
||||
idx: i32,
|
||||
_not_send: NotThreadSafe,
|
||||
}
|
||||
|
||||
impl<T: drm::Driver> Device<T, Registered> {
|
||||
/// Returns a reference to the registration data with lifetime shortened from `'static`.
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// The returned reference must not be exposed to code that can choose a concrete lifetime for
|
||||
/// it, as that would be unsound for types that are invariant over their lifetime parameter
|
||||
/// (e.g. it must be passed through an HRTB-bounded closure).
|
||||
#[inline]
|
||||
unsafe fn registration_data_unchecked(&self) -> &T::RegistrationData<'_> {
|
||||
// SAFETY:
|
||||
// - `Registered` guarantees the parent bus device is bound, hence the pointer is valid.
|
||||
// - The pointer cast from `Of<'static>` to `Of<'_>` is layout-compatible since lifetimes
|
||||
// are erased at runtime.
|
||||
// - Caller guarantees the reference is only used behind an HRTB, making the lifetime
|
||||
// shortening sound regardless of variance.
|
||||
unsafe { (*self.registration_data.get()).cast::<_>().as_ref() }
|
||||
}
|
||||
|
||||
/// Access the registration data through a closure, with the lifetime tied to the closure
|
||||
/// scope.
|
||||
///
|
||||
/// The data is owned by [`Registration`](drm::Registration) and is guaranteed to remain valid
|
||||
/// as long as the device is registered, since [`Registration`](drm::Registration)'s `drop`
|
||||
/// calls `drm_dev_unplug()` which waits for all `drm_dev_enter()` critical sections to
|
||||
/// complete.
|
||||
#[inline]
|
||||
pub fn registration_data_with<R, F>(&self, f: F) -> R
|
||||
where
|
||||
F: for<'a> FnOnce(&'a T::RegistrationData<'a>) -> R,
|
||||
{
|
||||
// SAFETY: `Registered` guarantees the device is registered and the parent bus device is
|
||||
// bound. The closure's HRTB `for<'a>` prevents the caller from smuggling in references
|
||||
// with a concrete short lifetime, satisfying the lifetime requirement of
|
||||
// `registration_data_unchecked`.
|
||||
f(unsafe { self.registration_data_unchecked() })
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: drm::Driver> Deref for RegistrationGuard<'_, T> {
|
||||
type Target = Device<T, Registered>;
|
||||
|
||||
#[inline]
|
||||
fn deref(&self) -> &Self::Target {
|
||||
self.dev
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: drm::Driver> Drop for RegistrationGuard<'_, T> {
|
||||
#[inline]
|
||||
fn drop(&mut self) {
|
||||
// SAFETY: `self.idx` was returned by a successful `drm_dev_enter()` call, as guaranteed
|
||||
// by the type invariants of `RegistrationGuard`.
|
||||
unsafe { bindings::drm_dev_exit(self.idx) };
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: drm::Driver> Deref for Device<T> {
|
||||
type Target = T::Data;
|
||||
|
||||
fn deref(&self) -> &Self::Target {
|
||||
|
|
@ -360,9 +455,31 @@ fn deref(&self) -> &Self::Target {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T: drm::Driver> Deref for Device<T, Registered> {
|
||||
type Target = Device<T>;
|
||||
|
||||
#[inline]
|
||||
fn deref(&self) -> &Self::Target {
|
||||
// SAFETY: The caller holds a `Device<T, Registered>`, which guarantees all invariants
|
||||
// of the weaker `Normal` context.
|
||||
unsafe { self.assume_ctx() }
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: drm::Driver> Deref for Device<T, Ioctl> {
|
||||
type Target = Device<T>;
|
||||
|
||||
#[inline]
|
||||
fn deref(&self) -> &Self::Target {
|
||||
// SAFETY: The caller holds a `Device<T, Ioctl>`, which guarantees all invariants
|
||||
// of the weaker `Normal` context.
|
||||
unsafe { self.assume_ctx() }
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: DRM device objects are always reference counted and the get/put functions
|
||||
// satisfy the requirements.
|
||||
unsafe impl<T: drm::Driver, C: DeviceContext> AlwaysRefCounted for Device<T, C> {
|
||||
unsafe impl<T: drm::Driver> AlwaysRefCounted for Device<T> {
|
||||
fn inc_ref(&self) {
|
||||
// SAFETY: The existence of a shared reference guarantees that the refcount is non-zero.
|
||||
unsafe { bindings::drm_dev_get(self.as_raw()) };
|
||||
|
|
@ -377,11 +494,29 @@ unsafe fn dec_ref(obj: NonNull<Self>) {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T: drm::Driver, C: DeviceContext> AsRef<device::Device> for Device<T, C> {
|
||||
fn as_ref(&self) -> &device::Device {
|
||||
impl<T: drm::Driver> AsRef<T::ParentDevice<device::Normal>> for Device<T> {
|
||||
fn as_ref(&self) -> &T::ParentDevice<device::Normal> {
|
||||
// SAFETY: `bindings::drm_device::dev` is valid as long as the DRM device itself is valid,
|
||||
// which is guaranteed by the type invariant.
|
||||
unsafe { device::Device::from_raw((*self.as_raw()).dev) }
|
||||
let dev = unsafe { device::Device::from_raw((*self.as_raw()).dev) };
|
||||
|
||||
// SAFETY: The DRM device was constructed in `UnregisteredDevice::new()` with a parent
|
||||
// device of type `T::ParentDevice`, hence `dev` is contained in a `T::ParentDevice`.
|
||||
unsafe { device::AsBusDevice::from_device(dev) }
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: drm::Driver> AsRef<T::ParentDevice<device::Bound>> for Device<T, Registered> {
|
||||
#[inline]
|
||||
fn as_ref(&self) -> &T::ParentDevice<device::Bound> {
|
||||
let dev = (**self).as_ref().as_ref();
|
||||
|
||||
// SAFETY: A `Device<T, Registered>` guarantees that the parent device is bound.
|
||||
let dev = unsafe { dev.as_bound() };
|
||||
|
||||
// SAFETY: The DRM device was constructed in `UnregisteredDevice::new()` with a parent
|
||||
// device of type `T::ParentDevice`, hence `dev` is contained in a `T::ParentDevice`.
|
||||
unsafe { device::AsBusDevice::from_device(dev) }
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -392,12 +527,10 @@ unsafe impl<T: drm::Driver, C: DeviceContext> Send for Device<T, C> {}
|
|||
// by the synchronization in `struct drm_device`.
|
||||
unsafe impl<T: drm::Driver, C: DeviceContext> Sync for Device<T, C> {}
|
||||
|
||||
impl<T, C, const ID: u64> WorkItem<ID> for Device<T, C>
|
||||
impl<T: drm::Driver, const ID: u64> WorkItem<ID> for Device<T>
|
||||
where
|
||||
T: drm::Driver,
|
||||
T::Data: WorkItem<ID, Pointer = ARef<Self>>,
|
||||
T::Data: HasWork<Self, ID>,
|
||||
C: DeviceContext,
|
||||
{
|
||||
type Pointer = ARef<Self>;
|
||||
|
||||
|
|
|
|||
|
|
@ -7,16 +7,12 @@
|
|||
use crate::{
|
||||
bindings,
|
||||
device,
|
||||
devres,
|
||||
drm,
|
||||
error::to_result,
|
||||
prelude::*,
|
||||
sync::aref::ARef, //
|
||||
};
|
||||
use core::{
|
||||
mem,
|
||||
ptr::NonNull, //
|
||||
};
|
||||
use core::ptr::NonNull;
|
||||
|
||||
/// Driver use the GEM memory manager. This should be set for all modern drivers.
|
||||
pub(crate) const FEAT_GEM: u32 = bindings::drm_driver_feature_DRIVER_GEM;
|
||||
|
|
@ -110,12 +106,23 @@ pub trait Driver {
|
|||
/// Context data associated with the DRM driver
|
||||
type Data: Sync + Send;
|
||||
|
||||
/// Data owned by the [`Registration`] and accessible within a
|
||||
/// [`RegistrationGuard`](drm::RegistrationGuard) critical section via
|
||||
/// [`Device::registration_data_with()`](drm::Device::registration_data_with).
|
||||
///
|
||||
/// The lifetime parameter is tied to the [`Registration`] scope, which is enclosed in the
|
||||
/// parent bus device binding scope but may be shorter.
|
||||
type RegistrationData<'a>: Send + Sync + 'a;
|
||||
|
||||
/// The type used to manage memory for this driver.
|
||||
type Object<Ctx: drm::DeviceContext>: AllocImpl;
|
||||
type Object: AllocImpl;
|
||||
|
||||
/// The type used to represent a DRM File (client)
|
||||
type File: drm::file::DriverFile;
|
||||
|
||||
/// The bus device type of the parent device that the DRM device is associated with.
|
||||
type ParentDevice<Ctx: device::DeviceContext>: device::AsBusDevice<Ctx>;
|
||||
|
||||
/// Driver metadata
|
||||
const INFO: DriverInfo;
|
||||
|
||||
|
|
@ -125,7 +132,7 @@ pub trait Driver {
|
|||
/// Sets the `DRIVER_RENDER` feature for this driver.
|
||||
///
|
||||
/// When enabled, the driver exposes `/dev/dri/renderDXX` render nodes to
|
||||
/// userspace. The render node is an alternate low-priviledge way to access
|
||||
/// userspace. The render node is an alternate low-privilege way to access
|
||||
/// the driver, which is enforced on a per-ioctl level. Userspace processes
|
||||
/// that open the render node can only invoke ioctls explicitly listed as
|
||||
/// usable from the render node (i.e. marked DRM_RENDER_ALLOW), whereas
|
||||
|
|
@ -136,68 +143,84 @@ pub trait Driver {
|
|||
/// The registration type of a `drm::Device`.
|
||||
///
|
||||
/// Once the `Registration` structure is dropped, the device is unregistered.
|
||||
pub struct Registration<T: Driver>(ARef<drm::Device<T>>);
|
||||
pub struct Registration<'a, T: Driver> {
|
||||
drm: ARef<drm::Device<T>>,
|
||||
_reg_data: Pin<KBox<T::RegistrationData<'a>>>,
|
||||
}
|
||||
|
||||
impl<T: Driver> Registration<T> {
|
||||
fn new(drm: drm::UnregisteredDevice<T>, flags: usize) -> Result<Self> {
|
||||
// SAFETY: `drm.as_raw()` is valid by the invariants of `drm::Device`.
|
||||
to_result(unsafe { bindings::drm_dev_register(drm.as_raw(), flags) })?;
|
||||
|
||||
// SAFETY: We just called `drm_dev_register` above
|
||||
let new = NonNull::from(unsafe { drm.assume_ctx() });
|
||||
|
||||
// Leak the ARef from UnregisteredDevice in preparation for transferring its ownership.
|
||||
mem::forget(drm);
|
||||
|
||||
// SAFETY: `drm`'s `Drop` constructor was never called, ensuring that there remains at least
|
||||
// one reference to the device - which we take ownership over here.
|
||||
let new = unsafe { ARef::from_raw(new) };
|
||||
|
||||
Ok(Self(new))
|
||||
}
|
||||
|
||||
/// Registers a new [`UnregisteredDevice`](drm::UnregisteredDevice) with userspace.
|
||||
impl<'a, T: Driver> Registration<'a, T> {
|
||||
/// Register a new [`UnregisteredDevice`](drm::UnregisteredDevice) with userspace.
|
||||
///
|
||||
/// Ownership of the [`Registration`] object is passed to [`devres::register`].
|
||||
pub fn new_foreign_owned<'a>(
|
||||
drm: drm::UnregisteredDevice<T>,
|
||||
/// # Safety
|
||||
///
|
||||
/// The caller must not `mem::forget()` the returned [`Registration`] or otherwise prevent its
|
||||
/// [`Drop`] implementation from running, since the registration data may contain borrowed
|
||||
/// references that become invalid after `'a` ends.
|
||||
pub unsafe fn new<E>(
|
||||
dev: &'a device::Device<device::Bound>,
|
||||
drm: drm::UnregisteredDevice<T>,
|
||||
reg_data: impl PinInit<T::RegistrationData<'a>, E>,
|
||||
flags: usize,
|
||||
) -> Result<&'a drm::Device<T>>
|
||||
) -> Result<Self>
|
||||
where
|
||||
T: 'static,
|
||||
Error: From<E>,
|
||||
{
|
||||
if drm.as_ref().as_raw() != dev.as_raw() {
|
||||
let parent = drm.as_ref();
|
||||
if parent.as_ref().as_raw() != dev.as_raw() {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let reg = Registration::<T>::new(drm, flags)?;
|
||||
let drm = NonNull::from(reg.device());
|
||||
let reg_data: Pin<KBox<T::RegistrationData<'a>>> = KBox::pin_init(reg_data, GFP_KERNEL)?;
|
||||
|
||||
devres::register(dev, reg, GFP_KERNEL)?;
|
||||
// Store the registration data pointer in the device before registration, so that it is
|
||||
// visible once ioctls can be called.
|
||||
let ptr: NonNull<T::RegistrationData<'static>> =
|
||||
NonNull::from(Pin::get_ref(reg_data.as_ref())).cast();
|
||||
|
||||
// SAFETY: Since `reg` was passed to devres::register(), the device now owns the lifetime
|
||||
// of the DRM registration - ensuring that this references lives for at least as long as 'a.
|
||||
Ok(unsafe { drm.as_ref() })
|
||||
// SAFETY: No concurrent access; the device is not yet registered.
|
||||
unsafe { *drm.registration_data.get() = ptr };
|
||||
|
||||
// SAFETY: `drm` is a valid, initialized but not yet registered DRM device.
|
||||
let ret = unsafe { bindings::drm_dev_register(drm.as_raw(), flags) };
|
||||
if let Err(e) = to_result(ret) {
|
||||
// SAFETY: `drm_dev_register()` synchronizes SRCU on failure, so no concurrent
|
||||
// access to `registration_data` is possible at this point.
|
||||
unsafe { *drm.registration_data.get() = NonNull::dangling() };
|
||||
return Err(e);
|
||||
}
|
||||
|
||||
Ok(Self {
|
||||
drm: (&*drm).into(),
|
||||
_reg_data: reg_data,
|
||||
})
|
||||
}
|
||||
|
||||
/// Returns a reference to the `Device` instance for this registration.
|
||||
pub fn device(&self) -> &drm::Device<T> {
|
||||
&self.0
|
||||
&self.drm
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: `Registration` doesn't offer any methods or access to fields when shared between
|
||||
// threads, hence it's safe to share it.
|
||||
unsafe impl<T: Driver> Sync for Registration<T> {}
|
||||
unsafe impl<T: Driver> Sync for Registration<'_, T> {}
|
||||
|
||||
// SAFETY: Registration with and unregistration from the DRM subsystem can happen from any thread.
|
||||
unsafe impl<T: Driver> Send for Registration<T> {}
|
||||
unsafe impl<T: Driver> Send for Registration<'_, T> {}
|
||||
|
||||
impl<T: Driver> Drop for Registration<T> {
|
||||
impl<T: Driver> Drop for Registration<'_, T> {
|
||||
fn drop(&mut self) {
|
||||
// Use `drm_dev_unplug` rather than `drm_dev_unregister` to ensure that existing
|
||||
// `drm_dev_enter()` critical sections complete before unregistration proceeds. This
|
||||
// is required for the safety of `RegistrationGuard`, which relies on the SRCU barrier in
|
||||
// `drm_dev_unplug()` to guarantee that the parent device is still bound within the
|
||||
// critical section.
|
||||
//
|
||||
// SAFETY: Safe by the invariant of `ARef<drm::Device<T>>`. The existence of this
|
||||
// `Registration` also guarantees the this `drm::Device` is actually registered.
|
||||
unsafe { bindings::drm_dev_unregister(self.0.as_raw()) };
|
||||
// `Registration` also guarantees that this `drm::Device` is actually registered.
|
||||
unsafe { bindings::drm_dev_unplug(self.drm.as_raw()) };
|
||||
// After drm_dev_unplug(), the SRCU barrier guarantees that all RegistrationGuard critical
|
||||
// sections have completed, so no one holds a reference to reg_data anymore.
|
||||
// reg_data is dropped here automatically.
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -10,7 +10,7 @@
|
|||
self,
|
||||
device::{
|
||||
DeviceContext,
|
||||
Registered, //
|
||||
Normal, //
|
||||
},
|
||||
driver::{
|
||||
AllocImpl,
|
||||
|
|
@ -81,11 +81,10 @@ unsafe fn dec_ref(obj: core::ptr::NonNull<Self>) {
|
|||
/// A type alias for retrieving the current [`AllocImpl`] for a given [`DriverObject`].
|
||||
///
|
||||
/// [`Driver`]: drm::Driver
|
||||
pub type DriverAllocImpl<T, Ctx = Registered> =
|
||||
<<T as DriverObject>::Driver as drm::Driver>::Object<Ctx>;
|
||||
pub type DriverAllocImpl<T> = <<T as DriverObject>::Driver as drm::Driver>::Object;
|
||||
|
||||
/// GEM object functions, which must be implemented by drivers.
|
||||
pub trait DriverObject: Sync + Send + Sized {
|
||||
pub trait DriverObject: Sync + Send + Sized + 'static {
|
||||
/// Parent `Driver` for this object.
|
||||
type Driver: drm::Driver;
|
||||
|
||||
|
|
@ -93,8 +92,8 @@ pub trait DriverObject: Sync + Send + Sized {
|
|||
type Args;
|
||||
|
||||
/// Create a new driver data object for a GEM object of a given size.
|
||||
fn new<Ctx: DeviceContext>(
|
||||
dev: &drm::Device<Self::Driver, Ctx>,
|
||||
fn new(
|
||||
dev: &drm::Device<Self::Driver>,
|
||||
size: usize,
|
||||
args: Self::Args,
|
||||
) -> impl PinInit<Self, Error>;
|
||||
|
|
@ -109,7 +108,7 @@ fn close(_obj: &DriverAllocImpl<Self>, _file: &DriverFile<Self>) {}
|
|||
}
|
||||
|
||||
/// Trait that represents a GEM object subtype
|
||||
pub trait IntoGEMObject: Sized + super::private::Sealed + AlwaysRefCounted {
|
||||
pub trait IntoGEMObject: Sized + super::private::Sealed {
|
||||
/// Returns a reference to the raw `drm_gem_object` structure, which must be valid as long as
|
||||
/// this owning object is valid.
|
||||
fn as_raw(&self) -> *mut bindings::drm_gem_object;
|
||||
|
|
@ -118,7 +117,8 @@ pub trait IntoGEMObject: Sized + super::private::Sealed + AlwaysRefCounted {
|
|||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// - `self_ptr` must be a valid pointer to `Self`.
|
||||
/// - `self_ptr` must be a valid pointer to the `struct drm_gem_object` embedded in a
|
||||
/// valid instance of `Self`.
|
||||
/// - The caller promises that holding the immutable reference returned by this function does
|
||||
/// not violate rust's data aliasing rules and remains valid throughout the lifetime of `'a`.
|
||||
unsafe fn from_raw<'a>(self_ptr: *mut bindings::drm_gem_object) -> &'a Self;
|
||||
|
|
@ -183,7 +183,7 @@ fn size(&self) -> usize {
|
|||
fn create_handle<D, F>(&self, file: &drm::File<F>) -> Result<u32>
|
||||
where
|
||||
Self: AllocImpl<Driver = D>,
|
||||
D: drm::Driver<Object<Registered> = Self, File = F>,
|
||||
D: drm::Driver<Object = Self, File = F>,
|
||||
F: drm::file::DriverFile<Driver = D>,
|
||||
{
|
||||
let mut handle: u32 = 0;
|
||||
|
|
@ -197,8 +197,8 @@ fn create_handle<D, F>(&self, file: &drm::File<F>) -> Result<u32>
|
|||
/// Looks up an object by its handle for a given `File`.
|
||||
fn lookup_handle<D, F>(file: &drm::File<F>, handle: u32) -> Result<ARef<Self>>
|
||||
where
|
||||
Self: AllocImpl<Driver = D>,
|
||||
D: drm::Driver<Object<Registered> = Self, File = F>,
|
||||
Self: AllocImpl<Driver = D> + AlwaysRefCounted,
|
||||
D: drm::Driver<Object = Self, File = F>,
|
||||
F: drm::file::DriverFile<Driver = D>,
|
||||
{
|
||||
// SAFETY: The arguments are all valid per the type invariants.
|
||||
|
|
@ -254,7 +254,7 @@ impl<T: IntoGEMObject> BaseObjectPrivate for T {}
|
|||
/// * Any type invariants of `Ctx` apply to the parent DRM device for this GEM object.
|
||||
#[repr(C)]
|
||||
#[pin_data]
|
||||
pub struct Object<T: DriverObject + Send + Sync, Ctx: DeviceContext = Registered> {
|
||||
pub struct Object<T: DriverObject + Send + Sync, Ctx: DeviceContext = Normal> {
|
||||
obj: Opaque<bindings::drm_gem_object>,
|
||||
#[pin]
|
||||
data: T,
|
||||
|
|
@ -280,48 +280,6 @@ impl<T: DriverObject, Ctx: DeviceContext> Object<T, Ctx> {
|
|||
rss: None,
|
||||
};
|
||||
|
||||
/// Create a new GEM object.
|
||||
pub fn new(
|
||||
dev: &drm::Device<T::Driver, Ctx>,
|
||||
size: usize,
|
||||
args: T::Args,
|
||||
) -> Result<ARef<Self>> {
|
||||
let obj: Pin<KBox<Self>> = KBox::pin_init(
|
||||
try_pin_init!(Self {
|
||||
obj: Opaque::new(bindings::drm_gem_object::default()),
|
||||
data <- T::new(dev, size, args),
|
||||
_ctx: PhantomData,
|
||||
}),
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
|
||||
// SAFETY: `obj.as_raw()` is guaranteed to be valid by the initialization above.
|
||||
unsafe { (*obj.as_raw()).funcs = &Self::OBJECT_FUNCS };
|
||||
|
||||
// INVARIANT: `dev` and the GEM object are in the same state at the moment, and upgrading
|
||||
// the typestate in `dev` will not carry over to the GEM object.
|
||||
if let Err(err) =
|
||||
// SAFETY: The arguments are all valid per the type invariants.
|
||||
to_result(unsafe {
|
||||
bindings::drm_gem_object_init(dev.as_raw(), obj.obj.get(), size)
|
||||
})
|
||||
{
|
||||
// SAFETY: `drm_gem_object_init()` initializes the private GEM object state before
|
||||
// failing, so `drm_gem_private_object_fini()` is the matching cleanup.
|
||||
unsafe { bindings::drm_gem_private_object_fini(obj.obj.get()) };
|
||||
return Err(err);
|
||||
}
|
||||
|
||||
// SAFETY: We will never move out of `Self` as `ARef<Self>` is always treated as pinned.
|
||||
let ptr = KBox::into_raw(unsafe { Pin::into_inner_unchecked(obj) });
|
||||
|
||||
// SAFETY: `ptr` comes from `KBox::into_raw` and hence can't be NULL.
|
||||
let ptr = unsafe { NonNull::new_unchecked(ptr) };
|
||||
|
||||
// SAFETY: We take over the initial reference count from `drm_gem_object_init()`.
|
||||
Ok(unsafe { ARef::from_raw(ptr) })
|
||||
}
|
||||
|
||||
/// Returns the `Device` that owns this GEM object.
|
||||
pub fn dev(&self) -> &drm::Device<T::Driver, Ctx> {
|
||||
// SAFETY:
|
||||
|
|
@ -356,11 +314,50 @@ extern "C" fn free_callback(obj: *mut bindings::drm_gem_object) {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T: DriverObject> Object<T> {
|
||||
/// Create a new GEM object.
|
||||
pub fn new(dev: &drm::Device<T::Driver>, size: usize, args: T::Args) -> Result<ARef<Self>> {
|
||||
let obj: Pin<KBox<Self>> = KBox::pin_init(
|
||||
try_pin_init!(Self {
|
||||
obj: Opaque::new(bindings::drm_gem_object::default()),
|
||||
data <- T::new(dev, size, args),
|
||||
_ctx: PhantomData,
|
||||
}),
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
|
||||
// SAFETY: `obj.as_raw()` is guaranteed to be valid by the initialization above.
|
||||
unsafe { (*obj.as_raw()).funcs = &Self::OBJECT_FUNCS };
|
||||
|
||||
// INVARIANT: `dev` and the GEM object are in the same state at the moment, and upgrading
|
||||
// the typestate in `dev` will not carry over to the GEM object.
|
||||
if let Err(err) =
|
||||
// SAFETY: The arguments are all valid per the type invariants.
|
||||
to_result(unsafe {
|
||||
bindings::drm_gem_object_init(dev.as_raw(), obj.obj.get(), size)
|
||||
})
|
||||
{
|
||||
// SAFETY: `drm_gem_object_init()` initializes the private GEM object state before
|
||||
// failing, so `drm_gem_private_object_fini()` is the matching cleanup.
|
||||
unsafe { bindings::drm_gem_private_object_fini(obj.obj.get()) };
|
||||
return Err(err);
|
||||
}
|
||||
|
||||
// SAFETY: We will never move out of `Self` as `ARef<Self>` is always treated as pinned.
|
||||
let ptr = KBox::into_raw(unsafe { Pin::into_inner_unchecked(obj) });
|
||||
|
||||
// SAFETY: `ptr` comes from `KBox::into_raw` and hence can't be NULL.
|
||||
let ptr = unsafe { NonNull::new_unchecked(ptr) };
|
||||
|
||||
// SAFETY: We take over the initial reference count from `drm_gem_object_init()`.
|
||||
Ok(unsafe { ARef::from_raw(ptr) })
|
||||
}
|
||||
}
|
||||
|
||||
impl_aref_for_gem_obj! {
|
||||
impl<T, C> for Object<T, C>
|
||||
impl<T> for Object<T>
|
||||
where
|
||||
T: DriverObject,
|
||||
C: DeviceContext
|
||||
T: DriverObject
|
||||
}
|
||||
|
||||
impl<T: DriverObject, Ctx: DeviceContext> super::private::Sealed for Object<T, Ctx> {}
|
||||
|
|
|
|||
|
|
@ -11,28 +11,57 @@
|
|||
|
||||
use crate::{
|
||||
container_of,
|
||||
device::{
|
||||
self,
|
||||
Bound, //
|
||||
},
|
||||
devres::*,
|
||||
drm::{
|
||||
driver,
|
||||
gem,
|
||||
private::Sealed,
|
||||
Device,
|
||||
DeviceContext,
|
||||
Registered, //
|
||||
Device, //
|
||||
},
|
||||
error::{
|
||||
from_err_ptr,
|
||||
to_result, //
|
||||
},
|
||||
io::{
|
||||
IoBase,
|
||||
Region,
|
||||
SysMem,
|
||||
SysMemBackend, //
|
||||
},
|
||||
error::to_result,
|
||||
prelude::*,
|
||||
sync::aref::ARef,
|
||||
types::Opaque, //
|
||||
scatterlist,
|
||||
sync::{
|
||||
aref::ARef,
|
||||
new_mutex,
|
||||
Mutex,
|
||||
SetOnce, //
|
||||
},
|
||||
types::{
|
||||
NotThreadSafe,
|
||||
Opaque, //
|
||||
},
|
||||
};
|
||||
use core::{
|
||||
marker::PhantomData,
|
||||
ffi::c_void,
|
||||
mem::{
|
||||
ManuallyDrop,
|
||||
MaybeUninit, //
|
||||
},
|
||||
ops::{
|
||||
Deref,
|
||||
DerefMut, //
|
||||
},
|
||||
ptr::NonNull, //
|
||||
ptr::{
|
||||
self,
|
||||
NonNull, //
|
||||
},
|
||||
};
|
||||
use gem::{
|
||||
BaseObject,
|
||||
BaseObjectPrivate,
|
||||
DriverObject,
|
||||
IntoGEMObject, //
|
||||
|
|
@ -42,15 +71,24 @@
|
|||
///
|
||||
/// This is used with [`Object::new()`] to control various properties that can only be set when
|
||||
/// initially creating a shmem-backed GEM object.
|
||||
#[derive(Default)]
|
||||
pub struct ObjectConfig<'a, T: DriverObject, C: DeviceContext = Registered> {
|
||||
pub struct ObjectConfig<'a, T: DriverObject> {
|
||||
/// Whether to set the write-combine map flag.
|
||||
pub map_wc: bool,
|
||||
|
||||
/// Reuse the DMA reservation from another GEM object.
|
||||
///
|
||||
/// The newly created [`Object`] will hold an owned refcount to `parent_resv_obj` if specified.
|
||||
pub parent_resv_obj: Option<&'a Object<T, C>>,
|
||||
pub parent_resv_obj: Option<&'a Object<T>>,
|
||||
}
|
||||
|
||||
impl<'a, T: DriverObject> Default for ObjectConfig<'a, T> {
|
||||
#[inline(always)]
|
||||
fn default() -> Self {
|
||||
Self {
|
||||
map_wc: false,
|
||||
parent_resv_obj: None,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A shmem-backed GEM object.
|
||||
|
|
@ -59,33 +97,35 @@ pub struct ObjectConfig<'a, T: DriverObject, C: DeviceContext = Registered> {
|
|||
///
|
||||
/// - `obj` contains a valid initialized `struct drm_gem_shmem_object` for the lifetime of this
|
||||
/// object.
|
||||
/// - Any type invariants of `C` apply to the parent DRM device for this GEM object.
|
||||
#[repr(C)]
|
||||
#[pin_data]
|
||||
pub struct Object<T: DriverObject, C: DeviceContext = Registered> {
|
||||
pub struct Object<T: DriverObject> {
|
||||
#[pin]
|
||||
obj: Opaque<bindings::drm_gem_shmem_object>,
|
||||
/// Parent object that owns this object's DMA reservation object.
|
||||
parent_resv_obj: Option<ARef<Object<T, C>>>,
|
||||
parent_resv_obj: Option<ARef<Object<T>>>,
|
||||
/// Devres object for unmapping any SGTable on driver-unbind.
|
||||
sgt_res: ManuallyDrop<SetOnce<Devres<SGTableMap<T>>>>,
|
||||
#[pin]
|
||||
/// Lock for protecting initialization of `sgt_res`.
|
||||
sgt_lock: Mutex<()>,
|
||||
#[pin]
|
||||
inner: T,
|
||||
_ctx: PhantomData<C>,
|
||||
}
|
||||
|
||||
super::impl_aref_for_gem_obj! {
|
||||
impl<T, C> for Object<T, C>
|
||||
impl<T> for Object<T>
|
||||
where
|
||||
T: DriverObject,
|
||||
C: DeviceContext
|
||||
T: DriverObject
|
||||
}
|
||||
|
||||
// SAFETY: All GEM objects are thread-safe.
|
||||
unsafe impl<T: DriverObject, C: DeviceContext> Send for Object<T, C> {}
|
||||
unsafe impl<T: DriverObject> Send for Object<T> {}
|
||||
|
||||
// SAFETY: All GEM objects are thread-safe.
|
||||
unsafe impl<T: DriverObject, C: DeviceContext> Sync for Object<T, C> {}
|
||||
unsafe impl<T: DriverObject> Sync for Object<T> {}
|
||||
|
||||
impl<T: DriverObject, C: DeviceContext> Object<T, C> {
|
||||
impl<T: DriverObject> Object<T> {
|
||||
/// `drm_gem_object_funcs` vtable suitable for GEM shmem objects.
|
||||
const VTABLE: bindings::drm_gem_object_funcs = bindings::drm_gem_object_funcs {
|
||||
free: Some(Self::free_callback),
|
||||
|
|
@ -112,21 +152,166 @@ fn as_raw_shmem(&self) -> *mut bindings::drm_gem_shmem_object {
|
|||
self.obj.get()
|
||||
}
|
||||
|
||||
/// Returns the `Device` that owns this GEM object.
|
||||
pub fn dev(&self) -> &Device<T::Driver> {
|
||||
// SAFETY: `dev` will have been initialized in `Self::new()` by `drm_gem_shmem_init()`.
|
||||
unsafe { Device::from_raw((*self.as_raw()).dev) }
|
||||
}
|
||||
|
||||
extern "C" fn free_callback(obj: *mut bindings::drm_gem_object) {
|
||||
// SAFETY:
|
||||
// - DRM always passes a valid gem object here
|
||||
// - We used drm_gem_shmem_create() in our create_gem_object callback, so we know that
|
||||
// `obj` is contained within a drm_gem_shmem_object
|
||||
let base = unsafe { container_of!(obj, bindings::drm_gem_shmem_object, base) };
|
||||
|
||||
// SAFETY:
|
||||
// - We verified above that `obj` is valid, which makes `this` valid
|
||||
// - This function is set in AllocOps, so we know that `this` is contained within an
|
||||
// `Object<T>`
|
||||
let this = unsafe { container_of!(Opaque::cast_from(base), Self, obj) }.cast_mut();
|
||||
|
||||
// We need to drop `sgt_res` first, since doing so requires that the GEM object is still
|
||||
// alive.
|
||||
// SAFETY:
|
||||
// - We verified above that `this` is valid.
|
||||
// - We are in free_callback, guaranteeing we have exclusive access to `this` and that
|
||||
// `sgt_res` will not be used after dropping it here.
|
||||
unsafe { ManuallyDrop::drop(&mut (*this).sgt_res) };
|
||||
|
||||
// SAFETY:
|
||||
// - We're in free_callback - so this function is safe to call.
|
||||
// - We won't be using the gem resources on `this` after this call.
|
||||
unsafe { bindings::drm_gem_shmem_release(base) };
|
||||
|
||||
// SAFETY: We're recovering the Kbox<> we created in gem_create_object()
|
||||
let _ = unsafe { KBox::from_raw(this) };
|
||||
}
|
||||
|
||||
/// Attempt to create a vmap from the gem object, and confirm the size of said vmap.
|
||||
fn make_vmap<'a, R, const SIZE: usize>(&'a self) -> Result<VMap<T, R, SIZE>>
|
||||
where
|
||||
R: Deref<Target = Self> + From<&'a Self>,
|
||||
{
|
||||
// INVARIANT: We check here that the gem object is at least as large as `SIZE`.
|
||||
if self.size() < SIZE {
|
||||
return Err(ENOSPC);
|
||||
}
|
||||
|
||||
let mut map: MaybeUninit<bindings::iosys_map> = MaybeUninit::uninit();
|
||||
let guard = DmaResvGuard::new(self);
|
||||
|
||||
// SAFETY: `drm_gem_shmem_vmap()` can be called with the DMA reservation lock held.
|
||||
to_result(unsafe {
|
||||
bindings::drm_gem_shmem_vmap_locked(self.as_raw_shmem(), map.as_mut_ptr())
|
||||
})?;
|
||||
|
||||
// Drop the guard explicitly here, since we may need to call `raw_vunmap()` (which
|
||||
// re-acquires the lock).
|
||||
drop(guard);
|
||||
|
||||
// SAFETY: The call to `drm_gem_shmem_vmap_locked()` succeeded above, so we are guaranteed
|
||||
// that map is properly initialized.
|
||||
let map = unsafe { map.assume_init() };
|
||||
|
||||
// XXX: We don't currently support iomem allocations
|
||||
if map.is_iomem {
|
||||
// SAFETY: The vmap operation above succeeded, guaranteeing that `map` points to a valid
|
||||
// memory mapping.
|
||||
unsafe { self.raw_vunmap(map) };
|
||||
|
||||
Err(ENOTSUPP)
|
||||
} else {
|
||||
Ok(VMap {
|
||||
// INVARIANT: `addr` remains valid for as long as `owner` does, which extends to the
|
||||
// lifetime of `VMap` itself.
|
||||
// SAFETY: We checked that this is not an iomem allocation, making it safe to read
|
||||
// vaddr.
|
||||
addr: unsafe { map.__bindgen_anon_1.vaddr },
|
||||
owner: self.into(),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// Unmap a vmap from the gem object.
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// - The caller promises that `map` is a valid vmap on this gem object.
|
||||
/// - The caller promises that the memory pointed to by map will no longer be accesed through
|
||||
/// this instance.
|
||||
unsafe fn raw_vunmap(&self, mut map: bindings::iosys_map) {
|
||||
let _guard = DmaResvGuard::new(self);
|
||||
|
||||
// SAFETY:
|
||||
// - This function is safe to call with the DMA reservation lock held.
|
||||
// - The caller promises that `map` is a valid vmap on this gem object.
|
||||
unsafe { bindings::drm_gem_shmem_vunmap_locked(self.as_raw_shmem(), &mut map) };
|
||||
}
|
||||
|
||||
/// Creates and returns a virtual kernel memory mapping for this object.
|
||||
#[inline]
|
||||
pub fn vmap<const SIZE: usize>(&self) -> Result<VMapRef<'_, T, SIZE>> {
|
||||
self.make_vmap()
|
||||
}
|
||||
|
||||
/// Creates (if necessary) and returns an immutable reference to a scatter-gather table of DMA
|
||||
/// pages for this object.
|
||||
///
|
||||
/// This will pin the object in memory. It is expected that `dev` should be a pointer to the
|
||||
/// same [`device::Device`] which `self` belongs to, otherwise this function will return
|
||||
/// `Err(EINVAL)`.
|
||||
pub fn sg_table<'a>(
|
||||
&'a self,
|
||||
dev: &'a device::Device<Bound>,
|
||||
) -> Result<&'a scatterlist::SGTable> {
|
||||
let parent = self.dev().as_ref();
|
||||
if dev.as_raw() != parent.as_ref().as_raw() {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let sgt_res = 'out: {
|
||||
// Fast path: sgt_res is already initialized
|
||||
if let Some(sgt_res) = self.sgt_res.as_ref() {
|
||||
break 'out sgt_res;
|
||||
}
|
||||
|
||||
// Slow path: Grab the lock and see if we need to initialize sgt_res.
|
||||
let _guard = self.sgt_lock.lock();
|
||||
|
||||
// If someone initialized it while we were waiting, we can exit early.
|
||||
if let Some(sgt_res) = self.sgt_res.as_ref() {
|
||||
break 'out sgt_res;
|
||||
}
|
||||
|
||||
// If not, finish initializing and return. `populate()` cannot return false, as
|
||||
// `sgt_res` must be unpopulated, and we must hold `sgt_lock` to reach this point.
|
||||
self.sgt_res
|
||||
.populate(Devres::new(dev, SGTableMap::new(self))?);
|
||||
|
||||
// SAFETY: We just populated sgt_res above.
|
||||
unsafe { self.sgt_res.as_ref().unwrap_unchecked() }
|
||||
};
|
||||
|
||||
Ok(sgt_res.access(dev)?)
|
||||
}
|
||||
|
||||
/// Create a new shmem-backed DRM object of the given size.
|
||||
///
|
||||
/// Additional config options can be specified using `config`.
|
||||
pub fn new(
|
||||
dev: &Device<T::Driver, C>,
|
||||
dev: &Device<T::Driver>,
|
||||
size: usize,
|
||||
config: ObjectConfig<'_, T, C>,
|
||||
config: ObjectConfig<'_, T>,
|
||||
args: T::Args,
|
||||
) -> Result<ARef<Self>> {
|
||||
let new: Pin<KBox<Self>> = KBox::try_pin_init(
|
||||
try_pin_init!(Self {
|
||||
obj <- Opaque::init_zeroed(),
|
||||
parent_resv_obj: config.parent_resv_obj.map(|p| p.into()),
|
||||
sgt_res: ManuallyDrop::new(SetOnce::new()),
|
||||
sgt_lock <- new_mutex!(()),
|
||||
inner <- T::new(dev, size, args),
|
||||
_ctx: PhantomData::<C>,
|
||||
}),
|
||||
GFP_KERNEL,
|
||||
)?;
|
||||
|
|
@ -158,36 +343,14 @@ pub fn new(
|
|||
Ok(obj)
|
||||
}
|
||||
|
||||
/// Returns the `Device` that owns this GEM object.
|
||||
pub fn dev(&self) -> &Device<T::Driver, C> {
|
||||
// SAFETY: `dev` will have been initialized in `Self::new()` by `drm_gem_shmem_init()`.
|
||||
unsafe { Device::from_raw((*self.as_raw()).dev) }
|
||||
}
|
||||
|
||||
extern "C" fn free_callback(obj: *mut bindings::drm_gem_object) {
|
||||
// SAFETY:
|
||||
// - DRM always passes a valid gem object here
|
||||
// - We used drm_gem_shmem_create() in our create_gem_object callback, so we know that
|
||||
// `obj` is contained within a drm_gem_shmem_object
|
||||
let this = unsafe { container_of!(obj, bindings::drm_gem_shmem_object, base) };
|
||||
|
||||
// SAFETY:
|
||||
// - We're in free_callback - so this function is safe to call.
|
||||
// - We won't be using the gem resources on `this` after this call.
|
||||
unsafe { bindings::drm_gem_shmem_release(this) };
|
||||
|
||||
// SAFETY:
|
||||
// - We verified above that `obj` is valid, which makes `this` valid
|
||||
// - This function is set in AllocOps, so we know that `this` is contained within a
|
||||
// `Object<T, C>`
|
||||
let this = unsafe { container_of!(Opaque::cast_from(this), Self, obj) }.cast_mut();
|
||||
|
||||
// SAFETY: We're recovering the Kbox<> we created in gem_create_object()
|
||||
let _ = unsafe { KBox::from_raw(this) };
|
||||
/// Creates and returns an owned reference to a virtual kernel memory mapping for this object.
|
||||
#[inline]
|
||||
pub fn owned_vmap<const SIZE: usize>(&self) -> Result<VMapOwned<T, SIZE>> {
|
||||
self.make_vmap()
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: DriverObject, C: DeviceContext> Deref for Object<T, C> {
|
||||
impl<T: DriverObject> Deref for Object<T> {
|
||||
type Target = T;
|
||||
|
||||
fn deref(&self) -> &Self::Target {
|
||||
|
|
@ -195,15 +358,15 @@ fn deref(&self) -> &Self::Target {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T: DriverObject, C: DeviceContext> DerefMut for Object<T, C> {
|
||||
impl<T: DriverObject> DerefMut for Object<T> {
|
||||
fn deref_mut(&mut self) -> &mut Self::Target {
|
||||
&mut self.inner
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: DriverObject, C: DeviceContext> Sealed for Object<T, C> {}
|
||||
impl<T: DriverObject> Sealed for Object<T> {}
|
||||
|
||||
impl<T: DriverObject, C: DeviceContext> gem::IntoGEMObject for Object<T, C> {
|
||||
impl<T: DriverObject> gem::IntoGEMObject for Object<T> {
|
||||
fn as_raw(&self) -> *mut bindings::drm_gem_object {
|
||||
// SAFETY:
|
||||
// - Our immutable reference is proof that this is safe to dereference.
|
||||
|
|
@ -222,7 +385,7 @@ unsafe fn from_raw<'a>(obj: *mut bindings::drm_gem_object) -> &'a Self {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T: DriverObject, C: DeviceContext> driver::AllocImpl for Object<T, C> {
|
||||
impl<T: DriverObject> driver::AllocImpl for Object<T> {
|
||||
type Driver = T::Driver;
|
||||
|
||||
const ALLOC_OPS: driver::AllocOps = driver::AllocOps {
|
||||
|
|
@ -235,3 +398,324 @@ impl<T: DriverObject, C: DeviceContext> driver::AllocImpl for Object<T, C> {
|
|||
dumb_map_offset: None,
|
||||
};
|
||||
}
|
||||
|
||||
/// Private helper-type for holding the `dma_resv` object for a GEM shmem object.
|
||||
///
|
||||
/// When this is dropped, the `dma_resv` lock is dropped as well.
|
||||
///
|
||||
// TODO: This should be replace with a WwMutex equivalent once we have such bindings in the kernel.
|
||||
struct DmaResvGuard<'a, T: DriverObject>(&'a Object<T>, NotThreadSafe);
|
||||
|
||||
impl<'a, T: DriverObject> DmaResvGuard<'a, T> {
|
||||
#[inline]
|
||||
fn new(obj: &'a Object<T>) -> Self {
|
||||
// SAFETY: This lock is initialized throughout the lifetime of `object`.
|
||||
unsafe { bindings::dma_resv_lock(obj.raw_dma_resv(), ptr::null_mut()) };
|
||||
|
||||
Self(obj, NotThreadSafe)
|
||||
}
|
||||
}
|
||||
|
||||
impl<'a, T: DriverObject> Drop for DmaResvGuard<'a, T> {
|
||||
#[inline]
|
||||
fn drop(&mut self) {
|
||||
// SAFETY: We are releasing the lock grabbed during the creation of this object.
|
||||
unsafe { bindings::dma_resv_unlock(self.0.raw_dma_resv()) };
|
||||
}
|
||||
}
|
||||
|
||||
/// A reference to a virtual mapping for an shmem-based GEM object in kernel address space.
|
||||
///
|
||||
/// # Invariants
|
||||
///
|
||||
/// - The size of `owner` is >= SIZE.
|
||||
/// - The memory pointed to by `addr` remains valid at least until this object is dropped.
|
||||
pub struct VMap<D, R, const SIZE: usize = 0>
|
||||
where
|
||||
D: DriverObject,
|
||||
R: Deref<Target = Object<D>>,
|
||||
{
|
||||
addr: *mut c_void,
|
||||
owner: R,
|
||||
}
|
||||
|
||||
/// An alias type for a reference to a shmem-based GEM object's VMap.
|
||||
pub type VMapRef<'a, D, const SIZE: usize = 0> = VMap<D, &'a Object<D>, SIZE>;
|
||||
|
||||
/// An alias type for an owned reference to a shmem-based GEM object's VMap.
|
||||
pub type VMapOwned<D, const SIZE: usize = 0> = VMap<D, ARef<Object<D>>, SIZE>;
|
||||
|
||||
impl<D, R, const SIZE: usize> VMap<D, R, SIZE>
|
||||
where
|
||||
D: DriverObject,
|
||||
R: Deref<Target = Object<D>>,
|
||||
{
|
||||
/// Borrows a reference to the object that owns this virtual mapping.
|
||||
#[inline]
|
||||
pub fn owner(&self) -> &Object<D> {
|
||||
&self.owner
|
||||
}
|
||||
}
|
||||
|
||||
impl<'a, D, R, const SIZE: usize> IoBase<'a> for &'a VMap<D, R, SIZE>
|
||||
where
|
||||
D: DriverObject,
|
||||
R: Deref<Target = Object<D>>,
|
||||
{
|
||||
type Backend = SysMemBackend;
|
||||
type Target = Region<SIZE>;
|
||||
|
||||
#[inline]
|
||||
fn as_view(self) -> SysMem<'a, Region<SIZE>> {
|
||||
let ptr = Region::ptr_from_raw_parts_mut(self.addr.cast(), self.owner.size());
|
||||
|
||||
// SAFETY: Per type invariants of `VMap`:
|
||||
// - `addr .. addr + owner.size()` is a valid kernel accessible memory region.
|
||||
// - `addr` is page-aligned, which satisfies `Region`'s 4-byte alignment requirement.
|
||||
// - The memory remains valid until this `VMap` is dropped; since `self` is `&'a VMap`,
|
||||
// the borrow prevents the `VMap` from being dropped for the lifetime `'a`.
|
||||
unsafe { SysMem::new(ptr) }
|
||||
}
|
||||
}
|
||||
|
||||
impl<D, R, const SIZE: usize> Drop for VMap<D, R, SIZE>
|
||||
where
|
||||
D: DriverObject,
|
||||
R: Deref<Target = Object<D>>,
|
||||
{
|
||||
#[inline]
|
||||
fn drop(&mut self) {
|
||||
// SAFETY:
|
||||
// - Our existence is proof that this map was previously created using self.owner.
|
||||
// - Since we are in Drop, we are guaranteed that no one will access the memory
|
||||
// through this mapping after calling this.
|
||||
unsafe {
|
||||
self.owner.raw_vunmap(bindings::iosys_map {
|
||||
is_iomem: false,
|
||||
__bindgen_anon_1: bindings::iosys_map__bindgen_ty_1 { vaddr: self.addr },
|
||||
})
|
||||
};
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: `addr` points to a valid memory address for as long as `owner` exists, meaning that so
|
||||
// long as `owner` is `Send` so is `VMap`.
|
||||
unsafe impl<D, R, const SIZE: usize> Send for VMap<D, R, SIZE>
|
||||
where
|
||||
D: DriverObject,
|
||||
R: Deref<Target = Object<D>> + Send,
|
||||
{
|
||||
}
|
||||
|
||||
// SAFETY: `addr` points to a valid memory address for as long as `owner` exists, meaning that so
|
||||
// long as `owner` is `Sync` so is `VMap`.
|
||||
unsafe impl<D, R, const SIZE: usize> Sync for VMap<D, R, SIZE>
|
||||
where
|
||||
D: DriverObject,
|
||||
R: Deref<Target = Object<D>> + Sync,
|
||||
{
|
||||
}
|
||||
|
||||
/// A reference to a GEM object that is known to have a mapped [`SGTable`].
|
||||
///
|
||||
/// This is used by the Rust bindings with [`Devres`] in order to ensure that mappings for SGTables
|
||||
/// on GEM shmem objects are revoked on driver-unbind.
|
||||
///
|
||||
/// # Invariants
|
||||
///
|
||||
/// - `self.obj` always points to a valid GEM object.
|
||||
/// - This object is proof that `self.obj.owner.sgt_res` has an initialized and valid pointer to an
|
||||
/// [`SGTable`].
|
||||
///
|
||||
/// [`SGTable`]: scatterlist::SGTable
|
||||
pub struct SGTableMap<T: DriverObject> {
|
||||
obj: NonNull<Object<T>>,
|
||||
}
|
||||
|
||||
impl<T: DriverObject> Deref for SGTableMap<T> {
|
||||
type Target = scatterlist::SGTable;
|
||||
|
||||
fn deref(&self) -> &Self::Target {
|
||||
// SAFETY:
|
||||
// - The NonNull is guaranteed to be valid via our type invariants.
|
||||
// - The sgt field is guaranteed to be initialized and valid via our type invariants.
|
||||
unsafe { scatterlist::SGTable::from_raw((*self.obj.as_ref().as_raw_shmem()).sgt) }
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: DriverObject> Drop for SGTableMap<T> {
|
||||
fn drop(&mut self) {
|
||||
// SAFETY: `obj` is always valid via our type invariants
|
||||
let obj = unsafe { self.obj.as_ref() };
|
||||
let _lock = DmaResvGuard::new(obj);
|
||||
|
||||
// SAFETY: We acquired the lock needed for calling this function above
|
||||
unsafe { bindings::__drm_gem_shmem_free_sgt_locked(obj.as_raw_shmem()) };
|
||||
}
|
||||
}
|
||||
|
||||
impl<T: DriverObject> SGTableMap<T> {
|
||||
fn new(obj: &Object<T>) -> impl Init<Self, Error> {
|
||||
// INVARIANT:
|
||||
// - We call drm_gem_shmem_get_pages_sgt below and check whether or not it succeeds,
|
||||
// fulfilling the invariant of SGTableMap that the object's `sgt` field is initialized.
|
||||
// SAFETY:
|
||||
// - `obj` is fully initialized, making this function safe to call.
|
||||
from_err_ptr(unsafe { bindings::drm_gem_shmem_get_pages_sgt(obj.as_raw_shmem()) })?;
|
||||
|
||||
Ok(Self { obj: obj.into() })
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: The NonNull in SGTableMap is guaranteed valid by our type invariants, and the GEM object
|
||||
// it points to is guaranteed to be thread-safe.
|
||||
unsafe impl<T: DriverObject> Send for SGTableMap<T> {}
|
||||
// SAFETY: The NonNull in SGTableMap is guaranteed valid by our type invariants, and the GEM object
|
||||
// it points to is guaranteed to be thread-safe.
|
||||
unsafe impl<T: DriverObject> Sync for SGTableMap<T> {}
|
||||
|
||||
#[kunit_tests(rust_drm_gem_shmem)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::{
|
||||
drm::{
|
||||
self,
|
||||
UnregisteredDevice, //
|
||||
},
|
||||
faux,
|
||||
io::Io,
|
||||
page::PAGE_SIZE, //
|
||||
};
|
||||
|
||||
// The bare minimum needed to create a fake drm driver for kunit
|
||||
|
||||
#[pin_data]
|
||||
struct KunitData {}
|
||||
struct KunitDriver;
|
||||
struct KunitFile;
|
||||
#[pin_data]
|
||||
struct KunitObject {}
|
||||
|
||||
const INFO: drm::DriverInfo = drm::DriverInfo {
|
||||
major: 0,
|
||||
minor: 0,
|
||||
patchlevel: 0,
|
||||
name: c"kunit",
|
||||
desc: c"Kunit",
|
||||
};
|
||||
|
||||
impl drm::file::DriverFile for KunitFile {
|
||||
type Driver = KunitDriver;
|
||||
|
||||
fn open(_dev: &drm::Device<KunitDriver>) -> Result<Pin<KBox<Self>>> {
|
||||
Ok(KBox::new(Self, GFP_KERNEL)?.into())
|
||||
}
|
||||
}
|
||||
|
||||
impl gem::DriverObject for KunitObject {
|
||||
type Driver = KunitDriver;
|
||||
type Args = ();
|
||||
|
||||
fn new(
|
||||
_dev: &drm::Device<KunitDriver>,
|
||||
_size: usize,
|
||||
_args: Self::Args,
|
||||
) -> impl PinInit<Self, Error> {
|
||||
try_pin_init!(KunitObject {})
|
||||
}
|
||||
}
|
||||
|
||||
#[vtable]
|
||||
impl drm::Driver for KunitDriver {
|
||||
type Data = KunitData;
|
||||
type RegistrationData<'a> = ();
|
||||
type File = KunitFile;
|
||||
type Object = Object<KunitObject>;
|
||||
type ParentDevice<Ctx: device::DeviceContext> = faux::Device<Ctx>;
|
||||
|
||||
const INFO: drm::DriverInfo = INFO;
|
||||
const IOCTLS: &'static [drm::ioctl::DrmIoctlDescriptor] = &[];
|
||||
}
|
||||
|
||||
fn create_drm_dev() -> Result<(faux::Registration, UnregisteredDevice<KunitDriver>)> {
|
||||
// Create a faux DRM device so we can test gem object creation.
|
||||
let data = try_pin_init!(KunitData {});
|
||||
let reg = faux::Registration::new(c"Kunit", None)?;
|
||||
let fdev = reg.as_ref();
|
||||
let drm = UnregisteredDevice::new(fdev, data)?;
|
||||
|
||||
Ok((reg, drm))
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn compile_time_vmap_sizes() -> Result {
|
||||
let (_dev, drm) = create_drm_dev()?;
|
||||
|
||||
let obj = Object::<KunitObject>::new(&drm, PAGE_SIZE, ObjectConfig::default(), ())?;
|
||||
|
||||
// Try creating a normal vmap
|
||||
obj.vmap::<PAGE_SIZE>()?;
|
||||
|
||||
// Try creating a vmap that's smaller then the size we specified
|
||||
let vmap = obj.vmap::<{ PAGE_SIZE - 100 }>()?;
|
||||
|
||||
// Verify the owner matches
|
||||
assert!(ptr::eq(vmap.owner(), obj.deref()));
|
||||
|
||||
// Verify the size matches the actual object size
|
||||
assert_eq!(vmap.size(), PAGE_SIZE);
|
||||
|
||||
// Make sure creating a vmap that's too large fails
|
||||
assert!(obj.vmap::<{ PAGE_SIZE + 200 }>().is_err());
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn vmap_io() -> Result {
|
||||
let (_dev, drm) = create_drm_dev()?;
|
||||
|
||||
let obj = Object::<KunitObject>::new(&drm, PAGE_SIZE, ObjectConfig::default(), ())?;
|
||||
|
||||
let vmap = obj.vmap::<PAGE_SIZE>()?;
|
||||
|
||||
vmap.write8(0xDE, 0x0);
|
||||
assert_eq!(vmap.read8(0x0), 0xDE);
|
||||
vmap.write32(0xFEDCBA98, 0x20);
|
||||
|
||||
assert_eq!(vmap.read32(0x20), 0xFEDCBA98);
|
||||
|
||||
// Ensure the ordering in memory is correct
|
||||
let expected = 0xFEDCBA98_u32.to_ne_bytes().into_iter();
|
||||
for (offset, expected) in (0x20..=0x23).zip(expected) {
|
||||
assert_eq!(vmap.try_read8(offset).unwrap(), expected);
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
// TODO: I would love to actually test the success paths of sg_table(), but that would require
|
||||
// also implementing dummy dma_ops so that trying to create a mapping doesn't explode. So, leave
|
||||
// that for someone else.
|
||||
|
||||
// Ensures that passing the wrong device to sg_table() fails as we expect, and also ensure it
|
||||
// skips initializing `sgt_res` since we could otherwise create `sgt_res` with the wrong device
|
||||
// bound to it.
|
||||
#[test]
|
||||
fn fail_sg_table_on_wrong_dev() -> Result {
|
||||
let (_dev, drm) = create_drm_dev()?;
|
||||
let reg = faux::Registration::new(c"EvilKunit", None)?;
|
||||
let wrong_dev = reg.as_ref();
|
||||
|
||||
let obj = Object::<KunitObject>::new(&drm, PAGE_SIZE, ObjectConfig::default(), ())?;
|
||||
|
||||
assert_eq!(obj.sg_table(wrong_dev.as_ref()).err().unwrap(), EINVAL);
|
||||
|
||||
// If sgt_res was not initialized mistakenly with the wrong device, this should still fail.
|
||||
assert_eq!(obj.sg_table(wrong_dev.as_ref()).err().unwrap(), EINVAL);
|
||||
|
||||
// TODO: Someday, we should test that creating an sg_table here still succeeds.
|
||||
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -72,10 +72,12 @@ pub struct GpuVm<T: DriverGpuVm> {
|
|||
data: UnsafeCell<T>,
|
||||
}
|
||||
|
||||
// SAFETY: The GPUVM api does not assume that it is tied to a specific thread. The destructor will
|
||||
// drop the `data` field, which is okay because it is guaranteed `Send` by the `DriverGpuVm` trait.
|
||||
// SAFETY: It is safe to send a `GpuVm<T>` to another thread: all data reachable through it
|
||||
// (`T`, `T::VmBoData`, and the GEM `T::Object`) is `Send` by the `DriverGpuVm` bounds.
|
||||
unsafe impl<T: DriverGpuVm> Send for GpuVm<T> {}
|
||||
// SAFETY: The GPUVM api is designed to allow &self methods to be called in parallel.
|
||||
// SAFETY: It is safe to share a `&GpuVm<T>` between threads: `&self` methods only alias data
|
||||
// that is `Sync` by the `DriverGpuVm` bounds, and any thread may drop that data, or upgrade the
|
||||
// reference and ultimately drop `T`, which the same bounds make `Send`.
|
||||
unsafe impl<T: DriverGpuVm> Sync for GpuVm<T> {}
|
||||
|
||||
// SAFETY: By type invariants, the allocation is managed by the refcount in `self.vm`.
|
||||
|
|
@ -116,9 +118,9 @@ const fn vtable() -> &'static bindings::drm_gpuvm_ops {
|
|||
|
||||
/// Creates a GPUVM instance.
|
||||
#[expect(clippy::new_ret_no_self)]
|
||||
pub fn new<E>(
|
||||
pub fn new<E, Ctx: drm::DeviceContext>(
|
||||
name: &'static CStr,
|
||||
dev: &drm::Device<T::Driver>,
|
||||
dev: &drm::Device<T::Driver, Ctx>,
|
||||
r_obj: &T::Object,
|
||||
range: Range<u64>,
|
||||
reserve_range: Range<u64>,
|
||||
|
|
@ -250,21 +252,27 @@ fn raw_resv(&self) -> *mut bindings::dma_resv {
|
|||
}
|
||||
|
||||
/// The manager for a GPUVM.
|
||||
pub trait DriverGpuVm: Sized + Send {
|
||||
pub trait DriverGpuVm: Sized + Send + Sync {
|
||||
/// Parent `Driver` for this object.
|
||||
type Driver: drm::Driver<Object = Self::Object>;
|
||||
type Driver: drm::Driver;
|
||||
|
||||
/// The kind of GEM object stored in this GPUVM.
|
||||
type Object: IntoGEMObject;
|
||||
type Object: drm::driver::AllocImpl<Driver = Self::Driver> + Send + Sync;
|
||||
|
||||
/// Data stored with each [`struct drm_gpuva`](struct@GpuVa).
|
||||
type VaData;
|
||||
///
|
||||
/// Only `Send` is required: the data has a single owner at all times, moving
|
||||
/// between threads by value (handed back as a [`GpuVaRemoved`]) but never
|
||||
/// accessed by two threads concurrently.
|
||||
type VaData: Send;
|
||||
|
||||
/// Data stored with each [`struct drm_gpuvm_bo`](struct@GpuVmBo).
|
||||
type VmBoData;
|
||||
type VmBoData: Send + Sync;
|
||||
|
||||
/// The private data passed to callbacks.
|
||||
type SmContext<'ctx>;
|
||||
type SmContext<'ctx>
|
||||
where
|
||||
Self: 'ctx;
|
||||
|
||||
/// Indicates that a new mapping should be created.
|
||||
fn sm_step_map<'op, 'ctx>(
|
||||
|
|
@ -296,12 +304,10 @@ fn sm_step_remap<'op, 'ctx>(
|
|||
/// # Invariants
|
||||
///
|
||||
/// Each `GpuVm` instance has at most one `UniqueRefGpuVm` reference.
|
||||
// `Send`/`Sync` derive from `ARef<GpuVm<T>>`; the trait bounds make them correct for the unique
|
||||
// handle's `&mut T` access.
|
||||
pub struct UniqueRefGpuVm<T: DriverGpuVm>(ARef<GpuVm<T>>);
|
||||
|
||||
// SAFETY: The GPUVM api is designed to allow &self methods to be called in parallel, and
|
||||
// concurrent access to `data` is safe due to the `T: Sync` requirement.
|
||||
unsafe impl<T: DriverGpuVm + Sync> Sync for UniqueRefGpuVm<T> {}
|
||||
|
||||
impl<T: DriverGpuVm> UniqueRefGpuVm<T> {
|
||||
/// Access the data owned by this `UniqueRefGpuVm` immutably.
|
||||
#[inline]
|
||||
|
|
|
|||
|
|
@ -3,7 +3,7 @@
|
|||
use super::*;
|
||||
|
||||
/// The actual data that gets threaded through the callbacks.
|
||||
struct SmData<'a, 'ctx, T: DriverGpuVm> {
|
||||
struct SmData<'a, 'ctx, T: DriverGpuVm + 'ctx> {
|
||||
gpuvm: &'a mut UniqueRefGpuVm<T>,
|
||||
user_context: &'a mut T::SmContext<'ctx>,
|
||||
}
|
||||
|
|
@ -20,7 +20,7 @@ struct SmMapData<'a, 'ctx, T: DriverGpuVm> {
|
|||
}
|
||||
|
||||
/// The argument for [`UniqueRefGpuVm::sm_map`].
|
||||
pub struct OpMapRequest<'a, 'ctx, T: DriverGpuVm> {
|
||||
pub struct OpMapRequest<'a, 'ctx, T: DriverGpuVm + 'ctx> {
|
||||
/// Address in GPU virtual address space.
|
||||
pub addr: u64,
|
||||
/// Length of mapping to create.
|
||||
|
|
|
|||
|
|
@ -104,6 +104,14 @@ pub fn vm_bo(&self) -> &GpuVmBo<T> {
|
|||
/// The memory is zeroed.
|
||||
pub struct GpuVaAlloc<T: DriverGpuVm>(KBox<MaybeUninit<GpuVa<T>>>);
|
||||
|
||||
// SAFETY: A `GpuVaAlloc` is an owned, uninitialised allocation with no live `T::VaData` and no
|
||||
// thread-bound state.
|
||||
unsafe impl<T: DriverGpuVm> Send for GpuVaAlloc<T> {}
|
||||
|
||||
// SAFETY: A `GpuVaAlloc` has no `&self` method that reaches its contents, so a shared
|
||||
// `&GpuVaAlloc` cannot access the allocation.
|
||||
unsafe impl<T: DriverGpuVm> Sync for GpuVaAlloc<T> {}
|
||||
|
||||
impl<T: DriverGpuVm> GpuVaAlloc<T> {
|
||||
/// Pre-allocate a [`GpuVa`] object.
|
||||
pub fn new(flags: AllocFlags) -> Result<GpuVaAlloc<T>, AllocError> {
|
||||
|
|
|
|||
|
|
@ -19,6 +19,15 @@ pub struct GpuVmBo<T: DriverGpuVm> {
|
|||
data: T::VmBoData,
|
||||
}
|
||||
|
||||
// SAFETY: It is safe to send a `GpuVmBo<T>` to another thread: dropping it there drops
|
||||
// `T::VmBoData` and the GEM `T::Object`, both `Send` by the `DriverGpuVm` bounds.
|
||||
unsafe impl<T: DriverGpuVm> Send for GpuVmBo<T> {}
|
||||
|
||||
// SAFETY: It is safe to share a `&GpuVmBo<T>` between threads: it effectively shares
|
||||
// `&T::VmBoData` and the GEM `&T::Object` (both `Sync`), and any thread may upgrade to an
|
||||
// `ARef` and ultimately drop them (both `Send`), per the `DriverGpuVm` bounds.
|
||||
unsafe impl<T: DriverGpuVm> Sync for GpuVmBo<T> {}
|
||||
|
||||
// SAFETY: By type invariants, the allocation is managed by the refcount in `self.inner`.
|
||||
unsafe impl<T: DriverGpuVm> AlwaysRefCounted for GpuVmBo<T> {
|
||||
fn inc_ref(&self) {
|
||||
|
|
|
|||
|
|
@ -70,6 +70,18 @@ pub mod internal {
|
|||
pub use bindings::drm_device;
|
||||
pub use bindings::drm_file;
|
||||
pub use bindings::drm_ioctl_desc;
|
||||
|
||||
/// Cast an [`Ioctl`] DRM device pointer to [`Registered`], preserving the driver type
|
||||
/// parameter `T`.
|
||||
///
|
||||
/// Used by [`declare_drm_ioctls!`] to anchor type inference.
|
||||
#[doc(hidden)]
|
||||
#[inline]
|
||||
pub const fn __dev_ctx_cast<T: crate::drm::Driver>(
|
||||
ptr: *const crate::drm::Device<T, crate::drm::Ioctl>,
|
||||
) -> *const crate::drm::Device<T, crate::drm::Registered> {
|
||||
ptr.cast()
|
||||
}
|
||||
}
|
||||
|
||||
/// Declare the DRM ioctls for a driver.
|
||||
|
|
@ -82,7 +94,8 @@ pub mod internal {
|
|||
/// `user_callback` should have the following prototype:
|
||||
///
|
||||
/// ```ignore
|
||||
/// fn foo(device: &kernel::drm::Device<Self>,
|
||||
/// fn foo(device: &kernel::drm::Device<Self, kernel::drm::Registered>,
|
||||
/// reg_data: &Self::RegistrationData<'_>,
|
||||
/// data: &mut uapi::argument_type,
|
||||
/// file: &kernel::drm::File<Self::File>,
|
||||
/// ) -> Result<u32>
|
||||
|
|
@ -131,10 +144,45 @@ macro_rules! declare_drm_ioctls {
|
|||
// - The DRM device must have been registered when we're called through
|
||||
// an IOCTL.
|
||||
//
|
||||
// INVARIANT: The `Ioctl` context requires that the device has been
|
||||
// registered via `drm_dev_register()` at some point; the DRM core
|
||||
// guarantees this for ioctl dispatch callbacks.
|
||||
//
|
||||
// FIXME: Currently there is nothing enforcing that the types of the
|
||||
// dev/file match the current driver these ioctls are being declared
|
||||
// for, and it's not clear how to enforce this within the type system.
|
||||
let dev = $crate::drm::device::Device::from_raw(raw_dev);
|
||||
let dev: &$crate::drm::device::Device<_, $crate::drm::Ioctl> =
|
||||
$crate::drm::device::Device::from_raw(raw_dev);
|
||||
|
||||
// Type-inference anchor: the closure is never called but ties `dev`'s
|
||||
// type to `$func`'s first parameter, which the compiler cannot infer
|
||||
// through method resolution and associated-type projections alone.
|
||||
#[allow(unreachable_code)]
|
||||
let _ = || {
|
||||
let __ptr = $crate::drm::ioctl::internal::__dev_ctx_cast(
|
||||
::core::ptr::from_ref(dev),
|
||||
);
|
||||
|
||||
$func(
|
||||
// SAFETY: This closure is never executed; the dereference
|
||||
// exists purely to unify the type parameter with `$func`.
|
||||
// The pointer is valid regardless.
|
||||
unsafe { &*__ptr },
|
||||
unreachable!(),
|
||||
unreachable!(),
|
||||
unreachable!(),
|
||||
)
|
||||
};
|
||||
|
||||
// Enforce that the handler accepts higher-ranked
|
||||
// lifetimes, preventing it from requiring 'static
|
||||
// references that could escape this scope.
|
||||
let _: for<'a> fn(&'a _, &'a _, &'a mut _, &'a _) -> _ = $func;
|
||||
|
||||
let Some(guard) = dev.registration_guard() else {
|
||||
return $crate::error::code::ENODEV.to_errno();
|
||||
};
|
||||
|
||||
// SAFETY: The ioctl argument has size `_IOC_SIZE(cmd)`, which we
|
||||
// asserted above matches the size of this type, and all bit patterns of
|
||||
// UAPI structs must be valid.
|
||||
|
|
@ -147,7 +195,9 @@ macro_rules! declare_drm_ioctls {
|
|||
// SAFETY: This is just the DRM file structure
|
||||
let file = unsafe { $crate::drm::File::from_raw(raw_file) };
|
||||
|
||||
match $func(dev, data, file) {
|
||||
match guard.registration_data_with(|reg_data| {
|
||||
$func(&*guard, reg_data, data, file)
|
||||
}) {
|
||||
Err(e) => e.to_errno(),
|
||||
Ok(i) => i.try_into()
|
||||
.unwrap_or($crate::error::code::ERANGE.to_errno()),
|
||||
|
|
|
|||
|
|
@ -11,8 +11,10 @@
|
|||
|
||||
pub use self::device::Device;
|
||||
pub use self::device::DeviceContext;
|
||||
pub use self::device::Ioctl;
|
||||
pub use self::device::Normal;
|
||||
pub use self::device::Registered;
|
||||
pub use self::device::Uninit;
|
||||
pub use self::device::RegistrationGuard;
|
||||
pub use self::device::UnregisteredDevice;
|
||||
pub use self::driver::Driver;
|
||||
pub use self::driver::DriverInfo;
|
||||
|
|
|
|||
|
|
@ -9,15 +9,63 @@
|
|||
use crate::{
|
||||
bindings,
|
||||
device,
|
||||
prelude::*, //
|
||||
prelude::*,
|
||||
types::Opaque, //
|
||||
};
|
||||
use core::ptr::{
|
||||
addr_of_mut,
|
||||
null,
|
||||
null_mut,
|
||||
NonNull, //
|
||||
use core::{
|
||||
marker::PhantomData,
|
||||
ptr::{
|
||||
null,
|
||||
null_mut,
|
||||
NonNull, //
|
||||
},
|
||||
};
|
||||
|
||||
/// A faux device.
|
||||
///
|
||||
/// A faux device is a virtual device backed by the faux bus, primarily used for scenarios where a
|
||||
/// real hardware device is not available or for testing.
|
||||
///
|
||||
/// # Invariants
|
||||
///
|
||||
/// The underlying `struct faux_device` is valid.
|
||||
#[repr(transparent)]
|
||||
pub struct Device<Ctx: device::DeviceContext = device::Normal>(
|
||||
Opaque<bindings::faux_device>,
|
||||
PhantomData<Ctx>,
|
||||
);
|
||||
|
||||
impl<Ctx: device::DeviceContext> Device<Ctx> {
|
||||
#[inline]
|
||||
fn as_raw(&self) -> *mut bindings::faux_device {
|
||||
self.0.get()
|
||||
}
|
||||
|
||||
/// # Safety
|
||||
///
|
||||
/// `ptr` must be a valid pointer to a `struct faux_device`.
|
||||
#[inline]
|
||||
unsafe fn from_raw<'a>(ptr: *mut bindings::faux_device) -> &'a Self {
|
||||
// SAFETY: `Device` is a transparent wrapper of `Opaque<bindings::faux_device>`.
|
||||
unsafe { &*ptr.cast() }
|
||||
}
|
||||
}
|
||||
|
||||
impl<Ctx: device::DeviceContext> AsRef<device::Device<Ctx>> for Device<Ctx> {
|
||||
#[inline]
|
||||
fn as_ref(&self) -> &device::Device<Ctx> {
|
||||
// SAFETY: By the type invariant of `Self`, `self.as_raw()` is a pointer to a valid
|
||||
// `struct faux_device`. `dev` points to a valid `struct device`.
|
||||
unsafe { device::Device::from_raw(&raw mut (*self.as_raw()).dev) }
|
||||
}
|
||||
}
|
||||
|
||||
// SAFETY: `faux::Device` is a transparent wrapper of `struct faux_device`.
|
||||
// The offset is guaranteed to point to a valid device field inside `faux::Device`.
|
||||
unsafe impl<Ctx: device::DeviceContext> device::AsBusDevice<Ctx> for Device<Ctx> {
|
||||
const OFFSET: usize = core::mem::offset_of!(bindings::faux_device, dev);
|
||||
}
|
||||
|
||||
/// The registration of a faux device.
|
||||
///
|
||||
/// This type represents the registration of a [`struct faux_device`]. When an instance of this type
|
||||
|
|
@ -25,7 +73,8 @@
|
|||
///
|
||||
/// # Invariants
|
||||
///
|
||||
/// `self.0` always holds a valid pointer to an initialized and registered [`struct faux_device`].
|
||||
/// - `self.0` always holds a valid pointer to an initialized and registered [`struct faux_device`].
|
||||
/// - This object is proof that the object described by this `Registration` is bound to a device.
|
||||
///
|
||||
/// [`struct faux_device`]: srctree/include/linux/device/faux.h
|
||||
pub struct Registration(NonNull<bindings::faux_device>);
|
||||
|
|
@ -59,11 +108,19 @@ fn as_raw(&self) -> *mut bindings::faux_device {
|
|||
}
|
||||
}
|
||||
|
||||
impl AsRef<device::Device> for Registration {
|
||||
fn as_ref(&self) -> &device::Device {
|
||||
// SAFETY: The underlying `device` in `faux_device` is guaranteed by the C API to be
|
||||
// a valid initialized `device`.
|
||||
unsafe { device::Device::from_raw(addr_of_mut!((*self.as_raw()).dev)) }
|
||||
impl AsRef<Device<device::Bound>> for Registration {
|
||||
#[inline]
|
||||
fn as_ref(&self) -> &Device<device::Bound> {
|
||||
// SAFETY:
|
||||
// - The underlying `struct faux_device` is guaranteed by the C API to be a valid
|
||||
// initialized `device`.
|
||||
// - `faux_match()` always returns 1, and probe runs synchronously
|
||||
// (PROBE_FORCE_SYNCHRONOUS).
|
||||
// - `suppress_bind_attrs = true` on faux_driver prevents userspace-triggered unbind via
|
||||
// sysfs.
|
||||
// - `mem::forget(Registration)` is not a problem; if the `Registration` is leaked, the faux
|
||||
// device stays bound forever.
|
||||
unsafe { Device::from_raw(self.as_raw()) }
|
||||
}
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -7,9 +7,9 @@
|
|||
use crate::{
|
||||
bindings,
|
||||
device::Device,
|
||||
error::Error,
|
||||
error::Result,
|
||||
error::to_result,
|
||||
ffi,
|
||||
prelude::*,
|
||||
str::{CStr, CStrExt as _},
|
||||
};
|
||||
use core::ptr::NonNull;
|
||||
|
|
@ -120,6 +120,48 @@ fn drop(&mut self) {
|
|||
}
|
||||
}
|
||||
|
||||
/// Load firmware directly into the caller-provided `buf`.
|
||||
///
|
||||
/// On success the firmware image has been copied into `buf`; the caller accesses the data
|
||||
/// through `buf` itself.
|
||||
///
|
||||
/// This is intentionally a stand-alone function rather than a `Firmware` constructor. For
|
||||
/// the `into_buf` path, the firmware data lives in the caller's `buf`, not in a
|
||||
/// kernel-owned buffer, so returning a `Firmware` would expose `Firmware::data()` as a
|
||||
/// second handle aliasing `buf` (and `release_firmware()` does not free `buf` anyway).
|
||||
pub fn request_into_buf(name: &CStr, dev: &Device, buf: &mut [u8]) -> Result {
|
||||
// `as_mut_ptr()` on an empty slice returns a non-NULL pointer to
|
||||
// memory which the loader does not own. Passing that pointer with `size == 0`
|
||||
// makes the loader believe that it is buffer it allocated itself, so when
|
||||
// `release_firmware()` is called, it will vfree the pointer and trigger a
|
||||
// bug. Reject empty slices to avoid this situation.
|
||||
if buf.is_empty() {
|
||||
return Err(EINVAL);
|
||||
}
|
||||
|
||||
let mut fw: *const bindings::firmware = core::ptr::null();
|
||||
|
||||
// SAFETY: `&raw mut fw` is a valid pointer to a NULL initialized `bindings::firmware` pointer.
|
||||
// `name` and `dev` are valid as by their type invariants. `buf` is a valid writable
|
||||
// buffer of `buf.len()` bytes.
|
||||
to_result(unsafe {
|
||||
bindings::request_firmware_into_buf(
|
||||
&raw mut fw,
|
||||
name.as_char_ptr(),
|
||||
dev.as_raw(),
|
||||
buf.as_mut_ptr().cast(),
|
||||
buf.len(),
|
||||
)
|
||||
})?;
|
||||
|
||||
// The firmware bytes are now in `buf`, which the caller owns, so we don't need
|
||||
// the kernel to hang on to it any more.
|
||||
// SAFETY: `fw` is a valid pointer returned by `request_firmware_into_buf`.
|
||||
unsafe { bindings::release_firmware(fw) };
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
// SAFETY: `Firmware` only holds a pointer to a C `struct firmware`, which is safe to be used from
|
||||
// any thread.
|
||||
unsafe impl Send for Firmware {}
|
||||
|
|
|
|||
1526
rust/kernel/io.rs
1526
rust/kernel/io.rs
File diff suppressed because it is too large
Load Diff
|
|
@ -2,8 +2,6 @@
|
|||
|
||||
//! Generic memory-mapped IO.
|
||||
|
||||
use core::ops::Deref;
|
||||
|
||||
use crate::{
|
||||
device::{
|
||||
Bound,
|
||||
|
|
@ -16,7 +14,9 @@
|
|||
Region,
|
||||
Resource, //
|
||||
},
|
||||
IoBase,
|
||||
Mmio,
|
||||
MmioBackend,
|
||||
MmioRaw, //
|
||||
},
|
||||
prelude::*,
|
||||
|
|
@ -210,11 +210,13 @@ pub fn into_devres(self) -> Result<Devres<ExclusiveIoMem<'static, SIZE>>> {
|
|||
}
|
||||
}
|
||||
|
||||
impl<const SIZE: usize> Deref for ExclusiveIoMem<'_, SIZE> {
|
||||
type Target = Mmio<SIZE>;
|
||||
impl<'a, const SIZE: usize> IoBase<'a> for &'a ExclusiveIoMem<'_, SIZE> {
|
||||
type Backend = MmioBackend;
|
||||
type Target = super::Region<SIZE>;
|
||||
|
||||
fn deref(&self) -> &Self::Target {
|
||||
&self.iomem
|
||||
#[inline]
|
||||
fn as_view(self) -> Mmio<'a, Self::Target> {
|
||||
self.iomem.as_view()
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -229,7 +231,7 @@ fn deref(&self) -> &Self::Target {
|
|||
/// start of the I/O memory mapped region.
|
||||
pub struct IoMem<'a, const SIZE: usize = 0> {
|
||||
dev: &'a Device<Bound>,
|
||||
io: MmioRaw<SIZE>,
|
||||
io: MmioRaw<super::Region<SIZE>>,
|
||||
}
|
||||
|
||||
impl<'a, const SIZE: usize> IoMem<'a, SIZE> {
|
||||
|
|
@ -264,8 +266,7 @@ fn ioremap(dev: &'a Device<Bound>, resource: &Resource) -> Result<Self> {
|
|||
return Err(ENOMEM);
|
||||
}
|
||||
|
||||
let io = MmioRaw::new(addr as usize, size)?;
|
||||
|
||||
let io = MmioRaw::new_region(addr as usize, size)?;
|
||||
Ok(IoMem { dev, io })
|
||||
}
|
||||
|
||||
|
|
@ -291,11 +292,13 @@ fn drop(&mut self) {
|
|||
}
|
||||
}
|
||||
|
||||
impl<const SIZE: usize> Deref for IoMem<'_, SIZE> {
|
||||
type Target = Mmio<SIZE>;
|
||||
impl<'a, const SIZE: usize> IoBase<'a> for &'a IoMem<'_, SIZE> {
|
||||
type Backend = MmioBackend;
|
||||
type Target = super::Region<SIZE>;
|
||||
|
||||
fn deref(&self) -> &Self::Target {
|
||||
#[inline]
|
||||
fn as_view(self) -> Mmio<'a, Self::Target> {
|
||||
// SAFETY: Safe as by the invariant of `IoMem`.
|
||||
unsafe { Mmio::from_raw(&self.io) }
|
||||
unsafe { Mmio::from_raw(self.io) }
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -48,13 +48,14 @@
|
|||
/// use kernel::io::{
|
||||
/// Io,
|
||||
/// Mmio,
|
||||
/// Region,
|
||||
/// poll::read_poll_timeout, //
|
||||
/// };
|
||||
/// use kernel::time::Delta;
|
||||
///
|
||||
/// const HW_READY: u16 = 0x01;
|
||||
///
|
||||
/// fn wait_for_hardware<const SIZE: usize>(io: &Mmio<SIZE>) -> Result {
|
||||
/// fn wait_for_hardware<const SIZE: usize>(io: Mmio<'_, Region<SIZE>>) -> Result {
|
||||
/// read_poll_timeout(
|
||||
/// // The `op` closure reads the value of a specific status register.
|
||||
/// || io.try_read16(0x1000),
|
||||
|
|
@ -135,13 +136,14 @@ pub fn read_poll_timeout<Op, Cond, T>(
|
|||
/// use kernel::io::{
|
||||
/// Io,
|
||||
/// Mmio,
|
||||
/// Region,
|
||||
/// poll::read_poll_timeout_atomic, //
|
||||
/// };
|
||||
/// use kernel::time::Delta;
|
||||
///
|
||||
/// const HW_READY: u16 = 0x01;
|
||||
///
|
||||
/// fn wait_for_hardware<const SIZE: usize>(io: &Mmio<SIZE>) -> Result {
|
||||
/// fn wait_for_hardware<const SIZE: usize>(io: Mmio<'_, Region<SIZE>>) -> Result {
|
||||
/// read_poll_timeout_atomic(
|
||||
/// // The `op` closure reads the value of a specific status register.
|
||||
/// || io.try_read16(0x1000),
|
||||
|
|
|
|||
|
|
@ -58,7 +58,7 @@
|
|||
//! },
|
||||
//! num::Bounded,
|
||||
//! };
|
||||
//! # use kernel::io::Mmio;
|
||||
//! # use kernel::io::{Mmio, Region};
|
||||
//! # register! {
|
||||
//! # pub BOOT_0(u32) @ 0x00000100 {
|
||||
//! # 15:8 vendor_id;
|
||||
|
|
@ -66,7 +66,7 @@
|
|||
//! # 3:0 minor_revision;
|
||||
//! # }
|
||||
//! # }
|
||||
//! # fn test(io: &Mmio<0x1000>) {
|
||||
//! # fn test(io: Mmio<'_, Region<0x1000>>) {
|
||||
//! # fn obtain_vendor_id() -> u8 { 0xff }
|
||||
//!
|
||||
//! // Read from the register's defined offset (0x100).
|
||||
|
|
@ -113,6 +113,8 @@
|
|||
io::IoLoc, //
|
||||
};
|
||||
|
||||
use super::Region;
|
||||
|
||||
/// Trait implemented by all registers.
|
||||
pub trait Register: Sized {
|
||||
/// Backing primitive type of the register.
|
||||
|
|
@ -129,7 +131,7 @@ pub trait FixedRegister: Register {}
|
|||
|
||||
/// Allows `()` to be used as the `location` parameter of [`Io::write`](super::Io::write) when
|
||||
/// passing a [`FixedRegister`] value.
|
||||
impl<T> IoLoc<T> for ()
|
||||
impl<const SIZE: usize, T> IoLoc<Region<SIZE>, T> for ()
|
||||
where
|
||||
T: FixedRegister,
|
||||
{
|
||||
|
|
@ -143,7 +145,7 @@ fn offset(self) -> usize {
|
|||
|
||||
/// A [`FixedRegister`] carries its location in its type. Thus `FixedRegister` values can be used
|
||||
/// as an [`IoLoc`].
|
||||
impl<T> IoLoc<T> for T
|
||||
impl<const SIZE: usize, T> IoLoc<Region<SIZE>, T> for T
|
||||
where
|
||||
T: FixedRegister,
|
||||
{
|
||||
|
|
@ -168,7 +170,7 @@ pub const fn new() -> Self {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T> IoLoc<T> for FixedRegisterLoc<T>
|
||||
impl<const SIZE: usize, T> IoLoc<Region<SIZE>, T> for FixedRegisterLoc<T>
|
||||
where
|
||||
T: FixedRegister,
|
||||
{
|
||||
|
|
@ -239,7 +241,7 @@ const fn offset(self) -> usize {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T, B> IoLoc<T> for RelativeRegisterLoc<T, B>
|
||||
impl<const SIZE: usize, T, B> IoLoc<Region<SIZE>, T> for RelativeRegisterLoc<T, B>
|
||||
where
|
||||
T: RelativeRegister,
|
||||
B: RegisterBase<T::BaseFamily> + ?Sized,
|
||||
|
|
@ -283,7 +285,7 @@ pub fn try_new(idx: usize) -> Option<Self> {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T> IoLoc<T> for RegisterArrayLoc<T>
|
||||
impl<const SIZE: usize, T> IoLoc<Region<SIZE>, T> for RegisterArrayLoc<T>
|
||||
where
|
||||
T: RegisterArray,
|
||||
{
|
||||
|
|
@ -370,7 +372,7 @@ pub fn try_at(self, idx: usize) -> Option<RelativeRegisterArrayLoc<T, B>> {
|
|||
}
|
||||
}
|
||||
|
||||
impl<T, B> IoLoc<T> for RelativeRegisterArrayLoc<T, B>
|
||||
impl<const SIZE: usize, T, B> IoLoc<Region<SIZE>, T> for RelativeRegisterArrayLoc<T, B>
|
||||
where
|
||||
T: RelativeRegisterArray,
|
||||
B: RegisterBase<T::BaseFamily> + ?Sized,
|
||||
|
|
@ -387,18 +389,18 @@ fn offset(self) -> usize {
|
|||
/// which to write it.
|
||||
///
|
||||
/// Implementors can be used with [`Io::write_reg`](super::Io::write_reg).
|
||||
pub trait LocatedRegister {
|
||||
pub trait LocatedRegister<Base: ?Sized> {
|
||||
/// Register value to write.
|
||||
type Value: Register;
|
||||
/// Full location information at which to write the value.
|
||||
type Location: IoLoc<Self::Value>;
|
||||
type Location: IoLoc<Base, Self::Value>;
|
||||
|
||||
/// Consumes `self` and returns a `(location, value)` tuple describing a valid I/O write
|
||||
/// operation.
|
||||
fn into_io_op(self) -> (Self::Location, Self::Value);
|
||||
}
|
||||
|
||||
impl<T> LocatedRegister for T
|
||||
impl<const SIZE: usize, T> LocatedRegister<Region<SIZE>> for T
|
||||
where
|
||||
T: FixedRegister,
|
||||
{
|
||||
|
|
@ -444,7 +446,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// Io,
|
||||
/// },
|
||||
/// };
|
||||
/// # use kernel::io::Mmio;
|
||||
/// # use kernel::io::{Mmio, Region};
|
||||
///
|
||||
/// register! {
|
||||
/// FIXED_REG(u32) @ 0x100 {
|
||||
|
|
@ -453,7 +455,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// }
|
||||
/// }
|
||||
///
|
||||
/// # fn test(io: &Mmio<0x1000>) {
|
||||
/// # fn test(io: Mmio<'_, Region<0x1000>>) {
|
||||
/// let val = io.read(FIXED_REG);
|
||||
///
|
||||
/// // Write from an already-existing value.
|
||||
|
|
@ -557,7 +559,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// Io,
|
||||
/// },
|
||||
/// };
|
||||
/// # use kernel::io::Mmio;
|
||||
/// # use kernel::io::{Mmio, Region};
|
||||
///
|
||||
/// // Type used to identify the base.
|
||||
/// pub struct CpuCtlBase;
|
||||
|
|
@ -582,7 +584,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// }
|
||||
/// }
|
||||
///
|
||||
/// # fn test(io: Mmio<0x1000>) {
|
||||
/// # fn test(io: Mmio<'_, Region<0x1000>>) {
|
||||
/// // Read the status of `Cpu0`.
|
||||
/// let cpu0_started = io.read(CPU_CTL::of::<Cpu0>());
|
||||
///
|
||||
|
|
@ -599,7 +601,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// }
|
||||
/// }
|
||||
///
|
||||
/// # fn test2(io: Mmio<0x1000>) {
|
||||
/// # fn test2(io: Mmio<'_, Region<0x1000>>) {
|
||||
/// // Start the aliased `CPU0`, leaving its other fields untouched.
|
||||
/// io.update(CPU_CTL_ALIAS::of::<Cpu0>(), |r| r.with_alias_start(true));
|
||||
/// # }
|
||||
|
|
@ -636,7 +638,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// Io,
|
||||
/// },
|
||||
/// };
|
||||
/// # use kernel::io::Mmio;
|
||||
/// # use kernel::io::{Mmio, Region};
|
||||
/// # fn get_scratch_idx() -> usize {
|
||||
/// # 0x15
|
||||
/// # }
|
||||
|
|
@ -649,7 +651,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// }
|
||||
/// }
|
||||
///
|
||||
/// # fn test(io: &Mmio<0x1000>)
|
||||
/// # fn test(io: Mmio<'_, Region<0x1000>>)
|
||||
/// # -> Result<(), Error>{
|
||||
/// // Read scratch register 0, i.e. I/O address `0x80`.
|
||||
/// let scratch_0 = io.read(SCRATCH::at(0)).value();
|
||||
|
|
@ -722,7 +724,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// Io,
|
||||
/// },
|
||||
/// };
|
||||
/// # use kernel::io::Mmio;
|
||||
/// # use kernel::io::{Mmio, Region};
|
||||
/// # fn get_scratch_idx() -> usize {
|
||||
/// # 0x15
|
||||
/// # }
|
||||
|
|
@ -750,7 +752,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// }
|
||||
/// }
|
||||
///
|
||||
/// # fn test(io: &Mmio<0x1000>) -> Result<(), Error> {
|
||||
/// # fn test(io: Mmio<'_, Region<0x1000>>) -> Result<(), Error> {
|
||||
/// // Read scratch register 0 of CPU0.
|
||||
/// let scratch = io.read(CPU_SCRATCH::of::<Cpu0>().at(0));
|
||||
///
|
||||
|
|
@ -792,7 +794,7 @@ fn into_io_op(self) -> (FixedRegisterLoc<T>, T) {
|
|||
/// }
|
||||
/// }
|
||||
///
|
||||
/// # fn test2(io: &Mmio<0x1000>) -> Result<(), Error> {
|
||||
/// # fn test2(io: Mmio<'_, Region<0x1000>>) -> Result<(), Error> {
|
||||
/// let cpu0_status = io.read(CPU_FIRMWARE_STATUS::of::<Cpu0>()).status();
|
||||
/// # Ok(())
|
||||
/// # }
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show More
Loading…
Reference in New Issue
Block a user