Unified configuration for metrics (#726)
* Show description of metrics during analysis
* Use --include-cols Description show the Description column in analyze mode (this is hidden by default)
* Remove tips field from analysis config
* Align metric names in analysis config and documentation
* Add unified config utils/unified_config.yaml
* Add python script utils/split_config.py to auto generate analysis configuration and documentation metrics description
* Add test case to ensure unified config is older than auto-generated config
* Auto generate analysis config and documentation metrics description
* Update CONTRIBUTING.md to add instructions to build documentation assets
* Add docker image and compose file to build documentation
* Update CHANGELOG and Documentation
* Use jinja template instead of hardcoding metric tables in documentation
[ROCm/rocprofiler-compute commit: bb44e90b2d]
This commit is contained in:
@@ -48,56 +48,8 @@ The L2 cache’s speed-of-light table contains a few key metrics about the
|
||||
performance of the L2 cache, aggregated over all the L2 channels, as a
|
||||
comparison with the peak achievable values of those metrics:
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
|
||||
* - Metric
|
||||
|
||||
- Description
|
||||
|
||||
- Unit
|
||||
|
||||
* - Utilization
|
||||
|
||||
- The ratio of the
|
||||
:ref:`number of cycles an L2 channel was active, summed over all L2 channels on the accelerator <total-active-l2-cycles>`
|
||||
over the :ref:`total L2 cycles <total-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Bandwidth
|
||||
|
||||
- The number of bytes looked up in the L2 cache, as a percent of the peak
|
||||
theoretical bandwidth achievable on the specific accelerator. The number
|
||||
of bytes is calculated as the number of cache lines requested multiplied
|
||||
by the cache line size. This value does not consider partial requests, so
|
||||
e.g., if only a single value is requested in a cache line, the data
|
||||
movement will still be counted as a full cache line.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Hit Rate
|
||||
|
||||
- The ratio of the number of L2 cache line requests that hit in the L2
|
||||
cache over the total number of incoming cache line requests to the L2
|
||||
cache.
|
||||
|
||||
- Percent
|
||||
|
||||
* - L2-Fabric Read BW
|
||||
|
||||
- The number of bytes read by the L2 over the
|
||||
:ref:`Infinity Fabric interface <l2-fabric>` per unit time.
|
||||
|
||||
- GB/s
|
||||
|
||||
* - L2-Fabric Write and Atomic BW
|
||||
|
||||
- The number of bytes sent by the L2 over the
|
||||
:ref:`Infinity Fabric interface <l2-fabric>` by write and atomic
|
||||
operations per unit time.
|
||||
|
||||
- GB/s
|
||||
.. jinja:: l2-sol
|
||||
:file: _templates/metrics_table.j2
|
||||
|
||||
.. note::
|
||||
|
||||
@@ -117,168 +69,8 @@ This section details the incoming requests to the L2 cache from the
|
||||
:doc:`vL1D <vector-l1-cache>` and other clients -- for instance, the
|
||||
:ref:`sL1D <desc-sL1D>` and :ref:`L1I <desc-l1i>` caches.
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
:widths: 13 70 17
|
||||
|
||||
* - Metric
|
||||
|
||||
- Description
|
||||
|
||||
- Unit
|
||||
|
||||
* - Bandwidth
|
||||
|
||||
- The number of bytes looked up in the L2 cache, per
|
||||
:ref:`normalization unit <normalization-units>`. The number of bytes is
|
||||
calculated as the number of cache lines requested multiplied by the cache
|
||||
line size. This value does not consider partial requests, so for example,
|
||||
if only a single value is requested in a cache line, the data movement
|
||||
will still be counted as a full cache line.
|
||||
|
||||
- Bytes per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Requests
|
||||
|
||||
- The total number of incoming requests to the L2 from all clients for all
|
||||
request types, per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Read Requests
|
||||
|
||||
- The total number of read requests to the L2 from all clients.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`
|
||||
|
||||
* - Write Requests
|
||||
|
||||
- The total number of write requests to the L2 from all clients.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`
|
||||
|
||||
* - Atomic Requests
|
||||
|
||||
- The total number of atomic requests (with and without return) to the L2
|
||||
from all clients.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`
|
||||
|
||||
* - Streaming Requests
|
||||
|
||||
- The total number of incoming requests to the L2 that are marked as
|
||||
*streaming*. The exact meaning of this may differ depending on the
|
||||
targeted accelerator, however on an :ref:`MI2XX <mixxx-note>` this
|
||||
corresponds to
|
||||
`non-temporal load or stores <https://clang.llvm.org/docs/LanguageExtensions.html#non-temporal-load-store-builtins>`_.
|
||||
The L2 cache attempts to evict *streaming* requests before normal
|
||||
requests when the L2 is at capacity.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`
|
||||
|
||||
* - Probe Requests
|
||||
|
||||
- The number of coherence probe requests made to the L2 cache from outside
|
||||
the accelerator. On an :ref:`MI2XX <mixxx-note>`, probe requests may be
|
||||
generated by, for example, writes to
|
||||
:ref:`fine-grained device <memory-type>` memory or by writes to
|
||||
:ref:`coarse-grained <memory-type>` device memory.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`
|
||||
|
||||
* - Hit Rate
|
||||
|
||||
- The ratio of the number of L2 cache line requests that hit in the L2
|
||||
cache over the total number of incoming cache line requests to the L2
|
||||
cache.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Hits
|
||||
|
||||
- The total number of requests to the L2 from all clients that hit in the
|
||||
cache. As noted in the :ref:`Speed-of-Light <l2-sol>` section, this
|
||||
includes hit-on-miss requests.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`
|
||||
|
||||
* - Misses
|
||||
|
||||
- The total number of requests to the L2 from all clients that miss in the
|
||||
cache. As noted in the :ref:`Speed-of-Light <l2-sol>` section, these do
|
||||
not include hit-on-miss requests.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`
|
||||
|
||||
* - Writebacks
|
||||
|
||||
- The total number of L2 cache lines written back to memory for any reason.
|
||||
Write-backs may occur due to user code (such as HIP kernel calls to
|
||||
``__threadfence_system`` or atomic built-ins) by the
|
||||
:doc:`command processor <command-processor>`'s memory acquire/release
|
||||
fences, or for other internal hardware reasons.
|
||||
|
||||
- Cache lines per :ref:`normalization unit <normalization-units>`
|
||||
|
||||
* - Writebacks (Internal)
|
||||
|
||||
- The total number of L2 cache lines written back to memory for internal
|
||||
hardware reasons, per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
- Cache lines per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Writebacks (vL1D Req)
|
||||
|
||||
- The total number of L2 cache lines written back to memory due to requests
|
||||
initiated by the :doc:`vL1D cache <vector-l1-cache>`, per
|
||||
:ref:`normalization unit <normalization-units>`.
|
||||
|
||||
- Cache lines per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Evictions (Normal)
|
||||
|
||||
- The total number of L2 cache lines evicted from the cache due to capacity
|
||||
limits, per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
- Cache lines per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Evictions (vL1D Req)
|
||||
|
||||
- The total number of L2 cache lines evicted from the cache due to
|
||||
invalidation requests initiated by the
|
||||
:doc:`vL1D cache <vector-l1-cache>`, per
|
||||
:ref:`normalization unit <normalization-units>`.
|
||||
|
||||
- Cache lines per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Non-hardware-Coherent Requests
|
||||
|
||||
- The total number of requests to the L2 to Not-hardware-Coherent (NC)
|
||||
memory allocations, per :ref:`normalization unit <normalization-units>`.
|
||||
See the :ref:`memory-type` for more information.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Uncached Requests
|
||||
|
||||
- The total number of requests to the L2 that go to Uncached (UC) memory
|
||||
allocations. See the :ref:`memory-type` for more information.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Coherently Cached Requests
|
||||
|
||||
- The total number of requests to the L2 that go to Coherently Cacheable (CC)
|
||||
memory allocations. See the :ref:`memory-type` for more information.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Read/Write Coherent Requests
|
||||
|
||||
- The total number of requests to the L2 that go to Read-Write coherent memory
|
||||
(RW) allocations. See the :ref:`memory-type` for more information.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
.. jinja:: l2-cache-accesses
|
||||
:file: _templates/metrics_table.j2
|
||||
|
||||
.. note::
|
||||
|
||||
@@ -300,7 +92,7 @@ is responsible for routing these memory requests/data to the correct
|
||||
location and returning any fetched data to the L2 cache. The
|
||||
:ref:`l2-request-flow` describes the flow of these requests through
|
||||
Infinity Fabric in more detail, as described by ROCm Compute Profiler metrics,
|
||||
while :ref:`l2-request-metrics` give detailed definitions of
|
||||
while :ref:`l2-fabric` give detailed definitions of
|
||||
individual metrics.
|
||||
|
||||
.. _l2-request-flow:
|
||||
@@ -363,176 +155,15 @@ to uncached memory (denoted by the dashed line), they will also be
|
||||
counted as *two* uncached read requests (that is, the request is split).
|
||||
|
||||
|
||||
.. _l2-request-metrics:
|
||||
.. _l2-fabric-metrics:
|
||||
|
||||
Metrics
|
||||
-------
|
||||
|
||||
The following metrics are reported for the L2-Fabric interface:
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
|
||||
* - Metric
|
||||
|
||||
- Description
|
||||
|
||||
- Unit
|
||||
|
||||
* - L2-Fabric Read Bandwidth
|
||||
|
||||
- The total number of bytes read by the L2 cache from Infinity Fabric per
|
||||
:ref:`normalization unit <normalization-units>`.
|
||||
|
||||
- Bytes per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - HBM Read Traffic
|
||||
|
||||
- The percent of read requests generated by the L2 cache that are routed to
|
||||
the accelerator's local high-bandwidth memory (HBM). This breakdown does
|
||||
not consider the *size* of the request (meaning that 32B and 64B requests
|
||||
are both counted as a single request), so this metric only *approximates*
|
||||
the percent of the L2-Fabric Read bandwidth directed to the local HBM.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Remote Read Traffic
|
||||
|
||||
- The percent of read requests generated by the L2 cache that are routed to
|
||||
any memory location other than the accelerator's local high-bandwidth
|
||||
memory (HBM) -- for example, the CPU's DRAM or a remote accelerator's
|
||||
HBM. This breakdown does not consider the *size* of the request (meaning
|
||||
that 32B and 64B requests are both counted as a single request), so this
|
||||
metric only *approximates* the percent of the L2-Fabric Read bandwidth
|
||||
directed to a remote location.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Uncached Read Traffic
|
||||
|
||||
- The percent of read requests generated by the L2 cache that are reading
|
||||
from an :ref:`uncached memory allocation <memory-type>`. Note, as
|
||||
described in the :ref:`request flow <l2-request-flow>` section, a single
|
||||
64B read request is typically counted as two uncached read requests. So,
|
||||
it is possible for the Uncached Read Traffic to reach up to 200% of the
|
||||
total number of read requests. This breakdown does not consider the
|
||||
*size* of the request (i.e., 32B and 64B requests are both counted as a
|
||||
single request), so this metric only *approximates* the percent of the
|
||||
L2-Fabric read bandwidth directed to an uncached memory location.
|
||||
|
||||
- Percent
|
||||
|
||||
* - L2-Fabric Write and Atomic Bandwidth
|
||||
|
||||
- The total number of bytes written by the L2 over Infinity Fabric by write
|
||||
and atomic operations per
|
||||
:ref:`normalization unit <normalization-units>`. Note that on current
|
||||
CDNA accelerators, such as the :ref:`MI2XX <mixxx-note>`, requests are
|
||||
only considered *atomic* by Infinity Fabric if they are targeted at
|
||||
non-write-cacheable memory, for example,
|
||||
:ref:`fine-grained memory <memory-type>` allocations or
|
||||
:ref:`uncached memory <memory-type>` allocations on the
|
||||
MI2XX.
|
||||
|
||||
- Bytes per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - HBM Write and Atomic Traffic
|
||||
|
||||
- The percent of write and atomic requests generated by the L2 cache that
|
||||
are routed to the accelerator's local high-bandwidth memory (HBM). This
|
||||
breakdown does not consider the *size* of the request (meaning that 32B
|
||||
and 64B requests are both counted as a single request), so this metric
|
||||
only *approximates* the percent of the L2-Fabric Write and Atomic
|
||||
bandwidth directed to the local HBM. Note that on current CDNA
|
||||
accelerators, such as the :ref:`MI2XX <mixxx-note>`, requests are only
|
||||
considered *atomic* by Infinity Fabric if they are targeted at
|
||||
:ref:`fine-grained memory <memory-type>` allocations or
|
||||
:ref:`uncached memory <memory-type>` allocations.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Remote Write and Atomic Traffic
|
||||
|
||||
- The percent of read requests generated by the L2 cache that are routed to
|
||||
any memory location other than the accelerator's local high-bandwidth
|
||||
memory (HBM) -- for example, the CPU's DRAM or a remote accelerator's
|
||||
HBM. This breakdown does not consider the *size* of the request (meaning
|
||||
that 32B and 64B requests are both counted as a single request), so this
|
||||
metric only *approximates* the percent of the L2-Fabric Read bandwidth
|
||||
directed to a remote location. Note that on current CDNA
|
||||
accelerators, such as the :ref:`MI2XX <mixxx-note>`, requests are only
|
||||
considered *atomic* by Infinity Fabric if they are targeted at
|
||||
:ref:`fine-grained memory <memory-type>` allocations or
|
||||
:ref:`uncached memory <memory-type>` allocations.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Atomic Traffic
|
||||
|
||||
- The percent of write requests generated by the L2 cache that are atomic
|
||||
requests to *any* memory location. This breakdown does not consider the
|
||||
*size* of the request (meaning that 32B and 64B requests are both counted
|
||||
as a single request), so this metric only *approximates* the percent of
|
||||
the L2-Fabric Read bandwidth directed to a remote location. Note that on
|
||||
current CDNA accelerators, such as the :ref:`MI2XX <mixxx-note>`,
|
||||
requests are only considered *atomic* by Infinity Fabric if they are
|
||||
targeted at :ref:`fine-grained memory <memory-type>` allocations or
|
||||
:ref:`uncached memory <memory-type>` allocations.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Uncached Write and Atomic Traffic
|
||||
|
||||
- The percent of write and atomic requests generated by the L2 cache that
|
||||
are targeting :ref:`uncached memory allocations <memory-type>`. This
|
||||
breakdown does not consider the *size* of the request (meaning that 32B
|
||||
and 64B requests are both counted as a single request), so this metric
|
||||
only *approximates* the percent of the L2-Fabric read bandwidth directed
|
||||
to uncached memory allocations.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Read Latency
|
||||
|
||||
- The time-averaged number of cycles read requests spent in Infinity Fabric
|
||||
before data was returned to the L2.
|
||||
|
||||
- Cycles
|
||||
|
||||
* - Write Latency
|
||||
|
||||
- The time-averaged number of cycles write requests spent in Infinity
|
||||
Fabric before a completion acknowledgement was returned to the L2.
|
||||
|
||||
- Cycles
|
||||
|
||||
* - Atomic Latency
|
||||
|
||||
- The time-averaged number of cycles atomic requests spent in Infinity
|
||||
Fabric before a completion acknowledgement (atomic without return value)
|
||||
or data (atomic with return value) was returned to the L2.
|
||||
|
||||
- Cycles
|
||||
|
||||
* - Read Stall
|
||||
|
||||
- The ratio of the total number of cycles the L2-Fabric interface was
|
||||
stalled on a read request to any destination (local HBM, remote PCIe®
|
||||
connected accelerator or CPU, or remote Infinity Fabric connected
|
||||
accelerator [#inf]_ or CPU) over the
|
||||
:ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Write Stall
|
||||
|
||||
- The ratio of the total number of cycles the L2-Fabric interface was
|
||||
stalled on a write or atomic request to any destination (local HBM,
|
||||
remote accelerator or CPU, PCIe connected accelerator or CPU, or remote
|
||||
Infinity Fabric connected accelerator [#inf]_ or CPU) over the
|
||||
:ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
.. jinja:: l2-fabric-metrics
|
||||
:file: _templates/metrics_table.j2
|
||||
|
||||
.. _l2-detailed-metrics:
|
||||
|
||||
@@ -542,121 +173,8 @@ Detailed transaction metrics
|
||||
The following metrics are available in the detailed L2-Fabric
|
||||
transaction breakdown table:
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
|
||||
* - Metric
|
||||
|
||||
- Description
|
||||
|
||||
- Unit
|
||||
|
||||
* - 32B Read Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to read 32B of data
|
||||
from any memory location, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail. Typically unused on CDNA
|
||||
accelerators.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Uncached Read Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to read
|
||||
:ref:`uncached data <memory-type>` from any memory location, per
|
||||
:ref:`normalization unit <normalization-units>`. 64B requests for
|
||||
uncached data are counted as two 32B uncached data requests. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - 64B Read Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to read 64B of data
|
||||
from any memory location, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - HBM Read Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to read 32B or 64B of
|
||||
data from the accelerator's local HBM, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Remote Read Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to read 32B or 64B of
|
||||
data from any source other than the accelerator's local HBM, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - 32B Write and Atomic Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B of data to any memory location, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Uncached Write and Atomic Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of :ref:`uncached data <memory-type>`, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - 64B Write and Atomic Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 64B of data in any memory location, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - HBM Write and Atomic Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of data in the accelerator's local HBM, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Remote Write and Atomic Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to write or atomically
|
||||
update 32B or 64B of data in any memory location other than the
|
||||
accelerator's local HBM, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
|
||||
* - Atomic Requests
|
||||
|
||||
- The total number of L2 requests to Infinity Fabric to atomically update
|
||||
32B or 64B of data in any memory location, per
|
||||
:ref:`normalization unit <normalization-units>`. See
|
||||
:ref:`l2-request-flow` for more detail. Note that on current CDNA
|
||||
accelerators, such as the :ref:`MI2XX <mixxx-note>`, requests are only
|
||||
considered *atomic* by Infinity Fabric if they are targeted at
|
||||
non-write-cacheable memory, such as
|
||||
:ref:`fine-grained memory <memory-type>` allocations or
|
||||
:ref:`uncached memory <memory-type>` allocations on the MI2XX.
|
||||
|
||||
- Requests per :ref:`normalization unit <normalization-units>`.
|
||||
.. jinja:: l2-detailed-metrics
|
||||
:file: _templates/metrics_table.j2
|
||||
|
||||
.. _l2-fabric-stalls:
|
||||
|
||||
@@ -670,72 +188,8 @@ what types of requests in a kernel caused a stall (like read versus write), and
|
||||
to which locations -- for instance, to the accelerator’s local memory, or to
|
||||
remote accelerators or CPUs.
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
|
||||
* - Metric
|
||||
|
||||
- Description
|
||||
|
||||
- Unit
|
||||
|
||||
* - Read - PCIe Stall
|
||||
|
||||
- The number of cycles the L2-Fabric interface was stalled on read requests
|
||||
to remote PCIe connected accelerators [#inf]_ or CPUs as a percent of the
|
||||
:ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Read - Infinity Fabric Stall
|
||||
|
||||
- The number of cycles the L2-Fabric interface was stalled on read requests
|
||||
to remote Infinity Fabric connected accelerators [#inf]_ or CPUs as a
|
||||
percent of the :ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Read - HBM Stall
|
||||
|
||||
- The number of cycles the L2-Fabric interface was stalled on read requests
|
||||
to the accelerator's local HBM as a percent of the
|
||||
:ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Write - PCIe Stall
|
||||
|
||||
- The number of cycles the L2-Fabric interface was stalled on write or
|
||||
atomic requests to remote PCIe connected accelerators [#inf]_ or CPUs as
|
||||
a percent of the :ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Write - Infinity Fabric Stall
|
||||
|
||||
- The number of cycles the L2-Fabric interface was stalled on write or
|
||||
atomic requests to remote Infinity Fabric connected accelerators [#inf]_
|
||||
or CPUs as a percent of the
|
||||
:ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Write - HBM Stall
|
||||
|
||||
- The number of cycles the L2-Fabric interface was stalled on write or
|
||||
atomic requests to accelerator's local HBM as a percent of the
|
||||
:ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
|
||||
* - Write - Credit Starvation
|
||||
|
||||
- The number of cycles the L2-Fabric interface was stalled on write or
|
||||
atomic requests to any memory location because too many write/atomic
|
||||
requests were currently in flight, as a percent of the
|
||||
:ref:`total active L2 cycles <total-active-l2-cycles>`.
|
||||
|
||||
- Percent
|
||||
.. jinja:: l2-fabric-stalls
|
||||
:file: _templates/metrics_table.j2
|
||||
|
||||
.. warning::
|
||||
|
||||
|
||||
Reference in New Issue
Block a user