diff --git a/projects/rocprofiler-compute/CHANGELOG.md b/projects/rocprofiler-compute/CHANGELOG.md index 89a960cdbd..5c3efb9a4c 100644 --- a/projects/rocprofiler-compute/CHANGELOG.md +++ b/projects/rocprofiler-compute/CHANGELOG.md @@ -41,6 +41,10 @@ Full documentation for ROCm Compute Profiler is available at [https://rocm.docs. * Corrected peak VALU Roofline profiling and analysis by removing `FP8` VALU and `BF16` VALU benchmarking. +### Removed + +* Removed "VL1 Lat" metric for AMD Instinct MI300 series GPUs, due to MI300 series not supporting TCP_TCP_LATENCY_sum counter. + ## ROCm Compute Profiler 3.4.0 for ROCm 7.2.0 ### Added diff --git a/projects/rocprofiler-compute/docs/how-to/analyze/cli.rst b/projects/rocprofiler-compute/docs/how-to/analyze/cli.rst index 1404fd4666..22a57c5408 100644 --- a/projects/rocprofiler-compute/docs/how-to/analyze/cli.rst +++ b/projects/rocprofiler-compute/docs/how-to/analyze/cli.rst @@ -562,12 +562,8 @@ Analysis database example DEBUG Applied analysis mode filters DEBUG Calculated dispatch data DEBUG Collected metrics data - WARNING Failed to evaluate expression for 3.1.25 - Value: to_round(to_avg( - (pmc_df.get("TCP_TCP_LATENCY_sum") / pmc_df.get("TCP_TA_TCP_STATE_READ_sum")).where((pmc_df.get("TCP_TA_TCP_STATE_READ_sum") != 0), None)), 0) - unsupported operand type(s) for /: 'NoneType' and 'float' WARNING Failed to evaluate expression for 3.1.39 - Value: to_round((to_avg( (pmc_df.get("pmc_perf_ACCUM") / pmc_df.get("SQC_ICACHE_REQ")).where((pmc_df.get("SQC_ICACHE_REQ") != 0), None)) * 100), 0) - unsupported operand type(s) for /: 'NoneType' and 'float' - WARNING Failed to evaluate expression for 3.1.25 - Value: to_round(to_avg( - (pmc_df.get("TCP_TCP_LATENCY_sum") / pmc_df.get("TCP_TA_TCP_STATE_READ_sum")).where((pmc_df.get("TCP_TA_TCP_STATE_READ_sum") != 0), None)), 0) - unsupported operand type(s) for /: 'NoneType' and 'float' WARNING Failed to evaluate expression for 3.1.39 - Value: to_round((to_avg( (pmc_df.get("pmc_perf_ACCUM") / pmc_df.get("SQC_ICACHE_REQ")).where((pmc_df.get("SQC_ICACHE_REQ") != 0), None)) * 100), 0) - unsupported operand type(s) for /: 'NoneType' and 'float' DEBUG Calculated metric values diff --git a/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx940/0300_memory_chart.yaml b/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx940/0300_memory_chart.yaml index 23464a9d46..9f78e7dd45 100644 --- a/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx940/0300_memory_chart.yaml +++ b/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx940/0300_memory_chart.yaml @@ -70,9 +70,6 @@ Panel Config: + TCP_TCC_ATOMIC_WITH_RET_REQ_sum) + TCP_TCC_ATOMIC_WITHOUT_RET_REQ_sum)) / TCP_TOTAL_CACHE_ACCESSES_sum)) if (TCP_TOTAL_CACHE_ACCESSES_sum != 0) else None )), 0) - VL1 Lat: - value: ROUND(AVG(((TCP_TCP_LATENCY_sum / TCP_TA_TCP_STATE_READ_sum) if (TCP_TA_TCP_STATE_READ_sum - != 0) else None)), 0) VL1 Coalesce: value: ROUND(AVG(((((TA_TOTAL_WAVEFRONTS_sum * 64) * 100) / (TCP_TOTAL_ACCESSES_sum * 4)) if (TCP_TOTAL_ACCESSES_sum != None) else 0)), 0) @@ -197,8 +194,6 @@ Panel Config: unit after coalescing per normalization unit VL1 Hit: The ratio of the number of vL1D cache line requests that hit in vL1D cache over the total number of cache line requests to the vL1D Cache RAM. - VL1 Lat: Calculated as the average number of cycles that a vL1D cache line request - spent in the vL1D cache pipeline. VL1 Coalesce: Indicates how well memory instructions were coalesced by the address processing unit, ranging from uncoalesced (25%) to fully coalesced (100%). Calculated as the average number of thread-requests generated per instruction divided by diff --git a/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx941/0300_memory_chart.yaml b/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx941/0300_memory_chart.yaml index 23464a9d46..9f78e7dd45 100644 --- a/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx941/0300_memory_chart.yaml +++ b/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx941/0300_memory_chart.yaml @@ -70,9 +70,6 @@ Panel Config: + TCP_TCC_ATOMIC_WITH_RET_REQ_sum) + TCP_TCC_ATOMIC_WITHOUT_RET_REQ_sum)) / TCP_TOTAL_CACHE_ACCESSES_sum)) if (TCP_TOTAL_CACHE_ACCESSES_sum != 0) else None )), 0) - VL1 Lat: - value: ROUND(AVG(((TCP_TCP_LATENCY_sum / TCP_TA_TCP_STATE_READ_sum) if (TCP_TA_TCP_STATE_READ_sum - != 0) else None)), 0) VL1 Coalesce: value: ROUND(AVG(((((TA_TOTAL_WAVEFRONTS_sum * 64) * 100) / (TCP_TOTAL_ACCESSES_sum * 4)) if (TCP_TOTAL_ACCESSES_sum != None) else 0)), 0) @@ -197,8 +194,6 @@ Panel Config: unit after coalescing per normalization unit VL1 Hit: The ratio of the number of vL1D cache line requests that hit in vL1D cache over the total number of cache line requests to the vL1D Cache RAM. - VL1 Lat: Calculated as the average number of cycles that a vL1D cache line request - spent in the vL1D cache pipeline. VL1 Coalesce: Indicates how well memory instructions were coalesced by the address processing unit, ranging from uncoalesced (25%) to fully coalesced (100%). Calculated as the average number of thread-requests generated per instruction divided by diff --git a/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx942/0300_memory_chart.yaml b/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx942/0300_memory_chart.yaml index 23464a9d46..9f78e7dd45 100644 --- a/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx942/0300_memory_chart.yaml +++ b/projects/rocprofiler-compute/src/rocprof_compute_soc/analysis_configs/gfx942/0300_memory_chart.yaml @@ -70,9 +70,6 @@ Panel Config: + TCP_TCC_ATOMIC_WITH_RET_REQ_sum) + TCP_TCC_ATOMIC_WITHOUT_RET_REQ_sum)) / TCP_TOTAL_CACHE_ACCESSES_sum)) if (TCP_TOTAL_CACHE_ACCESSES_sum != 0) else None )), 0) - VL1 Lat: - value: ROUND(AVG(((TCP_TCP_LATENCY_sum / TCP_TA_TCP_STATE_READ_sum) if (TCP_TA_TCP_STATE_READ_sum - != 0) else None)), 0) VL1 Coalesce: value: ROUND(AVG(((((TA_TOTAL_WAVEFRONTS_sum * 64) * 100) / (TCP_TOTAL_ACCESSES_sum * 4)) if (TCP_TOTAL_ACCESSES_sum != None) else 0)), 0) @@ -197,8 +194,6 @@ Panel Config: unit after coalescing per normalization unit VL1 Hit: The ratio of the number of vL1D cache line requests that hit in vL1D cache over the total number of cache line requests to the vL1D Cache RAM. - VL1 Lat: Calculated as the average number of cycles that a vL1D cache line request - spent in the vL1D cache pipeline. VL1 Coalesce: Indicates how well memory instructions were coalesced by the address processing unit, ranging from uncoalesced (25%) to fully coalesced (100%). Calculated as the average number of thread-requests generated per instruction divided by diff --git a/projects/rocprofiler-compute/src/rocprof_compute_soc/soc_base.py b/projects/rocprofiler-compute/src/rocprof_compute_soc/soc_base.py index 1d985e49f1..e3b7ede821 100644 --- a/projects/rocprofiler-compute/src/rocprof_compute_soc/soc_base.py +++ b/projects/rocprofiler-compute/src/rocprof_compute_soc/soc_base.py @@ -344,10 +344,6 @@ class OmniSoC_Base: """Filter default performance counter set based on user arguments""" counters, filter_blocks = self.detect_counters() - # TCP_TCP_LATENCY_sum not supported for MI300 (gfx940, gfx941, gfx942) - if self.__arch in ("gfx940", "gfx941", "gfx942"): - counters = counters - {"TCP_TCP_LATENCY_sum"} - # SQ_ACCUM_PREV_HIRES will be injected for level counters later on counters = counters - {"SQ_ACCUM_PREV_HIRES"} diff --git a/projects/rocprofiler-compute/tools/config_management/.config_hashes.json b/projects/rocprofiler-compute/tools/config_management/.config_hashes.json index fe569bc7bf..0b37655aed 100644 --- a/projects/rocprofiler-compute/tools/config_management/.config_hashes.json +++ b/projects/rocprofiler-compute/tools/config_management/.config_hashes.json @@ -52,7 +52,7 @@ "0000_top_stats.yaml": "2819d96f5b1c3704f2ac50868a246a7f", "0100_system_info.yaml": "cefae2b10db8cf4b0d3a971cff5e82c8", "0200_system_speed_of_light.yaml": "d5df0a2b701972fa08ba0a44e72aa752", - "0300_memory_chart.yaml": "40dc04c73c3cec3d0a93e26d2db8c6f3", + "0300_memory_chart.yaml": "0a57cdf55be606799ee8d7b42a993027", "0400_roofline.yaml": "d4650e008f2e3a7d28871e8518153575", "0500_command_processor_cpc_cpf.yaml": "d8f424ec3fcfa4b2fcee2ad5e6456531", "0600_workgroup_manager_spi.yaml": "8b6a89de516bed5821a9849627ad634a", @@ -75,7 +75,7 @@ "0000_top_stats.yaml": "2819d96f5b1c3704f2ac50868a246a7f", "0100_system_info.yaml": "cefae2b10db8cf4b0d3a971cff5e82c8", "0200_system_speed_of_light.yaml": "0ccc1a63ebe11079832741c6d86ec3aa", - "0300_memory_chart.yaml": "40dc04c73c3cec3d0a93e26d2db8c6f3", + "0300_memory_chart.yaml": "0a57cdf55be606799ee8d7b42a993027", "0400_roofline.yaml": "c066a19bc0e00e692c34998e44c62387", "0500_command_processor_cpc_cpf.yaml": "d8f424ec3fcfa4b2fcee2ad5e6456531", "0600_workgroup_manager_spi.yaml": "8b6a89de516bed5821a9849627ad634a", @@ -98,7 +98,7 @@ "0000_top_stats.yaml": "2819d96f5b1c3704f2ac50868a246a7f", "0100_system_info.yaml": "cefae2b10db8cf4b0d3a971cff5e82c8", "0200_system_speed_of_light.yaml": "d5df0a2b701972fa08ba0a44e72aa752", - "0300_memory_chart.yaml": "40dc04c73c3cec3d0a93e26d2db8c6f3", + "0300_memory_chart.yaml": "0a57cdf55be606799ee8d7b42a993027", "0400_roofline.yaml": "318c3e774d41a639628a7f72c2462375", "0500_command_processor_cpc_cpf.yaml": "d8f424ec3fcfa4b2fcee2ad5e6456531", "0600_workgroup_manager_spi.yaml": "8b6a89de516bed5821a9849627ad634a", diff --git a/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx940_metrics_description.yaml b/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx940_metrics_description.yaml index f208484a66..c9079a32ad 100644 --- a/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx940_metrics_description.yaml +++ b/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx940_metrics_description.yaml @@ -1,69 +1,58 @@ +# AUTOGENERATED FILE. Only edit for testing purposes, not for development. Generated by tools/config_management/metric_description_manager.py System Speed-of-Light: VALU FLOPs: - rst: >- - The total floating-point operations executed per second on the :ref:`VALU - `. This is also presented as a percent of the peak theoretical - FLOPs achievable on the specific accelerator. Note: this does not include - any floating-point operations from :ref:`MFMA ` instructions. + rst: 'The total floating-point operations executed per second on the :ref:`VALU + `. This is also presented as a percent of the peak theoretical FLOPs + achievable on the specific accelerator. Note: this does not include any floating-point + operations from :ref:`MFMA ` instructions.' unit: GFLOPs VALU IOPs: - rst: >- - The total integer operations executed per second on the :ref:`VALU `. + rst: 'The total integer operations executed per second on the :ref:`VALU `. This is also presented as a percent of the peak theoretical IOPs achievable on the specific accelerator. Note: this does not include any integer operations - from :ref:`MFMA ` instructions. + from :ref:`MFMA ` instructions.' unit: GOIPs MFMA FLOPs (F8): - rst: >- - The total number of 8-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. This - is also presented as a percent of the peak theoretical F8 MFMA operations - achievable on the specific accelerator. It is supported on AMD Instinct MI300 - series and later only. + rst: 'The total number of 8-bit brain floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit brain floating point + operations from :ref:`VALU ` instructions. This is also presented + as a percent of the peak theoretical F8 MFMA operations achievable on the specific + accelerator. It is supported on AMD Instinct MI300 series and later only.' unit: GFLOPs MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` + rst: 'The total number of 16-bit brain floating point :ref:`MFMA ` operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. This - is also presented as a percent of the peak theoretical BF16 MFMA operations - achievable on the specific accelerator. + floating point operations from :ref:`VALU ` instructions. This is + also presented as a percent of the peak theoretical BF16 MFMA operations achievable + on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F16 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 16-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F16 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F32 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 32-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 32-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F32 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F64): - rst: >- - The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F64 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 64-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA IOPs (Int8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. This is also presented as a percent - of the peak theoretical INT8 MFMA operations achievable on the specific accelerator. + rst: 'The total number of 8-bit integer :ref:`MFMA ` operations executed + per second. Note: this does not include any 8-bit integer operations from :ref:`VALU + ` instructions. This is also presented as a percent of the peak theoretical + INT8 MFMA operations achievable on the specific accelerator.' unit: GIOPs - Active CUs: + Active CUs (deprecated): rst: Total number of active compute units (CUs) on the accelerator during the - kernel execution. + kernel execution. (Deprecated - See CU Utilization instead) unit: Number SALU Utilization: rst: Indicates what percent of the kernel's duration the :ref:`SALU ` @@ -108,11 +97,10 @@ System Speed-of-Light: over the :ref:`total active CU cycles `. unit: Instructions per-cycle Wavefront Occupancy: - rst: >- - The time-averaged number of wavefronts resident on the accelerator over + rst: 'The time-averaged number of wavefronts resident on the accelerator over the lifetime of the kernel. Note: this metric may be inaccurate for short-running kernels (less than 1ms). This is also presented as a percent of the peak theoretical - occupancy achievable on the specific accelerator. + occupancy achievable on the specific accelerator.' unit: Wavefronts Theoretical LDS Bandwidth: rst: Indicates the maximum amount of bytes that could have been loaded from, stored @@ -153,10 +141,9 @@ System Speed-of-Light: peak theoretical bandwidth achievable on the specific accelerator. unit: GB/s L2-Fabric Read BW: - rst: >- - The number of bytes read by the L2 over the :ref:`Infinity Fabric\u2122 - interface ` per unit time. This is also presented as a percent - of the peak theoretical bandwidth achievable on the specific accelerator. + rst: The number of bytes read by the L2 over the :ref:`Infinity Fabric\u2122 interface + ` per unit time. This is also presented as a percent of the peak + theoretical bandwidth achievable on the specific accelerator. unit: GB/s L2-Fabric Write BW: rst: The number of bytes sent by the L2 over the :ref:`Infinity Fabric interface @@ -194,359 +181,211 @@ System Speed-of-Light: L1I Fetch Latency: rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. unit: Cycles -Memory Chart: + CU Utilization: + rst: The percent of :ref:`total SIMD cycles ` in the kernel + where any :ref:`SIMD ` on a CU was actively doing any work, summed + over all CUs. Low values (less than 100%) indicate that the accelerator was + not fully saturated by the kernel, or a potential load-imbalance issue. + unit: Percent +General: Wavefront Occupancy: - rst: Wavefronts per active CU. - unit: Wavefronts - Wave Life: - rst: Average number of cycles executing a wave. - unit: Cycles per wave - SALU: - rst: Total Number of SALU (Scalar ALU) instructions issued per normalization unit. - unit: Instructions per normalization unit - SMEM: - rst: Total number of SMEM (Scalar Memory Read) instructions issued normalization - unit. - unit: Instructions per normalization unit - VALU: - rst: The number of VALU (Vector ALU) instructions issued per normalization unit. - unit: Instructions per normalization unit - MFMA: - rst: Total number of MFMA (Matrix-Fused-Multiply-Add) instructions issued per - normalization unit. - unit: Instructions per normalization unit - VMEM: - rst: The number of VMEM (GPU Memory) read instructions issued (including FLAT/scratch - memory) per normalization unit. - unit: Instructions per normalization unit - LDS: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's __shfl instructions) executed per normalization unit. - unit: Instructions per normalization unit - GWS: - rst: Total number of GDS (global data sync) instructions issued per normalization - unit. - unit: Instructions per normalization unit - BR: - rst: Total number of BRANCH instructions issued per normalization unit. - unit: Instructions per normalization unit - Active CUs: - rst: Total number of active compute units (CUs) on the accelerator during the - kernel execution. - unit: CUs - Num CUs: - rst: Total number of compute units (CUs) on the accelerator. - unit: CUs - VGPR: - rst: >- - The number of architected vector general-purpose registers allocated for the - kernel, see :ref:`VALU `. Note: this may not exactly match the - number of VGPRs requested by the compiler due to allocation granularity. - unit: VGPRs - SGPR: - rst: >- - The number of scalar general-purpose registers allocated for the kernel, see - :ref:`SALU `. Note: this may not exactly match the number of - SGPRs requested by the compiler due to allocation granularity. - unit: SGPRs - LDS Allocation: - rst: >- - The number of bytes of :doc:`LDS ` memory (or, shared memory) - allocated for this kernel. Note: This may also be larger than what was requested - at compile time due to both allocation granularity and dynamic per-dispatch - LDS allocations. - unit: Bytes per workgroup - Scratch Allocation: - rst: The number of bytes of :ref:`scratch memory ` requested per - work-item for this kernel. Scratch memory is used for stack memory on the accelerator, - as well as for register spills and restores. - unit: Bytes per workgroup - Wavefronts: - rst: The total number of wavefronts, summed over all workgroups, forming this - kernel launch. - unit: Wavefronts - Workgroups: - rst: The total number of workgroups forming this kernel launch. - unit: Workgroups - LDS Req: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. - unit: Instructions per normalization unit - LDS Util: - rst: Indicates what percent of the kernel's duration the :ref:`LDS ` - was actively executing instructions (including, but not limited to, load, store, - atomic and HIP's ``__shfl`` operations). Calculated as the ratio of the total - number of cycles LDS was active over the :ref:`total CU cycles `. - unit: Percent - LDS Latency: - rst: The average number of round-trip cycles (i.e., from issue to data-return - / acknowledgment) required for an LDS instruction to complete. - unit: Cycles - VL1 Rd: - rst: The total number of incoming read requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Wr: - rst: The total number of incoming write requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Atomic: - rst: The total number of incoming atomic requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Hit: - rst: The ratio of the number of vL1D cache line requests that hit in vL1D cache - over the total number of cache line requests to the :ref:`vL1D Cache RAM `. - unit: Percent - VL1 Lat: - rst: Calculated as the average number of cycles that a vL1D cache line request - spent in the vL1D cache pipeline. - unit: Cycles - VL1 Coalesce: - rst: Indicates how well memory instructions were coalesced by the :ref:`address - processing unit `, ranging from uncoalesced (25%) to fully coalesced - (100%). Calculated as the average number of :ref:`thread-requests ` - generated per instruction divided by the ideal number of thread-requests per - instruction. - unit: Percent - VL1 Stall: - rst: The ratio of the number of cycles where the vL1D is stalled waiting to issue - a request for data to the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - VL1_L2 Rd: - rst: The number of read requests for a vL1D cache line that were not satisfied - by the vL1D and must be retrieved from the to the :doc:`L2 Cache ` - per :ref:`normalization unit `. - unit: Requests per normalization unit - VL1_L2 Wr: - rst: The number of write requests to a vL1D cache line that were sent through - the vL1D to the :doc:`L2 cache `, per :ref:`normalization unit `. - unit: Requests per normalization unit - VL1_L2 Atomic: - rst: The number of atomic requests that are sent through the vL1D to the :doc:`L2 - cache `, per :ref:`normalization unit `. This - includes requests for atomics with, and without return. - unit: Requests per normalization unit - sL1D Rd: - rst: The total number of requests, of any size or type, made to the sL1D per :ref:`normalization - unit `. - unit: Requests per normalization unit - sL1D Hit: - rst: The total number of sL1D requests that hit on a previously loaded cache line, - per :ref:`normalization unit `. - unit: Requests per normalization unit - sL1D Lat: rst: '' - unit: Unknown + Wave Life: + rst: '' + SALU: + rst: '' + SMEM: + rst: '' + VALU: + rst: '' + MFMA: + rst: '' + VMEM: + rst: '' + LDS: + rst: '' + GWS: + rst: '' + BR: + rst: '' + Active CUs (deprecated): + rst: '' + Num CUs: + rst: '' + VGPR: + rst: '' + SGPR: + rst: '' + LDS Allocation: + rst: '' + Scratch Allocation: + rst: '' + Wavefronts: + rst: '' + Workgroups: + rst: '' + LDS Req: + rst: '' + LDS Util: + rst: '' + LDS Latency: + rst: '' + VL1 Rd: + rst: '' + VL1 Wr: + rst: '' + VL1 Atomic: + rst: '' + VL1 Hit: + rst: '' + VL1 Coalesce: + rst: '' + VL1 Stall: + rst: '' + VL1_L2 Rd: + rst: '' + VL1_L2 Wr: + rst: '' + VL1_L2 Atomic: + rst: '' + sL1D Rd: + rst: '' + sL1D Hit: + rst: '' sL1D_L2 Rd: - rst: The total number of read requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. - unit: Requests per normalization unit + rst: '' sL1D_L2 Wr: - rst: The total number of write requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit + rst: '' sL1D_L2 Atomic: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit + rst: '' IL1 Fetch: - rst: The total number of requests made to the L1I per :ref:`normalization-unit - `. - unit: Requests per normalization unit + rst: '' IL1 Hit: - rst: The total number of L1I requests that hit on a previously loaded cache line, - per :ref:`normalization-unit `. - unit: Percent + rst: '' IL1 Lat: - rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. - unit: Cycles + rst: '' IL1_L2 Rd: - rst: The total number of requests across the L1I - L2 interface per normalization-unit. - unit: Requests per normalization unit + rst: '' L2 Rd: - rst: The total number of read requests to the L2 from all clients. - unit: Requests per normalization unit + rst: '' L2 Wr: - rst: The total number of write requests to the L2 from all clients. - unit: Requests per normalization unit + rst: '' L2 Atomic: - rst: The total number of atomic requests (with and without return) to the L2 from - all clients. - unit: Requests per normalization unit + rst: '' L2 Hit: - rst: The ratio of the number of L2 cache line requests that hit in the L2 cache - over the total number of incoming cache line requests to the L2 cache. - unit: Percent + rst: '' Fabric_L2 Rd: - rst: Number of L2 cache - Infinity Fabric read requests (either 32-byte or 64-byte) - summed over TCC instances per normalization unit. - unit: Requests per normalization unit + rst: '' Fabric_L2 Wr: - rst: Number of L2 cache - Infinity Fabric write requests (either 32-byte or 64-byte) - summed over TCC instances per normalization unit. - unit: Requests per normalization unit + rst: '' Fabric_L2 Atomic: - rst: Number of L2 cache - Infinity Fabric write requests (either 32-byte or 64-byte) - that are actually atomic requests summed over TCC instances per normalization - unit. - unit: Requests per normalization unit + rst: '' Fabric Rd Lat: - rst: The time-averaged number of cycles read requests spent in Infinity Fabric - before data was returned to the L2. - unit: Cycles + rst: '' Fabric Wr Lat: - rst: The time-averaged number of cycles write requests spent in Infinity Fabric - before a completion acknowledgement was returned to the L2. - unit: Cycles + rst: '' Fabric Atomic Lat: - rst: The time-averaged number of cycles atomic requests spent in Infinity Fabric - before a completion acknowledgement (atomic without return value) or data (atomic - with return value) was returned to the L2. - unit: Cycles + rst: '' HBM Rd: - rst: The total number of L2 requests to Infinity Fabric to read 32B or 64B of - data from the accelerator's local HBM, per :ref:`normalization unit `. - See :ref:`l2-request-flow` for more detail. - unit: Requests per normalization unit + rst: '' HBM Wr: - rst: The total number of L2 requests to Infinity Fabric to write 32B or 64B of - data from the accelerator's local HBM, per :ref:`normalization unit `. - See :ref:`l2-request-flow` for more detail. - unit: Requests per normalization unit -Roofline Performance Rates: + rst: '' VALU FLOPs (F16): - rst: >- - The total 16-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F16 FLOPs achievable - on the specific accelerator. Note: this does not include any F16 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs + rst: '' VALU FLOPs (F32): - rst: >- - The total 32-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F32 FLOPs achievable - on the specific accelerator. Note: this does not include any F32 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs + rst: '' VALU FLOPs (F64): - rst: >- - The total 64-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F64 FLOPs achievable - on the specific accelerator. Note: this does not include any F64 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs - MFMA FLOPs (F64): - rst: >- - The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F64 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F32 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F16 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. The - peak empirically measured BF16 MFMA operations achievable on the specific - accelerator is displayed alongside for comparison. - unit: GFLOPs + rst: '' MFMA FLOPs (F8): - rst: >- - The total number of 8-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. The - peak empirically measured F8 MFMA operations achievable on the specific - accelerator is displayed alongside for comparison. It is supported on AMD - Instinct MI300 series and later only. - unit: GFLOPs + rst: '' + MFMA FLOPs (BF16): + rst: '' + MFMA FLOPs (F16): + rst: '' + MFMA FLOPs (F32): + rst: '' + MFMA FLOPs (F64): + rst: '' MFMA IOPs (Int8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. The peak empirically measured INT8 MFMA - operations achievable on the specific accelerator is displayed alongside - for comparison. - unit: GIOPs + rst: '' HBM Bandwidth: - rst: >- - The total number of bytes read from and written to High-Bandwidth - Memory (HBM) per second. The peak empirically measured bandwidth achievable - on the specific accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' L2 Cache Bandwidth: - rst: The number of bytes looked up in the L2 cache per unit time. The number of - bytes is calculated as the number of cache lines requested multiplied by the - cache line size. This value does not consider partial requests, so e.g., if - only a single value is requested in a cache line, the data movement will still - be counted as a full cache line. The peak empirically measured bandwidth achievable - on the specific accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' L1 Cache Bandwidth: - rst: The number of bytes looked up in the vL1D cache as a result of :ref:`VMEM - ` instructions per unit time. The number of bytes is calculated as - the number of cache lines requested multiplied by the cache line size. This - value does not consider partial requests, so e.g., if only a single value is - requested in a cache line, the data movement will still be counted as a full - cache line. The peak empirically measured bandwidth achievable on the specific - accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' LDS Bandwidth: - rst: Indicates the maximum amount of bytes that could have been loaded from, stored - to, or atomically updated in the LDS per unit time (see :ref:`LDS Bandwidth - ` example for more detail). The peak empirically measured LDS - bandwidth achievable on the specific accelerator is displayed alongside for - comparison. - unit: GB/s -Roofline Plot Points: - AI HBM: - rst: >- - The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM). - It is the ratio of total floating-point operations (FLOPs) to total bytes - transferred between HBM and the L2 cache. This value is used as the x-coordinate - for the HBM roofline. - unit: FLOPs/Byte - AI L2: - rst: >- - The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio - of total floating-point operations (FLOPs) to total bytes transferred between - the L2 cache and the L1 cache. This value is used as the x-coordinate for - the L2 roofline. - unit: FLOPs/Byte + rst: '' AI L1: - rst: >- - The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio - of total floating-point operations (FLOPs) to total bytes transferred between - the L1 cache and the processing units. This value is used as the x-coordinate - for the L1 roofline. - unit: FLOPs/Byte + rst: '' + AI L2: + rst: '' + AI HBM: + rst: '' Performance (GFLOPs): - rst: >- - The overall achieved performance, measured in GigaFLOPs - per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point - operations divided by the total execution time. This value is used as the y-coordinate - for the kernel's point on the Roofline plot. - unit: GFLOP/s + rst: '' + Global/Generic Instr: + rst: '' + Global/Generic Read: + rst: '' + Global/Generic Write: + rst: '' + Global/Generic Atomic: + rst: '' + Spill/Stack Instr: + rst: '' + Spill/Stack Read: + rst: '' + Spill/Stack Write: + rst: '' + Spill/Stack Atomic: + rst: '' + Stalled on L2 Data: + rst: '' + Stalled on L2 Req: + rst: '' + Tag RAM Stall (Read): + rst: '' + Tag RAM Stall (Write): + rst: '' + Tag RAM Stall (Atomic): + rst: '' + NC - Read: + rst: '' + UC - Read: + rst: '' + CC - Read: + rst: '' + RW - Read: + rst: '' + RW - Write: + rst: '' + NC - Write: + rst: '' + UC - Write: + rst: '' + CC - Write: + rst: '' + NC - Atomic: + rst: '' + UC - Atomic: + rst: '' + CC - Atomic: + rst: '' + RW - Atomic: + rst: '' + Req: + rst: '' + Hit Ratio: + rst: '' + Hits: + rst: '' + Translation Misses: + rst: '' + Permission Misses: + rst: '' + L2 Cache Hit Rate: + rst: '' Command processor fetcher (CPF): CPF Utilization: rst: Percent of total cycles where the CPF was busy actively doing any work. The @@ -599,10 +438,9 @@ Workgroup manager utilizations: any work. unit: Percent Scheduler-Pipe Utilization: - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where the scheduler-pipes were actively doing any work. Note: this - value is expected to range between 0% and 25%. See :ref:`desc-spi`. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where the scheduler-pipes were actively doing any work. Note: this + value is expected to range between 0% and 25%. See :ref:`desc-spi`.' unit: Percent Workgroup Manager Utilization: rst: The percent of cycles in the kernel where the workgroup manager was actively @@ -637,30 +475,25 @@ Workgroup manager utilizations: unit: Cycles/wave Workgroup Manager - Resource Allocation: Not-scheduled Rate (Workgroup Manager): - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` - due to a bottleneck within the workgroup manager rather than a lack of a - CU or :ref:`SIMD ` with sufficient resources. Note: this value - is expected to range between 0-25%. See note in :ref:`workgroup manager ` - description. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` + due to a bottleneck within the workgroup manager rather than a lack of a CU + or :ref:`SIMD ` with sufficient resources. Note: this value is expected + to range between 0-25%. See note in :ref:`workgroup manager ` description.' unit: Percent Not-scheduled Rate (Scheduler-Pipe): - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` - due to a bottleneck within the scheduler-pipes rather than a lack of a CU - or :ref:`SIMD ` with sufficient resources. Note: this value is - expected to range between 0-25%, see note in :ref:`workgroup manager ` - description. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` + due to a bottleneck within the scheduler-pipes rather than a lack of a CU or + :ref:`SIMD ` with sufficient resources. Note: this value is expected + to range between 0-25%, see note in :ref:`workgroup manager ` description.' unit: Percent Scheduler-Pipe Stall Rate: - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` due to occupancy limitations (like a lack of a CU or :ref:`SIMD ` - with sufficient resources). Note: this value is expected to range between - 0-25%, see note in :ref:`workgroup manager ` description. + with sufficient resources). Note: this value is expected to range between 0-25%, + see note in :ref:`workgroup manager ` description.' unit: Percent Scratch Stall Rate: rst: The percent of :ref:`total shader-engine cycles ` in the @@ -707,7 +540,7 @@ Workgroup Manager - Resource Allocation: within the workgroup manager. This is expected to be always be zero on CDNA2 or newer accelerators (and small for previous accelerators). unit: Percent -Wavefront Launch Stats: +Wavefront launch stats: Grid Size: rst: The total number of work-items (or, threads) launched as a part of the kernel dispatch. In HIP, this is equivalent to the total grid size multiplied by the @@ -719,11 +552,10 @@ Wavefront Launch Stats: block size. unit: Work-Items Total Wavefronts: - rst: >- - The total number of wavefronts launched as part of the kernel dispatch. - On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront - size is always 64 work-items. Thus, the total number of wavefronts should - be equivalent to the ceiling of grid size divided by 64. + rst: The total number of wavefronts launched as part of the kernel dispatch. On + AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront + size is always 64 work-items. Thus, the total number of wavefronts should be + equivalent to the ceiling of grid size divided by 64. unit: Wavefronts Saved Wavefronts: rst: The total number of wavefronts saved at a context-save. See `cwsr_enable @@ -734,36 +566,32 @@ Wavefront Launch Stats: `_. unit: Wavefronts VGPRs: - rst: >- - The number of architected vector general-purpose registers allocated for the - kernel, see :ref:`VALU `. Note: this may not exactly match the - number of VGPRs requested by the compiler due to allocation granularity. + rst: 'The number of architected vector general-purpose registers allocated for + the kernel, see :ref:`VALU `. Note: this may not exactly match the + number of VGPRs requested by the compiler due to allocation granularity.' unit: VGPRs AGPRs: - rst: >- - The number of accumulation vector general-purpose registers allocated - for the kernel, see :ref:`AGPRs `. Note: this may not exactly match - the number of AGPRs requested by the compiler due to allocation granularity. + rst: 'The number of accumulation vector general-purpose registers allocated for + the kernel, see :ref:`AGPRs `. Note: this may not exactly match + the number of AGPRs requested by the compiler due to allocation granularity.' unit: AGPRs SGPRs: - rst: >- - The number of scalar general-purpose registers allocated for the kernel, see - :ref:`SALU `. Note: this may not exactly match the number of - SGPRs requested by the compiler due to allocation granularity. + rst: 'The number of scalar general-purpose registers allocated for the kernel, + see :ref:`SALU `. Note: this may not exactly match the number of + SGPRs requested by the compiler due to allocation granularity.' unit: SGPRs LDS Allocation: - rst: >- - The number of bytes of :doc:`LDS ` memory (or, shared memory) - allocated for this kernel. Note: This may also be larger than what was requested - at compile time due to both allocation granularity and dynamic per-dispatch - LDS allocations. + rst: 'The number of bytes of :doc:`LDS ` memory (or, shared + memory) allocated for this kernel. Note: This may also be larger than what was + requested at compile time due to both allocation granularity and dynamic per-dispatch + LDS allocations.' unit: Bytes per workgroup Scratch Allocation: rst: The number of bytes of :ref:`scratch memory ` requested per work-item for this kernel. Scratch memory is used for stack memory on the accelerator, as well as for register spills and restores. unit: Bytes per work-item -Wavefront Runtime Stats: +Wavefront runtime stats: Kernel Time: rst: The total duration of the executed kernel. unit: Nanoseconds @@ -775,11 +603,10 @@ Wavefront Runtime Stats: This is averaged over all wavefronts in a kernel dispatch. unit: Instructions per wavefront Wave Cycles: - rst: >- - The number of cycles a wavefront in the kernel dispatch spent resident - on a compute unit per :ref:`normalization unit `. This is - averaged over all wavefronts in a kernel dispatch. Note: this should not - be directly compared to the kernel cycles above. + rst: 'The number of cycles a wavefront in the kernel dispatch spent resident on + a compute unit per :ref:`normalization unit `. This is + averaged over all wavefronts in a kernel dispatch. Note: this should not be + directly compared to the kernel cycles above.' unit: Cycles per normalization unit Dependency Wait Cycles: rst: The number of cycles a wavefront in the kernel dispatch stalled waiting on @@ -813,12 +640,11 @@ Wavefront Runtime Stats: the total Wave Cycles metric. unit: Cycles per normalization unit Wavefront Occupancy: - rst: >- - The time-averaged number of wavefronts resident on the accelerator over the - lifetime of the kernel. Note: this metric may be inaccurate for short-running - kernels (less than 1ms). + rst: 'The time-averaged number of wavefronts resident on the accelerator over + the lifetime of the kernel. Note: this metric may be inaccurate for short-running + kernels (less than 1ms).' unit: Wavefronts -Overall Instruction Mix: +Overall instruction mix: VALU: rst: The total number of vector arithmetic logic unit (VALU) operations issued. These are the workhorses of the :doc:`compute unit `, and are @@ -854,7 +680,7 @@ Overall Instruction Mix: rst: The total number of branch operations issued. These typically consist of jump or branch operations and are used to implement control flow. unit: Instructions -VALU Arithmetic Instruction Mix: +VALU arithmetic instruction mix: INT32: rst: The total number of instructions operating on 32-bit integer operands issued to the VALU per :ref:`normalization unit `. @@ -915,53 +741,10 @@ VALU Arithmetic Instruction Mix: unit `. unit: Instructions per normalization unit Conversion: - rst: >- - The total number of type conversion instructions (such as converting data - to or from F32\u2194F64) issued to the VALU per :ref:`normalization unit - `. + rst: The total number of type conversion instructions (such as converting data + to or from F32\u2194F64) issued to the VALU per :ref:`normalization unit `. unit: Instructions per normalization unit -VMEM Instruction Mix: - Global/Generic Instr: - rst: The total number of global & generic memory instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Read: - rst: The total number of global & generic memory read instructions executed on - all :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Write: - rst: The total number of global & generic memory write instructions executed on - all :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Atomic: - rst: The total number of global & generic memory atomic (with and without return) - instructions executed on all :doc:`compute units ` on the accelerator, - per :ref:`normalization unit `. - unit: Instructions per normalization unit - Spill/Stack Instr: - rst: The total number of spill/stack memory instructions executed on all :doc:`compute - units ` on the accelerator, per :ref:`normalization unit `. - unit: Instructions per normalization unit - Spill/Stack Read: - rst: The total number of spill/stack memory read instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Spill/Stack Write: - rst: The total number of spill/stack memory write instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Spill/Stack Atomic: - rst: The total number of spill/stack memory atomic (with and without return) instructions - executed on all :doc:`compute units ` on the accelerator, per - :ref:`normalization unit `. Typically unused as these memory - operations are typically used to implement thread-local storage. - unit: Instructions per normalization unit -MFMA Arithmetic Instruction Mix: +MFMA instruction mix: MFMA-I8: rst: The total number of 8-bit integer :ref:`MFMA ` instructions issued per :ref:`normalization unit `. @@ -989,66 +772,53 @@ MFMA Arithmetic Instruction Mix: unit: Instructions per normalization unit Compute Speed-of-Light: VALU FLOPs: - rst: >- - The total floating-point operations executed per second on the :ref:`VALU - `. This is also presented as a percent of the peak theoretical - FLOPs achievable on the specific accelerator. Note: this does not include - any floating-point operations from :ref:`MFMA ` instructions. + rst: 'The total floating-point operations executed per second on the :ref:`VALU + `. This is also presented as a percent of the peak theoretical FLOPs + achievable on the specific accelerator. Note: this does not include any floating-point + operations from :ref:`MFMA ` instructions.' unit: GFLOPs VALU IOPs: - rst: >- - The total integer operations executed per second on the :ref:`VALU `. + rst: 'The total integer operations executed per second on the :ref:`VALU `. This is also presented as a percent of the peak theoretical IOPs achievable on the specific accelerator. Note: this does not include any integer operations - from :ref:`MFMA ` instructions. + from :ref:`MFMA ` instructions.' unit: GIOPs - MFMA FLOPs (F8): - rst: '' - unit: Unknown MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit brain floating - point operations from :ref:`VALU ` instructions. This is also - presented as a percent of the peak theoretical BF16 MFMA operations achievable - on the specific accelerator. + rst: 'The total number of 16-bit brain floating point :ref:`MFMA ` + operations executed per second. Note: this does not include any 16-bit brain + floating point operations from :ref:`VALU ` instructions. This is + also presented as a percent of the peak theoretical BF16 MFMA operations achievable + on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F16 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 16-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F16 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F32 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 32-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 32-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F32 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F64): - rst: >- + rst: 'The total number of 64-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator. The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F64 MFMA operations achievable on the - specific accelerator. The total number of 64-bit floating point :ref:`MFMA - ` operations executed per second. Note: this does not include - any 64-bit floating point operations from :ref:`VALU ` instructions. - This is also presented as a percent of the peak theoretical F64 MFMA operations - achievable on the specific accelerator. + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA IOPs (INT8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. This is also presented as a percent - of the peak theoretical INT8 MFMA operations achievable on the specific accelerator. + rst: 'The total number of 8-bit integer :ref:`MFMA ` operations executed + per second. Note: this does not include any 8-bit integer operations from :ref:`VALU + ` instructions. This is also presented as a percent of the peak theoretical + INT8 MFMA operations achievable on the specific accelerator.' unit: GFLOPs -Pipeline Statistics: +Pipeline statistics: IPC: rst: The ratio of the total number of instructions executed on the :doc:`CU ` over the :ref:`total active CU cycles `. @@ -1111,7 +881,7 @@ Pipeline Statistics: rst: The average number of round-trip cycles (that is, from issue to data return / acknowledgment) required for a SMEM instruction to complete. unit: Cycles -Arithmetic Operations: +Arithmetic operations: FLOPs (Total): rst: The total number of floating-point operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit @@ -1122,20 +892,16 @@ Arithmetic Operations: ` or :ref:`MFMA ` units, per :ref:`normalization unit `. unit: IOP per normalization unit - F8 OPs: - rst: '' - unit: Unknown F16 OPs: rst: The total number of 16-bit floating-point operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit `. unit: FLOP per normalization unit BF16 OPs: - rst: >- - The total number of 16-bit brain floating-point operations executed on - either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization - unit `. Note: on current CDNA accelerators, the VALU - has no native BF16 instructions. + rst: 'The total number of 16-bit brain floating-point operations executed on either + the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization + unit `. Note: on current CDNA accelerators, the VALU has + no native BF16 instructions.' unit: FLOP per normalization unit F32 OPs: rst: The total number of 32-bit floating-point operations executed on either the @@ -1148,11 +914,10 @@ Arithmetic Operations: unit `. unit: FLOP per normalization unit INT8 OPs: - rst: >- - The total number of 8-bit integer operations executed on either the :ref:`VALU + rst: 'The total number of 8-bit integer operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit - `. Note: on current CDNA accelerators, the VALU has - no native INT8 instructions. + `. Note: on current CDNA accelerators, the VALU has no + native INT8 instructions.' unit: IOP per normalization unit LDS Speed-of-Light: Utilization: @@ -1182,16 +947,16 @@ LDS Speed-of-Light: amount of data in an uncontended access. [#lds-bank-conflict]_ unit: Percent LDS Statistics: - LDS Instructions: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. - unit: Instructions per normalization unit Theoretical Bandwidth: rst: Indicates the maximum amount of bytes that could have been loaded from, stored to, or atomically updated in the LDS divided by total duration. Does *not* take into account the execution mask of the wavefront when the instruction was executed. See the :ref:`LDS bandwidth example ` for more detail. unit: Gbps + LDS Instructions: + rst: The total number of LDS instructions (including, but not limited to, read/write/atomics + and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. + unit: Instructions per normalization unit LDS Latency: rst: The average number of round-trip cycles (i.e., from issue to data-return acknowledgment) required for an LDS instruction to complete. @@ -1225,10 +990,9 @@ LDS Statistics: to stalls from non-dword aligned addresses per :ref:`normalization unit `. unit: Cycles per normalization unit Mem Violations: - rst: >- - The total number of out-of-bounds accesses made to the LDS, per :ref:`normalization - unit `. This is unused and expected to be zero in - most configurations for modern CDNA\u2122 accelerators. + rst: The total number of out-of-bounds accesses made to the LDS, per :ref:`normalization + unit `. This is unused and expected to be zero in most + configurations for modern CDNA\u2122 accelerators. unit: Accesses per normalization unit L1I Speed-of-Light: Bandwidth Utilization: @@ -1236,18 +1000,17 @@ L1I Speed-of-Light: theoretical bandwidth. Calculated as the ratio of L1I requests over the :ref:`total L1I cycles `. unit: Percent + L1I-L2 Bandwidth Utilization: + rst: The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth + achieved. Calculated as the ratio of the total number of requests from the L1I + to the L2 cache over the :ref:`total L1I-L2 interface cycles `. + unit: Percent +L1I cache accesses: Cache Hit Rate: rst: The percent of L1I requests that hit [#l1i-cache]_ on a previously loaded line the cache. Calculated as the ratio of the number of L1I requests that hit over the number of all L1I requests. unit: Percent - L1I-L2 Bandwidth Utilization: - rst: >- - The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth - achieved. Calculated as the ratio of the total number of requests from - the L1I to the L2 cache over the :ref:`total L1I-L2 interface cycles `. - unit: Percent -L1I cache accesses: Req: rst: The total number of requests made to the L1I per normalization-unit unit: Requests per normalization unit @@ -1265,11 +1028,6 @@ L1I cache accesses: already pending due to another request, per :ref:`normalization-unit `. See note in :ref:`desc-l1i-sol` for more detail. unit: Requests per normalization unit - Cache Hit Rate: - rst: The percent of L1I requests that hit [#l1i-cache]_ on a previously loaded - line the cache. Calculated as the ratio of the number of L1I requests that hit - over the number of all L1I requests. - unit: Percent Instruction Fetch Latency: rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. unit: Cycles @@ -1284,17 +1042,17 @@ Scalar L1D Speed-of-Light: theoretical bandwidth. Calculated as the ratio of sL1D requests over the :ref:`total sL1D cycles `. unit: Percent + sL1D-L2 BW Utilization: + rst: The percentage of the peak theoretical sL1D - L2 interface bandwidth acheived. + Calculated as total number of bytes read from, written to, or atomically updated + across the sL1D - L2 interface. + unit: Percent +Scalar L1D cache accesses: Cache Hit Rate: rst: Indicates the percent of sL1D requests that hit on a previously loaded line the cache. The ratio of the number of sL1D requests that hit [#sl1d-cache]_ over the number of all sL1D requests. unit: Percent - sL1D-L2 BW Utilization: - rst: The percentage of the peak theoretical sL1D - L2 interface bandwidth acheived. - Caclulated as total number of bytes read from, written to, or atomically updated - across the sL1D - L2 interface. - unit: Percent -Scalar L1D cache accesses: Req: rst: The total number of requests, of any size or type, made to the sL1D per :ref:`normalization unit `. @@ -1313,20 +1071,10 @@ Scalar L1D cache accesses: already pending due to another request, per :ref:`normalization unit `. See :ref:`desc-sl1d-sol` for more detail. unit: Requests per normalization unit - Cache Hit Rate: - rst: Indicates the percent of sL1D requests that hit on a previously loaded line - the cache. The ratio of the number of sL1D requests that hit [#sl1d-cache]_ - over the number of all sL1D requests. - unit: Percent Read Req (Total): rst: The total number of sL1D read requests of any size, per :ref:`normalization unit `. unit: Requests per normalization unit - Atomic Req: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit Read Req (1 DWord): rst: The total number of sL1D read requests made for a single dword of data (4B), per :ref:`normalization unit `. @@ -1349,13 +1097,17 @@ Scalar L1D cache accesses: unit: Requests per normalization unit Scalar L1D Cache - L2 Interface: sL1D-L2 BW: - rst: >- - The total number of bytes read from, written to, or atomically updated - across the sL1D\u2194:doc:`L2 ` interface, divided by total duration. - Note that sL1D writes and atomics are typically - unused on current CDNA accelerators, so in the majority of cases this can - be interpreted as an sL1D\u2192L2 read bandwidth. + rst: The total number of bytes read from, written to, or atomically updated across + the sL1D\u2194:doc:`L2 ` interface, divided by total duration. Note + that sL1D writes and atomics are typically unused on current CDNA accelerators, + so in the majority of cases this can be interpreted as an sL1D\u2192L2 read + bandwidth. unit: Gbps + Atomic Req: + rst: The total number of atomic requests from sL1D to the :doc:`L2 `, + per :ref:`normalization unit `. Typically unused on current + CDNA accelerators. + unit: Requests per normalization unit Read Req: rst: The total number of read requests from sL1D to the :doc:`L2 `, per :ref:`normalization unit `. @@ -1365,17 +1117,11 @@ Scalar L1D Cache - L2 Interface: per :ref:`normalization unit `. Typically unused on current CDNA accelerators. unit: Requests per normalization unit - Atomic Req: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit Stall Cycles: - rst: >- - The total number of cycles the sL1D\u2194 :doc:`L2 ` interface + rst: The total number of cycles the sL1D\u2194 :doc:`L2 ` interface was stalled, per :ref:`normalization unit `. unit: Cycles per normalization unit -Busy and stall metrics: +Busy / stall metrics: Address Processing Unit Busy: rst: Percent of the :ref:`total CU cycles ` the address processor was busy @@ -1388,19 +1134,10 @@ Busy and stall metrics: rst: Percent of the :ref:`total CU cycles ` the address processor was stalled from sending write/atomic data further into the vL1D pipeline unit: Percent - Data-Processor → Address Stall: + "Data-Processor \u2192 Address Stall": rst: Percent of :ref:`total CU cycles ` the address processor was stalled waiting to send command data to the :ref:`data processor ` unit: Percent - Sequencer → TA Address Stall: - rst: '' - unit: Unknown - Sequencer → TA Command Stall: - rst: '' - unit: Unknown - Sequencer → TA Data Stall: - rst: '' - unit: Unknown Instruction counts: Total Instructions: rst: The total number of memory instructions executed by the address processer @@ -1446,7 +1183,7 @@ Instruction counts: :ref:`normalization unit `. Typically unused as these memory operations are typically used to implement thread-local storage. unit: Instructions per normalization unit -Spill and stack metrics: +Spill / stack metrics: Spill/Stack Total Cycles: rst: The number of cycles the address processing unit spent working on spill/stack instructions, per :ref:`normalization unit `. @@ -1464,11 +1201,11 @@ Vector L1 data-return path or Texture Data (TD): rst: Percent of the :ref:`total CU cycles ` the data-return unit was busy processing or waiting on data to return to the :doc:`CU `. unit: Percent - Cache RAM → Data-Return Stall: + "Cache RAM \u2192 Data-Return Stall": rst: Percent of the :ref:`total CU cycles ` the data-return unit was stalled on data to be returned from the :ref:`vL1D Cache RAM `. unit: Percent - Workgroup manager → Data-Return Stall: + "Workgroup manager \u2192 Data-Return Stall": rst: Percent of the :ref:`total CU cycles ` the data-return unit was stalled by the :ref:`workgroup manager ` due to initialization of registers as a part of launching new workgroups. @@ -1525,32 +1262,6 @@ vL1D Speed-of-Light: generated per instruction divided by the ideal number of thread-requests per instruction. unit: Percent -vL1D cache stall metrics: - Stalled on L2 Data: - rst: The ratio of the number of cycles where the vL1D is stalled waiting for requested - data to return from the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Stalled on L2 Req: - rst: The ratio of the number of cycles where the vL1D is stalled waiting to issue - a request for data to the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Read): - rst: The ratio of the number of cycles where the vL1D is stalled due to Read requests - with conflicting tags being looked up concurrently, divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Write): - rst: The ratio of the number of cycles where the vL1D is stalled due to Write - requests with conflicting tags being looked up concurrently, divided by the - number of cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Atomic): - rst: The ratio of the number of cycles where the vL1D is stalled due to Atomic - requests with conflicting tags being looked up concurrently, divided by the - number of cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent vL1D cache access metrics: Total Req: rst: The total number of incoming requests from the :ref:`address processing unit @@ -1615,79 +1326,6 @@ vL1D cache access metrics: cache `, per :ref:`normalization unit `. This includes requests for atomics with, and without return. unit: Requests per normalization unit -L1D - L2 Transactions: - NC - Read: - rst: Total read requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Read: - rst: Total read requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Read: - rst: Total read requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Read: - rst: Total read requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Write: - rst: Total write requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - NC - Write: - rst: Total write requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Write: - rst: Total write requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Write: - rst: Total write requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - NC - Atomic: - rst: Total atomic requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Atomic: - rst: Total atomic requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Atomic: - rst: Total atomic requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Atomic: - rst: Total atomic requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit -L1 Unified Translation Cache (UTCL1): - Req: - rst: The number of translation requests made to the UTCL1 per normalization unit. - unit: Requests per normalization unit - Hit Ratio: - rst: The ratio of the number of translation requests that hit in the UTCL1 divided - by the total number of translation requests made to the UTCL1. - unit: Percent - Hits: - rst: The number of translation requests that hit in the UTCL1, and could be reused, - per normalization unit. - unit: Requests per normalization unit - Translation Misses: - rst: The total number of translation requests that missed in the UTCL1 due to - translation not being present in the cache, per :ref:`normalization unit `. - unit: unit - Permission Misses: - rst: >- - The total number of translation requests that missed in the UTCL1 due - to a permission error, per :ref:`normalization unit `. - This is unused and expected to be zero in most configurations for modern - CDNA\u2122 accelerators. - unit: Requests per normalization unit -L1D Addr Translation Stalls: {} L2 Speed-of-Light: Utilization: rst: The ratio of the :ref:`number of cycles an L2 channel was active, summed @@ -1809,7 +1447,7 @@ L2-Fabric interface metrics: before a completion acknowledgement (atomic without return value) or data (atomic with return value) was returned to the L2. unit: Cycles -L2 Cache Accesses: +L2 cache accesses: Bandwidth: rst: The number of bytes looked up in the L2 cache, divided by total duration. The number of bytes is calculated as the number of cache lines requested multiplied @@ -1900,7 +1538,6 @@ L2 Cache Accesses: rst: The total number of requests to the L2 that go to Read-Write coherent memory (RW) allocations. See the :ref:`memory-type` for more information. unit: Requests per normalization unit -L2 Cache Stalls: {} L2 - Fabric Interface stalls: Write - Credit Starvation: rst: The number of cycles the L2-Fabric interface was stalled on write or atomic @@ -1969,72 +1606,3 @@ L2 - Fabric interface detailed metrics: as :ref:`fine-grained memory ` allocations or :ref:`uncached memory ` allocations on the MI2XX. unit: Requests per normalization unit -Aggregate Stats (All channels): - L2 Cache Hit Rate: - rst: The total number of requests to the L2 from all clients that hit in the cache. - As noted in the :ref:`Speed-of-Light ` section, this includes hit-on-miss - requests. - unit: Percent -L2 Cache Hit Rate (pct): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2 Requests (per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Requests (per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Read Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Write and Atomic Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Atomic Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Read Stall (Cycles per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Write and Atomic Stall (Cycles per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric (128B read requests per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown diff --git a/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx941_metrics_description.yaml b/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx941_metrics_description.yaml index f208484a66..c9079a32ad 100644 --- a/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx941_metrics_description.yaml +++ b/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx941_metrics_description.yaml @@ -1,69 +1,58 @@ +# AUTOGENERATED FILE. Only edit for testing purposes, not for development. Generated by tools/config_management/metric_description_manager.py System Speed-of-Light: VALU FLOPs: - rst: >- - The total floating-point operations executed per second on the :ref:`VALU - `. This is also presented as a percent of the peak theoretical - FLOPs achievable on the specific accelerator. Note: this does not include - any floating-point operations from :ref:`MFMA ` instructions. + rst: 'The total floating-point operations executed per second on the :ref:`VALU + `. This is also presented as a percent of the peak theoretical FLOPs + achievable on the specific accelerator. Note: this does not include any floating-point + operations from :ref:`MFMA ` instructions.' unit: GFLOPs VALU IOPs: - rst: >- - The total integer operations executed per second on the :ref:`VALU `. + rst: 'The total integer operations executed per second on the :ref:`VALU `. This is also presented as a percent of the peak theoretical IOPs achievable on the specific accelerator. Note: this does not include any integer operations - from :ref:`MFMA ` instructions. + from :ref:`MFMA ` instructions.' unit: GOIPs MFMA FLOPs (F8): - rst: >- - The total number of 8-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. This - is also presented as a percent of the peak theoretical F8 MFMA operations - achievable on the specific accelerator. It is supported on AMD Instinct MI300 - series and later only. + rst: 'The total number of 8-bit brain floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit brain floating point + operations from :ref:`VALU ` instructions. This is also presented + as a percent of the peak theoretical F8 MFMA operations achievable on the specific + accelerator. It is supported on AMD Instinct MI300 series and later only.' unit: GFLOPs MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` + rst: 'The total number of 16-bit brain floating point :ref:`MFMA ` operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. This - is also presented as a percent of the peak theoretical BF16 MFMA operations - achievable on the specific accelerator. + floating point operations from :ref:`VALU ` instructions. This is + also presented as a percent of the peak theoretical BF16 MFMA operations achievable + on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F16 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 16-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F16 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F32 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 32-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 32-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F32 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F64): - rst: >- - The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F64 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 64-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA IOPs (Int8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. This is also presented as a percent - of the peak theoretical INT8 MFMA operations achievable on the specific accelerator. + rst: 'The total number of 8-bit integer :ref:`MFMA ` operations executed + per second. Note: this does not include any 8-bit integer operations from :ref:`VALU + ` instructions. This is also presented as a percent of the peak theoretical + INT8 MFMA operations achievable on the specific accelerator.' unit: GIOPs - Active CUs: + Active CUs (deprecated): rst: Total number of active compute units (CUs) on the accelerator during the - kernel execution. + kernel execution. (Deprecated - See CU Utilization instead) unit: Number SALU Utilization: rst: Indicates what percent of the kernel's duration the :ref:`SALU ` @@ -108,11 +97,10 @@ System Speed-of-Light: over the :ref:`total active CU cycles `. unit: Instructions per-cycle Wavefront Occupancy: - rst: >- - The time-averaged number of wavefronts resident on the accelerator over + rst: 'The time-averaged number of wavefronts resident on the accelerator over the lifetime of the kernel. Note: this metric may be inaccurate for short-running kernels (less than 1ms). This is also presented as a percent of the peak theoretical - occupancy achievable on the specific accelerator. + occupancy achievable on the specific accelerator.' unit: Wavefronts Theoretical LDS Bandwidth: rst: Indicates the maximum amount of bytes that could have been loaded from, stored @@ -153,10 +141,9 @@ System Speed-of-Light: peak theoretical bandwidth achievable on the specific accelerator. unit: GB/s L2-Fabric Read BW: - rst: >- - The number of bytes read by the L2 over the :ref:`Infinity Fabric\u2122 - interface ` per unit time. This is also presented as a percent - of the peak theoretical bandwidth achievable on the specific accelerator. + rst: The number of bytes read by the L2 over the :ref:`Infinity Fabric\u2122 interface + ` per unit time. This is also presented as a percent of the peak + theoretical bandwidth achievable on the specific accelerator. unit: GB/s L2-Fabric Write BW: rst: The number of bytes sent by the L2 over the :ref:`Infinity Fabric interface @@ -194,359 +181,211 @@ System Speed-of-Light: L1I Fetch Latency: rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. unit: Cycles -Memory Chart: + CU Utilization: + rst: The percent of :ref:`total SIMD cycles ` in the kernel + where any :ref:`SIMD ` on a CU was actively doing any work, summed + over all CUs. Low values (less than 100%) indicate that the accelerator was + not fully saturated by the kernel, or a potential load-imbalance issue. + unit: Percent +General: Wavefront Occupancy: - rst: Wavefronts per active CU. - unit: Wavefronts - Wave Life: - rst: Average number of cycles executing a wave. - unit: Cycles per wave - SALU: - rst: Total Number of SALU (Scalar ALU) instructions issued per normalization unit. - unit: Instructions per normalization unit - SMEM: - rst: Total number of SMEM (Scalar Memory Read) instructions issued normalization - unit. - unit: Instructions per normalization unit - VALU: - rst: The number of VALU (Vector ALU) instructions issued per normalization unit. - unit: Instructions per normalization unit - MFMA: - rst: Total number of MFMA (Matrix-Fused-Multiply-Add) instructions issued per - normalization unit. - unit: Instructions per normalization unit - VMEM: - rst: The number of VMEM (GPU Memory) read instructions issued (including FLAT/scratch - memory) per normalization unit. - unit: Instructions per normalization unit - LDS: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's __shfl instructions) executed per normalization unit. - unit: Instructions per normalization unit - GWS: - rst: Total number of GDS (global data sync) instructions issued per normalization - unit. - unit: Instructions per normalization unit - BR: - rst: Total number of BRANCH instructions issued per normalization unit. - unit: Instructions per normalization unit - Active CUs: - rst: Total number of active compute units (CUs) on the accelerator during the - kernel execution. - unit: CUs - Num CUs: - rst: Total number of compute units (CUs) on the accelerator. - unit: CUs - VGPR: - rst: >- - The number of architected vector general-purpose registers allocated for the - kernel, see :ref:`VALU `. Note: this may not exactly match the - number of VGPRs requested by the compiler due to allocation granularity. - unit: VGPRs - SGPR: - rst: >- - The number of scalar general-purpose registers allocated for the kernel, see - :ref:`SALU `. Note: this may not exactly match the number of - SGPRs requested by the compiler due to allocation granularity. - unit: SGPRs - LDS Allocation: - rst: >- - The number of bytes of :doc:`LDS ` memory (or, shared memory) - allocated for this kernel. Note: This may also be larger than what was requested - at compile time due to both allocation granularity and dynamic per-dispatch - LDS allocations. - unit: Bytes per workgroup - Scratch Allocation: - rst: The number of bytes of :ref:`scratch memory ` requested per - work-item for this kernel. Scratch memory is used for stack memory on the accelerator, - as well as for register spills and restores. - unit: Bytes per workgroup - Wavefronts: - rst: The total number of wavefronts, summed over all workgroups, forming this - kernel launch. - unit: Wavefronts - Workgroups: - rst: The total number of workgroups forming this kernel launch. - unit: Workgroups - LDS Req: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. - unit: Instructions per normalization unit - LDS Util: - rst: Indicates what percent of the kernel's duration the :ref:`LDS ` - was actively executing instructions (including, but not limited to, load, store, - atomic and HIP's ``__shfl`` operations). Calculated as the ratio of the total - number of cycles LDS was active over the :ref:`total CU cycles `. - unit: Percent - LDS Latency: - rst: The average number of round-trip cycles (i.e., from issue to data-return - / acknowledgment) required for an LDS instruction to complete. - unit: Cycles - VL1 Rd: - rst: The total number of incoming read requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Wr: - rst: The total number of incoming write requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Atomic: - rst: The total number of incoming atomic requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Hit: - rst: The ratio of the number of vL1D cache line requests that hit in vL1D cache - over the total number of cache line requests to the :ref:`vL1D Cache RAM `. - unit: Percent - VL1 Lat: - rst: Calculated as the average number of cycles that a vL1D cache line request - spent in the vL1D cache pipeline. - unit: Cycles - VL1 Coalesce: - rst: Indicates how well memory instructions were coalesced by the :ref:`address - processing unit `, ranging from uncoalesced (25%) to fully coalesced - (100%). Calculated as the average number of :ref:`thread-requests ` - generated per instruction divided by the ideal number of thread-requests per - instruction. - unit: Percent - VL1 Stall: - rst: The ratio of the number of cycles where the vL1D is stalled waiting to issue - a request for data to the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - VL1_L2 Rd: - rst: The number of read requests for a vL1D cache line that were not satisfied - by the vL1D and must be retrieved from the to the :doc:`L2 Cache ` - per :ref:`normalization unit `. - unit: Requests per normalization unit - VL1_L2 Wr: - rst: The number of write requests to a vL1D cache line that were sent through - the vL1D to the :doc:`L2 cache `, per :ref:`normalization unit `. - unit: Requests per normalization unit - VL1_L2 Atomic: - rst: The number of atomic requests that are sent through the vL1D to the :doc:`L2 - cache `, per :ref:`normalization unit `. This - includes requests for atomics with, and without return. - unit: Requests per normalization unit - sL1D Rd: - rst: The total number of requests, of any size or type, made to the sL1D per :ref:`normalization - unit `. - unit: Requests per normalization unit - sL1D Hit: - rst: The total number of sL1D requests that hit on a previously loaded cache line, - per :ref:`normalization unit `. - unit: Requests per normalization unit - sL1D Lat: rst: '' - unit: Unknown + Wave Life: + rst: '' + SALU: + rst: '' + SMEM: + rst: '' + VALU: + rst: '' + MFMA: + rst: '' + VMEM: + rst: '' + LDS: + rst: '' + GWS: + rst: '' + BR: + rst: '' + Active CUs (deprecated): + rst: '' + Num CUs: + rst: '' + VGPR: + rst: '' + SGPR: + rst: '' + LDS Allocation: + rst: '' + Scratch Allocation: + rst: '' + Wavefronts: + rst: '' + Workgroups: + rst: '' + LDS Req: + rst: '' + LDS Util: + rst: '' + LDS Latency: + rst: '' + VL1 Rd: + rst: '' + VL1 Wr: + rst: '' + VL1 Atomic: + rst: '' + VL1 Hit: + rst: '' + VL1 Coalesce: + rst: '' + VL1 Stall: + rst: '' + VL1_L2 Rd: + rst: '' + VL1_L2 Wr: + rst: '' + VL1_L2 Atomic: + rst: '' + sL1D Rd: + rst: '' + sL1D Hit: + rst: '' sL1D_L2 Rd: - rst: The total number of read requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. - unit: Requests per normalization unit + rst: '' sL1D_L2 Wr: - rst: The total number of write requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit + rst: '' sL1D_L2 Atomic: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit + rst: '' IL1 Fetch: - rst: The total number of requests made to the L1I per :ref:`normalization-unit - `. - unit: Requests per normalization unit + rst: '' IL1 Hit: - rst: The total number of L1I requests that hit on a previously loaded cache line, - per :ref:`normalization-unit `. - unit: Percent + rst: '' IL1 Lat: - rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. - unit: Cycles + rst: '' IL1_L2 Rd: - rst: The total number of requests across the L1I - L2 interface per normalization-unit. - unit: Requests per normalization unit + rst: '' L2 Rd: - rst: The total number of read requests to the L2 from all clients. - unit: Requests per normalization unit + rst: '' L2 Wr: - rst: The total number of write requests to the L2 from all clients. - unit: Requests per normalization unit + rst: '' L2 Atomic: - rst: The total number of atomic requests (with and without return) to the L2 from - all clients. - unit: Requests per normalization unit + rst: '' L2 Hit: - rst: The ratio of the number of L2 cache line requests that hit in the L2 cache - over the total number of incoming cache line requests to the L2 cache. - unit: Percent + rst: '' Fabric_L2 Rd: - rst: Number of L2 cache - Infinity Fabric read requests (either 32-byte or 64-byte) - summed over TCC instances per normalization unit. - unit: Requests per normalization unit + rst: '' Fabric_L2 Wr: - rst: Number of L2 cache - Infinity Fabric write requests (either 32-byte or 64-byte) - summed over TCC instances per normalization unit. - unit: Requests per normalization unit + rst: '' Fabric_L2 Atomic: - rst: Number of L2 cache - Infinity Fabric write requests (either 32-byte or 64-byte) - that are actually atomic requests summed over TCC instances per normalization - unit. - unit: Requests per normalization unit + rst: '' Fabric Rd Lat: - rst: The time-averaged number of cycles read requests spent in Infinity Fabric - before data was returned to the L2. - unit: Cycles + rst: '' Fabric Wr Lat: - rst: The time-averaged number of cycles write requests spent in Infinity Fabric - before a completion acknowledgement was returned to the L2. - unit: Cycles + rst: '' Fabric Atomic Lat: - rst: The time-averaged number of cycles atomic requests spent in Infinity Fabric - before a completion acknowledgement (atomic without return value) or data (atomic - with return value) was returned to the L2. - unit: Cycles + rst: '' HBM Rd: - rst: The total number of L2 requests to Infinity Fabric to read 32B or 64B of - data from the accelerator's local HBM, per :ref:`normalization unit `. - See :ref:`l2-request-flow` for more detail. - unit: Requests per normalization unit + rst: '' HBM Wr: - rst: The total number of L2 requests to Infinity Fabric to write 32B or 64B of - data from the accelerator's local HBM, per :ref:`normalization unit `. - See :ref:`l2-request-flow` for more detail. - unit: Requests per normalization unit -Roofline Performance Rates: + rst: '' VALU FLOPs (F16): - rst: >- - The total 16-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F16 FLOPs achievable - on the specific accelerator. Note: this does not include any F16 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs + rst: '' VALU FLOPs (F32): - rst: >- - The total 32-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F32 FLOPs achievable - on the specific accelerator. Note: this does not include any F32 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs + rst: '' VALU FLOPs (F64): - rst: >- - The total 64-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F64 FLOPs achievable - on the specific accelerator. Note: this does not include any F64 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs - MFMA FLOPs (F64): - rst: >- - The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F64 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F32 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F16 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. The - peak empirically measured BF16 MFMA operations achievable on the specific - accelerator is displayed alongside for comparison. - unit: GFLOPs + rst: '' MFMA FLOPs (F8): - rst: >- - The total number of 8-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. The - peak empirically measured F8 MFMA operations achievable on the specific - accelerator is displayed alongside for comparison. It is supported on AMD - Instinct MI300 series and later only. - unit: GFLOPs + rst: '' + MFMA FLOPs (BF16): + rst: '' + MFMA FLOPs (F16): + rst: '' + MFMA FLOPs (F32): + rst: '' + MFMA FLOPs (F64): + rst: '' MFMA IOPs (Int8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. The peak empirically measured INT8 MFMA - operations achievable on the specific accelerator is displayed alongside - for comparison. - unit: GIOPs + rst: '' HBM Bandwidth: - rst: >- - The total number of bytes read from and written to High-Bandwidth - Memory (HBM) per second. The peak empirically measured bandwidth achievable - on the specific accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' L2 Cache Bandwidth: - rst: The number of bytes looked up in the L2 cache per unit time. The number of - bytes is calculated as the number of cache lines requested multiplied by the - cache line size. This value does not consider partial requests, so e.g., if - only a single value is requested in a cache line, the data movement will still - be counted as a full cache line. The peak empirically measured bandwidth achievable - on the specific accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' L1 Cache Bandwidth: - rst: The number of bytes looked up in the vL1D cache as a result of :ref:`VMEM - ` instructions per unit time. The number of bytes is calculated as - the number of cache lines requested multiplied by the cache line size. This - value does not consider partial requests, so e.g., if only a single value is - requested in a cache line, the data movement will still be counted as a full - cache line. The peak empirically measured bandwidth achievable on the specific - accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' LDS Bandwidth: - rst: Indicates the maximum amount of bytes that could have been loaded from, stored - to, or atomically updated in the LDS per unit time (see :ref:`LDS Bandwidth - ` example for more detail). The peak empirically measured LDS - bandwidth achievable on the specific accelerator is displayed alongside for - comparison. - unit: GB/s -Roofline Plot Points: - AI HBM: - rst: >- - The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM). - It is the ratio of total floating-point operations (FLOPs) to total bytes - transferred between HBM and the L2 cache. This value is used as the x-coordinate - for the HBM roofline. - unit: FLOPs/Byte - AI L2: - rst: >- - The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio - of total floating-point operations (FLOPs) to total bytes transferred between - the L2 cache and the L1 cache. This value is used as the x-coordinate for - the L2 roofline. - unit: FLOPs/Byte + rst: '' AI L1: - rst: >- - The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio - of total floating-point operations (FLOPs) to total bytes transferred between - the L1 cache and the processing units. This value is used as the x-coordinate - for the L1 roofline. - unit: FLOPs/Byte + rst: '' + AI L2: + rst: '' + AI HBM: + rst: '' Performance (GFLOPs): - rst: >- - The overall achieved performance, measured in GigaFLOPs - per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point - operations divided by the total execution time. This value is used as the y-coordinate - for the kernel's point on the Roofline plot. - unit: GFLOP/s + rst: '' + Global/Generic Instr: + rst: '' + Global/Generic Read: + rst: '' + Global/Generic Write: + rst: '' + Global/Generic Atomic: + rst: '' + Spill/Stack Instr: + rst: '' + Spill/Stack Read: + rst: '' + Spill/Stack Write: + rst: '' + Spill/Stack Atomic: + rst: '' + Stalled on L2 Data: + rst: '' + Stalled on L2 Req: + rst: '' + Tag RAM Stall (Read): + rst: '' + Tag RAM Stall (Write): + rst: '' + Tag RAM Stall (Atomic): + rst: '' + NC - Read: + rst: '' + UC - Read: + rst: '' + CC - Read: + rst: '' + RW - Read: + rst: '' + RW - Write: + rst: '' + NC - Write: + rst: '' + UC - Write: + rst: '' + CC - Write: + rst: '' + NC - Atomic: + rst: '' + UC - Atomic: + rst: '' + CC - Atomic: + rst: '' + RW - Atomic: + rst: '' + Req: + rst: '' + Hit Ratio: + rst: '' + Hits: + rst: '' + Translation Misses: + rst: '' + Permission Misses: + rst: '' + L2 Cache Hit Rate: + rst: '' Command processor fetcher (CPF): CPF Utilization: rst: Percent of total cycles where the CPF was busy actively doing any work. The @@ -599,10 +438,9 @@ Workgroup manager utilizations: any work. unit: Percent Scheduler-Pipe Utilization: - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where the scheduler-pipes were actively doing any work. Note: this - value is expected to range between 0% and 25%. See :ref:`desc-spi`. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where the scheduler-pipes were actively doing any work. Note: this + value is expected to range between 0% and 25%. See :ref:`desc-spi`.' unit: Percent Workgroup Manager Utilization: rst: The percent of cycles in the kernel where the workgroup manager was actively @@ -637,30 +475,25 @@ Workgroup manager utilizations: unit: Cycles/wave Workgroup Manager - Resource Allocation: Not-scheduled Rate (Workgroup Manager): - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` - due to a bottleneck within the workgroup manager rather than a lack of a - CU or :ref:`SIMD ` with sufficient resources. Note: this value - is expected to range between 0-25%. See note in :ref:`workgroup manager ` - description. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` + due to a bottleneck within the workgroup manager rather than a lack of a CU + or :ref:`SIMD ` with sufficient resources. Note: this value is expected + to range between 0-25%. See note in :ref:`workgroup manager ` description.' unit: Percent Not-scheduled Rate (Scheduler-Pipe): - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` - due to a bottleneck within the scheduler-pipes rather than a lack of a CU - or :ref:`SIMD ` with sufficient resources. Note: this value is - expected to range between 0-25%, see note in :ref:`workgroup manager ` - description. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` + due to a bottleneck within the scheduler-pipes rather than a lack of a CU or + :ref:`SIMD ` with sufficient resources. Note: this value is expected + to range between 0-25%, see note in :ref:`workgroup manager ` description.' unit: Percent Scheduler-Pipe Stall Rate: - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` due to occupancy limitations (like a lack of a CU or :ref:`SIMD ` - with sufficient resources). Note: this value is expected to range between - 0-25%, see note in :ref:`workgroup manager ` description. + with sufficient resources). Note: this value is expected to range between 0-25%, + see note in :ref:`workgroup manager ` description.' unit: Percent Scratch Stall Rate: rst: The percent of :ref:`total shader-engine cycles ` in the @@ -707,7 +540,7 @@ Workgroup Manager - Resource Allocation: within the workgroup manager. This is expected to be always be zero on CDNA2 or newer accelerators (and small for previous accelerators). unit: Percent -Wavefront Launch Stats: +Wavefront launch stats: Grid Size: rst: The total number of work-items (or, threads) launched as a part of the kernel dispatch. In HIP, this is equivalent to the total grid size multiplied by the @@ -719,11 +552,10 @@ Wavefront Launch Stats: block size. unit: Work-Items Total Wavefronts: - rst: >- - The total number of wavefronts launched as part of the kernel dispatch. - On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront - size is always 64 work-items. Thus, the total number of wavefronts should - be equivalent to the ceiling of grid size divided by 64. + rst: The total number of wavefronts launched as part of the kernel dispatch. On + AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront + size is always 64 work-items. Thus, the total number of wavefronts should be + equivalent to the ceiling of grid size divided by 64. unit: Wavefronts Saved Wavefronts: rst: The total number of wavefronts saved at a context-save. See `cwsr_enable @@ -734,36 +566,32 @@ Wavefront Launch Stats: `_. unit: Wavefronts VGPRs: - rst: >- - The number of architected vector general-purpose registers allocated for the - kernel, see :ref:`VALU `. Note: this may not exactly match the - number of VGPRs requested by the compiler due to allocation granularity. + rst: 'The number of architected vector general-purpose registers allocated for + the kernel, see :ref:`VALU `. Note: this may not exactly match the + number of VGPRs requested by the compiler due to allocation granularity.' unit: VGPRs AGPRs: - rst: >- - The number of accumulation vector general-purpose registers allocated - for the kernel, see :ref:`AGPRs `. Note: this may not exactly match - the number of AGPRs requested by the compiler due to allocation granularity. + rst: 'The number of accumulation vector general-purpose registers allocated for + the kernel, see :ref:`AGPRs `. Note: this may not exactly match + the number of AGPRs requested by the compiler due to allocation granularity.' unit: AGPRs SGPRs: - rst: >- - The number of scalar general-purpose registers allocated for the kernel, see - :ref:`SALU `. Note: this may not exactly match the number of - SGPRs requested by the compiler due to allocation granularity. + rst: 'The number of scalar general-purpose registers allocated for the kernel, + see :ref:`SALU `. Note: this may not exactly match the number of + SGPRs requested by the compiler due to allocation granularity.' unit: SGPRs LDS Allocation: - rst: >- - The number of bytes of :doc:`LDS ` memory (or, shared memory) - allocated for this kernel. Note: This may also be larger than what was requested - at compile time due to both allocation granularity and dynamic per-dispatch - LDS allocations. + rst: 'The number of bytes of :doc:`LDS ` memory (or, shared + memory) allocated for this kernel. Note: This may also be larger than what was + requested at compile time due to both allocation granularity and dynamic per-dispatch + LDS allocations.' unit: Bytes per workgroup Scratch Allocation: rst: The number of bytes of :ref:`scratch memory ` requested per work-item for this kernel. Scratch memory is used for stack memory on the accelerator, as well as for register spills and restores. unit: Bytes per work-item -Wavefront Runtime Stats: +Wavefront runtime stats: Kernel Time: rst: The total duration of the executed kernel. unit: Nanoseconds @@ -775,11 +603,10 @@ Wavefront Runtime Stats: This is averaged over all wavefronts in a kernel dispatch. unit: Instructions per wavefront Wave Cycles: - rst: >- - The number of cycles a wavefront in the kernel dispatch spent resident - on a compute unit per :ref:`normalization unit `. This is - averaged over all wavefronts in a kernel dispatch. Note: this should not - be directly compared to the kernel cycles above. + rst: 'The number of cycles a wavefront in the kernel dispatch spent resident on + a compute unit per :ref:`normalization unit `. This is + averaged over all wavefronts in a kernel dispatch. Note: this should not be + directly compared to the kernel cycles above.' unit: Cycles per normalization unit Dependency Wait Cycles: rst: The number of cycles a wavefront in the kernel dispatch stalled waiting on @@ -813,12 +640,11 @@ Wavefront Runtime Stats: the total Wave Cycles metric. unit: Cycles per normalization unit Wavefront Occupancy: - rst: >- - The time-averaged number of wavefronts resident on the accelerator over the - lifetime of the kernel. Note: this metric may be inaccurate for short-running - kernels (less than 1ms). + rst: 'The time-averaged number of wavefronts resident on the accelerator over + the lifetime of the kernel. Note: this metric may be inaccurate for short-running + kernels (less than 1ms).' unit: Wavefronts -Overall Instruction Mix: +Overall instruction mix: VALU: rst: The total number of vector arithmetic logic unit (VALU) operations issued. These are the workhorses of the :doc:`compute unit `, and are @@ -854,7 +680,7 @@ Overall Instruction Mix: rst: The total number of branch operations issued. These typically consist of jump or branch operations and are used to implement control flow. unit: Instructions -VALU Arithmetic Instruction Mix: +VALU arithmetic instruction mix: INT32: rst: The total number of instructions operating on 32-bit integer operands issued to the VALU per :ref:`normalization unit `. @@ -915,53 +741,10 @@ VALU Arithmetic Instruction Mix: unit `. unit: Instructions per normalization unit Conversion: - rst: >- - The total number of type conversion instructions (such as converting data - to or from F32\u2194F64) issued to the VALU per :ref:`normalization unit - `. + rst: The total number of type conversion instructions (such as converting data + to or from F32\u2194F64) issued to the VALU per :ref:`normalization unit `. unit: Instructions per normalization unit -VMEM Instruction Mix: - Global/Generic Instr: - rst: The total number of global & generic memory instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Read: - rst: The total number of global & generic memory read instructions executed on - all :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Write: - rst: The total number of global & generic memory write instructions executed on - all :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Atomic: - rst: The total number of global & generic memory atomic (with and without return) - instructions executed on all :doc:`compute units ` on the accelerator, - per :ref:`normalization unit `. - unit: Instructions per normalization unit - Spill/Stack Instr: - rst: The total number of spill/stack memory instructions executed on all :doc:`compute - units ` on the accelerator, per :ref:`normalization unit `. - unit: Instructions per normalization unit - Spill/Stack Read: - rst: The total number of spill/stack memory read instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Spill/Stack Write: - rst: The total number of spill/stack memory write instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Spill/Stack Atomic: - rst: The total number of spill/stack memory atomic (with and without return) instructions - executed on all :doc:`compute units ` on the accelerator, per - :ref:`normalization unit `. Typically unused as these memory - operations are typically used to implement thread-local storage. - unit: Instructions per normalization unit -MFMA Arithmetic Instruction Mix: +MFMA instruction mix: MFMA-I8: rst: The total number of 8-bit integer :ref:`MFMA ` instructions issued per :ref:`normalization unit `. @@ -989,66 +772,53 @@ MFMA Arithmetic Instruction Mix: unit: Instructions per normalization unit Compute Speed-of-Light: VALU FLOPs: - rst: >- - The total floating-point operations executed per second on the :ref:`VALU - `. This is also presented as a percent of the peak theoretical - FLOPs achievable on the specific accelerator. Note: this does not include - any floating-point operations from :ref:`MFMA ` instructions. + rst: 'The total floating-point operations executed per second on the :ref:`VALU + `. This is also presented as a percent of the peak theoretical FLOPs + achievable on the specific accelerator. Note: this does not include any floating-point + operations from :ref:`MFMA ` instructions.' unit: GFLOPs VALU IOPs: - rst: >- - The total integer operations executed per second on the :ref:`VALU `. + rst: 'The total integer operations executed per second on the :ref:`VALU `. This is also presented as a percent of the peak theoretical IOPs achievable on the specific accelerator. Note: this does not include any integer operations - from :ref:`MFMA ` instructions. + from :ref:`MFMA ` instructions.' unit: GIOPs - MFMA FLOPs (F8): - rst: '' - unit: Unknown MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit brain floating - point operations from :ref:`VALU ` instructions. This is also - presented as a percent of the peak theoretical BF16 MFMA operations achievable - on the specific accelerator. + rst: 'The total number of 16-bit brain floating point :ref:`MFMA ` + operations executed per second. Note: this does not include any 16-bit brain + floating point operations from :ref:`VALU ` instructions. This is + also presented as a percent of the peak theoretical BF16 MFMA operations achievable + on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F16 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 16-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F16 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F32 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 32-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 32-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F32 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F64): - rst: >- + rst: 'The total number of 64-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator. The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F64 MFMA operations achievable on the - specific accelerator. The total number of 64-bit floating point :ref:`MFMA - ` operations executed per second. Note: this does not include - any 64-bit floating point operations from :ref:`VALU ` instructions. - This is also presented as a percent of the peak theoretical F64 MFMA operations - achievable on the specific accelerator. + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA IOPs (INT8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. This is also presented as a percent - of the peak theoretical INT8 MFMA operations achievable on the specific accelerator. + rst: 'The total number of 8-bit integer :ref:`MFMA ` operations executed + per second. Note: this does not include any 8-bit integer operations from :ref:`VALU + ` instructions. This is also presented as a percent of the peak theoretical + INT8 MFMA operations achievable on the specific accelerator.' unit: GFLOPs -Pipeline Statistics: +Pipeline statistics: IPC: rst: The ratio of the total number of instructions executed on the :doc:`CU ` over the :ref:`total active CU cycles `. @@ -1111,7 +881,7 @@ Pipeline Statistics: rst: The average number of round-trip cycles (that is, from issue to data return / acknowledgment) required for a SMEM instruction to complete. unit: Cycles -Arithmetic Operations: +Arithmetic operations: FLOPs (Total): rst: The total number of floating-point operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit @@ -1122,20 +892,16 @@ Arithmetic Operations: ` or :ref:`MFMA ` units, per :ref:`normalization unit `. unit: IOP per normalization unit - F8 OPs: - rst: '' - unit: Unknown F16 OPs: rst: The total number of 16-bit floating-point operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit `. unit: FLOP per normalization unit BF16 OPs: - rst: >- - The total number of 16-bit brain floating-point operations executed on - either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization - unit `. Note: on current CDNA accelerators, the VALU - has no native BF16 instructions. + rst: 'The total number of 16-bit brain floating-point operations executed on either + the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization + unit `. Note: on current CDNA accelerators, the VALU has + no native BF16 instructions.' unit: FLOP per normalization unit F32 OPs: rst: The total number of 32-bit floating-point operations executed on either the @@ -1148,11 +914,10 @@ Arithmetic Operations: unit `. unit: FLOP per normalization unit INT8 OPs: - rst: >- - The total number of 8-bit integer operations executed on either the :ref:`VALU + rst: 'The total number of 8-bit integer operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit - `. Note: on current CDNA accelerators, the VALU has - no native INT8 instructions. + `. Note: on current CDNA accelerators, the VALU has no + native INT8 instructions.' unit: IOP per normalization unit LDS Speed-of-Light: Utilization: @@ -1182,16 +947,16 @@ LDS Speed-of-Light: amount of data in an uncontended access. [#lds-bank-conflict]_ unit: Percent LDS Statistics: - LDS Instructions: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. - unit: Instructions per normalization unit Theoretical Bandwidth: rst: Indicates the maximum amount of bytes that could have been loaded from, stored to, or atomically updated in the LDS divided by total duration. Does *not* take into account the execution mask of the wavefront when the instruction was executed. See the :ref:`LDS bandwidth example ` for more detail. unit: Gbps + LDS Instructions: + rst: The total number of LDS instructions (including, but not limited to, read/write/atomics + and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. + unit: Instructions per normalization unit LDS Latency: rst: The average number of round-trip cycles (i.e., from issue to data-return acknowledgment) required for an LDS instruction to complete. @@ -1225,10 +990,9 @@ LDS Statistics: to stalls from non-dword aligned addresses per :ref:`normalization unit `. unit: Cycles per normalization unit Mem Violations: - rst: >- - The total number of out-of-bounds accesses made to the LDS, per :ref:`normalization - unit `. This is unused and expected to be zero in - most configurations for modern CDNA\u2122 accelerators. + rst: The total number of out-of-bounds accesses made to the LDS, per :ref:`normalization + unit `. This is unused and expected to be zero in most + configurations for modern CDNA\u2122 accelerators. unit: Accesses per normalization unit L1I Speed-of-Light: Bandwidth Utilization: @@ -1236,18 +1000,17 @@ L1I Speed-of-Light: theoretical bandwidth. Calculated as the ratio of L1I requests over the :ref:`total L1I cycles `. unit: Percent + L1I-L2 Bandwidth Utilization: + rst: The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth + achieved. Calculated as the ratio of the total number of requests from the L1I + to the L2 cache over the :ref:`total L1I-L2 interface cycles `. + unit: Percent +L1I cache accesses: Cache Hit Rate: rst: The percent of L1I requests that hit [#l1i-cache]_ on a previously loaded line the cache. Calculated as the ratio of the number of L1I requests that hit over the number of all L1I requests. unit: Percent - L1I-L2 Bandwidth Utilization: - rst: >- - The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth - achieved. Calculated as the ratio of the total number of requests from - the L1I to the L2 cache over the :ref:`total L1I-L2 interface cycles `. - unit: Percent -L1I cache accesses: Req: rst: The total number of requests made to the L1I per normalization-unit unit: Requests per normalization unit @@ -1265,11 +1028,6 @@ L1I cache accesses: already pending due to another request, per :ref:`normalization-unit `. See note in :ref:`desc-l1i-sol` for more detail. unit: Requests per normalization unit - Cache Hit Rate: - rst: The percent of L1I requests that hit [#l1i-cache]_ on a previously loaded - line the cache. Calculated as the ratio of the number of L1I requests that hit - over the number of all L1I requests. - unit: Percent Instruction Fetch Latency: rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. unit: Cycles @@ -1284,17 +1042,17 @@ Scalar L1D Speed-of-Light: theoretical bandwidth. Calculated as the ratio of sL1D requests over the :ref:`total sL1D cycles `. unit: Percent + sL1D-L2 BW Utilization: + rst: The percentage of the peak theoretical sL1D - L2 interface bandwidth acheived. + Calculated as total number of bytes read from, written to, or atomically updated + across the sL1D - L2 interface. + unit: Percent +Scalar L1D cache accesses: Cache Hit Rate: rst: Indicates the percent of sL1D requests that hit on a previously loaded line the cache. The ratio of the number of sL1D requests that hit [#sl1d-cache]_ over the number of all sL1D requests. unit: Percent - sL1D-L2 BW Utilization: - rst: The percentage of the peak theoretical sL1D - L2 interface bandwidth acheived. - Caclulated as total number of bytes read from, written to, or atomically updated - across the sL1D - L2 interface. - unit: Percent -Scalar L1D cache accesses: Req: rst: The total number of requests, of any size or type, made to the sL1D per :ref:`normalization unit `. @@ -1313,20 +1071,10 @@ Scalar L1D cache accesses: already pending due to another request, per :ref:`normalization unit `. See :ref:`desc-sl1d-sol` for more detail. unit: Requests per normalization unit - Cache Hit Rate: - rst: Indicates the percent of sL1D requests that hit on a previously loaded line - the cache. The ratio of the number of sL1D requests that hit [#sl1d-cache]_ - over the number of all sL1D requests. - unit: Percent Read Req (Total): rst: The total number of sL1D read requests of any size, per :ref:`normalization unit `. unit: Requests per normalization unit - Atomic Req: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit Read Req (1 DWord): rst: The total number of sL1D read requests made for a single dword of data (4B), per :ref:`normalization unit `. @@ -1349,13 +1097,17 @@ Scalar L1D cache accesses: unit: Requests per normalization unit Scalar L1D Cache - L2 Interface: sL1D-L2 BW: - rst: >- - The total number of bytes read from, written to, or atomically updated - across the sL1D\u2194:doc:`L2 ` interface, divided by total duration. - Note that sL1D writes and atomics are typically - unused on current CDNA accelerators, so in the majority of cases this can - be interpreted as an sL1D\u2192L2 read bandwidth. + rst: The total number of bytes read from, written to, or atomically updated across + the sL1D\u2194:doc:`L2 ` interface, divided by total duration. Note + that sL1D writes and atomics are typically unused on current CDNA accelerators, + so in the majority of cases this can be interpreted as an sL1D\u2192L2 read + bandwidth. unit: Gbps + Atomic Req: + rst: The total number of atomic requests from sL1D to the :doc:`L2 `, + per :ref:`normalization unit `. Typically unused on current + CDNA accelerators. + unit: Requests per normalization unit Read Req: rst: The total number of read requests from sL1D to the :doc:`L2 `, per :ref:`normalization unit `. @@ -1365,17 +1117,11 @@ Scalar L1D Cache - L2 Interface: per :ref:`normalization unit `. Typically unused on current CDNA accelerators. unit: Requests per normalization unit - Atomic Req: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit Stall Cycles: - rst: >- - The total number of cycles the sL1D\u2194 :doc:`L2 ` interface + rst: The total number of cycles the sL1D\u2194 :doc:`L2 ` interface was stalled, per :ref:`normalization unit `. unit: Cycles per normalization unit -Busy and stall metrics: +Busy / stall metrics: Address Processing Unit Busy: rst: Percent of the :ref:`total CU cycles ` the address processor was busy @@ -1388,19 +1134,10 @@ Busy and stall metrics: rst: Percent of the :ref:`total CU cycles ` the address processor was stalled from sending write/atomic data further into the vL1D pipeline unit: Percent - Data-Processor → Address Stall: + "Data-Processor \u2192 Address Stall": rst: Percent of :ref:`total CU cycles ` the address processor was stalled waiting to send command data to the :ref:`data processor ` unit: Percent - Sequencer → TA Address Stall: - rst: '' - unit: Unknown - Sequencer → TA Command Stall: - rst: '' - unit: Unknown - Sequencer → TA Data Stall: - rst: '' - unit: Unknown Instruction counts: Total Instructions: rst: The total number of memory instructions executed by the address processer @@ -1446,7 +1183,7 @@ Instruction counts: :ref:`normalization unit `. Typically unused as these memory operations are typically used to implement thread-local storage. unit: Instructions per normalization unit -Spill and stack metrics: +Spill / stack metrics: Spill/Stack Total Cycles: rst: The number of cycles the address processing unit spent working on spill/stack instructions, per :ref:`normalization unit `. @@ -1464,11 +1201,11 @@ Vector L1 data-return path or Texture Data (TD): rst: Percent of the :ref:`total CU cycles ` the data-return unit was busy processing or waiting on data to return to the :doc:`CU `. unit: Percent - Cache RAM → Data-Return Stall: + "Cache RAM \u2192 Data-Return Stall": rst: Percent of the :ref:`total CU cycles ` the data-return unit was stalled on data to be returned from the :ref:`vL1D Cache RAM `. unit: Percent - Workgroup manager → Data-Return Stall: + "Workgroup manager \u2192 Data-Return Stall": rst: Percent of the :ref:`total CU cycles ` the data-return unit was stalled by the :ref:`workgroup manager ` due to initialization of registers as a part of launching new workgroups. @@ -1525,32 +1262,6 @@ vL1D Speed-of-Light: generated per instruction divided by the ideal number of thread-requests per instruction. unit: Percent -vL1D cache stall metrics: - Stalled on L2 Data: - rst: The ratio of the number of cycles where the vL1D is stalled waiting for requested - data to return from the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Stalled on L2 Req: - rst: The ratio of the number of cycles where the vL1D is stalled waiting to issue - a request for data to the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Read): - rst: The ratio of the number of cycles where the vL1D is stalled due to Read requests - with conflicting tags being looked up concurrently, divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Write): - rst: The ratio of the number of cycles where the vL1D is stalled due to Write - requests with conflicting tags being looked up concurrently, divided by the - number of cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Atomic): - rst: The ratio of the number of cycles where the vL1D is stalled due to Atomic - requests with conflicting tags being looked up concurrently, divided by the - number of cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent vL1D cache access metrics: Total Req: rst: The total number of incoming requests from the :ref:`address processing unit @@ -1615,79 +1326,6 @@ vL1D cache access metrics: cache `, per :ref:`normalization unit `. This includes requests for atomics with, and without return. unit: Requests per normalization unit -L1D - L2 Transactions: - NC - Read: - rst: Total read requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Read: - rst: Total read requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Read: - rst: Total read requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Read: - rst: Total read requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Write: - rst: Total write requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - NC - Write: - rst: Total write requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Write: - rst: Total write requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Write: - rst: Total write requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - NC - Atomic: - rst: Total atomic requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Atomic: - rst: Total atomic requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Atomic: - rst: Total atomic requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Atomic: - rst: Total atomic requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit -L1 Unified Translation Cache (UTCL1): - Req: - rst: The number of translation requests made to the UTCL1 per normalization unit. - unit: Requests per normalization unit - Hit Ratio: - rst: The ratio of the number of translation requests that hit in the UTCL1 divided - by the total number of translation requests made to the UTCL1. - unit: Percent - Hits: - rst: The number of translation requests that hit in the UTCL1, and could be reused, - per normalization unit. - unit: Requests per normalization unit - Translation Misses: - rst: The total number of translation requests that missed in the UTCL1 due to - translation not being present in the cache, per :ref:`normalization unit `. - unit: unit - Permission Misses: - rst: >- - The total number of translation requests that missed in the UTCL1 due - to a permission error, per :ref:`normalization unit `. - This is unused and expected to be zero in most configurations for modern - CDNA\u2122 accelerators. - unit: Requests per normalization unit -L1D Addr Translation Stalls: {} L2 Speed-of-Light: Utilization: rst: The ratio of the :ref:`number of cycles an L2 channel was active, summed @@ -1809,7 +1447,7 @@ L2-Fabric interface metrics: before a completion acknowledgement (atomic without return value) or data (atomic with return value) was returned to the L2. unit: Cycles -L2 Cache Accesses: +L2 cache accesses: Bandwidth: rst: The number of bytes looked up in the L2 cache, divided by total duration. The number of bytes is calculated as the number of cache lines requested multiplied @@ -1900,7 +1538,6 @@ L2 Cache Accesses: rst: The total number of requests to the L2 that go to Read-Write coherent memory (RW) allocations. See the :ref:`memory-type` for more information. unit: Requests per normalization unit -L2 Cache Stalls: {} L2 - Fabric Interface stalls: Write - Credit Starvation: rst: The number of cycles the L2-Fabric interface was stalled on write or atomic @@ -1969,72 +1606,3 @@ L2 - Fabric interface detailed metrics: as :ref:`fine-grained memory ` allocations or :ref:`uncached memory ` allocations on the MI2XX. unit: Requests per normalization unit -Aggregate Stats (All channels): - L2 Cache Hit Rate: - rst: The total number of requests to the L2 from all clients that hit in the cache. - As noted in the :ref:`Speed-of-Light ` section, this includes hit-on-miss - requests. - unit: Percent -L2 Cache Hit Rate (pct): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2 Requests (per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Requests (per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Read Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Write and Atomic Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Atomic Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Read Stall (Cycles per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Write and Atomic Stall (Cycles per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric (128B read requests per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown diff --git a/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx942_metrics_description.yaml b/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx942_metrics_description.yaml index be9b4719e0..c9079a32ad 100644 --- a/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx942_metrics_description.yaml +++ b/projects/rocprofiler-compute/tools/per_arch_metric_definitions/gfx942_metrics_description.yaml @@ -1,69 +1,58 @@ +# AUTOGENERATED FILE. Only edit for testing purposes, not for development. Generated by tools/config_management/metric_description_manager.py System Speed-of-Light: VALU FLOPs: - rst: >- - The total floating-point operations executed per second on the :ref:`VALU - `. This is also presented as a percent of the peak theoretical - FLOPs achievable on the specific accelerator. Note: this does not include - any floating-point operations from :ref:`MFMA ` instructions. + rst: 'The total floating-point operations executed per second on the :ref:`VALU + `. This is also presented as a percent of the peak theoretical FLOPs + achievable on the specific accelerator. Note: this does not include any floating-point + operations from :ref:`MFMA ` instructions.' unit: GFLOPs VALU IOPs: - rst: >- - The total integer operations executed per second on the :ref:`VALU `. + rst: 'The total integer operations executed per second on the :ref:`VALU `. This is also presented as a percent of the peak theoretical IOPs achievable on the specific accelerator. Note: this does not include any integer operations - from :ref:`MFMA ` instructions. + from :ref:`MFMA ` instructions.' unit: GOIPs MFMA FLOPs (F8): - rst: >- - The total number of 8-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. This - is also presented as a percent of the peak theoretical F8 MFMA operations - achievable on the specific accelerator. It is supported on AMD Instinct MI300 - series and later only. + rst: 'The total number of 8-bit brain floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit brain floating point + operations from :ref:`VALU ` instructions. This is also presented + as a percent of the peak theoretical F8 MFMA operations achievable on the specific + accelerator. It is supported on AMD Instinct MI300 series and later only.' unit: GFLOPs MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` + rst: 'The total number of 16-bit brain floating point :ref:`MFMA ` operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. This - is also presented as a percent of the peak theoretical BF16 MFMA operations - achievable on the specific accelerator. + floating point operations from :ref:`VALU ` instructions. This is + also presented as a percent of the peak theoretical BF16 MFMA operations achievable + on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F16 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 16-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F16 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F32 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 32-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 32-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F32 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F64): - rst: >- - The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F64 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 64-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA IOPs (Int8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. This is also presented as a percent - of the peak theoretical INT8 MFMA operations achievable on the specific accelerator. + rst: 'The total number of 8-bit integer :ref:`MFMA ` operations executed + per second. Note: this does not include any 8-bit integer operations from :ref:`VALU + ` instructions. This is also presented as a percent of the peak theoretical + INT8 MFMA operations achievable on the specific accelerator.' unit: GIOPs - Active CUs: + Active CUs (deprecated): rst: Total number of active compute units (CUs) on the accelerator during the - kernel execution. + kernel execution. (Deprecated - See CU Utilization instead) unit: Number SALU Utilization: rst: Indicates what percent of the kernel's duration the :ref:`SALU ` @@ -108,11 +97,10 @@ System Speed-of-Light: over the :ref:`total active CU cycles `. unit: Instructions per-cycle Wavefront Occupancy: - rst: >- - The time-averaged number of wavefronts resident on the accelerator over + rst: 'The time-averaged number of wavefronts resident on the accelerator over the lifetime of the kernel. Note: this metric may be inaccurate for short-running kernels (less than 1ms). This is also presented as a percent of the peak theoretical - occupancy achievable on the specific accelerator. + occupancy achievable on the specific accelerator.' unit: Wavefronts Theoretical LDS Bandwidth: rst: Indicates the maximum amount of bytes that could have been loaded from, stored @@ -153,10 +141,9 @@ System Speed-of-Light: peak theoretical bandwidth achievable on the specific accelerator. unit: GB/s L2-Fabric Read BW: - rst: >- - The number of bytes read by the L2 over the :ref:`Infinity Fabric\u2122 - interface ` per unit time. This is also presented as a percent - of the peak theoretical bandwidth achievable on the specific accelerator. + rst: The number of bytes read by the L2 over the :ref:`Infinity Fabric\u2122 interface + ` per unit time. This is also presented as a percent of the peak + theoretical bandwidth achievable on the specific accelerator. unit: GB/s L2-Fabric Write BW: rst: The number of bytes sent by the L2 over the :ref:`Infinity Fabric interface @@ -194,359 +181,211 @@ System Speed-of-Light: L1I Fetch Latency: rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. unit: Cycles -Memory Chart: + CU Utilization: + rst: The percent of :ref:`total SIMD cycles ` in the kernel + where any :ref:`SIMD ` on a CU was actively doing any work, summed + over all CUs. Low values (less than 100%) indicate that the accelerator was + not fully saturated by the kernel, or a potential load-imbalance issue. + unit: Percent +General: Wavefront Occupancy: - rst: Wavefronts per active CU. - unit: Wavefronts - Wave Life: - rst: Average number of cycles executing a wave. - unit: Cycles per wave - SALU: - rst: Total Number of SALU (Scalar ALU) instructions issued per normalization unit. - unit: Instructions per normalization unit - SMEM: - rst: Total number of SMEM (Scalar Memory Read) instructions issued normalization - unit. - unit: Instructions per normalization unit - VALU: - rst: The number of VALU (Vector ALU) instructions issued per normalization unit. - unit: Instructions per normalization unit - MFMA: - rst: Total number of MFMA (Matrix-Fused-Multiply-Add) instructions issued per - normalization unit. - unit: Instructions per normalization unit - VMEM: - rst: The number of VMEM (GPU Memory) read instructions issued (including FLAT/scratch - memory) per normalization unit. - unit: Instructions per normalization unit - LDS: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's __shfl instructions) executed per normalization unit. - unit: Instructions per normalization unit - GWS: - rst: Total number of GDS (global data sync) instructions issued per normalization - unit. - unit: Instructions per normalization unit - BR: - rst: Total number of BRANCH instructions issued per normalization unit. - unit: Instructions per normalization unit - Active CUs: - rst: Total number of active compute units (CUs) on the accelerator during the - kernel execution. - unit: CUs - Num CUs: - rst: Total number of compute units (CUs) on the accelerator. - unit: CUs - VGPR: - rst: >- - The number of architected vector general-purpose registers allocated for the - kernel, see :ref:`VALU `. Note: this may not exactly match the - number of VGPRs requested by the compiler due to allocation granularity. - unit: VGPRs - SGPR: - rst: >- - The number of scalar general-purpose registers allocated for the kernel, see - :ref:`SALU `. Note: this may not exactly match the number of - SGPRs requested by the compiler due to allocation granularity. - unit: SGPRs - LDS Allocation: - rst: >- - The number of bytes of :doc:`LDS ` memory (or, shared memory) - allocated for this kernel. Note: This may also be larger than what was requested - at compile time due to both allocation granularity and dynamic per-dispatch - LDS allocations. - unit: Bytes per workgroup - Scratch Allocation: - rst: The number of bytes of :ref:`scratch memory ` requested per - work-item for this kernel. Scratch memory is used for stack memory on the accelerator, - as well as for register spills and restores. - unit: Bytes per workgroup - Wavefronts: - rst: The total number of wavefronts, summed over all workgroups, forming this - kernel launch. - unit: Wavefronts - Workgroups: - rst: The total number of workgroups forming this kernel launch. - unit: Workgroups - LDS Req: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. - unit: Instructions per normalization unit - LDS Util: - rst: Indicates what percent of the kernel's duration the :ref:`LDS ` - was actively executing instructions (including, but not limited to, load, store, - atomic and HIP's ``__shfl`` operations). Calculated as the ratio of the total - number of cycles LDS was active over the :ref:`total CU cycles `. - unit: Percent - LDS Latency: - rst: The average number of round-trip cycles (i.e., from issue to data-return - / acknowledgment) required for an LDS instruction to complete. - unit: Cycles - VL1 Rd: - rst: The total number of incoming read requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Wr: - rst: The total number of incoming write requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Atomic: - rst: The total number of incoming atomic requests from the :ref:`address processing - unit ` after coalescing per :ref:`normalization unit ` - unit: Requests per normalization unit - VL1 Hit: - rst: The ratio of the number of vL1D cache line requests that hit in vL1D cache - over the total number of cache line requests to the :ref:`vL1D Cache RAM `. - unit: Percent - VL1 Lat: - rst: Calculated as the average number of cycles that a vL1D cache line request - spent in the vL1D cache pipeline. - unit: Cycles - VL1 Coalesce: - rst: Indicates how well memory instructions were coalesced by the :ref:`address - processing unit `, ranging from uncoalesced (25%) to fully coalesced - (100%). Calculated as the average number of :ref:`thread-requests ` - generated per instruction divided by the ideal number of thread-requests per - instruction. - unit: Percent - VL1 Stall: - rst: The ratio of the number of cycles where the vL1D is stalled waiting to issue - a request for data to the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - VL1_L2 Rd: - rst: The number of read requests for a vL1D cache line that were not satisfied - by the vL1D and must be retrieved from the to the :doc:`L2 Cache ` - per :ref:`normalization unit `. - unit: Requests per normalization unit - VL1_L2 Wr: - rst: The number of write requests to a vL1D cache line that were sent through - the vL1D to the :doc:`L2 cache `, per :ref:`normalization unit `. - unit: Requests per normalization unit - VL1_L2 Atomic: - rst: The number of atomic requests that are sent through the vL1D to the :doc:`L2 - cache `, per :ref:`normalization unit `. This - includes requests for atomics with, and without return. - unit: Requests per normalization unit - sL1D Rd: - rst: The total number of requests, of any size or type, made to the sL1D per :ref:`normalization - unit `. - unit: Requests per normalization unit - sL1D Hit: - rst: The total number of sL1D requests that hit on a previously loaded cache line, - per :ref:`normalization unit `. - unit: Requests per normalization unit - sL1D Lat: rst: '' - unit: Unknown + Wave Life: + rst: '' + SALU: + rst: '' + SMEM: + rst: '' + VALU: + rst: '' + MFMA: + rst: '' + VMEM: + rst: '' + LDS: + rst: '' + GWS: + rst: '' + BR: + rst: '' + Active CUs (deprecated): + rst: '' + Num CUs: + rst: '' + VGPR: + rst: '' + SGPR: + rst: '' + LDS Allocation: + rst: '' + Scratch Allocation: + rst: '' + Wavefronts: + rst: '' + Workgroups: + rst: '' + LDS Req: + rst: '' + LDS Util: + rst: '' + LDS Latency: + rst: '' + VL1 Rd: + rst: '' + VL1 Wr: + rst: '' + VL1 Atomic: + rst: '' + VL1 Hit: + rst: '' + VL1 Coalesce: + rst: '' + VL1 Stall: + rst: '' + VL1_L2 Rd: + rst: '' + VL1_L2 Wr: + rst: '' + VL1_L2 Atomic: + rst: '' + sL1D Rd: + rst: '' + sL1D Hit: + rst: '' sL1D_L2 Rd: - rst: The total number of read requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. - unit: Requests per normalization unit + rst: '' sL1D_L2 Wr: - rst: The total number of write requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit + rst: '' sL1D_L2 Atomic: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit + rst: '' IL1 Fetch: - rst: The total number of requests made to the L1I per :ref:`normalization-unit - `. - unit: Requests per normalization unit + rst: '' IL1 Hit: - rst: The total number of L1I requests that hit on a previously loaded cache line, - per :ref:`normalization-unit `. - unit: Percent + rst: '' IL1 Lat: - rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. - unit: Cycles + rst: '' IL1_L2 Rd: - rst: The total number of requests across the L1I - L2 interface per normalization-unit. - unit: Requests per normalization unit + rst: '' L2 Rd: - rst: The total number of read requests to the L2 from all clients. - unit: Requests per normalization unit + rst: '' L2 Wr: - rst: The total number of write requests to the L2 from all clients. - unit: Requests per normalization unit + rst: '' L2 Atomic: - rst: The total number of atomic requests (with and without return) to the L2 from - all clients. - unit: Requests per normalization unit + rst: '' L2 Hit: - rst: The ratio of the number of L2 cache line requests that hit in the L2 cache - over the total number of incoming cache line requests to the L2 cache. - unit: Percent + rst: '' Fabric_L2 Rd: - rst: Number of L2 cache - Infinity Fabric read requests (either 32-byte or 64-byte) - summed over TCC instances per normalization unit. - unit: Requests per normalization unit + rst: '' Fabric_L2 Wr: - rst: Number of L2 cache - Infinity Fabric write requests (either 32-byte or 64-byte) - summed over TCC instances per normalization unit. - unit: Requests per normalization unit + rst: '' Fabric_L2 Atomic: - rst: Number of L2 cache - Infinity Fabric write requests (either 32-byte or 64-byte) - that are actually atomic requests summed over TCC instances per normalization - unit. - unit: Requests per normalization unit + rst: '' Fabric Rd Lat: - rst: The time-averaged number of cycles read requests spent in Infinity Fabric - before data was returned to the L2. - unit: Cycles + rst: '' Fabric Wr Lat: - rst: The time-averaged number of cycles write requests spent in Infinity Fabric - before a completion acknowledgement was returned to the L2. - unit: Cycles + rst: '' Fabric Atomic Lat: - rst: The time-averaged number of cycles atomic requests spent in Infinity Fabric - before a completion acknowledgement (atomic without return value) or data (atomic - with return value) was returned to the L2. - unit: Cycles + rst: '' HBM Rd: - rst: The total number of L2 requests to Infinity Fabric to read 32B or 64B of - data from the accelerator's local HBM, per :ref:`normalization unit `. - See :ref:`l2-request-flow` for more detail. - unit: Requests per normalization unit + rst: '' HBM Wr: - rst: The total number of L2 requests to Infinity Fabric to write 32B or 64B of - data from the accelerator's local HBM, per :ref:`normalization unit `. - See :ref:`l2-request-flow` for more detail. - unit: Requests per normalization unit -Roofline Performance Rates: + rst: '' VALU FLOPs (F16): - rst: >- - The total 16-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F16 FLOPs achievable - on the specific accelerator. Note: this does not include any F16 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs + rst: '' VALU FLOPs (F32): - rst: >- - The total 32-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F32 FLOPs achievable - on the specific accelerator. Note: this does not include any F32 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs + rst: '' VALU FLOPs (F64): - rst: >- - The total 64-bit floating-point operations executed per second on the :ref:`VALU - `. This is presented with the value of the peak empirical F64 FLOPs achievable - on the specific accelerator. Note: this does not include any F64 operations - from :ref:`MFMA ` instructions. - unit: GFLOPs - MFMA FLOPs (F64): - rst: >- - The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F64 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F32 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. The peak empirically - measured F16 MFMA operations achievable on the specific accelerator is - displayed alongside for comparison. - unit: GFLOPs - MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. The - peak empirically measured BF16 MFMA operations achievable on the specific - accelerator is displayed alongside for comparison. - unit: GFLOPs + rst: '' MFMA FLOPs (F8): - rst: >- - The total number of 8-bit brain floating point :ref:`MFMA ` - operations executed per second. Note: this does not include any 16-bit brain - floating point operations from :ref:`VALU ` instructions. The - peak empirically measured F8 MFMA operations achievable on the specific - accelerator is displayed alongside for comparison. It is supported on AMD - Instinct MI300 series and later only. - unit: GFLOPs + rst: '' + MFMA FLOPs (BF16): + rst: '' + MFMA FLOPs (F16): + rst: '' + MFMA FLOPs (F32): + rst: '' + MFMA FLOPs (F64): + rst: '' MFMA IOPs (Int8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. The peak empirically measured INT8 MFMA - operations achievable on the specific accelerator is displayed alongside - for comparison. - unit: GIOPs + rst: '' HBM Bandwidth: - rst: >- - The total number of bytes read from and written to High-Bandwidth - Memory (HBM) per second. The peak empirically measured bandwidth achievable - on the specific accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' L2 Cache Bandwidth: - rst: The number of bytes looked up in the L2 cache per unit time. The number of - bytes is calculated as the number of cache lines requested multiplied by the - cache line size. This value does not consider partial requests, so e.g., if - only a single value is requested in a cache line, the data movement will still - be counted as a full cache line. The peak empirically measured bandwidth achievable - on the specific accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' L1 Cache Bandwidth: - rst: The number of bytes looked up in the vL1D cache as a result of :ref:`VMEM - ` instructions per unit time. The number of bytes is calculated as - the number of cache lines requested multiplied by the cache line size. This - value does not consider partial requests, so e.g., if only a single value is - requested in a cache line, the data movement will still be counted as a full - cache line. The peak empirically measured bandwidth achievable on the specific - accelerator is displayed alongside for comparison. - unit: GB/s + rst: '' LDS Bandwidth: - rst: Indicates the maximum amount of bytes that could have been loaded from, stored - to, or atomically updated in the LDS per unit time (see :ref:`LDS Bandwidth - ` example for more detail). The peak empirically measured LDS - bandwidth achievable on the specific accelerator is displayed alongside for - comparison. - unit: GB/s -Roofline Plot Points: - AI HBM: - rst: >- - The Arithmetic Intensity (AI) relative to High-Bandwidth Memory (HBM). - It is the ratio of total floating-point operations (FLOPs) to total bytes - transferred between HBM and the L2 cache. This value is used as the x-coordinate - for the HBM roofline. - unit: FLOPs/Byte - AI L2: - rst: >- - The Arithmetic Intensity (AI) relative to the L2 Cache. It is the ratio - of total floating-point operations (FLOPs) to total bytes transferred between - the L2 cache and the L1 cache. This value is used as the x-coordinate for - the L2 roofline. - unit: FLOPs/Byte + rst: '' AI L1: - rst: >- - The Arithmetic Intensity (AI) relative to the L1 Cache. It is the ratio - of total floating-point operations (FLOPs) to total bytes transferred between - the L1 cache and the processing units. This value is used as the x-coordinate - for the L1 roofline. - unit: FLOPs/Byte + rst: '' + AI L2: + rst: '' + AI HBM: + rst: '' Performance (GFLOPs): - rst: >- - The overall achieved performance, measured in GigaFLOPs - per second (GFLOP/s). This is calculated as the sum of all VALU and MFMA floating-point - operations divided by the total execution time. This value is used as the y-coordinate - for the kernel's point on the Roofline plot. - unit: GFLOP/s + rst: '' + Global/Generic Instr: + rst: '' + Global/Generic Read: + rst: '' + Global/Generic Write: + rst: '' + Global/Generic Atomic: + rst: '' + Spill/Stack Instr: + rst: '' + Spill/Stack Read: + rst: '' + Spill/Stack Write: + rst: '' + Spill/Stack Atomic: + rst: '' + Stalled on L2 Data: + rst: '' + Stalled on L2 Req: + rst: '' + Tag RAM Stall (Read): + rst: '' + Tag RAM Stall (Write): + rst: '' + Tag RAM Stall (Atomic): + rst: '' + NC - Read: + rst: '' + UC - Read: + rst: '' + CC - Read: + rst: '' + RW - Read: + rst: '' + RW - Write: + rst: '' + NC - Write: + rst: '' + UC - Write: + rst: '' + CC - Write: + rst: '' + NC - Atomic: + rst: '' + UC - Atomic: + rst: '' + CC - Atomic: + rst: '' + RW - Atomic: + rst: '' + Req: + rst: '' + Hit Ratio: + rst: '' + Hits: + rst: '' + Translation Misses: + rst: '' + Permission Misses: + rst: '' + L2 Cache Hit Rate: + rst: '' Command processor fetcher (CPF): CPF Utilization: rst: Percent of total cycles where the CPF was busy actively doing any work. The @@ -599,10 +438,9 @@ Workgroup manager utilizations: any work. unit: Percent Scheduler-Pipe Utilization: - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where the scheduler-pipes were actively doing any work. Note: this - value is expected to range between 0% and 25%. See :ref:`desc-spi`. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where the scheduler-pipes were actively doing any work. Note: this + value is expected to range between 0% and 25%. See :ref:`desc-spi`.' unit: Percent Workgroup Manager Utilization: rst: The percent of cycles in the kernel where the workgroup manager was actively @@ -637,30 +475,25 @@ Workgroup manager utilizations: unit: Cycles/wave Workgroup Manager - Resource Allocation: Not-scheduled Rate (Workgroup Manager): - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` - due to a bottleneck within the workgroup manager rather than a lack of a - CU or :ref:`SIMD ` with sufficient resources. Note: this value - is expected to range between 0-25%. See note in :ref:`workgroup manager ` - description. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` + due to a bottleneck within the workgroup manager rather than a lack of a CU + or :ref:`SIMD ` with sufficient resources. Note: this value is expected + to range between 0-25%. See note in :ref:`workgroup manager ` description.' unit: Percent Not-scheduled Rate (Scheduler-Pipe): - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` - due to a bottleneck within the scheduler-pipes rather than a lack of a CU - or :ref:`SIMD ` with sufficient resources. Note: this value is - expected to range between 0-25%, see note in :ref:`workgroup manager ` - description. + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` + due to a bottleneck within the scheduler-pipes rather than a lack of a CU or + :ref:`SIMD ` with sufficient resources. Note: this value is expected + to range between 0-25%, see note in :ref:`workgroup manager ` description.' unit: Percent Scheduler-Pipe Stall Rate: - rst: >- - The percent of :ref:`total scheduler-pipe cycles ` - in the kernel where a workgroup could not be scheduled to a :doc:`CU ` + rst: 'The percent of :ref:`total scheduler-pipe cycles ` in + the kernel where a workgroup could not be scheduled to a :doc:`CU ` due to occupancy limitations (like a lack of a CU or :ref:`SIMD ` - with sufficient resources). Note: this value is expected to range between - 0-25%, see note in :ref:`workgroup manager ` description. + with sufficient resources). Note: this value is expected to range between 0-25%, + see note in :ref:`workgroup manager ` description.' unit: Percent Scratch Stall Rate: rst: The percent of :ref:`total shader-engine cycles ` in the @@ -707,7 +540,7 @@ Workgroup Manager - Resource Allocation: within the workgroup manager. This is expected to be always be zero on CDNA2 or newer accelerators (and small for previous accelerators). unit: Percent -Wavefront Launch Stats: +Wavefront launch stats: Grid Size: rst: The total number of work-items (or, threads) launched as a part of the kernel dispatch. In HIP, this is equivalent to the total grid size multiplied by the @@ -719,11 +552,10 @@ Wavefront Launch Stats: block size. unit: Work-Items Total Wavefronts: - rst: >- - The total number of wavefronts launched as part of the kernel dispatch. - On AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront - size is always 64 work-items. Thus, the total number of wavefronts should - be equivalent to the ceiling of grid size divided by 64. + rst: The total number of wavefronts launched as part of the kernel dispatch. On + AMD Instinct\u2122 CDNA\u2122 accelerators and GCN\u2122 GPUs, the wavefront + size is always 64 work-items. Thus, the total number of wavefronts should be + equivalent to the ceiling of grid size divided by 64. unit: Wavefronts Saved Wavefronts: rst: The total number of wavefronts saved at a context-save. See `cwsr_enable @@ -734,36 +566,32 @@ Wavefront Launch Stats: `_. unit: Wavefronts VGPRs: - rst: >- - The number of architected vector general-purpose registers allocated for the - kernel, see :ref:`VALU `. Note: this may not exactly match the - number of VGPRs requested by the compiler due to allocation granularity. + rst: 'The number of architected vector general-purpose registers allocated for + the kernel, see :ref:`VALU `. Note: this may not exactly match the + number of VGPRs requested by the compiler due to allocation granularity.' unit: VGPRs AGPRs: - rst: >- - The number of accumulation vector general-purpose registers allocated - for the kernel, see :ref:`AGPRs `. Note: this may not exactly match - the number of AGPRs requested by the compiler due to allocation granularity. + rst: 'The number of accumulation vector general-purpose registers allocated for + the kernel, see :ref:`AGPRs `. Note: this may not exactly match + the number of AGPRs requested by the compiler due to allocation granularity.' unit: AGPRs SGPRs: - rst: >- - The number of scalar general-purpose registers allocated for the kernel, see - :ref:`SALU `. Note: this may not exactly match the number of - SGPRs requested by the compiler due to allocation granularity. + rst: 'The number of scalar general-purpose registers allocated for the kernel, + see :ref:`SALU `. Note: this may not exactly match the number of + SGPRs requested by the compiler due to allocation granularity.' unit: SGPRs LDS Allocation: - rst: >- - The number of bytes of :doc:`LDS ` memory (or, shared memory) - allocated for this kernel. Note: This may also be larger than what was requested - at compile time due to both allocation granularity and dynamic per-dispatch - LDS allocations. + rst: 'The number of bytes of :doc:`LDS ` memory (or, shared + memory) allocated for this kernel. Note: This may also be larger than what was + requested at compile time due to both allocation granularity and dynamic per-dispatch + LDS allocations.' unit: Bytes per workgroup Scratch Allocation: rst: The number of bytes of :ref:`scratch memory ` requested per work-item for this kernel. Scratch memory is used for stack memory on the accelerator, as well as for register spills and restores. unit: Bytes per work-item -Wavefront Runtime Stats: +Wavefront runtime stats: Kernel Time: rst: The total duration of the executed kernel. unit: Nanoseconds @@ -775,11 +603,10 @@ Wavefront Runtime Stats: This is averaged over all wavefronts in a kernel dispatch. unit: Instructions per wavefront Wave Cycles: - rst: >- - The number of cycles a wavefront in the kernel dispatch spent resident - on a compute unit per :ref:`normalization unit `. This is - averaged over all wavefronts in a kernel dispatch. Note: this should not - be directly compared to the kernel cycles above. + rst: 'The number of cycles a wavefront in the kernel dispatch spent resident on + a compute unit per :ref:`normalization unit `. This is + averaged over all wavefronts in a kernel dispatch. Note: this should not be + directly compared to the kernel cycles above.' unit: Cycles per normalization unit Dependency Wait Cycles: rst: The number of cycles a wavefront in the kernel dispatch stalled waiting on @@ -813,12 +640,11 @@ Wavefront Runtime Stats: the total Wave Cycles metric. unit: Cycles per normalization unit Wavefront Occupancy: - rst: >- - The time-averaged number of wavefronts resident on the accelerator over the - lifetime of the kernel. Note: this metric may be inaccurate for short-running - kernels (less than 1ms). + rst: 'The time-averaged number of wavefronts resident on the accelerator over + the lifetime of the kernel. Note: this metric may be inaccurate for short-running + kernels (less than 1ms).' unit: Wavefronts -Overall Instruction Mix: +Overall instruction mix: VALU: rst: The total number of vector arithmetic logic unit (VALU) operations issued. These are the workhorses of the :doc:`compute unit `, and are @@ -854,7 +680,7 @@ Overall Instruction Mix: rst: The total number of branch operations issued. These typically consist of jump or branch operations and are used to implement control flow. unit: Instructions -VALU Arithmetic Instruction Mix: +VALU arithmetic instruction mix: INT32: rst: The total number of instructions operating on 32-bit integer operands issued to the VALU per :ref:`normalization unit `. @@ -915,53 +741,10 @@ VALU Arithmetic Instruction Mix: unit `. unit: Instructions per normalization unit Conversion: - rst: >- - The total number of type conversion instructions (such as converting data - to or from F32\u2194F64) issued to the VALU per :ref:`normalization unit - `. + rst: The total number of type conversion instructions (such as converting data + to or from F32\u2194F64) issued to the VALU per :ref:`normalization unit `. unit: Instructions per normalization unit -VMEM Instruction Mix: - Global/Generic Instr: - rst: The total number of global & generic memory instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Read: - rst: The total number of global & generic memory read instructions executed on - all :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Write: - rst: The total number of global & generic memory write instructions executed on - all :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Global/Generic Atomic: - rst: The total number of global & generic memory atomic (with and without return) - instructions executed on all :doc:`compute units ` on the accelerator, - per :ref:`normalization unit `. - unit: Instructions per normalization unit - Spill/Stack Instr: - rst: The total number of spill/stack memory instructions executed on all :doc:`compute - units ` on the accelerator, per :ref:`normalization unit `. - unit: Instructions per normalization unit - Spill/Stack Read: - rst: The total number of spill/stack memory read instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Spill/Stack Write: - rst: The total number of spill/stack memory write instructions executed on all - :doc:`compute units ` on the accelerator, per :ref:`normalization - unit `. - unit: Instructions per normalization unit - Spill/Stack Atomic: - rst: The total number of spill/stack memory atomic (with and without return) instructions - executed on all :doc:`compute units ` on the accelerator, per - :ref:`normalization unit `. Typically unused as these memory - operations are typically used to implement thread-local storage. - unit: Instructions per normalization unit -MFMA Arithmetic Instruction Mix: +MFMA instruction mix: MFMA-I8: rst: The total number of 8-bit integer :ref:`MFMA ` instructions issued per :ref:`normalization unit `. @@ -989,66 +772,53 @@ MFMA Arithmetic Instruction Mix: unit: Instructions per normalization unit Compute Speed-of-Light: VALU FLOPs: - rst: >- - The total floating-point operations executed per second on the :ref:`VALU - `. This is also presented as a percent of the peak theoretical - FLOPs achievable on the specific accelerator. Note: this does not include - any floating-point operations from :ref:`MFMA ` instructions. + rst: 'The total floating-point operations executed per second on the :ref:`VALU + `. This is also presented as a percent of the peak theoretical FLOPs + achievable on the specific accelerator. Note: this does not include any floating-point + operations from :ref:`MFMA ` instructions.' unit: GFLOPs VALU IOPs: - rst: >- - The total integer operations executed per second on the :ref:`VALU `. + rst: 'The total integer operations executed per second on the :ref:`VALU `. This is also presented as a percent of the peak theoretical IOPs achievable on the specific accelerator. Note: this does not include any integer operations - from :ref:`MFMA ` instructions. + from :ref:`MFMA ` instructions.' unit: GIOPs - MFMA FLOPs (F8): - rst: '' - unit: Unknown MFMA FLOPs (BF16): - rst: >- - The total number of 16-bit brain floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit brain floating - point operations from :ref:`VALU ` instructions. This is also - presented as a percent of the peak theoretical BF16 MFMA operations achievable - on the specific accelerator. + rst: 'The total number of 16-bit brain floating point :ref:`MFMA ` + operations executed per second. Note: this does not include any 16-bit brain + floating point operations from :ref:`VALU ` instructions. This is + also presented as a percent of the peak theoretical BF16 MFMA operations achievable + on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F16): - rst: >- - The total number of 16-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 16-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F16 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 16-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 16-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F16 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F32): - rst: >- - The total number of 32-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 32-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F32 MFMA operations achievable on the - specific accelerator. + rst: 'The total number of 32-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 32-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F32 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA FLOPs (F64): - rst: >- + rst: 'The total number of 64-bit floating point :ref:`MFMA ` operations + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator. The total number of 64-bit floating point :ref:`MFMA ` operations - executed per second. Note: this does not include any 64-bit floating point - operations from :ref:`VALU ` instructions. This is also presented - as a percent of the peak theoretical F64 MFMA operations achievable on the - specific accelerator. The total number of 64-bit floating point :ref:`MFMA - ` operations executed per second. Note: this does not include - any 64-bit floating point operations from :ref:`VALU ` instructions. - This is also presented as a percent of the peak theoretical F64 MFMA operations - achievable on the specific accelerator. + executed per second. Note: this does not include any 64-bit floating point operations + from :ref:`VALU ` instructions. This is also presented as a percent + of the peak theoretical F64 MFMA operations achievable on the specific accelerator.' unit: GFLOPs MFMA IOPs (INT8): - rst: >- - The total number of 8-bit integer :ref:`MFMA ` operations executed - per second. Note: this does not include any 8-bit integer operations from - :ref:`VALU ` instructions. This is also presented as a percent - of the peak theoretical INT8 MFMA operations achievable on the specific accelerator. + rst: 'The total number of 8-bit integer :ref:`MFMA ` operations executed + per second. Note: this does not include any 8-bit integer operations from :ref:`VALU + ` instructions. This is also presented as a percent of the peak theoretical + INT8 MFMA operations achievable on the specific accelerator.' unit: GFLOPs -Pipeline Statistics: +Pipeline statistics: IPC: rst: The ratio of the total number of instructions executed on the :doc:`CU ` over the :ref:`total active CU cycles `. @@ -1111,7 +881,7 @@ Pipeline Statistics: rst: The average number of round-trip cycles (that is, from issue to data return / acknowledgment) required for a SMEM instruction to complete. unit: Cycles -Arithmetic Operations: +Arithmetic operations: FLOPs (Total): rst: The total number of floating-point operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit @@ -1122,20 +892,16 @@ Arithmetic Operations: ` or :ref:`MFMA ` units, per :ref:`normalization unit `. unit: IOP per normalization unit - F8 OPs: - rst: '' - unit: Unknown F16 OPs: rst: The total number of 16-bit floating-point operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit `. unit: FLOP per normalization unit BF16 OPs: - rst: >- - The total number of 16-bit brain floating-point operations executed on - either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization - unit `. Note: on current CDNA accelerators, the VALU - has no native BF16 instructions. + rst: 'The total number of 16-bit brain floating-point operations executed on either + the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization + unit `. Note: on current CDNA accelerators, the VALU has + no native BF16 instructions.' unit: FLOP per normalization unit F32 OPs: rst: The total number of 32-bit floating-point operations executed on either the @@ -1148,11 +914,10 @@ Arithmetic Operations: unit `. unit: FLOP per normalization unit INT8 OPs: - rst: >- - The total number of 8-bit integer operations executed on either the :ref:`VALU + rst: 'The total number of 8-bit integer operations executed on either the :ref:`VALU ` or :ref:`MFMA ` units, per :ref:`normalization unit - `. Note: on current CDNA accelerators, the VALU has - no native INT8 instructions. + `. Note: on current CDNA accelerators, the VALU has no + native INT8 instructions.' unit: IOP per normalization unit LDS Speed-of-Light: Utilization: @@ -1182,16 +947,16 @@ LDS Speed-of-Light: amount of data in an uncontended access. [#lds-bank-conflict]_ unit: Percent LDS Statistics: - LDS Instructions: - rst: The total number of LDS instructions (including, but not limited to, read/write/atomics - and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. - unit: Instructions per normalization unit Theoretical Bandwidth: rst: Indicates the maximum amount of bytes that could have been loaded from, stored to, or atomically updated in the LDS divided by total duration. Does *not* take into account the execution mask of the wavefront when the instruction was executed. See the :ref:`LDS bandwidth example ` for more detail. unit: Gbps + LDS Instructions: + rst: The total number of LDS instructions (including, but not limited to, read/write/atomics + and HIP's ``__shfl`` instructions) executed per :ref:`normalization unit `. + unit: Instructions per normalization unit LDS Latency: rst: The average number of round-trip cycles (i.e., from issue to data-return acknowledgment) required for an LDS instruction to complete. @@ -1225,10 +990,9 @@ LDS Statistics: to stalls from non-dword aligned addresses per :ref:`normalization unit `. unit: Cycles per normalization unit Mem Violations: - rst: >- - The total number of out-of-bounds accesses made to the LDS, per :ref:`normalization - unit `. This is unused and expected to be zero in - most configurations for modern CDNA\u2122 accelerators. + rst: The total number of out-of-bounds accesses made to the LDS, per :ref:`normalization + unit `. This is unused and expected to be zero in most + configurations for modern CDNA\u2122 accelerators. unit: Accesses per normalization unit L1I Speed-of-Light: Bandwidth Utilization: @@ -1236,18 +1000,17 @@ L1I Speed-of-Light: theoretical bandwidth. Calculated as the ratio of L1I requests over the :ref:`total L1I cycles `. unit: Percent + L1I-L2 Bandwidth Utilization: + rst: The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth + achieved. Calculated as the ratio of the total number of requests from the L1I + to the L2 cache over the :ref:`total L1I-L2 interface cycles `. + unit: Percent +L1I cache accesses: Cache Hit Rate: rst: The percent of L1I requests that hit [#l1i-cache]_ on a previously loaded line the cache. Calculated as the ratio of the number of L1I requests that hit over the number of all L1I requests. unit: Percent - L1I-L2 Bandwidth Utilization: - rst: >- - The percent of the peak theoretical L1I \u2192 L2 cache request bandwidth - achieved. Calculated as the ratio of the total number of requests from - the L1I to the L2 cache over the :ref:`total L1I-L2 interface cycles `. - unit: Percent -L1I cache accesses: Req: rst: The total number of requests made to the L1I per normalization-unit unit: Requests per normalization unit @@ -1265,11 +1028,6 @@ L1I cache accesses: already pending due to another request, per :ref:`normalization-unit `. See note in :ref:`desc-l1i-sol` for more detail. unit: Requests per normalization unit - Cache Hit Rate: - rst: The percent of L1I requests that hit [#l1i-cache]_ on a previously loaded - line the cache. Calculated as the ratio of the number of L1I requests that hit - over the number of all L1I requests. - unit: Percent Instruction Fetch Latency: rst: The average number of cycles spent to fetch instructions to a :doc:`CU `. unit: Cycles @@ -1284,17 +1042,17 @@ Scalar L1D Speed-of-Light: theoretical bandwidth. Calculated as the ratio of sL1D requests over the :ref:`total sL1D cycles `. unit: Percent + sL1D-L2 BW Utilization: + rst: The percentage of the peak theoretical sL1D - L2 interface bandwidth acheived. + Calculated as total number of bytes read from, written to, or atomically updated + across the sL1D - L2 interface. + unit: Percent +Scalar L1D cache accesses: Cache Hit Rate: rst: Indicates the percent of sL1D requests that hit on a previously loaded line the cache. The ratio of the number of sL1D requests that hit [#sl1d-cache]_ over the number of all sL1D requests. unit: Percent - sL1D-L2 BW Utilization: - rst: The percentage of the peak theoretical sL1D - L2 interface bandwidth acheived. - Caclulated as total number of bytes read from, written to, or atomically updated - across the sL1D - L2 interface. - unit: Percent -Scalar L1D cache accesses: Req: rst: The total number of requests, of any size or type, made to the sL1D per :ref:`normalization unit `. @@ -1313,20 +1071,10 @@ Scalar L1D cache accesses: already pending due to another request, per :ref:`normalization unit `. See :ref:`desc-sl1d-sol` for more detail. unit: Requests per normalization unit - Cache Hit Rate: - rst: Indicates the percent of sL1D requests that hit on a previously loaded line - the cache. The ratio of the number of sL1D requests that hit [#sl1d-cache]_ - over the number of all sL1D requests. - unit: Percent Read Req (Total): rst: The total number of sL1D read requests of any size, per :ref:`normalization unit `. unit: Requests per normalization unit - Atomic Req: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit Read Req (1 DWord): rst: The total number of sL1D read requests made for a single dword of data (4B), per :ref:`normalization unit `. @@ -1349,13 +1097,17 @@ Scalar L1D cache accesses: unit: Requests per normalization unit Scalar L1D Cache - L2 Interface: sL1D-L2 BW: - rst: >- - The total number of bytes read from, written to, or atomically updated - across the sL1D\u2194:doc:`L2 ` interface, divided by total duration. - Note that sL1D writes and atomics are typically - unused on current CDNA accelerators, so in the majority of cases this can - be interpreted as an sL1D\u2192L2 read bandwidth. + rst: The total number of bytes read from, written to, or atomically updated across + the sL1D\u2194:doc:`L2 ` interface, divided by total duration. Note + that sL1D writes and atomics are typically unused on current CDNA accelerators, + so in the majority of cases this can be interpreted as an sL1D\u2192L2 read + bandwidth. unit: Gbps + Atomic Req: + rst: The total number of atomic requests from sL1D to the :doc:`L2 `, + per :ref:`normalization unit `. Typically unused on current + CDNA accelerators. + unit: Requests per normalization unit Read Req: rst: The total number of read requests from sL1D to the :doc:`L2 `, per :ref:`normalization unit `. @@ -1365,17 +1117,11 @@ Scalar L1D Cache - L2 Interface: per :ref:`normalization unit `. Typically unused on current CDNA accelerators. unit: Requests per normalization unit - Atomic Req: - rst: The total number of atomic requests from sL1D to the :doc:`L2 `, - per :ref:`normalization unit `. Typically unused on current - CDNA accelerators. - unit: Requests per normalization unit Stall Cycles: - rst: >- - The total number of cycles the sL1D\u2194 :doc:`L2 ` interface + rst: The total number of cycles the sL1D\u2194 :doc:`L2 ` interface was stalled, per :ref:`normalization unit `. unit: Cycles per normalization unit -Busy and stall metrics: +Busy / stall metrics: Address Processing Unit Busy: rst: Percent of the :ref:`total CU cycles ` the address processor was busy @@ -1388,19 +1134,10 @@ Busy and stall metrics: rst: Percent of the :ref:`total CU cycles ` the address processor was stalled from sending write/atomic data further into the vL1D pipeline unit: Percent - Data-Processor → Address Stall: + "Data-Processor \u2192 Address Stall": rst: Percent of :ref:`total CU cycles ` the address processor was stalled waiting to send command data to the :ref:`data processor ` unit: Percent - Sequencer → TA Address Stall: - rst: '' - unit: Unknown - Sequencer → TA Command Stall: - rst: '' - unit: Unknown - Sequencer → TA Data Stall: - rst: '' - unit: Unknown Instruction counts: Total Instructions: rst: The total number of memory instructions executed by the address processer @@ -1446,7 +1183,7 @@ Instruction counts: :ref:`normalization unit `. Typically unused as these memory operations are typically used to implement thread-local storage. unit: Instructions per normalization unit -Spill and stack metrics: +Spill / stack metrics: Spill/Stack Total Cycles: rst: The number of cycles the address processing unit spent working on spill/stack instructions, per :ref:`normalization unit `. @@ -1464,11 +1201,11 @@ Vector L1 data-return path or Texture Data (TD): rst: Percent of the :ref:`total CU cycles ` the data-return unit was busy processing or waiting on data to return to the :doc:`CU `. unit: Percent - Cache RAM → Data-Return Stall: + "Cache RAM \u2192 Data-Return Stall": rst: Percent of the :ref:`total CU cycles ` the data-return unit was stalled on data to be returned from the :ref:`vL1D Cache RAM `. unit: Percent - Workgroup manager → Data-Return Stall: + "Workgroup manager \u2192 Data-Return Stall": rst: Percent of the :ref:`total CU cycles ` the data-return unit was stalled by the :ref:`workgroup manager ` due to initialization of registers as a part of launching new workgroups. @@ -1525,32 +1262,6 @@ vL1D Speed-of-Light: generated per instruction divided by the ideal number of thread-requests per instruction. unit: Percent -vL1D cache stall metrics: - Stalled on L2 Data: - rst: The ratio of the number of cycles where the vL1D is stalled waiting for requested - data to return from the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Stalled on L2 Req: - rst: The ratio of the number of cycles where the vL1D is stalled waiting to issue - a request for data to the :doc:`L2 cache ` divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Read): - rst: The ratio of the number of cycles where the vL1D is stalled due to Read requests - with conflicting tags being looked up concurrently, divided by the number of - cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Write): - rst: The ratio of the number of cycles where the vL1D is stalled due to Write - requests with conflicting tags being looked up concurrently, divided by the - number of cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent - Tag RAM Stall (Atomic): - rst: The ratio of the number of cycles where the vL1D is stalled due to Atomic - requests with conflicting tags being looked up concurrently, divided by the - number of cycles where the vL1D is active [#vl1d-activity]_. - unit: Percent vL1D cache access metrics: Total Req: rst: The total number of incoming requests from the :ref:`address processing unit @@ -1615,79 +1326,6 @@ vL1D cache access metrics: cache `, per :ref:`normalization unit `. This includes requests for atomics with, and without return. unit: Requests per normalization unit -L1D - L2 Transactions: - NC - Read: - rst: Total read requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Read: - rst: Total read requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Read: - rst: Total read requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Read: - rst: Total read requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Write: - rst: Total write requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - NC - Write: - rst: Total write requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Write: - rst: Total write requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Write: - rst: Total write requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - NC - Atomic: - rst: Total atomic requests with NC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - UC - Atomic: - rst: Total atomic requests with UC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - CC - Atomic: - rst: Total atomic requests with CC mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit - RW - Atomic: - rst: Total atomic requests with RW mtype from this TCP to all TCCs Sum over TCP - instances per normalization unit. - unit: Requests per normalization unit -L1 Unified Translation Cache (UTCL1): - Req: - rst: The number of translation requests made to the UTCL1 per normalization unit. - unit: Requests per normalization unit - Hit Ratio: - rst: The ratio of the number of translation requests that hit in the UTCL1 divided - by the total number of translation requests made to the UTCL1. - unit: Percent - Hits: - rst: The number of translation requests that hit in the UTCL1, and could be reused, - per normalization unit. - unit: Requests per normalization unit - Translation Misses: - rst: The total number of translation requests that missed in the UTCL1 due to - translation not being present in the cache, per :ref:`normalization unit `. - unit: unit - Permission Misses: - rst: >- - The total number of translation requests that missed in the UTCL1 due - to a permission error, per :ref:`normalization unit `. - This is unused and expected to be zero in most configurations for modern - CDNA\u2122 accelerators. - unit: Requests per normalization unit -L1D Addr Translation Stalls: {} L2 Speed-of-Light: Utilization: rst: The ratio of the :ref:`number of cycles an L2 channel was active, summed @@ -1809,7 +1447,7 @@ L2-Fabric interface metrics: before a completion acknowledgement (atomic without return value) or data (atomic with return value) was returned to the L2. unit: Cycles -L2 Cache Accesses: +L2 cache accesses: Bandwidth: rst: The number of bytes looked up in the L2 cache, divided by total duration. The number of bytes is calculated as the number of cache lines requested multiplied @@ -1900,7 +1538,6 @@ L2 Cache Accesses: rst: The total number of requests to the L2 that go to Read-Write coherent memory (RW) allocations. See the :ref:`memory-type` for more information. unit: Requests per normalization unit -L2 Cache Stalls: {} L2 - Fabric Interface stalls: Write - Credit Starvation: rst: The number of cycles the L2-Fabric interface was stalled on write or atomic @@ -1918,9 +1555,6 @@ L2 - Fabric interface detailed metrics: any memory location, per :ref:`normalization unit `. See :ref:`l2-request-flow` for more detail. unit: Requests per normalization unit - Read (128B): - rst: '' - unit: Unknown Read (Uncached): rst: The total number of L2 requests to Infinity Fabric to read :ref:`uncached data ` from any memory location, per :ref:`normalization unit `. @@ -1972,72 +1606,3 @@ L2 - Fabric interface detailed metrics: as :ref:`fine-grained memory ` allocations or :ref:`uncached memory ` allocations on the MI2XX. unit: Requests per normalization unit -Aggregate Stats (All channels): - L2 Cache Hit Rate: - rst: The total number of requests to the L2 from all clients that hit in the cache. - As noted in the :ref:`Speed-of-Light ` section, this includes hit-on-miss - requests. - unit: Percent -L2 Cache Hit Rate (pct): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2 Requests (per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Requests (per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Read Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Write and Atomic Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Atomic Latency (Cycles): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Read Stall (Cycles per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric Write and Atomic Stall (Cycles per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown -L2-Fabric (128B read requests per normUnit): - ::_1: - rst: '' - unit: Unknown - placeholder_range: - rst: '' - unit: Unknown