[rocprofiler-compute] Remove TCP_TCP_LATENCY_sum counter for MI300 (#2174)

* Remove TCP_TCP_LATENCY_sum counter for MI300

* Remove TCP_TCP_LATENCY_sum counter which is unsupported for MI300 per register specification

* Remove VL1 Lat metric from memory chart section (block 3) for MI 300
  since it uses TCP_TCP_LATENCY_sum counter which is unsupported

* Remove references to TCP_TCP_LATENCY_sum

* Update CHANGELOG

* reword changelog
This commit is contained in:
vedithal-amd
2025-12-10 09:41:46 -05:00
کامیت شده توسط GitHub
والد 9d34098350
کامیت 252a5e8146
10فایلهای تغییر یافته به همراه1012 افزوده شده و 2330 حذف شده
@@ -41,6 +41,10 @@ Full documentation for ROCm Compute Profiler is available at [https://rocm.docs.
* Corrected peak VALU Roofline profiling and analysis by removing `FP8` VALU and `BF16` VALU benchmarking.
### Removed
* Removed "VL1 Lat" metric for AMD Instinct MI300 series GPUs, due to MI300 series not supporting TCP_TCP_LATENCY_sum counter.
## ROCm Compute Profiler 3.4.0 for ROCm 7.2.0
### Added
@@ -562,12 +562,8 @@ Analysis database example
DEBUG Applied analysis mode filters
DEBUG Calculated dispatch data
DEBUG Collected metrics data
WARNING Failed to evaluate expression for 3.1.25 - Value: to_round(to_avg(
(pmc_df.get("TCP_TCP_LATENCY_sum") / pmc_df.get("TCP_TA_TCP_STATE_READ_sum")).where((pmc_df.get("TCP_TA_TCP_STATE_READ_sum") != 0), None)), 0) - unsupported operand type(s) for /: 'NoneType' and 'float'
WARNING Failed to evaluate expression for 3.1.39 - Value: to_round((to_avg(
(pmc_df.get("pmc_perf_ACCUM") / pmc_df.get("SQC_ICACHE_REQ")).where((pmc_df.get("SQC_ICACHE_REQ") != 0), None)) * 100), 0) - unsupported operand type(s) for /: 'NoneType' and 'float'
WARNING Failed to evaluate expression for 3.1.25 - Value: to_round(to_avg(
(pmc_df.get("TCP_TCP_LATENCY_sum") / pmc_df.get("TCP_TA_TCP_STATE_READ_sum")).where((pmc_df.get("TCP_TA_TCP_STATE_READ_sum") != 0), None)), 0) - unsupported operand type(s) for /: 'NoneType' and 'float'
WARNING Failed to evaluate expression for 3.1.39 - Value: to_round((to_avg(
(pmc_df.get("pmc_perf_ACCUM") / pmc_df.get("SQC_ICACHE_REQ")).where((pmc_df.get("SQC_ICACHE_REQ") != 0), None)) * 100), 0) - unsupported operand type(s) for /: 'NoneType' and 'float'
DEBUG Calculated metric values
@@ -70,9 +70,6 @@ Panel Config:
+ TCP_TCC_ATOMIC_WITH_RET_REQ_sum) + TCP_TCC_ATOMIC_WITHOUT_RET_REQ_sum))
/ TCP_TOTAL_CACHE_ACCESSES_sum)) if (TCP_TOTAL_CACHE_ACCESSES_sum != 0)
else None )), 0)
VL1 Lat:
value: ROUND(AVG(((TCP_TCP_LATENCY_sum / TCP_TA_TCP_STATE_READ_sum) if (TCP_TA_TCP_STATE_READ_sum
!= 0) else None)), 0)
VL1 Coalesce:
value: ROUND(AVG(((((TA_TOTAL_WAVEFRONTS_sum * 64) * 100) / (TCP_TOTAL_ACCESSES_sum
* 4)) if (TCP_TOTAL_ACCESSES_sum != None) else 0)), 0)
@@ -197,8 +194,6 @@ Panel Config:
unit after coalescing per normalization unit
VL1 Hit: The ratio of the number of vL1D cache line requests that hit in vL1D
cache over the total number of cache line requests to the vL1D Cache RAM.
VL1 Lat: Calculated as the average number of cycles that a vL1D cache line request
spent in the vL1D cache pipeline.
VL1 Coalesce: Indicates how well memory instructions were coalesced by the address
processing unit, ranging from uncoalesced (25%) to fully coalesced (100%). Calculated
as the average number of thread-requests generated per instruction divided by
@@ -70,9 +70,6 @@ Panel Config:
+ TCP_TCC_ATOMIC_WITH_RET_REQ_sum) + TCP_TCC_ATOMIC_WITHOUT_RET_REQ_sum))
/ TCP_TOTAL_CACHE_ACCESSES_sum)) if (TCP_TOTAL_CACHE_ACCESSES_sum != 0)
else None )), 0)
VL1 Lat:
value: ROUND(AVG(((TCP_TCP_LATENCY_sum / TCP_TA_TCP_STATE_READ_sum) if (TCP_TA_TCP_STATE_READ_sum
!= 0) else None)), 0)
VL1 Coalesce:
value: ROUND(AVG(((((TA_TOTAL_WAVEFRONTS_sum * 64) * 100) / (TCP_TOTAL_ACCESSES_sum
* 4)) if (TCP_TOTAL_ACCESSES_sum != None) else 0)), 0)
@@ -197,8 +194,6 @@ Panel Config:
unit after coalescing per normalization unit
VL1 Hit: The ratio of the number of vL1D cache line requests that hit in vL1D
cache over the total number of cache line requests to the vL1D Cache RAM.
VL1 Lat: Calculated as the average number of cycles that a vL1D cache line request
spent in the vL1D cache pipeline.
VL1 Coalesce: Indicates how well memory instructions were coalesced by the address
processing unit, ranging from uncoalesced (25%) to fully coalesced (100%). Calculated
as the average number of thread-requests generated per instruction divided by
@@ -70,9 +70,6 @@ Panel Config:
+ TCP_TCC_ATOMIC_WITH_RET_REQ_sum) + TCP_TCC_ATOMIC_WITHOUT_RET_REQ_sum))
/ TCP_TOTAL_CACHE_ACCESSES_sum)) if (TCP_TOTAL_CACHE_ACCESSES_sum != 0)
else None )), 0)
VL1 Lat:
value: ROUND(AVG(((TCP_TCP_LATENCY_sum / TCP_TA_TCP_STATE_READ_sum) if (TCP_TA_TCP_STATE_READ_sum
!= 0) else None)), 0)
VL1 Coalesce:
value: ROUND(AVG(((((TA_TOTAL_WAVEFRONTS_sum * 64) * 100) / (TCP_TOTAL_ACCESSES_sum
* 4)) if (TCP_TOTAL_ACCESSES_sum != None) else 0)), 0)
@@ -197,8 +194,6 @@ Panel Config:
unit after coalescing per normalization unit
VL1 Hit: The ratio of the number of vL1D cache line requests that hit in vL1D
cache over the total number of cache line requests to the vL1D Cache RAM.
VL1 Lat: Calculated as the average number of cycles that a vL1D cache line request
spent in the vL1D cache pipeline.
VL1 Coalesce: Indicates how well memory instructions were coalesced by the address
processing unit, ranging from uncoalesced (25%) to fully coalesced (100%). Calculated
as the average number of thread-requests generated per instruction divided by
@@ -344,10 +344,6 @@ class OmniSoC_Base:
"""Filter default performance counter set based on user arguments"""
counters, filter_blocks = self.detect_counters()
# TCP_TCP_LATENCY_sum not supported for MI300 (gfx940, gfx941, gfx942)
if self.__arch in ("gfx940", "gfx941", "gfx942"):
counters = counters - {"TCP_TCP_LATENCY_sum"}
# SQ_ACCUM_PREV_HIRES will be injected for level counters later on
counters = counters - {"SQ_ACCUM_PREV_HIRES"}
@@ -52,7 +52,7 @@
"0000_top_stats.yaml": "2819d96f5b1c3704f2ac50868a246a7f",
"0100_system_info.yaml": "cefae2b10db8cf4b0d3a971cff5e82c8",
"0200_system_speed_of_light.yaml": "d5df0a2b701972fa08ba0a44e72aa752",
"0300_memory_chart.yaml": "40dc04c73c3cec3d0a93e26d2db8c6f3",
"0300_memory_chart.yaml": "0a57cdf55be606799ee8d7b42a993027",
"0400_roofline.yaml": "d4650e008f2e3a7d28871e8518153575",
"0500_command_processor_cpc_cpf.yaml": "d8f424ec3fcfa4b2fcee2ad5e6456531",
"0600_workgroup_manager_spi.yaml": "8b6a89de516bed5821a9849627ad634a",
@@ -75,7 +75,7 @@
"0000_top_stats.yaml": "2819d96f5b1c3704f2ac50868a246a7f",
"0100_system_info.yaml": "cefae2b10db8cf4b0d3a971cff5e82c8",
"0200_system_speed_of_light.yaml": "0ccc1a63ebe11079832741c6d86ec3aa",
"0300_memory_chart.yaml": "40dc04c73c3cec3d0a93e26d2db8c6f3",
"0300_memory_chart.yaml": "0a57cdf55be606799ee8d7b42a993027",
"0400_roofline.yaml": "c066a19bc0e00e692c34998e44c62387",
"0500_command_processor_cpc_cpf.yaml": "d8f424ec3fcfa4b2fcee2ad5e6456531",
"0600_workgroup_manager_spi.yaml": "8b6a89de516bed5821a9849627ad634a",
@@ -98,7 +98,7 @@
"0000_top_stats.yaml": "2819d96f5b1c3704f2ac50868a246a7f",
"0100_system_info.yaml": "cefae2b10db8cf4b0d3a971cff5e82c8",
"0200_system_speed_of_light.yaml": "d5df0a2b701972fa08ba0a44e72aa752",
"0300_memory_chart.yaml": "40dc04c73c3cec3d0a93e26d2db8c6f3",
"0300_memory_chart.yaml": "0a57cdf55be606799ee8d7b42a993027",
"0400_roofline.yaml": "318c3e774d41a639628a7f72c2462375",
"0500_command_processor_cpc_cpf.yaml": "d8f424ec3fcfa4b2fcee2ad5e6456531",
"0600_workgroup_manager_spi.yaml": "8b6a89de516bed5821a9849627ad634a",
تفاوت فایلی نمایش داده نمی شود زیرا این فایل بسیار بزرگ است Diff را بارگزاری کن
تفاوت فایلی نمایش داده نمی شود زیرا این فایل بسیار بزرگ است Diff را بارگزاری کن
تفاوت فایلی نمایش داده نمی شود زیرا این فایل بسیار بزرگ است Diff را بارگزاری کن