Support MI 350 profiling (#632)
* Add MI 350 hardware information
* Refactor MI GPU YAML file and corresponding interface
* Add SoC file for gfx950 architecture
* Add analysis report configs for MI 350 containing existing metrics
* Add placeholder None valued metrics for previous architectures to make
baseline comparison work
* Enable testing on MI 350
* Analysis config metric changes
- SPI changes
- Update metric formula for default SPI pipe counter
- Use efficiently collected pipe wise SPI counters
- Add SPI Wave Occupancy
- Add Scheduler-Pipe Wave Utilization
- Update formula for VGPR Writes
- Add Scheduler-Pipe FIFO Full Rate
- CPC changes
- Add CPC SYNC FIFO Full Rate
- Add CPC CANE Stall Rate
- Add CPC ADC Utilization
- SQ changes
- Add VALU co-issue efficiency
- Add F6F4 datatype metrics
- Update formula for total FLOPs by adding F6F4 counters
- Add LDS STORE / LOAD / ATOMIC metrics
- Add LDS STORE / LOAD / ATOMIC bandwidth
- Add LDS FIFO and TA ADDR / CMD / DATA FIFO full rates
* Collect TCP_TCP_LATENCY_sum only for gfx950 (MI 350)
* Do not inject SQ_ACCUM_PREV_HIRES unnecesarily
* Do not hardcode memory and shader clock speeds
* Write num_hbm_channels to sysinfo.csv instead of hbm_bw while profiling
* Move generate sysinfo.csv to pre processing step of profiling
* Add warnings to use --specs-correction for missing sysinfo.csv values during analysis phase
* Update CHANGELOG
* Analysis phase warning to use --specs-correction when needed
[ROCm/rocprofiler-compute commit: f9aa7be97c]
This commit is contained in:
@@ -22,17 +22,33 @@ Full documentation for ROCm Compute Profiler is available at [https://rocm.docs.
|
||||
|
||||
* Support host-trap PC Sampling on CLI (beta version)
|
||||
|
||||
* Add support for tuned performance counters for gfx950 GPUs
|
||||
* Add L1 latencies
|
||||
* Add L2 latencies
|
||||
* Add L2 to EA stalls
|
||||
* Add L2 to EA stalls per channel
|
||||
* Support for AMD Instinct MI350 series GPUs with the addition of the following counters:
|
||||
* VALU co-issue (Two VALUs are issued instructions) efficiency
|
||||
* Stream Processor Instruction (SPI) Wave Occupancy
|
||||
* Scheduler-Pipe Wave Utilization
|
||||
* Scheduler FIFO Full Rate
|
||||
* CPC ADC Utilization
|
||||
* F6F4 datatype metrics
|
||||
* Update formula for total FLOPs while taking into account F6F4 ops
|
||||
* LDS STORE, LDS LOAD, LDS ATOMIC instruction count metrics
|
||||
* LDS STORE, LDS LOAD, LDS ATOMIC bandwidth metrics
|
||||
* LDS FIFO full rate
|
||||
* Sequencer -> TA ADDR Stall rates
|
||||
* Sequencer -> TA CMD Stall rates
|
||||
* Sequencer -> TA DATA Stall rates
|
||||
* L1 latencies
|
||||
* L2 latencies
|
||||
* L2 to EA stalls
|
||||
* L2 to EA stalls per channel
|
||||
|
||||
### Changed
|
||||
|
||||
* Change normal_unit default to per_kernel
|
||||
* Change dependency from rocm-smi to amd-smi
|
||||
* Decrease profiling time by not collecting counters not used in post analysis
|
||||
* Update definition of following metrics for MI 350:
|
||||
* VGPR Writes
|
||||
* Total FLOPs (consider fp6 and fp4 ops)
|
||||
|
||||
### Resolved issues
|
||||
|
||||
@@ -44,6 +60,14 @@ Full documentation for ROCm Compute Profiler is available at [https://rocm.docs.
|
||||
|
||||
* GPU id filtering is not supported when using rocprof v3
|
||||
|
||||
* Analysis of previously collected workload data will not work due to sysinfo.csv schema change
|
||||
* As a workaround, run the profiling operation again for the workload and interrupt the process after ten seconds.
|
||||
Followed by copying the `sysinfo.csv` file from the new data folder to the old one.
|
||||
This assumes your system specification hasn't changed since the creation of the previous workload data.
|
||||
|
||||
* Analysis of new workloads might require providing shader/memory clock speed using
|
||||
--specs-correction operation if `amd-smi` or `rocminfo` does not provide clock speeds.
|
||||
|
||||
## ROCm Compute Profiler 3.1.0 for ROCm 6.4.0
|
||||
|
||||
### Added
|
||||
|
||||
Reference in New Issue
Block a user