SWDEV-470612 - Add the optimized multistream path

- Added the optimized multi stream path in graph execution. It uses a fixed number of async streams in the execution
- Optimize the launch latency, where commands
creation and execution is done at the same time
- Optimize the scheduling to use less barriers and waiting signals if
the same queue  can be detected
- The new path is controlled by  DEBUG_HIP_FORCE_GRAPH_QUEUES
environment variable, where 0 will use the original path and any other
value will force the number of asynchronous queues for execution
- DEBUG_HIP_FORCE_ASYNC_QUEUE can force single queue async
execution in graphs(applicable for Navi families only)

Change-Id: I7eb40bc15c45f508d6911868a6f6d4c3598d380e


[ROCm/clr commit: 9db52f9a46]
This commit is contained in:
German Andryeyev
2024-07-29 11:08:51 -04:00
parent 31927fefd6
commit 35c7a87014
6 ha cambiato i file con 367 aggiunte e 35 eliminazioni
@@ -1384,6 +1384,9 @@ bool VirtualGPU::initPool(size_t kernarg_pool_size) {
roc_device_.info().largeBar_) {
kernarg_pool_base_ =
reinterpret_cast<address>(roc_device_.deviceLocalAlloc(kernarg_pool_size_));
// @note Workaround first access penalty.
// KFD may update CPU page tables on the first CPU access
*kernarg_pool_base_ = 0;
} else {
kernarg_pool_base_ = reinterpret_cast<address>(roc_device_.hostAlloc(kernarg_pool_size_, 0,
Device::MemorySegment::kKernArg));