* Add dlmalloc_strat allocator strategy
- Use mspace variant to ease encapsulation
- Make pow2bins and dlmalloc cmake selectable
* Add unit tester for dlmalloc, rework single_heap, pow2bins unit testers
accordingly
- add dlmalloc get_used/get_avail, and have all strats allocators also have a get_used
- Rework memallocator unit tests: bin size is per strat, alignment is verified in singleheap
* bugfix: dlmalloc exposed that the pingpong test would write past end of
allocation with -w 32
* iostream leakage/mixed usage of cerr and fprintf(stderr
---------
Signed-off-by: Aurelien Bouteiller <aurelien.bouteiller@amd.com>
* Remove dev_mono_linear (followup to removal of slab_heap)
* cleanup: use CHECK_HIP rather than ad-hoc error checking
---------
Signed-off-by: Aurelien Bouteiller <aurelien.bouteiller@amd.com>
* Update backend to use provided MPI communicator during library initialization, default to `MPI_COMM_WORLD`
* Update `rocshmem_my_pe` and `rocshmem_n_pes` host APIs
- Return values from backend if initialized; otherwise, fallback to MPI_Singleton.
* Remove unused forward_list
* Remove unused __read_clock function
* Replace wallClk code with hip function
* Remove unused unit test for ipc
* Remove slab heap
* Remove unused EBO spinlock
* Update HIP version check for compatibility with versions >= 5.5
* Update memory allocator for context BlockHandle
- Replaced `HIPAllocator` with `HIPDefaultFinegrainedAllocator` for context `BlockHandle`.
* Update run commands for `rocshmem_g` and `rocshmem_p` functional tests
* Update(DeviceProxy): Dynamically Determine Memory Allocation Size & Remove Compile-Time size Calculations
- Modified the Device proxy class to determine memory allocation size at runtime.
- Updated all classes that include the Device proxy to use dynamic memory allocation.
- Removed compile-time memory size calculations.
- Ensured the allocated number of backend queue data structures matches the number of RO device contexts.