Commit gráf

39 Commit-ok

Szerző SHA1 Üzenet Dátum
Aaron Enye Shi 890beb81d6 Guard rcp rounded implementation as well
Since rcp implementations of non-default rounded versions are not correct or supported in OCML, guard them using the same macro OCML_BASIC_ROUNDED_OPERATIONS. Also update the docs and tests.


[ROCm/clr commit: 7b3bbc85c5]
2018-11-06 19:53:28 +00:00
Aaron Enye Shi 4480bb6d06 Guard the OCML rounded operations instead
Instead of commenting all these functions out, guard the functions with a macro OCML_BASIC_ROUNDED_OPERATIONS.


[ROCm/clr commit: 9aa92238ab]
2018-11-06 16:32:14 +00:00
Aaron Enye Shi 5c1dc7a071 Remove non-working non-default-rounded math apis
In ROCm-Device-Libs, they have dropped the non-default-rounded versions of add, sub, mul, div, sqrt and fma. Therefore, ocml has removed the rte, rtp, rtn, and rtz counterparts. This will remove the same math APIs in HIP for _ru, _rd, _rn, and _rz.


[ROCm/clr commit: cef6e8ef1f]
2018-11-05 22:34:16 +00:00
Maneesh Gupta b859ab46df Merge pull request #705 from ROCm-Developer-Tools/feature_minimal_changes_for_hc_next
Feature minimal changes for hc next

[ROCm/clr commit: 407e092a13]
2018-10-19 06:58:31 +05:30
Alex Voicu 35e9dfc593 Dumb workaround is still needed, so add it back.
[ROCm/clr commit: 3678063598]
2018-10-18 15:33:46 +01:00
Alex Voicu 50388a3bd1 Re-sync with upstream.
[ROCm/clr commit: 9ec697c620]
2018-10-10 11:43:49 +01:00
Aaron Enye Shi 4d59fbbcbe Use sinf and cosf from ocml device libs
Using llvm_amdgcn builtin fails to produce accurate values, we should move to using the ocml device library versions.


[ROCm/clr commit: eeb6c11050]
2018-09-25 19:31:39 +00:00
Maneesh Gupta f648d28bcb Merge pull request #614 from ROCm-Developer-Tools/fma
Add overloading resolution functions for fma

[ROCm/clr commit: 7cec85e1ae]
2018-09-20 13:38:03 +05:30
Yaxun Sam Liu 2f6bee46f9 Silent warnings about duplicate static keyword
static is already in __DEVICE__, so should be removed.


[ROCm/clr commit: cfa71293ce]
2018-09-19 10:39:45 -04:00
Yaxun Sam Liu a7b03bee08 Add fma function with float and _Float16 arguments
[ROCm/clr commit: a8bed200b7]
2018-09-19 09:59:33 -04:00
Yaxun Sam Liu 445697cfe3 Fix build failure of hipTestHalf and hipTestIncludeMath for hip-clang
[ROCm/clr commit: 781f0d97b5]
2018-09-18 21:00:15 -04:00
Alex Voicu f5d204450c Align with HC Next.
[ROCm/clr commit: 26c26f63cb]
2018-09-17 11:50:29 +03:00
Aaron Enye Shi 70d1143f62 Fix Tensorflow ambiguous min issue
[ROCm/clr commit: 6fa3ba3e12]
2018-09-13 23:16:20 +00:00
Aaron Enye Shi 96fe74cd2b Avoid AMP-retrict call to CPU-restrict
[ROCm/clr commit: ebf565819e]
2018-09-12 14:54:31 +00:00
Aaron Enye Shi 031932e509 Avoid host min func conflict with gcc min
[ROCm/clr commit: 7aa282897f]
2018-09-11 18:48:31 +00:00
Aaron Enye Shi 17d6266150 Use templates for min to prevent ambiguity
[ROCm/clr commit: 5f0838300e]
2018-09-11 18:21:54 +00:00
Yaxun Sam Liu d3295e61e2 Fix __HIP_ARCH_* not defined after including math_functions.h
hcc_detail/math_functions.h used to include hcc_detail/hip_runtime.h.

Removing it has caused regression in TensorFlow 1.8.

Put it back for backward compatibiliity.


[ROCm/clr commit: 87de95975a]
2018-08-08 08:55:28 -04:00
Yaxun Sam Liu 6df5aef807 Fix declaration conflict when hip/math_functions.h is included first
This fixes build failure in TensorFlow 1.8 for HCC


[ROCm/clr commit: 69bbf45b44]
2018-08-07 15:44:59 -04:00
Yaxun (Sam) Liu dc1257a16c Support std::complex for hip-clang
[ROCm/clr commit: a8dc1257df]
2018-07-18 00:08:04 -04:00
Alex Voicu 4fcfb64a88 Removes use of unimplemented OCML functionality.
[ROCm/clr commit: 9452298fbe]
2018-06-25 19:16:27 +01:00
Alex Voicu 8d86f5c7aa Rename for minimal confusion.
[ROCm/clr commit: 3a17e2ad06]
2018-06-01 22:55:33 +01:00
Alex Voicu 4306856ec9 Fix typos / address review comments.
[ROCm/clr commit: 91d9ec75d7]
2018-06-01 16:20:21 +01:00
Alex Voicu a9cabfc8df Re-sync with upstream. Add integer abs.
[ROCm/clr commit: e03ca1a72e]
2018-05-31 16:38:00 +01:00
Alex Voicu 2924783ee4 Switch to using ROCDL directly, as opposed to via HC. Add missing bits.
[ROCm/clr commit: 14e6a04387]
2018-05-31 03:17:26 +01:00
Jorghi12 8a2d7a25b1 Update math_functions.h
CUDA also has a function named labs.

[ROCm/clr commit: b33876ee4d]
2018-05-26 16:22:10 -04:00
Jorghi12 43d55a79e4 Adding double/long int signatures for abs
Adding overloads for abs that are found in cuda's math_functions.

[ROCm/clr commit: b127488ea9]
2018-05-26 00:41:24 -04:00
Deven Desai 13ab5368d0 Checkin to fix bugs in math functions.
This change fixes the following bugs that were discovered while debuggnig TF unit test failures (cwise_ops_test)

1. __hisinf and __hisnan routines
   Both had incorrect implementations.

2. abs
   A "long long" (64bit int) version was missing, resulting in the 32bit version being used for 64bit ints (which resulted in incorrect results, when the value passed in was outside the 32bit int range)

3. lgamma
  We seemed to have a custom version for the 'double' datatype (which was giving incorrect results). Replaced it with a call to the 'double' version of the underlying 'hc::precision_math::lgamma'


[ROCm/clr commit: 65a90c55e7]
2018-04-24 18:10:07 +00:00
Maneesh Gupta 46ddefedee Apply .clangformat to all repo source files
Change-Id: I7e79c6058f0303f9a98911e3b7dd2e8596079344


[ROCm/clr commit: 9e47fccc89]
2018-03-12 11:29:03 +05:30
Aaron En Ye Shi 880cd966de Fix ilogb/ilogbf functions to return int
This patch will fix hipDoublePrecisionMathDevice test on ThinLTO, which uncovered that hip math_function's ilogb/ilogbf should return type int instead of double. This will match rocdl.


[ROCm/clr commit: b439b45641]
2017-12-05 23:14:10 +00:00
Ben Sander e5bb8f73d0 Fix math ordering for --genco mode.
[ROCm/clr commit: 2e9b78c274]
2017-10-02 21:52:16 +00:00
Maneesh Gupta f29245a846 Renable frexp(f) device math function
Change-Id: I53c022b8ddf38cd17ddb42eba457b9020db66395


[ROCm/clr commit: 1bd1c6bea8]
2017-07-20 14:41:30 +05:30
Maneesh Gupta 0f0f7c8e61 Merge branch 'amd-develop' into amd-master
Change-Id: I05572d2b32f1df70b54e2efeb32c8a4d8055912d
(cherry picked from commit 3a56e5c09b)


[ROCm/clr commit: cc4a73f215]
2017-04-13 03:36:11 -05:00
Aditya Atluri e54234667c fixed header names
Change-Id: I21650d6398187d3767b28e8ac81b2642d3b89a0e


[ROCm/clr commit: 614e7db513]
2017-03-31 12:18:55 -05:00
Aditya Atluri 35d829773d added support for lgammaf and lgamma
1. Implementation inside HIP

Change-Id: I657263b7276a57c56081d3336fef816b5f204eff


[ROCm/clr commit: 52859a8a40]
2017-03-17 18:26:10 -05:00
Aditya Atluri 4085e01263 removed host math functions from math_functions.h
Change-Id: I90d8784e2d6b58c6fade9f0fa12c0db3ee417d3e


[ROCm/clr commit: a5d017e406]
2017-01-27 17:38:43 -06:00
Aditya Atluri b6f4fedaaf fixed compilation issues for vector types and math functions
1. Added math_functions.h to hip_runtime.h
2. Changed operator overloading classifier static to static inline
3. Added vector types test for gpu
4. Seperated __host__ and __device__ for math functions in headers

Change-Id: I499862fad5d7b10da686da9011d7ecefe523f8e2


[ROCm/clr commit: 02190736e3]
2017-01-20 09:49:11 -06:00
Aditya Atluri 5ea40f27b3 Added script for generating math api docs
1. Commented out unsupported device math functions
2. Moved function signatures to the top of implementation snippets
3. Added script to generate markdown documentation for device math apis
4. Added the generated file from the script which should be present everytime

Change-Id: Ic579dd8b8fdffa6e1b4d4f5f3fd8a803f4dcaac7


[ROCm/clr commit: 3d4dcee35d]
2017-01-18 14:40:50 -06:00
Aditya Atluri 69903887a4 fixed compilation issues
1. Fixed compilation issues for tests
2. Added missing intrinsics + math functions
3. Disabled some device functions as they are causing linking error with HCC

Change-Id: I79d52c4c7a539cc8ef40580247ad97ffcb975f09


[ROCm/clr commit: 41a46effef]
2017-01-18 11:53:47 -06:00
Aditya Atluri 42c627fbe8 Moved device code to mimic cuda header behavior
1. All fp32, fp64 math device/host functions should be in math_functions.h/.cpp
2. All fp32, fp64 fast math intrinsics for device/host functions should be in device_functions.h/.cpp
3. All the device code implementations should be in device_util.h/.cpp
4. Hence, made changes appropriately by moving code and creating new header files
5. Added math_functions.cpp/.h
6. Changed #ifndef signature to make sure no conflicts between headers with same names in hip/hip_runtime.h and hip/hcc_detail/hip_runtime.h
7. Changed tests to fit the code changes, making them to include appropriate headers
8. Added math_functions.cpp to CMakeLists.txt
9. Some of the tests are still broken, mostly host math functions will fix them in next commit
10. TODO: FIX compilation issues for host math functions

Change-Id: I7a17637d7e294a7d224ffba932c1a08668febd26


[ROCm/clr commit: d23b6b8694]
2017-01-17 14:57:51 -06:00