llvm-project

Author	SHA1	Message	Date
Shilei Tian	fbcce33706	[OpenMP] Honor `thread_limit` value when choosing grid size D152014 introduced an optimization that favors more smaller blocks over fewer larger blocks, even if user sets `thread_limit` explicitly. This patch changes the behavior to honor user value. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D158802	2023-08-26 22:17:49 -04:00
Joseph Huber	aa78e94b0b	[Libomptarget] Support mapping indirect host calls to device functions The changes in D157738 allowed for us to emit stub globals on the device in the offloading entry section. These globals contain the addresses of device functions and allow us to map host functions to their corresponding device equivalent. This patch provides the initial support required to build a table on the device to lookup the associated value. This is done by finding these entries and creating a global table on the device that can be searched with a simple binary search. This requires an allocation, which supposedly should be automatically freed at plugin shutdown. This includes a basic test which looks up device pointers via a host pointer using the added function. This will need to be built upon to provide full support for these calls in the runtime. To support reverse offloading it would also be useful to provide a reverse table that allows us to get host functions from device stubs. Depends on D157738 Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D157918	2023-08-25 18:51:56 -05:00
Kevin Sala	b8e297d1af	[OpenMP][libomptarget] Improve kernel initialization in plugins This patch modifies the plugins so that the initialization of KernelTy objects is done in the init method. Part of the initialization was done in the constructKernelEntry method. Now this method is called constructKernel and only allocates and constructs a KernelTy object. This patch prepares the kernel class for the new implementation of device reductions. Differential Revision: https://reviews.llvm.org/D156917	2023-08-06 11:53:58 +02:00
Shilei Tian	14d57545b2	[NFC][OpenMP] Fix compile warnings introduced in recent patches	2023-08-05 19:38:45 -04:00
koparasy	73cb01dc8a	[OpenMP] Support for OpenMP-Offload Record Replay Enable record-replay for OpenMP offload kernels. On recording the initialization is performed on device initialization by reading env variables. (This is similar to the way rr used to operate). The primary change takes place in the replay phase with the replay tool explicitly initializing the record-replay functionality. Differential Revision: https://reviews.llvm.org/D156174 Fix	2023-08-05 00:46:06 -07:00
Joseph Huber	c96cba3aea	[Libomptarget] Fix compilation of libomptarget with old GCC Summary: Older gcc can't figure out the copy elision and needs an explicit move.	2023-08-03 10:49:35 -05:00
Kevin Sala	4f46a48aaf	[OpenMP][libomptarget] Remove unused virtual functions in GenericKernelTy The virtual functions getDefaultNumBlocks and getDefaultNumThreads from the kernels are only forwarding the call to the generic device's ones. This patch removes those two functions from the kernels (and their derived ones). Now calls are made to the device's functions directly. Differential Revision: https://reviews.llvm.org/D156905	2023-08-02 17:18:50 +02:00
Michael Halkenhaeuser	5b19f42b63	[OpenMP][AMDGPU] Single eager resource init + HSA queue utilization tracking This patch lazily initializes queues/streams/events since their initialization might come at a cost even if we do not use them. To further benefit from this, AMDGPU/HSA queue management is moved into the AMDGPUStreamManager of an AMDGPUDevice. Streams may now use different HSA queues during their lifetime and identify busy queues. When a Stream is requested from the resource manager, it will search for and try to assign an idle queue. During the search for an idle queue the manager may initialize more queues, up to the set maximum (default: 4). When no idle queue could be found: resort to round robin selection. With contributions from Johannes Doerfert <johannes@jdoerfert.de> Depends on D156245 Reviewed By: kevinsala Differential Revision: https://reviews.llvm.org/D154523	2023-08-02 08:22:26 -04:00
Shilei Tian	10068cd654	[OpenMP] Introduce kernel environment This patch introduces per kernel environment. Previously, flags such as execution mode are set through global variables with name like `__kernel_name_exec_mode`. They are accessible on the host by reading the corresponding global variable, but not from the device. Besides, some assumptions, such as no nested parallelism, are not per kernel basis, preventing us applying per kernel optimization in the device runtime. This is a combination and refinement of patch series D116908, D116909, and D116910. Depend on D155886. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D142569	2023-07-26 13:35:14 -04:00
Shilei Tian	6bd74fd65f	Revert commits for kernel environment This reverts commits for kernel environments as they causes issues in AMD BB.	2023-07-23 23:32:31 -04:00
Shilei Tian	c5c8040390	[OpenMP] Introduce kernel environment This patch introduces per kernel environment. Previously, flags such as execution mode are set through global variables with name like `__kernel_name_exec_mode`. They are accessible on the host by reading the corresponding global variable, but not from the device. Besides, some assumptions, such as no nested parallelism, are not per kernel basis, preventing us applying per kernel optimization in the device runtime. This is a combination and refinement of patch series D116908, D116909, and D116910. Depend on D155886. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D142569	2023-07-23 18:36:01 -04:00
Michael Halkenhaeuser	d82eace1c9	[OpenMP][OMPT] Add 'Initialized' flag We observed some overhead and unnecessary debug output. This can be alleviated by (re-)introduction of a boolean that indicates, if the OMPT initialization has been performed. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D155186	2023-07-21 08:19:03 -04:00
Joseph Huber	8a0763f19c	[Libomptarget] Remove RPCHandleTy indirection The 'RPCHandleTy' was intended to capture the intention that a specific device owns its slot in the RPC server. However, this required creating a temporary store to hold these pointers. This was causing really weird spurious failure due to undefined behaviour in the order of library teardown. For example, the x64 plugin would be torn down, set this to some invalid memory, and then the CUDA plugin would crash. Rather than spend the time to fully diagnose this problem I found it pertinent to simply remove the failure mode. This patch removes this indirection so now the usage of the RPC server must always be done with the intended device. This just requires some extra handling for the AMDGPU indirection where we need to store a reference to the device. Reviewed By: JonChesterfield Differential Revision: https://reviews.llvm.org/D154971	2023-07-11 10:54:40 -05:00
Michael Halkenhaeuser	142faf56f5	[OpenMP] [OMPT] [amdgpu] [5/8] Implemented device init/fini/load callbacks Added support in the generic plugin to invoke registered callbacks. Depends on D124070 Patch from John Mellor-Crummey <johnmc@rice.edu> (With contributions from Dhruva Chakrabarti <Dhruva.Chakrabarti@amd.com>) Differential Revision: https://reviews.llvm.org/D124652	2023-07-11 07:13:22 -04:00
Shao-Ce SUN	048423702d	[OpenMP] Fix build warnings ``` llvm-project/openmp/libomptarget/src/private.h:260:9: warning: 'DEBUG_PREFIX' macro redefined [-Wmacro-redefined] #define DEBUG_PREFIX GETNAME(TARGET_NAME) ^ llvm-project/openmp/libomptarget/include/ompt_device_callbacks.h:22:9: note: previous definition is here #define DEBUG_PREFIX "OMPT" ^ 1 warning generated. ``` ``` llvm-project/openmp/libomptarget/plugins-nextgen/common/PluginInterface/PluginInterface.cpp:458:14: warning: moving a local object in a return statement prevents copy elision [-Wpessimizing-move] return std::move(Err); ^ llvm-project/openmp/libomptarget/plugins-nextgen/common/PluginInterface/PluginInterface.cpp:458:14: note: remove std::move call here return std::move(Err); ^~~~~~~~~~ ~ llvm-project/openmp/libomptarget/plugins-nextgen/common/PluginInterface/PluginInterface.cpp:552:12: warning: moving a local object in a return statement prevents copy elision [-Wpessimizing-move] return std::move(Err); ^ llvm-project/openmp/libomptarget/plugins-nextgen/common/PluginInterface/PluginInterface.cpp:552:12: note: remove std::move call here return std::move(Err); ^~~~~~~~~~ ~ 2 warnings generated. ``` Reviewed By: jhuber6 Differential Revision: https://reviews.llvm.org/D154787	2023-07-09 22:12:23 +08:00
Joseph Huber	691dc2d10d	[Libomptarget] Begin implementing support for RPC services This patch adds the intial support for running an RPC server in libomptarget to handle host services. We interface with the library provided by the `libc` project to stand up a basic server. We introduce a new type that is controlled by the plugin and has each device intialize its interface. We then run a basic server to check the RPC buffer. This patch does not fully implement the interface. In the future each plugin will want to define special handlers via the interface to support things like malloc or H2D copies coming from RPC. We will also want to allow the plugin to specify t he number of ports. This is currently capped in the implementation but will be adjusted soon. Right now running the server is handled by whatever thread ends up doing the waiting. This is probably not a completely sound solution but I am not overly familiar with the behaviour of OpenMP tasks and what would be required here. This works okay with synchrnous regions, and somewhat fine with `nowait` regions, but I've observed some weird behavior when one of those regions calls `exit`. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D154312	2023-07-07 12:36:46 -05:00
Joseph Huber	6764301a6b	[Libomptarget] Correctly implement `getWTime` on AMDGPU AMDGPU provides a fixed frequency clock since some generations back. However, the frequency is variable by card and must be looked up at runtime. This patch adds a new device environment line for the clock frequency so that we can use it in the same way as NVPTX. This is the correct implementation and the version in ASO should be replaced. Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D154456	2023-07-04 21:50:43 -05:00
Johannes Doerfert	6629a96a8c	[OpenMP] Improve default block count selection fow low block counts If a combined loop has insufficient parallelism (= low trip count), we might end up with too few teams/blocks. To counter that we can reduce the number of threads per team we use. This patch implements a heuristic and exposes a new environment variable to control the minimum of threads to be employed in this case. Issue reported by: Felipe Cabarcas Jaramillo <cabarcas@udel.edu> (@fel-cab). Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D152014	2023-06-05 16:35:44 -07:00
Kevin Sala	843f496b71	[OpenMP][libomptarget] Improve device info printing in NextGen plugins This patch improves the device info printing in the NextGen plugins. The device info properties are composed of keys, values and units (if necessary). These properties are pushed into a queue by each vendor-specifc plugin, and later, these properties are printed processed and printed by the common Plugin Interface. The printing format is common across the different plugins. Differential Revision: https://reviews.llvm.org/D148178	2023-05-09 15:34:15 +02:00
Shilei Tian	d4ecd1241c	Revert "[OpenMP] Introduce kernel environment" This reverts commit 35cfadfbe2decd9633560b3046fa6c17523b2fa9. It makes a couple of buildbots unhappy because of the following test failures: - `Transforms/OpenMP/add_attributes.ll'` - `mapping/declare_mapper_target_data.cpp` on AMDGPU	2023-04-22 20:56:35 -04:00
Shilei Tian	35cfadfbe2	[OpenMP] Introduce kernel environment This patch introduces per kernel environment. Previously, flags such as execution mode are set through global variables with name like `__kernel_name_exec_mode`. They are accessible on the host by reading the corresponding global variable, but not from the device. Besides, some assumptions, such as no nested parallelism, are not per kernel basis, preventing us applying per kernel optimization in the device runtime. This is a combination and refinement of patch series D116908, D116909, and D116910. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D142569	2023-04-22 20:46:38 -04:00
Kevin Sala	221350965a	[OpenMP][libomptarget][NFC] Remove error data member from AsyncInfoWrapperTy This patch removes the Err data member from the AsyncInfoWrapperTy class. Now the error is stored externally, in the caller side, and it is explicitly passed to the AsyncInfoWrapperTy::finalize() function as a reference. Differential Revision: https://reviews.llvm.org/D148027	2023-04-18 18:52:01 +02:00
Johannes Doerfert	110cf873ad	[OpenMP][NFC] Silence warning	2023-04-17 15:57:10 -07:00
Kevin Sala	8dad7f4953	[OpenMP][libomptarget] Do not rely on AsyncInfoWrapperTy's destructor	2023-04-04 17:51:28 +02:00
Kevin Sala	48cd8b54d1	[NFC][OpenMP][libomptarget] Remove unnecessary AsyncInfoWrapperTy parameter	2023-03-28 17:28:12 +02:00
Joseph Huber	edc0355006	[Libomptarget] Add missing explicit moves on llvm::Error Summary: Some older compilers, which we still support, have problems handling the copy elision that allows us to directly move an `Error` to an `Expected`. This patch adds explicit moves to remove the error.	2023-03-20 11:49:59 -05:00
Joseph Huber	48d5ad93cd	[OpenMP][NFC] Clean up Twines and other issues in plugins Summary: Tihs patch is mostly NFC to fix some warning currently present in OpenMP offloading plugins. Specifically this mostly removes the use of Twine variables in favor of LLVM's small string. Twine variables are prone to use-after-free and this is a cleaner way to concatenate a string.	2023-03-01 15:03:21 -06:00
Joseph Huber	656378085e	[Libomptarget] Fix block and thread limit environment variables not being respected The next-gen plugins did not properly set the values from `OMP_NUM_TEAMS` and `OMP_TEAMS_THREAD_LIMIT`. This is because these maximum values are set by each plugin to its hardware maximum. This happens after the previous initialization. Move it to the correct place and then add a test. Fixes https://github.com/llvm/llvm-project/issues/61082 Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D145105	2023-03-01 14:12:46 -06:00
JP Lehr	b82ac74f7e	[OpenMP][AMDGPU] More detail in AMDGPU kernel launch info Makes the info that is printed for kernel launches configurable for different plugins. Adds all machinery to print the detailed launch info that the current AMD plugin provides and includes e.g. register spill counts. The files msgpack.cpp, msgpack.def, and msgpack.h are copied from the old plugin and are untouched. The contents of UtilitiesHSA.cpp and .h are copied together from various files from the old plugin. The code was originally written by Jon Chesterfield. I updated the function and type names visible to the outside, i.e. in headers, to respect the LLVM conventions. Reviewed By: jhuber6 Differential Revision: https://reviews.llvm.org/D144521	2023-02-28 07:41:48 -05:00
Joseph Huber	9b8e4b4f96	[Libomptarget] Remove unused image argument from global handler function Summary: A previous patch got rid of the use of this image but forgot to remove it from this function. Simply remove it as it is unused now.	2023-02-24 07:24:29 -06:00
Kevin Sala	6ca034644d	[OpenMP][libomptarget] Notify the plugins regarding new mapping/unmappings The NextGen plugins use the information regarding new mapping/unmappings to lock/unlock the corresponding host buffer and speed up the host-device memory transfers involving those buffers. The locking/unlocking is disabled by default and can be enabled by the LIBOMPTARGET_LOCK_MAPPED_HOST_BUFFERS envar. The envar accepts boolean values (on/off) and a special option: - off: Do not lock mapped host buffers (default). - on: Lock mapped host buffers automatically, but do not report lock failures if the plugin fails to lock them. - mandatory: Lock mapped host buffers automatically and treat locking failures in the plugins as fatal errors. This option may be useful for debugging purposes. Differential Revision: https://reviews.llvm.org/D142514	2023-02-06 10:09:35 +01:00
Kevin Sala	2a539ee17d	[OpenMP][libomptarget] Implement memory lock/unlock API in NextGen plugins This patch implements the memory lock/unlock API, introduced in patch https://reviews.llvm.org/D139208, in the NextGen plugins. Locked buffers feature reference counting and we allow certain overlapping. Given an already locked buffer A, other buffers that are fully contained inside A can be locked again, even if they are smaller than A. In this case, the reference count of locked buffer A will be incremented. However, extending an existing locked buffer is not allowed. The original buffer is actually unlocked once all its users have released the locked buffer and sub-buffers (i.e., the reference counter becomes zero). Differential Revision: https://reviews.llvm.org/D141227	2023-01-25 00:11:38 +01:00
Joseph Huber	b280e12a3d	[Libomptarget][NFC] Address a few warnings in libomptarget Summary: Fix a few minor warnings that show up in `libomptarget`.	2023-01-23 08:56:03 -06:00
Johannes Doerfert	40f9bf082f	[OpenMP] Introduce the `ompx_dyn_cgroup_mem(<N>)` clause Dynamic memory allows users to allocate fast shared memory when a kernel is launched. We support a single size for all kernels via the `LIBOMPTARGET_SHARED_MEMORY_SIZE` environment variable but now we can control it per kernel invocation, hence allow computed values. Note: Only the nextgen plugins will allocate memory based on the clause, the old plugins will silently miscompile. Differential Revision: https://reviews.llvm.org/D141233	2023-01-21 18:46:36 -08:00
Johannes Doerfert	16a385ba21	[OpenMP] Modernize the kernel launching interface and APIs We already created a versioned `__tgt_kernel_arguments` struct but it was only briefly used and its content was passed in isolation anyway. This makes it hard to add more information in the future. With this patch we fully embrace the struct as means to pass information from the compiler to the plugin as part of a kernel launch. The patch also extends and renames the struct, bumping the version number to 2. Version 1 entries are auto-upgraded. This is in preparation for "bare" kernel launches, per kernel dynamic shared memory, CUDA/HIP lowering, etc. The `__tgt_target_kernel_nowait` interface was deprecated as it was unused. Once we actually implement support for something like that, we can add an appropriate API. Note: Only plugins with the `launch_kernel` interface are now supported. That means that a new clang won't be able to use an old runtime. An old clang can still use the new runtime since the libomptarget interface did not change. Differential Revision: https://reviews.llvm.org/D141232	2023-01-21 11:16:21 -08:00
Giorgis Georgakoudis	0f4b4e8e4d	[OpenMP] RecordReplay saves bitcode when JIT-ing This patch enables to store bitcode images when JIT is enabled for the record-and-replay functionality (see https://reviews.llvm.org/D138931). Credits to @jdoerfert for refactoring the code. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D141986	2023-01-18 11:25:25 -08:00
Giorgis Georgakoudis	94c772dc92	[OpenMP] Support kernel record and replay This patch adds functionality for recording and replaying the execution of OpenMP offload kernels, based on an original implementation by Steve Rangel. The patch extends libomptarget to extract a json description of the kernel, the device image binary, and a device memory snapshot before and after the execution of a recorded kernel. Kernel recording/replaying in libomptarget is controlled through env vars (LIBOMPTARGET_RECORD, LIBOMPTARGET_REPLAY). It provides a tool, llvm-omp-kernel-replay, for replaying a kernel using the extracted information with the ability to verify replayed execution using the post-execution device memory snapshot, also supporting changing the number of teams/threads for replaying. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D138931	2023-01-17 16:29:03 -08:00
Joseph Huber	566ecc2231	[Libomptarget][NFC] Rename device environment variable This variable is used by the runtime. Before kernel launch we set it to indicate several configuration options from the host. This patch renames it to be more in-line with the rest of the named exported from the runtime. This is better because this is the only symbol visible to the host from the runtime, so it should have a reserved name. Reviewed By: tianshilei1992 Differential Revision: https://reviews.llvm.org/D141960	2023-01-17 14:28:04 -06:00
Johannes Doerfert	f8e094be81	[OpenMP][JIT] Cleanup JIT interface, caching, and races The JIT interface was somewhat irregular as it used multiple global functions. It also did not cache the results of the JIT, hence multiple GPU systems would perform the work multiple times. Finally, there might have been races on the state if we have multi-threaded initialization of different embedded images, or one image initialized on multiple devices. This patch tries to rectify all of the above. The JITEngine is now a part of the GenericPluginTy and tied to one target triple. To support multiple "ComputeUnitKind"s (previously confusingly called Arch or [M]CPU) and to avoid re-jitting for the same ComputeUnitKind, we keep a map of JIT results per ComputeUnitKind. All interaction with the JIT happens through the JITEngine directly, two functions are exposed. Both use (shared) locks to avoid races and cache the result. All JIT-related environment variables are now defined together. Differential Revision: https://reviews.llvm.org/D141081	2023-01-15 11:43:50 -08:00
Ron Lieberman	d179dfe8ce	fix : add missing open brace [OpenMP][FIX] Avoid using an Error object after a std::move.	2023-01-09 19:52:16 -05:00
Johannes Doerfert	bdcbf6c85d	[OpenMP][FIX] Avoid using an Error object after a std::move. The error was always a success even if the error case happened as the std::move reseted the error object.	2023-01-09 16:03:52 -08:00
Johannes Doerfert	ccc1324120	Introduce environment variables to deal with JIT IR We can now dump the IR before and after JIT optimizations into the files passed via `LIBOMPTARGET_JIT_PRE_OPT_IR_MODULE` and `LIBOMPTARGET_JIT_POST_OPT_IR_MODULE`, respectively. Similarly, users can set `LIBOMPTARGET_JIT_REPLACEMENT_MODULE` to replace the IR in the image with a custom IR module in a file. All options take file paths, documentation was added. Reviewed by: tianshilei1992 Differential revision: https://reviews.llvm.org/D140945	2023-01-05 00:17:46 -08:00
Johannes Doerfert	428bc510bf	[OpenMP] Unify "exec_mode" query code and default to SPMD Defaulting to Generic mode doesn't make much sense as the kernel needs to be prepared for it. SPMD mode is the "native" execution, e.g., for "bare" kernels. It also is the execution method for constructors and destructors (as we might otherwise throw an extra warp onto them). Differential Revision: https://reviews.llvm.org/D140718	2023-01-03 16:58:13 -08:00
Shilei Tian	5a3a527f8a	[OpenMP] Introduce basic JIT support to OpenMP target offloading This patch adds the basic JIT support for OpenMP. Currently it only works on Nvidia GPUs. The support for AMDGPU can be extended easily by just implementing three interface functions. However, the infrastructure requires a small extra extension (add a pre process hook) to support portability for AMDGPU because the AMDGPU backend reads target features of functions. `02bc7effcc (diff-321c2038035972ad4994ff9d85b29950ba72c08a79891db5048b8f5d46915314R432)` shows how it roughly works. As for the test, even though I added the corresponding code in CMake files, the test still cannot be triggered because some code is missing in the new plugin CMake file, which has nothing to do with this patch. It will be fixed later. In order to enable JIT mode, when compiling, `-foffload-lto` is needed, and when linking, `-foffload-lto -Wl,--embed-bitcode` is needed. That implies that, LTO is required to enable JIT mode. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D139287	2022-12-27 22:19:05 -05:00
Shilei Tian	95956bd896	Revert "[OpenMP] Introduce basic JIT support to OpenMP target offloading" This reverts commit 58906e4901ec5b7ed230d7fa96123654f6a974af because it breaks AMD's buildbot.	2022-12-27 21:52:07 -05:00
Shilei Tian	58906e4901	[OpenMP] Introduce basic JIT support to OpenMP target offloading This patch adds the basic JIT support for OpenMP. Currently it only works on Nvidia GPUs. The support for AMDGPU can be extended easily by just implementing three interface functions. However, the infrastructure requires a small extra extension (add a pre process hook) to support portability for AMDGPU because the AMDGPU backend reads target features of functions. `02bc7effcc (diff-321c2038035972ad4994ff9d85b29950ba72c08a79891db5048b8f5d46915314R432)` shows how it roughly works. As for the test, even though I added the corresponding code in CMake files, the test still cannot be triggered because some code is missing in the new plugin CMake file, which has nothing to do with this patch. It will be fixed later. In order to enable JIT mode, when compiling, `-foffload-lto` is needed, and when linking, `-foffload-lto -Wl,--embed-bitcode` is needed. That implies that, LTO is required to enable JIT mode. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D139287	2022-12-27 19:07:32 -05:00
Shilei Tian	a82e5825e0	[NFC][OpenMP] Fix compile warning caused by using `std::move` on a local object on a `return` statement	2022-12-23 10:42:29 -05:00
Kevin Sala	e5354a2bfa	[OpenMP][libomptarget] Centralize host pinned buffers map to NextGen's PluginInterface This patch moves the management/tracking of host pinned buffers to the common PluginInterface in NextGen plugins. For the moment, the management consists of tracking the host pinned allocations into a map in each device. Differential Revision: https://reviews.llvm.org/D140502	2022-12-22 02:11:05 +01:00
Johannes Doerfert	fb2c42df41	[OpenMP] Improve AMDGPU Plugin With this patch we: - pick more sensible defaults for the number of teams, inspired by the old plugin, and configured via LIBOMPTARGET_AMDGPU_TEAMS_PER_CU. - check the input signal of a kernel launch late, after the queue lock was taken, to avoid a barrier packet more often. - copy the kernel arguments in one swoop into the appropriate memory. - manually specialize the callbacks to avoid potential indirect calls.	2022-12-19 19:09:43 -08:00
Ye Luo	ee3d9ee49c	[OpenMP] Change the nextgen plugin kernel thread count scheme as old plugins' Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D140352	2022-12-19 18:27:02 -06:00

1 2

57 Commits