Go to file

Leandro Lacerda 75bf739208

[libc][gpu] Disable loop unrolling in the throughput benchmark loop (#153971 )

This patch makes GPU throughput benchmark results more comparable across
targets by disabling loop unrolling in the benchmark loop.

Motivation:
* PTX (post-LTO) evidence on NVPTX: for libc `sin`, the generated PTX
shows the `throughput` loop unrolled 8x at `N=128` (one iteration
advances the input pointer by 64 bytes = 8 doubles), interleaving eight
independent chains before the back-edge. This hides latency and
significantly reduces cycles/call as the batch size `N` grows.
* Observed scaling (NVPTX measurements): with unrolling enabled, `sin`
dropped from ~3,100 cycles/call at `N=1` to ~360 at `N=128`. After
enforcing `#pragma clang loop unroll(disable)`, results stabilized
(e.g., from ~3100 cycles/call at `N=1` to ~2700 at `N=128`).
* libdevice contrast: the libdevice `sin` path did not exhibit a similar
drop in our measurements, and the PTX appears as compact internal calls
rather than a long FMA chain, leaving less ILP for the outer loop to
extract.

What this change does:
* Applies `#pragma clang loop unroll(disable)` to the GPU `throughput()`
loop in both NVPTX and AMDGPU backends.

Leaving unrolling entirely to the optimizer makes apples-to-apples
comparisons uneven (e.g., libc vs. vendor). Disabling unrolling yields
fairer, more consistent numbers.

2025-08-16 20:14:26 +00:00

.ci

[CI][Github] Fix Premerge Summary for Build Failures (#153076 )

2025-08-11 15:37:55 -07:00

.github

github: Add llvm:mc label for generic MC interface (#153737 )

2025-08-15 18:23:24 -07:00

bolt

[BOLT] Do not use HLT as split point when build the CFG (#150963 )

2025-08-15 14:35:13 -07:00

clang

[clang][bytecode] Prefer ParmVarDecls as function parameters (#153952 )

2025-08-16 17:22:14 +02:00

clang-tools-extra

[clang] don't create type source info for vardecl created for structured bindings (#153923 )

2025-08-16 02:04:31 -03:00

cmake

[libc] Change LIBC_THREAD_LOCAL to be dependent on LIBC_THREAD_MODE (#151527 )

2025-08-06 15:04:26 +01:00

compiler-rt

[msan] Reland with even more improvement: Improve packed multiply-add instrumentation (#153353 )

2025-08-15 16:35:42 -07:00

cross-project-tests

[Dexter] Add DAP stepNext and stepOut support (#152717 )

2025-08-13 13:57:53 +01:00

flang

Revert "[flang] Lower EOSHIFT into hlfir.eoshift." (#153907 )

2025-08-15 17:48:40 -07:00

flang-rt

[flang/flang-rt] Add -isysroot flag only to tests really requiring (#152914 )

2025-08-13 21:43:53 +00:00

libc

[libc][gpu] Disable loop unrolling in the throughput benchmark loop (#153971 )

2025-08-16 20:14:26 +00:00

libclc

[libclc] Add __attribute__((const)) to functions that don't access memory (#152456 )

2025-08-12 17:19:08 +08:00

libcxx

[libc++][jthread] LWG3788: jthread::operator=(jthread&&) postconditions are unimplementable under self-assignment (#153758 )

2025-08-16 11:17:30 +08:00

libcxxabi

[libc++][hardening] Introduce assertion semantics. (#149459 )

2025-07-29 00:19:15 -07:00

libsycl

[SYCL] Add libsycl, a SYCL RT library implementation project (#144372 )

2025-07-31 11:28:39 -07:00

libunwind

[libunwind] Fix return type of DwarfFDECache::findFDE() in definition (#146308 )

2025-07-23 00:03:19 +02:00

lld

[LLVM][utils] Add script which clears release notes (#153593 )

2025-08-15 19:00:08 +08:00

lldb

[lldb][nfc] Update docstring of StackFrame "get variable" methods. (#153728 )

2025-08-15 16:47:31 -07:00

llvm

[TableGen] Remove redundant variable (NFC)

2025-08-16 23:11:53 +03:00

llvm-libgcc

[runtimes] Correctly apply libdir subdir for multilib (#93354 )

2024-05-31 11:48:45 -07:00

mlir

[mlir][SparseTensor] Simplify pipeline (#152908 )

2025-08-16 18:45:26 +02:00

offload

[Offload] Introduce dataFence plugin interface. (#153793 )

2025-08-15 11:49:35 -07:00

openmp

[OpenMP] Update ompdModule.c printf to match argument type (#152785 )

2025-08-15 14:30:47 -05:00

polly

Slightly improve the getenv("bar") linking problem

2025-07-23 12:14:51 +01:00

runtimes

[runtimes] Append -nostd*++ flags only for Clang (#151930 )

2025-08-12 15:14:30 +02:00

third-party

[win][arm64ec] Fixes to unblock building LLVM and Clang as Arm64EC (#150068 )

2025-07-31 09:30:05 -07:00

utils/bazel

[bazel] Fix //mlir:XeGPUDialect compilation. (#153904 )

2025-08-15 23:45:32 +00:00

.clang-format

…

.clang-format-ignore

Add empty top level .clang-format-ignore (#136022 )

2025-04-16 15:48:30 -07:00

.clang-tidy

Format root clang-tidy config (NFC) (#147902 )

2025-07-11 07:18:29 +03:00

.git-blame-ignore-revs

Revert "Update .git-blame-ignore-revs for Pack/Unpack move (#152469 )" (#152661 )

2025-08-10 17:02:35 +01:00

.gitattributes

Revert "Finally formalise our defacto line-ending policy"

2024-10-18 21:16:24 +01:00

.gitignore

[llvm] Ignore coding assistant artifacts (#153853 )

2025-08-15 12:27:54 -07:00

.mailmap

[mailmap] Update my name

2025-04-14 16:54:14 +08:00

CODE_OF_CONDUCT.md

…

CONTRIBUTING.md

[www][docs] Remove last mentions of IRC (#139076 )

2025-05-08 09:40:33 -04:00

LICENSE.TXT

…

pyproject.toml

[Py Reformat] Exclude third-party from reformat (#83491 )

2024-03-02 14:51:06 -08:00

README.md

[docs] README: Switch link to clang.llvm.org to use HTTPS.

2024-02-17 12:28:31 +01:00

SECURITY.md

…

README.md

The LLVM Compiler Infrastructure

Welcome to the LLVM project!

This repository contains the source code for LLVM, a toolkit for the construction of highly optimized compilers, optimizers, and run-time environments.

The LLVM project has multiple components. The core of the project is itself called "LLVM". This contains all of the tools, libraries, and header files needed to process intermediate representations and convert them into object files. Tools include an assembler, disassembler, bitcode analyzer, and bitcode optimizer.

C-like languages use the Clang frontend. This component compiles C, C++, Objective-C, and Objective-C++ code into LLVM bitcode -- and from there into object files, using LLVM.

Other components include: the libc++ C++ standard library, the LLD linker, and more.

Getting the Source Code and Building LLVM

Consult the Getting Started with LLVM page for information on building and running LLVM.

For information on how to contribute to the LLVM project, please take a look at the Contributing to LLVM guide.

Getting in touch

Join the LLVM Discourse forums, Discord chat, LLVM Office Hours or Regular sync-ups.

The LLVM project has adopted a code of conduct for participants to all modes of communication within the project.

Languages

LLVM 42%

C++ 30.8%

C 13%

Assembly 9.5%

MLIR 1.4%

Other 2.9%