This is a first attempt at a constant value consecutive store merging pass, a counterpart to the DAGCombiner's store merging optimization. The high level goals of this pass: * Have a simple and efficient algorithm. As close to linear time as we can get. Thus, prioritizing scalability of the algorithm over merging every corner case we can find. The DAGCombiner's store merging code has been the source of compile time and complexity issues in the past and I wanted to avoid that. * Don't introduce any new data structures for ordering memory operations. In MIR, we don't have the concept of chains like we do in the DAG, and the instruction order is stricter than enforcing ordering with graph edges. Although I considered adding something similar, I couldn't justify the overhead. The pass is current split into 3 main parts. The main store merging code focuses on identifying candidate stores and managing the candidate group that's under consideration for merging. Analyzing addressing of stores is a potentially complex part and for now there's just a basic implementation to identify easy cases. Finally, the other main bit of complexity is the alias analysis, which tries to follow the same logic as the DAG's AA. Currently this implementation only supports merging of constant stores. Stores of arbitrary variables are technically possible with a very small change, but the DAG chooses not to do this. Doing so here makes most code worse since there's extra overhead in merging values into wider registers. On AArch64 -Os, this optimization results in very minor savings on CTMark. Differential Revision: https://reviews.llvm.org/D109131
93 lines
4.0 KiB
LLVM
93 lines
4.0 KiB
LLVM
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 -O0 \
|
|
; RUN: | FileCheck %s --check-prefixes=ENABLED,FALLBACK
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs -O0 \
|
|
; RUN: | FileCheck %s --check-prefixes=ENABLED,FALLBACK,VERIFY,VERIFY-O0
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 -O0 -aarch64-enable-global-isel-at-O=0 -global-isel-abort=1 \
|
|
; RUN: | FileCheck %s --check-prefixes=ENABLED,NOFALLBACK
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 -O0 -aarch64-enable-global-isel-at-O=0 -global-isel-abort=2 \
|
|
; RUN: | FileCheck %s --check-prefixes=ENABLED,FALLBACK
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 -global-isel \
|
|
; RUN: | FileCheck %s --check-prefix ENABLED --check-prefix NOFALLBACK --check-prefix ENABLED-O1
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 -global-isel -global-isel-abort=2 \
|
|
; RUN: | FileCheck %s --check-prefix ENABLED --check-prefix FALLBACK --check-prefix ENABLED-O1
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 -O1 -aarch64-enable-global-isel-at-O=3 \
|
|
; RUN: | FileCheck %s --check-prefix ENABLED --check-prefix ENABLED-O1
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 -O1 -aarch64-enable-global-isel-at-O=0 \
|
|
; RUN: | FileCheck %s --check-prefix DISABLED
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 -aarch64-enable-global-isel-at-O=-1 \
|
|
; RUN: | FileCheck %s --check-prefix DISABLED
|
|
|
|
; RUN: llc -mtriple=aarch64-- -debug-pass=Structure %s -o /dev/null 2>&1 \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -verify-machineinstrs=0 | FileCheck %s --check-prefix DISABLED
|
|
|
|
; RUN: llc -mtriple=aarch64-- -fast-isel=0 -global-isel=false \
|
|
; RUN: --debugify-and-strip-all-safe=0 \
|
|
; RUN: -debug-pass=Structure %s -o /dev/null 2>&1 -verify-machineinstrs=0 \
|
|
; RUN: | FileCheck %s --check-prefix DISABLED
|
|
|
|
; ENABLED: IRTranslator
|
|
; VERIFY-NEXT: Verify generated machine code
|
|
; ENABLED-NEXT: Analysis for ComputingKnownBits
|
|
; ENABLED-O1-NEXT: MachineDominator Tree Construction
|
|
; ENABLED-O1-NEXT: Analysis containing CSE Info
|
|
; ENABLED-O1-NEXT: PreLegalizerCombiner
|
|
; VERIFY-O0-NEXT: AArch64O0PreLegalizerCombiner
|
|
; VERIFY-NEXT: Verify generated machine code
|
|
; ENABLED-O1-NEXT: Basic Alias Analysis (stateless AA impl)
|
|
; ENABLED-O1-NEXT: Function Alias Analysis Results
|
|
; ENABLED-O1-NEXT: LoadStoreOpt
|
|
; ENABLED-O1-NEXT: Analysis containing CSE Info
|
|
; VERIFY-O0-NEXT: Analysis containing CSE Info
|
|
; ENABLED-NEXT: Legalizer
|
|
; VERIFY-NEXT: Verify generated machine code
|
|
; ENABLED: RegBankSelect
|
|
; VERIFY-NEXT: Verify generated machine code
|
|
; ENABLED-NEXT: Localizer
|
|
; VERIFY-O0-NEXT: Verify generated machine code
|
|
; ENABLED-O1-NEXT: Analysis for ComputingKnownBits
|
|
; ENABLED-O1-NEXT: Lazy Branch Probability Analysis
|
|
; ENABLED-O1-NEXT: Lazy Block Frequency Analysis
|
|
; ENABLED-NEXT: InstructionSelect
|
|
; ENABLED-O1-NEXT: AArch64 Post Select Optimizer
|
|
; VERIFY-NEXT: Verify generated machine code
|
|
; ENABLED-NEXT: ResetMachineFunction
|
|
|
|
; FALLBACK: AArch64 Instruction Selection
|
|
; NOFALLBACK-NOT: AArch64 Instruction Selection
|
|
|
|
; DISABLED-NOT: IRTranslator
|
|
|
|
; DISABLED: AArch64 Instruction Selection
|
|
; DISABLED: Finalize ISel and expand pseudo-instructions
|
|
|
|
define void @empty() {
|
|
ret void
|
|
}
|