3184 Commits

Author SHA1 Message Date
Nemanja Ivanovic
563720c3be
[RISCV] Fix lowering of negative zero with Zdinx 32-bit (#71869)
The compiler currently abends with an impossible reg-to-reg copy when
producing a negative zero FP immediate on RV32 with -Zdinx. This is
because we emit a negation that uses FP registers. Emit the right node
to produce correct code.
2023-11-13 07:38:14 +01:00
Craig Topper
ee95819503 [RISCV][GISel] Legalize G_FSHL/G_FSHR. 2023-11-11 20:23:29 -08:00
Craig Topper
b0e97c7757 [RISCV][GISel] Legalize G_ROTL/G_ROTR. 2023-11-11 20:12:07 -08:00
Craig Topper
7965a21f7a [RISCV] Add more packh patterns. 2023-11-11 19:31:23 -08:00
Craig Topper
bfb7843580 [RISCV] Add packw/packh patterns for -riscv-experimental-rv64-legal-i32 2023-11-11 17:52:22 -08:00
Craig Topper
6b9752cc72 [RISCV] Add rv64zbkb.ll test for -riscv-experimental-rv64-legal-i32. NFC 2023-11-11 17:52:22 -08:00
Craig Topper
fdc904e568 [RISCV] Add isel pattern to turn (or (zext X), Y) into add.uw when X and Y are disjoint.
Improve code for -riscv-experimental-rv64-legal-i32.
2023-11-11 15:51:38 -08:00
Craig Topper
bf0963620c [RISCV] Add (shl (zext GPR:), uimm5:) pattern for -riscv-experimental-rv64-legal-i32. 2023-11-11 15:14:02 -08:00
Craig Topper
994d882e15 [RISCV] Add an slli.uw pattern using zext for -riscv-experimental-rv64-legal-i32
We already had the pattern for GlobalISel. Move it over to SelectionDAG.
2023-11-11 14:41:56 -08:00
Craig Topper
bab2cf2d01 [RISCV][GISel] Promote s32 constant shift amounts to s64 on RV64.
This allows us to reuse isel patterns from SelectionDAG.

This is similar to what is done on AArch64.
2023-11-10 23:07:00 -08:00
Craig Topper
647c490f8a [RISCV] Add an add.uw pattern using zext for -riscv-experimental-rv64-legal-i32 and global isel 2023-11-10 21:36:29 -08:00
Craig Topper
7e0bae5b34 [RISCV][GISel] Add isel patterns for SHXADD with s32 type on RV64. 2023-11-10 19:52:57 -08:00
Craig Topper
a93dfb589d [RISCV] Peek through zext in selectShiftMask.
This improves the code for -riscv-experimental-rv64-legal-i32
2023-11-10 19:02:14 -08:00
Craig Topper
83cc24e598 [RISCV] Add test case showing unnecessary zext of shift amounts with -riscv-experimental-rv64-legal-i32. NFC 2023-11-10 19:02:13 -08:00
Craig Topper
ca603343db
[RISCV][GISel] Legalizer and register bank selection for G_JUMP_TABLE and G_BRJT (#71970)
Testing together since they should come paired.

Instruction selection will be a separate PR.
2023-11-10 13:09:24 -08:00
Yingwei Zheng
650026897c
[RISCV][SDAG] Prefer ShortForwardBranch to lower sdiv by pow2 (#67364)
This patch lowers `sdiv x, +/-2**k` to `add + select + shift` when the
short forward branch optimization is enabled. The latter inst seq
performs faster than the seq generated by target-independent
DAGCombiner. This algorithm is described in ***Hacker's Delight***.

This patch also removes duplicate logic in the X86 and AArch64 backend.
But we cannot do this for the PowerPC backend since it generates a
special instruction `addze`.
2023-11-10 21:38:47 +08:00
Wang Pengcheng
9bb69c1d96
[RISCV] Enable LoopDataPrefetch pass (#66201)
So that we can benefit from data prefetch when `Zicbop` extension is
supported.

Tune information for data prefetching are added in `RISCVTuneInfo`.
2023-11-10 15:39:58 +08:00
Craig Topper
fdbff88196
[RISCV][GISel] Add support for G_FCMP with F and D extensions. (#70624)
We only have instructions for OEQ, OLT, and OLE. We need to convert
other comparison codes into those.

I think we'll likely want to split this up in the future to support
optimizations. Maybe do some of it in the legalizer or in a new post
legalizer lowering pass. So this patch is just enough to get something
working without adding 11 additional patterns to tablegen for each type.
2023-11-09 20:45:35 -08:00
Craig Topper
aae30f9e2c [RISCV] Use Align(8) for the stack temporary created for SPLAT_VECTOR_SPLIT_I64_VL.
The value needs to be read as an 8 byte vector element which requires
the pointer to be 8 byte aligned according to the vector spec.

Fixes #71787
2023-11-09 20:43:22 -08:00
Craig Topper
247eb13fab [RISCV][GISel] Legalize G_BITREVERSE. 2023-11-09 16:27:21 -08:00
Craig Topper
8b98d5b813 [RISCV][GISel] Enable libcall expansion for G_FCEIL and G_FFLOOR. 2023-11-09 13:14:42 -08:00
Craig Topper
679cc16c99 [RISCV] Disable early promotion for Zbs in performANDCombine with riscv-experimental-rv64-legal-i32
We can match this directly in isel with the i32 type being legal.

The generic DAG combine will unpromote part of the pattern and
prevent it from being matched in isel.
2023-11-09 09:51:31 -08:00
Craig Topper
24577bd089 [RISCV] Add BSET/BCLR/BINV/BEXT patterns for riscv-experimental-rv64-legal-i32. 2023-11-09 09:17:22 -08:00
Philip Reames
7ac8486e54
[RISCVInsertVSETVLI] Allow PRE with non-immediate AVLs (#71728)
Extend our PRE logic to cover non-immediate AVL values. This covers
large constant AVLs (which must be materialized in registers), and may
help some code written explicitly with intrinsics.

Looking at the existing code, I can't entirely figure out why I thought
we needed VL == AVL to perform the PRE. My best guess is that I was
worried about the VLMAX < VL < 2 * VLMAX case, but the spec explicitly
says that vsetvli must be determinist on any particular AVL value.

That case was, possibly by accident, covering another legality
precondition. Specifically, by only returning true for immediate and
VLMAX AVL values, we didn't encounter the case where the AVL was a
register and that register wasn't available in the predecessor (e.g. if
AVL is a load in the MBB block itself).

---------

Co-authored-by: Luke Lau <luke_lau@icloud.com>
2023-11-09 08:03:13 -08:00
Craig Topper
e3c120a585 [RISCV] Add a Zbb+Zbs command line to rv*zbs.ll to get coverage on an existing isel pattern. NFC
This pattern wasn't tested

def : Pat<(XLenVT (and (rotl -2, (XLenVT GPR:$rs2)), GPR:$rs1)),
          (BCLR GPR:$rs1, GPR:$rs2)>;1
2023-11-08 22:31:49 -08:00
Jianjian Guan
d36eb79ccc
[RISCV] Support Strict FP arithmetic Op when only have Zvfhmin (#68867)
Include: STRICT_FADD, STRICT_FSUB, STRICT_FMUL, STRICT_FDIV,
STRICT_FSQRT and STRICT_FMA.
2023-11-09 09:55:48 +08:00
Craig Topper
24b11ba24d [RISCV][GISel] Use default lowering for G_DYN_STACKALLOC. 2023-11-07 23:59:27 -08:00
Luke Lau
11c182740a
[RISCV] Use masked pseudo peephole for reduction pseudos (#71508)
After #71483 we now have a way of marking masked pseudos as having an
unmasked
equivalent, but their mask shouldn't be folded unless it's all ones
since it
would affect the result.

This patch uses it to mark the pseudos for vredsum and friends, which in
turn
allows us to remove the unmasked patterns, and catch some other forms of
vmerge.
2023-11-08 12:46:06 +08:00
Craig Topper
a6c80c4f70 [RISCV][GISel] Add support for G_SITOFP/G_UITOFP with F and D extensions. 2023-11-07 16:40:58 -08:00
Michael Maitland
ac4ff6168a
[CodeGen][MachineVerifier] Use TypeSize instead of unsigned for getRe… (#70881)
…gSizeInBits

This patch changes getRegSizeInBits to return a TypeSize instead of an
unsigned in the case that a virtual register has a scalable LLT. In the
case that register is physical, a Fixed TypeSize is returned.

The MachineVerifier pass is updated to allow copies between fixed and
scalable operands as long as the Src size will fit into the Dest size.

This is a precommit which will be stacked on by a change to GISel to
generate COPYs with a scalable destination but a fixed size source.

This patch is stacked on https://github.com/llvm/llvm-project/pull/70893
for the ability to use scalable vector types in MIR tests.
2023-11-07 14:38:46 -05:00
Craig Topper
374fb4126f [RISCV][GISel] Add support for G_FPTOSI/G_FPTOUI with F and D extensions. 2023-11-07 10:14:37 -08:00
Luke Lau
fd4804423b [RISCV] Add tests for pseudos that shouldn't have vmerge folded into them. NFC 2023-11-07 18:25:37 +08:00
Jim Lin
4306cfd40e
[RISCV] Fix using undefined variable %pt2 in mask-reg-alloc.mir testcase (#70764)
First PseudoVMERGE_VIM_M1 should use %pt1 as its operand instead of
%pt2.

I found this error when I add LiveIntervals analysis pass in my
downstream. And it crashes with the message:

```
Use of %7 does not have a corresponding definition on every path:
112r %6:vrnov0 = PseudoVMERGE_VIM_M1 %pt2:vrnov0(tied-def 0), %2:vr, 1, %4:vmv0, 1, 3
LLVM ERROR: Use not jointly dominated by defs.
```
2023-11-07 17:05:03 +08:00
Yeting Kuo
a5c1ecada2
[RISCV] Disable performCombineVMergeAndVOps for PseduoVIOTA_M. (#71483)
This transformation might be illegal for `PseduoVIOTA_M`. The value of
`viota.m vd, vs2` is the prefix sum of vd2 and adding mask for it may
cause wrong prefix sum.
Take an example, the result of following expression is `{5, 5, 5, 3}`,
```
; v4 = {1, 1, 1, 1}
viota.m v1, v4
; v0 = {0, 0, 0, 1}, v1 = {0, 1, 2, 3}, v8 = {5, 5, 5, 5}
vmerge.vvm v8, v8, v1, v0.t
; v8 = {5, 5, 5, 3}
```
but if we merge them to `viota.m v8, v4, v0.t`, then the result of is
`{5, 5, 5, 0}`.
Also, we still does `performCombineVMergeAndVOps` for `voita.m` when
mask of `vmerge.vvm` is a true mask.
2023-11-07 16:21:35 +08:00
Luke Lau
5ed0b21580 [RISCV] Add FileCheck prefixes for test where RV32/RV64 output differs. NFC 2023-11-06 19:32:07 +08:00
Craig Topper
90f768440d
[VP][RISCV] Add llvm.experimental.vp.reverse. (#70405)
This is similar to vector.reverse, but only reverses the first EVL
elements.

I extracted this code from our downstream. Some of it may have come from
https://repo.hca.bsc.es/gitlab/rferrer/llvm-epi/ originally.
2023-11-05 22:39:27 -08:00
Craig Topper
9f6010d09e [RISCV][GISel] Emit ADJCALLSTACKDOWN/UP instructions in RISCVCallLowering::lowerCall.
This is needed to mark the stack usage for the call.

With this change, I'm now able to succesfully execute spec2006int
compiled with -O0. There are still many fallbacks that need to be
addressed.
2023-11-05 18:26:45 -08:00
Craig Topper
f4bc18916a [RISCV][GISel] Pass the IsFixed flag into CC_RISCV for outgoing arguments.
This is needed to make FP values be pased in a GPR as required by
the variadic function ABI.
2023-11-05 13:06:23 -08:00
Craig Topper
5d2f9aee66 [RISCV][GISel] Add call preserved regmask to calls created by RISCVCallLowering::lowerCall. 2023-11-05 11:45:35 -08:00
Craig Topper
39edac23df [RISCV][GISel] Fix incorrect call to getGlobalAddress in selectGlobalValue.
RISCVII::MO_HI was being passed to the offset argument instead of
the flags argument.

Adjust some other calls to not pass an explicit 0 to the offset argument
since it already has a default value of 0.
2023-11-04 23:44:10 -07:00
Craig Topper
422ffc525a [RISCV][GISel] Add instruction selection for G_FCONSTANT using integer materialization.
This supports any G_FCONSTANT for F and D extensions.

This builds the constant in the integer domain and moves it to FP
using either FMV or the stack.

Eventually we should use the constant pool for some constants that
require many instructions, but this is a good starting point to
get something working.
2023-11-04 11:40:52 -07:00
Yeting Kuo
af4abc4fa7
[RISCV] Remove experimental- prefix for smaia and ssaia. (#71172)
Since smaia and ssaia are ratified now, we could remove their
experimental- prefix.
2023-11-04 08:16:55 +08:00
Ramkumar Ramachandra
fd887a3633
LegalizeVectorTypes: fix bug in widening of vec result in xrint (#71198)
Fix a bug introduced in 98c90a1 (ISel: introduce vector ISD::LRINT,
ISD::LLRINT; custom RISCV lowering), where ISD::LRINT and ISD::LLRINT
used WidenVecRes_Unary to widen the vector result. This leads to
incorrect CodeGen for RISC-V fixed-vectors of length 3, and a crash in
SelectionDAG when we try to lower llvm.lrint.vxi32.vxf64 on i686. Fix
the bug by implementing a correct WidenVecRes_XRINT.

Fixes #71187.
2023-11-03 21:04:09 +00:00
Min-Yih Hsu
1e39575a98
[RISCV] CSE by swapping conditional branches (#71111)
DAGCombiner, as well as InstCombine, tend to canonicalize GE/LE into
GT/LT, namely:
```
X >= C --> X > (C - 1)
```
Which sometime generates off-by-one constants that could have been CSE'd
with surrounding constants.
Instead of changing such canonicalization, this patch tries to swap
those branch conditions post-isel, in the hope of resurfacing more
constant CSE opportunities. More specifically, it performs the following
optimization:

For two constants C0 and C1 from
```
li Y, C0
li Z, C1
```
To remove redundnat `li Y, C0`,
 1. if C1 = C0 + 1 we can turn: 
    (a) blt Y, X -> bge X, Z
    (b) bge Y, X -> blt X, Z
 2. if C1 = C0 - 1 we can turn: 
    (a) blt X, Y -> bge Z, X
    (b) bge X, Y -> blt Z, X

This optimization will be done by PeepholeOptimizer through
RISCVInstrInfo::optimizeCondBranch.
2023-11-03 09:03:52 -07:00
Brandon Wu
74f38df1d1
[RISCV] Support Xsfvfnrclipxfqf extensions (#68297)
FP32-to-int8 Ranged Clip Instructions

https://sifive.cdn.prismic.io/sifive/0aacff47-f530-43dc-8446-5caa2260ece0_xsfvfnrclipxfqf-spec.pdf
2023-11-03 10:52:37 +08:00
Brandon Wu
945d2e6e60
[RISCV] Support Xsfvfwmaccqqq extensions (#68296)
Bfloat16 Matrix Multiply Accumulate Instruction

https://sifive.cdn.prismic.io/sifive/c391d53e-ffcf-4091-82f6-c37bf3e883ed_xsfvfwmaccqqq-spec.pdf
2023-11-03 10:08:26 +08:00
Brandon Wu
65dc96c2cf
[RISCV] Fix wrong implication for zvknhb. (#66860) 2023-11-03 09:32:21 +08:00
Craig Topper
9769026858 [RISCV] Add (i32 (and GPR:, TrailingOnesMask:)) pattern for RV64 with legal i32. 2023-11-02 15:03:05 -07:00
Craig Topper
014390d937
[RISCV] Implement cross basic block VXRM write insertion. (#70382)
This adds a new pass to insert VXRM writes for vector instructions. With
the goal of avoiding redundant writes.

The pass does 2 dataflow algorithms. The first is a forward data flow to
calculate where a VXRM value is available. The second is a backwards
dataflow to determine where a VXRM value is anticipated.

Finally, we use the results of these two dataflows to insert VXRM writes
where a value is anticipated, but not available.

The pass does not split critical edges so we aren't always able to
eliminate all redundancy.

The pass will only insert vxrm writes on paths that always require it.
2023-11-02 14:09:27 -07:00
Amara Emerson
d62c6ad2b0 Fix more RISCV GISel tests using -march instead of -mtriple 2023-11-02 12:42:00 -07:00