aarch64: skip the extend of a scalar comparison's result - #14082
Conversation
|
We've generally tried to avoid lowering rules like this historically where correctness relies on other lowering rules in the system, although we also have some already for x64 so it's not the most principled stance per se. I'd be surprised though if this uextend+icmp showed up too too often in terms of materializing the result of a comparison vs feeding the uextend+icmp into a branch/select/trap/etc. Is this perhaps something where we could get the lion's share of the benefit by shifting around these rules to where conditions are lowered or similar? |
`icmp` and `fcmp` on scalars both lower through `lower_cond_result_bool`, whose every arm already clears upper bits, so extending it is a no-op. This removes ~18,000 `uxtb` instructions emitted from a PCA-based subset of Sightglass (this is ~87% of `uxtb` instructions emitted and 0.220% of *all* instructions emitted). It also results in an average 0.42% faster execution (significant) in terms of cycles.
17a7e37 to
b3faf3c
Compare
I pushed another commit that folds this into the matching rule, but unfortunately, this has the effect of re-lowering the |
|
With optimizations enabled I think that would be resolved with GVN though, right? |
Yeah, and in fact I read the diff backwards (d'oh) so the filetests are asserting that the duplication doesn't happen right now, so we should be good to merge this (assuming you think it looks good) |
icmpandfcmpon scalars both lower throughlower_cond_result_bool, whose every arm already clears upper bits, so extending it is a no-op.This removes ~18,000
uxtbinstructions emitted from a PCA-based subset of Sightglass (this is ~87% ofuxtbinstructions emitted and 0.220% of all instructions emitted). It also results in an average 0.42% faster execution (significant) in terms of cycles.