| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
This resolves a dead code warning when building without the `compiler`
feature:
```
warning: method `as_u8` is never used
--> crates/hashx/src/register.rs:39:19
|
27 | impl RegisterId {
| --------------- method in this implementation
...
39 | pub(crate) fn as_u8(&self) -> u8 {
| ^^^^^
|
= note: `#[warn(dead_code)]` on by default
```
|
| |
|
|
| |
Run maint/add_warning
|
| | |
|
| |
|
|
|
| |
The new conversion mechanisms in dynasm 4.0 make clippy unhappy
under aarch64.
|
| |
|
|
|
| |
(Starting with version 4, dynasm wants something
that implements Into<u8>.)
|
| | |
|
| |
|
|
|
| |
Some of these lints are in macros in a way that seems to make them
impossible to avoid, or at least, I can't figure out how to avoid them.
|
| |
|
|
| |
This feature has been removed from nightly, in favor of doc_cfg.
|
| |
|
|
|
|
| |
This is now a reserved identifier. The automatic migration
changed it to a raw identifier (`r#gen`), but it's better to use a
different name.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
| |
First, run
```
git grep -l "^edition =" |
xargs perl -i -pe 's/^edition *=.*/edition = "2024"/;'
```
Second, manually verify that all Cargo.toml files have changed,
and nothing else has changed.
Third, run cargo fmt again.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
1. Run cargo fix --edition
2. Selectively revert the "if let"->"match" changes.
These changes are meant to protect us from the lifetime changes
for "if let" bindings in Rust 2024.
But we're not actually relying on the old lifetime rules
anywhere, and the match syntax here is quite ugly.
3. Automatically revert `$pat:expr_2021` to `$pat:expr`.
(We don't actually want to restrict the expression syntax
that our macros accept).
Done with
`git grep -l expr_2021 | xargs perl -i -pe 's/expr_2021/expr/g;'`
4. Run cargo fmt.
|
| |
|
|
| |
Made with https://crates.io/crates/typos-cli
|
| |
|
|
| |
See #2060.
|
| |
|
|
|
|
|
|
|
|
| |
- Replaced single quotes in [`hashx::rand`] documentation link to
[`hashx::rand::RngBuffer`] with backticks so it actually works.
- Added a missing backtick to the link to
[`tor_persist::state_dir::StateDirectory::instance_peek_storage`] in the
documentation of
[`tor_persist::state_dir::StateDirectory::with_instance_path_pieces`].
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Example:
```text
warning: struct pattern is not needed for a unit variant
--> crates/hashx/src/program.rs:165:32
|
165 | Instruction::Target { .. } => Opcode::Target,
| ^^^^^^^ help: remove the struct pattern
|
= help: for further information visit https://rust-lang.github.io/rust-clippy/master/index.html#unneeded_struct_pattern
note: the lint level is defined here
--> crates/hashx/src/lib.rs:9:9
|
9 | #![warn(clippy::all)]
| ^^^^^^^^^^^
= note: `#[warn(clippy::unneeded_struct_pattern)]` implied by `#[warn(clippy::all)]`
```
|
| |
|
|
| |
- `try_fill_bytes()` is no longer a member of RngCore.
|
| |\
| |
| |
| |
| | |
clippy: deny `mod_module_files`
See merge request tpo/core/arti!2689
|
| | |
| |
| |
| |
| |
| | |
Denies 'mod.rs' files for consistency.
https://rust-lang.github.io/rust-clippy/master/index.html#mod_module_files
|
| |/ |
|
| |
|
|
|
|
|
|
| |
In 1.83, this warning triggers on many of our crates.
We're thinking of fixing them all, but for now,
we're going to disable the warning.
This is part of #1765.
|
| |
|
|
| |
This commit is automatically generated.
|
| |
|
|
|
|
| |
Nightly clippy doesn't like using "expr as T" when the conversion is
lossless; it prefers "T::from(expr)" so that if we later change the
type of T to something with a lossy conversion, we'll know.
|
| | |
|
| | |
|
| | |
|
| |
|
|
|
|
|
|
|
| |
This patch tries to make some of the expressions around NUM_INSTRUCTIONS
more convenient. We can import it directly where it's needed, but most
uses are replaced by new type aliases for InstructionArray and
InstructionVec.
No change to any hashx_cachegrind iai benchmarks
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
This is a very simple change, just passing 'self' by reference instead of
value. The by-value version generates a memcpy of the entire temporary
program buffer which doesn't optimize out like I expected it would.
The juicy impact here is a much lower cache footprint for compilation,
since we avoid having yet another temporary storage location for the program
data.
generate_compiled_1000x
Instructions: 271682605 (-0.627292%)
L1 Accesses: 341834751 (-0.813903%)
L2 Accesses: 56420 (-39.48365%)
RAM Accesses: 618 (-20.25806%)
Estimated Cycles: 342138481 (-0.867660%)
|
| | |
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
generate_interp_1000x
Instructions: 219216169 (No change)
L1 Accesses: 278017243 (-0.000441%)
L2 Accesses: 1257 (+4389.286%)
RAM Accesses: 415 (-0.479616%)
Estimated Cycles: 278038053 (+0.001744%)
generate_interp_1000x_c
Instructions: 272748034 (No change)
L1 Accesses: 349932964 (No change)
L2 Accesses: 76 (-1.298701%)
RAM Accesses: 411 (+0.243902%)
Estimated Cycles: 349947729 (+0.000009%)
generate_compiled_1000x
Instructions: 256896028 (+0.175731%)
L1 Accesses: 342543838 (+0.131802%)
L2 Accesses: 149273 (-10.34924%)
RAM Accesses: 810 (-0.246305%)
Estimated Cycles: 343318553 (+0.106328%)
generate_compiled_1000x_c
Instructions: 281855218 (No change)
L1 Accesses: 362569035 (-0.000001%)
L2 Accesses: 88 (+1.149425%)
RAM Accesses: 473 (+0.211864%)
Estimated Cycles: 362586030 (+0.000010%)
interp_u64_hash_1000x
Instructions: 13450926 (No change)
L1 Accesses: 16622561 (+0.000024%)
L2 Accesses: 28 (No change)
RAM Accesses: 390 (-1.015228%)
Estimated Cycles: 16636351 (-0.000817%)
interp_8b_hash_1000x_c
Instructions: 8618541 (No change)
L1 Accesses: 12316160 (-0.000008%)
L2 Accesses: 80 (No change)
RAM Accesses: 433 (+0.231481%)
Estimated Cycles: 12331715 (+0.000276%)
compiled_u64_hash_100000x
Instructions: 87311792 (+0.000520%)
L1 Accesses: 94396598 (+0.000463%)
L2 Accesses: 215 (+2.380952%)
RAM Accesses: 774 (-0.641849%)
Estimated Cycles: 94424763 (+0.000304%)
compiled_8b_hash_100000x_c
Instructions: 91547640 (No change)
L1 Accesses: 98838166 (-0.000007%)
L2 Accesses: 137 (+3.007519%)
RAM Accesses: 488 (+0.618557%)
Estimated Cycles: 98855931 (+0.000119%)
|
| |
|
|
| |
Yet another new module concurrently with a new lint.
|
| | |
|
| | |
|
| | |
|
| |
|
|
| |
These pass miri too.
|
| |
|
|
|
|
| |
To support testing.
No change to iai benchmarks.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
generate_interp_1000x
Instructions: 219216169 (-0.408535%)
L1 Accesses: 278018470 (-0.322766%)
L2 Accesses: 28 (+12.00000%)
RAM Accesses: 417 (+0.724638%)
Estimated Cycles: 278033205 (-0.322706%)
generate_interp_1000x_c
Instructions: 272748034 (No change)
L1 Accesses: 349932964 (-0.000002%)
L2 Accesses: 77 (+1.315789%)
RAM Accesses: 410 (+1.234568%)
Estimated Cycles: 349947699 (+0.000050%)
generate_compiled_1000x
Instructions: 256445375 (-0.349434%)
L1 Accesses: 342092951 (-0.266541%)
L2 Accesses: 166505 (+9.182175%)
RAM Accesses: 812 (+0.370828%)
Estimated Cycles: 342953896 (-0.245532%)
generate_compiled_1000x_c
Instructions: 281855218 (No change)
L1 Accesses: 362569037 (-0.000002%)
L2 Accesses: 87 (No change)
RAM Accesses: 472 (+1.287554%)
Estimated Cycles: 362585992 (+0.000056%)
interp_u64_hash_1000x
Instructions: 13450926 (-0.006631%)
L1 Accesses: 16622557 (-0.005384%)
L2 Accesses: 28 (No change)
RAM Accesses: 394 (+0.510204%)
Estimated Cycles: 16636487 (-0.004959%)
interp_8b_hash_1000x_c
Instructions: 8618541 (No change)
L1 Accesses: 12316161 (-0.000032%)
L2 Accesses: 80 (No change)
RAM Accesses: 432 (+0.934579%)
Estimated Cycles: 12331681 (+0.001103%)
compiled_u64_hash_100000x
Instructions: 87311338 (-0.001022%)
L1 Accesses: 94396161 (-0.000947%)
L2 Accesses: 210 (-0.943396%)
RAM Accesses: 779 (+0.386598%)
Estimated Cycles: 94424476 (-0.000846%)
compiled_8b_hash_100000x_c
Instructions: 91547640 (No change)
L1 Accesses: 98838173 (-0.000003%)
L2 Accesses: 133 (-0.746269%)
RAM Accesses: 485 (+0.831601%)
Estimated Cycles: 98855813 (+0.000134%)
|
| | |
|
| | |
|
| |
|
|
|
| |
This version pushes the panic into the call site, which seems much
better. No change to the iai results.
|
| | |
|
| | |
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
This is quite shoddy. It shouldn't be merged without some tidying up
and unit tests and so on. Also I am confused about the difference
between NUM_INSTRUCTIONS and model::REQUIRED_INSTRUCTIONS.
However:
generate_interp_1000x
Instructions: 220115418 (-1.147100%)
L1 Accesses: 278918725 (-0.936103%)
L2 Accesses: 25 (-98.46248%)
RAM Accesses: 414 (-0.956938%)
Estimated Cycles: 278933340 (-0.938920%)
generate_interp_1000x_c
Instructions: 272748034 (No change)
L1 Accesses: 349932970 (+0.000001%)
L2 Accesses: 76 (-1.298701%)
RAM Accesses: 405 (-0.491400%)
Estimated Cycles: 349947525 (-0.000021%)
generate_compiled_1000x
Instructions: 257344624 (-0.982784%)
L1 Accesses: 343007206 (-0.753942%)
L2 Accesses: 152502 (-17.12839%)
RAM Accesses: 809 (-0.369458%)
Estimated Cycles: 343798031 (-0.797384%)
generate_compiled_1000x_c
Instructions: 281855218 (No change)
L1 Accesses: 362569043 (+0.000002%)
L2 Accesses: 87 (-4.395604%)
RAM Accesses: 466 (-0.427350%)
Estimated Cycles: 362585788 (-0.000023%)
interp_u64_hash_1000x
Instructions: 13451818 (-0.100680%)
L1 Accesses: 16623452 (-0.105967%)
L2 Accesses: 28 (No change)
RAM Accesses: 392 (-1.507538%)
Estimated Cycles: 16637312 (-0.107138%)
interp_8b_hash_1000x_c
Instructions: 8618541 (No change)
L1 Accesses: 12316165 (+0.000032%)
L2 Accesses: 80 (-2.439024%)
RAM Accesses: 428 (-0.465116%)
Estimated Cycles: 12331545 (-0.000616%)
compiled_u64_hash_100000x
Instructions: 87312230 (-1.358594%)
L1 Accesses: 94397055 (-1.669415%)
L2 Accesses: 212 (-0.469484%)
RAM Accesses: 776 (-0.767263%)
Estimated Cycles: 94425275 (-1.669144%)
compiled_8b_hash_100000x_c
Instructions: 91547640 (No change)
L1 Accesses: 98838176 (+0.000009%)
L2 Accesses: 134 (-4.964539%)
RAM Accesses: 481 (-0.414079%)
Estimated Cycles: 98855681 (-0.000097%)
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
generate_interp_1000x
Instructions: 222669662 (-0.146601%)
L1 Accesses: 281554367 (+0.069275%)
L2 Accesses: 1623 (-90.13494%)
RAM Accesses: 418 (No change)
Estimated Cycles: 281577112 (+0.042908%)
generate_interp_1000x_c
Instructions: 272748034 (No change)
L1 Accesses: 349932970 (+0.000001%)
L2 Accesses: 74 (No change)
RAM Accesses: 407 (-0.731707%)
Estimated Cycles: 349947585 (-0.000029%)
generate_compiled_1000x
Instructions: 259898868 (-0.124860%)
L1 Accesses: 345603941 (+0.055045%)
L2 Accesses: 193008 (-4.001910%)
RAM Accesses: 812 (-0.490196%)
Estimated Cycles: 346597401 (+0.043228%)
generate_compiled_1000x_c
Instructions: 281855218 (No change)
L1 Accesses: 362569040 (+0.000000%)
L2 Accesses: 88 (+2.325581%)
RAM Accesses: 468 (-0.636943%)
Estimated Cycles: 362585860 (-0.000026%)
interp_u64_hash_1000x
Instructions: 13465375 (-0.024687%)
L1 Accesses: 16641089 (-0.028974%)
L2 Accesses: 25 (+8.695652%)
RAM Accesses: 398 (+0.505051%)
Estimated Cycles: 16655144 (-0.028470%)
interp_8b_hash_1000x_c
Instructions: 8618541 (No change)
L1 Accesses: 12316165 (+0.000016%)
L2 Accesses: 78 (No change)
RAM Accesses: 430 (-0.462963%)
Estimated Cycles: 12331605 (-0.000551%)
compiled_u64_hash_100000x
Instructions: 88514787 (-0.000365%)
L1 Accesses: 95999693 (+0.000196%)
L2 Accesses: 208 (-0.952381%)
RAM Accesses: 782 (-0.255102%)
Estimated Cycles: 96028103 (+0.000112%)
compiled_8b_hash_100000x_c
Instructions: 91547640 (No change)
L1 Accesses: 98838171 (-0.000002%)
L2 Accesses: 137 (+3.007519%)
RAM Accesses: 483 (-0.412371%)
Estimated Cycles: 98855761 (-0.000053%)
|
| |
|
|
|
|
|
|
|
| |
This is a very small change that converts our Vec cheaply into a boxed
slice during program generation. Program generation speed shows no
changes, and there's no change when using compiled hashes, but is a
surprisingly effective 10% speedup to interpreted hash execution.
Signed-off-by: Micah Elizabeth Scott <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
I was looking for ways to optimize out the many redundant capacity
checks in the Assembler. I didn't find any promising approaches, but
I also saw no evidence that it was an important bottleneck. (A simple
unsafe fix didn't improve any important metrics)
While I was in there, I tightened up the buffer size definitions for
both x86_64 and aarch64, and added assertions to test the limits we
set for the size of prologue, epilogue, and single instructions.
I kept some of the inlining and data type tweaks, even though benchmarks
show no difference. They seem like a step in the right direction, from
the disassembly at least.
Signed-off-by: Micah Elizabeth Scott <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
| |
This is a very simple change that avoids a surprising performance
pitfall: using the code() method on an enum from another crate
caused a non-inlined function call in code where we otherwise expect
a high level of compiler optimization. Replacing code() with a cast
to u8 avoids this function call and allows more intensive optimization
at the call site.
Signed-off-by: Micah Elizabeth Scott <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
This hoists a few decisions out of the innermost portions of
choose_dst_reg, by moving what we can out of dst_register_allowed.
Wallclock time benchmarks:
generate-interp improves, -6.0%
Cachegrind benchmarks:
generate_interp_1000x, -5.0% instructions, -11.6% L2 access, -6% RAM
Signed-off-by: Micah Elizabeth Scott <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
I was trying to eliminate all the places where we copied a Program
(about 4100 bytes) except for the one final copy into a Box; but that
approach was proving too annoying. Even returning a Program via Result
will cause multiple unnecessary copies that don't optimize out.
This patch switches approaches, and instead allocates a Vec<Instruction>
presized to the correct capacity. This allocation is made as early as
possible and retained for the lifetime of the program if necessary.
This means we'll never avoid a heap allocation, but we can always
avoid extra copies and we don't need a separate Box for interpreted
programs.
Performance effects are subtle. Overall wallclock time doesn't change
much. Cachegrind shows some accesses moving up from RAM to L2 cache.
Using GDB to probe memcpy sizes shows that large (>1024b) memcpy are now
totally gone in the generate-interp test.
Signed-off-by: Micah Elizabeth Scott <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Closer inspection of the CPU counters showed that the branching in
RegisterSet::index() was a big problem, contributing to the overall
CPU frontend stall bottleneck in program generation.
This new version is less general, and closer to the appraoch used by
the original C implementation. We store a sorted ArrayVec of in-set
registers, and most operations construct the RegisterSet only once
using a combined filter predicate.
Choosing a register from a set is now cheaper in branches, instructions,
and L1 cache space. We now very rarely manipulate an entire RegisterSet
in any way other than by selecting a register randomly. (Just for the
register R5 special case.)
Wallclock time benchmarks:
generate-interp improves, -7.0%
generate-x86_64 improves, -7.2%
Cachegrind benchmarks:
generate_interp_1000x, more total instructions run but a large
decrease in frontend cache misses. +4.6% instructions, +11% L1
accesses, -99% L2 access, -40% RAM access.
generate_compiled_100x, +4.0% instructions, +9.4% L1 access.
cache miss improvements: -57% L2 access, -25% RAM access.
Signed-off-by: Micah Elizabeth Scott <[email protected]>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
There was a special case in writer_pair_allowed for making add and
subtract equivalent. This patch changes RegisterWriter's encoding, using
per-opcode variants instead of per-format variants. The Add/Sub merge
can now happen earlier, when RegisterWriter is constructed.
Before and after RegisterWriter sizes are the same, at 8 bytes.
This patch removes many uses of Option<RegisterWriter> in favor
of using a new RegisterWriter::None default, and passes by value
rather than by reference.
Wallclock time benchmarks:
generate-interp improves, -7.5%
generate-x86_64 improves, -5.3%
Cachegrind benchmarks:
generate_interp_1000x, negligible change in total instructions,
improvement in cache footprint: -22.8% L2 accesses
Signed-off-by: Micah Elizabeth Scott <[email protected]>
|