mirror of
https://github.com/bitcoin/bitcoin.git
synced 2026-09-12 05:32:22 +02:00
3bfdcbd7eecoins: reuse cache hasher for txid set (Lőrinc)2beab94896coins: use SipHash-1-3-UJ for `CCoinsMap` (Lőrinc)7ff55cc650bench: add fixed-width SipHash benchmarks (Lőrinc)3aea85411ftest: add SipHash-1-3-UJ coverage (Pieter Wuille)a0ccd4ad17crypto: add fixed-width SipHash-1-3-UJ (Pieter Wuille)c2d7931b5ccrypto: add generic SipHash-1-3-UJ (Pieter Wuille)25bfca06d6refactor: simplify adding SipHash-1-3-UJ (Lőrinc)af50ba8500test: add shared SipHash vectors (Lőrinc) Pull request description: **Problem:** The in-memory UTXO cache hashes `COutPoint` keys containing a 32-byte txid and a 32-bit output index. SipHash-2-4 processes the txid as four independent 64-bit blocks, so its optimized 32-byte and 36-byte paths both take 14 SipRounds. This also matters for hash-prefix index work such as [#35531](https://github.com/bitcoin/bitcoin/pull/35531): once a persisted key format chooses a hash function, changing it later requires reindexing. **Fix:** Add `SipHasher13UJ`, a custom block-oriented variant combining Pieter Wuille's jumbo-block suggestion with SipHash-1-3, the reduced-round variant discussed in the [SipHash analysis](https://eprint.iacr.org/2012/351.pdf). It provides inline `Hash` overloads for the fixed-width inputs used here. Use a dedicated `SaltedCoinsCacheHasher` for `CCoinsMap` and `CoinsViewOverlay`'s temporary earlier-txid set, while other outpoint tables remain on SipHash-2-4. The salted hash values vary between restarts and are never persisted or sent over the network. **Design:** `SipHasher13UJ` accepts normal 64-bit blocks and 256-bit jumbo blocks. For hash-table use, cryptographic hash outputs must make up all but a small bounded number of retained jumbo blocks. The construction mixes all four limbs around one SipRound, omits byte-oriented padding, and uses an `"unpadded"` finalizer distinct from standard SipHash-1-3. The fixed-width paths take four rounds for one `uint256` jumbo block and five when followed by one normal block. For outpoints, the 32-bit output index is zero-extended into a normal 64-bit block. Retained `CCoinsMap` entries identify real transaction outputs, so their keys contain computed txids. Missing-input validation may probe arbitrary claimed prevouts, but `FetchCoin()` immediately erases their temporary entries when the backend lookup fails, so non-hash keys cannot accumulate. The assumeutxo loader assumes snapshot txids are valid while loading and verifies the complete snapshot's content hash before activation. Every entry in the temporary earlier-txid set is a computed transaction hash, and the set is bounded by the block's transaction count. This construction is limited to local hash tables and is not a general-purpose or protocol SipHash replacement. Pieter discussed the construction with [SipHash co-author Jean-Philippe Aumasson](https://github.com/bitcoin/bitcoin/pull/35215#issuecomment-4385336928), whose preliminary analysis did not find an easier collision construction and supported SipHash-1-3 for this hash-table use. <img width="2100" height="860" alt="siphash_compare_updated" src="https://github.com/user-attachments/assets/cefec6f8-5ec0-450a-a0a2-f946de9ef36d" /> **Structure:** Shared vectors first cover the existing generic and fixed SipHash-2-4 paths in C++, the generic path in Python, and their randomized equivalence in the fuzzer. A behavior-neutral refactor then moves the round, compression, and finalization logic into inline `SipHashState` methods; assembly inspection shows that the fixed-width paths retain their instruction counts, while the generic byte loop retains its prior code generation through a local state copy. Three Pieter-authored commits add the generic UJ specification, fixed-width implementation, and shared correctness coverage. Benchmarks follow that coverage, then separate commits change `CCoinsMap`'s hasher and reuse it for the temporary earlier-txid set. **Tests:** The shared JSON supplies the same byte sequences to the generic C++ and Python SipHash-2-4 implementations, with applicable fixed-width paths checked against the same expected output. The SipHash-2-4 rows include the 64 official vectors for inputs from 0 to 63 bytes and cases that vary input chunking. The UJ outputs were generated by an independent implementation and are checked using normal blocks, equivalent zero-extended jumbo blocks, and applicable fixed-width `Hash` overloads. The integer fuzzer extends these comparisons to arbitrary values and mixed normal/jumbo block encodings. [Counting the dbcache buckets](https://gist.github.com/l0rinc/d68f56c3ed89f76f56da6632ef6f2d92) indicates the new outpoint hasher retains the uniform bucket distribution expected by `CCoinsMap`: <img width="1200" height="750" alt="ccoinsmap-collisions" src="https://github.com/user-attachments/assets/eeedec81-acdc-4adf-a9c8-bfce089700da" /> **Benchmarks:** Fixed-width microbenchmarks compare SipHash-2-4 with SipHash-1-3-UJ for 32-byte hashes and inputs containing a 32-byte hash plus a 32-bit index. Reported aarch64 measurements and an [independent x86_64 run](https://github.com/bitcoin/bitcoin/pull/35215#issuecomment-4400609637) show the outpoint path is about 2x faster. <details><summary>Benchmark runner</summary> ```bash for COMPILER in gcc clang; do \ if [ "$COMPILER" = gcc ]; then CC=gcc; CXX=g++; else CC=clang; CXX=clang++; fi; \ cmake -B "build-bench-$COMPILER" -DCMAKE_BUILD_TYPE=Release -DBUILD_BENCH=ON -DBUILD_TESTS=OFF -DBUILD_GUI=OFF -DENABLE_WALLET=OFF -DCMAKE_C_COMPILER="$CC" -DCMAKE_CXX_COMPILER="$CXX" >/dev/null 2>&1 && \ cmake --build "build-bench-$COMPILER" --target bench_bitcoin -j"$(nproc)" >/dev/null 2>&1 && \ echo "" && echo "$(date -I) | SipHash fixed-width microbench | $("$CXX" --version | head -1) | $(hostname) | $(uname -m) | $(lscpu | awk -F: '/Model name/{print $2; exit}' | xargs) | $(nproc) cores | $(free -h | awk '/^Mem:/{print $2}') RAM" && \ "build-bench-$COMPILER/bin/bench_bitcoin" -filter='SipHash.*32b|SipHash.*36b' -min-time=10000; \ done ``` </details> A two-run GCC `-reindex-chainstate` comparison of the same `CCoinsMap` hot path through height 957,759 with `-dbcache=2000` on a Ryzen 7 3700X/SSD reduced mean wall time from 11,278 s to 10,759 s, a ~5% validation speedup. ACKs for top commit: achow101: light ACK3bfdcbd7eesipa: ACK3bfdcbd7ee(to the extent the code/ideas aren't my own) andrewtoth: ACK3bfdcbd7eeoptout21: ACK3bfdcbd7eeTree-SHA512: c3c66051cb1ebdb0cddbc8b8bed2297c524842f92960531de23d9586c5bb950b302c33d06fc15c2320cbbf97273b22ec3f17fd0e7ddcc7df4a8132a47a606277