mirror of
https://github.com/facebookresearch/faiss.git
synced 2026-10-11 22:50:00 +00:00
Summary: Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5587 **TL;DR:** Adds a batched 20-byte Hamming kernel that measures eight codes per call, the only thing that helps the 160-bit width, worth 1.21x end to end. A 20-byte code has no vector kernel at any x86 SIMD level. `HammingComputer20` measures one distance with three scalar popcounts at every level, including `AVX512_VPOPCNT`, so a scan of 160-bit codes runs the same code everywhere. Vectorizing a single 20-byte distance does not help. A masked 256-bit load plus `vpopcntq` plus a horizontal reduction measures the same as the scalar body, because the reduction costs more than the three popcounts it replaces. Measuring several codes per call does help, because it amortizes that reduction. This adds `hamming_batch()` to `HammingComputer20` at `AVX512_VPOPCNT`, which measures eight codes at once: - Eight codes span 160 bytes. `_mm512_popcnt_epi8` gives a count per byte. - `_mm512_sad_epu8` sums each aligned 8-byte group. - A 20-byte code covers two whole groups plus half of a group it shares with its neighbour, so each code needs one scalar popcount for its half. `IVFBinaryScannerL2::scan_codes` and `IVFBinaryScannerL2::scan_codes_range` call `hamming_batch()` when the computer declares a `batch_size`, then measure the remaining codes one at a time. A computer without a `batch_size` keeps the existing loop, so no other code size changes. The scanner holds the tiled query, and the constructor rejects a code size that does not match the computer. A caller therefore cannot reach the batch kernel with a stride the kernel does not read. The byte-wise popcount needs `AVX512_BITALG`, which is a separate CPUID bit from `AVX512_VPOPCNTDQ`. This adds the flag to the `AVX512_VPOPCNT` level and requires both bits to select it. Every CPU that has VPOPCNTDQ also has BITALG, so no CPU loses the level. An alternative byte-wise popcount built from `vpshufb`, which needs neither bit and would run at the plain `AVX512` level, measures only 1.08x against 1.54x, so it is not worth the wider reach. Reviewed By: alibeklfc Differential Revision: D118971285 fbshipit-source-id: 5fd4637c972e5c55e2695b32b9b15fba8181a242