mirror of
https://github.com/facebookresearch/faiss.git
synced 2026-10-11 22:50:00 +00:00
main
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cec5702b0b |
DD: give fvec_L2sqr_ny_nearest_y_transposed_D internal linkage (#5720)
Summary: Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5720 `fvec_L2sqr_ny_nearest_y_transposed_D<DIM>` is a template defined at namespace scope in both `distances_avx2.cpp` (`-mavx2`) and `distances_avx512.cpp` (`-mavx512*`), with the same signature, so both TUs emit the same weak symbols and the linker keeps one copy per `DIM`. Whichever copy it keeps is used by both levels. In a current DD build the AVX512 copies of `D<1>`, `D<2>`, `D<4>` were kept (72-95 `zmm` instructions each) and the AVX2 copy of `D<8>`, so the AVX512 path ran AVX2 code for `D<8>`. Nothing on the AVX2 path calls `D<1/2/4>` today, so there is no crash now; but a link that keeps the AVX512 `D<8>` would make the AVX2 path execute AVX512 instructions and raise SIGILL on CPUs without AVX512 (e.g. AMD Milan). Same class as D124399620 (`PQCodeDistanceScalar`), found by scanning the DD binary for functions that contain `zmm` instructions but are not marked as an AVX512 level and have external linkage. The only other hits were `DCBF16_IP` / `DCBF16_L2`, which are defined only in `sq-avx512-spr.cpp` and reached only through AVX512_SPR dispatch. Fix: put the template in an anonymous namespace in each TU, so each level keeps its own copy. Reviewed By: mnorris11 Differential Revision: D124411046 fbshipit-source-id: 9165061e67f1476ff038c173aa837f7e0959e920 |
||
|
|
5d635f354c |
DD: give PQCodeDistanceScalar a SIMD level so the AVX2 path cannot run AVX512 code (#5719)
Summary: Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5719 `PQCodeDistanceScalar<PQDecoderT>` is defined in a header and instantiated in every per-SIMD translation unit (`avx2.cpp` with `-mavx2`, `avx512.cpp` with `-mavx512*`, the generic TU). All copies have the same symbol, so the linker keeps one. In a dynamic-dispatch x86 build it kept the `-mavx512*` copy: `PQCodeDistanceScalar<PQDecoderGeneric>::distance_four_codes` contained 93 `zmm` instructions and was called from `PQCodeDistance<PQDecoderGeneric, SIMDLevel::AVX2>`. On a CPU without AVX512 (AMD Milan, T1_MLN), IVFPQ search with a non-8-bit PQ (`PQDecoderGeneric`, also `PQDecoder16`) on the AVX2 path raises SIGILL. Found by running the faiss C++ tests under the vendored `qemu-x86_64`, which has no AVX512: `IVFPQFastScan.*` (`test_fast_scan_distance_to_code`) died with SIGILL at `vmovdqa64 ...,%zmm0` in that function. Same class as S672695 (AVX512 code reached on a CPU without it); found before any non-AVX512 host was added to the Cogwheel suite. Fix: template `PQCodeDistanceScalar` on `SIMDLevel` as well, as the DD migration guide prescribes for code compiled per level, so each TU's instantiation is a distinct symbol. `PQCodeDistance<D, SL>` passes its level; the NONE and ARM_NEON specializations in `pq_code_distance-generic.h` pass theirs. Applied to both definitions (`impl/pq_code_distance/pq_code_distance-inl.h` and `utils/pq_code_distance.h`). `SL` defaults to `SIMDLevel::NONE` so code outside faiss that names only the decoder keeps compiling: `unicorn/features/encoders/PQFSTable.h` uses `faiss::PQCodeDistanceScalar<PQDecoderT>`, and without the default 15 Unicorn targets failed to build on this diff. Reviewed By: mnorris11 Differential Revision: D124399620 fbshipit-source-id: ece1baf0e1eba39712f881976dea33fe22b265c1 |
||
|
|
f46050eb56 |
Deflake test_io TestIVFPQRead.test_reader (#5718)
Summary: Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5718 `test_reader` searches one IVFPQ index loaded twice, with and without the precomputed table, and required the neighbor ids to be identical. The two paths add the distance terms in a different order, so near-tied neighbors can swap. With unseeded data it failed in 7 of 400 random draws, and once in seven runs of the Cogwheel faiss suite on T1_BGM (2 of 10000 ids). Compare with `check_ref_knn_with_draws`, which allows swaps only among tied distances, and seed the data so a failure reproduces. `test_io` moves to `py_tests_with_contrib` for the helper. Reviewed By: mnorris11 Differential Revision: D124383290 fbshipit-source-id: ceb81262ffaba81ac23e5a592e1658741192060d |
||
|
|
db6aa06d96 |
Hoist per-list allocations out of the IVF fast-scan search loops (#5717)
Summary: Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5717 `IndexIVFFastScan::search_implem_12` and `search_implem_14` allocate a fresh `AlignedTable<uint8_t> LUT`, `q_map` and `lut_entries` for every inverted list they visit, and call `get_block_stride()`, which heap-allocates a `CodePacker` on each call; `search_implem_10` makes the same `get_block_stride()` call per list. These loops run on OpenMP worker threads, so with jemalloc heap profiling on (the production default for task 0 of every job) on aarch64 every sampled allocation pays libunwind's slow fallback at the `__kmp_invoke_microtask` frame. gdb thread snapshots of the D124204736 build under profiling (`hr_ivf1024_rabitqfs4`, Grace): of 206 threads caught in jemalloc's `prof_backtrace`, 163 were in `posix_memalign` called from `search_implem_12` (the per-list LUT), 6 in its per-list vectors and 4 in `get_block_stride()`. This allocates those buffers once per call / per thread, sized for the largest batch (`qbs2` queries x `dim12`), and computes the block stride once. Each iteration still uses only the first `nc` entries, so the search is unchanged. Reviewed By: junjieqi Differential Revision: D124387842 fbshipit-source-id: 6bcfe66270fcc0b208b0c88aaf578fea952f6174 |
||
|
|
44afa7729d |
Stop allocating per (query, probe) on OpenMP workers in IVF RaBitQ fast-scan (#5716)
Summary: Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5716 `IndexIVFRaBitQFastScan` heap-allocated up to three times per (query, probe) inside the OpenMP parallel regions of `compute_LUT` / `compute_LUT_uint8`: - `compute_residual_LUT` built a fresh `std::vector<uint8_t> rotated_qq(d)` on every call. - For multi-bit indexes, it copied `rotated_q` (d floats) into a temporary `QueryFactorsData`. - `compute_LUT_uint8` then copied that temporary into `context.query_factors[ij]` instead of moving it. Under jemalloc heap profiling on aarch64, which Tupperware enables by default in task 0 of every job, each sampled allocation on an OpenMP worker unwinds through libomp's `__kmp_invoke_microtask`. That frame has no CFI, so libunwind's fallback maps and unmaps a whole ELF image per sample, serialised on `mmap_lock` (same mechanism as D123990523; details in that diff). After D123990523, a 74-scenario ServiceLab sweep with profiling on still showed IVF RaBitQ fast-scan at up to 30k (1-bit) and 57k (multi-bit) minor faults per pass, with up to 58 of 72 threads blocked. gdb stack snapshots of a profiled `hr_ivf1024_rabitqfs4` run (348 threads caught inside `je_prof_backtrace`) attributed 149 samples to the `rotated_q` copy and 31 to `rotated_qq` inside `compute_residual_LUT`. The second copy, in `compute_LUT_uint8`, is visible in the code; its frames were not symbolised in the snapshot. This diff: - makes `rotated_qq` a per-thread buffer, like `rotated_q` and `centroid_buf` already are; - stores multi-bit rotated queries in one block allocated once per search in `search_preassigned`, outside the parallel region and uninitialised. Each (query, probe) writes its slice and the handler reads it by storage index, via a new `FastScanDistancePostProcessing::rotated_q` that is offset per thread slice like `query_factors`; - moves rather than copies into `context.query_factors[ij]`; - keeps the capacity of a reused `QueryFactorsData` in `compute_residual_LUT`, so the single-query `IVFRaBitQFastScanScanner` no longer reallocates `rotated_q` on every `set_list`. Memory is unchanged: the same n × nprobe × d floats as before, now contiguous instead of in n × nprobe separate vectors. **Effect** (devbig334 Grace, Tupperware-matched environment, heap profiling on as in production, nq=10000, single runs on a busy shared host; recall unchanged): | hr_ivf1024 | nprobe | QPS before → after | minor faults per pass before → after | |---|---|---|---| | rabitqfs4 (multi-bit) | 16 | 27,032 → 36,069 | 4,006 → 1,911 | | rabitqfs4 (multi-bit) | 64 | 10,168 → 23,295 | 14,932 → 5,822 | | rabitqfs4 (multi-bit) | 256 | 2,883 → 6,438 | 49,841 → 20,371 | | rabitqfs (1-bit) | 256 | 3,179 → 6,854 | 23,935 → 20,654 | Profiling off, rabitqfs4 at nprobe=256: 14,337 → 13,151 QPS, within this host's noise. **What this does not fix.** With these allocations gone, 1-bit and multi-bit converge on about 20k faults per pass at nprobe=256. gdb attributes 163 of the 206 remaining samples to `posix_memalign` in `IndexIVFFastScan::search_implem_12` on worker threads: the LUT tables (`dis_tables`, n × nprobe × LUT size) that each thread slice allocates. Those allocations are needed, and jemalloc samples them by bytes. The remedy for that remainder is unwind info for libomp's aarch64 `__kmp_invoke_microtask`, being raised separately, rather than more FAISS changes. Reviewed By: alibeklfc Differential Revision: D124204736 fbshipit-source-id: 28be93c1cabbcb679bd5328808207fd8e24828b8 |
||
|
|
6692d7ab37 |
Stop copying QueryFactorsData per block in RaBitQHeapHandler
Summary: `RaBitQHeapHandler::handle()` in `IndexRaBitQFastScan.h` copied `rabitq_utils::QueryFactorsData` by value once per (query, 32-vector block). The struct owns `std::vector<float> rotated_q`, so each copy is a heap malloc + free inside the scan's innermost loop: about 312M allocations per 10,000-query pass on sift-1M. It now binds a `const&` to `context->query_factors[q]`, or to a function-local `static const` empty instance when there is no context. `compute_1bit_adjusted_distance` already takes `const&`, so nothing downstream changes. **Why this showed up as an ARM "page-fault storm"** The copy is wasteful on every platform. It becomes catastrophic with jemalloc heap profiling on aarch64, when the allocation happens on an OpenMP worker thread: - Heap profiling is on by default in task 0 of every Tupperware job: the TW scheduler appends `prof:true,prof_final:false,prof_prefix:/tmp/jeprof` when `taskID % contHeapEnabledRatio == 0` (`tupperware/scheduler/JobSpecUtil.cpp`). jemalloc then backtraces one allocation per 512 KiB on average (`lg_prof_sample` 19), through libunwind on aarch64 (`config.prof_libunwind=1`). - On OpenMP workers the stack runs through `__kmp_invoke_microtask`, libomp's assembly entry, which has no usable CFI. DWARF unwinding fails there, and libunwind's aarch64 fallback (`get_frame_state`) calls `get_proc_name`, which maps the frame's whole ELF image and unmaps it on every call. Every map and unmap serialises on `mmap_lock`. - gdb on a reproducer without FAISS: all 95 `get_proc_name` fallbacks hit `__kmp_invoke_microtask`; the same allocations on plain pthreads hit none. - The fault count tracks the sampling rate exactly: 3,053,990 / 1,525,976 / 764,008 faults per pass at `lg_prof_sample` 18 / 19 / 20. That is about 305k samples per pass (160 GB of 512 B copies ÷ 512 KiB) at roughly 5 faults each. **Numbers**, hr_rabitqfs, sift-1M, recall 0.3567 throughout. ServiceLab Grace (72 threads, heap profiling on as in production), nq=10000: | | QPS | minor faults / pass | |---|---|---| | before | 91 | 1.53M | | after (E2362759441136789) | 17,186 | 858 | devbig334 Neoverse-V2, environment matched to Tupperware (no `KMP_*` variables), nq=10000, single runs on a busy shared host, so treat as indicative: | | QPS | kcycles/query | minor faults / pass | |---|---|---|---| | before, profiling off | 9,342 | 13,223 | 322 | | after, profiling off | 11,083 | 11,831 | 341 | | before, profiling on | 93 | 33,563 | 1,525,976 | | after, profiling on | 10,970 | 11,756 | 840 | Without profiling the copy alone costs about 10% more cycles per query. **Not fixed here: `IndexIVFRaBitQFastScan`.** Its scan already reads query factors by reference, but a ServiceLab sweep with this fix and profiling on still shows up to 30k (1-bit) and 57k (multi-bit) faults per pass with up to 58 threads blocked. The allocation site is being located and will be fixed in a separate diff. Reviewed By: mnorris11 Differential Revision: D123990523 fbshipit-source-id: 0112bdbe78ac0fd647ea6199b00000193f96788e |