6 Commits
Author SHA1 Message Date
Weixiang Hong cec5702b0b DD: give fvec_L2sqr_ny_nearest_y_transposed_D internal linkage (#5720)
Summary:
Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5720

`fvec_L2sqr_ny_nearest_y_transposed_D<DIM>` is a template defined at namespace scope in both `distances_avx2.cpp` (`-mavx2`) and `distances_avx512.cpp` (`-mavx512*`), with the same signature, so both TUs emit the same weak symbols and the linker keeps one copy per `DIM`. Whichever copy it keeps is used by both levels. In a current DD build the AVX512 copies of `D<1>`, `D<2>`, `D<4>` were kept (72-95 `zmm` instructions each) and the AVX2 copy of `D<8>`, so the AVX512 path ran AVX2 code for `D<8>`. Nothing on the AVX2 path calls `D<1/2/4>` today, so there is no crash now; but a link that keeps the AVX512 `D<8>` would make the AVX2 path execute AVX512 instructions and raise SIGILL on CPUs without AVX512 (e.g. AMD Milan).

Same class as D124399620 (`PQCodeDistanceScalar`), found by scanning the DD binary for functions that contain `zmm` instructions but are not marked as an AVX512 level and have external linkage. The only other hits were `DCBF16_IP` / `DCBF16_L2`, which are defined only in `sq-avx512-spr.cpp` and reached only through AVX512_SPR dispatch.

Fix: put the template in an anonymous namespace in each TU, so each level keeps its own copy.

Reviewed By: mnorris11

Differential Revision: D124411046

fbshipit-source-id: 9165061e67f1476ff038c173aa837f7e0959e920
2026-10-10 01:42:32 -07:00
Weixiang Hong 5d635f354c DD: give PQCodeDistanceScalar a SIMD level so the AVX2 path cannot run AVX512 code (#5719)
Summary:
Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5719

`PQCodeDistanceScalar<PQDecoderT>` is defined in a header and instantiated in every per-SIMD translation unit (`avx2.cpp` with `-mavx2`, `avx512.cpp` with `-mavx512*`, the generic TU). All copies have the same symbol, so the linker keeps one. In a dynamic-dispatch x86 build it kept the `-mavx512*` copy: `PQCodeDistanceScalar<PQDecoderGeneric>::distance_four_codes` contained 93 `zmm` instructions and was called from `PQCodeDistance<PQDecoderGeneric, SIMDLevel::AVX2>`. On a CPU without AVX512 (AMD Milan, T1_MLN), IVFPQ search with a non-8-bit PQ (`PQDecoderGeneric`, also `PQDecoder16`) on the AVX2 path raises SIGILL.

Found by running the faiss C++ tests under the vendored `qemu-x86_64`, which has no AVX512: `IVFPQFastScan.*` (`test_fast_scan_distance_to_code`) died with SIGILL at `vmovdqa64 ...,%zmm0` in that function. Same class as S672695 (AVX512 code reached on a CPU without it); found before any non-AVX512 host was added to the Cogwheel suite.

Fix: template `PQCodeDistanceScalar` on `SIMDLevel` as well, as the DD migration guide prescribes for code compiled per level, so each TU's instantiation is a distinct symbol. `PQCodeDistance<D, SL>` passes its level; the NONE and ARM_NEON specializations in `pq_code_distance-generic.h` pass theirs. Applied to both definitions (`impl/pq_code_distance/pq_code_distance-inl.h` and `utils/pq_code_distance.h`). `SL` defaults to `SIMDLevel::NONE` so code outside faiss that names only the decoder keeps compiling: `unicorn/features/encoders/PQFSTable.h` uses `faiss::PQCodeDistanceScalar<PQDecoderT>`, and without the default 15 Unicorn targets failed to build on this diff.

Reviewed By: mnorris11

Differential Revision: D124399620

fbshipit-source-id: ece1baf0e1eba39712f881976dea33fe22b265c1
2026-10-10 01:36:08 -07:00
Weixiang Hong f46050eb56 Deflake test_io TestIVFPQRead.test_reader (#5718)
Summary:
Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5718

`test_reader` searches one IVFPQ index loaded twice, with and without the precomputed table, and required the neighbor ids to be identical. The two paths add the distance terms in a different order, so near-tied neighbors can swap. With unseeded data it failed in 7 of 400 random draws, and once in seven runs of the Cogwheel faiss suite on T1_BGM (2 of 10000 ids).

Compare with `check_ref_knn_with_draws`, which allows swaps only among tied distances, and seed the data so a failure reproduces. `test_io` moves to `py_tests_with_contrib` for the helper.

Reviewed By: mnorris11

Differential Revision: D124383290

fbshipit-source-id: ceb81262ffaba81ac23e5a592e1658741192060d
2026-10-09 16:25:26 -07:00
Weixiang Hong db6aa06d96 Hoist per-list allocations out of the IVF fast-scan search loops (#5717)
Summary:
Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5717

`IndexIVFFastScan::search_implem_12` and `search_implem_14` allocate a fresh `AlignedTable<uint8_t> LUT`, `q_map` and `lut_entries` for every inverted list they visit, and call `get_block_stride()`, which heap-allocates a `CodePacker` on each call; `search_implem_10` makes the same `get_block_stride()` call per list. These loops run on OpenMP worker threads, so with jemalloc heap profiling on (the production default for task 0 of every job) on aarch64 every sampled allocation pays libunwind's slow fallback at the `__kmp_invoke_microtask` frame.

gdb thread snapshots of the D124204736 build under profiling (`hr_ivf1024_rabitqfs4`, Grace): of 206 threads caught in jemalloc's `prof_backtrace`, 163 were in `posix_memalign` called from `search_implem_12` (the per-list LUT), 6 in its per-list vectors and 4 in `get_block_stride()`.

This allocates those buffers once per call / per thread, sized for the largest batch (`qbs2` queries x `dim12`), and computes the block stride once. Each iteration still uses only the first `nc` entries, so the search is unchanged.

Reviewed By: junjieqi

Differential Revision: D124387842

fbshipit-source-id: 6bcfe66270fcc0b208b0c88aaf578fea952f6174
2026-10-09 14:58:46 -07:00
Weixiang Hong 44afa7729d Stop allocating per (query, probe) on OpenMP workers in IVF RaBitQ fast-scan (#5716)
Summary:
Pull Request resolved: https://github.com/facebookresearch/faiss/pull/5716

`IndexIVFRaBitQFastScan` heap-allocated up to three times per (query, probe) inside the OpenMP parallel regions of `compute_LUT` / `compute_LUT_uint8`:
- `compute_residual_LUT` built a fresh `std::vector<uint8_t> rotated_qq(d)` on every call.
- For multi-bit indexes, it copied `rotated_q` (d floats) into a temporary `QueryFactorsData`.
- `compute_LUT_uint8` then copied that temporary into `context.query_factors[ij]` instead of moving it.

Under jemalloc heap profiling on aarch64, which Tupperware enables by default in task 0 of every job, each sampled allocation on an OpenMP worker unwinds through libomp's `__kmp_invoke_microtask`. That frame has no CFI, so libunwind's fallback maps and unmaps a whole ELF image per sample, serialised on `mmap_lock` (same mechanism as D123990523; details in that diff). After D123990523, a 74-scenario ServiceLab sweep with profiling on still showed IVF RaBitQ fast-scan at up to 30k (1-bit) and 57k (multi-bit) minor faults per pass, with up to 58 of 72 threads blocked.

gdb stack snapshots of a profiled `hr_ivf1024_rabitqfs4` run (348 threads caught inside `je_prof_backtrace`) attributed 149 samples to the `rotated_q` copy and 31 to `rotated_qq` inside `compute_residual_LUT`. The second copy, in `compute_LUT_uint8`, is visible in the code; its frames were not symbolised in the snapshot.

This diff:
- makes `rotated_qq` a per-thread buffer, like `rotated_q` and `centroid_buf` already are;
- stores multi-bit rotated queries in one block allocated once per search in `search_preassigned`, outside the parallel region and uninitialised. Each (query, probe) writes its slice and the handler reads it by storage index, via a new `FastScanDistancePostProcessing::rotated_q` that is offset per thread slice like `query_factors`;
- moves rather than copies into `context.query_factors[ij]`;
- keeps the capacity of a reused `QueryFactorsData` in `compute_residual_LUT`, so the single-query `IVFRaBitQFastScanScanner` no longer reallocates `rotated_q` on every `set_list`.

Memory is unchanged: the same n × nprobe × d floats as before, now contiguous instead of in n × nprobe separate vectors.

**Effect** (devbig334 Grace, Tupperware-matched environment, heap profiling on as in production, nq=10000, single runs on a busy shared host; recall unchanged):

| hr_ivf1024 | nprobe | QPS before → after | minor faults per pass before → after |
|---|---|---|---|
| rabitqfs4 (multi-bit) | 16 | 27,032 → 36,069 | 4,006 → 1,911 |
| rabitqfs4 (multi-bit) | 64 | 10,168 → 23,295 | 14,932 → 5,822 |
| rabitqfs4 (multi-bit) | 256 | 2,883 → 6,438 | 49,841 → 20,371 |
| rabitqfs (1-bit) | 256 | 3,179 → 6,854 | 23,935 → 20,654 |

Profiling off, rabitqfs4 at nprobe=256: 14,337 → 13,151 QPS, within this host's noise.

**What this does not fix.** With these allocations gone, 1-bit and multi-bit converge on about 20k faults per pass at nprobe=256. gdb attributes 163 of the 206 remaining samples to `posix_memalign` in `IndexIVFFastScan::search_implem_12` on worker threads: the LUT tables (`dis_tables`, n × nprobe × LUT size) that each thread slice allocates. Those allocations are needed, and jemalloc samples them by bytes. The remedy for that remainder is unwind info for libomp's aarch64 `__kmp_invoke_microtask`, being raised separately, rather than more FAISS changes.

Reviewed By: alibeklfc

Differential Revision: D124204736

fbshipit-source-id: 28be93c1cabbcb679bd5328808207fd8e24828b8
2026-10-09 11:11:55 -07:00
Weixiang Hong 6692d7ab37 Stop copying QueryFactorsData per block in RaBitQHeapHandler
Summary:
`RaBitQHeapHandler::handle()` in `IndexRaBitQFastScan.h` copied `rabitq_utils::QueryFactorsData` by value once per (query, 32-vector block). The struct owns `std::vector<float> rotated_q`, so each copy is a heap malloc + free inside the scan's innermost loop: about 312M allocations per 10,000-query pass on sift-1M.

It now binds a `const&` to `context->query_factors[q]`, or to a function-local `static const` empty instance when there is no context. `compute_1bit_adjusted_distance` already takes `const&`, so nothing downstream changes.

**Why this showed up as an ARM "page-fault storm"**
The copy is wasteful on every platform. It becomes catastrophic with jemalloc heap profiling on aarch64, when the allocation happens on an OpenMP worker thread:
- Heap profiling is on by default in task 0 of every Tupperware job: the TW scheduler appends `prof:true,prof_final:false,prof_prefix:/tmp/jeprof` when `taskID % contHeapEnabledRatio == 0` (`tupperware/scheduler/JobSpecUtil.cpp`). jemalloc then backtraces one allocation per 512 KiB on average (`lg_prof_sample` 19), through libunwind on aarch64 (`config.prof_libunwind=1`).
- On OpenMP workers the stack runs through `__kmp_invoke_microtask`, libomp's assembly entry, which has no usable CFI. DWARF unwinding fails there, and libunwind's aarch64 fallback (`get_frame_state`) calls `get_proc_name`, which maps the frame's whole ELF image and unmaps it on every call. Every map and unmap serialises on `mmap_lock`.
- gdb on a reproducer without FAISS: all 95 `get_proc_name` fallbacks hit `__kmp_invoke_microtask`; the same allocations on plain pthreads hit none.
- The fault count tracks the sampling rate exactly: 3,053,990 / 1,525,976 / 764,008 faults per pass at `lg_prof_sample` 18 / 19 / 20. That is about 305k samples per pass (160 GB of 512 B copies ÷ 512 KiB) at roughly 5 faults each.

**Numbers**, hr_rabitqfs, sift-1M, recall 0.3567 throughout.

ServiceLab Grace (72 threads, heap profiling on as in production), nq=10000:

| | QPS | minor faults / pass |
|---|---|---|
| before | 91 | 1.53M |
| after (E2362759441136789) | 17,186 | 858 |

devbig334 Neoverse-V2, environment matched to Tupperware (no `KMP_*` variables), nq=10000, single runs on a busy shared host, so treat as indicative:

| | QPS | kcycles/query | minor faults / pass |
|---|---|---|---|
| before, profiling off | 9,342 | 13,223 | 322 |
| after, profiling off | 11,083 | 11,831 | 341 |
| before, profiling on | 93 | 33,563 | 1,525,976 |
| after, profiling on | 10,970 | 11,756 | 840 |

Without profiling the copy alone costs about 10% more cycles per query.

**Not fixed here: `IndexIVFRaBitQFastScan`.** Its scan already reads query factors by reference, but a ServiceLab sweep with this fix and profiling on still shows up to 30k (1-bit) and 57k (multi-bit) faults per pass with up to 58 threads blocked. The allocation site is being located and will be fixed in a separate diff.

Reviewed By: mnorris11

Differential Revision: D123990523

fbshipit-source-id: 0112bdbe78ac0fd647ea6199b00000193f96788e
2026-10-08 13:19:13 -07:00