* ci: add Mac Intel (x86_64) build support * fix: auto-detect Homebrew path for Intel vs Apple Silicon Macs This fixes the hardcoded /opt/homebrew path which only works on Apple Silicon Macs. Intel Macs use /usr/local as the Homebrew prefix. * fix: auto-detect Homebrew paths for both DiskANN and HNSW backends - Fix DiskANN CMakeLists.txt path reference - Add macOS environment variable detection for OpenMP_ROOT - Support both Intel (/usr/local) and Apple Silicon (/opt/homebrew) paths * fix: improve macOS build reliability with proper OpenMP path detection - Add proper CMAKE_PREFIX_PATH and OpenMP_ROOT detection for both Intel and Apple Silicon Macs - Set LDFLAGS and CPPFLAGS for all Homebrew packages to ensure CMake can find them - Apply CMAKE_ARGS to both HNSW and DiskANN backends for consistent builds - Fix hardcoded paths that caused build failures on Intel Macs (macos-13) 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: add abseil library path for protobuf compilation on macOS - Include abseil in CMAKE_PREFIX_PATH for both Intel and Apple Silicon Macs - Add explicit absl_DIR CMake variable to help find abseil for protobuf - Fixes 'absl/log/absl_log.h' file not found error during compilation 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: add abseil include path to CPPFLAGS for both Intel and Apple Silicon - Add -I/opt/homebrew/opt/abseil/include to CPPFLAGS for Apple Silicon - Add -I/usr/local/opt/abseil/include to CPPFLAGS for Intel - Fixes 'absl/log/absl_log.h' file not found by ensuring abseil headers are in compiler include path Root cause: CMAKE_PREFIX_PATH alone wasn't sufficient - compiler needs explicit -I flags 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: clean build system and Python 3.9 compatibility Build system improvements: - Simplify macOS environment detection using brew --prefix - Remove complex hardcoded paths and CMAKE_ARGS - Let CMake automatically find Homebrew packages via CMAKE_PREFIX_PATH - Clean separation between Intel (/usr/local) and Apple Silicon (/opt/homebrew) Python 3.9 compatibility: - Set ruff target-version to py39 to match project requirements - Replace str | None with Union[str, None] in type annotations - Add Union imports where needed - Fix core interface, CLI, chat, and embedding server files 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: type * fix: ensure CMAKE_PREFIX_PATH is passed to backend builds - Add CMAKE_ARGS with CMAKE_PREFIX_PATH and OpenMP_ROOT for both HNSW and DiskANN backends - This ensures CMake can find Homebrew packages on both Intel (/usr/local) and Apple Silicon (/opt/homebrew) - Fixes the issue where CMake was still looking for hardcoded paths instead of using detected ones 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: configure CMake paths in pyproject.toml for proper Homebrew detection - Add CMAKE_PREFIX_PATH and OpenMP_ROOT environment variable mapping in both backends - Remove CMAKE_ARGS from GitHub Actions workflow (cleaner separation) - Ensure scikit-build-core correctly uses environment variables for CMake configuration - This should fix the hardcoded /opt/homebrew paths on Intel Macs 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: remove hardcoded /opt/homebrew paths from DiskANN CMake - Auto-detect Homebrew libomp path using OpenMP_ROOT environment variable - Fallback to CMAKE_PREFIX_PATH/opt/libomp if OpenMP_ROOT not set - Final fallback to brew --prefix libomp for auto-detection - Maintains backwards compatibility with old hardcoded path - Fixes Intel Mac builds that were failing due to hardcoded Apple Silicon paths 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: update DiskANN submodule with macOS Intel/Apple Silicon compatibility fixes - Auto-detect Homebrew libomp path using OpenMP_ROOT environment variable - Exclude mkl_set_num_threads on macOS (uses Accelerate framework instead of MKL) - Fixes compilation on Intel Macs by using correct /usr/local paths 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: update DiskANN submodule with SIMD function name corrections - Fix _mm128_loadu_ps to _mm_loadu_ps (and similar functions) - This is a known issue in upstream DiskANN code where incorrect function names were used - Resolves compilation errors on macOS Intel builds References: Known DiskANN issue with SIMD intrinsics naming 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: update DiskANN submodule with type cast fix for signed char templates - Add missing type casts (float*)a and (float*)b in SSE2 version - This matches the existing type casts in the AVX version - Fixes compilation error when instantiating DistanceInnerProduct<int8_t> - Resolves "cannot initialize const float* with const signed char*" error 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: update Faiss submodule with override keyword fix - Add missing override keyword to IDSelectorModulo::is_member function - Fixes C++ compilation warning that was treated as error due to -Werror flag - Resolves "warning: 'is_member' overrides a member function but is not marked 'override'" - Improves code conformance to modern C++ best practices 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: update Faiss submodule with override keyword fix * fix: update DiskANN submodule with additional type cast fix - Add missing type cast in DistanceFastL2::norm function SSE2 version - Fixes const float* = const signed char* compilation error - Ensures consistent type casting across all SIMD code paths - Resolves template instantiation error for DistanceFastL2<int8_t> 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * debug: simplify wheel compatibility checking - Fix YAML syntax error in debug step - Use simpler approach to show platform tags and wheel names - This will help identify platform tag compatibility issues 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: use correct Python version for wheel builds - Replace --python python with --python ${{ matrix.python }} - This ensures wheels are built for the correct Python version in each matrix job - Fixes Python version mismatch where cp39 wheels were used in cp311 environments 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: resolve wheel installation conflicts in CI matrix builds Fix issue where multiple Python versions' wheels in the same dist directory caused installation conflicts during CI testing. The problem occurred when matrix builds for different Python versions accumulated wheels in shared directories, and uv pip install would find incompatible wheels. Changes: - Add Python version detection using matrix.python variable - Convert Python version to wheel tag format (e.g., 3.11 -> cp311) - Use find with version-specific pattern matching to select correct wheels - Add explicit error handling if no matching wheel is found This ensures each CI job installs only wheels compatible with its specific Python version, preventing "A path dependency is incompatible with the current platform" errors. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: ensure virtual environment uses correct Python version in CI Fix issue where uv venv was creating virtual environments with a different Python version than specified in the matrix, causing wheel compatibility errors. The problem occurred when the system had multiple Python versions and uv venv defaulted to a different version than intended. Changes: - Add --python ${{ matrix.python }} flag to uv venv command - Ensures virtual environment matches the matrix-specified Python version - Fixes "The wheel is compatible with CPython 3.X but you're using CPython 3.Y" errors This ensures wheel installation selects and installs the correctly built wheels that match the runtime Python version. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: complete Python 3.9 type annotation compatibility fixes Fix remaining Python 3.9 incompatible type annotations throughout the leann-core package that were causing test failures in CI. The union operator (|) syntax for type hints was introduced in Python 3.10 and causes "TypeError: unsupported operand type(s) for |" errors in Python 3.9. Changes: - Convert dict[str, Any] | None to Optional[dict[str, Any]] - Convert int | None to Optional[int] - Convert subprocess.Popen | None to Optional[subprocess.Popen] - Convert LeannBackendFactoryInterface | None to Optional[LeannBackendFactoryInterface] - Add missing Optional imports to all affected files This resolves all test failures related to type annotation syntax and ensures compatibility with Python 3.9 as specified in pyproject.toml. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: complete Python 3.9 type annotation fixes in backend packages Fix remaining Python 3.9 incompatible type annotations in backend packages that were causing test failures. The union operator (|) syntax for type hints was introduced in Python 3.10 and causes "TypeError: unsupported operand type(s) for |" errors in Python 3.9. Changes in leann-backend-diskann: - Convert zmq_port: int | None to Optional[int] in diskann_backend.py - Convert passages_file: str | None to Optional[str] in diskann_embedding_server.py - Add Optional imports to both files Changes in leann-backend-hnsw: - Convert zmq_port: int | None to Optional[int] in hnsw_backend.py - Add Optional import This resolves the final test failures related to type annotation syntax and ensures full Python 3.9 compatibility across all packages. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: remove Python 3.10+ zip strict parameter for Python 3.9 compatibility Remove the strict=False parameter from zip() call in api.py as it was introduced in Python 3.10 and causes "TypeError: zip() takes no keyword arguments" in Python 3.9. The strict parameter controls whether zip() raises an exception when the iterables have different lengths. Since we're not relying on this behavior and the code works correctly without it, removing it maintains the same functionality while ensuring Python 3.9 compatibility. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: ensure leann-core package is built on all platforms, not just Ubuntu This fixes the issue where CI was installing leann-core from PyPI instead of using locally built package with Python 3.9 compatibility fixes. * fix: build and install leann meta package on all platforms The leann meta package is pure Python and platform-independent, so there's no reason to restrict it to Ubuntu only. This ensures all platforms use consistent local builds instead of falling back to PyPI versions. * fix: restrict MLX dependencies to Apple Silicon Macs only MLX framework only supports Apple Silicon (ARM64) Macs, not Intel x86_64. Add platform_machine == 'arm64' condition to prevent installation failures on Intel Macs (macos-13). * cleanup: simplify CI configuration - Remove debug step with non-existent 'uv pip debug' command - Simplify wheel installation logic - let uv handle compatibility - Use -e .[test] instead of manually listing all test dependencies * fix: install backend wheels before meta packages Install backend wheels first to ensure they're available when core/meta packages are installed, preventing uv from trying to resolve backend dependencies from PyPI. * fix: use local leann-core when building backend packages Add --find-links to backend builds to ensure they use the locally built leann-core with fixed MLX dependencies instead of downloading from PyPI. Also bump leann-core version to 0.2.8 to ensure clean dependency resolution. * fix: use absolute path for find-links and upgrade backend version - Use GITHUB_WORKSPACE for absolute path to ensure find-links works - Upgrade leann-backend-hnsw to 0.2.8 to match leann-core version * fix: use absolute path for find-links and upgrade backend version - Use GITHUB_WORKSPACE for absolute path to ensure find-links works - Upgrade leann-backend-hnsw to 0.2.8 to match leann-core version * fix: correct version consistency for --find-links to work properly - All packages now use version 0.2.7 consistently - Backend packages can find exact leann-core==0.2.7 from local build - This ensures --find-links works during CI builds instead of falling back to PyPI 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: revert all packages to consistent version 0.2.7 - This PR should not bump versions, only fix Intel Mac build - Version bumps should be done in release_manual workflow - All packages now use 0.2.7 consistently for --find-links to work 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: use --find-links during package installation to avoid PyPI MLX conflicts - Backend wheels contain Requires-Dist: leann-core==0.2.7 - Without --find-links, uv resolves this from PyPI which has MLX for all Darwin - With --find-links, uv uses local leann-core with proper platform restrictions - Root cause: dependency resolution happens at install time, not just build time - Local test confirms this fixes Intel Mac MLX dependency issues 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: restrict MLX dependencies to ARM64 Macs in workspace pyproject.toml - Root pyproject.toml also had MLX dependencies without platform_machine restriction - This caused test dependency installation to fail on Intel Macs - Now consistent with packages/leann-core/pyproject.toml platform restrictions 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * chore: cleanup unused files and fix GitHub Actions warnings - Remove unused packages/leann-backend-diskann/CMakeLists.txt (DiskANN uses cmake.source-dir=third_party/DiskANN instead) - Replace macos-latest with macos-14 to avoid migration warnings (macos-latest will migrate to macOS 15 on August 4, 2025) - Keep packages/leann-backend-hnsw/CMakeLists.txt (needed for Faiss config) 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: properly handle Python 3.13 support with PyTorch compatibility - Support Python 3.13 on most platforms (Ubuntu, ARM64 Mac) - Exclude Intel Mac + Python 3.13 combination due to PyTorch wheel availability - PyTorch <2.5 supports Intel Mac but not Python 3.13 - PyTorch 2.5+ supports Python 3.13 but not Intel Mac x86_64 - Document limitation in CI configuration comments - Update README badges with detailed Python version support and CI status 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
248 lines
9.4 KiB
Python
248 lines
9.4 KiB
Python
import logging
|
|
import os
|
|
import shutil
|
|
from pathlib import Path
|
|
from typing import Any, Literal, Optional
|
|
|
|
import numpy as np
|
|
from leann.interface import (
|
|
LeannBackendBuilderInterface,
|
|
LeannBackendFactoryInterface,
|
|
LeannBackendSearcherInterface,
|
|
)
|
|
from leann.registry import register_backend
|
|
from leann.searcher_base import BaseSearcher
|
|
|
|
from .convert_to_csr import convert_hnsw_graph_to_csr
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
def get_metric_map():
|
|
from . import faiss # type: ignore
|
|
|
|
return {
|
|
"mips": faiss.METRIC_INNER_PRODUCT,
|
|
"l2": faiss.METRIC_L2,
|
|
"cosine": faiss.METRIC_INNER_PRODUCT,
|
|
}
|
|
|
|
|
|
def normalize_l2(data: np.ndarray) -> np.ndarray:
|
|
norms = np.linalg.norm(data, axis=1, keepdims=True)
|
|
norms[norms == 0] = 1 # Avoid division by zero
|
|
return data / norms
|
|
|
|
|
|
@register_backend("hnsw")
|
|
class HNSWBackend(LeannBackendFactoryInterface):
|
|
@staticmethod
|
|
def builder(**kwargs) -> LeannBackendBuilderInterface:
|
|
return HNSWBuilder(**kwargs)
|
|
|
|
@staticmethod
|
|
def searcher(index_path: str, **kwargs) -> LeannBackendSearcherInterface:
|
|
return HNSWSearcher(index_path, **kwargs)
|
|
|
|
|
|
class HNSWBuilder(LeannBackendBuilderInterface):
|
|
def __init__(self, **kwargs):
|
|
self.build_params = kwargs.copy()
|
|
self.is_compact = self.build_params.setdefault("is_compact", True)
|
|
self.is_recompute = self.build_params.setdefault("is_recompute", True)
|
|
self.M = self.build_params.setdefault("M", 32)
|
|
self.efConstruction = self.build_params.setdefault("efConstruction", 200)
|
|
self.distance_metric = self.build_params.setdefault("distance_metric", "mips")
|
|
self.dimensions = self.build_params.get("dimensions")
|
|
if not self.is_recompute:
|
|
if self.is_compact:
|
|
# TODO: support this case @andy
|
|
raise ValueError(
|
|
"is_recompute is False, but is_compact is True. This is not compatible now. change is compact to False and you can use the original HNSW index."
|
|
)
|
|
|
|
def build(self, data: np.ndarray, ids: list[str], index_path: str, **kwargs):
|
|
from . import faiss # type: ignore
|
|
|
|
path = Path(index_path)
|
|
index_dir = path.parent
|
|
index_prefix = path.stem
|
|
index_dir.mkdir(parents=True, exist_ok=True)
|
|
|
|
if data.dtype != np.float32:
|
|
logger.warning(f"Converting data to float32, shape: {data.shape}")
|
|
data = data.astype(np.float32)
|
|
|
|
metric_enum = get_metric_map().get(self.distance_metric.lower())
|
|
if metric_enum is None:
|
|
raise ValueError(f"Unsupported distance_metric '{self.distance_metric}'.")
|
|
|
|
dim = self.dimensions or data.shape[1]
|
|
index = faiss.IndexHNSWFlat(dim, self.M, metric_enum)
|
|
index.hnsw.efConstruction = self.efConstruction
|
|
|
|
if self.distance_metric.lower() == "cosine":
|
|
data = normalize_l2(data)
|
|
|
|
index.add(data.shape[0], faiss.swig_ptr(data))
|
|
index_file = index_dir / f"{index_prefix}.index"
|
|
faiss.write_index(index, str(index_file))
|
|
|
|
if self.is_compact:
|
|
self._convert_to_csr(index_file)
|
|
|
|
def _convert_to_csr(self, index_file: Path):
|
|
"""Convert built index to CSR format"""
|
|
mode_str = "CSR-pruned" if self.is_recompute else "CSR-standard"
|
|
logger.info(f"INFO: Converting HNSW index to {mode_str} format...")
|
|
|
|
csr_temp_file = index_file.with_suffix(".csr.tmp")
|
|
|
|
success = convert_hnsw_graph_to_csr(
|
|
str(index_file), str(csr_temp_file), prune_embeddings=self.is_recompute
|
|
)
|
|
|
|
if success:
|
|
logger.info("✅ CSR conversion successful.")
|
|
# index_file_old = index_file.with_suffix(".old")
|
|
# shutil.move(str(index_file), str(index_file_old))
|
|
shutil.move(str(csr_temp_file), str(index_file))
|
|
logger.info(f"INFO: Replaced original index with {mode_str} version at '{index_file}'")
|
|
else:
|
|
# Clean up and fail fast
|
|
if csr_temp_file.exists():
|
|
os.remove(csr_temp_file)
|
|
raise RuntimeError("CSR conversion failed - cannot proceed with compact format")
|
|
|
|
|
|
class HNSWSearcher(BaseSearcher):
|
|
def __init__(self, index_path: str, **kwargs):
|
|
super().__init__(
|
|
index_path,
|
|
backend_module_name="leann_backend_hnsw.hnsw_embedding_server",
|
|
**kwargs,
|
|
)
|
|
from . import faiss # type: ignore
|
|
|
|
self.distance_metric = (
|
|
self.meta.get("backend_kwargs", {}).get("distance_metric", "mips").lower()
|
|
)
|
|
metric_enum = get_metric_map().get(self.distance_metric)
|
|
if metric_enum is None:
|
|
raise ValueError(f"Unsupported distance_metric '{self.distance_metric}'.")
|
|
|
|
self.is_compact, self.is_pruned = (
|
|
self.meta.get("is_compact", True),
|
|
self.meta.get("is_pruned", True),
|
|
)
|
|
|
|
index_file = self.index_dir / f"{self.index_path.stem}.index"
|
|
if not index_file.exists():
|
|
raise FileNotFoundError(f"HNSW index file not found at {index_file}")
|
|
|
|
hnsw_config = faiss.HNSWIndexConfig()
|
|
hnsw_config.is_compact = self.is_compact
|
|
hnsw_config.is_recompute = (
|
|
self.is_pruned
|
|
) # In C++ code, it's called is_recompute, but it's only for loading IIUC.
|
|
|
|
self._index = faiss.read_index(str(index_file), faiss.IO_FLAG_MMAP, hnsw_config)
|
|
|
|
def search(
|
|
self,
|
|
query: np.ndarray,
|
|
top_k: int,
|
|
zmq_port: Optional[int] = None,
|
|
complexity: int = 64,
|
|
beam_width: int = 1,
|
|
prune_ratio: float = 0.0,
|
|
recompute_embeddings: bool = True,
|
|
pruning_strategy: Literal["global", "local", "proportional"] = "global",
|
|
batch_size: int = 0,
|
|
**kwargs,
|
|
) -> dict[str, Any]:
|
|
"""
|
|
Search for nearest neighbors using HNSW index.
|
|
|
|
Args:
|
|
query: Query vectors (B, D) where B is batch size, D is dimension
|
|
top_k: Number of nearest neighbors to return
|
|
complexity: Search complexity/efSearch, higher = more accurate but slower
|
|
beam_width: Number of parallel search paths/beam_size
|
|
prune_ratio: Ratio of neighbors to prune via PQ (0.0-1.0)
|
|
recompute_embeddings: Whether to fetch fresh embeddings from server
|
|
pruning_strategy: PQ candidate selection strategy:
|
|
- "global": Use global PQ queue size for selection (default)
|
|
- "local": Local pruning, sort and select best candidates
|
|
- "proportional": Base selection on new neighbor count ratio
|
|
zmq_port: ZMQ port for embedding server communication. Must be provided if recompute_embeddings is True.
|
|
batch_size: Neighbor processing batch size, 0=disabled (HNSW-specific)
|
|
**kwargs: Additional HNSW-specific parameters (for legacy compatibility)
|
|
|
|
Returns:
|
|
Dict with 'labels' (list of lists) and 'distances' (ndarray)
|
|
"""
|
|
from . import faiss # type: ignore
|
|
|
|
if not recompute_embeddings:
|
|
if self.is_pruned:
|
|
raise RuntimeError("Recompute is required for pruned index.")
|
|
if recompute_embeddings:
|
|
if zmq_port is None:
|
|
raise ValueError("zmq_port must be provided if recompute_embeddings is True")
|
|
|
|
if query.dtype != np.float32:
|
|
query = query.astype(np.float32)
|
|
if self.distance_metric == "cosine":
|
|
query = normalize_l2(query)
|
|
|
|
params = faiss.SearchParametersHNSW()
|
|
if zmq_port is not None:
|
|
params.zmq_port = zmq_port # C++ code won't use this if recompute_embeddings is False
|
|
params.efSearch = complexity
|
|
params.beam_size = beam_width
|
|
|
|
# For OpenAI embeddings with cosine distance, disable relative distance check
|
|
# This prevents early termination when all scores are in a narrow range
|
|
embedding_model = self.meta.get("embedding_model", "").lower()
|
|
if self.distance_metric == "cosine" and any(
|
|
openai_model in embedding_model for openai_model in ["text-embedding", "openai"]
|
|
):
|
|
params.check_relative_distance = False
|
|
else:
|
|
params.check_relative_distance = True
|
|
|
|
# PQ pruning: direct mapping to HNSW's pq_pruning_ratio
|
|
params.pq_pruning_ratio = prune_ratio
|
|
|
|
# Map pruning_strategy to HNSW parameters
|
|
if pruning_strategy == "local":
|
|
params.local_prune = True
|
|
params.send_neigh_times_ratio = 0.0
|
|
elif pruning_strategy == "proportional":
|
|
params.local_prune = False
|
|
params.send_neigh_times_ratio = 1.0 # Any value > 1e-6 triggers proportional mode
|
|
else: # "global"
|
|
params.local_prune = False
|
|
params.send_neigh_times_ratio = 0.0
|
|
|
|
# HNSW-specific batch processing parameter
|
|
params.batch_size = batch_size
|
|
|
|
batch_size_query = query.shape[0]
|
|
distances = np.empty((batch_size_query, top_k), dtype=np.float32)
|
|
labels = np.empty((batch_size_query, top_k), dtype=np.int64)
|
|
|
|
self._index.search(
|
|
query.shape[0],
|
|
faiss.swig_ptr(query),
|
|
top_k,
|
|
faiss.swig_ptr(distances),
|
|
faiss.swig_ptr(labels),
|
|
params,
|
|
)
|
|
|
|
string_labels = [[str(int_label) for int_label in batch_labels] for batch_labels in labels]
|
|
|
|
return {"labels": string_labels, "distances": distances}
|