Skip to content
This repository was archived by the owner on Oct 6, 2026. It is now read-only.

filter - #154

Closed
shaleenji wants to merge 116 commits into
masterfrom
filter_pass
Closed

filter#154
shaleenji wants to merge 116 commits into
masterfrom
filter_pass

Conversation

@shaleenji

Copy link
Copy Markdown
Collaborator

No description provided.

shaleengarg and others added 30 commits January 24, 2026 21:04
Update install.sh
…registry image. (#6)

Co-authored-by: Burhan Kapdawala <burhan@endee.io>
…d_correction

Fix: Delete vectors from sparse storage when deleting by filter
feat: Implement backup download and upload API endpoints
quantization optimization and bug fixes
clean of DISABLE_HYBRID_QUANTIZATION and fixing of binary index bug
* first commit

* readme update

---------

Co-authored-by: Shaleen Garg <shaleengarg.in@gmail.com>
* filter aware search

* multiple entry points to base layer

* progressive filtering approach

* budget based pruning

* filter_params
* fix:sparse storage db path correction

* Apply suggestion from @Copilot

change variable name

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: rajeshkomaravelli <rajeshkomuravelli333@gmail.com>
Co-authored-by: Vineet Dwivedi <164136199+vindwid@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* backfill
* vector cache - dynamic hot cache
add meta and filter in application/json in insert api
Co-authored-by: rajeshkomaravelli <rajeshkomuravelli333@gmail.com>
@shaleenji shaleenji mentioned this pull request May 8, 2026
@shaleenji

Copy link
Copy Markdown
Collaborator Author
  1. Numeric bucket: fix the duplicate-heavy cliff bug.

When >65,536 ids shared the same numeric filter value, Bucket::serialize
truncated the on-disk count to a uint16_t and corrupted the bucket --
recall collapsed to zero past that boundary (reproducible with
tests/repo_filter.py).

The fix has four parts in src/filter/numeric_index.hpp:

  • Bucket::add now caps deltas/ids at MAX_SIZE for delta_32==0
    duplicates and routes the excess id into summary_bitmap only.
    Cardinality is preserved, on-disk arrays no longer grow without
    bound.

  • The on-disk count field is removed from Bucket::serialize entirely.
    Bucket::deserialize derives nr_array_entries from the residual
    bytes after the bitmap, so there is no count to overflow.

  • Bucket::deserialize stays backward compatible with the old on-disk
    format via a modulus check on the residual: existing DBs written
    by the old code keep working with no migration.

  • range() gains two correctness branches: a legacy-salvage branch
    for cliff-corrupted buckets (ids.empty() but bitmap non-empty),
    and a bitmap-only-inclusion branch in the partial-overlap path
    for buckets with cardinality > ids.size().

  • range() also gains a coarse full-coverage fast path: when a
    bucket's [base, base+MAX_DELTA] extent is wholly inside the query
    range, skip the deltas/ids deserialize and union just the bitmap
    via Bucket::read_summary_bitmap.


  1. Build: pin -falign-functions=64 in the release flags (CMakeLists.txt).

Editing any header transitively included by ndd.hpp (filter.hpp,
numeric_index.hpp, vector_storage.hpp) was producing 10-30% QPS
swings on the int-filter HTTP bench with no algorithmic change,
because the HNSW search loop is sensitive to function placement
relative to cache lines. Microbenches of range() and bitmap.contains()
were byte-identical between affected builds; the cost lived in
i-cache effects on the surrounding HNSW inner loop. Forcing 64-byte
function alignment removes the variance so future header touches
don't masquerade as perf regressions.

@shaleenji

Copy link
Copy Markdown
Collaborator Author

Adjacent issue not addressed here: the slide-split LEFT-bucket rebuild
in add_to_buckets() rebuilds summary_bitmap from ids only, which
silently drops bitmap-only entries. Saturation handling is therefore
not durable across splits; tracked separately.

@shaleenji

Copy link
Copy Markdown
Collaborator Author

THIS IS A Breaking change. Need to reindex

@shaleenji

Copy link
Copy Markdown
Collaborator Author
  1. Now doing metadata updates when filters are updated
  2. Stale entries are removed when updating filters

@shaleenji

shaleenji commented May 12, 2026 •

Copy link
Copy Markdown
Collaborator Author
  • add explicit Roaring bitmap payload validation for category filter reads
  • replace unsafe filter bitmap reads with bounded readSafe + exact byte-count checks
  • validate deserialized Roaring internals before using stored bitmaps
  • apply the same hardening to numeric bucket bitmap payloads, including the range fast path
  • return/propagate corruption as OperationResult code 200 instead of trusting malformed payloads
  • add regression coverage for valid, truncated, trailing-byte, and garbage bitmap payloads

@shaleenji

Copy link
Copy Markdown
Collaborator Author

solves #237 , #238 , #239 , #240 , #241 , #242

partially #244 and #25

@shaleenji

Copy link
Copy Markdown
Collaborator Author

Requires reindexing

…y checks ids.empty() in numeric_index.cpp (line 306), removal deletes the bucket on that basis in numeric_index.cpp (line 412), and range skips ids.empty() buckets in numeric_index.cpp (line 1004). Also, split rebuilds the left bitmap only from ids in numeric_index.cpp (line 626), dropping bitmap-only duplicate IDs.
@shaleenji

Copy link
Copy Markdown
Collaborator Author

Server:
8 CPUs
32GB Memory
100GB NVME SSD

Label Filter

Screenshot 2026-05-15 at 08 24 19

Int Filter

Screenshot 2026-05-15 at 08 25 06

Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.