Skip to content

fix: hybridRank stack overflow on large candidate sets#62

Open
mags-sully wants to merge 1 commit into
JuliusBrussee:mainfrom
mags-sully:fix/hybridrank-spread-stack-overflow
Open

fix: hybridRank stack overflow on large candidate sets#62
mags-sully wants to merge 1 commit into
JuliusBrussee:mainfrom
mags-sully:fix/hybridrank-spread-stack-overflow

Conversation

@mags-sully

Copy link
Copy Markdown

Problem

hybridRank computes score extrema with Math.min(...bm25s) / Math.max(...bm25s). Spread passes every element as a call argument, and past V8's argument limit (~100k) that throws RangeError: Maximum call stack size exceeded.

Because MemoryStore.search() feeds the entire embedding corpus into hybridRank (all rows from allEmbeddings() merged with keyword hits), any store past ~100k observations crashes on every search — CLI and MCP alike. Hit in the wild on a real store: 480MB / 138k observations after ~3 weeks of heavy Claude Code use.

Fix

Compute extrema with a single loop over items. Identical behavior (including the empty-input case: ranges default via || 1 as before), no argument-count ceiling, and one pass instead of two intermediate arrays.

Evidence

New regression test with 150k candidates:

  • before: RangeError: Maximum call stack size exceeded
  • after: 4/4 tests pass (pnpm vitest run test/ranker.test.ts)

Possible follow-ups (not in this PR)

  • search() is O(corpus) per query — cosine-scoring every embedding then ranking the full merged set. A top-k cap before hybridRank would bound both latency and memory as stores grow.
  • There's currently no retention/pruning mechanism, so data.db grows without bound (~7k observations/day under heavy agent use). A retention window setting would pair well with the cap.

Happy to take a swing at either if you're interested.

Math.min(...arr)/Math.max(...arr) spreads every candidate as a call
argument; past V8's argument limit (~100k items) search() crashes with
'RangeError: Maximum call stack size exceeded'. search() ranks the
entire embedding corpus, so any store past ~100k observations hits
this on every query. Compute extrema with a loop instead — identical
behavior, no argument-count ceiling.

Repro: 480MB real-world store (138k observations) crashed on every
'cavemem search'; regression test with 150k candidates fails with
RangeError before this change and passes after.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant