Local code search for AI coding agents. Six fast, purpose-built tools that hand Claude Code, Codex & friends ranked answers, not raw grep. Zero API keys, 100% on-device.

Maybe grep isn't all you need… 🍬 Every coding agent today reaches for grep + Read by reflex. sweet-search challenges the narrative. 😎

npm GitHub stars license node platforms inference


✨ Highlights

  • Hybrid retrieval — one of the six tools uses BM25F lexical + dense semantic + structural graph signals, fused per query and reranked by late-interaction
  • Agent-native by design — token-budgeted output tiers, an optional MCP server (and default zero-overhead CLI), and a GEPA-evolved system prompt — one init installs it into Claude Code by default (Codex, Gemini CLI, and Cursor via flags)
  • Indexed grep, ~10× faster than ripgrep — a sparse n-gram prefilter skips the files that provably can't match
  • ColBERT-style reranking, locally — per-token MaxSim late interaction on hand-written SIMD kernels
  • GPU-accelerated indexing — Apple Metal, CUDA, CoreML Neural Engine, or plain CPU via ORT; same engine, auto-selected
  • Never stale — incremental indexing keeps the index aligned with your working tree, uncommitted edits included
  • No storage hassle — indexed artifacts maximally optimized without any accuracy tradeoff; up to INT4 quantization
  • Local-first — all models run on-device; nothing is sent anywhere, ever. CPU-inference supported for all models

📚 Table of Contents

GET STARTED

🚀 Quickstart three commands to a searchable repo

🖥️ Platform Support macOS · Linux · WASM fallback

USE IT

🧰 The Six Tools search · grep · find · semantic · trace · read

🧠 The Evolved Agent Prompt GEPA-optimized search discipline

🔌 Works With Your Agent MCP · Claude Code · Codex · Gemini · Cursor

UNDER THE HOOD

⚡ GPU-Accelerated Indexing candle · fused kernels · cAST chunking

🔄 An Index That Never Goes Stale reconcile daemon tracks your working tree

🦀 The Native Engine Room four Rust crates + INT4 LI compression

THE RECEIPTS

📊 Benchmarks agent cost savings · engine speed · full-corpus MRR

🧭 Where sweet-search Fits honest wins & trade-offs vs peers

🙏 Prior Art & Acknowledgements the shoulders we stand on

📄 License Apache-2.0

🚀 Quickstart

Requirements: Node.js 18+ on macOS (Apple silicon or Intel) or Linux (x64 or ARM64). On Windows, run sweet-search inside WSL2.

npm install -g sweet-search

cd your-repo
sweet-search init     # one-time: downloads local models, wires up your agent
sweet-search index    # builds the index — GPU-accelerated where available

sweet-search "where do we validate JWT tokens?"

That's it. init is idempotent and SHA256-verifies every model binary; re-running it is always safe. From then on, the index stays up to date automatically as you work.

sweet-search stores its index in SQLite via better-sqlite3, which downloads its prebuilt native binary from an install script. Recent npm versions block those scripts by default and only print a warning, which leaves the binding missing and makes indexing fail. sweet-search init detects this and stops with instructions rather than reporting success.

For a global install, approve the scripts once:

npm install -g sweet-search --allow-scripts=sweet-search,better-sqlite3

For a project-local install, add the allowlist to that project's package.json (npm rejects the flag for project-scoped installs), then reinstall:

"allowScripts": { "sweet-search": true, "better-sqlite3": true }

Verify with sweet-search init — the native:sqlite check must pass.

For Claude Code, init automatically installs and activates the sweet-search output style, which adds the compact routing override at system-prompt priority. Start a new Claude session or run /clear after init, and keep that output style selected for reliable ss-* routing.

To uninstall 😢:

sweet-search uninstall --keep-models  # current repo only; keeps shared models and the global CLI
sweet-search uninstall                # current repo plus its shared model downloads; keeps the global CLI
npm uninstall -g sweet-search         # global CLI only; does not clean initialized repos

Run sweet-search uninstall inside each initialized repo before removing the global CLI.

sweet-search init --wizard          # interactive: shows your hardware, recommends a model tier
sweet-search init --profile core    # lexical-only, no model downloads (CI-friendly)
sweet-search init --li-model edge   # compact late-interaction model for constrained machines
sweet-search init --agents          # also configure AGENTS.md for Codex/OpenCode
sweet-search init --no-claude --agents # streamlined AGENTS-only configuration
sweet-search uninstall --dry-run    # preview cleanup for the current repo
  • Footprint: CPU-only hosts download a few hundred MB of INT8 models; GPU hosts add ~1.2 GB of FP32 backbones (skipped automatically where they'd be useless); M3+ Macs can additionally fetch a ~3.2 GB CoreML cascade for Neural Engine acceleration. Everything lands in ~/.cache/sweet-search/models/ and is used strictly on-device.
  • Claude Code wiring (default): init leaves CLAUDE.md untouched, writes the verbatim evolved guide to .claude/rules/sweet-search.md, installs .claude/output-styles/sweet-search.md, and selects it in .claude/settings.json. It also registers a session-start prewarm hook and installs the /sweet-index skill.
  • Output-style conflicts: init never silently replaces another selected style. It still installs the Sweet Search style so it appears under /config, then emits a warning. A higher-priority .claude/settings.local.json selection is also detected and reported. Select sweet-search, then run /clear or restart Claude Code.
  • Codex/OpenCode wiring: pass --agents to place the same verbatim guide directly in AGENTS.md. Use --no-claude --agents when AGENTS.md is the only integration you want; --codex additionally installs Codex's project prewarm hook.
  • What gets indexed: what you'd expect — .gitignore is respected, node_modules/build dirs/minified artifacts are denied, files over 1 MB skipped, with a .sweet-search-ignore for extra rules.

Init flags

Flag Behavior
--profile core|full Select lexical-only core or the full model-backed profile.
--li-model standard|edge|none Select the late-interaction model tier or disable it.
--search-reranking auto|on|off Control search-time late-interaction reranking.
--wizard Choose model and reranking settings interactively.
--verify-deep Load modules and verify checksums after setup.
--force Re-download models even when cached.
--build-coreml-cascade Build the optional CoreML cascade locally on eligible Apple silicon.
--skip-coreml-cascade Skip fetching or building the CoreML cascade.
--skip-dedup Skip near-duplicate-detection readiness checks.
--skip-cuda Disable the CUDA backend even when available.
--skip-prewarm-hook Do not register the Claude/Codex session-start prewarm hook.
--agents Also write the verbatim guide to AGENTS.md for Codex/OpenCode.
--codex Add the Codex prewarm hook and project feature flag; implies --agents.
--codex-enable-global-hooks Advanced opt-in: also enable hooks in the user-level Codex config.
--no-claude Write nothing under .claude/; combine with --agents for AGENTS-only setup.
--gemini Also write GEMINI.md (sharing AGENTS.md when enabled).
--cursor Also write .cursor/rules/sweet-search.mdc.
--symlink-instruction-files Explicitly use the default GEMINI.md symlink behavior.
--no-symlink-instruction-files Use a regular GEMINI.md import instead of a symlink.
--no-agent-instructions Skip all agent policy and Claude output-style installation.
--mcp Also register the project MCP server; the CLI remains the default contact surface.
--no-cli With --mcp, give agents the MCP-specific guide and remove the CLI-specific Claude output style.
--enforce-tools Optional strict Claude mode: deny native Grep and hint native Read.
--verbose, -v Print additional setup diagnostics.
--help, -h Show the complete CLI help.

Uninstall flags and cleanup

Flag Behavior
--dry-run Preview all detected removals.
--keep-models Preserve shared model and CoreML caches.
--purge Also uninstall the npm package and Sweet Search native packages.
--force Skip the confirmation prompt.
--help, -h Show uninstall help.

sweet-search uninstall removes Sweet Search's rule, output style and active selection, AGENTS/GEMINI/Cursor instruction blocks, project hooks and skill, MCP registration, enforcement/reminder artifacts, and .sweet-search/. It preserves user-authored content. Generic Codex [features] hooks = true flags are left in place because other tools may share them, and an otherwise-empty settings file may remain as {}.

📊 Benchmarks

We measure sweet-search four ways — from how much it helps a real agent down to raw engine throughput:

🤖 Code-retrieval (agent-in-the-loop) Does it make a real coding agent cheaper and more useful when it searches your repo? Paired against each model's own grep-and-read loop.

🚧 Task-completion (coming soon) Does cheaper, denser context compound into a higher resolve-rate on multi-step engineering tasks? Harness in progress.

📄 Paper-type IR (academic) The standard NL→code retrieval suites (GCSN, M2CRB, CoSQA…), full-corpus MRR@10.

Engine speed Raw systems numbers — grep throughput, query latency, rerank kernels, HNSW.


🤖 1. Code-retrieval benchmarks — the agent-in-the-loop test

One variable changes: how the agent searches a real repository.

  • 🍬 sweet-search: the model gets our GEPA-evolved search discipline and search tools.
  • 🐌 Native: the same model uses its built-in grep-and-read loop.

Same tasks, same judge, paired probe-for-probe.

sealed vault · exact paired results · five representative profiles · full 11-cell matrix and held-out/OOD replication below

The headline, in four claims:

  • 💰 Cheaper where the agent thrashes — up to −34% realized cost on Codex; −18 to −32% across the GPT-5.5 / opencode / bare-API harnesses.
  • 🔧 Fewer round-trips — up to −56% tool calls, significant on 9 of 11 cells.
  • More useful per response+0.18 to +0.31 on a 5-dimension usefulness score, and still denser when length-matched (significant on 8 of 11 cells).
  • 🎯 Accuracy held — and lifted on the weak — a statistical tie on flagship models (saturated at 0.94–0.99), and +3 pp (up to +8 pp out-of-distribution) on weaker models like GLM-5.1 and DeepSeek.

The win is harness-adaptive: where the native loop is disciplined (Claude Code) it shows up as denser, more useful context per token; where it thrashes (Codex floods 30k+ tokens of its own grep output into context) it shows up as a large cost and tool-call cut. Either way, final-answer accuracy never significantly regresses.

🧰 Native agent harness 💰 Realized cost 🔧 Tool calls ✨ Useful content / response 🎯 Final accuracy
🤖 Codex (GPT-5.5) −30 to −34% −44 to −56% +0.06 → +0.17 ↑ tie (saturated)
🐚 opencode (GPT-5.5 / GLM-5.1) −18 to −22% −15 to −49% +0.23 to +0.31 tie
🔌 bare API (GPT-5.5 / GLM / DeepSeek) −15 to −32% ᵃ −15 to −33% +0.08 to +0.24 ↑ tie · +3 pp on weak models
🟣 Claude Code (Sonnet / Opus) −10% to +14% ᵇ −5 to −33% +0.18 to +0.29 ↑ tie

↑ "Useful content / response" is the per-response delta on a 5-dimension usefulness score (answer-grounding · workable-code · navigability · edit-locality · sufficiency), 0–1 scale. "tie" = final-answer correctness statistically indistinguishable (saturated in the 0.94–0.99 band on flagships).ᵃ the two cheapest bare models cost fractions of a cent either way (GLM +27% of $0.008; DeepSeek −15% of $0.004). ᵇ Opus −5/−10%; Sonnet +8–14%, which is ≈1¢ on a flat-rate subscription for a richer answer.

Denser, not just longer. The usefulness lift survives length-matching — comparing sweet-search and native responses of equal token length, sweet-search's content is significantly higher on 8 of 11 cells. The validated single-number usefulness composite (grounding × content × density) is significant on all 11 sealed cells.

  • What's being compared: the installed sweet-search agent prompt + tools vs. the same model using only its built-in file-reading and shell-grep tools. Not a different model — the same model, with and without sweet-search.
  • Design: 11 model×harness cells. Sealed vault (n=60/arm, the pre-registered primary) opened once; plus held-out (n=30) and out-of-distribution (n=40) sets for generalization. Stratified, fixed-seed splits.
  • Judging: 3-judge panel (DeepSeek-V4-flash + Gemini-3.1-flash-lite + MiniMax-M2.7), paired by probe, 20k-sample bootstrap CIs, Benjamini–Hochberg FDR multiplicity correction across each metric family. We report family-level survival counts, never a single cherry-picked cell.
  • What survives FDR (vault): useful-content 10/11, density-composite 11/11, length-matched content 8/11, fewer-tool-calls 9/11. Generalization (held-out + OOD): content 17–18/20, fewer calls 14/20.
  • The token fact that drives everything: sweet-search's footprint is nearly constant (~1.3k–3.3k tokens) because the tool responses are capped; native's footprint is whatever the model decides to grep — up to 37k tokens on Codex. That single fact is what drives the cost and tool-call gaps.
  • Honest caveats we keep attached: (1) accuracy ties on flagship models — it is not an accuracy win there, it's saturated; the accuracy gains are real only on weaker models. (2) The two weakest cells for length-matched density (Codex-low, DeepSeek) are correct-sign but underpowered — Codex's responses are so token-divergent that too few equal-length pairs exist to reach significance, and DeepSeek is simply under-powered. Those are honest non-victories, not wins.
  • Full methodology and per-cell tables: docs/PHASE7.md.

🚧 2. Task-completion benchmarks — coming soon

Retrieval quality is necessary but not sufficient. Cheaper, denser context only matters if it compounds across a real, multi-step engineering task — finding the code, understanding it, changing it, and not breaking anything. The next suite measures exactly that: resolve-rate on SWE-bench-style multi-file tasks, sweet-search-wired vs. native, on the same paired, multiplicity-controlled bar as above. Harness and pilot are in progress — numbers land here when they clear that bar, and not before.


📄 3. Paper-type retrieval benchmarks — academic NL→code IR

Every number below is the ss-search pipeline end-to-end — the same binary you install — run against the full benchmark corpus (no 99-distractor shortcuts), zero-shot (we never fine-tune on these tasks). Where a benchmark's queries are docstrings, we strip the docstring out of the indexed code so the query can't trivially match itself — the standard retrieval protocol.

We're SOTA in June 2026 on 3/4 attempted benchmarks at HARDER settings (running on full pool) than most other attempts!

📚 Benchmark 🔍 What it tests # Queries 📂 Pool 🎯 MRR@10 🏆 SOTA?
🌐 GenCodeSearchNet NL→code, 6 languages 6,000 full 6,000 86.6 YES ✅
🐍 CoSQA web queries → Python 500 full 6,267 65.5 ✅ (zero-shot)
🗺️ M2CRB multilingual NL→code (ES/PT/DE/FR → Py/Java/JS) 5,795 full 5,795 54.0 YES ✅
🛡️ AdvTest adversarial, identifier-obfuscated Python 19,210 full 19,210 51.4 NO ❌

SOTA = best result we can find in the published literature as of June 2026; cross-metric/protocol comparisons are spelled out per benchmark below.

🌐 GenCodeSearchNet → 86.6  ·  🏆 SOTA in June 2026

  • The BEST PUBLISHED number we can find, anywhere
  • The benchmark's own paper caps at MRR ≤ 0.42 for fine-tuned baselines (≤ 0.10 cross-lingual); even zero-shot OpenAI Ada-2 reaches 0.79–0.94 — but all of it against a tiny 99-distractor pool.
  • We score 0.866 against the entire 6,000-document corpusa strictly harder setting — and zero-shot. 🔥

🐍 CoSQA → 65.5  ·  🥇 Zero-shot SOTA in June 2026

  • Beats EVERY PUBLISHED zero-shot model
  • Canonical setup: 500 real web queries → the fixed 6,267-code database, no fine-tuning.
  • Clears the strongest zero-shot results out there — CodeSage-Large 47.5 · OpenAI text-embedding-3-large 55.4 · OASIS 55.8 — and goes toe-to-toe with fine-tuned CodeBERT / GraphCodeBERT (64.7 / 67.5). 💪
  • CoSQA has known label noise, so we read the absolute height with a pinch of salt.

🗺️ M2CRB → 54.0  ·  🏆 SOTA in June 2026

  • the BEST PUBLISHED number we can find, anywhere — and zero-shot
  • 🇪🇸 Spanish · 🇵🇹 Portuguese · 🇩🇪 German · 🇫🇷 French → Python / Java / JavaScript.
  • The paper's best — a CodeBERT fine-tuned on the task — reaches 52.7 auMRRc, a metric that averages over easier, smaller pools (so auMRRc ≥ full-pool MRR for any model). Our 54.0 is full-pool MRR@10 over all 5,795 functions in one pool — a strictly harder measure, cleared with no fine-tuning. 🔥

🛡️ AdvTest → 51.4  ·  🧪 our honest worst case — and we publish it anyway

  • Adversarial obfuscation (def Func(arg_0):) deletes the lexical + graph signals our hybrid feeds on — yet we still beat the classic fine-tuned baselines (CodeBERT 27 · GraphCodeBERT 35 · UniXcoder 41), and our stack still lifts our own encoder ~3pp even here.
  • 🔍 Full transparency: we could not reproduce the often-cited 59.5 for the bare CodeRankEmbed encoder — the reference FP32 model scores 54.7 on our leak-free corpus, our shipped INT8 build 51.4. The gap is stricter preprocessing + INT8 quantization, not the retrieval pipeline. We report exactly what we measured.
  • Reproduction: result artifacts live in eval/results/; rerun via eval/run_all.js. The canonical full-pool loaders are in eval/download_data.py.
  • Full corpus, not distractors. Published baselines for GCSN- and CoSQA-style benchmarks typically rank the gold against 99 sampled distractors; every number here ranks against the benchmark's full corpus (6k–19k candidates) — strictly harder.
  • Zero-shot + docstring-stripped. We never fine-tune on these tasks. For docstring-derived benchmarks (AdvTest, M2CRB) we strip the docstring from the indexed code — otherwise the NL query matches itself verbatim (a no-strip AdvTest run scores a meaningless 0.98). This is the standard protocol; it is also why our AdvTest is lower than naïve setups that leave the docstring in.
  • Dev/held-out split. Our ranking work iterates against a fixed dev split of each benchmark (e.g. GCSN: 600 dev + 400 held-out per language, stratified, seed=42) and treats the remainder as held-out, inspected aggregate-only at milestones. The table figures are full-corpus runs — they include the dev portion we tuned against; a held-out-only breakdown ships with the next results refresh.
  • What we deliberately don't claim yet. CoIR (official metric NDCG@10 over per-subtask corpora up to ~1M docs), CoSQA+ (multi-positive, MAP-primary), and CLARC (per-group pools) use protocols and metrics our single-pool MRR@10 harness doesn't currently match. Rather than publish apples-to-oranges numbers, we omit them; faithful per-subtask CoIR (NDCG@10) runs are queued.
  • M2CRB — the paper's metric is auMRRc (area under the MRR-vs-pool-size curve; best published 52.7, fine-tuned). Because that area averages over easier small pools, auMRRc ≥ full-pool MRR for any model — so our 54.0 full-pool MRR@10 (all 5,795 functions, zero-shot) clears their best on a strictly harder measure. No one publishes a plain full-corpus MRR@10 on M2CRB, so ours is the best available.
  • AdvTest honesty note. We could not reproduce the commonly-cited 59.5 for the bare CodeRankEmbed encoder on our corpus: the reference FP32 model scores 54.7 on our leak-free, docstring-stripped, full-19,210 setup, and our shipped INT8 build 51.4. We report our measured numbers and the reference check rather than the leaderboard figure.
  • Honesty corner: CrossCodeEval — cross-file completion-context retrieval, a different task than NL search — sits at 0.12. We don't optimize for it and report it anyway.

⚡ 4. Engine speed — systems benchmarks, measured in-repo

10.2× ripgrep's median grep  ·  2.9 ms warm queries  ·  47× MaxSim kernels  ·  −33% HNSW search p50

⚙️ What 📈 Result 📄 Source
⚡ Indexed grep vs ripgrep 10.2× faster at the median (8.5–17.7× across 5 repos, 353 realistic queries, 1 ms p50 — identical match counts on every query) docs/GREP_INDEXING_STRATEGY.md
⏱️ Warm query latency (native CLI) 2.9 ms warm · 108 ms cold docs/INIT_STRATEGY.md
🧮 MaxSim rerank kernels 1.26 s → 27 ms for a 231-candidate pass (47× native Rust; 16× WASM SIMD) docs/MAXSIM_OPTIMIZATION.md
🧠 HNSW tuning for code −33% search p50, +5.9 pp recall@200 docs/HNSW_APPROACH.md
💾 Indexing memory peak JS heap 785 MB → 213 MB docs/DISK_FLUSHING_STRATEGY.md
🍏 CoreML cascade (M3 Max) 18% faster full indexing vs the Metal baseline docs/INIT_STRATEGY.md

🧭 Where sweet-search Fits

Code search is a crowded space. Here's an honest read on where sweet-search wins and where it gives ground, against the trending leaders and our closest local peers.

Capability sweet-search claude-context Cursor index codebase-memory SocratiCode
100% local — code never leaves your machine ✅¹
Works with zero API keys ✅¹
No external service to run (vector DB · Ollama · Docker) ❌ Milvus ❌ cloud ⚠️⁵
ColBERT late-interaction rerank
Faster-than-ripgrep exact grep ✅⁷
Call-graph trace (callers · callees · impact)
Drives any terminal agent (Claude Code · Codex · Gemini CLI) ❌²
Published NL→code retrieval benchmarks ⚠️³ ⚠️³ ⚠️³
…and where sweet-search gives ground
Native Windows ❌⁴ ⚠️⁸
Deep-AST language coverage ⚠️ 14 (+70 via regex) ⚠️ ⚠️ ✅ 158 ⚠️
In-editor GUI · writes & edits code ❌⁶
Org-wide, multi-repo scale ⚠️ ⚠️ ⚠️

✅ yes · ⚠️ partial / with caveats · ❌ no. Verified June 2026; capabilities drift. ¹ claude-context's local path (Milvus Lite + Ollama embeddings) needs no API key, but it defaults to OpenAI/Voyage embeddings + Zilliz Cloud — and still runs Milvus + Ollama either way. ² Cursor's index is editor-locked — external terminal agents can't query it. ³ Reports token-reduction / efficiency, not a public NL→code retrieval-quality leaderboard. ⁴ Runs on Windows via WSL2. ⁵ SocratiCode manages a bundled Qdrant for you, but uses an auto-detected Ollama for local embeddings. ⁶ Ships an interactive HTML graph viewer, but doesn't edit code. ⁷ Cursor's local Instant Grep — a literal + regex index it benchmarks at ripgrep 16.8 s → 13 ms (the post that inspired our own n-gram prefilter). ⁸ SocratiCode runs on Windows via Docker only — no native binary, and no GPU there.

Where we lose, plainly: no native Windows yet, no editor GUI, and we index one repo at a time. If you need org-wide search across many repos and branches, that's where SocratiCode and Sourcegraph are built to win. If you live inside one editor, Cursor's index is already there. sweet-search is for the terminal agent that wants the best local retrieval on the repo in front of it. No one else combines all of it: ColBERT late-interaction reranking and faster-than-grep search, fully on-device, with nothing to sign up for.

Also in the space: Sourcegraph/Cody (org-scale, server-based), Continue.dev (local-default RAG), Serena (LSP symbol search, no embeddings), grepai (local CLI + trace), and cocoindex-code (embedded AST search).

🧰 The Six Tools

Six small tools, one shared index. Each returns ranked, deduplicated, token-budgeted output designed to be consumed by an agent — a useful answer, not a wall of matches to scroll through.

Tool What you give it What you get back
1. ss-search a natural-language query ranked, self-contained code blocks
2. ss-grep an exact regex/literal every file:line hit, ripgrep-identical
3. ss-find a regex + a query regex matches, semantically re-ranked, as code blocks
4. ss-semantic a file + a question just the relevant spans of that file
5. ss-trace a symbol callers + callees + impact, in one call
6. ss-read a file (± line range) exact bytes + symbol metadata

1. 🔍 ss-search — hybrid search powerhouse

A hybrid search pipeline with late interaction reranking that returns actual code blocks.

Leading published-benchmark results — strongest we can find on GenCodeSearchNet, and above every published zero-shot model on CoSQA. See benchmarks.

flowchart TD
    Q(["🔍  natural-language query"]) --> ROUTE{{"🧭 WASM CatBoost router · lexical / hybrid"}}

    ROUTE --> BM["📑 <b>BM25F</b><br/>field-weighted FTS5"]
    ROUTE --> ANN

    subgraph ANN ["🧬 three-stage ANN cascade"]
        direction LR
        BIN["binary <b>HNSW</b><br/>Hamming · ~100µs"] --> INT["INT8<br/>rescore"] --> FL["float32<br/>mmap sidecar"]
    end

    BM  --> FUSE
    ANN --> FUSE
    FUSE["🔀 <b>CCFusion</b><br/>convex combo · RRF fallback"] --> ROW1

    subgraph ROW1 [" "]
        direction LR
        IAR["⚓ <b>IAR</b><br/>exact-symbol injection"] --> INTENT["🎯 intent rerank<br/>demote docs · tests · config"]
    end

    ROW1 --> ROW2

    subgraph ROW2 [" "]
        direction LR
        GRAPH["🕸️ graph expansion<br/>typed edges · 1–2 hops · <b>PathRAG</b>"] --> MAXSIM["🧮 <b>Late-Interaction Rerank</b><br/>⚡ native Rust MaxSim kernel"] --> OUT(["🏁 <b>self-contained code blocks</b><br/>whole functions · 3k/8k/12k budget"])
    end

    classDef io    fill:#fde68a,stroke:#f59e0b,color:#000;
    classDef out   fill:#bbf7d0,stroke:#15803d,color:#000,stroke-width:3px;
    classDef route fill:#e0e7ff,stroke:#818cf8,color:#000;
    classDef lex   fill:#dbeafe,stroke:#60a5fa,color:#000;
    classDef fuse  fill:#f3e8ff,stroke:#c084fc,color:#000;
    classDef rank  fill:#ffe4e6,stroke:#fb7185,color:#000;

    class Q io;
    class OUT out;
    class ROUTE route;
    class BM,BIN,INT,FL lex;
    class FUSE,IAR fuse;
    class INTENT,GRAPH,MAXSIM rank;

    style ANN  fill:#eff6ff,stroke:#93c5fd,color:#000;
    style ROW1 fill:none,stroke:none;
    style ROW2 fill:none,stroke:none;

↑ The diagram traces the hybrid route. A pure-lexical query — or a literal file path — short-circuits at the router straight to BM25F, skipping the vector cascade and fusion.

Stage What it actually does
🧭 Route WASM-exported CatBoost · lexical / hybrid · ~10 µs routing · low-confidence → max-recall hybrid
🧬 Retrieve LexicalBM25F over field-weighted FTS5 (name 10× · signature 5× · alias 4× · doc 1×)• Embed — query vectorized by the local CodeRankEmbed model (swappable for Voyage / Jina / Codestral)• Vector cascade — binary HNSW (Hamming, 64-byte, ~100 µs) → INT8 rescore → exact float32 from a memory-mapped sidecar
🔀 Fuse CCFusion — convex-combine both rankings · per-route weights · quantile-normalized• MMR (λ=0.9) diversity pass over the fused list• auto RRF (k=60) fallback on degenerate score distributions
Anchor IAR (Identifier Anchor Retrieval) — a real symbol in the query fires an exact-name code-graph lookup that injects that entity, even when the encoder ranked it too low
🎯 Intent Rerank • demote docs / tests / config when you want implementation• log-scaled call-site boosts surface the most-referenced function
🕸️ Graph Expansion • typed-edge walks (imports/extends/calls/uses) · adaptive 2-hop on the AST graph · edges picked by intent• PathRAG flow pruning + degree normalization → hubs can't dominate
🧮 Late interaction Rerank • Query embedded per-token by LateOn-Code (149M; a 17M edge variant auto-selected on low-RAM hosts)• MaxSim against the pre-indexed quantized token vectors• native Rust+Rayon MaxSim kernel ⚡ · WASM-SIMD fallback (1.26 s → 27 ms on a 231-candidate rerank)
📦 Package • entity-aware expansion → whole functions (imports, docstrings, decorators)• same-file overlap demotion → diverse, non-overlapping spans• symbol-family completion (agent mode) — generated/width families surface as a compact indexed manifest instead of truncating silently, inside the same budget• auto-selected 3k / 8k / 12k token budget

🧠 The HNSW, in full (full writeup). Stage 1 is a from-scratch binary HNSW, and every "advanced" trick ships on by default:

  • Heuristic neighbor selection (HNSW Algorithm 4) + M0 = 2M on layer 0 — a real graph backbone, not naïve closest-M
  • Shuffled insertion order — no filesystem-ordering bias baked into the highway structure
  • Discovery-rate adaptive early termination + adaptive ef — easy queries stop early, hard ones keep their budget
  • A denser graph than most vendors ship (M=64 · efC=800 · efS=400) — which broke an 80.6 % → 86.5 % recall@200 plateau and cut p50 latency ~33 %
  • Zero-GC search: typed-array heaps + generation-stamped visited lists — no per-query allocation
  • 64-byte sign-bit vectors (Hamming) → INT8 → exact float32 from a memory-mapped sidecar

⚡ Why it's quick. A native Rust + Rayon MaxSim kernel (47× over scalar; 16× WASM-SIMD fallback) · int4-quantized, binary-packed token vectors (plain INT4 is the shipped path — the full TurboQuant algorithm is researched but deferred; binary packing alone cut the LI index ~3.4×, 1.34 GiB → ~396 MiB) · a memory-mapped float32 sidecar that skips SQL on the rescore hot path · score-spread adaptive pooling (decisive queries shrink the rescore pool, ambiguous ones widen it) · and a warm daemon that answers in a single NAPI call — no process is ever forked.

🎛️ Priors & structure.

  • Quality priors: every chunk carries a 0–1 prior from test proximity, git recency, symbol centrality (PageRank), comment density, and complexity — production code surfaces, stale fixtures sink.
  • Community structure: a canonical Leiden pass detects code communities on the entity graph at index time, feeding vocabulary prewarming and structural signals — it understands your modules, not just your directories.
  • Multilingual: 14 languages get full tree-sitter AST treatment; a 39-config registry covers 70+ extensions beyond that. Router features handle camelCase/snake_case, CJK density, and German compounds.
  • Format-gated signals: structure-aware boosts and demotions (symbol-exact, path-token, anomalous-chunk, mega-entity) fire only in agent mode — they help agent-shaped queries and would hurt plain NL, so they stay gated by default.

🛟 Rescues & honest trade-offs.

  • Long-query rescue: wordy NL queries that FTS5 would tokenize into an unsatisfiable AND fall back to multi-query BM25F + RRF — one query per content keyword, fused.
  • Near-duplicate dedup: a SimHash + MinHash-LSH pass (Jaccard τ=0.9) clusters copy-paste and vendored code at index time; aliases reuse their exemplar's vectors and skip both the bi-encoder and late-interaction encoding.
  • A negative result we ship anyway: we built a full cross-encoder rerank cascade behind an adaptive confidence gate, measured it on our eval sets — and it didn't beat MaxSim at 3× the latency. So it ships disabled (SWEET_SEARCH_CASCADE_ENABLED=true to try it). We'd rather ship the faster path than a fancier diagram.
  • Budget tiers: the expensive 8k/12k tiers fire on ~1–5 % of queries — the default stays cheap. Force one with --full / --xl, or pick a mode with --mode lexical|semantic|hybrid|pattern.

Also available as sweet-search "<query>" on the CLI and the search MCP tool.


2. ⚡ ss-grep — grep, minus every wasted millisecond

10.2× faster than ripgrep end-to-end at the median — measured across 353 realistic queries on 5 real repos (range 8.5–17.7× per repo, 1 ms p50), with identical match counts on every single query. Three things buy that:

  • A sparse n-gram index (inspired by Cursor's fast-regex-search and GitHub's Blackbird): instead of a fixed trigram table, gram boundaries adapt to your codebase's character-pair frequencies, so common trigrams get absorbed into longer, more selective grams.
  • Regex-AST literal extraction + SIMD intersection: required substrings are pulled from the pattern's syntax tree, posting lists are intersected with NEON/SSE2 block merges (galloping search for skewed sizes), and only the files that can match — typically 0.1–5% of the corpus — see the real regex.
  • Fully in-process: verification runs on Rust's regex crate with Rayon across all cores, inside the warm daemon, in a single NAPI call. No child process is ever spawned — zero fork/exec, zero pipe I/O, zero JSON re-parsing.

Every match comes back in stable file:line order — ripgrep-identical counts, optional context lines — with no relevance guessing, no subprocess, in one warm call.

  • Full methodology, per-repo table, and the optimization log: docs/GREP_INDEXING_STRATEGY.md.
  • Regexes with no extractable literals fall back to native grep over the indexed file set; fixed-string and glob queries use a ripgrep fallback.
  • Dialect recovery (agent mode): patterns written in GNU-grep BRE muscle memory (foo\|bar, \(group\)) are literals in Rust's regex dialect and used to silently match nothing — a zero-hit exact search now gets one gated auto-retry with the translated pattern instead of a false "no matches".

3. ss-find — ColGrep, on a faster engine

ss-find "token refresh logic" --regex "refresh.*[Tt]oken"

Inspired by LightOn's ColGrep — regex precision, semantically ranked — but rebuilt on our own substrate:

  • The regex stage runs on the same indexed sparse-gram engine as ss-grep (in-process, no subprocess), not a filesystem scan.
  • The ranking stage scores candidates with per-token MaxSim over pre-indexed late-interaction embeddings — no model inference over documents at query time — on our custom kernels: native R