Learned-sparse first-stage retriever for BrainAPI POST /retrieve/search. It keeps its own inverted index (plugin-local, in memory) and registers channel plugin:splade. It does not run on /retrieve/context. Core hybrid (passages / BM25 / dense) works if this plugin is absent.
Unknown or missing plugin:splade is 400, never a silent BM25 fallback.
| Registry name | search-splade |
| Version | 0.1.0 |
| BrainAPI | >=2.17.0 |
| Channel | plugin:splade |
| Default model | naver/splade-cocondenser-ensembledistil |
| Index | POST /search-splade/index |
| Health | GET /search-splade/health |
git clone https://github.com/Lumen-Labs/brainapi-plugin-search-splade.git plugins/search-spladeOr:
./bin/brainapi install search-spladeRestart the API. Encoding needs torch and transformers in the BrainAPI environment. The model is lazy-loaded on first encode.
Index text chunks already stored in the brain, then search:
curl -X POST "$BRAINAPI_URL/search-splade/index" \
-H "Content-Type: application/json" \
-H "BrainPAT: $BRAINPAT_TOKEN" \
-d '{"brain_id": "searchbenchsmoke", "limit": 1000}'
curl -X POST "$BRAINAPI_URL/retrieve/search" \
-H "Content-Type: application/json" \
-H "BrainPAT: $BRAINPAT_TOKEN" \
-H "X-Brain-ID: searchbenchsmoke" \
-d '{
"query": "navy wool coat",
"k": 50,
"channels": ["plugin:splade"]
}'Fuse with core passages:
{ "channels": ["passages", "plugin:splade"] }Benchmark harness: --channels plugin:splade after indexing.
POST /search-splade/indexpagesget_text_chunks(up tolimit, max 20 000) and SPLADE-encodes each doc to a sparse{token: weight}vector.- An inverted index maps term →
[(chunk_id, weight), …]. - A query is encoded the same way. Score is the dot product over overlapping terms.
- Top
kids are returned to/retrieve/searchas channelplugin:splade.
The index lives in process. Restarting the API clears it — you must re-index. index_chunks replaces the brain’s index (reset then rebuild).
| Env | Default |
|---|---|
SEARCH_SPLADE_MODEL |
naver/splade-cocondenser-ensembledistil |
Encoding: MLM logits → log1p(relu) → max-pool over sequence → nonzero vocab weights. Special tokens (cls / sep / pad / unk) are dropped. Max length 256.
Tests can inject set_encoder(fn).
{
"plugin": "search-splade",
"channel": "plugin:splade",
"model": "naver/splade-cocondenser-ensembledistil",
"loaded": false,
"error": null,
"index": { "brain_id": "searchbenchsmoke", "n_docs": 2043, "n_terms": 12000 }
}index is included only when brain_id is passed.
{ "brain_id": "searchbenchsmoke", "limit": 1000 }limit is 1…20000 (default 1000). Returns { brain_id, n_docs, n_terms }.
search-splade/
plugin.yaml
main.py # register_search_retriever("splade", …)
encode.py # SPLADE sparse encoder
index.py # inverted index + retrieve
routes.py # health + index
Pushes to main publish to the BrainAPI registry via GitHub Actions.
Apache License, Version 2.0. See LICENSE.
- search-colbert
- search-rerank
- BrainAPI
docs/research/18-search-eval-protocol.mdon brainapi2